JIT: don't kill FP/SIMD/mask regs across x64 write barriers - #128778

Merged
EgorBo merged 3 commits into
dotnet:mainfrom
EgorBo:jit-wb-preserve-fp
Jun 2, 2026
Merged

JIT: don't kill FP/SIMD/mask regs across x64 write barriers#128778
EgorBo merged 3 commits into
dotnet:mainfrom
EgorBo:jit-wb-preserve-fp

Conversation

@EgorBo

@EgorBoEgorBo commented May 29, 2026

Copy link
Copy Markdown
Member
{7366C871-08EB-48F1-897F-021AE10303FC}

What I deliberately did not do

  • Did not exclude RCX/RDI (dst) from the kill set — the asm helpers shift dst in place (shr rcx, 0Bh), so it must be considered clobbered. Preserving dst would require rewriting all ~20 asm variants.
  • Did not exclude RDX/RSI (src) — the Region variants destroy src.
  • Did not exclude R10/R11 on Windows even though the default patched-slot path doesn't touch them. Kept conservative to cover _DEBUG runtime variant (JIT_WriteBarrier_Debug) and the DOTNET_UseGCWriteBarrierCopy=0 config (RhpAssignRef), both of which touch R10/R11.

PS: We probably can do this for APX cc @dotnet/intel

I think it would be nice to preserve dst register unchanged (we do that for arm64), but that requires some changes in the ASM helpers.

The x64 JIT_WriteBarrier / JIT_CheckedWriteBarrier helpers (and all the
patched-slot variants: PreGrow/PostGrow/SVR/Region + WriteWatch flavors)
never execute any SSE/AVX/AVX-512/EVEX-mask instruction. So XMM/YMM/ZMM
and K mask registers can stay live across a write barrier call.
Narrow RBM_CALLEE_TRASH_WRITEBARRIER (and the matching GCTRASH mask)
from the full RBM_CALLEE_TRASH down to RBM_INT_CALLEE_TRASH_INIT on
amd64. Integer callee-trash regs remain conservatively in the kill set
to keep _DEBUG runtime builds and DOTNET_UseGCWriteBarrierCopy=0 working.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings May 29, 2026 16:56
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 29, 2026
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR narrows the amd64 JIT write-barrier register kill masks so FP/SIMD/mask registers can remain live across CORINFO_HELP_ASSIGN_REF and CORINFO_HELP_CHECKED_ASSIGN_REF, matching the documented helper ABI and reducing unnecessary spills.

Changes:

  • Documents the amd64 write-barrier helper register preservation/clobbering contract.
  • Changes amd64 write-barrier trash and GC-trash masks from full callee-trash to integer-only initial callee-trash registers.
  • Keeps conservative integer clobbers for runtime/debug write-barrier variants.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@MihuBot

@VSadov

Copy link
Copy Markdown
Member

Will this require write barriers that need to call into C++ code on some paths to spill more around the calls?

For context:
https://github.com/dotnet/runtimelab/blob/97a242e8b9a66df1f9de9ff16fc24c42035bcc98/src/coreclr/runtime/amd64/WriteBarriers.asm#L517-L526

@EgorBo

EgorBo commented May 30, 2026

Copy link
Copy Markdown
MemberAuthor

Will this require write barriers that need to call into C++ code on some paths to spill more around the calls?

For context: https://github.com/dotnet/runtimelab/blob/97a242e8b9a66df1f9de9ff16fc24c42035bcc98/src/coreclr/runtime/amd64/WriteBarriers.asm#L517-L526

Probably not since we cannot guarantee what is happening in a code compiled by C++. But if that is an issue we need a new JIT-EE API, because we already precisely track registers used by WB on arm64 today (and just removing that will lead to quite big regressions) and we've always been doing that for byref WB that I deleted recently

CopilotAI review requested due to automatic review settings May 30, 2026 01:16

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 1 out of 1 changed files in this pull request and generated 1 comment.

Comment threadsrc/coreclr/jit/targetamd64.h
Comment threadsrc/coreclr/jit/targetamd64.h
Comment threadsrc/coreclr/jit/targetamd64.h
kg
kg approved these changes Jun 1, 2026

@kgkg left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

vsadov's comment occurred to me as well, but otherwise this looks okay to me

@VSadov

Copy link
Copy Markdown
Member

What bothers me a bit is that this just encodes the current implementation that happens to be. If the eventual goal is to be able to switch GCs without rebuilding the entire runtime, the barrier calling conventions should be more "designed".

We already have to spill a lot more on arm64:
https://github.com/dotnet/runtimelab/blob/97a242e8b9a66df1f9de9ff16fc24c42035bcc98/src/coreclr/runtime/arm64/WriteBarriers.asm#L598-L609

 ;; because of the barrier call convention ;; we need to preserve caller-saved x0 through x15 and x29/x30 stp x29,x30,[sp,-16*9]! stp x0, x1,[sp,16*1] stp x2, x3,[sp,16*2] stp x4, x5,[sp,16*3] stp x6, x7,[sp,16*4] stp x8, x9,[sp,16*5] stp x10,x11,[sp,16*6] stp x12,x13,[sp,16*7] stp x14,x15,[sp,16*8]

On x64 the spill used to be fairly minimal in comparison.
Either way kind of worked, but which way is better? Are we trading binary size for throughput here?

Calls into C++ code are not on common hot paths, but even on uncommon paths spilling too much may become measurable.

I think we may want to revisit the barriers calling convention is at some point.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@VSadov I think we just need a JIT-EE API that will tell the JIT about the calling convention for write barriers (e.g. eithe just normal or custom) and Satori can return "normal" (on both x64 and arm64) and remove that spilling logic you have.
BTW, are you sure you also don't need to spill SIMD regs there? Because I think arm64 JIT assumes they're not killed too today, but it's possible that your C++ call just never use them either (but not something you can rely on).

Custom calling convention seems to be very proftiable, e.g. in this PR I tried just to avoid touching RCX/RDI reg (typically holding this) and got massive diffs in SPMI (-2.0Mb on SysV and -1.3Mb on win-x64).

@VSadov

VSadov commented Jun 1, 2026

Copy link
Copy Markdown
Member

@VSadov I think we just need a JIT-EE API that will tell the JIT about the calling convention for write barriers (e.g. eithe just normal or custom) and Satori can return "normal" (on both x64 and arm64) and remove that spilling logic you have. BTW, are you sure you also don't need to spill SIMD regs there? Because I think arm64 JIT assumes they're not killed too today, but it's possible that your C++ call just never use them either (but not something you can rely on).

Custom calling convention seems to be very proftiable, e.g. in this PR I tried just to avoid touching RCX/RDI reg (typically holding this) and got massive diffs in SPMI (-2.0Mb on SysV and -1.3Mb on win-x64).

I think standard calling convention may be too generous towards GC barriers. I understand the desire to leave registers for JIT to use.

However, the current conventions could be:

  • too tight that the asm part of the barrier may need to push/pop for extra temps.
  • unfriendly to C++ code.

It may be a good idea to make it configurable via JIT-EE API.
And I assume crossgen could just use the most conservative "common denominator" combination.

BTW, are you sure you also don't need to spill SIMD regs there?

We rely on C++ code that we call to not trash SIMD (by manual examination). On x64 we spill XMM0, because it does.
It is a bit of playing with fire.

SIMD is probably where I'd put the line. Spilling the whole SIMD context could be expensive, while chance of it used while calling a barrier are not high.

To be fair - the effects were never a target of too much research/measuring. Perhaps the current scheme is even completely OK.

@EgorBo

EgorBo commented Jun 1, 2026

Copy link
Copy Markdown
MemberAuthor

@VSadov It should be trivial to implement the JIT-EE I think, I can give it a try if Satori is anywhere near to be productized/merged into dotnet/runtime, something like CORINFO_WRITEBARRIER_CALLC getWriteBarrierCallingConvention();

The reason I decided to tune the registers for WB are potential regressions from my other change to remove CORINFO_HELP_ASSIGN_BYREF (#128542 + #128687) because that WB had a very compact calling convention with no spills between calls at all

@jkotas

Copy link
Copy Markdown
Member

We rely on C++ code that we call to not trash SIMD (by manual examination).

We used to do that in a few places and learned hard way that it is a bad idea...

@EgorBo
EgorBo enabled auto-merge (squash) June 2, 2026 21:01
@EgorBo
EgorBo merged commit db87610 into dotnet:mainJun 2, 2026
131 of 139 checks passed
@EgorBo

Copy link
Copy Markdown
MemberAuthor

@VSadov let me know if you want me to introduce an API that can switch CC for WB between this custom and standard (if we sure we'll merge Satori into runtime eventually)

I'm merging it to match the behavior with arm64

@VSadov

Copy link
Copy Markdown
Member

@VSadov let me know if you want me to introduce an API that can switch CC for WB between this custom and standard (if we sure we'll merge Satori into runtime eventually)

I'm merging it to match the behavior with arm64

I'll let you know. Right now Satori is a bit behind main. I think we will need to catch up to the state just before this PR first. (deleting byref barriers might be a good thing in terms of reducing complexity).

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@EgorBo@VSadov@jkotas@kg@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

JIT: don't kill FP/SIMD/mask regs across x64 write barriers - #128778

Merged
EgorBo merged 3 commits into
dotnet:mainfrom
EgorBo:jit-wb-preserve-fp
Jun 2, 2026
Merged

JIT: don't kill FP/SIMD/mask regs across x64 write barriers#128778
EgorBo merged 3 commits into
dotnet:mainfrom
EgorBo:jit-wb-preserve-fp

Conversation

@EgorBo

@EgorBoEgorBo commented May 29, 2026

Copy link
Copy Markdown
Member
{7366C871-08EB-48F1-897F-021AE10303FC}

What I deliberately did not do

  • Did not exclude RCX/RDI (dst) from the kill set — the asm helpers shift dst in place (shr rcx, 0Bh), so it must be considered clobbered. Preserving dst would require rewriting all ~20 asm variants.
  • Did not exclude RDX/RSI (src) — the Region variants destroy src.
  • Did not exclude R10/R11 on Windows even though the default patched-slot path doesn't touch them. Kept conservative to cover _DEBUG runtime variant (JIT_WriteBarrier_Debug) and the DOTNET_UseGCWriteBarrierCopy=0 config (RhpAssignRef), both of which touch R10/R11.

PS: We probably can do this for APX cc @dotnet/intel

I think it would be nice to preserve dst register unchanged (we do that for arm64), but that requires some changes in the ASM helpers.

The x64 JIT_WriteBarrier / JIT_CheckedWriteBarrier helpers (and all the
patched-slot variants: PreGrow/PostGrow/SVR/Region + WriteWatch flavors)
never execute any SSE/AVX/AVX-512/EVEX-mask instruction. So XMM/YMM/ZMM
and K mask registers can stay live across a write barrier call.
Narrow RBM_CALLEE_TRASH_WRITEBARRIER (and the matching GCTRASH mask)
from the full RBM_CALLEE_TRASH down to RBM_INT_CALLEE_TRASH_INIT on
amd64. Integer callee-trash regs remain conservatively in the kill set
to keep _DEBUG runtime builds and DOTNET_UseGCWriteBarrierCopy=0 working.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings May 29, 2026 16:56
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 29, 2026
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR narrows the amd64 JIT write-barrier register kill masks so FP/SIMD/mask registers can remain live across CORINFO_HELP_ASSIGN_REF and CORINFO_HELP_CHECKED_ASSIGN_REF, matching the documented helper ABI and reducing unnecessary spills.

Changes:

  • Documents the amd64 write-barrier helper register preservation/clobbering contract.
  • Changes amd64 write-barrier trash and GC-trash masks from full callee-trash to integer-only initial callee-trash registers.
  • Keeps conservative integer clobbers for runtime/debug write-barrier variants.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@MihuBot

@VSadov

Copy link
Copy Markdown
Member

Will this require write barriers that need to call into C++ code on some paths to spill more around the calls?

For context:
https://github.com/dotnet/runtimelab/blob/97a242e8b9a66df1f9de9ff16fc24c42035bcc98/src/coreclr/runtime/amd64/WriteBarriers.asm#L517-L526

@EgorBo

EgorBo commented May 30, 2026

Copy link
Copy Markdown
MemberAuthor

Will this require write barriers that need to call into C++ code on some paths to spill more around the calls?

For context: https://github.com/dotnet/runtimelab/blob/97a242e8b9a66df1f9de9ff16fc24c42035bcc98/src/coreclr/runtime/amd64/WriteBarriers.asm#L517-L526

Probably not since we cannot guarantee what is happening in a code compiled by C++. But if that is an issue we need a new JIT-EE API, because we already precisely track registers used by WB on arm64 today (and just removing that will lead to quite big regressions) and we've always been doing that for byref WB that I deleted recently

CopilotAI review requested due to automatic review settings May 30, 2026 01:16

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 1 out of 1 changed files in this pull request and generated 1 comment.

Comment threadsrc/coreclr/jit/targetamd64.h
Comment threadsrc/coreclr/jit/targetamd64.h
Comment threadsrc/coreclr/jit/targetamd64.h
kg
kg approved these changes Jun 1, 2026

@kgkg left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

vsadov's comment occurred to me as well, but otherwise this looks okay to me

@VSadov

Copy link
Copy Markdown
Member

What bothers me a bit is that this just encodes the current implementation that happens to be. If the eventual goal is to be able to switch GCs without rebuilding the entire runtime, the barrier calling conventions should be more "designed".

We already have to spill a lot more on arm64:
https://github.com/dotnet/runtimelab/blob/97a242e8b9a66df1f9de9ff16fc24c42035bcc98/src/coreclr/runtime/arm64/WriteBarriers.asm#L598-L609

 ;; because of the barrier call convention ;; we need to preserve caller-saved x0 through x15 and x29/x30 stp x29,x30,[sp,-16*9]! stp x0, x1,[sp,16*1] stp x2, x3,[sp,16*2] stp x4, x5,[sp,16*3] stp x6, x7,[sp,16*4] stp x8, x9,[sp,16*5] stp x10,x11,[sp,16*6] stp x12,x13,[sp,16*7] stp x14,x15,[sp,16*8]

On x64 the spill used to be fairly minimal in comparison.
Either way kind of worked, but which way is better? Are we trading binary size for throughput here?

Calls into C++ code are not on common hot paths, but even on uncommon paths spilling too much may become measurable.

I think we may want to revisit the barriers calling convention is at some point.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@VSadov I think we just need a JIT-EE API that will tell the JIT about the calling convention for write barriers (e.g. eithe just normal or custom) and Satori can return "normal" (on both x64 and arm64) and remove that spilling logic you have.
BTW, are you sure you also don't need to spill SIMD regs there? Because I think arm64 JIT assumes they're not killed too today, but it's possible that your C++ call just never use them either (but not something you can rely on).

Custom calling convention seems to be very proftiable, e.g. in this PR I tried just to avoid touching RCX/RDI reg (typically holding this) and got massive diffs in SPMI (-2.0Mb on SysV and -1.3Mb on win-x64).

@VSadov

VSadov commented Jun 1, 2026

Copy link
Copy Markdown
Member

@VSadov I think we just need a JIT-EE API that will tell the JIT about the calling convention for write barriers (e.g. eithe just normal or custom) and Satori can return "normal" (on both x64 and arm64) and remove that spilling logic you have. BTW, are you sure you also don't need to spill SIMD regs there? Because I think arm64 JIT assumes they're not killed too today, but it's possible that your C++ call just never use them either (but not something you can rely on).

Custom calling convention seems to be very proftiable, e.g. in this PR I tried just to avoid touching RCX/RDI reg (typically holding this) and got massive diffs in SPMI (-2.0Mb on SysV and -1.3Mb on win-x64).

I think standard calling convention may be too generous towards GC barriers. I understand the desire to leave registers for JIT to use.

However, the current conventions could be:

  • too tight that the asm part of the barrier may need to push/pop for extra temps.
  • unfriendly to C++ code.

It may be a good idea to make it configurable via JIT-EE API.
And I assume crossgen could just use the most conservative "common denominator" combination.

BTW, are you sure you also don't need to spill SIMD regs there?

We rely on C++ code that we call to not trash SIMD (by manual examination). On x64 we spill XMM0, because it does.
It is a bit of playing with fire.

SIMD is probably where I'd put the line. Spilling the whole SIMD context could be expensive, while chance of it used while calling a barrier are not high.

To be fair - the effects were never a target of too much research/measuring. Perhaps the current scheme is even completely OK.

@EgorBo

EgorBo commented Jun 1, 2026

Copy link
Copy Markdown
MemberAuthor

@VSadov It should be trivial to implement the JIT-EE I think, I can give it a try if Satori is anywhere near to be productized/merged into dotnet/runtime, something like CORINFO_WRITEBARRIER_CALLC getWriteBarrierCallingConvention();

The reason I decided to tune the registers for WB are potential regressions from my other change to remove CORINFO_HELP_ASSIGN_BYREF (#128542 + #128687) because that WB had a very compact calling convention with no spills between calls at all

@jkotas

Copy link
Copy Markdown
Member

We rely on C++ code that we call to not trash SIMD (by manual examination).

We used to do that in a few places and learned hard way that it is a bad idea...

@EgorBo
EgorBo enabled auto-merge (squash) June 2, 2026 21:01
@EgorBo
EgorBo merged commit db87610 into dotnet:mainJun 2, 2026
131 of 139 checks passed
@EgorBo

Copy link
Copy Markdown
MemberAuthor

@VSadov let me know if you want me to introduce an API that can switch CC for WB between this custom and standard (if we sure we'll merge Satori into runtime eventually)

I'm merging it to match the behavior with arm64

@VSadov

Copy link
Copy Markdown
Member

@VSadov let me know if you want me to introduce an API that can switch CC for WB between this custom and standard (if we sure we'll merge Satori into runtime eventually)

I'm merging it to match the behavior with arm64

I'll let you know. Right now Satori is a bit behind main. I think we will need to catch up to the state just before this PR first. (deleting byref barriers might be a good thing in terms of reducing complexity).

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@EgorBo@VSadov@jkotas@kg@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: don't kill FP/SIMD/mask regs across x64 write barriers - #128778

Merged
EgorBo merged 3 commits into
dotnet:mainfrom
EgorBo:jit-wb-preserve-fp
Jun 2, 2026
Merged

JIT: don't kill FP/SIMD/mask regs across x64 write barriers#128778
EgorBo merged 3 commits into
dotnet:mainfrom
EgorBo:jit-wb-preserve-fp

Conversation

@EgorBo

@EgorBoEgorBo commented May 29, 2026

Copy link
Copy Markdown
Member
{7366C871-08EB-48F1-897F-021AE10303FC}

What I deliberately did not do

  • Did not exclude RCX/RDI (dst) from the kill set — the asm helpers shift dst in place (shr rcx, 0Bh), so it must be considered clobbered. Preserving dst would require rewriting all ~20 asm variants.
  • Did not exclude RDX/RSI (src) — the Region variants destroy src.
  • Did not exclude R10/R11 on Windows even though the default patched-slot path doesn't touch them. Kept conservative to cover _DEBUG runtime variant (JIT_WriteBarrier_Debug) and the DOTNET_UseGCWriteBarrierCopy=0 config (RhpAssignRef), both of which touch R10/R11.

PS: We probably can do this for APX cc @dotnet/intel

I think it would be nice to preserve dst register unchanged (we do that for arm64), but that requires some changes in the ASM helpers.

The x64 JIT_WriteBarrier / JIT_CheckedWriteBarrier helpers (and all the
patched-slot variants: PreGrow/PostGrow/SVR/Region + WriteWatch flavors)
never execute any SSE/AVX/AVX-512/EVEX-mask instruction. So XMM/YMM/ZMM
and K mask registers can stay live across a write barrier call.
Narrow RBM_CALLEE_TRASH_WRITEBARRIER (and the matching GCTRASH mask)
from the full RBM_CALLEE_TRASH down to RBM_INT_CALLEE_TRASH_INIT on
amd64. Integer callee-trash regs remain conservatively in the kill set
to keep _DEBUG runtime builds and DOTNET_UseGCWriteBarrierCopy=0 working.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings May 29, 2026 16:56
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 29, 2026
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR narrows the amd64 JIT write-barrier register kill masks so FP/SIMD/mask registers can remain live across CORINFO_HELP_ASSIGN_REF and CORINFO_HELP_CHECKED_ASSIGN_REF, matching the documented helper ABI and reducing unnecessary spills.

Changes:

  • Documents the amd64 write-barrier helper register preservation/clobbering contract.
  • Changes amd64 write-barrier trash and GC-trash masks from full callee-trash to integer-only initial callee-trash registers.
  • Keeps conservative integer clobbers for runtime/debug write-barrier variants.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@MihuBot

@VSadov

Copy link
Copy Markdown
Member

Will this require write barriers that need to call into C++ code on some paths to spill more around the calls?

For context:
https://github.com/dotnet/runtimelab/blob/97a242e8b9a66df1f9de9ff16fc24c42035bcc98/src/coreclr/runtime/amd64/WriteBarriers.asm#L517-L526

@EgorBo

EgorBo commented May 30, 2026

Copy link
Copy Markdown
MemberAuthor

Will this require write barriers that need to call into C++ code on some paths to spill more around the calls?

For context: https://github.com/dotnet/runtimelab/blob/97a242e8b9a66df1f9de9ff16fc24c42035bcc98/src/coreclr/runtime/amd64/WriteBarriers.asm#L517-L526

Probably not since we cannot guarantee what is happening in a code compiled by C++. But if that is an issue we need a new JIT-EE API, because we already precisely track registers used by WB on arm64 today (and just removing that will lead to quite big regressions) and we've always been doing that for byref WB that I deleted recently

CopilotAI review requested due to automatic review settings May 30, 2026 01:16

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 1 out of 1 changed files in this pull request and generated 1 comment.

Comment threadsrc/coreclr/jit/targetamd64.h
Comment threadsrc/coreclr/jit/targetamd64.h
Comment threadsrc/coreclr/jit/targetamd64.h
kg
kg approved these changes Jun 1, 2026

@kgkg left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

vsadov's comment occurred to me as well, but otherwise this looks okay to me

@VSadov

Copy link
Copy Markdown
Member

What bothers me a bit is that this just encodes the current implementation that happens to be. If the eventual goal is to be able to switch GCs without rebuilding the entire runtime, the barrier calling conventions should be more "designed".

We already have to spill a lot more on arm64:
https://github.com/dotnet/runtimelab/blob/97a242e8b9a66df1f9de9ff16fc24c42035bcc98/src/coreclr/runtime/arm64/WriteBarriers.asm#L598-L609

 ;; because of the barrier call convention ;; we need to preserve caller-saved x0 through x15 and x29/x30 stp x29,x30,[sp,-16*9]! stp x0, x1,[sp,16*1] stp x2, x3,[sp,16*2] stp x4, x5,[sp,16*3] stp x6, x7,[sp,16*4] stp x8, x9,[sp,16*5] stp x10,x11,[sp,16*6] stp x12,x13,[sp,16*7] stp x14,x15,[sp,16*8]

On x64 the spill used to be fairly minimal in comparison.
Either way kind of worked, but which way is better? Are we trading binary size for throughput here?

Calls into C++ code are not on common hot paths, but even on uncommon paths spilling too much may become measurable.

I think we may want to revisit the barriers calling convention is at some point.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@VSadov I think we just need a JIT-EE API that will tell the JIT about the calling convention for write barriers (e.g. eithe just normal or custom) and Satori can return "normal" (on both x64 and arm64) and remove that spilling logic you have.
BTW, are you sure you also don't need to spill SIMD regs there? Because I think arm64 JIT assumes they're not killed too today, but it's possible that your C++ call just never use them either (but not something you can rely on).

Custom calling convention seems to be very proftiable, e.g. in this PR I tried just to avoid touching RCX/RDI reg (typically holding this) and got massive diffs in SPMI (-2.0Mb on SysV and -1.3Mb on win-x64).

@VSadov

VSadov commented Jun 1, 2026

Copy link
Copy Markdown
Member

@VSadov I think we just need a JIT-EE API that will tell the JIT about the calling convention for write barriers (e.g. eithe just normal or custom) and Satori can return "normal" (on both x64 and arm64) and remove that spilling logic you have. BTW, are you sure you also don't need to spill SIMD regs there? Because I think arm64 JIT assumes they're not killed too today, but it's possible that your C++ call just never use them either (but not something you can rely on).

Custom calling convention seems to be very proftiable, e.g. in this PR I tried just to avoid touching RCX/RDI reg (typically holding this) and got massive diffs in SPMI (-2.0Mb on SysV and -1.3Mb on win-x64).

I think standard calling convention may be too generous towards GC barriers. I understand the desire to leave registers for JIT to use.

However, the current conventions could be:

  • too tight that the asm part of the barrier may need to push/pop for extra temps.
  • unfriendly to C++ code.

It may be a good idea to make it configurable via JIT-EE API.
And I assume crossgen could just use the most conservative "common denominator" combination.

BTW, are you sure you also don't need to spill SIMD regs there?

We rely on C++ code that we call to not trash SIMD (by manual examination). On x64 we spill XMM0, because it does.
It is a bit of playing with fire.

SIMD is probably where I'd put the line. Spilling the whole SIMD context could be expensive, while chance of it used while calling a barrier are not high.

To be fair - the effects were never a target of too much research/measuring. Perhaps the current scheme is even completely OK.

@EgorBo

EgorBo commented Jun 1, 2026

Copy link
Copy Markdown
MemberAuthor

@VSadov It should be trivial to implement the JIT-EE I think, I can give it a try if Satori is anywhere near to be productized/merged into dotnet/runtime, something like CORINFO_WRITEBARRIER_CALLC getWriteBarrierCallingConvention();

The reason I decided to tune the registers for WB are potential regressions from my other change to remove CORINFO_HELP_ASSIGN_BYREF (#128542 + #128687) because that WB had a very compact calling convention with no spills between calls at all

@jkotas

Copy link
Copy Markdown
Member

We rely on C++ code that we call to not trash SIMD (by manual examination).

We used to do that in a few places and learned hard way that it is a bad idea...

@EgorBo
EgorBo enabled auto-merge (squash) June 2, 2026 21:01
@EgorBo
EgorBo merged commit db87610 into dotnet:mainJun 2, 2026
131 of 139 checks passed
@EgorBo

Copy link
Copy Markdown
MemberAuthor

@VSadov let me know if you want me to introduce an API that can switch CC for WB between this custom and standard (if we sure we'll merge Satori into runtime eventually)

I'm merging it to match the behavior with arm64

@VSadov

Copy link
Copy Markdown
Member

@VSadov let me know if you want me to introduce an API that can switch CC for WB between this custom and standard (if we sure we'll merge Satori into runtime eventually)

I'm merging it to match the behavior with arm64

I'll let you know. Right now Satori is a bit behind main. I think we will need to catch up to the state just before this PR first. (deleting byref barriers might be a good thing in terms of reducing complexity).

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@EgorBo@VSadov@jkotas@kg@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: don't kill FP/SIMD/mask regs across x64 write barriers - #128778

Merged
EgorBo merged 3 commits into
dotnet:mainfrom
EgorBo:jit-wb-preserve-fp
Jun 2, 2026
Merged

JIT: don't kill FP/SIMD/mask regs across x64 write barriers#128778
EgorBo merged 3 commits into
dotnet:mainfrom
EgorBo:jit-wb-preserve-fp

Conversation

@EgorBo

@EgorBoEgorBo commented May 29, 2026

Copy link
Copy Markdown
Member
{7366C871-08EB-48F1-897F-021AE10303FC}

What I deliberately did not do

  • Did not exclude RCX/RDI (dst) from the kill set — the asm helpers shift dst in place (shr rcx, 0Bh), so it must be considered clobbered. Preserving dst would require rewriting all ~20 asm variants.
  • Did not exclude RDX/RSI (src) — the Region variants destroy src.
  • Did not exclude R10/R11 on Windows even though the default patched-slot path doesn't touch them. Kept conservative to cover _DEBUG runtime variant (JIT_WriteBarrier_Debug) and the DOTNET_UseGCWriteBarrierCopy=0 config (RhpAssignRef), both of which touch R10/R11.

PS: We probably can do this for APX cc @dotnet/intel

I think it would be nice to preserve dst register unchanged (we do that for arm64), but that requires some changes in the ASM helpers.

The x64 JIT_WriteBarrier / JIT_CheckedWriteBarrier helpers (and all the
patched-slot variants: PreGrow/PostGrow/SVR/Region + WriteWatch flavors)
never execute any SSE/AVX/AVX-512/EVEX-mask instruction. So XMM/YMM/ZMM
and K mask registers can stay live across a write barrier call.
Narrow RBM_CALLEE_TRASH_WRITEBARRIER (and the matching GCTRASH mask)
from the full RBM_CALLEE_TRASH down to RBM_INT_CALLEE_TRASH_INIT on
amd64. Integer callee-trash regs remain conservatively in the kill set
to keep _DEBUG runtime builds and DOTNET_UseGCWriteBarrierCopy=0 working.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings May 29, 2026 16:56
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 29, 2026
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR narrows the amd64 JIT write-barrier register kill masks so FP/SIMD/mask registers can remain live across CORINFO_HELP_ASSIGN_REF and CORINFO_HELP_CHECKED_ASSIGN_REF, matching the documented helper ABI and reducing unnecessary spills.

Changes:

  • Documents the amd64 write-barrier helper register preservation/clobbering contract.
  • Changes amd64 write-barrier trash and GC-trash masks from full callee-trash to integer-only initial callee-trash registers.
  • Keeps conservative integer clobbers for runtime/debug write-barrier variants.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@MihuBot

@VSadov

Copy link
Copy Markdown
Member

Will this require write barriers that need to call into C++ code on some paths to spill more around the calls?

For context:
https://github.com/dotnet/runtimelab/blob/97a242e8b9a66df1f9de9ff16fc24c42035bcc98/src/coreclr/runtime/amd64/WriteBarriers.asm#L517-L526

@EgorBo

EgorBo commented May 30, 2026

Copy link
Copy Markdown
MemberAuthor

Will this require write barriers that need to call into C++ code on some paths to spill more around the calls?

For context: https://github.com/dotnet/runtimelab/blob/97a242e8b9a66df1f9de9ff16fc24c42035bcc98/src/coreclr/runtime/amd64/WriteBarriers.asm#L517-L526

Probably not since we cannot guarantee what is happening in a code compiled by C++. But if that is an issue we need a new JIT-EE API, because we already precisely track registers used by WB on arm64 today (and just removing that will lead to quite big regressions) and we've always been doing that for byref WB that I deleted recently

CopilotAI review requested due to automatic review settings May 30, 2026 01:16

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 1 out of 1 changed files in this pull request and generated 1 comment.

Comment threadsrc/coreclr/jit/targetamd64.h
Comment threadsrc/coreclr/jit/targetamd64.h
Comment threadsrc/coreclr/jit/targetamd64.h
kg
kg approved these changes Jun 1, 2026

@kgkg left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

vsadov's comment occurred to me as well, but otherwise this looks okay to me

@VSadov

Copy link
Copy Markdown
Member

What bothers me a bit is that this just encodes the current implementation that happens to be. If the eventual goal is to be able to switch GCs without rebuilding the entire runtime, the barrier calling conventions should be more "designed".

We already have to spill a lot more on arm64:
https://github.com/dotnet/runtimelab/blob/97a242e8b9a66df1f9de9ff16fc24c42035bcc98/src/coreclr/runtime/arm64/WriteBarriers.asm#L598-L609

 ;; because of the barrier call convention ;; we need to preserve caller-saved x0 through x15 and x29/x30 stp x29,x30,[sp,-16*9]! stp x0, x1,[sp,16*1] stp x2, x3,[sp,16*2] stp x4, x5,[sp,16*3] stp x6, x7,[sp,16*4] stp x8, x9,[sp,16*5] stp x10,x11,[sp,16*6] stp x12,x13,[sp,16*7] stp x14,x15,[sp,16*8]

On x64 the spill used to be fairly minimal in comparison.
Either way kind of worked, but which way is better? Are we trading binary size for throughput here?

Calls into C++ code are not on common hot paths, but even on uncommon paths spilling too much may become measurable.

I think we may want to revisit the barriers calling convention is at some point.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@VSadov I think we just need a JIT-EE API that will tell the JIT about the calling convention for write barriers (e.g. eithe just normal or custom) and Satori can return "normal" (on both x64 and arm64) and remove that spilling logic you have.
BTW, are you sure you also don't need to spill SIMD regs there? Because I think arm64 JIT assumes they're not killed too today, but it's possible that your C++ call just never use them either (but not something you can rely on).

Custom calling convention seems to be very proftiable, e.g. in this PR I tried just to avoid touching RCX/RDI reg (typically holding this) and got massive diffs in SPMI (-2.0Mb on SysV and -1.3Mb on win-x64).

@VSadov

VSadov commented Jun 1, 2026

Copy link
Copy Markdown
Member

@VSadov I think we just need a JIT-EE API that will tell the JIT about the calling convention for write barriers (e.g. eithe just normal or custom) and Satori can return "normal" (on both x64 and arm64) and remove that spilling logic you have. BTW, are you sure you also don't need to spill SIMD regs there? Because I think arm64 JIT assumes they're not killed too today, but it's possible that your C++ call just never use them either (but not something you can rely on).

Custom calling convention seems to be very proftiable, e.g. in this PR I tried just to avoid touching RCX/RDI reg (typically holding this) and got massive diffs in SPMI (-2.0Mb on SysV and -1.3Mb on win-x64).

I think standard calling convention may be too generous towards GC barriers. I understand the desire to leave registers for JIT to use.

However, the current conventions could be:

  • too tight that the asm part of the barrier may need to push/pop for extra temps.
  • unfriendly to C++ code.

It may be a good idea to make it configurable via JIT-EE API.
And I assume crossgen could just use the most conservative "common denominator" combination.

BTW, are you sure you also don't need to spill SIMD regs there?

We rely on C++ code that we call to not trash SIMD (by manual examination). On x64 we spill XMM0, because it does.
It is a bit of playing with fire.

SIMD is probably where I'd put the line. Spilling the whole SIMD context could be expensive, while chance of it used while calling a barrier are not high.

To be fair - the effects were never a target of too much research/measuring. Perhaps the current scheme is even completely OK.

@EgorBo

EgorBo commented Jun 1, 2026

Copy link
Copy Markdown
MemberAuthor

@VSadov It should be trivial to implement the JIT-EE I think, I can give it a try if Satori is anywhere near to be productized/merged into dotnet/runtime, something like CORINFO_WRITEBARRIER_CALLC getWriteBarrierCallingConvention();

The reason I decided to tune the registers for WB are potential regressions from my other change to remove CORINFO_HELP_ASSIGN_BYREF (#128542 + #128687) because that WB had a very compact calling convention with no spills between calls at all

@jkotas

Copy link
Copy Markdown
Member

We rely on C++ code that we call to not trash SIMD (by manual examination).

We used to do that in a few places and learned hard way that it is a bad idea...

@EgorBo
EgorBo enabled auto-merge (squash) June 2, 2026 21:01
@EgorBo
EgorBo merged commit db87610 into dotnet:mainJun 2, 2026
131 of 139 checks passed
@EgorBo

Copy link
Copy Markdown
MemberAuthor

@VSadov let me know if you want me to introduce an API that can switch CC for WB between this custom and standard (if we sure we'll merge Satori into runtime eventually)

I'm merging it to match the behavior with arm64

@VSadov

Copy link
Copy Markdown
Member

@VSadov let me know if you want me to introduce an API that can switch CC for WB between this custom and standard (if we sure we'll merge Satori into runtime eventually)

I'm merging it to match the behavior with arm64

I'll let you know. Right now Satori is a bit behind main. I think we will need to catch up to the state just before this PR first. (deleting byref barriers might be a good thing in terms of reducing complexity).

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@EgorBo@VSadov@jkotas@kg@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

JIT: don't kill FP/SIMD/mask regs across x64 write barriers - #128778

Merged
EgorBo merged 3 commits into
dotnet:mainfrom
EgorBo:jit-wb-preserve-fp
Jun 2, 2026
Merged

JIT: don't kill FP/SIMD/mask regs across x64 write barriers#128778
EgorBo merged 3 commits into
dotnet:mainfrom
EgorBo:jit-wb-preserve-fp

Conversation

@EgorBo

@EgorBoEgorBo commented May 29, 2026

Copy link
Copy Markdown
Member
{7366C871-08EB-48F1-897F-021AE10303FC}

What I deliberately did not do

  • Did not exclude RCX/RDI (dst) from the kill set — the asm helpers shift dst in place (shr rcx, 0Bh), so it must be considered clobbered. Preserving dst would require rewriting all ~20 asm variants.
  • Did not exclude RDX/RSI (src) — the Region variants destroy src.
  • Did not exclude R10/R11 on Windows even though the default patched-slot path doesn't touch them. Kept conservative to cover _DEBUG runtime variant (JIT_WriteBarrier_Debug) and the DOTNET_UseGCWriteBarrierCopy=0 config (RhpAssignRef), both of which touch R10/R11.

PS: We probably can do this for APX cc @dotnet/intel

I think it would be nice to preserve dst register unchanged (we do that for arm64), but that requires some changes in the ASM helpers.

The x64 JIT_WriteBarrier / JIT_CheckedWriteBarrier helpers (and all the
patched-slot variants: PreGrow/PostGrow/SVR/Region + WriteWatch flavors)
never execute any SSE/AVX/AVX-512/EVEX-mask instruction. So XMM/YMM/ZMM
and K mask registers can stay live across a write barrier call.
Narrow RBM_CALLEE_TRASH_WRITEBARRIER (and the matching GCTRASH mask)
from the full RBM_CALLEE_TRASH down to RBM_INT_CALLEE_TRASH_INIT on
amd64. Integer callee-trash regs remain conservatively in the kill set
to keep _DEBUG runtime builds and DOTNET_UseGCWriteBarrierCopy=0 working.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings May 29, 2026 16:56
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 29, 2026
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR narrows the amd64 JIT write-barrier register kill masks so FP/SIMD/mask registers can remain live across CORINFO_HELP_ASSIGN_REF and CORINFO_HELP_CHECKED_ASSIGN_REF, matching the documented helper ABI and reducing unnecessary spills.

Changes:

  • Documents the amd64 write-barrier helper register preservation/clobbering contract.
  • Changes amd64 write-barrier trash and GC-trash masks from full callee-trash to integer-only initial callee-trash registers.
  • Keeps conservative integer clobbers for runtime/debug write-barrier variants.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@MihuBot

@VSadov

Copy link
Copy Markdown
Member

Will this require write barriers that need to call into C++ code on some paths to spill more around the calls?

For context:
https://github.com/dotnet/runtimelab/blob/97a242e8b9a66df1f9de9ff16fc24c42035bcc98/src/coreclr/runtime/amd64/WriteBarriers.asm#L517-L526

@EgorBo

EgorBo commented May 30, 2026

Copy link
Copy Markdown
MemberAuthor

Will this require write barriers that need to call into C++ code on some paths to spill more around the calls?

For context: https://github.com/dotnet/runtimelab/blob/97a242e8b9a66df1f9de9ff16fc24c42035bcc98/src/coreclr/runtime/amd64/WriteBarriers.asm#L517-L526

Probably not since we cannot guarantee what is happening in a code compiled by C++. But if that is an issue we need a new JIT-EE API, because we already precisely track registers used by WB on arm64 today (and just removing that will lead to quite big regressions) and we've always been doing that for byref WB that I deleted recently

CopilotAI review requested due to automatic review settings May 30, 2026 01:16

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 1 out of 1 changed files in this pull request and generated 1 comment.

Comment threadsrc/coreclr/jit/targetamd64.h
Comment threadsrc/coreclr/jit/targetamd64.h
Comment threadsrc/coreclr/jit/targetamd64.h
kg
kg approved these changes Jun 1, 2026

@kgkg left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

vsadov's comment occurred to me as well, but otherwise this looks okay to me

@VSadov

Copy link
Copy Markdown
Member

What bothers me a bit is that this just encodes the current implementation that happens to be. If the eventual goal is to be able to switch GCs without rebuilding the entire runtime, the barrier calling conventions should be more "designed".

We already have to spill a lot more on arm64:
https://github.com/dotnet/runtimelab/blob/97a242e8b9a66df1f9de9ff16fc24c42035bcc98/src/coreclr/runtime/arm64/WriteBarriers.asm#L598-L609

 ;; because of the barrier call convention ;; we need to preserve caller-saved x0 through x15 and x29/x30 stp x29,x30,[sp,-16*9]! stp x0, x1,[sp,16*1] stp x2, x3,[sp,16*2] stp x4, x5,[sp,16*3] stp x6, x7,[sp,16*4] stp x8, x9,[sp,16*5] stp x10,x11,[sp,16*6] stp x12,x13,[sp,16*7] stp x14,x15,[sp,16*8]

On x64 the spill used to be fairly minimal in comparison.
Either way kind of worked, but which way is better? Are we trading binary size for throughput here?

Calls into C++ code are not on common hot paths, but even on uncommon paths spilling too much may become measurable.

I think we may want to revisit the barriers calling convention is at some point.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@VSadov I think we just need a JIT-EE API that will tell the JIT about the calling convention for write barriers (e.g. eithe just normal or custom) and Satori can return "normal" (on both x64 and arm64) and remove that spilling logic you have.
BTW, are you sure you also don't need to spill SIMD regs there? Because I think arm64 JIT assumes they're not killed too today, but it's possible that your C++ call just never use them either (but not something you can rely on).

Custom calling convention seems to be very proftiable, e.g. in this PR I tried just to avoid touching RCX/RDI reg (typically holding this) and got massive diffs in SPMI (-2.0Mb on SysV and -1.3Mb on win-x64).

@VSadov

VSadov commented Jun 1, 2026

Copy link
Copy Markdown
Member

@VSadov I think we just need a JIT-EE API that will tell the JIT about the calling convention for write barriers (e.g. eithe just normal or custom) and Satori can return "normal" (on both x64 and arm64) and remove that spilling logic you have. BTW, are you sure you also don't need to spill SIMD regs there? Because I think arm64 JIT assumes they're not killed too today, but it's possible that your C++ call just never use them either (but not something you can rely on).

Custom calling convention seems to be very proftiable, e.g. in this PR I tried just to avoid touching RCX/RDI reg (typically holding this) and got massive diffs in SPMI (-2.0Mb on SysV and -1.3Mb on win-x64).

I think standard calling convention may be too generous towards GC barriers. I understand the desire to leave registers for JIT to use.

However, the current conventions could be:

  • too tight that the asm part of the barrier may need to push/pop for extra temps.
  • unfriendly to C++ code.

It may be a good idea to make it configurable via JIT-EE API.
And I assume crossgen could just use the most conservative "common denominator" combination.

BTW, are you sure you also don't need to spill SIMD regs there?

We rely on C++ code that we call to not trash SIMD (by manual examination). On x64 we spill XMM0, because it does.
It is a bit of playing with fire.

SIMD is probably where I'd put the line. Spilling the whole SIMD context could be expensive, while chance of it used while calling a barrier are not high.

To be fair - the effects were never a target of too much research/measuring. Perhaps the current scheme is even completely OK.

@EgorBo

EgorBo commented Jun 1, 2026

Copy link
Copy Markdown
MemberAuthor

@VSadov It should be trivial to implement the JIT-EE I think, I can give it a try if Satori is anywhere near to be productized/merged into dotnet/runtime, something like CORINFO_WRITEBARRIER_CALLC getWriteBarrierCallingConvention();

The reason I decided to tune the registers for WB are potential regressions from my other change to remove CORINFO_HELP_ASSIGN_BYREF (#128542 + #128687) because that WB had a very compact calling convention with no spills between calls at all

@jkotas

Copy link
Copy Markdown
Member

We rely on C++ code that we call to not trash SIMD (by manual examination).

We used to do that in a few places and learned hard way that it is a bad idea...

@EgorBo
EgorBo enabled auto-merge (squash) June 2, 2026 21:01
@EgorBo
EgorBo merged commit db87610 into dotnet:mainJun 2, 2026
131 of 139 checks passed
@EgorBo

Copy link
Copy Markdown
MemberAuthor

@VSadov let me know if you want me to introduce an API that can switch CC for WB between this custom and standard (if we sure we'll merge Satori into runtime eventually)

I'm merging it to match the behavior with arm64

@VSadov

Copy link
Copy Markdown
Member

@VSadov let me know if you want me to introduce an API that can switch CC for WB between this custom and standard (if we sure we'll merge Satori into runtime eventually)

I'm merging it to match the behavior with arm64

I'll let you know. Right now Satori is a bit behind main. I think we will need to catch up to the state just before this PR first. (deleting byref barriers might be a good thing in terms of reducing complexity).

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@EgorBo@VSadov@jkotas@kg@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: don't kill FP/SIMD/mask regs across x64 write barriers - #128778

Merged
EgorBo merged 3 commits into
dotnet:mainfrom
EgorBo:jit-wb-preserve-fp
Jun 2, 2026
Merged

JIT: don't kill FP/SIMD/mask regs across x64 write barriers#128778
EgorBo merged 3 commits into
dotnet:mainfrom
EgorBo:jit-wb-preserve-fp

Conversation

@EgorBo

@EgorBoEgorBo commented May 29, 2026

Copy link
Copy Markdown
Member
{7366C871-08EB-48F1-897F-021AE10303FC}

What I deliberately did not do

  • Did not exclude RCX/RDI (dst) from the kill set — the asm helpers shift dst in place (shr rcx, 0Bh), so it must be considered clobbered. Preserving dst would require rewriting all ~20 asm variants.
  • Did not exclude RDX/RSI (src) — the Region variants destroy src.
  • Did not exclude R10/R11 on Windows even though the default patched-slot path doesn't touch them. Kept conservative to cover _DEBUG runtime variant (JIT_WriteBarrier_Debug) and the DOTNET_UseGCWriteBarrierCopy=0 config (RhpAssignRef), both of which touch R10/R11.

PS: We probably can do this for APX cc @dotnet/intel

I think it would be nice to preserve dst register unchanged (we do that for arm64), but that requires some changes in the ASM helpers.

The x64 JIT_WriteBarrier / JIT_CheckedWriteBarrier helpers (and all the
patched-slot variants: PreGrow/PostGrow/SVR/Region + WriteWatch flavors)
never execute any SSE/AVX/AVX-512/EVEX-mask instruction. So XMM/YMM/ZMM
and K mask registers can stay live across a write barrier call.
Narrow RBM_CALLEE_TRASH_WRITEBARRIER (and the matching GCTRASH mask)
from the full RBM_CALLEE_TRASH down to RBM_INT_CALLEE_TRASH_INIT on
amd64. Integer callee-trash regs remain conservatively in the kill set
to keep _DEBUG runtime builds and DOTNET_UseGCWriteBarrierCopy=0 working.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings May 29, 2026 16:56
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 29, 2026
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR narrows the amd64 JIT write-barrier register kill masks so FP/SIMD/mask registers can remain live across CORINFO_HELP_ASSIGN_REF and CORINFO_HELP_CHECKED_ASSIGN_REF, matching the documented helper ABI and reducing unnecessary spills.

Changes:

  • Documents the amd64 write-barrier helper register preservation/clobbering contract.
  • Changes amd64 write-barrier trash and GC-trash masks from full callee-trash to integer-only initial callee-trash registers.
  • Keeps conservative integer clobbers for runtime/debug write-barrier variants.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@MihuBot

@VSadov

Copy link
Copy Markdown
Member

Will this require write barriers that need to call into C++ code on some paths to spill more around the calls?

For context:
https://github.com/dotnet/runtimelab/blob/97a242e8b9a66df1f9de9ff16fc24c42035bcc98/src/coreclr/runtime/amd64/WriteBarriers.asm#L517-L526

@EgorBo

EgorBo commented May 30, 2026

Copy link
Copy Markdown
MemberAuthor

Will this require write barriers that need to call into C++ code on some paths to spill more around the calls?

For context: https://github.com/dotnet/runtimelab/blob/97a242e8b9a66df1f9de9ff16fc24c42035bcc98/src/coreclr/runtime/amd64/WriteBarriers.asm#L517-L526

Probably not since we cannot guarantee what is happening in a code compiled by C++. But if that is an issue we need a new JIT-EE API, because we already precisely track registers used by WB on arm64 today (and just removing that will lead to quite big regressions) and we've always been doing that for byref WB that I deleted recently

CopilotAI review requested due to automatic review settings May 30, 2026 01:16

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 1 out of 1 changed files in this pull request and generated 1 comment.

Comment threadsrc/coreclr/jit/targetamd64.h
Comment threadsrc/coreclr/jit/targetamd64.h
Comment threadsrc/coreclr/jit/targetamd64.h
kg
kg approved these changes Jun 1, 2026

@kgkg left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

vsadov's comment occurred to me as well, but otherwise this looks okay to me

@VSadov

Copy link
Copy Markdown
Member

What bothers me a bit is that this just encodes the current implementation that happens to be. If the eventual goal is to be able to switch GCs without rebuilding the entire runtime, the barrier calling conventions should be more "designed".

We already have to spill a lot more on arm64:
https://github.com/dotnet/runtimelab/blob/97a242e8b9a66df1f9de9ff16fc24c42035bcc98/src/coreclr/runtime/arm64/WriteBarriers.asm#L598-L609

 ;; because of the barrier call convention ;; we need to preserve caller-saved x0 through x15 and x29/x30 stp x29,x30,[sp,-16*9]! stp x0, x1,[sp,16*1] stp x2, x3,[sp,16*2] stp x4, x5,[sp,16*3] stp x6, x7,[sp,16*4] stp x8, x9,[sp,16*5] stp x10,x11,[sp,16*6] stp x12,x13,[sp,16*7] stp x14,x15,[sp,16*8]

On x64 the spill used to be fairly minimal in comparison.
Either way kind of worked, but which way is better? Are we trading binary size for throughput here?

Calls into C++ code are not on common hot paths, but even on uncommon paths spilling too much may become measurable.

I think we may want to revisit the barriers calling convention is at some point.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@VSadov I think we just need a JIT-EE API that will tell the JIT about the calling convention for write barriers (e.g. eithe just normal or custom) and Satori can return "normal" (on both x64 and arm64) and remove that spilling logic you have.
BTW, are you sure you also don't need to spill SIMD regs there? Because I think arm64 JIT assumes they're not killed too today, but it's possible that your C++ call just never use them either (but not something you can rely on).

Custom calling convention seems to be very proftiable, e.g. in this PR I tried just to avoid touching RCX/RDI reg (typically holding this) and got massive diffs in SPMI (-2.0Mb on SysV and -1.3Mb on win-x64).

@VSadov

VSadov commented Jun 1, 2026

Copy link
Copy Markdown
Member

@VSadov I think we just need a JIT-EE API that will tell the JIT about the calling convention for write barriers (e.g. eithe just normal or custom) and Satori can return "normal" (on both x64 and arm64) and remove that spilling logic you have. BTW, are you sure you also don't need to spill SIMD regs there? Because I think arm64 JIT assumes they're not killed too today, but it's possible that your C++ call just never use them either (but not something you can rely on).

Custom calling convention seems to be very proftiable, e.g. in this PR I tried just to avoid touching RCX/RDI reg (typically holding this) and got massive diffs in SPMI (-2.0Mb on SysV and -1.3Mb on win-x64).

I think standard calling convention may be too generous towards GC barriers. I understand the desire to leave registers for JIT to use.

However, the current conventions could be:

  • too tight that the asm part of the barrier may need to push/pop for extra temps.
  • unfriendly to C++ code.

It may be a good idea to make it configurable via JIT-EE API.
And I assume crossgen could just use the most conservative "common denominator" combination.

BTW, are you sure you also don't need to spill SIMD regs there?

We rely on C++ code that we call to not trash SIMD (by manual examination). On x64 we spill XMM0, because it does.
It is a bit of playing with fire.

SIMD is probably where I'd put the line. Spilling the whole SIMD context could be expensive, while chance of it used while calling a barrier are not high.

To be fair - the effects were never a target of too much research/measuring. Perhaps the current scheme is even completely OK.

@EgorBo

EgorBo commented Jun 1, 2026

Copy link
Copy Markdown
MemberAuthor

@VSadov It should be trivial to implement the JIT-EE I think, I can give it a try if Satori is anywhere near to be productized/merged into dotnet/runtime, something like CORINFO_WRITEBARRIER_CALLC getWriteBarrierCallingConvention();

The reason I decided to tune the registers for WB are potential regressions from my other change to remove CORINFO_HELP_ASSIGN_BYREF (#128542 + #128687) because that WB had a very compact calling convention with no spills between calls at all

@jkotas

Copy link
Copy Markdown
Member

We rely on C++ code that we call to not trash SIMD (by manual examination).

We used to do that in a few places and learned hard way that it is a bad idea...

@EgorBo
EgorBo enabled auto-merge (squash) June 2, 2026 21:01
@EgorBo
EgorBo merged commit db87610 into dotnet:mainJun 2, 2026
131 of 139 checks passed
@EgorBo

Copy link
Copy Markdown
MemberAuthor

@VSadov let me know if you want me to introduce an API that can switch CC for WB between this custom and standard (if we sure we'll merge Satori into runtime eventually)

I'm merging it to match the behavior with arm64

@VSadov

Copy link
Copy Markdown
Member

@VSadov let me know if you want me to introduce an API that can switch CC for WB between this custom and standard (if we sure we'll merge Satori into runtime eventually)

I'm merging it to match the behavior with arm64

I'll let you know. Right now Satori is a bit behind main. I think we will need to catch up to the state just before this PR first. (deleting byref barriers might be a good thing in terms of reducing complexity).

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@EgorBo@VSadov@jkotas@kg@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: don't kill FP/SIMD/mask regs across x64 write barriers - #128778

Merged
EgorBo merged 3 commits into
dotnet:mainfrom
EgorBo:jit-wb-preserve-fp
Jun 2, 2026
Merged

JIT: don't kill FP/SIMD/mask regs across x64 write barriers#128778
EgorBo merged 3 commits into
dotnet:mainfrom
EgorBo:jit-wb-preserve-fp

Conversation

@EgorBo

@EgorBoEgorBo commented May 29, 2026

Copy link
Copy Markdown
Member
{7366C871-08EB-48F1-897F-021AE10303FC}

What I deliberately did not do

  • Did not exclude RCX/RDI (dst) from the kill set — the asm helpers shift dst in place (shr rcx, 0Bh), so it must be considered clobbered. Preserving dst would require rewriting all ~20 asm variants.
  • Did not exclude RDX/RSI (src) — the Region variants destroy src.
  • Did not exclude R10/R11 on Windows even though the default patched-slot path doesn't touch them. Kept conservative to cover _DEBUG runtime variant (JIT_WriteBarrier_Debug) and the DOTNET_UseGCWriteBarrierCopy=0 config (RhpAssignRef), both of which touch R10/R11.

PS: We probably can do this for APX cc @dotnet/intel

I think it would be nice to preserve dst register unchanged (we do that for arm64), but that requires some changes in the ASM helpers.

The x64 JIT_WriteBarrier / JIT_CheckedWriteBarrier helpers (and all the
patched-slot variants: PreGrow/PostGrow/SVR/Region + WriteWatch flavors)
never execute any SSE/AVX/AVX-512/EVEX-mask instruction. So XMM/YMM/ZMM
and K mask registers can stay live across a write barrier call.
Narrow RBM_CALLEE_TRASH_WRITEBARRIER (and the matching GCTRASH mask)
from the full RBM_CALLEE_TRASH down to RBM_INT_CALLEE_TRASH_INIT on
amd64. Integer callee-trash regs remain conservatively in the kill set
to keep _DEBUG runtime builds and DOTNET_UseGCWriteBarrierCopy=0 working.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings May 29, 2026 16:56
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 29, 2026
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR narrows the amd64 JIT write-barrier register kill masks so FP/SIMD/mask registers can remain live across CORINFO_HELP_ASSIGN_REF and CORINFO_HELP_CHECKED_ASSIGN_REF, matching the documented helper ABI and reducing unnecessary spills.

Changes:

  • Documents the amd64 write-barrier helper register preservation/clobbering contract.
  • Changes amd64 write-barrier trash and GC-trash masks from full callee-trash to integer-only initial callee-trash registers.
  • Keeps conservative integer clobbers for runtime/debug write-barrier variants.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@MihuBot

@VSadov

Copy link
Copy Markdown
Member

Will this require write barriers that need to call into C++ code on some paths to spill more around the calls?

For context:
https://github.com/dotnet/runtimelab/blob/97a242e8b9a66df1f9de9ff16fc24c42035bcc98/src/coreclr/runtime/amd64/WriteBarriers.asm#L517-L526

@EgorBo

EgorBo commented May 30, 2026

Copy link
Copy Markdown
MemberAuthor

Will this require write barriers that need to call into C++ code on some paths to spill more around the calls?

For context: https://github.com/dotnet/runtimelab/blob/97a242e8b9a66df1f9de9ff16fc24c42035bcc98/src/coreclr/runtime/amd64/WriteBarriers.asm#L517-L526

Probably not since we cannot guarantee what is happening in a code compiled by C++. But if that is an issue we need a new JIT-EE API, because we already precisely track registers used by WB on arm64 today (and just removing that will lead to quite big regressions) and we've always been doing that for byref WB that I deleted recently

CopilotAI review requested due to automatic review settings May 30, 2026 01:16

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 1 out of 1 changed files in this pull request and generated 1 comment.

Comment threadsrc/coreclr/jit/targetamd64.h
Comment threadsrc/coreclr/jit/targetamd64.h
Comment threadsrc/coreclr/jit/targetamd64.h
kg
kg approved these changes Jun 1, 2026

@kgkg left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

vsadov's comment occurred to me as well, but otherwise this looks okay to me

@VSadov

Copy link
Copy Markdown
Member

What bothers me a bit is that this just encodes the current implementation that happens to be. If the eventual goal is to be able to switch GCs without rebuilding the entire runtime, the barrier calling conventions should be more "designed".

We already have to spill a lot more on arm64:
https://github.com/dotnet/runtimelab/blob/97a242e8b9a66df1f9de9ff16fc24c42035bcc98/src/coreclr/runtime/arm64/WriteBarriers.asm#L598-L609

 ;; because of the barrier call convention ;; we need to preserve caller-saved x0 through x15 and x29/x30 stp x29,x30,[sp,-16*9]! stp x0, x1,[sp,16*1] stp x2, x3,[sp,16*2] stp x4, x5,[sp,16*3] stp x6, x7,[sp,16*4] stp x8, x9,[sp,16*5] stp x10,x11,[sp,16*6] stp x12,x13,[sp,16*7] stp x14,x15,[sp,16*8]

On x64 the spill used to be fairly minimal in comparison.
Either way kind of worked, but which way is better? Are we trading binary size for throughput here?

Calls into C++ code are not on common hot paths, but even on uncommon paths spilling too much may become measurable.

I think we may want to revisit the barriers calling convention is at some point.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@VSadov I think we just need a JIT-EE API that will tell the JIT about the calling convention for write barriers (e.g. eithe just normal or custom) and Satori can return "normal" (on both x64 and arm64) and remove that spilling logic you have.
BTW, are you sure you also don't need to spill SIMD regs there? Because I think arm64 JIT assumes they're not killed too today, but it's possible that your C++ call just never use them either (but not something you can rely on).

Custom calling convention seems to be very proftiable, e.g. in this PR I tried just to avoid touching RCX/RDI reg (typically holding this) and got massive diffs in SPMI (-2.0Mb on SysV and -1.3Mb on win-x64).

@VSadov

VSadov commented Jun 1, 2026

Copy link
Copy Markdown
Member

@VSadov I think we just need a JIT-EE API that will tell the JIT about the calling convention for write barriers (e.g. eithe just normal or custom) and Satori can return "normal" (on both x64 and arm64) and remove that spilling logic you have. BTW, are you sure you also don't need to spill SIMD regs there? Because I think arm64 JIT assumes they're not killed too today, but it's possible that your C++ call just never use them either (but not something you can rely on).

Custom calling convention seems to be very proftiable, e.g. in this PR I tried just to avoid touching RCX/RDI reg (typically holding this) and got massive diffs in SPMI (-2.0Mb on SysV and -1.3Mb on win-x64).

I think standard calling convention may be too generous towards GC barriers. I understand the desire to leave registers for JIT to use.

However, the current conventions could be:

  • too tight that the asm part of the barrier may need to push/pop for extra temps.
  • unfriendly to C++ code.

It may be a good idea to make it configurable via JIT-EE API.
And I assume crossgen could just use the most conservative "common denominator" combination.

BTW, are you sure you also don't need to spill SIMD regs there?

We rely on C++ code that we call to not trash SIMD (by manual examination). On x64 we spill XMM0, because it does.
It is a bit of playing with fire.

SIMD is probably where I'd put the line. Spilling the whole SIMD context could be expensive, while chance of it used while calling a barrier are not high.

To be fair - the effects were never a target of too much research/measuring. Perhaps the current scheme is even completely OK.

@EgorBo

EgorBo commented Jun 1, 2026

Copy link
Copy Markdown
MemberAuthor

@VSadov It should be trivial to implement the JIT-EE I think, I can give it a try if Satori is anywhere near to be productized/merged into dotnet/runtime, something like CORINFO_WRITEBARRIER_CALLC getWriteBarrierCallingConvention();

The reason I decided to tune the registers for WB are potential regressions from my other change to remove CORINFO_HELP_ASSIGN_BYREF (#128542 + #128687) because that WB had a very compact calling convention with no spills between calls at all

@jkotas

Copy link
Copy Markdown
Member

We rely on C++ code that we call to not trash SIMD (by manual examination).

We used to do that in a few places and learned hard way that it is a bad idea...

@EgorBo
EgorBo enabled auto-merge (squash) June 2, 2026 21:01
@EgorBo
EgorBo merged commit db87610 into dotnet:mainJun 2, 2026
131 of 139 checks passed
@EgorBo

Copy link
Copy Markdown
MemberAuthor

@VSadov let me know if you want me to introduce an API that can switch CC for WB between this custom and standard (if we sure we'll merge Satori into runtime eventually)

I'm merging it to match the behavior with arm64

@VSadov

Copy link
Copy Markdown
Member

@VSadov let me know if you want me to introduce an API that can switch CC for WB between this custom and standard (if we sure we'll merge Satori into runtime eventually)

I'm merging it to match the behavior with arm64

I'll let you know. Right now Satori is a bit behind main. I think we will need to catch up to the state just before this PR first. (deleting byref barriers might be a good thing in terms of reducing complexity).

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@EgorBo@VSadov@jkotas@kg@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

JIT: don't kill FP/SIMD/mask regs across x64 write barriers - #128778

Merged
EgorBo merged 3 commits into
dotnet:mainfrom
EgorBo:jit-wb-preserve-fp
Jun 2, 2026
Merged

JIT: don't kill FP/SIMD/mask regs across x64 write barriers#128778
EgorBo merged 3 commits into
dotnet:mainfrom
EgorBo:jit-wb-preserve-fp

Conversation

@EgorBo

@EgorBoEgorBo commented May 29, 2026

Copy link
Copy Markdown
Member
{7366C871-08EB-48F1-897F-021AE10303FC}

What I deliberately did not do

  • Did not exclude RCX/RDI (dst) from the kill set — the asm helpers shift dst in place (shr rcx, 0Bh), so it must be considered clobbered. Preserving dst would require rewriting all ~20 asm variants.
  • Did not exclude RDX/RSI (src) — the Region variants destroy src.
  • Did not exclude R10/R11 on Windows even though the default patched-slot path doesn't touch them. Kept conservative to cover _DEBUG runtime variant (JIT_WriteBarrier_Debug) and the DOTNET_UseGCWriteBarrierCopy=0 config (RhpAssignRef), both of which touch R10/R11.

PS: We probably can do this for APX cc @dotnet/intel

I think it would be nice to preserve dst register unchanged (we do that for arm64), but that requires some changes in the ASM helpers.

The x64 JIT_WriteBarrier / JIT_CheckedWriteBarrier helpers (and all the
patched-slot variants: PreGrow/PostGrow/SVR/Region + WriteWatch flavors)
never execute any SSE/AVX/AVX-512/EVEX-mask instruction. So XMM/YMM/ZMM
and K mask registers can stay live across a write barrier call.
Narrow RBM_CALLEE_TRASH_WRITEBARRIER (and the matching GCTRASH mask)
from the full RBM_CALLEE_TRASH down to RBM_INT_CALLEE_TRASH_INIT on
amd64. Integer callee-trash regs remain conservatively in the kill set
to keep _DEBUG runtime builds and DOTNET_UseGCWriteBarrierCopy=0 working.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
CopilotAI review requested due to automatic review settings May 29, 2026 16:56
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 29, 2026
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR narrows the amd64 JIT write-barrier register kill masks so FP/SIMD/mask registers can remain live across CORINFO_HELP_ASSIGN_REF and CORINFO_HELP_CHECKED_ASSIGN_REF, matching the documented helper ABI and reducing unnecessary spills.

Changes:

  • Documents the amd64 write-barrier helper register preservation/clobbering contract.
  • Changes amd64 write-barrier trash and GC-trash masks from full callee-trash to integer-only initial callee-trash registers.
  • Keeps conservative integer clobbers for runtime/debug write-barrier variants.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@MihuBot

@VSadov

Copy link
Copy Markdown
Member

Will this require write barriers that need to call into C++ code on some paths to spill more around the calls?

For context:
https://github.com/dotnet/runtimelab/blob/97a242e8b9a66df1f9de9ff16fc24c42035bcc98/src/coreclr/runtime/amd64/WriteBarriers.asm#L517-L526

@EgorBo

EgorBo commented May 30, 2026

Copy link
Copy Markdown
MemberAuthor

Will this require write barriers that need to call into C++ code on some paths to spill more around the calls?

For context: https://github.com/dotnet/runtimelab/blob/97a242e8b9a66df1f9de9ff16fc24c42035bcc98/src/coreclr/runtime/amd64/WriteBarriers.asm#L517-L526

Probably not since we cannot guarantee what is happening in a code compiled by C++. But if that is an issue we need a new JIT-EE API, because we already precisely track registers used by WB on arm64 today (and just removing that will lead to quite big regressions) and we've always been doing that for byref WB that I deleted recently

CopilotAI review requested due to automatic review settings May 30, 2026 01:16

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 1 out of 1 changed files in this pull request and generated 1 comment.

Comment threadsrc/coreclr/jit/targetamd64.h
Comment threadsrc/coreclr/jit/targetamd64.h
Comment threadsrc/coreclr/jit/targetamd64.h
kg
kg approved these changes Jun 1, 2026

@kgkg left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

vsadov's comment occurred to me as well, but otherwise this looks okay to me

@VSadov

Copy link
Copy Markdown
Member

What bothers me a bit is that this just encodes the current implementation that happens to be. If the eventual goal is to be able to switch GCs without rebuilding the entire runtime, the barrier calling conventions should be more "designed".

We already have to spill a lot more on arm64:
https://github.com/dotnet/runtimelab/blob/97a242e8b9a66df1f9de9ff16fc24c42035bcc98/src/coreclr/runtime/arm64/WriteBarriers.asm#L598-L609

 ;; because of the barrier call convention ;; we need to preserve caller-saved x0 through x15 and x29/x30 stp x29,x30,[sp,-16*9]! stp x0, x1,[sp,16*1] stp x2, x3,[sp,16*2] stp x4, x5,[sp,16*3] stp x6, x7,[sp,16*4] stp x8, x9,[sp,16*5] stp x10,x11,[sp,16*6] stp x12,x13,[sp,16*7] stp x14,x15,[sp,16*8]

On x64 the spill used to be fairly minimal in comparison.
Either way kind of worked, but which way is better? Are we trading binary size for throughput here?

Calls into C++ code are not on common hot paths, but even on uncommon paths spilling too much may become measurable.

I think we may want to revisit the barriers calling convention is at some point.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@VSadov I think we just need a JIT-EE API that will tell the JIT about the calling convention for write barriers (e.g. eithe just normal or custom) and Satori can return "normal" (on both x64 and arm64) and remove that spilling logic you have.
BTW, are you sure you also don't need to spill SIMD regs there? Because I think arm64 JIT assumes they're not killed too today, but it's possible that your C++ call just never use them either (but not something you can rely on).

Custom calling convention seems to be very proftiable, e.g. in this PR I tried just to avoid touching RCX/RDI reg (typically holding this) and got massive diffs in SPMI (-2.0Mb on SysV and -1.3Mb on win-x64).

@VSadov

VSadov commented Jun 1, 2026

Copy link
Copy Markdown
Member

@VSadov I think we just need a JIT-EE API that will tell the JIT about the calling convention for write barriers (e.g. eithe just normal or custom) and Satori can return "normal" (on both x64 and arm64) and remove that spilling logic you have. BTW, are you sure you also don't need to spill SIMD regs there? Because I think arm64 JIT assumes they're not killed too today, but it's possible that your C++ call just never use them either (but not something you can rely on).

Custom calling convention seems to be very proftiable, e.g. in this PR I tried just to avoid touching RCX/RDI reg (typically holding this) and got massive diffs in SPMI (-2.0Mb on SysV and -1.3Mb on win-x64).

I think standard calling convention may be too generous towards GC barriers. I understand the desire to leave registers for JIT to use.

However, the current conventions could be:

  • too tight that the asm part of the barrier may need to push/pop for extra temps.
  • unfriendly to C++ code.

It may be a good idea to make it configurable via JIT-EE API.
And I assume crossgen could just use the most conservative "common denominator" combination.

BTW, are you sure you also don't need to spill SIMD regs there?

We rely on C++ code that we call to not trash SIMD (by manual examination). On x64 we spill XMM0, because it does.
It is a bit of playing with fire.

SIMD is probably where I'd put the line. Spilling the whole SIMD context could be expensive, while chance of it used while calling a barrier are not high.

To be fair - the effects were never a target of too much research/measuring. Perhaps the current scheme is even completely OK.

@EgorBo

EgorBo commented Jun 1, 2026

Copy link
Copy Markdown
MemberAuthor

@VSadov It should be trivial to implement the JIT-EE I think, I can give it a try if Satori is anywhere near to be productized/merged into dotnet/runtime, something like CORINFO_WRITEBARRIER_CALLC getWriteBarrierCallingConvention();

The reason I decided to tune the registers for WB are potential regressions from my other change to remove CORINFO_HELP_ASSIGN_BYREF (#128542 + #128687) because that WB had a very compact calling convention with no spills between calls at all

@jkotas

Copy link
Copy Markdown
Member

We rely on C++ code that we call to not trash SIMD (by manual examination).

We used to do that in a few places and learned hard way that it is a bad idea...

@EgorBo
EgorBo enabled auto-merge (squash) June 2, 2026 21:01
@EgorBo
EgorBo merged commit db87610 into dotnet:mainJun 2, 2026
131 of 139 checks passed
@EgorBo

Copy link
Copy Markdown
MemberAuthor

@VSadov let me know if you want me to introduce an API that can switch CC for WB between this custom and standard (if we sure we'll merge Satori into runtime eventually)

I'm merging it to match the behavior with arm64

@VSadov

Copy link
Copy Markdown
Member

@VSadov let me know if you want me to introduce an API that can switch CC for WB between this custom and standard (if we sure we'll merge Satori into runtime eventually)

I'm merging it to match the behavior with arm64

I'll let you know. Right now Satori is a bit behind main. I think we will need to catch up to the state just before this PR first. (deleting byref barriers might be a good thing in terms of reducing complexity).

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@EgorBo@VSadov@jkotas@kg@tannergooding