Batched write barrier for byrefs - #99096

Closed
EgorBo wants to merge 15 commits into
dotnet:mainfrom
EgorBo:batched-byref-wb
Closed

Batched write barrier for byrefs#99096
EgorBo wants to merge 15 commits into
dotnet:mainfrom
EgorBo:batched-byref-wb

Conversation

@EgorBo

@EgorBoEgorBo commented Feb 29, 2024

Copy link
Copy Markdown
Member

Closes#8627

Testing on CI, for now - win-x64 coreclr only. Presumably, the asm version can be better than just a copy of JIT_ByRefWriteBarrier with a loop over R8 (length):

  • If first iteration finds out that the destination doesn't belong to any heap, we should skip all checks for the rest inside the batch
  • TODO: what else
  • Do we need debug-only WRITE_BARRIER_CHECK? we can use the non-batched version for DEBUG/CHECKED
  • REPRET looks like some legacy we can replace with just a normal ret

jit-diffs (seems like a few hits in Roslyn). Size regressions are from the cases when batchSize is 2, two calls are smaller than the batched one + mov r8d, <size>. But the perf should still be better with the batch one.

Example:

structMyStruct{objecta1;objecta2;objecta3;stringa4;stringa5;stringa6;}staticMyStructTest(MyStructms)=>ms;// simple copy of a struct with gc fields

Current codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 49

New codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxmovr8d,6call CORINFO_HELP_ASSIGN_BYREF_BATCHmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 30

Benchmark:

structMyStruct{objecta1;objecta2;objecta3;objecta4;}[Benchmark]publicvoidStructCopy(){MyStructms=default;Test(ms);[MethodImpl(MethodImplOptions.NoInlining)]MyStructTest(MyStructms)=>ms;}
MethodToolchainMeanErrorStdDevRatio
StructCopyCore_Root_Main\corerun.exe4.655 ns0.0097 ns0.0090 ns2.05
StructCopyCore_Root_PR\corerun.exe2.269 ns0.0097 ns0.0086 ns1.00

When destination is on the heap, the perf is also improved

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Feb 29, 2024
@ghostghost assigned EgorBoFeb 29, 2024
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

Testing on CI, for now - x64 nonAOT only. Presumably, the asm version can be better than just a loop over R8 (length), e.g. when dest is not in heap then other iterations can be faster.

Example:

structMyStruct{objecta1;objecta2;objecta3;stringa4;stringa5;stringa6;}staticMyStructTest(MyStructms)=>ms;// simple copy of a struct with gc fields

Current codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 49

New codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxmovr8d,6call CORINFO_HELP_ASSIGN_BYREF_BATCHmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 30
Author:EgorBo
Assignees:EgorBo
Labels:

area-CodeGen-coreclr

Milestone:-

add rdi, 8h
add rsi, 8h
dec r8d
jne NextByref

@jkotasjkotasFeb 29, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This loop can hang GC thread suspension.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you please elaborate? Is it because of shadow gc?

@jkotasjkotasFeb 29, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We won't be able to suspend the thread for GC while this helper is running. If this helper is called for large block repeatedly and there is not much managed code running between subsequent invocations of the helper, the GC thread suspension can be very slow or hang.

The same problem existed in BulkMoveWithWriteBarrier. It was fixed in dotnet/coreclr#27776 . You should be able to modify the repro from that PR to use InlineArray to hit the hang with this helper.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I can limit the max batch size. e.g. 1024 gc slots will be handled as 16 call CORINFO_HELP_ASSIGN_BYREF_BATCH batches

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would measure the worst-case scenario, e.g. call MyStruct Test(MyStruct ms) => ms; in a loop on one thread and measure average GC pause time on a second thread.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

From what I see from the diffs, we're mostly dealing with small values like 2 (most popular) and less than 10

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok, then the math does not work - there is something else going on. (For some reason, I thought that the batch size is 16 in your current change.)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If I change the microbenchmark so that the write barrier has to do actual work, I am seeing 15sec - 30sec pause times with your current change: https://gist.github.com/jkotas/d9b48e48935bcc5a6f3dc87db086196d

@EgorBoEgorBoMar 1, 2024

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Don't know why, but this benchmark shows better numbers in my branch than in Main (with 16 slots max)

@EgorBoEgorBoMar 1, 2024

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe because method is just too big in Main and that made gcinfo slower to process? (; Total bytes of code 5012)

@jkotas

jkotas commented Feb 29, 2024

Copy link
Copy Markdown
Member

Observation: The helper is functional equivalent of BulkMoveWithWriteBarrier, with different perf trade-off. Compared toBulkMoveWithWriteBarrier , the helper as currently implemented in the PR tries to be more precise with setting the cards so it is slower but leaves less work to be done at GC time.

@EgorBo

EgorBo commented Feb 29, 2024

Copy link
Copy Markdown
MemberAuthor

Observation: The helper is functional equivalent of BulkMoveWithWriteBarrier, with different perf trade-off. Compared toBulkMoveWithWriteBarrier , the helper as currently implemented in the PR tries to be more precise with setting the cards so it is slower but leaves less work to be done at GC time.

Thanks! Presumably it's easier to just call BulkMoveWithWriteBarrier so then we don't have to duplicate these helpers for all platforms. Do you have an opinion on which way is generally better?

One thing that motivated me is that we've seen write barriers on hot paths in several real world benchmarks (e.g. @kunalspathak recently did). I am not sure this specific one was the culprit (they're generally less advanced on arm64), just looked like a low-hanging fruit to me.

@jkotas

jkotas commented Feb 29, 2024

Copy link
Copy Markdown
Member

Presumably it's easier to just call BulkMoveWithWriteBarrier

BulkMoveWithWriteBarrier has higher fixed overhead than the hand-written assembly helper. I expect that it would be profitable to call it only for blocks above certain size.

we've seen write barriers on hot paths in several real world benchmarks

Is this addressing the hot paths that we have seen?

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Is this addressing the hot paths that we have seen?

Not sure yet, but jitdiffs imply they're not rare (jitdiffs are usually way smaller than SPMI because we only run it for framework libs, no tests, benchmarks, apsnet, etc).

@kunalspathak

Copy link
Copy Markdown
Contributor

on hot paths in several real world benchmarks (e.g. @kunalspathak recently did)

We saw it more on arm64 and we do not know yet if too many samples of JIT_WriteBarrier on the call stack was because the arm64 version is slower or it was due to lack of batching. I see that this one is just for amd64, so won't help the arm64 case that I was seeing.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

I see that this one is just for amd64, so won't help the arm64 case that I was seeing.

It's just for the draft, I planned to support x64 (windows, unix) and arm64 (windows, unix) for both CoreCLR and NativeAOT

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Here is the current diff between non-batched (left) and batched (right) version: https://www.diffchecker.com/clp44lCx/

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@MihuBot

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@jkotas Does it look good now? I changed to 16 slots max and implemented Windows-x64+Unix-x64 on both CoreCLR and NAOT. I'd like to handle ARM64 in a separate PR if you don't mind.

To simplify code-review of the newly added helpers, I made diffs for each against their non-batch version

  1. src/coreclr/nativeaot/Runtime/amd64/WriteBarriers.S: diff
  2. src/coreclr/nativeaot/Runtime/amd64/WriteBarriers.asm: diff
  3. src/coreclr/vm/amd64/JitHelpers_Fast.asm: diff <--- the least verbose diff
  4. src/coreclr/vm/amd64/jithelpers_fast.S: diff

@EgorBo
EgorBo marked this pull request as ready for review March 1, 2024 00:54
@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

This comment was marked as resolved.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr outerloop, runtime-coreclr jitstressregs, runtime-coreclr gcstress0x3-gcstress0xc, runtime-nativeaot-outerloop

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 4 pipeline(s).

@jkotas

Copy link
Copy Markdown
Member

@VSadov What do you think about impact of this change on GC thread suspension? Have you seen bad cases where the current write barriers are stalling GC thread suspension?

I am wondering where we should set the bar for the GC thread suspension impact on changes like this one. There are multiple ways to make it more friendly to GC thread suspension - explicitly poll in the helper, explicit GC polls in the caller, enable hijacking of return address from these helpers, ... .

@VSadov

VSadov commented Mar 1, 2024

Copy link
Copy Markdown
Member

My first thought about this is that the barriers are the kind of platform- and achitecture-specific pieces of assembly, that we would usually avoid for portability reasons. With Core/AOT + 6 architecture + Win/Unix - how many combinations will need a copy of the new barrier? Do we get enough savings for that?

It looks like it saves a call. Also hoists “not in heap” check, although that will help only if it is really not in heap and short-circuits, otherwise that check is cheap. The rest of the cost - other checks, lock or, the actual assignment are still there.
Considering the barrier cost is typically in 1-5% range and that this will be hit in fairly rare cases in the code (“few cases in Roslyn”?), I wonder how much of an improvement is expected here in general?

What happens to the benchmark if the write is to a bunch of array elements instead of a single stack location?
(ideally on a Server GC)

@VSadov

Copy link
Copy Markdown
Member

I kind of suspect the original issue expected more “batching” - like hoisting other checks, setting cards in bulk, etc… But that will turn the barrier into BulkMoveWithWriteBarrier, which we already have.

;; r8: number of byrefs
;;
;; On exit:
;; rdi, rsi are incremented by 8,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shouldn't them be increased by the bytes processed? And also comments in other asm code.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

What happens to the benchmark if the write is to a bunch of array elements instead of a single stack location?
(ideally on a Server GC)

Not much indeed, I don't have a strong opinion here, since it raises concerns/doubts I am fine in closing it before I invest my time into arm side of it 🤷‍♂️ The case where I noticed this issue were tuples, e.g.:

(string,string)StructCopy((string,string)s)=>s;

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Closing, will try to benchmark/compare arm64 write barriers (not only byref one) against x64 to see if I can contribute there instead

@EgorBoEgorBo closed this Mar 1, 2024
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Apr 1, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Consider creating a ranged version of CORINFO_HELP_ASSIGN_BYREF

5 participants

@EgorBo@jkotas@kunalspathak@VSadov@huoyaoyuan
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Batched write barrier for byrefs - #99096

Closed
EgorBo wants to merge 15 commits into
dotnet:mainfrom
EgorBo:batched-byref-wb
Closed

Batched write barrier for byrefs#99096
EgorBo wants to merge 15 commits into
dotnet:mainfrom
EgorBo:batched-byref-wb

Conversation

@EgorBo

@EgorBoEgorBo commented Feb 29, 2024

Copy link
Copy Markdown
Member

Closes#8627

Testing on CI, for now - win-x64 coreclr only. Presumably, the asm version can be better than just a copy of JIT_ByRefWriteBarrier with a loop over R8 (length):

  • If first iteration finds out that the destination doesn't belong to any heap, we should skip all checks for the rest inside the batch
  • TODO: what else
  • Do we need debug-only WRITE_BARRIER_CHECK? we can use the non-batched version for DEBUG/CHECKED
  • REPRET looks like some legacy we can replace with just a normal ret

jit-diffs (seems like a few hits in Roslyn). Size regressions are from the cases when batchSize is 2, two calls are smaller than the batched one + mov r8d, <size>. But the perf should still be better with the batch one.

Example:

structMyStruct{objecta1;objecta2;objecta3;stringa4;stringa5;stringa6;}staticMyStructTest(MyStructms)=>ms;// simple copy of a struct with gc fields

Current codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 49

New codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxmovr8d,6call CORINFO_HELP_ASSIGN_BYREF_BATCHmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 30

Benchmark:

structMyStruct{objecta1;objecta2;objecta3;objecta4;}[Benchmark]publicvoidStructCopy(){MyStructms=default;Test(ms);[MethodImpl(MethodImplOptions.NoInlining)]MyStructTest(MyStructms)=>ms;}
MethodToolchainMeanErrorStdDevRatio
StructCopyCore_Root_Main\corerun.exe4.655 ns0.0097 ns0.0090 ns2.05
StructCopyCore_Root_PR\corerun.exe2.269 ns0.0097 ns0.0086 ns1.00

When destination is on the heap, the perf is also improved

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Feb 29, 2024
@ghostghost assigned EgorBoFeb 29, 2024
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

Testing on CI, for now - x64 nonAOT only. Presumably, the asm version can be better than just a loop over R8 (length), e.g. when dest is not in heap then other iterations can be faster.

Example:

structMyStruct{objecta1;objecta2;objecta3;stringa4;stringa5;stringa6;}staticMyStructTest(MyStructms)=>ms;// simple copy of a struct with gc fields

Current codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 49

New codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxmovr8d,6call CORINFO_HELP_ASSIGN_BYREF_BATCHmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 30
Author:EgorBo
Assignees:EgorBo
Labels:

area-CodeGen-coreclr

Milestone:-

add rdi, 8h
add rsi, 8h
dec r8d
jne NextByref

@jkotasjkotasFeb 29, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This loop can hang GC thread suspension.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you please elaborate? Is it because of shadow gc?

@jkotasjkotasFeb 29, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We won't be able to suspend the thread for GC while this helper is running. If this helper is called for large block repeatedly and there is not much managed code running between subsequent invocations of the helper, the GC thread suspension can be very slow or hang.

The same problem existed in BulkMoveWithWriteBarrier. It was fixed in dotnet/coreclr#27776 . You should be able to modify the repro from that PR to use InlineArray to hit the hang with this helper.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I can limit the max batch size. e.g. 1024 gc slots will be handled as 16 call CORINFO_HELP_ASSIGN_BYREF_BATCH batches

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would measure the worst-case scenario, e.g. call MyStruct Test(MyStruct ms) => ms; in a loop on one thread and measure average GC pause time on a second thread.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

From what I see from the diffs, we're mostly dealing with small values like 2 (most popular) and less than 10

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok, then the math does not work - there is something else going on. (For some reason, I thought that the batch size is 16 in your current change.)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If I change the microbenchmark so that the write barrier has to do actual work, I am seeing 15sec - 30sec pause times with your current change: https://gist.github.com/jkotas/d9b48e48935bcc5a6f3dc87db086196d

@EgorBoEgorBoMar 1, 2024

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Don't know why, but this benchmark shows better numbers in my branch than in Main (with 16 slots max)

@EgorBoEgorBoMar 1, 2024

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe because method is just too big in Main and that made gcinfo slower to process? (; Total bytes of code 5012)

@jkotas

jkotas commented Feb 29, 2024

Copy link
Copy Markdown
Member

Observation: The helper is functional equivalent of BulkMoveWithWriteBarrier, with different perf trade-off. Compared toBulkMoveWithWriteBarrier , the helper as currently implemented in the PR tries to be more precise with setting the cards so it is slower but leaves less work to be done at GC time.

@EgorBo

EgorBo commented Feb 29, 2024

Copy link
Copy Markdown
MemberAuthor

Observation: The helper is functional equivalent of BulkMoveWithWriteBarrier, with different perf trade-off. Compared toBulkMoveWithWriteBarrier , the helper as currently implemented in the PR tries to be more precise with setting the cards so it is slower but leaves less work to be done at GC time.

Thanks! Presumably it's easier to just call BulkMoveWithWriteBarrier so then we don't have to duplicate these helpers for all platforms. Do you have an opinion on which way is generally better?

One thing that motivated me is that we've seen write barriers on hot paths in several real world benchmarks (e.g. @kunalspathak recently did). I am not sure this specific one was the culprit (they're generally less advanced on arm64), just looked like a low-hanging fruit to me.

@jkotas

jkotas commented Feb 29, 2024

Copy link
Copy Markdown
Member

Presumably it's easier to just call BulkMoveWithWriteBarrier

BulkMoveWithWriteBarrier has higher fixed overhead than the hand-written assembly helper. I expect that it would be profitable to call it only for blocks above certain size.

we've seen write barriers on hot paths in several real world benchmarks

Is this addressing the hot paths that we have seen?

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Is this addressing the hot paths that we have seen?

Not sure yet, but jitdiffs imply they're not rare (jitdiffs are usually way smaller than SPMI because we only run it for framework libs, no tests, benchmarks, apsnet, etc).

@kunalspathak

Copy link
Copy Markdown
Contributor

on hot paths in several real world benchmarks (e.g. @kunalspathak recently did)

We saw it more on arm64 and we do not know yet if too many samples of JIT_WriteBarrier on the call stack was because the arm64 version is slower or it was due to lack of batching. I see that this one is just for amd64, so won't help the arm64 case that I was seeing.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

I see that this one is just for amd64, so won't help the arm64 case that I was seeing.

It's just for the draft, I planned to support x64 (windows, unix) and arm64 (windows, unix) for both CoreCLR and NativeAOT

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Here is the current diff between non-batched (left) and batched (right) version: https://www.diffchecker.com/clp44lCx/

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@MihuBot

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@jkotas Does it look good now? I changed to 16 slots max and implemented Windows-x64+Unix-x64 on both CoreCLR and NAOT. I'd like to handle ARM64 in a separate PR if you don't mind.

To simplify code-review of the newly added helpers, I made diffs for each against their non-batch version

  1. src/coreclr/nativeaot/Runtime/amd64/WriteBarriers.S: diff
  2. src/coreclr/nativeaot/Runtime/amd64/WriteBarriers.asm: diff
  3. src/coreclr/vm/amd64/JitHelpers_Fast.asm: diff <--- the least verbose diff
  4. src/coreclr/vm/amd64/jithelpers_fast.S: diff

@EgorBo
EgorBo marked this pull request as ready for review March 1, 2024 00:54
@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

This comment was marked as resolved.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr outerloop, runtime-coreclr jitstressregs, runtime-coreclr gcstress0x3-gcstress0xc, runtime-nativeaot-outerloop

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 4 pipeline(s).

@jkotas

Copy link
Copy Markdown
Member

@VSadov What do you think about impact of this change on GC thread suspension? Have you seen bad cases where the current write barriers are stalling GC thread suspension?

I am wondering where we should set the bar for the GC thread suspension impact on changes like this one. There are multiple ways to make it more friendly to GC thread suspension - explicitly poll in the helper, explicit GC polls in the caller, enable hijacking of return address from these helpers, ... .

@VSadov

VSadov commented Mar 1, 2024

Copy link
Copy Markdown
Member

My first thought about this is that the barriers are the kind of platform- and achitecture-specific pieces of assembly, that we would usually avoid for portability reasons. With Core/AOT + 6 architecture + Win/Unix - how many combinations will need a copy of the new barrier? Do we get enough savings for that?

It looks like it saves a call. Also hoists “not in heap” check, although that will help only if it is really not in heap and short-circuits, otherwise that check is cheap. The rest of the cost - other checks, lock or, the actual assignment are still there.
Considering the barrier cost is typically in 1-5% range and that this will be hit in fairly rare cases in the code (“few cases in Roslyn”?), I wonder how much of an improvement is expected here in general?

What happens to the benchmark if the write is to a bunch of array elements instead of a single stack location?
(ideally on a Server GC)

@VSadov

Copy link
Copy Markdown
Member

I kind of suspect the original issue expected more “batching” - like hoisting other checks, setting cards in bulk, etc… But that will turn the barrier into BulkMoveWithWriteBarrier, which we already have.

;; r8: number of byrefs
;;
;; On exit:
;; rdi, rsi are incremented by 8,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shouldn't them be increased by the bytes processed? And also comments in other asm code.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

What happens to the benchmark if the write is to a bunch of array elements instead of a single stack location?
(ideally on a Server GC)

Not much indeed, I don't have a strong opinion here, since it raises concerns/doubts I am fine in closing it before I invest my time into arm side of it 🤷‍♂️ The case where I noticed this issue were tuples, e.g.:

(string,string)StructCopy((string,string)s)=>s;

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Closing, will try to benchmark/compare arm64 write barriers (not only byref one) against x64 to see if I can contribute there instead

@EgorBoEgorBo closed this Mar 1, 2024
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Apr 1, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Consider creating a ranged version of CORINFO_HELP_ASSIGN_BYREF

5 participants

@EgorBo@jkotas@kunalspathak@VSadov@huoyaoyuan
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Batched write barrier for byrefs - #99096

Closed
EgorBo wants to merge 15 commits into
dotnet:mainfrom
EgorBo:batched-byref-wb
Closed

Batched write barrier for byrefs#99096
EgorBo wants to merge 15 commits into
dotnet:mainfrom
EgorBo:batched-byref-wb

Conversation

@EgorBo

@EgorBoEgorBo commented Feb 29, 2024

Copy link
Copy Markdown
Member

Closes#8627

Testing on CI, for now - win-x64 coreclr only. Presumably, the asm version can be better than just a copy of JIT_ByRefWriteBarrier with a loop over R8 (length):

  • If first iteration finds out that the destination doesn't belong to any heap, we should skip all checks for the rest inside the batch
  • TODO: what else
  • Do we need debug-only WRITE_BARRIER_CHECK? we can use the non-batched version for DEBUG/CHECKED
  • REPRET looks like some legacy we can replace with just a normal ret

jit-diffs (seems like a few hits in Roslyn). Size regressions are from the cases when batchSize is 2, two calls are smaller than the batched one + mov r8d, <size>. But the perf should still be better with the batch one.

Example:

structMyStruct{objecta1;objecta2;objecta3;stringa4;stringa5;stringa6;}staticMyStructTest(MyStructms)=>ms;// simple copy of a struct with gc fields

Current codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 49

New codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxmovr8d,6call CORINFO_HELP_ASSIGN_BYREF_BATCHmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 30

Benchmark:

structMyStruct{objecta1;objecta2;objecta3;objecta4;}[Benchmark]publicvoidStructCopy(){MyStructms=default;Test(ms);[MethodImpl(MethodImplOptions.NoInlining)]MyStructTest(MyStructms)=>ms;}
MethodToolchainMeanErrorStdDevRatio
StructCopyCore_Root_Main\corerun.exe4.655 ns0.0097 ns0.0090 ns2.05
StructCopyCore_Root_PR\corerun.exe2.269 ns0.0097 ns0.0086 ns1.00

When destination is on the heap, the perf is also improved

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Feb 29, 2024
@ghostghost assigned EgorBoFeb 29, 2024
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

Testing on CI, for now - x64 nonAOT only. Presumably, the asm version can be better than just a loop over R8 (length), e.g. when dest is not in heap then other iterations can be faster.

Example:

structMyStruct{objecta1;objecta2;objecta3;stringa4;stringa5;stringa6;}staticMyStructTest(MyStructms)=>ms;// simple copy of a struct with gc fields

Current codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 49

New codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxmovr8d,6call CORINFO_HELP_ASSIGN_BYREF_BATCHmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 30
Author:EgorBo
Assignees:EgorBo
Labels:

area-CodeGen-coreclr

Milestone:-

add rdi, 8h
add rsi, 8h
dec r8d
jne NextByref

@jkotasjkotasFeb 29, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This loop can hang GC thread suspension.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you please elaborate? Is it because of shadow gc?

@jkotasjkotasFeb 29, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We won't be able to suspend the thread for GC while this helper is running. If this helper is called for large block repeatedly and there is not much managed code running between subsequent invocations of the helper, the GC thread suspension can be very slow or hang.

The same problem existed in BulkMoveWithWriteBarrier. It was fixed in dotnet/coreclr#27776 . You should be able to modify the repro from that PR to use InlineArray to hit the hang with this helper.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I can limit the max batch size. e.g. 1024 gc slots will be handled as 16 call CORINFO_HELP_ASSIGN_BYREF_BATCH batches

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would measure the worst-case scenario, e.g. call MyStruct Test(MyStruct ms) => ms; in a loop on one thread and measure average GC pause time on a second thread.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

From what I see from the diffs, we're mostly dealing with small values like 2 (most popular) and less than 10

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok, then the math does not work - there is something else going on. (For some reason, I thought that the batch size is 16 in your current change.)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If I change the microbenchmark so that the write barrier has to do actual work, I am seeing 15sec - 30sec pause times with your current change: https://gist.github.com/jkotas/d9b48e48935bcc5a6f3dc87db086196d

@EgorBoEgorBoMar 1, 2024

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Don't know why, but this benchmark shows better numbers in my branch than in Main (with 16 slots max)

@EgorBoEgorBoMar 1, 2024

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe because method is just too big in Main and that made gcinfo slower to process? (; Total bytes of code 5012)

@jkotas

jkotas commented Feb 29, 2024

Copy link
Copy Markdown
Member

Observation: The helper is functional equivalent of BulkMoveWithWriteBarrier, with different perf trade-off. Compared toBulkMoveWithWriteBarrier , the helper as currently implemented in the PR tries to be more precise with setting the cards so it is slower but leaves less work to be done at GC time.

@EgorBo

EgorBo commented Feb 29, 2024

Copy link
Copy Markdown
MemberAuthor

Observation: The helper is functional equivalent of BulkMoveWithWriteBarrier, with different perf trade-off. Compared toBulkMoveWithWriteBarrier , the helper as currently implemented in the PR tries to be more precise with setting the cards so it is slower but leaves less work to be done at GC time.

Thanks! Presumably it's easier to just call BulkMoveWithWriteBarrier so then we don't have to duplicate these helpers for all platforms. Do you have an opinion on which way is generally better?

One thing that motivated me is that we've seen write barriers on hot paths in several real world benchmarks (e.g. @kunalspathak recently did). I am not sure this specific one was the culprit (they're generally less advanced on arm64), just looked like a low-hanging fruit to me.

@jkotas

jkotas commented Feb 29, 2024

Copy link
Copy Markdown
Member

Presumably it's easier to just call BulkMoveWithWriteBarrier

BulkMoveWithWriteBarrier has higher fixed overhead than the hand-written assembly helper. I expect that it would be profitable to call it only for blocks above certain size.

we've seen write barriers on hot paths in several real world benchmarks

Is this addressing the hot paths that we have seen?

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Is this addressing the hot paths that we have seen?

Not sure yet, but jitdiffs imply they're not rare (jitdiffs are usually way smaller than SPMI because we only run it for framework libs, no tests, benchmarks, apsnet, etc).

@kunalspathak

Copy link
Copy Markdown
Contributor

on hot paths in several real world benchmarks (e.g. @kunalspathak recently did)

We saw it more on arm64 and we do not know yet if too many samples of JIT_WriteBarrier on the call stack was because the arm64 version is slower or it was due to lack of batching. I see that this one is just for amd64, so won't help the arm64 case that I was seeing.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

I see that this one is just for amd64, so won't help the arm64 case that I was seeing.

It's just for the draft, I planned to support x64 (windows, unix) and arm64 (windows, unix) for both CoreCLR and NativeAOT

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Here is the current diff between non-batched (left) and batched (right) version: https://www.diffchecker.com/clp44lCx/

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@MihuBot

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@jkotas Does it look good now? I changed to 16 slots max and implemented Windows-x64+Unix-x64 on both CoreCLR and NAOT. I'd like to handle ARM64 in a separate PR if you don't mind.

To simplify code-review of the newly added helpers, I made diffs for each against their non-batch version

  1. src/coreclr/nativeaot/Runtime/amd64/WriteBarriers.S: diff
  2. src/coreclr/nativeaot/Runtime/amd64/WriteBarriers.asm: diff
  3. src/coreclr/vm/amd64/JitHelpers_Fast.asm: diff <--- the least verbose diff
  4. src/coreclr/vm/amd64/jithelpers_fast.S: diff

@EgorBo
EgorBo marked this pull request as ready for review March 1, 2024 00:54
@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

This comment was marked as resolved.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr outerloop, runtime-coreclr jitstressregs, runtime-coreclr gcstress0x3-gcstress0xc, runtime-nativeaot-outerloop

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 4 pipeline(s).

@jkotas

Copy link
Copy Markdown
Member

@VSadov What do you think about impact of this change on GC thread suspension? Have you seen bad cases where the current write barriers are stalling GC thread suspension?

I am wondering where we should set the bar for the GC thread suspension impact on changes like this one. There are multiple ways to make it more friendly to GC thread suspension - explicitly poll in the helper, explicit GC polls in the caller, enable hijacking of return address from these helpers, ... .

@VSadov

VSadov commented Mar 1, 2024

Copy link
Copy Markdown
Member

My first thought about this is that the barriers are the kind of platform- and achitecture-specific pieces of assembly, that we would usually avoid for portability reasons. With Core/AOT + 6 architecture + Win/Unix - how many combinations will need a copy of the new barrier? Do we get enough savings for that?

It looks like it saves a call. Also hoists “not in heap” check, although that will help only if it is really not in heap and short-circuits, otherwise that check is cheap. The rest of the cost - other checks, lock or, the actual assignment are still there.
Considering the barrier cost is typically in 1-5% range and that this will be hit in fairly rare cases in the code (“few cases in Roslyn”?), I wonder how much of an improvement is expected here in general?

What happens to the benchmark if the write is to a bunch of array elements instead of a single stack location?
(ideally on a Server GC)

@VSadov

Copy link
Copy Markdown
Member

I kind of suspect the original issue expected more “batching” - like hoisting other checks, setting cards in bulk, etc… But that will turn the barrier into BulkMoveWithWriteBarrier, which we already have.

;; r8: number of byrefs
;;
;; On exit:
;; rdi, rsi are incremented by 8,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shouldn't them be increased by the bytes processed? And also comments in other asm code.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

What happens to the benchmark if the write is to a bunch of array elements instead of a single stack location?
(ideally on a Server GC)

Not much indeed, I don't have a strong opinion here, since it raises concerns/doubts I am fine in closing it before I invest my time into arm side of it 🤷‍♂️ The case where I noticed this issue were tuples, e.g.:

(string,string)StructCopy((string,string)s)=>s;

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Closing, will try to benchmark/compare arm64 write barriers (not only byref one) against x64 to see if I can contribute there instead

@EgorBoEgorBo closed this Mar 1, 2024
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Apr 1, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Consider creating a ranged version of CORINFO_HELP_ASSIGN_BYREF

5 participants

@EgorBo@jkotas@kunalspathak@VSadov@huoyaoyuan
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Batched write barrier for byrefs - #99096

Closed
EgorBo wants to merge 15 commits into
dotnet:mainfrom
EgorBo:batched-byref-wb
Closed

Batched write barrier for byrefs#99096
EgorBo wants to merge 15 commits into
dotnet:mainfrom
EgorBo:batched-byref-wb

Conversation

@EgorBo

@EgorBoEgorBo commented Feb 29, 2024

Copy link
Copy Markdown
Member

Closes#8627

Testing on CI, for now - win-x64 coreclr only. Presumably, the asm version can be better than just a copy of JIT_ByRefWriteBarrier with a loop over R8 (length):

  • If first iteration finds out that the destination doesn't belong to any heap, we should skip all checks for the rest inside the batch
  • TODO: what else
  • Do we need debug-only WRITE_BARRIER_CHECK? we can use the non-batched version for DEBUG/CHECKED
  • REPRET looks like some legacy we can replace with just a normal ret

jit-diffs (seems like a few hits in Roslyn). Size regressions are from the cases when batchSize is 2, two calls are smaller than the batched one + mov r8d, <size>. But the perf should still be better with the batch one.

Example:

structMyStruct{objecta1;objecta2;objecta3;stringa4;stringa5;stringa6;}staticMyStructTest(MyStructms)=>ms;// simple copy of a struct with gc fields

Current codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 49

New codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxmovr8d,6call CORINFO_HELP_ASSIGN_BYREF_BATCHmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 30

Benchmark:

structMyStruct{objecta1;objecta2;objecta3;objecta4;}[Benchmark]publicvoidStructCopy(){MyStructms=default;Test(ms);[MethodImpl(MethodImplOptions.NoInlining)]MyStructTest(MyStructms)=>ms;}
MethodToolchainMeanErrorStdDevRatio
StructCopyCore_Root_Main\corerun.exe4.655 ns0.0097 ns0.0090 ns2.05
StructCopyCore_Root_PR\corerun.exe2.269 ns0.0097 ns0.0086 ns1.00

When destination is on the heap, the perf is also improved

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Feb 29, 2024
@ghostghost assigned EgorBoFeb 29, 2024
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

Testing on CI, for now - x64 nonAOT only. Presumably, the asm version can be better than just a loop over R8 (length), e.g. when dest is not in heap then other iterations can be faster.

Example:

structMyStruct{objecta1;objecta2;objecta3;stringa4;stringa5;stringa6;}staticMyStructTest(MyStructms)=>ms;// simple copy of a struct with gc fields

Current codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 49

New codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxmovr8d,6call CORINFO_HELP_ASSIGN_BYREF_BATCHmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 30
Author:EgorBo
Assignees:EgorBo
Labels:

area-CodeGen-coreclr

Milestone:-

add rdi, 8h
add rsi, 8h
dec r8d
jne NextByref

@jkotasjkotasFeb 29, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This loop can hang GC thread suspension.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you please elaborate? Is it because of shadow gc?

@jkotasjkotasFeb 29, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We won't be able to suspend the thread for GC while this helper is running. If this helper is called for large block repeatedly and there is not much managed code running between subsequent invocations of the helper, the GC thread suspension can be very slow or hang.

The same problem existed in BulkMoveWithWriteBarrier. It was fixed in dotnet/coreclr#27776 . You should be able to modify the repro from that PR to use InlineArray to hit the hang with this helper.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I can limit the max batch size. e.g. 1024 gc slots will be handled as 16 call CORINFO_HELP_ASSIGN_BYREF_BATCH batches

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would measure the worst-case scenario, e.g. call MyStruct Test(MyStruct ms) => ms; in a loop on one thread and measure average GC pause time on a second thread.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

From what I see from the diffs, we're mostly dealing with small values like 2 (most popular) and less than 10

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok, then the math does not work - there is something else going on. (For some reason, I thought that the batch size is 16 in your current change.)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If I change the microbenchmark so that the write barrier has to do actual work, I am seeing 15sec - 30sec pause times with your current change: https://gist.github.com/jkotas/d9b48e48935bcc5a6f3dc87db086196d

@EgorBoEgorBoMar 1, 2024

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Don't know why, but this benchmark shows better numbers in my branch than in Main (with 16 slots max)

@EgorBoEgorBoMar 1, 2024

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe because method is just too big in Main and that made gcinfo slower to process? (; Total bytes of code 5012)

@jkotas

jkotas commented Feb 29, 2024

Copy link
Copy Markdown
Member

Observation: The helper is functional equivalent of BulkMoveWithWriteBarrier, with different perf trade-off. Compared toBulkMoveWithWriteBarrier , the helper as currently implemented in the PR tries to be more precise with setting the cards so it is slower but leaves less work to be done at GC time.

@EgorBo

EgorBo commented Feb 29, 2024

Copy link
Copy Markdown
MemberAuthor

Observation: The helper is functional equivalent of BulkMoveWithWriteBarrier, with different perf trade-off. Compared toBulkMoveWithWriteBarrier , the helper as currently implemented in the PR tries to be more precise with setting the cards so it is slower but leaves less work to be done at GC time.

Thanks! Presumably it's easier to just call BulkMoveWithWriteBarrier so then we don't have to duplicate these helpers for all platforms. Do you have an opinion on which way is generally better?

One thing that motivated me is that we've seen write barriers on hot paths in several real world benchmarks (e.g. @kunalspathak recently did). I am not sure this specific one was the culprit (they're generally less advanced on arm64), just looked like a low-hanging fruit to me.

@jkotas

jkotas commented Feb 29, 2024

Copy link
Copy Markdown
Member

Presumably it's easier to just call BulkMoveWithWriteBarrier

BulkMoveWithWriteBarrier has higher fixed overhead than the hand-written assembly helper. I expect that it would be profitable to call it only for blocks above certain size.

we've seen write barriers on hot paths in several real world benchmarks

Is this addressing the hot paths that we have seen?

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Is this addressing the hot paths that we have seen?

Not sure yet, but jitdiffs imply they're not rare (jitdiffs are usually way smaller than SPMI because we only run it for framework libs, no tests, benchmarks, apsnet, etc).

@kunalspathak

Copy link
Copy Markdown
Contributor

on hot paths in several real world benchmarks (e.g. @kunalspathak recently did)

We saw it more on arm64 and we do not know yet if too many samples of JIT_WriteBarrier on the call stack was because the arm64 version is slower or it was due to lack of batching. I see that this one is just for amd64, so won't help the arm64 case that I was seeing.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

I see that this one is just for amd64, so won't help the arm64 case that I was seeing.

It's just for the draft, I planned to support x64 (windows, unix) and arm64 (windows, unix) for both CoreCLR and NativeAOT

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Here is the current diff between non-batched (left) and batched (right) version: https://www.diffchecker.com/clp44lCx/

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@MihuBot

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@jkotas Does it look good now? I changed to 16 slots max and implemented Windows-x64+Unix-x64 on both CoreCLR and NAOT. I'd like to handle ARM64 in a separate PR if you don't mind.

To simplify code-review of the newly added helpers, I made diffs for each against their non-batch version

  1. src/coreclr/nativeaot/Runtime/amd64/WriteBarriers.S: diff
  2. src/coreclr/nativeaot/Runtime/amd64/WriteBarriers.asm: diff
  3. src/coreclr/vm/amd64/JitHelpers_Fast.asm: diff <--- the least verbose diff
  4. src/coreclr/vm/amd64/jithelpers_fast.S: diff

@EgorBo
EgorBo marked this pull request as ready for review March 1, 2024 00:54
@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

This comment was marked as resolved.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr outerloop, runtime-coreclr jitstressregs, runtime-coreclr gcstress0x3-gcstress0xc, runtime-nativeaot-outerloop

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 4 pipeline(s).

@jkotas

Copy link
Copy Markdown
Member

@VSadov What do you think about impact of this change on GC thread suspension? Have you seen bad cases where the current write barriers are stalling GC thread suspension?

I am wondering where we should set the bar for the GC thread suspension impact on changes like this one. There are multiple ways to make it more friendly to GC thread suspension - explicitly poll in the helper, explicit GC polls in the caller, enable hijacking of return address from these helpers, ... .

@VSadov

VSadov commented Mar 1, 2024

Copy link
Copy Markdown
Member

My first thought about this is that the barriers are the kind of platform- and achitecture-specific pieces of assembly, that we would usually avoid for portability reasons. With Core/AOT + 6 architecture + Win/Unix - how many combinations will need a copy of the new barrier? Do we get enough savings for that?

It looks like it saves a call. Also hoists “not in heap” check, although that will help only if it is really not in heap and short-circuits, otherwise that check is cheap. The rest of the cost - other checks, lock or, the actual assignment are still there.
Considering the barrier cost is typically in 1-5% range and that this will be hit in fairly rare cases in the code (“few cases in Roslyn”?), I wonder how much of an improvement is expected here in general?

What happens to the benchmark if the write is to a bunch of array elements instead of a single stack location?
(ideally on a Server GC)

@VSadov

Copy link
Copy Markdown
Member

I kind of suspect the original issue expected more “batching” - like hoisting other checks, setting cards in bulk, etc… But that will turn the barrier into BulkMoveWithWriteBarrier, which we already have.

;; r8: number of byrefs
;;
;; On exit:
;; rdi, rsi are incremented by 8,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shouldn't them be increased by the bytes processed? And also comments in other asm code.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

What happens to the benchmark if the write is to a bunch of array elements instead of a single stack location?
(ideally on a Server GC)

Not much indeed, I don't have a strong opinion here, since it raises concerns/doubts I am fine in closing it before I invest my time into arm side of it 🤷‍♂️ The case where I noticed this issue were tuples, e.g.:

(string,string)StructCopy((string,string)s)=>s;

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Closing, will try to benchmark/compare arm64 write barriers (not only byref one) against x64 to see if I can contribute there instead

@EgorBoEgorBo closed this Mar 1, 2024
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Apr 1, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Consider creating a ranged version of CORINFO_HELP_ASSIGN_BYREF

5 participants

@EgorBo@jkotas@kunalspathak@VSadov@huoyaoyuan
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Batched write barrier for byrefs - #99096

Closed
EgorBo wants to merge 15 commits into
dotnet:mainfrom
EgorBo:batched-byref-wb
Closed

Batched write barrier for byrefs#99096
EgorBo wants to merge 15 commits into
dotnet:mainfrom
EgorBo:batched-byref-wb

Conversation

@EgorBo

@EgorBoEgorBo commented Feb 29, 2024

Copy link
Copy Markdown
Member

Closes#8627

Testing on CI, for now - win-x64 coreclr only. Presumably, the asm version can be better than just a copy of JIT_ByRefWriteBarrier with a loop over R8 (length):

  • If first iteration finds out that the destination doesn't belong to any heap, we should skip all checks for the rest inside the batch
  • TODO: what else
  • Do we need debug-only WRITE_BARRIER_CHECK? we can use the non-batched version for DEBUG/CHECKED
  • REPRET looks like some legacy we can replace with just a normal ret

jit-diffs (seems like a few hits in Roslyn). Size regressions are from the cases when batchSize is 2, two calls are smaller than the batched one + mov r8d, <size>. But the perf should still be better with the batch one.

Example:

structMyStruct{objecta1;objecta2;objecta3;stringa4;stringa5;stringa6;}staticMyStructTest(MyStructms)=>ms;// simple copy of a struct with gc fields

Current codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 49

New codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxmovr8d,6call CORINFO_HELP_ASSIGN_BYREF_BATCHmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 30

Benchmark:

structMyStruct{objecta1;objecta2;objecta3;objecta4;}[Benchmark]publicvoidStructCopy(){MyStructms=default;Test(ms);[MethodImpl(MethodImplOptions.NoInlining)]MyStructTest(MyStructms)=>ms;}
MethodToolchainMeanErrorStdDevRatio
StructCopyCore_Root_Main\corerun.exe4.655 ns0.0097 ns0.0090 ns2.05
StructCopyCore_Root_PR\corerun.exe2.269 ns0.0097 ns0.0086 ns1.00

When destination is on the heap, the perf is also improved

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Feb 29, 2024
@ghostghost assigned EgorBoFeb 29, 2024
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

Testing on CI, for now - x64 nonAOT only. Presumably, the asm version can be better than just a loop over R8 (length), e.g. when dest is not in heap then other iterations can be faster.

Example:

structMyStruct{objecta1;objecta2;objecta3;stringa4;stringa5;stringa6;}staticMyStructTest(MyStructms)=>ms;// simple copy of a struct with gc fields

Current codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 49

New codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxmovr8d,6call CORINFO_HELP_ASSIGN_BYREF_BATCHmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 30
Author:EgorBo
Assignees:EgorBo
Labels:

area-CodeGen-coreclr

Milestone:-

add rdi, 8h
add rsi, 8h
dec r8d
jne NextByref

@jkotasjkotasFeb 29, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This loop can hang GC thread suspension.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you please elaborate? Is it because of shadow gc?

@jkotasjkotasFeb 29, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We won't be able to suspend the thread for GC while this helper is running. If this helper is called for large block repeatedly and there is not much managed code running between subsequent invocations of the helper, the GC thread suspension can be very slow or hang.

The same problem existed in BulkMoveWithWriteBarrier. It was fixed in dotnet/coreclr#27776 . You should be able to modify the repro from that PR to use InlineArray to hit the hang with this helper.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I can limit the max batch size. e.g. 1024 gc slots will be handled as 16 call CORINFO_HELP_ASSIGN_BYREF_BATCH batches

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would measure the worst-case scenario, e.g. call MyStruct Test(MyStruct ms) => ms; in a loop on one thread and measure average GC pause time on a second thread.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

From what I see from the diffs, we're mostly dealing with small values like 2 (most popular) and less than 10

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok, then the math does not work - there is something else going on. (For some reason, I thought that the batch size is 16 in your current change.)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If I change the microbenchmark so that the write barrier has to do actual work, I am seeing 15sec - 30sec pause times with your current change: https://gist.github.com/jkotas/d9b48e48935bcc5a6f3dc87db086196d

@EgorBoEgorBoMar 1, 2024

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Don't know why, but this benchmark shows better numbers in my branch than in Main (with 16 slots max)

@EgorBoEgorBoMar 1, 2024

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe because method is just too big in Main and that made gcinfo slower to process? (; Total bytes of code 5012)

@jkotas

jkotas commented Feb 29, 2024

Copy link
Copy Markdown
Member

Observation: The helper is functional equivalent of BulkMoveWithWriteBarrier, with different perf trade-off. Compared toBulkMoveWithWriteBarrier , the helper as currently implemented in the PR tries to be more precise with setting the cards so it is slower but leaves less work to be done at GC time.

@EgorBo

EgorBo commented Feb 29, 2024

Copy link
Copy Markdown
MemberAuthor

Observation: The helper is functional equivalent of BulkMoveWithWriteBarrier, with different perf trade-off. Compared toBulkMoveWithWriteBarrier , the helper as currently implemented in the PR tries to be more precise with setting the cards so it is slower but leaves less work to be done at GC time.

Thanks! Presumably it's easier to just call BulkMoveWithWriteBarrier so then we don't have to duplicate these helpers for all platforms. Do you have an opinion on which way is generally better?

One thing that motivated me is that we've seen write barriers on hot paths in several real world benchmarks (e.g. @kunalspathak recently did). I am not sure this specific one was the culprit (they're generally less advanced on arm64), just looked like a low-hanging fruit to me.

@jkotas

jkotas commented Feb 29, 2024

Copy link
Copy Markdown
Member

Presumably it's easier to just call BulkMoveWithWriteBarrier

BulkMoveWithWriteBarrier has higher fixed overhead than the hand-written assembly helper. I expect that it would be profitable to call it only for blocks above certain size.

we've seen write barriers on hot paths in several real world benchmarks

Is this addressing the hot paths that we have seen?

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Is this addressing the hot paths that we have seen?

Not sure yet, but jitdiffs imply they're not rare (jitdiffs are usually way smaller than SPMI because we only run it for framework libs, no tests, benchmarks, apsnet, etc).

@kunalspathak

Copy link
Copy Markdown
Contributor

on hot paths in several real world benchmarks (e.g. @kunalspathak recently did)

We saw it more on arm64 and we do not know yet if too many samples of JIT_WriteBarrier on the call stack was because the arm64 version is slower or it was due to lack of batching. I see that this one is just for amd64, so won't help the arm64 case that I was seeing.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

I see that this one is just for amd64, so won't help the arm64 case that I was seeing.

It's just for the draft, I planned to support x64 (windows, unix) and arm64 (windows, unix) for both CoreCLR and NativeAOT

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Here is the current diff between non-batched (left) and batched (right) version: https://www.diffchecker.com/clp44lCx/

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@MihuBot

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@jkotas Does it look good now? I changed to 16 slots max and implemented Windows-x64+Unix-x64 on both CoreCLR and NAOT. I'd like to handle ARM64 in a separate PR if you don't mind.

To simplify code-review of the newly added helpers, I made diffs for each against their non-batch version

  1. src/coreclr/nativeaot/Runtime/amd64/WriteBarriers.S: diff
  2. src/coreclr/nativeaot/Runtime/amd64/WriteBarriers.asm: diff
  3. src/coreclr/vm/amd64/JitHelpers_Fast.asm: diff <--- the least verbose diff
  4. src/coreclr/vm/amd64/jithelpers_fast.S: diff

@EgorBo
EgorBo marked this pull request as ready for review March 1, 2024 00:54
@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

This comment was marked as resolved.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr outerloop, runtime-coreclr jitstressregs, runtime-coreclr gcstress0x3-gcstress0xc, runtime-nativeaot-outerloop

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 4 pipeline(s).

@jkotas

Copy link
Copy Markdown
Member

@VSadov What do you think about impact of this change on GC thread suspension? Have you seen bad cases where the current write barriers are stalling GC thread suspension?

I am wondering where we should set the bar for the GC thread suspension impact on changes like this one. There are multiple ways to make it more friendly to GC thread suspension - explicitly poll in the helper, explicit GC polls in the caller, enable hijacking of return address from these helpers, ... .

@VSadov

VSadov commented Mar 1, 2024

Copy link
Copy Markdown
Member

My first thought about this is that the barriers are the kind of platform- and achitecture-specific pieces of assembly, that we would usually avoid for portability reasons. With Core/AOT + 6 architecture + Win/Unix - how many combinations will need a copy of the new barrier? Do we get enough savings for that?

It looks like it saves a call. Also hoists “not in heap” check, although that will help only if it is really not in heap and short-circuits, otherwise that check is cheap. The rest of the cost - other checks, lock or, the actual assignment are still there.
Considering the barrier cost is typically in 1-5% range and that this will be hit in fairly rare cases in the code (“few cases in Roslyn”?), I wonder how much of an improvement is expected here in general?

What happens to the benchmark if the write is to a bunch of array elements instead of a single stack location?
(ideally on a Server GC)

@VSadov

Copy link
Copy Markdown
Member

I kind of suspect the original issue expected more “batching” - like hoisting other checks, setting cards in bulk, etc… But that will turn the barrier into BulkMoveWithWriteBarrier, which we already have.

;; r8: number of byrefs
;;
;; On exit:
;; rdi, rsi are incremented by 8,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shouldn't them be increased by the bytes processed? And also comments in other asm code.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

What happens to the benchmark if the write is to a bunch of array elements instead of a single stack location?
(ideally on a Server GC)

Not much indeed, I don't have a strong opinion here, since it raises concerns/doubts I am fine in closing it before I invest my time into arm side of it 🤷‍♂️ The case where I noticed this issue were tuples, e.g.:

(string,string)StructCopy((string,string)s)=>s;

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Closing, will try to benchmark/compare arm64 write barriers (not only byref one) against x64 to see if I can contribute there instead

@EgorBoEgorBo closed this Mar 1, 2024
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Apr 1, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Consider creating a ranged version of CORINFO_HELP_ASSIGN_BYREF

5 participants

@EgorBo@jkotas@kunalspathak@VSadov@huoyaoyuan
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Batched write barrier for byrefs - #99096

Closed
EgorBo wants to merge 15 commits into
dotnet:mainfrom
EgorBo:batched-byref-wb
Closed

Batched write barrier for byrefs#99096
EgorBo wants to merge 15 commits into
dotnet:mainfrom
EgorBo:batched-byref-wb

Conversation

@EgorBo

@EgorBoEgorBo commented Feb 29, 2024

Copy link
Copy Markdown
Member

Closes#8627

Testing on CI, for now - win-x64 coreclr only. Presumably, the asm version can be better than just a copy of JIT_ByRefWriteBarrier with a loop over R8 (length):

  • If first iteration finds out that the destination doesn't belong to any heap, we should skip all checks for the rest inside the batch
  • TODO: what else
  • Do we need debug-only WRITE_BARRIER_CHECK? we can use the non-batched version for DEBUG/CHECKED
  • REPRET looks like some legacy we can replace with just a normal ret

jit-diffs (seems like a few hits in Roslyn). Size regressions are from the cases when batchSize is 2, two calls are smaller than the batched one + mov r8d, <size>. But the perf should still be better with the batch one.

Example:

structMyStruct{objecta1;objecta2;objecta3;stringa4;stringa5;stringa6;}staticMyStructTest(MyStructms)=>ms;// simple copy of a struct with gc fields

Current codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 49

New codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxmovr8d,6call CORINFO_HELP_ASSIGN_BYREF_BATCHmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 30

Benchmark:

structMyStruct{objecta1;objecta2;objecta3;objecta4;}[Benchmark]publicvoidStructCopy(){MyStructms=default;Test(ms);[MethodImpl(MethodImplOptions.NoInlining)]MyStructTest(MyStructms)=>ms;}
MethodToolchainMeanErrorStdDevRatio
StructCopyCore_Root_Main\corerun.exe4.655 ns0.0097 ns0.0090 ns2.05
StructCopyCore_Root_PR\corerun.exe2.269 ns0.0097 ns0.0086 ns1.00

When destination is on the heap, the perf is also improved

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Feb 29, 2024
@ghostghost assigned EgorBoFeb 29, 2024
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

Testing on CI, for now - x64 nonAOT only. Presumably, the asm version can be better than just a loop over R8 (length), e.g. when dest is not in heap then other iterations can be faster.

Example:

structMyStruct{objecta1;objecta2;objecta3;stringa4;stringa5;stringa6;}staticMyStructTest(MyStructms)=>ms;// simple copy of a struct with gc fields

Current codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 49

New codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxmovr8d,6call CORINFO_HELP_ASSIGN_BYREF_BATCHmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 30
Author:EgorBo
Assignees:EgorBo
Labels:

area-CodeGen-coreclr

Milestone:-

add rdi, 8h
add rsi, 8h
dec r8d
jne NextByref

@jkotasjkotasFeb 29, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This loop can hang GC thread suspension.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you please elaborate? Is it because of shadow gc?

@jkotasjkotasFeb 29, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We won't be able to suspend the thread for GC while this helper is running. If this helper is called for large block repeatedly and there is not much managed code running between subsequent invocations of the helper, the GC thread suspension can be very slow or hang.

The same problem existed in BulkMoveWithWriteBarrier. It was fixed in dotnet/coreclr#27776 . You should be able to modify the repro from that PR to use InlineArray to hit the hang with this helper.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I can limit the max batch size. e.g. 1024 gc slots will be handled as 16 call CORINFO_HELP_ASSIGN_BYREF_BATCH batches

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would measure the worst-case scenario, e.g. call MyStruct Test(MyStruct ms) => ms; in a loop on one thread and measure average GC pause time on a second thread.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

From what I see from the diffs, we're mostly dealing with small values like 2 (most popular) and less than 10

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok, then the math does not work - there is something else going on. (For some reason, I thought that the batch size is 16 in your current change.)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If I change the microbenchmark so that the write barrier has to do actual work, I am seeing 15sec - 30sec pause times with your current change: https://gist.github.com/jkotas/d9b48e48935bcc5a6f3dc87db086196d

@EgorBoEgorBoMar 1, 2024

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Don't know why, but this benchmark shows better numbers in my branch than in Main (with 16 slots max)

@EgorBoEgorBoMar 1, 2024

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe because method is just too big in Main and that made gcinfo slower to process? (; Total bytes of code 5012)

@jkotas

jkotas commented Feb 29, 2024

Copy link
Copy Markdown
Member

Observation: The helper is functional equivalent of BulkMoveWithWriteBarrier, with different perf trade-off. Compared toBulkMoveWithWriteBarrier , the helper as currently implemented in the PR tries to be more precise with setting the cards so it is slower but leaves less work to be done at GC time.

@EgorBo

EgorBo commented Feb 29, 2024

Copy link
Copy Markdown
MemberAuthor

Observation: The helper is functional equivalent of BulkMoveWithWriteBarrier, with different perf trade-off. Compared toBulkMoveWithWriteBarrier , the helper as currently implemented in the PR tries to be more precise with setting the cards so it is slower but leaves less work to be done at GC time.

Thanks! Presumably it's easier to just call BulkMoveWithWriteBarrier so then we don't have to duplicate these helpers for all platforms. Do you have an opinion on which way is generally better?

One thing that motivated me is that we've seen write barriers on hot paths in several real world benchmarks (e.g. @kunalspathak recently did). I am not sure this specific one was the culprit (they're generally less advanced on arm64), just looked like a low-hanging fruit to me.

@jkotas

jkotas commented Feb 29, 2024

Copy link
Copy Markdown
Member

Presumably it's easier to just call BulkMoveWithWriteBarrier

BulkMoveWithWriteBarrier has higher fixed overhead than the hand-written assembly helper. I expect that it would be profitable to call it only for blocks above certain size.

we've seen write barriers on hot paths in several real world benchmarks

Is this addressing the hot paths that we have seen?

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Is this addressing the hot paths that we have seen?

Not sure yet, but jitdiffs imply they're not rare (jitdiffs are usually way smaller than SPMI because we only run it for framework libs, no tests, benchmarks, apsnet, etc).

@kunalspathak

Copy link
Copy Markdown
Contributor

on hot paths in several real world benchmarks (e.g. @kunalspathak recently did)

We saw it more on arm64 and we do not know yet if too many samples of JIT_WriteBarrier on the call stack was because the arm64 version is slower or it was due to lack of batching. I see that this one is just for amd64, so won't help the arm64 case that I was seeing.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

I see that this one is just for amd64, so won't help the arm64 case that I was seeing.

It's just for the draft, I planned to support x64 (windows, unix) and arm64 (windows, unix) for both CoreCLR and NativeAOT

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Here is the current diff between non-batched (left) and batched (right) version: https://www.diffchecker.com/clp44lCx/

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@MihuBot

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@jkotas Does it look good now? I changed to 16 slots max and implemented Windows-x64+Unix-x64 on both CoreCLR and NAOT. I'd like to handle ARM64 in a separate PR if you don't mind.

To simplify code-review of the newly added helpers, I made diffs for each against their non-batch version

  1. src/coreclr/nativeaot/Runtime/amd64/WriteBarriers.S: diff
  2. src/coreclr/nativeaot/Runtime/amd64/WriteBarriers.asm: diff
  3. src/coreclr/vm/amd64/JitHelpers_Fast.asm: diff <--- the least verbose diff
  4. src/coreclr/vm/amd64/jithelpers_fast.S: diff

@EgorBo
EgorBo marked this pull request as ready for review March 1, 2024 00:54
@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

This comment was marked as resolved.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr outerloop, runtime-coreclr jitstressregs, runtime-coreclr gcstress0x3-gcstress0xc, runtime-nativeaot-outerloop

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 4 pipeline(s).

@jkotas

Copy link
Copy Markdown
Member

@VSadov What do you think about impact of this change on GC thread suspension? Have you seen bad cases where the current write barriers are stalling GC thread suspension?

I am wondering where we should set the bar for the GC thread suspension impact on changes like this one. There are multiple ways to make it more friendly to GC thread suspension - explicitly poll in the helper, explicit GC polls in the caller, enable hijacking of return address from these helpers, ... .

@VSadov

VSadov commented Mar 1, 2024

Copy link
Copy Markdown
Member

My first thought about this is that the barriers are the kind of platform- and achitecture-specific pieces of assembly, that we would usually avoid for portability reasons. With Core/AOT + 6 architecture + Win/Unix - how many combinations will need a copy of the new barrier? Do we get enough savings for that?

It looks like it saves a call. Also hoists “not in heap” check, although that will help only if it is really not in heap and short-circuits, otherwise that check is cheap. The rest of the cost - other checks, lock or, the actual assignment are still there.
Considering the barrier cost is typically in 1-5% range and that this will be hit in fairly rare cases in the code (“few cases in Roslyn”?), I wonder how much of an improvement is expected here in general?

What happens to the benchmark if the write is to a bunch of array elements instead of a single stack location?
(ideally on a Server GC)

@VSadov

Copy link
Copy Markdown
Member

I kind of suspect the original issue expected more “batching” - like hoisting other checks, setting cards in bulk, etc… But that will turn the barrier into BulkMoveWithWriteBarrier, which we already have.

;; r8: number of byrefs
;;
;; On exit:
;; rdi, rsi are incremented by 8,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shouldn't them be increased by the bytes processed? And also comments in other asm code.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

What happens to the benchmark if the write is to a bunch of array elements instead of a single stack location?
(ideally on a Server GC)

Not much indeed, I don't have a strong opinion here, since it raises concerns/doubts I am fine in closing it before I invest my time into arm side of it 🤷‍♂️ The case where I noticed this issue were tuples, e.g.:

(string,string)StructCopy((string,string)s)=>s;

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Closing, will try to benchmark/compare arm64 write barriers (not only byref one) against x64 to see if I can contribute there instead

@EgorBoEgorBo closed this Mar 1, 2024
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Apr 1, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Consider creating a ranged version of CORINFO_HELP_ASSIGN_BYREF

5 participants

@EgorBo@jkotas@kunalspathak@VSadov@huoyaoyuan
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Batched write barrier for byrefs - #99096

Closed
EgorBo wants to merge 15 commits into
dotnet:mainfrom
EgorBo:batched-byref-wb
Closed

Batched write barrier for byrefs#99096
EgorBo wants to merge 15 commits into
dotnet:mainfrom
EgorBo:batched-byref-wb

Conversation

@EgorBo

@EgorBoEgorBo commented Feb 29, 2024

Copy link
Copy Markdown
Member

Closes#8627

Testing on CI, for now - win-x64 coreclr only. Presumably, the asm version can be better than just a copy of JIT_ByRefWriteBarrier with a loop over R8 (length):

  • If first iteration finds out that the destination doesn't belong to any heap, we should skip all checks for the rest inside the batch
  • TODO: what else
  • Do we need debug-only WRITE_BARRIER_CHECK? we can use the non-batched version for DEBUG/CHECKED
  • REPRET looks like some legacy we can replace with just a normal ret

jit-diffs (seems like a few hits in Roslyn). Size regressions are from the cases when batchSize is 2, two calls are smaller than the batched one + mov r8d, <size>. But the perf should still be better with the batch one.

Example:

structMyStruct{objecta1;objecta2;objecta3;stringa4;stringa5;stringa6;}staticMyStructTest(MyStructms)=>ms;// simple copy of a struct with gc fields

Current codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 49

New codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxmovr8d,6call CORINFO_HELP_ASSIGN_BYREF_BATCHmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 30

Benchmark:

structMyStruct{objecta1;objecta2;objecta3;objecta4;}[Benchmark]publicvoidStructCopy(){MyStructms=default;Test(ms);[MethodImpl(MethodImplOptions.NoInlining)]MyStructTest(MyStructms)=>ms;}
MethodToolchainMeanErrorStdDevRatio
StructCopyCore_Root_Main\corerun.exe4.655 ns0.0097 ns0.0090 ns2.05
StructCopyCore_Root_PR\corerun.exe2.269 ns0.0097 ns0.0086 ns1.00

When destination is on the heap, the perf is also improved

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Feb 29, 2024
@ghostghost assigned EgorBoFeb 29, 2024
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

Testing on CI, for now - x64 nonAOT only. Presumably, the asm version can be better than just a loop over R8 (length), e.g. when dest is not in heap then other iterations can be faster.

Example:

structMyStruct{objecta1;objecta2;objecta3;stringa4;stringa5;stringa6;}staticMyStructTest(MyStructms)=>ms;// simple copy of a struct with gc fields

Current codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 49

New codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxmovr8d,6call CORINFO_HELP_ASSIGN_BYREF_BATCHmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 30
Author:EgorBo
Assignees:EgorBo
Labels:

area-CodeGen-coreclr

Milestone:-

add rdi, 8h
add rsi, 8h
dec r8d
jne NextByref

@jkotasjkotasFeb 29, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This loop can hang GC thread suspension.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you please elaborate? Is it because of shadow gc?

@jkotasjkotasFeb 29, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We won't be able to suspend the thread for GC while this helper is running. If this helper is called for large block repeatedly and there is not much managed code running between subsequent invocations of the helper, the GC thread suspension can be very slow or hang.

The same problem existed in BulkMoveWithWriteBarrier. It was fixed in dotnet/coreclr#27776 . You should be able to modify the repro from that PR to use InlineArray to hit the hang with this helper.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I can limit the max batch size. e.g. 1024 gc slots will be handled as 16 call CORINFO_HELP_ASSIGN_BYREF_BATCH batches

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would measure the worst-case scenario, e.g. call MyStruct Test(MyStruct ms) => ms; in a loop on one thread and measure average GC pause time on a second thread.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

From what I see from the diffs, we're mostly dealing with small values like 2 (most popular) and less than 10

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok, then the math does not work - there is something else going on. (For some reason, I thought that the batch size is 16 in your current change.)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If I change the microbenchmark so that the write barrier has to do actual work, I am seeing 15sec - 30sec pause times with your current change: https://gist.github.com/jkotas/d9b48e48935bcc5a6f3dc87db086196d

@EgorBoEgorBoMar 1, 2024

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Don't know why, but this benchmark shows better numbers in my branch than in Main (with 16 slots max)

@EgorBoEgorBoMar 1, 2024

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe because method is just too big in Main and that made gcinfo slower to process? (; Total bytes of code 5012)

@jkotas

jkotas commented Feb 29, 2024

Copy link
Copy Markdown
Member

Observation: The helper is functional equivalent of BulkMoveWithWriteBarrier, with different perf trade-off. Compared toBulkMoveWithWriteBarrier , the helper as currently implemented in the PR tries to be more precise with setting the cards so it is slower but leaves less work to be done at GC time.

@EgorBo

EgorBo commented Feb 29, 2024

Copy link
Copy Markdown
MemberAuthor

Observation: The helper is functional equivalent of BulkMoveWithWriteBarrier, with different perf trade-off. Compared toBulkMoveWithWriteBarrier , the helper as currently implemented in the PR tries to be more precise with setting the cards so it is slower but leaves less work to be done at GC time.

Thanks! Presumably it's easier to just call BulkMoveWithWriteBarrier so then we don't have to duplicate these helpers for all platforms. Do you have an opinion on which way is generally better?

One thing that motivated me is that we've seen write barriers on hot paths in several real world benchmarks (e.g. @kunalspathak recently did). I am not sure this specific one was the culprit (they're generally less advanced on arm64), just looked like a low-hanging fruit to me.

@jkotas

jkotas commented Feb 29, 2024

Copy link
Copy Markdown
Member

Presumably it's easier to just call BulkMoveWithWriteBarrier

BulkMoveWithWriteBarrier has higher fixed overhead than the hand-written assembly helper. I expect that it would be profitable to call it only for blocks above certain size.

we've seen write barriers on hot paths in several real world benchmarks

Is this addressing the hot paths that we have seen?

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Is this addressing the hot paths that we have seen?

Not sure yet, but jitdiffs imply they're not rare (jitdiffs are usually way smaller than SPMI because we only run it for framework libs, no tests, benchmarks, apsnet, etc).

@kunalspathak

Copy link
Copy Markdown
Contributor

on hot paths in several real world benchmarks (e.g. @kunalspathak recently did)

We saw it more on arm64 and we do not know yet if too many samples of JIT_WriteBarrier on the call stack was because the arm64 version is slower or it was due to lack of batching. I see that this one is just for amd64, so won't help the arm64 case that I was seeing.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

I see that this one is just for amd64, so won't help the arm64 case that I was seeing.

It's just for the draft, I planned to support x64 (windows, unix) and arm64 (windows, unix) for both CoreCLR and NativeAOT

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Here is the current diff between non-batched (left) and batched (right) version: https://www.diffchecker.com/clp44lCx/

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@MihuBot

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@jkotas Does it look good now? I changed to 16 slots max and implemented Windows-x64+Unix-x64 on both CoreCLR and NAOT. I'd like to handle ARM64 in a separate PR if you don't mind.

To simplify code-review of the newly added helpers, I made diffs for each against their non-batch version

  1. src/coreclr/nativeaot/Runtime/amd64/WriteBarriers.S: diff
  2. src/coreclr/nativeaot/Runtime/amd64/WriteBarriers.asm: diff
  3. src/coreclr/vm/amd64/JitHelpers_Fast.asm: diff <--- the least verbose diff
  4. src/coreclr/vm/amd64/jithelpers_fast.S: diff

@EgorBo
EgorBo marked this pull request as ready for review March 1, 2024 00:54
@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

This comment was marked as resolved.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr outerloop, runtime-coreclr jitstressregs, runtime-coreclr gcstress0x3-gcstress0xc, runtime-nativeaot-outerloop

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 4 pipeline(s).

@jkotas

Copy link
Copy Markdown
Member

@VSadov What do you think about impact of this change on GC thread suspension? Have you seen bad cases where the current write barriers are stalling GC thread suspension?

I am wondering where we should set the bar for the GC thread suspension impact on changes like this one. There are multiple ways to make it more friendly to GC thread suspension - explicitly poll in the helper, explicit GC polls in the caller, enable hijacking of return address from these helpers, ... .

@VSadov

VSadov commented Mar 1, 2024

Copy link
Copy Markdown
Member

My first thought about this is that the barriers are the kind of platform- and achitecture-specific pieces of assembly, that we would usually avoid for portability reasons. With Core/AOT + 6 architecture + Win/Unix - how many combinations will need a copy of the new barrier? Do we get enough savings for that?

It looks like it saves a call. Also hoists “not in heap” check, although that will help only if it is really not in heap and short-circuits, otherwise that check is cheap. The rest of the cost - other checks, lock or, the actual assignment are still there.
Considering the barrier cost is typically in 1-5% range and that this will be hit in fairly rare cases in the code (“few cases in Roslyn”?), I wonder how much of an improvement is expected here in general?

What happens to the benchmark if the write is to a bunch of array elements instead of a single stack location?
(ideally on a Server GC)

@VSadov

Copy link
Copy Markdown
Member

I kind of suspect the original issue expected more “batching” - like hoisting other checks, setting cards in bulk, etc… But that will turn the barrier into BulkMoveWithWriteBarrier, which we already have.

;; r8: number of byrefs
;;
;; On exit:
;; rdi, rsi are incremented by 8,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shouldn't them be increased by the bytes processed? And also comments in other asm code.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

What happens to the benchmark if the write is to a bunch of array elements instead of a single stack location?
(ideally on a Server GC)

Not much indeed, I don't have a strong opinion here, since it raises concerns/doubts I am fine in closing it before I invest my time into arm side of it 🤷‍♂️ The case where I noticed this issue were tuples, e.g.:

(string,string)StructCopy((string,string)s)=>s;

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Closing, will try to benchmark/compare arm64 write barriers (not only byref one) against x64 to see if I can contribute there instead

@EgorBoEgorBo closed this Mar 1, 2024
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Apr 1, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Consider creating a ranged version of CORINFO_HELP_ASSIGN_BYREF

5 participants

@EgorBo@jkotas@kunalspathak@VSadov@huoyaoyuan
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Batched write barrier for byrefs - #99096

Closed
EgorBo wants to merge 15 commits into
dotnet:mainfrom
EgorBo:batched-byref-wb
Closed

Batched write barrier for byrefs#99096
EgorBo wants to merge 15 commits into
dotnet:mainfrom
EgorBo:batched-byref-wb

Conversation

@EgorBo

@EgorBoEgorBo commented Feb 29, 2024

Copy link
Copy Markdown
Member

Closes#8627

Testing on CI, for now - win-x64 coreclr only. Presumably, the asm version can be better than just a copy of JIT_ByRefWriteBarrier with a loop over R8 (length):

  • If first iteration finds out that the destination doesn't belong to any heap, we should skip all checks for the rest inside the batch
  • TODO: what else
  • Do we need debug-only WRITE_BARRIER_CHECK? we can use the non-batched version for DEBUG/CHECKED
  • REPRET looks like some legacy we can replace with just a normal ret

jit-diffs (seems like a few hits in Roslyn). Size regressions are from the cases when batchSize is 2, two calls are smaller than the batched one + mov r8d, <size>. But the perf should still be better with the batch one.

Example:

structMyStruct{objecta1;objecta2;objecta3;stringa4;stringa5;stringa6;}staticMyStructTest(MyStructms)=>ms;// simple copy of a struct with gc fields

Current codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 49

New codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxmovr8d,6call CORINFO_HELP_ASSIGN_BYREF_BATCHmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 30

Benchmark:

structMyStruct{objecta1;objecta2;objecta3;objecta4;}[Benchmark]publicvoidStructCopy(){MyStructms=default;Test(ms);[MethodImpl(MethodImplOptions.NoInlining)]MyStructTest(MyStructms)=>ms;}
MethodToolchainMeanErrorStdDevRatio
StructCopyCore_Root_Main\corerun.exe4.655 ns0.0097 ns0.0090 ns2.05
StructCopyCore_Root_PR\corerun.exe2.269 ns0.0097 ns0.0086 ns1.00

When destination is on the heap, the perf is also improved

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Feb 29, 2024
@ghostghost assigned EgorBoFeb 29, 2024
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

Testing on CI, for now - x64 nonAOT only. Presumably, the asm version can be better than just a loop over R8 (length), e.g. when dest is not in heap then other iterations can be faster.

Example:

structMyStruct{objecta1;objecta2;objecta3;stringa4;stringa5;stringa6;}staticMyStructTest(MyStructms)=>ms;// simple copy of a struct with gc fields

Current codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFcall CORINFO_HELP_ASSIGN_BYREFmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 49

New codegen:

; Assembly listing for method Prog:Test(MyStruct):MyStructG_M32391_IG01:pushrdipushrsipushrbxmovrbx,rcxG_M32391_IG02:movrdi,rbxmovrsi,rdxmovr8d,6call CORINFO_HELP_ASSIGN_BYREF_BATCHmovrax,rbxG_M32391_IG03:poprbxpoprsipoprdiret; Total bytes of code 30
Author:EgorBo
Assignees:EgorBo
Labels:

area-CodeGen-coreclr

Milestone:-

add rdi, 8h
add rsi, 8h
dec r8d
jne NextByref

@jkotasjkotasFeb 29, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This loop can hang GC thread suspension.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you please elaborate? Is it because of shadow gc?

@jkotasjkotasFeb 29, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We won't be able to suspend the thread for GC while this helper is running. If this helper is called for large block repeatedly and there is not much managed code running between subsequent invocations of the helper, the GC thread suspension can be very slow or hang.

The same problem existed in BulkMoveWithWriteBarrier. It was fixed in dotnet/coreclr#27776 . You should be able to modify the repro from that PR to use InlineArray to hit the hang with this helper.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I can limit the max batch size. e.g. 1024 gc slots will be handled as 16 call CORINFO_HELP_ASSIGN_BYREF_BATCH batches

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would measure the worst-case scenario, e.g. call MyStruct Test(MyStruct ms) => ms; in a loop on one thread and measure average GC pause time on a second thread.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

From what I see from the diffs, we're mostly dealing with small values like 2 (most popular) and less than 10

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok, then the math does not work - there is something else going on. (For some reason, I thought that the batch size is 16 in your current change.)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If I change the microbenchmark so that the write barrier has to do actual work, I am seeing 15sec - 30sec pause times with your current change: https://gist.github.com/jkotas/d9b48e48935bcc5a6f3dc87db086196d

@EgorBoEgorBoMar 1, 2024

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Don't know why, but this benchmark shows better numbers in my branch than in Main (with 16 slots max)

@EgorBoEgorBoMar 1, 2024

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe because method is just too big in Main and that made gcinfo slower to process? (; Total bytes of code 5012)

@jkotas

jkotas commented Feb 29, 2024

Copy link
Copy Markdown
Member

Observation: The helper is functional equivalent of BulkMoveWithWriteBarrier, with different perf trade-off. Compared toBulkMoveWithWriteBarrier , the helper as currently implemented in the PR tries to be more precise with setting the cards so it is slower but leaves less work to be done at GC time.

@EgorBo

EgorBo commented Feb 29, 2024

Copy link
Copy Markdown
MemberAuthor

Observation: The helper is functional equivalent of BulkMoveWithWriteBarrier, with different perf trade-off. Compared toBulkMoveWithWriteBarrier , the helper as currently implemented in the PR tries to be more precise with setting the cards so it is slower but leaves less work to be done at GC time.

Thanks! Presumably it's easier to just call BulkMoveWithWriteBarrier so then we don't have to duplicate these helpers for all platforms. Do you have an opinion on which way is generally better?

One thing that motivated me is that we've seen write barriers on hot paths in several real world benchmarks (e.g. @kunalspathak recently did). I am not sure this specific one was the culprit (they're generally less advanced on arm64), just looked like a low-hanging fruit to me.

@jkotas

jkotas commented Feb 29, 2024

Copy link
Copy Markdown
Member

Presumably it's easier to just call BulkMoveWithWriteBarrier

BulkMoveWithWriteBarrier has higher fixed overhead than the hand-written assembly helper. I expect that it would be profitable to call it only for blocks above certain size.

we've seen write barriers on hot paths in several real world benchmarks

Is this addressing the hot paths that we have seen?

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Is this addressing the hot paths that we have seen?

Not sure yet, but jitdiffs imply they're not rare (jitdiffs are usually way smaller than SPMI because we only run it for framework libs, no tests, benchmarks, apsnet, etc).

@kunalspathak

Copy link
Copy Markdown
Contributor

on hot paths in several real world benchmarks (e.g. @kunalspathak recently did)

We saw it more on arm64 and we do not know yet if too many samples of JIT_WriteBarrier on the call stack was because the arm64 version is slower or it was due to lack of batching. I see that this one is just for amd64, so won't help the arm64 case that I was seeing.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

I see that this one is just for amd64, so won't help the arm64 case that I was seeing.

It's just for the draft, I planned to support x64 (windows, unix) and arm64 (windows, unix) for both CoreCLR and NativeAOT

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Here is the current diff between non-batched (left) and batched (right) version: https://www.diffchecker.com/clp44lCx/

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@MihuBot

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@jkotas Does it look good now? I changed to 16 slots max and implemented Windows-x64+Unix-x64 on both CoreCLR and NAOT. I'd like to handle ARM64 in a separate PR if you don't mind.

To simplify code-review of the newly added helpers, I made diffs for each against their non-batch version

  1. src/coreclr/nativeaot/Runtime/amd64/WriteBarriers.S: diff
  2. src/coreclr/nativeaot/Runtime/amd64/WriteBarriers.asm: diff
  3. src/coreclr/vm/amd64/JitHelpers_Fast.asm: diff <--- the least verbose diff
  4. src/coreclr/vm/amd64/jithelpers_fast.S: diff

@EgorBo
EgorBo marked this pull request as ready for review March 1, 2024 00:54
@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

This comment was marked as resolved.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr outerloop, runtime-coreclr jitstressregs, runtime-coreclr gcstress0x3-gcstress0xc, runtime-nativeaot-outerloop

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 4 pipeline(s).

@jkotas

Copy link
Copy Markdown
Member

@VSadov What do you think about impact of this change on GC thread suspension? Have you seen bad cases where the current write barriers are stalling GC thread suspension?

I am wondering where we should set the bar for the GC thread suspension impact on changes like this one. There are multiple ways to make it more friendly to GC thread suspension - explicitly poll in the helper, explicit GC polls in the caller, enable hijacking of return address from these helpers, ... .

@VSadov

VSadov commented Mar 1, 2024

Copy link
Copy Markdown
Member

My first thought about this is that the barriers are the kind of platform- and achitecture-specific pieces of assembly, that we would usually avoid for portability reasons. With Core/AOT + 6 architecture + Win/Unix - how many combinations will need a copy of the new barrier? Do we get enough savings for that?

It looks like it saves a call. Also hoists “not in heap” check, although that will help only if it is really not in heap and short-circuits, otherwise that check is cheap. The rest of the cost - other checks, lock or, the actual assignment are still there.
Considering the barrier cost is typically in 1-5% range and that this will be hit in fairly rare cases in the code (“few cases in Roslyn”?), I wonder how much of an improvement is expected here in general?

What happens to the benchmark if the write is to a bunch of array elements instead of a single stack location?
(ideally on a Server GC)

@VSadov

Copy link
Copy Markdown
Member

I kind of suspect the original issue expected more “batching” - like hoisting other checks, setting cards in bulk, etc… But that will turn the barrier into BulkMoveWithWriteBarrier, which we already have.

;; r8: number of byrefs
;;
;; On exit:
;; rdi, rsi are incremented by 8,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shouldn't them be increased by the bytes processed? And also comments in other asm code.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

What happens to the benchmark if the write is to a bunch of array elements instead of a single stack location?
(ideally on a Server GC)

Not much indeed, I don't have a strong opinion here, since it raises concerns/doubts I am fine in closing it before I invest my time into arm side of it 🤷‍♂️ The case where I noticed this issue were tuples, e.g.:

(string,string)StructCopy((string,string)s)=>s;

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Closing, will try to benchmark/compare arm64 write barriers (not only byref one) against x64 to see if I can contribute there instead

@EgorBoEgorBo closed this Mar 1, 2024
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Apr 1, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Consider creating a ranged version of CORINFO_HELP_ASSIGN_BYREF

5 participants

@EgorBo@jkotas@kunalspathak@VSadov@huoyaoyuan