Align data in Buffer.Memmove for arm64 - #93214

Merged
EgorBo merged 6 commits into
dotnet:mainfrom
EgorBo:align16-memmove
Oct 9, 2023
Merged

Align data in Buffer.Memmove for arm64#93214
EgorBo merged 6 commits into
dotnet:mainfrom
EgorBo:align16-memmove

Conversation

@EgorBo

@EgorBoEgorBo commented Oct 9, 2023

Copy link
Copy Markdown
Member

Currently, we never switch to the native memmove on ARM64 for Mac/Linux. Instead of enabling it back, I tested manual alignment for dest and for most cases saw small wins or the same performance as the native memmove on Apple M2 (macOS) and Ampere (Linux).

What's more important, it shows nice wins compared to the current impl when data is not 16 bytes aligned (even just 8-byte alignment is expensive) which is quite a common case since GC only offers 8-byte alignment and e.g. String's data is always 4-byte aligned, etc. A benchmark on Apple M2:

staticbyte[]_src=newbyte[10000];staticbyte[]_dst=newbyte[10000];[Benchmark]publicvoidCopyTo_align4()=>_src.AsSpan(4).CopyTo(_dst.AsSpan(4));

Presumably, it should also help with the noise in microbenchmarks since GC gives us 8-byte alignment which can be sometimes naturally aligned to 16 bytes or may be not, depends on the Moon's phase.

MethodToolchainMeanErrorStdDevRatio
CopyTo_align4/Core_Root_base/corerun195.1 ns0.75 ns0.70 ns1.68
CopyTo_align4/Core_Root_pr/corerun116.2 ns0.54 ns0.50 ns1.00

I was using the following benchmark to get exact alignment in my tests:

staticBenchmarks(){byte*align64_src=(byte*)NativeMemory.AlignedAlloc(1_000_000,64);byte*align64_dst=(byte*)NativeMemory.AlignedAlloc(1_000_000,64);_align4_src=align64_src+4;_align4_dst=align64_dst+4;}[Benchmark][Arguments(...)]publicvoidCopyTo_align4(intlen)=>newSpan<byte>(_align4_src,len).CopyTo(newSpan<byte>(_align4_dst,len));

@ghostghost assigned EgorBoOct 9, 2023
@ghostghost added the needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners label Oct 9, 2023
@EgorBoEgorBo added area-System.Buffers and removed needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners labels Oct 9, 2023
@ghost

ghost commented Oct 9, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/area-system-buffers
See info in area-owners.md if you want to be subscribed.

Issue Details

Currently, we never switch to the native memmove on ARM64 for Mac/Linux. Instead of enabling it back, I tested manual alignment for dest and for most cases saw wins over the native memmove on Apple M1 (macOS) and Ampere (Linux).

What's more important, it shows nice wins compared to the current impl when data is not 16 bytes aligned (even just 8-byte alignment is expensive) which is quite a common case since GC only offers 8-byte alignment and e.g. String's data is always 4-byte aligned, etc. A benchmark on Apple M2:

staticbyte[]_src=newbyte[10000];staticbyte[]_dst=newbyte[10000];[Benchmark]publicvoidCopyTo_align4()=>_src.AsSpan(4).CopyTo(_dst.AsSpan(4));

Presumably, it should also help with the noise in microbenchmarks since GC gives us 8-byte alignment which can be sometimes naturally aligned to 16 bytes or may be not, depends on the Moon's phase.

MethodToolchainMeanErrorStdDevRatio
CopyTo_align4/Core_Root_base/corerun195.1 ns0.75 ns0.70 ns1.68
CopyTo_align4/Core_Root_pr/corerun116.2 ns0.54 ns0.50 ns1.00
Author:EgorBo
Assignees:EgorBo
Labels:

area-System.Buffers

Milestone:-

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@jkotas PTAL

@jkotas

Copy link
Copy Markdown
Member

// TODO-ARM64-UNIX-OPT revisit when glibc optimized memmove is in Linux distros

Would it be better to fix this TODO instead? What is the perf of your microbenchmark with MemmoveNativeThreshold = 2048?

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

// TODO-ARM64-UNIX-OPT revisit when glibc optimized memmove is in Linux distros

Would it be better to fix this TODO instead? What is the perf of your microbenchmark with MemmoveNativeThreshold = 2048?

Do you mean to jump into native for len>512 uncoditionally? I assume it was disabled for a reason and on some platforms it's slow?

@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

What is the perf of your microbenchmark with MemmoveNativeThreshold = 2048

pretty much the same, sometimes the managed impl is a nanosec or two faster. But it's on Apple M2 where memmove is quite optimized, I am not so sure about other platforms/distros, e.g. I've seen a glibc version that didn't try to align data.

@jkotas

Copy link
Copy Markdown
Member

I assume it was disabled for a reason and on some platforms it's slow?

Yes, it was slow during the initial Arm64 bring up when we run on glibc without any Arm64 specific optimizations. I would expect that memcpy is optimized on all current Arm64 systems that we run on. It is what the TODO alludes to.

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
@EgorBo

Copy link
Copy Markdown
MemberAuthor

Enabled for XARCH as well since the suggested way to align is cheaper. I am seeing wins even on XArch now for misaligned access:

MethodJobToolchainlenMean
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe2562.284 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe2562.299 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe5122.986 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe5125.534 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe10245.717 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe10248.299 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe204811.466 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe204816.798 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe409626.465 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe409627.076 ns

Len = 4096 is done in native.

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

For the reference, codegen of block copy:

[StructLayout(LayoutKind.Sequential,Size=64)]privatestructBlock64{}staticvoidCopyBlock64(refbytedest,refbytesrc){Unsafe.As<byte,Block64>(refdest)= Unsafe.As<byte,Block64>(refsrc);}

AVX512 CPU:

vmovdqu32zmm0, zmmword ptr [rdx]vmovdqu32 zmmword ptr [rcx],zmm0

AVX CPU:

 vmovdqu ymm0, ymmword ptr [rdx] vmovdqu ymmword ptr [rcx],ymm0 vmovdqu ymm0, ymmword ptr [rdx+0x20] vmovdqu ymmword ptr [rcx+0x20],ymm0

SSE2 CPU:

movupsxmm0, xmmword ptr [rdx]movups xmmword ptr [rcx],xmm0movupsxmm0, xmmword ptr [rdx+0x10]movups xmmword ptr [rcx+0x10],xmm0movupsxmm0, xmmword ptr [rdx+0x20]movups xmmword ptr [rcx+0x20],xmm0movupsxmm0, xmmword ptr [rdx+0x30]movups xmmword ptr [rcx+0x30],xmm0

ARM64 NEON CPU:

 ldp q16, q17,[x1] stp q16, q17,[x0] ldp q16, q17,[x1, #0x20] stp q16, q17,[x0, #0x20]

Two notes:

  1. it probably makes sense for JIT to use at least 2 regs for BLK unrolling for better pipelining (or even 4 on some x-archs).
    Although, I wasn't able to detect wins from that on XARCH, but there were wins on ARM64
  2. AVX-512 CPUs might benefit from Block128 copy or even Block256 (My Ryzen with AVX512 won't benefit from it because its AVX-512 is basically 2xAVX2)

@jkotasjkotas left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks

@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

Hopefully, will make benchmarks like this more stable:

image

@EgorBo
EgorBo merged commit f20b2b5 into dotnet:mainOct 9, 2023
@EgorBo
EgorBo deleted the align16-memmove branch October 9, 2023 21:53
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@EgorBo@jkotas@MichalPetryka
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Align data in Buffer.Memmove for arm64 - #93214

Merged
EgorBo merged 6 commits into
dotnet:mainfrom
EgorBo:align16-memmove
Oct 9, 2023
Merged

Align data in Buffer.Memmove for arm64#93214
EgorBo merged 6 commits into
dotnet:mainfrom
EgorBo:align16-memmove

Conversation

@EgorBo

@EgorBoEgorBo commented Oct 9, 2023

Copy link
Copy Markdown
Member

Currently, we never switch to the native memmove on ARM64 for Mac/Linux. Instead of enabling it back, I tested manual alignment for dest and for most cases saw small wins or the same performance as the native memmove on Apple M2 (macOS) and Ampere (Linux).

What's more important, it shows nice wins compared to the current impl when data is not 16 bytes aligned (even just 8-byte alignment is expensive) which is quite a common case since GC only offers 8-byte alignment and e.g. String's data is always 4-byte aligned, etc. A benchmark on Apple M2:

staticbyte[]_src=newbyte[10000];staticbyte[]_dst=newbyte[10000];[Benchmark]publicvoidCopyTo_align4()=>_src.AsSpan(4).CopyTo(_dst.AsSpan(4));

Presumably, it should also help with the noise in microbenchmarks since GC gives us 8-byte alignment which can be sometimes naturally aligned to 16 bytes or may be not, depends on the Moon's phase.

MethodToolchainMeanErrorStdDevRatio
CopyTo_align4/Core_Root_base/corerun195.1 ns0.75 ns0.70 ns1.68
CopyTo_align4/Core_Root_pr/corerun116.2 ns0.54 ns0.50 ns1.00

I was using the following benchmark to get exact alignment in my tests:

staticBenchmarks(){byte*align64_src=(byte*)NativeMemory.AlignedAlloc(1_000_000,64);byte*align64_dst=(byte*)NativeMemory.AlignedAlloc(1_000_000,64);_align4_src=align64_src+4;_align4_dst=align64_dst+4;}[Benchmark][Arguments(...)]publicvoidCopyTo_align4(intlen)=>newSpan<byte>(_align4_src,len).CopyTo(newSpan<byte>(_align4_dst,len));

@ghostghost assigned EgorBoOct 9, 2023
@ghostghost added the needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners label Oct 9, 2023
@EgorBoEgorBo added area-System.Buffers and removed needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners labels Oct 9, 2023
@ghost

ghost commented Oct 9, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/area-system-buffers
See info in area-owners.md if you want to be subscribed.

Issue Details

Currently, we never switch to the native memmove on ARM64 for Mac/Linux. Instead of enabling it back, I tested manual alignment for dest and for most cases saw wins over the native memmove on Apple M1 (macOS) and Ampere (Linux).

What's more important, it shows nice wins compared to the current impl when data is not 16 bytes aligned (even just 8-byte alignment is expensive) which is quite a common case since GC only offers 8-byte alignment and e.g. String's data is always 4-byte aligned, etc. A benchmark on Apple M2:

staticbyte[]_src=newbyte[10000];staticbyte[]_dst=newbyte[10000];[Benchmark]publicvoidCopyTo_align4()=>_src.AsSpan(4).CopyTo(_dst.AsSpan(4));

Presumably, it should also help with the noise in microbenchmarks since GC gives us 8-byte alignment which can be sometimes naturally aligned to 16 bytes or may be not, depends on the Moon's phase.

MethodToolchainMeanErrorStdDevRatio
CopyTo_align4/Core_Root_base/corerun195.1 ns0.75 ns0.70 ns1.68
CopyTo_align4/Core_Root_pr/corerun116.2 ns0.54 ns0.50 ns1.00
Author:EgorBo
Assignees:EgorBo
Labels:

area-System.Buffers

Milestone:-

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@jkotas PTAL

@jkotas

Copy link
Copy Markdown
Member

// TODO-ARM64-UNIX-OPT revisit when glibc optimized memmove is in Linux distros

Would it be better to fix this TODO instead? What is the perf of your microbenchmark with MemmoveNativeThreshold = 2048?

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

// TODO-ARM64-UNIX-OPT revisit when glibc optimized memmove is in Linux distros

Would it be better to fix this TODO instead? What is the perf of your microbenchmark with MemmoveNativeThreshold = 2048?

Do you mean to jump into native for len>512 uncoditionally? I assume it was disabled for a reason and on some platforms it's slow?

@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

What is the perf of your microbenchmark with MemmoveNativeThreshold = 2048

pretty much the same, sometimes the managed impl is a nanosec or two faster. But it's on Apple M2 where memmove is quite optimized, I am not so sure about other platforms/distros, e.g. I've seen a glibc version that didn't try to align data.

@jkotas

Copy link
Copy Markdown
Member

I assume it was disabled for a reason and on some platforms it's slow?

Yes, it was slow during the initial Arm64 bring up when we run on glibc without any Arm64 specific optimizations. I would expect that memcpy is optimized on all current Arm64 systems that we run on. It is what the TODO alludes to.

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
@EgorBo

Copy link
Copy Markdown
MemberAuthor

Enabled for XARCH as well since the suggested way to align is cheaper. I am seeing wins even on XArch now for misaligned access:

MethodJobToolchainlenMean
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe2562.284 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe2562.299 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe5122.986 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe5125.534 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe10245.717 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe10248.299 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe204811.466 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe204816.798 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe409626.465 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe409627.076 ns

Len = 4096 is done in native.

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

For the reference, codegen of block copy:

[StructLayout(LayoutKind.Sequential,Size=64)]privatestructBlock64{}staticvoidCopyBlock64(refbytedest,refbytesrc){Unsafe.As<byte,Block64>(refdest)= Unsafe.As<byte,Block64>(refsrc);}

AVX512 CPU:

vmovdqu32zmm0, zmmword ptr [rdx]vmovdqu32 zmmword ptr [rcx],zmm0

AVX CPU:

 vmovdqu ymm0, ymmword ptr [rdx] vmovdqu ymmword ptr [rcx],ymm0 vmovdqu ymm0, ymmword ptr [rdx+0x20] vmovdqu ymmword ptr [rcx+0x20],ymm0

SSE2 CPU:

movupsxmm0, xmmword ptr [rdx]movups xmmword ptr [rcx],xmm0movupsxmm0, xmmword ptr [rdx+0x10]movups xmmword ptr [rcx+0x10],xmm0movupsxmm0, xmmword ptr [rdx+0x20]movups xmmword ptr [rcx+0x20],xmm0movupsxmm0, xmmword ptr [rdx+0x30]movups xmmword ptr [rcx+0x30],xmm0

ARM64 NEON CPU:

 ldp q16, q17,[x1] stp q16, q17,[x0] ldp q16, q17,[x1, #0x20] stp q16, q17,[x0, #0x20]

Two notes:

  1. it probably makes sense for JIT to use at least 2 regs for BLK unrolling for better pipelining (or even 4 on some x-archs).
    Although, I wasn't able to detect wins from that on XARCH, but there were wins on ARM64
  2. AVX-512 CPUs might benefit from Block128 copy or even Block256 (My Ryzen with AVX512 won't benefit from it because its AVX-512 is basically 2xAVX2)

@jkotasjkotas left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks

@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

Hopefully, will make benchmarks like this more stable:

image

@EgorBo
EgorBo merged commit f20b2b5 into dotnet:mainOct 9, 2023
@EgorBo
EgorBo deleted the align16-memmove branch October 9, 2023 21:53
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@EgorBo@jkotas@MichalPetryka
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Align data in Buffer.Memmove for arm64 - #93214

Merged
EgorBo merged 6 commits into
dotnet:mainfrom
EgorBo:align16-memmove
Oct 9, 2023
Merged

Align data in Buffer.Memmove for arm64#93214
EgorBo merged 6 commits into
dotnet:mainfrom
EgorBo:align16-memmove

Conversation

@EgorBo

@EgorBoEgorBo commented Oct 9, 2023

Copy link
Copy Markdown
Member

Currently, we never switch to the native memmove on ARM64 for Mac/Linux. Instead of enabling it back, I tested manual alignment for dest and for most cases saw small wins or the same performance as the native memmove on Apple M2 (macOS) and Ampere (Linux).

What's more important, it shows nice wins compared to the current impl when data is not 16 bytes aligned (even just 8-byte alignment is expensive) which is quite a common case since GC only offers 8-byte alignment and e.g. String's data is always 4-byte aligned, etc. A benchmark on Apple M2:

staticbyte[]_src=newbyte[10000];staticbyte[]_dst=newbyte[10000];[Benchmark]publicvoidCopyTo_align4()=>_src.AsSpan(4).CopyTo(_dst.AsSpan(4));

Presumably, it should also help with the noise in microbenchmarks since GC gives us 8-byte alignment which can be sometimes naturally aligned to 16 bytes or may be not, depends on the Moon's phase.

MethodToolchainMeanErrorStdDevRatio
CopyTo_align4/Core_Root_base/corerun195.1 ns0.75 ns0.70 ns1.68
CopyTo_align4/Core_Root_pr/corerun116.2 ns0.54 ns0.50 ns1.00

I was using the following benchmark to get exact alignment in my tests:

staticBenchmarks(){byte*align64_src=(byte*)NativeMemory.AlignedAlloc(1_000_000,64);byte*align64_dst=(byte*)NativeMemory.AlignedAlloc(1_000_000,64);_align4_src=align64_src+4;_align4_dst=align64_dst+4;}[Benchmark][Arguments(...)]publicvoidCopyTo_align4(intlen)=>newSpan<byte>(_align4_src,len).CopyTo(newSpan<byte>(_align4_dst,len));

@ghostghost assigned EgorBoOct 9, 2023
@ghostghost added the needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners label Oct 9, 2023
@EgorBoEgorBo added area-System.Buffers and removed needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners labels Oct 9, 2023
@ghost

ghost commented Oct 9, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/area-system-buffers
See info in area-owners.md if you want to be subscribed.

Issue Details

Currently, we never switch to the native memmove on ARM64 for Mac/Linux. Instead of enabling it back, I tested manual alignment for dest and for most cases saw wins over the native memmove on Apple M1 (macOS) and Ampere (Linux).

What's more important, it shows nice wins compared to the current impl when data is not 16 bytes aligned (even just 8-byte alignment is expensive) which is quite a common case since GC only offers 8-byte alignment and e.g. String's data is always 4-byte aligned, etc. A benchmark on Apple M2:

staticbyte[]_src=newbyte[10000];staticbyte[]_dst=newbyte[10000];[Benchmark]publicvoidCopyTo_align4()=>_src.AsSpan(4).CopyTo(_dst.AsSpan(4));

Presumably, it should also help with the noise in microbenchmarks since GC gives us 8-byte alignment which can be sometimes naturally aligned to 16 bytes or may be not, depends on the Moon's phase.

MethodToolchainMeanErrorStdDevRatio
CopyTo_align4/Core_Root_base/corerun195.1 ns0.75 ns0.70 ns1.68
CopyTo_align4/Core_Root_pr/corerun116.2 ns0.54 ns0.50 ns1.00
Author:EgorBo
Assignees:EgorBo
Labels:

area-System.Buffers

Milestone:-

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@jkotas PTAL

@jkotas

Copy link
Copy Markdown
Member

// TODO-ARM64-UNIX-OPT revisit when glibc optimized memmove is in Linux distros

Would it be better to fix this TODO instead? What is the perf of your microbenchmark with MemmoveNativeThreshold = 2048?

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

// TODO-ARM64-UNIX-OPT revisit when glibc optimized memmove is in Linux distros

Would it be better to fix this TODO instead? What is the perf of your microbenchmark with MemmoveNativeThreshold = 2048?

Do you mean to jump into native for len>512 uncoditionally? I assume it was disabled for a reason and on some platforms it's slow?

@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

What is the perf of your microbenchmark with MemmoveNativeThreshold = 2048

pretty much the same, sometimes the managed impl is a nanosec or two faster. But it's on Apple M2 where memmove is quite optimized, I am not so sure about other platforms/distros, e.g. I've seen a glibc version that didn't try to align data.

@jkotas

Copy link
Copy Markdown
Member

I assume it was disabled for a reason and on some platforms it's slow?

Yes, it was slow during the initial Arm64 bring up when we run on glibc without any Arm64 specific optimizations. I would expect that memcpy is optimized on all current Arm64 systems that we run on. It is what the TODO alludes to.

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
@EgorBo

Copy link
Copy Markdown
MemberAuthor

Enabled for XARCH as well since the suggested way to align is cheaper. I am seeing wins even on XArch now for misaligned access:

MethodJobToolchainlenMean
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe2562.284 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe2562.299 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe5122.986 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe5125.534 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe10245.717 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe10248.299 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe204811.466 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe204816.798 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe409626.465 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe409627.076 ns

Len = 4096 is done in native.

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

For the reference, codegen of block copy:

[StructLayout(LayoutKind.Sequential,Size=64)]privatestructBlock64{}staticvoidCopyBlock64(refbytedest,refbytesrc){Unsafe.As<byte,Block64>(refdest)= Unsafe.As<byte,Block64>(refsrc);}

AVX512 CPU:

vmovdqu32zmm0, zmmword ptr [rdx]vmovdqu32 zmmword ptr [rcx],zmm0

AVX CPU:

 vmovdqu ymm0, ymmword ptr [rdx] vmovdqu ymmword ptr [rcx],ymm0 vmovdqu ymm0, ymmword ptr [rdx+0x20] vmovdqu ymmword ptr [rcx+0x20],ymm0

SSE2 CPU:

movupsxmm0, xmmword ptr [rdx]movups xmmword ptr [rcx],xmm0movupsxmm0, xmmword ptr [rdx+0x10]movups xmmword ptr [rcx+0x10],xmm0movupsxmm0, xmmword ptr [rdx+0x20]movups xmmword ptr [rcx+0x20],xmm0movupsxmm0, xmmword ptr [rdx+0x30]movups xmmword ptr [rcx+0x30],xmm0

ARM64 NEON CPU:

 ldp q16, q17,[x1] stp q16, q17,[x0] ldp q16, q17,[x1, #0x20] stp q16, q17,[x0, #0x20]

Two notes:

  1. it probably makes sense for JIT to use at least 2 regs for BLK unrolling for better pipelining (or even 4 on some x-archs).
    Although, I wasn't able to detect wins from that on XARCH, but there were wins on ARM64
  2. AVX-512 CPUs might benefit from Block128 copy or even Block256 (My Ryzen with AVX512 won't benefit from it because its AVX-512 is basically 2xAVX2)

@jkotasjkotas left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks

@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

Hopefully, will make benchmarks like this more stable:

image

@EgorBo
EgorBo merged commit f20b2b5 into dotnet:mainOct 9, 2023
@EgorBo
EgorBo deleted the align16-memmove branch October 9, 2023 21:53
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@EgorBo@jkotas@MichalPetryka
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Align data in Buffer.Memmove for arm64 - #93214

Merged
EgorBo merged 6 commits into
dotnet:mainfrom
EgorBo:align16-memmove
Oct 9, 2023
Merged

Align data in Buffer.Memmove for arm64#93214
EgorBo merged 6 commits into
dotnet:mainfrom
EgorBo:align16-memmove

Conversation

@EgorBo

@EgorBoEgorBo commented Oct 9, 2023

Copy link
Copy Markdown
Member

Currently, we never switch to the native memmove on ARM64 for Mac/Linux. Instead of enabling it back, I tested manual alignment for dest and for most cases saw small wins or the same performance as the native memmove on Apple M2 (macOS) and Ampere (Linux).

What's more important, it shows nice wins compared to the current impl when data is not 16 bytes aligned (even just 8-byte alignment is expensive) which is quite a common case since GC only offers 8-byte alignment and e.g. String's data is always 4-byte aligned, etc. A benchmark on Apple M2:

staticbyte[]_src=newbyte[10000];staticbyte[]_dst=newbyte[10000];[Benchmark]publicvoidCopyTo_align4()=>_src.AsSpan(4).CopyTo(_dst.AsSpan(4));

Presumably, it should also help with the noise in microbenchmarks since GC gives us 8-byte alignment which can be sometimes naturally aligned to 16 bytes or may be not, depends on the Moon's phase.

MethodToolchainMeanErrorStdDevRatio
CopyTo_align4/Core_Root_base/corerun195.1 ns0.75 ns0.70 ns1.68
CopyTo_align4/Core_Root_pr/corerun116.2 ns0.54 ns0.50 ns1.00

I was using the following benchmark to get exact alignment in my tests:

staticBenchmarks(){byte*align64_src=(byte*)NativeMemory.AlignedAlloc(1_000_000,64);byte*align64_dst=(byte*)NativeMemory.AlignedAlloc(1_000_000,64);_align4_src=align64_src+4;_align4_dst=align64_dst+4;}[Benchmark][Arguments(...)]publicvoidCopyTo_align4(intlen)=>newSpan<byte>(_align4_src,len).CopyTo(newSpan<byte>(_align4_dst,len));

@ghostghost assigned EgorBoOct 9, 2023
@ghostghost added the needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners label Oct 9, 2023
@EgorBoEgorBo added area-System.Buffers and removed needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners labels Oct 9, 2023
@ghost

ghost commented Oct 9, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/area-system-buffers
See info in area-owners.md if you want to be subscribed.

Issue Details

Currently, we never switch to the native memmove on ARM64 for Mac/Linux. Instead of enabling it back, I tested manual alignment for dest and for most cases saw wins over the native memmove on Apple M1 (macOS) and Ampere (Linux).

What's more important, it shows nice wins compared to the current impl when data is not 16 bytes aligned (even just 8-byte alignment is expensive) which is quite a common case since GC only offers 8-byte alignment and e.g. String's data is always 4-byte aligned, etc. A benchmark on Apple M2:

staticbyte[]_src=newbyte[10000];staticbyte[]_dst=newbyte[10000];[Benchmark]publicvoidCopyTo_align4()=>_src.AsSpan(4).CopyTo(_dst.AsSpan(4));

Presumably, it should also help with the noise in microbenchmarks since GC gives us 8-byte alignment which can be sometimes naturally aligned to 16 bytes or may be not, depends on the Moon's phase.

MethodToolchainMeanErrorStdDevRatio
CopyTo_align4/Core_Root_base/corerun195.1 ns0.75 ns0.70 ns1.68
CopyTo_align4/Core_Root_pr/corerun116.2 ns0.54 ns0.50 ns1.00
Author:EgorBo
Assignees:EgorBo
Labels:

area-System.Buffers

Milestone:-

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@jkotas PTAL

@jkotas

Copy link
Copy Markdown
Member

// TODO-ARM64-UNIX-OPT revisit when glibc optimized memmove is in Linux distros

Would it be better to fix this TODO instead? What is the perf of your microbenchmark with MemmoveNativeThreshold = 2048?

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

// TODO-ARM64-UNIX-OPT revisit when glibc optimized memmove is in Linux distros

Would it be better to fix this TODO instead? What is the perf of your microbenchmark with MemmoveNativeThreshold = 2048?

Do you mean to jump into native for len>512 uncoditionally? I assume it was disabled for a reason and on some platforms it's slow?

@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

What is the perf of your microbenchmark with MemmoveNativeThreshold = 2048

pretty much the same, sometimes the managed impl is a nanosec or two faster. But it's on Apple M2 where memmove is quite optimized, I am not so sure about other platforms/distros, e.g. I've seen a glibc version that didn't try to align data.

@jkotas

Copy link
Copy Markdown
Member

I assume it was disabled for a reason and on some platforms it's slow?

Yes, it was slow during the initial Arm64 bring up when we run on glibc without any Arm64 specific optimizations. I would expect that memcpy is optimized on all current Arm64 systems that we run on. It is what the TODO alludes to.

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
@EgorBo

Copy link
Copy Markdown
MemberAuthor

Enabled for XARCH as well since the suggested way to align is cheaper. I am seeing wins even on XArch now for misaligned access:

MethodJobToolchainlenMean
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe2562.284 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe2562.299 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe5122.986 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe5125.534 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe10245.717 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe10248.299 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe204811.466 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe204816.798 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe409626.465 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe409627.076 ns

Len = 4096 is done in native.

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

For the reference, codegen of block copy:

[StructLayout(LayoutKind.Sequential,Size=64)]privatestructBlock64{}staticvoidCopyBlock64(refbytedest,refbytesrc){Unsafe.As<byte,Block64>(refdest)= Unsafe.As<byte,Block64>(refsrc);}

AVX512 CPU:

vmovdqu32zmm0, zmmword ptr [rdx]vmovdqu32 zmmword ptr [rcx],zmm0

AVX CPU:

 vmovdqu ymm0, ymmword ptr [rdx] vmovdqu ymmword ptr [rcx],ymm0 vmovdqu ymm0, ymmword ptr [rdx+0x20] vmovdqu ymmword ptr [rcx+0x20],ymm0

SSE2 CPU:

movupsxmm0, xmmword ptr [rdx]movups xmmword ptr [rcx],xmm0movupsxmm0, xmmword ptr [rdx+0x10]movups xmmword ptr [rcx+0x10],xmm0movupsxmm0, xmmword ptr [rdx+0x20]movups xmmword ptr [rcx+0x20],xmm0movupsxmm0, xmmword ptr [rdx+0x30]movups xmmword ptr [rcx+0x30],xmm0

ARM64 NEON CPU:

 ldp q16, q17,[x1] stp q16, q17,[x0] ldp q16, q17,[x1, #0x20] stp q16, q17,[x0, #0x20]

Two notes:

  1. it probably makes sense for JIT to use at least 2 regs for BLK unrolling for better pipelining (or even 4 on some x-archs).
    Although, I wasn't able to detect wins from that on XARCH, but there were wins on ARM64
  2. AVX-512 CPUs might benefit from Block128 copy or even Block256 (My Ryzen with AVX512 won't benefit from it because its AVX-512 is basically 2xAVX2)

@jkotasjkotas left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks

@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

Hopefully, will make benchmarks like this more stable:

image

@EgorBo
EgorBo merged commit f20b2b5 into dotnet:mainOct 9, 2023
@EgorBo
EgorBo deleted the align16-memmove branch October 9, 2023 21:53
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@EgorBo@jkotas@MichalPetryka
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Align data in Buffer.Memmove for arm64 - #93214

Merged
EgorBo merged 6 commits into
dotnet:mainfrom
EgorBo:align16-memmove
Oct 9, 2023
Merged

Align data in Buffer.Memmove for arm64#93214
EgorBo merged 6 commits into
dotnet:mainfrom
EgorBo:align16-memmove

Conversation

@EgorBo

@EgorBoEgorBo commented Oct 9, 2023

Copy link
Copy Markdown
Member

Currently, we never switch to the native memmove on ARM64 for Mac/Linux. Instead of enabling it back, I tested manual alignment for dest and for most cases saw small wins or the same performance as the native memmove on Apple M2 (macOS) and Ampere (Linux).

What's more important, it shows nice wins compared to the current impl when data is not 16 bytes aligned (even just 8-byte alignment is expensive) which is quite a common case since GC only offers 8-byte alignment and e.g. String's data is always 4-byte aligned, etc. A benchmark on Apple M2:

staticbyte[]_src=newbyte[10000];staticbyte[]_dst=newbyte[10000];[Benchmark]publicvoidCopyTo_align4()=>_src.AsSpan(4).CopyTo(_dst.AsSpan(4));

Presumably, it should also help with the noise in microbenchmarks since GC gives us 8-byte alignment which can be sometimes naturally aligned to 16 bytes or may be not, depends on the Moon's phase.

MethodToolchainMeanErrorStdDevRatio
CopyTo_align4/Core_Root_base/corerun195.1 ns0.75 ns0.70 ns1.68
CopyTo_align4/Core_Root_pr/corerun116.2 ns0.54 ns0.50 ns1.00

I was using the following benchmark to get exact alignment in my tests:

staticBenchmarks(){byte*align64_src=(byte*)NativeMemory.AlignedAlloc(1_000_000,64);byte*align64_dst=(byte*)NativeMemory.AlignedAlloc(1_000_000,64);_align4_src=align64_src+4;_align4_dst=align64_dst+4;}[Benchmark][Arguments(...)]publicvoidCopyTo_align4(intlen)=>newSpan<byte>(_align4_src,len).CopyTo(newSpan<byte>(_align4_dst,len));

@ghostghost assigned EgorBoOct 9, 2023
@ghostghost added the needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners label Oct 9, 2023
@EgorBoEgorBo added area-System.Buffers and removed needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners labels Oct 9, 2023
@ghost

ghost commented Oct 9, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/area-system-buffers
See info in area-owners.md if you want to be subscribed.

Issue Details

Currently, we never switch to the native memmove on ARM64 for Mac/Linux. Instead of enabling it back, I tested manual alignment for dest and for most cases saw wins over the native memmove on Apple M1 (macOS) and Ampere (Linux).

What's more important, it shows nice wins compared to the current impl when data is not 16 bytes aligned (even just 8-byte alignment is expensive) which is quite a common case since GC only offers 8-byte alignment and e.g. String's data is always 4-byte aligned, etc. A benchmark on Apple M2:

staticbyte[]_src=newbyte[10000];staticbyte[]_dst=newbyte[10000];[Benchmark]publicvoidCopyTo_align4()=>_src.AsSpan(4).CopyTo(_dst.AsSpan(4));

Presumably, it should also help with the noise in microbenchmarks since GC gives us 8-byte alignment which can be sometimes naturally aligned to 16 bytes or may be not, depends on the Moon's phase.

MethodToolchainMeanErrorStdDevRatio
CopyTo_align4/Core_Root_base/corerun195.1 ns0.75 ns0.70 ns1.68
CopyTo_align4/Core_Root_pr/corerun116.2 ns0.54 ns0.50 ns1.00
Author:EgorBo
Assignees:EgorBo
Labels:

area-System.Buffers

Milestone:-

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@jkotas PTAL

@jkotas

Copy link
Copy Markdown
Member

// TODO-ARM64-UNIX-OPT revisit when glibc optimized memmove is in Linux distros

Would it be better to fix this TODO instead? What is the perf of your microbenchmark with MemmoveNativeThreshold = 2048?

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

// TODO-ARM64-UNIX-OPT revisit when glibc optimized memmove is in Linux distros

Would it be better to fix this TODO instead? What is the perf of your microbenchmark with MemmoveNativeThreshold = 2048?

Do you mean to jump into native for len>512 uncoditionally? I assume it was disabled for a reason and on some platforms it's slow?

@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

What is the perf of your microbenchmark with MemmoveNativeThreshold = 2048

pretty much the same, sometimes the managed impl is a nanosec or two faster. But it's on Apple M2 where memmove is quite optimized, I am not so sure about other platforms/distros, e.g. I've seen a glibc version that didn't try to align data.

@jkotas

Copy link
Copy Markdown
Member

I assume it was disabled for a reason and on some platforms it's slow?

Yes, it was slow during the initial Arm64 bring up when we run on glibc without any Arm64 specific optimizations. I would expect that memcpy is optimized on all current Arm64 systems that we run on. It is what the TODO alludes to.

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
@EgorBo

Copy link
Copy Markdown
MemberAuthor

Enabled for XARCH as well since the suggested way to align is cheaper. I am seeing wins even on XArch now for misaligned access:

MethodJobToolchainlenMean
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe2562.284 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe2562.299 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe5122.986 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe5125.534 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe10245.717 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe10248.299 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe204811.466 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe204816.798 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe409626.465 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe409627.076 ns

Len = 4096 is done in native.

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

For the reference, codegen of block copy:

[StructLayout(LayoutKind.Sequential,Size=64)]privatestructBlock64{}staticvoidCopyBlock64(refbytedest,refbytesrc){Unsafe.As<byte,Block64>(refdest)= Unsafe.As<byte,Block64>(refsrc);}

AVX512 CPU:

vmovdqu32zmm0, zmmword ptr [rdx]vmovdqu32 zmmword ptr [rcx],zmm0

AVX CPU:

 vmovdqu ymm0, ymmword ptr [rdx] vmovdqu ymmword ptr [rcx],ymm0 vmovdqu ymm0, ymmword ptr [rdx+0x20] vmovdqu ymmword ptr [rcx+0x20],ymm0

SSE2 CPU:

movupsxmm0, xmmword ptr [rdx]movups xmmword ptr [rcx],xmm0movupsxmm0, xmmword ptr [rdx+0x10]movups xmmword ptr [rcx+0x10],xmm0movupsxmm0, xmmword ptr [rdx+0x20]movups xmmword ptr [rcx+0x20],xmm0movupsxmm0, xmmword ptr [rdx+0x30]movups xmmword ptr [rcx+0x30],xmm0

ARM64 NEON CPU:

 ldp q16, q17,[x1] stp q16, q17,[x0] ldp q16, q17,[x1, #0x20] stp q16, q17,[x0, #0x20]

Two notes:

  1. it probably makes sense for JIT to use at least 2 regs for BLK unrolling for better pipelining (or even 4 on some x-archs).
    Although, I wasn't able to detect wins from that on XARCH, but there were wins on ARM64
  2. AVX-512 CPUs might benefit from Block128 copy or even Block256 (My Ryzen with AVX512 won't benefit from it because its AVX-512 is basically 2xAVX2)

@jkotasjkotas left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks

@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

Hopefully, will make benchmarks like this more stable:

image

@EgorBo
EgorBo merged commit f20b2b5 into dotnet:mainOct 9, 2023
@EgorBo
EgorBo deleted the align16-memmove branch October 9, 2023 21:53
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@EgorBo@jkotas@MichalPetryka
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Align data in Buffer.Memmove for arm64 - #93214

Merged
EgorBo merged 6 commits into
dotnet:mainfrom
EgorBo:align16-memmove
Oct 9, 2023
Merged

Align data in Buffer.Memmove for arm64#93214
EgorBo merged 6 commits into
dotnet:mainfrom
EgorBo:align16-memmove

Conversation

@EgorBo

@EgorBoEgorBo commented Oct 9, 2023

Copy link
Copy Markdown
Member

Currently, we never switch to the native memmove on ARM64 for Mac/Linux. Instead of enabling it back, I tested manual alignment for dest and for most cases saw small wins or the same performance as the native memmove on Apple M2 (macOS) and Ampere (Linux).

What's more important, it shows nice wins compared to the current impl when data is not 16 bytes aligned (even just 8-byte alignment is expensive) which is quite a common case since GC only offers 8-byte alignment and e.g. String's data is always 4-byte aligned, etc. A benchmark on Apple M2:

staticbyte[]_src=newbyte[10000];staticbyte[]_dst=newbyte[10000];[Benchmark]publicvoidCopyTo_align4()=>_src.AsSpan(4).CopyTo(_dst.AsSpan(4));

Presumably, it should also help with the noise in microbenchmarks since GC gives us 8-byte alignment which can be sometimes naturally aligned to 16 bytes or may be not, depends on the Moon's phase.

MethodToolchainMeanErrorStdDevRatio
CopyTo_align4/Core_Root_base/corerun195.1 ns0.75 ns0.70 ns1.68
CopyTo_align4/Core_Root_pr/corerun116.2 ns0.54 ns0.50 ns1.00

I was using the following benchmark to get exact alignment in my tests:

staticBenchmarks(){byte*align64_src=(byte*)NativeMemory.AlignedAlloc(1_000_000,64);byte*align64_dst=(byte*)NativeMemory.AlignedAlloc(1_000_000,64);_align4_src=align64_src+4;_align4_dst=align64_dst+4;}[Benchmark][Arguments(...)]publicvoidCopyTo_align4(intlen)=>newSpan<byte>(_align4_src,len).CopyTo(newSpan<byte>(_align4_dst,len));

@ghostghost assigned EgorBoOct 9, 2023
@ghostghost added the needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners label Oct 9, 2023
@EgorBoEgorBo added area-System.Buffers and removed needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners labels Oct 9, 2023
@ghost

ghost commented Oct 9, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/area-system-buffers
See info in area-owners.md if you want to be subscribed.

Issue Details

Currently, we never switch to the native memmove on ARM64 for Mac/Linux. Instead of enabling it back, I tested manual alignment for dest and for most cases saw wins over the native memmove on Apple M1 (macOS) and Ampere (Linux).

What's more important, it shows nice wins compared to the current impl when data is not 16 bytes aligned (even just 8-byte alignment is expensive) which is quite a common case since GC only offers 8-byte alignment and e.g. String's data is always 4-byte aligned, etc. A benchmark on Apple M2:

staticbyte[]_src=newbyte[10000];staticbyte[]_dst=newbyte[10000];[Benchmark]publicvoidCopyTo_align4()=>_src.AsSpan(4).CopyTo(_dst.AsSpan(4));

Presumably, it should also help with the noise in microbenchmarks since GC gives us 8-byte alignment which can be sometimes naturally aligned to 16 bytes or may be not, depends on the Moon's phase.

MethodToolchainMeanErrorStdDevRatio
CopyTo_align4/Core_Root_base/corerun195.1 ns0.75 ns0.70 ns1.68
CopyTo_align4/Core_Root_pr/corerun116.2 ns0.54 ns0.50 ns1.00
Author:EgorBo
Assignees:EgorBo
Labels:

area-System.Buffers

Milestone:-

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@jkotas PTAL

@jkotas

Copy link
Copy Markdown
Member

// TODO-ARM64-UNIX-OPT revisit when glibc optimized memmove is in Linux distros

Would it be better to fix this TODO instead? What is the perf of your microbenchmark with MemmoveNativeThreshold = 2048?

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

// TODO-ARM64-UNIX-OPT revisit when glibc optimized memmove is in Linux distros

Would it be better to fix this TODO instead? What is the perf of your microbenchmark with MemmoveNativeThreshold = 2048?

Do you mean to jump into native for len>512 uncoditionally? I assume it was disabled for a reason and on some platforms it's slow?

@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

What is the perf of your microbenchmark with MemmoveNativeThreshold = 2048

pretty much the same, sometimes the managed impl is a nanosec or two faster. But it's on Apple M2 where memmove is quite optimized, I am not so sure about other platforms/distros, e.g. I've seen a glibc version that didn't try to align data.

@jkotas

Copy link
Copy Markdown
Member

I assume it was disabled for a reason and on some platforms it's slow?

Yes, it was slow during the initial Arm64 bring up when we run on glibc without any Arm64 specific optimizations. I would expect that memcpy is optimized on all current Arm64 systems that we run on. It is what the TODO alludes to.

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
@EgorBo

Copy link
Copy Markdown
MemberAuthor

Enabled for XARCH as well since the suggested way to align is cheaper. I am seeing wins even on XArch now for misaligned access:

MethodJobToolchainlenMean
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe2562.284 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe2562.299 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe5122.986 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe5125.534 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe10245.717 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe10248.299 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe204811.466 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe204816.798 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe409626.465 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe409627.076 ns

Len = 4096 is done in native.

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

For the reference, codegen of block copy:

[StructLayout(LayoutKind.Sequential,Size=64)]privatestructBlock64{}staticvoidCopyBlock64(refbytedest,refbytesrc){Unsafe.As<byte,Block64>(refdest)= Unsafe.As<byte,Block64>(refsrc);}

AVX512 CPU:

vmovdqu32zmm0, zmmword ptr [rdx]vmovdqu32 zmmword ptr [rcx],zmm0

AVX CPU:

 vmovdqu ymm0, ymmword ptr [rdx] vmovdqu ymmword ptr [rcx],ymm0 vmovdqu ymm0, ymmword ptr [rdx+0x20] vmovdqu ymmword ptr [rcx+0x20],ymm0

SSE2 CPU:

movupsxmm0, xmmword ptr [rdx]movups xmmword ptr [rcx],xmm0movupsxmm0, xmmword ptr [rdx+0x10]movups xmmword ptr [rcx+0x10],xmm0movupsxmm0, xmmword ptr [rdx+0x20]movups xmmword ptr [rcx+0x20],xmm0movupsxmm0, xmmword ptr [rdx+0x30]movups xmmword ptr [rcx+0x30],xmm0

ARM64 NEON CPU:

 ldp q16, q17,[x1] stp q16, q17,[x0] ldp q16, q17,[x1, #0x20] stp q16, q17,[x0, #0x20]

Two notes:

  1. it probably makes sense for JIT to use at least 2 regs for BLK unrolling for better pipelining (or even 4 on some x-archs).
    Although, I wasn't able to detect wins from that on XARCH, but there were wins on ARM64
  2. AVX-512 CPUs might benefit from Block128 copy or even Block256 (My Ryzen with AVX512 won't benefit from it because its AVX-512 is basically 2xAVX2)

@jkotasjkotas left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks

@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

Hopefully, will make benchmarks like this more stable:

image

@EgorBo
EgorBo merged commit f20b2b5 into dotnet:mainOct 9, 2023
@EgorBo
EgorBo deleted the align16-memmove branch October 9, 2023 21:53
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@EgorBo@jkotas@MichalPetryka
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Align data in Buffer.Memmove for arm64 - #93214

Merged
EgorBo merged 6 commits into
dotnet:mainfrom
EgorBo:align16-memmove
Oct 9, 2023
Merged

Align data in Buffer.Memmove for arm64#93214
EgorBo merged 6 commits into
dotnet:mainfrom
EgorBo:align16-memmove

Conversation

@EgorBo

@EgorBoEgorBo commented Oct 9, 2023

Copy link
Copy Markdown
Member

Currently, we never switch to the native memmove on ARM64 for Mac/Linux. Instead of enabling it back, I tested manual alignment for dest and for most cases saw small wins or the same performance as the native memmove on Apple M2 (macOS) and Ampere (Linux).

What's more important, it shows nice wins compared to the current impl when data is not 16 bytes aligned (even just 8-byte alignment is expensive) which is quite a common case since GC only offers 8-byte alignment and e.g. String's data is always 4-byte aligned, etc. A benchmark on Apple M2:

staticbyte[]_src=newbyte[10000];staticbyte[]_dst=newbyte[10000];[Benchmark]publicvoidCopyTo_align4()=>_src.AsSpan(4).CopyTo(_dst.AsSpan(4));

Presumably, it should also help with the noise in microbenchmarks since GC gives us 8-byte alignment which can be sometimes naturally aligned to 16 bytes or may be not, depends on the Moon's phase.

MethodToolchainMeanErrorStdDevRatio
CopyTo_align4/Core_Root_base/corerun195.1 ns0.75 ns0.70 ns1.68
CopyTo_align4/Core_Root_pr/corerun116.2 ns0.54 ns0.50 ns1.00

I was using the following benchmark to get exact alignment in my tests:

staticBenchmarks(){byte*align64_src=(byte*)NativeMemory.AlignedAlloc(1_000_000,64);byte*align64_dst=(byte*)NativeMemory.AlignedAlloc(1_000_000,64);_align4_src=align64_src+4;_align4_dst=align64_dst+4;}[Benchmark][Arguments(...)]publicvoidCopyTo_align4(intlen)=>newSpan<byte>(_align4_src,len).CopyTo(newSpan<byte>(_align4_dst,len));

@ghostghost assigned EgorBoOct 9, 2023
@ghostghost added the needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners label Oct 9, 2023
@EgorBoEgorBo added area-System.Buffers and removed needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners labels Oct 9, 2023
@ghost

ghost commented Oct 9, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/area-system-buffers
See info in area-owners.md if you want to be subscribed.

Issue Details

Currently, we never switch to the native memmove on ARM64 for Mac/Linux. Instead of enabling it back, I tested manual alignment for dest and for most cases saw wins over the native memmove on Apple M1 (macOS) and Ampere (Linux).

What's more important, it shows nice wins compared to the current impl when data is not 16 bytes aligned (even just 8-byte alignment is expensive) which is quite a common case since GC only offers 8-byte alignment and e.g. String's data is always 4-byte aligned, etc. A benchmark on Apple M2:

staticbyte[]_src=newbyte[10000];staticbyte[]_dst=newbyte[10000];[Benchmark]publicvoidCopyTo_align4()=>_src.AsSpan(4).CopyTo(_dst.AsSpan(4));

Presumably, it should also help with the noise in microbenchmarks since GC gives us 8-byte alignment which can be sometimes naturally aligned to 16 bytes or may be not, depends on the Moon's phase.

MethodToolchainMeanErrorStdDevRatio
CopyTo_align4/Core_Root_base/corerun195.1 ns0.75 ns0.70 ns1.68
CopyTo_align4/Core_Root_pr/corerun116.2 ns0.54 ns0.50 ns1.00
Author:EgorBo
Assignees:EgorBo
Labels:

area-System.Buffers

Milestone:-

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@jkotas PTAL

@jkotas

Copy link
Copy Markdown
Member

// TODO-ARM64-UNIX-OPT revisit when glibc optimized memmove is in Linux distros

Would it be better to fix this TODO instead? What is the perf of your microbenchmark with MemmoveNativeThreshold = 2048?

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

// TODO-ARM64-UNIX-OPT revisit when glibc optimized memmove is in Linux distros

Would it be better to fix this TODO instead? What is the perf of your microbenchmark with MemmoveNativeThreshold = 2048?

Do you mean to jump into native for len>512 uncoditionally? I assume it was disabled for a reason and on some platforms it's slow?

@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

What is the perf of your microbenchmark with MemmoveNativeThreshold = 2048

pretty much the same, sometimes the managed impl is a nanosec or two faster. But it's on Apple M2 where memmove is quite optimized, I am not so sure about other platforms/distros, e.g. I've seen a glibc version that didn't try to align data.

@jkotas

Copy link
Copy Markdown
Member

I assume it was disabled for a reason and on some platforms it's slow?

Yes, it was slow during the initial Arm64 bring up when we run on glibc without any Arm64 specific optimizations. I would expect that memcpy is optimized on all current Arm64 systems that we run on. It is what the TODO alludes to.

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
@EgorBo

Copy link
Copy Markdown
MemberAuthor

Enabled for XARCH as well since the suggested way to align is cheaper. I am seeing wins even on XArch now for misaligned access:

MethodJobToolchainlenMean
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe2562.284 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe2562.299 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe5122.986 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe5125.534 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe10245.717 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe10248.299 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe204811.466 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe204816.798 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe409626.465 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe409627.076 ns

Len = 4096 is done in native.

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

For the reference, codegen of block copy:

[StructLayout(LayoutKind.Sequential,Size=64)]privatestructBlock64{}staticvoidCopyBlock64(refbytedest,refbytesrc){Unsafe.As<byte,Block64>(refdest)= Unsafe.As<byte,Block64>(refsrc);}

AVX512 CPU:

vmovdqu32zmm0, zmmword ptr [rdx]vmovdqu32 zmmword ptr [rcx],zmm0

AVX CPU:

 vmovdqu ymm0, ymmword ptr [rdx] vmovdqu ymmword ptr [rcx],ymm0 vmovdqu ymm0, ymmword ptr [rdx+0x20] vmovdqu ymmword ptr [rcx+0x20],ymm0

SSE2 CPU:

movupsxmm0, xmmword ptr [rdx]movups xmmword ptr [rcx],xmm0movupsxmm0, xmmword ptr [rdx+0x10]movups xmmword ptr [rcx+0x10],xmm0movupsxmm0, xmmword ptr [rdx+0x20]movups xmmword ptr [rcx+0x20],xmm0movupsxmm0, xmmword ptr [rdx+0x30]movups xmmword ptr [rcx+0x30],xmm0

ARM64 NEON CPU:

 ldp q16, q17,[x1] stp q16, q17,[x0] ldp q16, q17,[x1, #0x20] stp q16, q17,[x0, #0x20]

Two notes:

  1. it probably makes sense for JIT to use at least 2 regs for BLK unrolling for better pipelining (or even 4 on some x-archs).
    Although, I wasn't able to detect wins from that on XARCH, but there were wins on ARM64
  2. AVX-512 CPUs might benefit from Block128 copy or even Block256 (My Ryzen with AVX512 won't benefit from it because its AVX-512 is basically 2xAVX2)

@jkotasjkotas left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks

@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

Hopefully, will make benchmarks like this more stable:

image

@EgorBo
EgorBo merged commit f20b2b5 into dotnet:mainOct 9, 2023
@EgorBo
EgorBo deleted the align16-memmove branch October 9, 2023 21:53
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@EgorBo@jkotas@MichalPetryka
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Align data in Buffer.Memmove for arm64 - #93214

Merged
EgorBo merged 6 commits into
dotnet:mainfrom
EgorBo:align16-memmove
Oct 9, 2023
Merged

Align data in Buffer.Memmove for arm64#93214
EgorBo merged 6 commits into
dotnet:mainfrom
EgorBo:align16-memmove

Conversation

@EgorBo

@EgorBoEgorBo commented Oct 9, 2023

Copy link
Copy Markdown
Member

Currently, we never switch to the native memmove on ARM64 for Mac/Linux. Instead of enabling it back, I tested manual alignment for dest and for most cases saw small wins or the same performance as the native memmove on Apple M2 (macOS) and Ampere (Linux).

What's more important, it shows nice wins compared to the current impl when data is not 16 bytes aligned (even just 8-byte alignment is expensive) which is quite a common case since GC only offers 8-byte alignment and e.g. String's data is always 4-byte aligned, etc. A benchmark on Apple M2:

staticbyte[]_src=newbyte[10000];staticbyte[]_dst=newbyte[10000];[Benchmark]publicvoidCopyTo_align4()=>_src.AsSpan(4).CopyTo(_dst.AsSpan(4));

Presumably, it should also help with the noise in microbenchmarks since GC gives us 8-byte alignment which can be sometimes naturally aligned to 16 bytes or may be not, depends on the Moon's phase.

MethodToolchainMeanErrorStdDevRatio
CopyTo_align4/Core_Root_base/corerun195.1 ns0.75 ns0.70 ns1.68
CopyTo_align4/Core_Root_pr/corerun116.2 ns0.54 ns0.50 ns1.00

I was using the following benchmark to get exact alignment in my tests:

staticBenchmarks(){byte*align64_src=(byte*)NativeMemory.AlignedAlloc(1_000_000,64);byte*align64_dst=(byte*)NativeMemory.AlignedAlloc(1_000_000,64);_align4_src=align64_src+4;_align4_dst=align64_dst+4;}[Benchmark][Arguments(...)]publicvoidCopyTo_align4(intlen)=>newSpan<byte>(_align4_src,len).CopyTo(newSpan<byte>(_align4_dst,len));

@ghostghost assigned EgorBoOct 9, 2023
@ghostghost added the needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners label Oct 9, 2023
@EgorBoEgorBo added area-System.Buffers and removed needs-area-label An area label is needed to ensure this gets routed to the appropriate area owners labels Oct 9, 2023
@ghost

ghost commented Oct 9, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/area-system-buffers
See info in area-owners.md if you want to be subscribed.

Issue Details

Currently, we never switch to the native memmove on ARM64 for Mac/Linux. Instead of enabling it back, I tested manual alignment for dest and for most cases saw wins over the native memmove on Apple M1 (macOS) and Ampere (Linux).

What's more important, it shows nice wins compared to the current impl when data is not 16 bytes aligned (even just 8-byte alignment is expensive) which is quite a common case since GC only offers 8-byte alignment and e.g. String's data is always 4-byte aligned, etc. A benchmark on Apple M2:

staticbyte[]_src=newbyte[10000];staticbyte[]_dst=newbyte[10000];[Benchmark]publicvoidCopyTo_align4()=>_src.AsSpan(4).CopyTo(_dst.AsSpan(4));

Presumably, it should also help with the noise in microbenchmarks since GC gives us 8-byte alignment which can be sometimes naturally aligned to 16 bytes or may be not, depends on the Moon's phase.

MethodToolchainMeanErrorStdDevRatio
CopyTo_align4/Core_Root_base/corerun195.1 ns0.75 ns0.70 ns1.68
CopyTo_align4/Core_Root_pr/corerun116.2 ns0.54 ns0.50 ns1.00
Author:EgorBo
Assignees:EgorBo
Labels:

area-System.Buffers

Milestone:-

@EgorBo

Copy link
Copy Markdown
MemberAuthor

@jkotas PTAL

@jkotas

Copy link
Copy Markdown
Member

// TODO-ARM64-UNIX-OPT revisit when glibc optimized memmove is in Linux distros

Would it be better to fix this TODO instead? What is the perf of your microbenchmark with MemmoveNativeThreshold = 2048?

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

// TODO-ARM64-UNIX-OPT revisit when glibc optimized memmove is in Linux distros

Would it be better to fix this TODO instead? What is the perf of your microbenchmark with MemmoveNativeThreshold = 2048?

Do you mean to jump into native for len>512 uncoditionally? I assume it was disabled for a reason and on some platforms it's slow?

@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

What is the perf of your microbenchmark with MemmoveNativeThreshold = 2048

pretty much the same, sometimes the managed impl is a nanosec or two faster. But it's on Apple M2 where memmove is quite optimized, I am not so sure about other platforms/distros, e.g. I've seen a glibc version that didn't try to align data.

@jkotas

Copy link
Copy Markdown
Member

I assume it was disabled for a reason and on some platforms it's slow?

Yes, it was slow during the initial Arm64 bring up when we run on glibc without any Arm64 specific optimizations. I would expect that memcpy is optimized on all current Arm64 systems that we run on. It is what the TODO alludes to.

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
@EgorBo

Copy link
Copy Markdown
MemberAuthor

Enabled for XARCH as well since the suggested way to align is cheaper. I am seeing wins even on XArch now for misaligned access:

MethodJobToolchainlenMean
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe2562.284 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe2562.299 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe5122.986 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe5125.534 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe10245.717 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe10248.299 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe204811.466 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe204816.798 ns
CopyTo_align4Job-VLSCTC\Core_Root\corerun.exe409626.465 ns
CopyTo_align4Job-BCQRRZ\Core_Root_base\corerun.exe409627.076 ns

Len = 4096 is done in native.

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Buffer.cs Outdated
@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

For the reference, codegen of block copy:

[StructLayout(LayoutKind.Sequential,Size=64)]privatestructBlock64{}staticvoidCopyBlock64(refbytedest,refbytesrc){Unsafe.As<byte,Block64>(refdest)= Unsafe.As<byte,Block64>(refsrc);}

AVX512 CPU:

vmovdqu32zmm0, zmmword ptr [rdx]vmovdqu32 zmmword ptr [rcx],zmm0

AVX CPU:

 vmovdqu ymm0, ymmword ptr [rdx] vmovdqu ymmword ptr [rcx],ymm0 vmovdqu ymm0, ymmword ptr [rdx+0x20] vmovdqu ymmword ptr [rcx+0x20],ymm0

SSE2 CPU:

movupsxmm0, xmmword ptr [rdx]movups xmmword ptr [rcx],xmm0movupsxmm0, xmmword ptr [rdx+0x10]movups xmmword ptr [rcx+0x10],xmm0movupsxmm0, xmmword ptr [rdx+0x20]movups xmmword ptr [rcx+0x20],xmm0movupsxmm0, xmmword ptr [rdx+0x30]movups xmmword ptr [rcx+0x30],xmm0

ARM64 NEON CPU:

 ldp q16, q17,[x1] stp q16, q17,[x0] ldp q16, q17,[x1, #0x20] stp q16, q17,[x0, #0x20]

Two notes:

  1. it probably makes sense for JIT to use at least 2 regs for BLK unrolling for better pipelining (or even 4 on some x-archs).
    Although, I wasn't able to detect wins from that on XARCH, but there were wins on ARM64
  2. AVX-512 CPUs might benefit from Block128 copy or even Block256 (My Ryzen with AVX512 won't benefit from it because its AVX-512 is basically 2xAVX2)

@jkotasjkotas left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks

@EgorBo

EgorBo commented Oct 9, 2023

Copy link
Copy Markdown
MemberAuthor

Hopefully, will make benchmarks like this more stable:

image

@EgorBo
EgorBo merged commit f20b2b5 into dotnet:mainOct 9, 2023
@EgorBo
EgorBo deleted the align16-memmove branch October 9, 2023 21:53
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@EgorBo@jkotas@MichalPetryka