Accelerate Vector128<long>::op_Multiply on x64 - #103555

Merged
EgorBo merged 21 commits into
dotnet:mainfrom
EgorBo:arm-mul-64bit
Jun 28, 2024
Merged

Accelerate Vector128<long>::op_Multiply on x64#103555
EgorBo merged 21 commits into
dotnet:mainfrom
EgorBo:arm-mul-64bit

Conversation

@EgorBo

@EgorBoEgorBo commented Jun 17, 2024

Copy link
Copy Markdown
Member

This PR optimizes Vector128 and Vector256 multiplication for long/ulong when AVX512 is not presented in the system. It makes XxHash128 faster, see #103555 (comment)

publicVector128<long>Foo(Vector128<long>a,Vector128<long>b)=>a*b;

Current codegen on x64 cpu without AVX512:

; Method MyBench:Foopushrsipushrbxsubrsp,104movrbx,rdxmovrdx, qword ptr [r8]mov qword ptr [rsp+0x58],rdxmovrdx, qword ptr [r9]mov qword ptr [rsp+0x50],rdxmovrdx, qword ptr [rsp+0x58]imulrdx, qword ptr [rsp+0x50]mov qword ptr [rsp+0x60],rdxmovrsi, qword ptr [rsp+0x60]movrdx, qword ptr [r8+0x08]mov qword ptr [rsp+0x40],rdxmovrdx, qword ptr [r9+0x08]mov qword ptr [rsp+0x38],rdxmovrcx, qword ptr [rsp+0x40]movrdx, qword ptr [rsp+0x38]call[System.Runtime.Intrinsics.Scalar`1[long]:Multiply(long,long):long] ;;; not inlined call!mov qword ptr [rsp+0x48],raxmovrax, qword ptr [rsp+0x48]mov qword ptr [rsp+0x20],rsimov qword ptr [rsp+0x28],rax vmovaps xmm0, xmmword ptr [rsp+0x20]vmovups xmmword ptr [rbx],xmm0movrax,rbxaddrsp,104poprbxpoprsiret; Total bytes of code: 120

New codegen:

; Method MyBench:Foovmovupsxmm0, xmmword ptr [r8]vmovupsxmm1, xmmword ptr [r9] vpmuludq xmm2,xmm1,xmm0 vpshufd xmm1,xmm1,-79 vpmulld xmm0,xmm1,xmm0 vxorps xmm1,xmm1,xmm1vphadddxmm0,xmm0,xmm1 vpshufd xmm0,xmm0,115vpaddqxmm0,xmm0,xmm2vmovups xmmword ptr [rdx],xmm0movrax,rdxret; Total bytes of code: 50

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-runtime-intrinsics
See info in area-owners.md if you want to be subscribed.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Note: results should be better if we do it in JIT, it will enable loop hoisting, cse, etc for MUL

@neon-sunset

Copy link
Copy Markdown
Contributor

Note #103539 (comment) (and https://godbolt.org/z/eqsrf341M) from xxHash128 issue.

EgorBoand others added 2 commits June 17, 2024 17:01
…sics/Vector128_1.cs
Co-authored-by: Tanner Gooding <tagoo@outlook.com>
@dotnetdotnet deleted a comment from EgorBotJun 20, 2024
@dotnetdotnet deleted a comment from EgorBotJun 20, 2024
@EgorBo

Copy link
Copy Markdown
MemberAuthor

@EgorBot -amd -intel -arm64 -profiler --envvars DOTNET_PreferredVectorBitWidth:128

usingSystem.IO.Hashing;usingBenchmarkDotNet.Attributes;publicclassBench{staticreadonlybyte[]Data=newbyte[1000000];[Benchmark]publicbyte[]BenchXxHash128(){XxHash128hash=new();hash.Append(Data);returnhash.GetHashAndReset();}}

@EgorBot

Copy link
Copy Markdown
Benchmark results on Intel
BenchmarkDotNet v0.13.12, Ubuntu 22.04.4 LTS (Jammy Jellyfish)
Intel Xeon Platinum 8370C CPU 2.80GHz, 1 CPU, 16 logical and 8 physical cores
Job-ITXSAG : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX-512F+CD+BW+DQ+VL+VBMI
Job-XSORFZ : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX-512F+CD+BW+DQ+VL+VBMI
EnvironmentVariables=DOTNET_PreferredVectorBitWidth=128
MethodToolchainMeanErrorRatio
BenchXxHash128Main43.41 μs0.087 μs1.00
BenchXxHash128PR43.33 μs0.009 μs1.00

BDN_Artifacts.zip

Flame graphs: Main vs PR 🔥
Hot asm: Main vs PR
Hot functions: Main vs PR

For clean perf results, make sure you have just one [Benchmark] in your app.

@EgorBot

Copy link
Copy Markdown
Benchmark results on Amd
BenchmarkDotNet v0.13.12, Ubuntu 22.04.4 LTS (Jammy Jellyfish)
AMD EPYC 7763, 1 CPU, 16 logical and 8 physical cores
Job-SUBLYH : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX2
Job-OPUYDY : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX2
EnvironmentVariables=DOTNET_PreferredVectorBitWidth=128
MethodToolchainMeanErrorRatio
BenchXxHash128Main71.20 μs0.022 μs1.00
BenchXxHash128PR43.84 μs0.013 μs0.62

BDN_Artifacts.zip

Flame graphs: Main vs PR 🔥
Hot asm: Main vs PR
Hot functions: Main vs PR

For clean perf results, make sure you have just one [Benchmark] in your app.

@EgorBot

Copy link
Copy Markdown
Benchmark results on Arm64
BenchmarkDotNet v0.13.12, Ubuntu 22.04.4 LTS (Jammy Jellyfish)
Unknown processor
Job-EDPWDU : .NET 9.0.0 (42.42.42.42424), Arm64 RyuJIT AdvSIMD
Job-TIALUR : .NET 9.0.0 (42.42.42.42424), Arm64 RyuJIT AdvSIMD
EnvironmentVariables=DOTNET_PreferredVectorBitWidth=128
MethodToolchainMeanErrorRatio
BenchXxHash128Main116.9 μs0.11 μs1.00
BenchXxHash128PR116.8 μs0.07 μs1.00

BDN_Artifacts.zip

Flame graphs: Main vs PR 🔥
Hot asm: Main vs PR
Hot functions: Main vs PR

For clean perf results, make sure you have just one [Benchmark] in your app.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

This comment was marked as resolved.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr jitstress-isas-x86

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@EgorBo

EgorBo commented Jun 21, 2024

Copy link
Copy Markdown
MemberAuthor

@tannergooding PTAL, I'll add arm64 separately, need to test different impls.
I've expanded it in importer similar to existing op_Multiply expansions

Benchmark improvement: #103555 (comment)

@EgorBo
EgorBo marked this pull request as ready for review June 24, 2024 14:26
Comment threadsrc/coreclr/jit/gentree.cpp Outdated
Comment on lines +21627 to +21631
// Vector256<int> tmp3 = Avx2.HorizontalAdd(tmp2.AsInt32(), Vector256<int>.Zero);
GenTreeHWIntrinsic* tmp3 =
gtNewSimdHWIntrinsicNode(type, tmp2, gtNewZeroConNode(type),
is256 ? NI_AVX2_HorizontalAdd : NI_SSSE3_HorizontalAdd,
CORINFO_TYPE_UINT, simdSize);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I know in other places we've started avoiding hadd in favor of shuffle+add, might be worth seeing if that's appropriate here too (low priority, non blocking)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I tried to benchmark different implementations for it and they all were equaly fast e.g. #99871 (comment)

if (TARGET_POINTER_SIZE == 4)
{
// TODO-XARCH-CQ: We should support long/ulong multiplication
// TODO-XARCH-CQ: 32bit support

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What's blocking 32-bit support? It doesn't look like we're using any _X64 intrinsics in the fallback logic?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure to be honest, that check was pre-existing, I only changed comment

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@EgorBo@neon-sunset@EgorBot@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Accelerate Vector128<long>::op_Multiply on x64 - #103555

Merged
EgorBo merged 21 commits into
dotnet:mainfrom
EgorBo:arm-mul-64bit
Jun 28, 2024
Merged

Accelerate Vector128<long>::op_Multiply on x64#103555
EgorBo merged 21 commits into
dotnet:mainfrom
EgorBo:arm-mul-64bit

Conversation

@EgorBo

@EgorBoEgorBo commented Jun 17, 2024

Copy link
Copy Markdown
Member

This PR optimizes Vector128 and Vector256 multiplication for long/ulong when AVX512 is not presented in the system. It makes XxHash128 faster, see #103555 (comment)

publicVector128<long>Foo(Vector128<long>a,Vector128<long>b)=>a*b;

Current codegen on x64 cpu without AVX512:

; Method MyBench:Foopushrsipushrbxsubrsp,104movrbx,rdxmovrdx, qword ptr [r8]mov qword ptr [rsp+0x58],rdxmovrdx, qword ptr [r9]mov qword ptr [rsp+0x50],rdxmovrdx, qword ptr [rsp+0x58]imulrdx, qword ptr [rsp+0x50]mov qword ptr [rsp+0x60],rdxmovrsi, qword ptr [rsp+0x60]movrdx, qword ptr [r8+0x08]mov qword ptr [rsp+0x40],rdxmovrdx, qword ptr [r9+0x08]mov qword ptr [rsp+0x38],rdxmovrcx, qword ptr [rsp+0x40]movrdx, qword ptr [rsp+0x38]call[System.Runtime.Intrinsics.Scalar`1[long]:Multiply(long,long):long] ;;; not inlined call!mov qword ptr [rsp+0x48],raxmovrax, qword ptr [rsp+0x48]mov qword ptr [rsp+0x20],rsimov qword ptr [rsp+0x28],rax vmovaps xmm0, xmmword ptr [rsp+0x20]vmovups xmmword ptr [rbx],xmm0movrax,rbxaddrsp,104poprbxpoprsiret; Total bytes of code: 120

New codegen:

; Method MyBench:Foovmovupsxmm0, xmmword ptr [r8]vmovupsxmm1, xmmword ptr [r9] vpmuludq xmm2,xmm1,xmm0 vpshufd xmm1,xmm1,-79 vpmulld xmm0,xmm1,xmm0 vxorps xmm1,xmm1,xmm1vphadddxmm0,xmm0,xmm1 vpshufd xmm0,xmm0,115vpaddqxmm0,xmm0,xmm2vmovups xmmword ptr [rdx],xmm0movrax,rdxret; Total bytes of code: 50

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-runtime-intrinsics
See info in area-owners.md if you want to be subscribed.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Note: results should be better if we do it in JIT, it will enable loop hoisting, cse, etc for MUL

@neon-sunset

Copy link
Copy Markdown
Contributor

Note #103539 (comment) (and https://godbolt.org/z/eqsrf341M) from xxHash128 issue.

EgorBoand others added 2 commits June 17, 2024 17:01
…sics/Vector128_1.cs
Co-authored-by: Tanner Gooding <tagoo@outlook.com>
@dotnetdotnet deleted a comment from EgorBotJun 20, 2024
@dotnetdotnet deleted a comment from EgorBotJun 20, 2024
@EgorBo

Copy link
Copy Markdown
MemberAuthor

@EgorBot -amd -intel -arm64 -profiler --envvars DOTNET_PreferredVectorBitWidth:128

usingSystem.IO.Hashing;usingBenchmarkDotNet.Attributes;publicclassBench{staticreadonlybyte[]Data=newbyte[1000000];[Benchmark]publicbyte[]BenchXxHash128(){XxHash128hash=new();hash.Append(Data);returnhash.GetHashAndReset();}}

@EgorBot

Copy link
Copy Markdown
Benchmark results on Intel
BenchmarkDotNet v0.13.12, Ubuntu 22.04.4 LTS (Jammy Jellyfish)
Intel Xeon Platinum 8370C CPU 2.80GHz, 1 CPU, 16 logical and 8 physical cores
Job-ITXSAG : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX-512F+CD+BW+DQ+VL+VBMI
Job-XSORFZ : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX-512F+CD+BW+DQ+VL+VBMI
EnvironmentVariables=DOTNET_PreferredVectorBitWidth=128
MethodToolchainMeanErrorRatio
BenchXxHash128Main43.41 μs0.087 μs1.00
BenchXxHash128PR43.33 μs0.009 μs1.00

BDN_Artifacts.zip

Flame graphs: Main vs PR 🔥
Hot asm: Main vs PR
Hot functions: Main vs PR

For clean perf results, make sure you have just one [Benchmark] in your app.

@EgorBot

Copy link
Copy Markdown
Benchmark results on Amd
BenchmarkDotNet v0.13.12, Ubuntu 22.04.4 LTS (Jammy Jellyfish)
AMD EPYC 7763, 1 CPU, 16 logical and 8 physical cores
Job-SUBLYH : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX2
Job-OPUYDY : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX2
EnvironmentVariables=DOTNET_PreferredVectorBitWidth=128
MethodToolchainMeanErrorRatio
BenchXxHash128Main71.20 μs0.022 μs1.00
BenchXxHash128PR43.84 μs0.013 μs0.62

BDN_Artifacts.zip

Flame graphs: Main vs PR 🔥
Hot asm: Main vs PR
Hot functions: Main vs PR

For clean perf results, make sure you have just one [Benchmark] in your app.

@EgorBot

Copy link
Copy Markdown
Benchmark results on Arm64
BenchmarkDotNet v0.13.12, Ubuntu 22.04.4 LTS (Jammy Jellyfish)
Unknown processor
Job-EDPWDU : .NET 9.0.0 (42.42.42.42424), Arm64 RyuJIT AdvSIMD
Job-TIALUR : .NET 9.0.0 (42.42.42.42424), Arm64 RyuJIT AdvSIMD
EnvironmentVariables=DOTNET_PreferredVectorBitWidth=128
MethodToolchainMeanErrorRatio
BenchXxHash128Main116.9 μs0.11 μs1.00
BenchXxHash128PR116.8 μs0.07 μs1.00

BDN_Artifacts.zip

Flame graphs: Main vs PR 🔥
Hot asm: Main vs PR
Hot functions: Main vs PR

For clean perf results, make sure you have just one [Benchmark] in your app.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

This comment was marked as resolved.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr jitstress-isas-x86

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@EgorBo

EgorBo commented Jun 21, 2024

Copy link
Copy Markdown
MemberAuthor

@tannergooding PTAL, I'll add arm64 separately, need to test different impls.
I've expanded it in importer similar to existing op_Multiply expansions

Benchmark improvement: #103555 (comment)

@EgorBo
EgorBo marked this pull request as ready for review June 24, 2024 14:26
Comment threadsrc/coreclr/jit/gentree.cpp Outdated
Comment on lines +21627 to +21631
// Vector256<int> tmp3 = Avx2.HorizontalAdd(tmp2.AsInt32(), Vector256<int>.Zero);
GenTreeHWIntrinsic* tmp3 =
gtNewSimdHWIntrinsicNode(type, tmp2, gtNewZeroConNode(type),
is256 ? NI_AVX2_HorizontalAdd : NI_SSSE3_HorizontalAdd,
CORINFO_TYPE_UINT, simdSize);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I know in other places we've started avoiding hadd in favor of shuffle+add, might be worth seeing if that's appropriate here too (low priority, non blocking)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I tried to benchmark different implementations for it and they all were equaly fast e.g. #99871 (comment)

if (TARGET_POINTER_SIZE == 4)
{
// TODO-XARCH-CQ: We should support long/ulong multiplication
// TODO-XARCH-CQ: 32bit support

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What's blocking 32-bit support? It doesn't look like we're using any _X64 intrinsics in the fallback logic?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure to be honest, that check was pre-existing, I only changed comment

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@EgorBo@neon-sunset@EgorBot@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Accelerate Vector128<long>::op_Multiply on x64 - #103555

Merged
EgorBo merged 21 commits into
dotnet:mainfrom
EgorBo:arm-mul-64bit
Jun 28, 2024
Merged

Accelerate Vector128<long>::op_Multiply on x64#103555
EgorBo merged 21 commits into
dotnet:mainfrom
EgorBo:arm-mul-64bit

Conversation

@EgorBo

@EgorBoEgorBo commented Jun 17, 2024

Copy link
Copy Markdown
Member

This PR optimizes Vector128 and Vector256 multiplication for long/ulong when AVX512 is not presented in the system. It makes XxHash128 faster, see #103555 (comment)

publicVector128<long>Foo(Vector128<long>a,Vector128<long>b)=>a*b;

Current codegen on x64 cpu without AVX512:

; Method MyBench:Foopushrsipushrbxsubrsp,104movrbx,rdxmovrdx, qword ptr [r8]mov qword ptr [rsp+0x58],rdxmovrdx, qword ptr [r9]mov qword ptr [rsp+0x50],rdxmovrdx, qword ptr [rsp+0x58]imulrdx, qword ptr [rsp+0x50]mov qword ptr [rsp+0x60],rdxmovrsi, qword ptr [rsp+0x60]movrdx, qword ptr [r8+0x08]mov qword ptr [rsp+0x40],rdxmovrdx, qword ptr [r9+0x08]mov qword ptr [rsp+0x38],rdxmovrcx, qword ptr [rsp+0x40]movrdx, qword ptr [rsp+0x38]call[System.Runtime.Intrinsics.Scalar`1[long]:Multiply(long,long):long] ;;; not inlined call!mov qword ptr [rsp+0x48],raxmovrax, qword ptr [rsp+0x48]mov qword ptr [rsp+0x20],rsimov qword ptr [rsp+0x28],rax vmovaps xmm0, xmmword ptr [rsp+0x20]vmovups xmmword ptr [rbx],xmm0movrax,rbxaddrsp,104poprbxpoprsiret; Total bytes of code: 120

New codegen:

; Method MyBench:Foovmovupsxmm0, xmmword ptr [r8]vmovupsxmm1, xmmword ptr [r9] vpmuludq xmm2,xmm1,xmm0 vpshufd xmm1,xmm1,-79 vpmulld xmm0,xmm1,xmm0 vxorps xmm1,xmm1,xmm1vphadddxmm0,xmm0,xmm1 vpshufd xmm0,xmm0,115vpaddqxmm0,xmm0,xmm2vmovups xmmword ptr [rdx],xmm0movrax,rdxret; Total bytes of code: 50

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-runtime-intrinsics
See info in area-owners.md if you want to be subscribed.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Note: results should be better if we do it in JIT, it will enable loop hoisting, cse, etc for MUL

@neon-sunset

Copy link
Copy Markdown
Contributor

Note #103539 (comment) (and https://godbolt.org/z/eqsrf341M) from xxHash128 issue.

EgorBoand others added 2 commits June 17, 2024 17:01
…sics/Vector128_1.cs
Co-authored-by: Tanner Gooding <tagoo@outlook.com>
@dotnetdotnet deleted a comment from EgorBotJun 20, 2024
@dotnetdotnet deleted a comment from EgorBotJun 20, 2024
@EgorBo

Copy link
Copy Markdown
MemberAuthor

@EgorBot -amd -intel -arm64 -profiler --envvars DOTNET_PreferredVectorBitWidth:128

usingSystem.IO.Hashing;usingBenchmarkDotNet.Attributes;publicclassBench{staticreadonlybyte[]Data=newbyte[1000000];[Benchmark]publicbyte[]BenchXxHash128(){XxHash128hash=new();hash.Append(Data);returnhash.GetHashAndReset();}}

@EgorBot

Copy link
Copy Markdown
Benchmark results on Intel
BenchmarkDotNet v0.13.12, Ubuntu 22.04.4 LTS (Jammy Jellyfish)
Intel Xeon Platinum 8370C CPU 2.80GHz, 1 CPU, 16 logical and 8 physical cores
Job-ITXSAG : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX-512F+CD+BW+DQ+VL+VBMI
Job-XSORFZ : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX-512F+CD+BW+DQ+VL+VBMI
EnvironmentVariables=DOTNET_PreferredVectorBitWidth=128
MethodToolchainMeanErrorRatio
BenchXxHash128Main43.41 μs0.087 μs1.00
BenchXxHash128PR43.33 μs0.009 μs1.00

BDN_Artifacts.zip

Flame graphs: Main vs PR 🔥
Hot asm: Main vs PR
Hot functions: Main vs PR

For clean perf results, make sure you have just one [Benchmark] in your app.

@EgorBot

Copy link
Copy Markdown
Benchmark results on Amd
BenchmarkDotNet v0.13.12, Ubuntu 22.04.4 LTS (Jammy Jellyfish)
AMD EPYC 7763, 1 CPU, 16 logical and 8 physical cores
Job-SUBLYH : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX2
Job-OPUYDY : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX2
EnvironmentVariables=DOTNET_PreferredVectorBitWidth=128
MethodToolchainMeanErrorRatio
BenchXxHash128Main71.20 μs0.022 μs1.00
BenchXxHash128PR43.84 μs0.013 μs0.62

BDN_Artifacts.zip

Flame graphs: Main vs PR 🔥
Hot asm: Main vs PR
Hot functions: Main vs PR

For clean perf results, make sure you have just one [Benchmark] in your app.

@EgorBot

Copy link
Copy Markdown
Benchmark results on Arm64
BenchmarkDotNet v0.13.12, Ubuntu 22.04.4 LTS (Jammy Jellyfish)
Unknown processor
Job-EDPWDU : .NET 9.0.0 (42.42.42.42424), Arm64 RyuJIT AdvSIMD
Job-TIALUR : .NET 9.0.0 (42.42.42.42424), Arm64 RyuJIT AdvSIMD
EnvironmentVariables=DOTNET_PreferredVectorBitWidth=128
MethodToolchainMeanErrorRatio
BenchXxHash128Main116.9 μs0.11 μs1.00
BenchXxHash128PR116.8 μs0.07 μs1.00

BDN_Artifacts.zip

Flame graphs: Main vs PR 🔥
Hot asm: Main vs PR
Hot functions: Main vs PR

For clean perf results, make sure you have just one [Benchmark] in your app.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

This comment was marked as resolved.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr jitstress-isas-x86

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@EgorBo

EgorBo commented Jun 21, 2024

Copy link
Copy Markdown
MemberAuthor

@tannergooding PTAL, I'll add arm64 separately, need to test different impls.
I've expanded it in importer similar to existing op_Multiply expansions

Benchmark improvement: #103555 (comment)

@EgorBo
EgorBo marked this pull request as ready for review June 24, 2024 14:26
Comment threadsrc/coreclr/jit/gentree.cpp Outdated
Comment on lines +21627 to +21631
// Vector256<int> tmp3 = Avx2.HorizontalAdd(tmp2.AsInt32(), Vector256<int>.Zero);
GenTreeHWIntrinsic* tmp3 =
gtNewSimdHWIntrinsicNode(type, tmp2, gtNewZeroConNode(type),
is256 ? NI_AVX2_HorizontalAdd : NI_SSSE3_HorizontalAdd,
CORINFO_TYPE_UINT, simdSize);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I know in other places we've started avoiding hadd in favor of shuffle+add, might be worth seeing if that's appropriate here too (low priority, non blocking)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I tried to benchmark different implementations for it and they all were equaly fast e.g. #99871 (comment)

if (TARGET_POINTER_SIZE == 4)
{
// TODO-XARCH-CQ: We should support long/ulong multiplication
// TODO-XARCH-CQ: 32bit support

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What's blocking 32-bit support? It doesn't look like we're using any _X64 intrinsics in the fallback logic?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure to be honest, that check was pre-existing, I only changed comment

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@EgorBo@neon-sunset@EgorBot@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Accelerate Vector128<long>::op_Multiply on x64 - #103555

Merged
EgorBo merged 21 commits into
dotnet:mainfrom
EgorBo:arm-mul-64bit
Jun 28, 2024
Merged

Accelerate Vector128<long>::op_Multiply on x64#103555
EgorBo merged 21 commits into
dotnet:mainfrom
EgorBo:arm-mul-64bit

Conversation

@EgorBo

@EgorBoEgorBo commented Jun 17, 2024

Copy link
Copy Markdown
Member

This PR optimizes Vector128 and Vector256 multiplication for long/ulong when AVX512 is not presented in the system. It makes XxHash128 faster, see #103555 (comment)

publicVector128<long>Foo(Vector128<long>a,Vector128<long>b)=>a*b;

Current codegen on x64 cpu without AVX512:

; Method MyBench:Foopushrsipushrbxsubrsp,104movrbx,rdxmovrdx, qword ptr [r8]mov qword ptr [rsp+0x58],rdxmovrdx, qword ptr [r9]mov qword ptr [rsp+0x50],rdxmovrdx, qword ptr [rsp+0x58]imulrdx, qword ptr [rsp+0x50]mov qword ptr [rsp+0x60],rdxmovrsi, qword ptr [rsp+0x60]movrdx, qword ptr [r8+0x08]mov qword ptr [rsp+0x40],rdxmovrdx, qword ptr [r9+0x08]mov qword ptr [rsp+0x38],rdxmovrcx, qword ptr [rsp+0x40]movrdx, qword ptr [rsp+0x38]call[System.Runtime.Intrinsics.Scalar`1[long]:Multiply(long,long):long] ;;; not inlined call!mov qword ptr [rsp+0x48],raxmovrax, qword ptr [rsp+0x48]mov qword ptr [rsp+0x20],rsimov qword ptr [rsp+0x28],rax vmovaps xmm0, xmmword ptr [rsp+0x20]vmovups xmmword ptr [rbx],xmm0movrax,rbxaddrsp,104poprbxpoprsiret; Total bytes of code: 120

New codegen:

; Method MyBench:Foovmovupsxmm0, xmmword ptr [r8]vmovupsxmm1, xmmword ptr [r9] vpmuludq xmm2,xmm1,xmm0 vpshufd xmm1,xmm1,-79 vpmulld xmm0,xmm1,xmm0 vxorps xmm1,xmm1,xmm1vphadddxmm0,xmm0,xmm1 vpshufd xmm0,xmm0,115vpaddqxmm0,xmm0,xmm2vmovups xmmword ptr [rdx],xmm0movrax,rdxret; Total bytes of code: 50

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-runtime-intrinsics
See info in area-owners.md if you want to be subscribed.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Note: results should be better if we do it in JIT, it will enable loop hoisting, cse, etc for MUL

@neon-sunset

Copy link
Copy Markdown
Contributor

Note #103539 (comment) (and https://godbolt.org/z/eqsrf341M) from xxHash128 issue.

EgorBoand others added 2 commits June 17, 2024 17:01
…sics/Vector128_1.cs
Co-authored-by: Tanner Gooding <tagoo@outlook.com>
@dotnetdotnet deleted a comment from EgorBotJun 20, 2024
@dotnetdotnet deleted a comment from EgorBotJun 20, 2024
@EgorBo

Copy link
Copy Markdown
MemberAuthor

@EgorBot -amd -intel -arm64 -profiler --envvars DOTNET_PreferredVectorBitWidth:128

usingSystem.IO.Hashing;usingBenchmarkDotNet.Attributes;publicclassBench{staticreadonlybyte[]Data=newbyte[1000000];[Benchmark]publicbyte[]BenchXxHash128(){XxHash128hash=new();hash.Append(Data);returnhash.GetHashAndReset();}}

@EgorBot

Copy link
Copy Markdown
Benchmark results on Intel
BenchmarkDotNet v0.13.12, Ubuntu 22.04.4 LTS (Jammy Jellyfish)
Intel Xeon Platinum 8370C CPU 2.80GHz, 1 CPU, 16 logical and 8 physical cores
Job-ITXSAG : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX-512F+CD+BW+DQ+VL+VBMI
Job-XSORFZ : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX-512F+CD+BW+DQ+VL+VBMI
EnvironmentVariables=DOTNET_PreferredVectorBitWidth=128
MethodToolchainMeanErrorRatio
BenchXxHash128Main43.41 μs0.087 μs1.00
BenchXxHash128PR43.33 μs0.009 μs1.00

BDN_Artifacts.zip

Flame graphs: Main vs PR 🔥
Hot asm: Main vs PR
Hot functions: Main vs PR

For clean perf results, make sure you have just one [Benchmark] in your app.

@EgorBot

Copy link
Copy Markdown
Benchmark results on Amd
BenchmarkDotNet v0.13.12, Ubuntu 22.04.4 LTS (Jammy Jellyfish)
AMD EPYC 7763, 1 CPU, 16 logical and 8 physical cores
Job-SUBLYH : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX2
Job-OPUYDY : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX2
EnvironmentVariables=DOTNET_PreferredVectorBitWidth=128
MethodToolchainMeanErrorRatio
BenchXxHash128Main71.20 μs0.022 μs1.00
BenchXxHash128PR43.84 μs0.013 μs0.62

BDN_Artifacts.zip

Flame graphs: Main vs PR 🔥
Hot asm: Main vs PR
Hot functions: Main vs PR

For clean perf results, make sure you have just one [Benchmark] in your app.

@EgorBot

Copy link
Copy Markdown
Benchmark results on Arm64
BenchmarkDotNet v0.13.12, Ubuntu 22.04.4 LTS (Jammy Jellyfish)
Unknown processor
Job-EDPWDU : .NET 9.0.0 (42.42.42.42424), Arm64 RyuJIT AdvSIMD
Job-TIALUR : .NET 9.0.0 (42.42.42.42424), Arm64 RyuJIT AdvSIMD
EnvironmentVariables=DOTNET_PreferredVectorBitWidth=128
MethodToolchainMeanErrorRatio
BenchXxHash128Main116.9 μs0.11 μs1.00
BenchXxHash128PR116.8 μs0.07 μs1.00

BDN_Artifacts.zip

Flame graphs: Main vs PR 🔥
Hot asm: Main vs PR
Hot functions: Main vs PR

For clean perf results, make sure you have just one [Benchmark] in your app.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

This comment was marked as resolved.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr jitstress-isas-x86

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@EgorBo

EgorBo commented Jun 21, 2024

Copy link
Copy Markdown
MemberAuthor

@tannergooding PTAL, I'll add arm64 separately, need to test different impls.
I've expanded it in importer similar to existing op_Multiply expansions

Benchmark improvement: #103555 (comment)

@EgorBo
EgorBo marked this pull request as ready for review June 24, 2024 14:26
Comment threadsrc/coreclr/jit/gentree.cpp Outdated
Comment on lines +21627 to +21631
// Vector256<int> tmp3 = Avx2.HorizontalAdd(tmp2.AsInt32(), Vector256<int>.Zero);
GenTreeHWIntrinsic* tmp3 =
gtNewSimdHWIntrinsicNode(type, tmp2, gtNewZeroConNode(type),
is256 ? NI_AVX2_HorizontalAdd : NI_SSSE3_HorizontalAdd,
CORINFO_TYPE_UINT, simdSize);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I know in other places we've started avoiding hadd in favor of shuffle+add, might be worth seeing if that's appropriate here too (low priority, non blocking)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I tried to benchmark different implementations for it and they all were equaly fast e.g. #99871 (comment)

if (TARGET_POINTER_SIZE == 4)
{
// TODO-XARCH-CQ: We should support long/ulong multiplication
// TODO-XARCH-CQ: 32bit support

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What's blocking 32-bit support? It doesn't look like we're using any _X64 intrinsics in the fallback logic?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure to be honest, that check was pre-existing, I only changed comment

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@EgorBo@neon-sunset@EgorBot@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Accelerate Vector128<long>::op_Multiply on x64 - #103555

Merged
EgorBo merged 21 commits into
dotnet:mainfrom
EgorBo:arm-mul-64bit
Jun 28, 2024
Merged

Accelerate Vector128<long>::op_Multiply on x64#103555
EgorBo merged 21 commits into
dotnet:mainfrom
EgorBo:arm-mul-64bit

Conversation

@EgorBo

@EgorBoEgorBo commented Jun 17, 2024

Copy link
Copy Markdown
Member

This PR optimizes Vector128 and Vector256 multiplication for long/ulong when AVX512 is not presented in the system. It makes XxHash128 faster, see #103555 (comment)

publicVector128<long>Foo(Vector128<long>a,Vector128<long>b)=>a*b;

Current codegen on x64 cpu without AVX512:

; Method MyBench:Foopushrsipushrbxsubrsp,104movrbx,rdxmovrdx, qword ptr [r8]mov qword ptr [rsp+0x58],rdxmovrdx, qword ptr [r9]mov qword ptr [rsp+0x50],rdxmovrdx, qword ptr [rsp+0x58]imulrdx, qword ptr [rsp+0x50]mov qword ptr [rsp+0x60],rdxmovrsi, qword ptr [rsp+0x60]movrdx, qword ptr [r8+0x08]mov qword ptr [rsp+0x40],rdxmovrdx, qword ptr [r9+0x08]mov qword ptr [rsp+0x38],rdxmovrcx, qword ptr [rsp+0x40]movrdx, qword ptr [rsp+0x38]call[System.Runtime.Intrinsics.Scalar`1[long]:Multiply(long,long):long] ;;; not inlined call!mov qword ptr [rsp+0x48],raxmovrax, qword ptr [rsp+0x48]mov qword ptr [rsp+0x20],rsimov qword ptr [rsp+0x28],rax vmovaps xmm0, xmmword ptr [rsp+0x20]vmovups xmmword ptr [rbx],xmm0movrax,rbxaddrsp,104poprbxpoprsiret; Total bytes of code: 120

New codegen:

; Method MyBench:Foovmovupsxmm0, xmmword ptr [r8]vmovupsxmm1, xmmword ptr [r9] vpmuludq xmm2,xmm1,xmm0 vpshufd xmm1,xmm1,-79 vpmulld xmm0,xmm1,xmm0 vxorps xmm1,xmm1,xmm1vphadddxmm0,xmm0,xmm1 vpshufd xmm0,xmm0,115vpaddqxmm0,xmm0,xmm2vmovups xmmword ptr [rdx],xmm0movrax,rdxret; Total bytes of code: 50

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-runtime-intrinsics
See info in area-owners.md if you want to be subscribed.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Note: results should be better if we do it in JIT, it will enable loop hoisting, cse, etc for MUL

@neon-sunset

Copy link
Copy Markdown
Contributor

Note #103539 (comment) (and https://godbolt.org/z/eqsrf341M) from xxHash128 issue.

EgorBoand others added 2 commits June 17, 2024 17:01
…sics/Vector128_1.cs
Co-authored-by: Tanner Gooding <tagoo@outlook.com>
@dotnetdotnet deleted a comment from EgorBotJun 20, 2024
@dotnetdotnet deleted a comment from EgorBotJun 20, 2024
@EgorBo

Copy link
Copy Markdown
MemberAuthor

@EgorBot -amd -intel -arm64 -profiler --envvars DOTNET_PreferredVectorBitWidth:128

usingSystem.IO.Hashing;usingBenchmarkDotNet.Attributes;publicclassBench{staticreadonlybyte[]Data=newbyte[1000000];[Benchmark]publicbyte[]BenchXxHash128(){XxHash128hash=new();hash.Append(Data);returnhash.GetHashAndReset();}}

@EgorBot

Copy link
Copy Markdown
Benchmark results on Intel
BenchmarkDotNet v0.13.12, Ubuntu 22.04.4 LTS (Jammy Jellyfish)
Intel Xeon Platinum 8370C CPU 2.80GHz, 1 CPU, 16 logical and 8 physical cores
Job-ITXSAG : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX-512F+CD+BW+DQ+VL+VBMI
Job-XSORFZ : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX-512F+CD+BW+DQ+VL+VBMI
EnvironmentVariables=DOTNET_PreferredVectorBitWidth=128
MethodToolchainMeanErrorRatio
BenchXxHash128Main43.41 μs0.087 μs1.00
BenchXxHash128PR43.33 μs0.009 μs1.00

BDN_Artifacts.zip

Flame graphs: Main vs PR 🔥
Hot asm: Main vs PR
Hot functions: Main vs PR

For clean perf results, make sure you have just one [Benchmark] in your app.

@EgorBot

Copy link
Copy Markdown
Benchmark results on Amd
BenchmarkDotNet v0.13.12, Ubuntu 22.04.4 LTS (Jammy Jellyfish)
AMD EPYC 7763, 1 CPU, 16 logical and 8 physical cores
Job-SUBLYH : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX2
Job-OPUYDY : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX2
EnvironmentVariables=DOTNET_PreferredVectorBitWidth=128
MethodToolchainMeanErrorRatio
BenchXxHash128Main71.20 μs0.022 μs1.00
BenchXxHash128PR43.84 μs0.013 μs0.62

BDN_Artifacts.zip

Flame graphs: Main vs PR 🔥
Hot asm: Main vs PR
Hot functions: Main vs PR

For clean perf results, make sure you have just one [Benchmark] in your app.

@EgorBot

Copy link
Copy Markdown
Benchmark results on Arm64
BenchmarkDotNet v0.13.12, Ubuntu 22.04.4 LTS (Jammy Jellyfish)
Unknown processor
Job-EDPWDU : .NET 9.0.0 (42.42.42.42424), Arm64 RyuJIT AdvSIMD
Job-TIALUR : .NET 9.0.0 (42.42.42.42424), Arm64 RyuJIT AdvSIMD
EnvironmentVariables=DOTNET_PreferredVectorBitWidth=128
MethodToolchainMeanErrorRatio
BenchXxHash128Main116.9 μs0.11 μs1.00
BenchXxHash128PR116.8 μs0.07 μs1.00

BDN_Artifacts.zip

Flame graphs: Main vs PR 🔥
Hot asm: Main vs PR
Hot functions: Main vs PR

For clean perf results, make sure you have just one [Benchmark] in your app.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

This comment was marked as resolved.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr jitstress-isas-x86

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@EgorBo

EgorBo commented Jun 21, 2024

Copy link
Copy Markdown
MemberAuthor

@tannergooding PTAL, I'll add arm64 separately, need to test different impls.
I've expanded it in importer similar to existing op_Multiply expansions

Benchmark improvement: #103555 (comment)

@EgorBo
EgorBo marked this pull request as ready for review June 24, 2024 14:26
Comment threadsrc/coreclr/jit/gentree.cpp Outdated
Comment on lines +21627 to +21631
// Vector256<int> tmp3 = Avx2.HorizontalAdd(tmp2.AsInt32(), Vector256<int>.Zero);
GenTreeHWIntrinsic* tmp3 =
gtNewSimdHWIntrinsicNode(type, tmp2, gtNewZeroConNode(type),
is256 ? NI_AVX2_HorizontalAdd : NI_SSSE3_HorizontalAdd,
CORINFO_TYPE_UINT, simdSize);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I know in other places we've started avoiding hadd in favor of shuffle+add, might be worth seeing if that's appropriate here too (low priority, non blocking)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I tried to benchmark different implementations for it and they all were equaly fast e.g. #99871 (comment)

if (TARGET_POINTER_SIZE == 4)
{
// TODO-XARCH-CQ: We should support long/ulong multiplication
// TODO-XARCH-CQ: 32bit support

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What's blocking 32-bit support? It doesn't look like we're using any _X64 intrinsics in the fallback logic?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure to be honest, that check was pre-existing, I only changed comment

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@EgorBo@neon-sunset@EgorBot@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Accelerate Vector128<long>::op_Multiply on x64 - #103555

Merged
EgorBo merged 21 commits into
dotnet:mainfrom
EgorBo:arm-mul-64bit
Jun 28, 2024
Merged

Accelerate Vector128<long>::op_Multiply on x64#103555
EgorBo merged 21 commits into
dotnet:mainfrom
EgorBo:arm-mul-64bit

Conversation

@EgorBo

@EgorBoEgorBo commented Jun 17, 2024

Copy link
Copy Markdown
Member

This PR optimizes Vector128 and Vector256 multiplication for long/ulong when AVX512 is not presented in the system. It makes XxHash128 faster, see #103555 (comment)

publicVector128<long>Foo(Vector128<long>a,Vector128<long>b)=>a*b;

Current codegen on x64 cpu without AVX512:

; Method MyBench:Foopushrsipushrbxsubrsp,104movrbx,rdxmovrdx, qword ptr [r8]mov qword ptr [rsp+0x58],rdxmovrdx, qword ptr [r9]mov qword ptr [rsp+0x50],rdxmovrdx, qword ptr [rsp+0x58]imulrdx, qword ptr [rsp+0x50]mov qword ptr [rsp+0x60],rdxmovrsi, qword ptr [rsp+0x60]movrdx, qword ptr [r8+0x08]mov qword ptr [rsp+0x40],rdxmovrdx, qword ptr [r9+0x08]mov qword ptr [rsp+0x38],rdxmovrcx, qword ptr [rsp+0x40]movrdx, qword ptr [rsp+0x38]call[System.Runtime.Intrinsics.Scalar`1[long]:Multiply(long,long):long] ;;; not inlined call!mov qword ptr [rsp+0x48],raxmovrax, qword ptr [rsp+0x48]mov qword ptr [rsp+0x20],rsimov qword ptr [rsp+0x28],rax vmovaps xmm0, xmmword ptr [rsp+0x20]vmovups xmmword ptr [rbx],xmm0movrax,rbxaddrsp,104poprbxpoprsiret; Total bytes of code: 120

New codegen:

; Method MyBench:Foovmovupsxmm0, xmmword ptr [r8]vmovupsxmm1, xmmword ptr [r9] vpmuludq xmm2,xmm1,xmm0 vpshufd xmm1,xmm1,-79 vpmulld xmm0,xmm1,xmm0 vxorps xmm1,xmm1,xmm1vphadddxmm0,xmm0,xmm1 vpshufd xmm0,xmm0,115vpaddqxmm0,xmm0,xmm2vmovups xmmword ptr [rdx],xmm0movrax,rdxret; Total bytes of code: 50

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-runtime-intrinsics
See info in area-owners.md if you want to be subscribed.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Note: results should be better if we do it in JIT, it will enable loop hoisting, cse, etc for MUL

@neon-sunset

Copy link
Copy Markdown
Contributor

Note #103539 (comment) (and https://godbolt.org/z/eqsrf341M) from xxHash128 issue.

EgorBoand others added 2 commits June 17, 2024 17:01
…sics/Vector128_1.cs
Co-authored-by: Tanner Gooding <tagoo@outlook.com>
@dotnetdotnet deleted a comment from EgorBotJun 20, 2024
@dotnetdotnet deleted a comment from EgorBotJun 20, 2024
@EgorBo

Copy link
Copy Markdown
MemberAuthor

@EgorBot -amd -intel -arm64 -profiler --envvars DOTNET_PreferredVectorBitWidth:128

usingSystem.IO.Hashing;usingBenchmarkDotNet.Attributes;publicclassBench{staticreadonlybyte[]Data=newbyte[1000000];[Benchmark]publicbyte[]BenchXxHash128(){XxHash128hash=new();hash.Append(Data);returnhash.GetHashAndReset();}}

@EgorBot

Copy link
Copy Markdown
Benchmark results on Intel
BenchmarkDotNet v0.13.12, Ubuntu 22.04.4 LTS (Jammy Jellyfish)
Intel Xeon Platinum 8370C CPU 2.80GHz, 1 CPU, 16 logical and 8 physical cores
Job-ITXSAG : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX-512F+CD+BW+DQ+VL+VBMI
Job-XSORFZ : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX-512F+CD+BW+DQ+VL+VBMI
EnvironmentVariables=DOTNET_PreferredVectorBitWidth=128
MethodToolchainMeanErrorRatio
BenchXxHash128Main43.41 μs0.087 μs1.00
BenchXxHash128PR43.33 μs0.009 μs1.00

BDN_Artifacts.zip

Flame graphs: Main vs PR 🔥
Hot asm: Main vs PR
Hot functions: Main vs PR

For clean perf results, make sure you have just one [Benchmark] in your app.

@EgorBot

Copy link
Copy Markdown
Benchmark results on Amd
BenchmarkDotNet v0.13.12, Ubuntu 22.04.4 LTS (Jammy Jellyfish)
AMD EPYC 7763, 1 CPU, 16 logical and 8 physical cores
Job-SUBLYH : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX2
Job-OPUYDY : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX2
EnvironmentVariables=DOTNET_PreferredVectorBitWidth=128
MethodToolchainMeanErrorRatio
BenchXxHash128Main71.20 μs0.022 μs1.00
BenchXxHash128PR43.84 μs0.013 μs0.62

BDN_Artifacts.zip

Flame graphs: Main vs PR 🔥
Hot asm: Main vs PR
Hot functions: Main vs PR

For clean perf results, make sure you have just one [Benchmark] in your app.

@EgorBot

Copy link
Copy Markdown
Benchmark results on Arm64
BenchmarkDotNet v0.13.12, Ubuntu 22.04.4 LTS (Jammy Jellyfish)
Unknown processor
Job-EDPWDU : .NET 9.0.0 (42.42.42.42424), Arm64 RyuJIT AdvSIMD
Job-TIALUR : .NET 9.0.0 (42.42.42.42424), Arm64 RyuJIT AdvSIMD
EnvironmentVariables=DOTNET_PreferredVectorBitWidth=128
MethodToolchainMeanErrorRatio
BenchXxHash128Main116.9 μs0.11 μs1.00
BenchXxHash128PR116.8 μs0.07 μs1.00

BDN_Artifacts.zip

Flame graphs: Main vs PR 🔥
Hot asm: Main vs PR
Hot functions: Main vs PR

For clean perf results, make sure you have just one [Benchmark] in your app.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

This comment was marked as resolved.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr jitstress-isas-x86

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@EgorBo

EgorBo commented Jun 21, 2024

Copy link
Copy Markdown
MemberAuthor

@tannergooding PTAL, I'll add arm64 separately, need to test different impls.
I've expanded it in importer similar to existing op_Multiply expansions

Benchmark improvement: #103555 (comment)

@EgorBo
EgorBo marked this pull request as ready for review June 24, 2024 14:26
Comment threadsrc/coreclr/jit/gentree.cpp Outdated
Comment on lines +21627 to +21631
// Vector256<int> tmp3 = Avx2.HorizontalAdd(tmp2.AsInt32(), Vector256<int>.Zero);
GenTreeHWIntrinsic* tmp3 =
gtNewSimdHWIntrinsicNode(type, tmp2, gtNewZeroConNode(type),
is256 ? NI_AVX2_HorizontalAdd : NI_SSSE3_HorizontalAdd,
CORINFO_TYPE_UINT, simdSize);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I know in other places we've started avoiding hadd in favor of shuffle+add, might be worth seeing if that's appropriate here too (low priority, non blocking)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I tried to benchmark different implementations for it and they all were equaly fast e.g. #99871 (comment)

if (TARGET_POINTER_SIZE == 4)
{
// TODO-XARCH-CQ: We should support long/ulong multiplication
// TODO-XARCH-CQ: 32bit support

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What's blocking 32-bit support? It doesn't look like we're using any _X64 intrinsics in the fallback logic?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure to be honest, that check was pre-existing, I only changed comment

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@EgorBo@neon-sunset@EgorBot@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Accelerate Vector128<long>::op_Multiply on x64 - #103555

Merged
EgorBo merged 21 commits into
dotnet:mainfrom
EgorBo:arm-mul-64bit
Jun 28, 2024
Merged

Accelerate Vector128<long>::op_Multiply on x64#103555
EgorBo merged 21 commits into
dotnet:mainfrom
EgorBo:arm-mul-64bit

Conversation

@EgorBo

@EgorBoEgorBo commented Jun 17, 2024

Copy link
Copy Markdown
Member

This PR optimizes Vector128 and Vector256 multiplication for long/ulong when AVX512 is not presented in the system. It makes XxHash128 faster, see #103555 (comment)

publicVector128<long>Foo(Vector128<long>a,Vector128<long>b)=>a*b;

Current codegen on x64 cpu without AVX512:

; Method MyBench:Foopushrsipushrbxsubrsp,104movrbx,rdxmovrdx, qword ptr [r8]mov qword ptr [rsp+0x58],rdxmovrdx, qword ptr [r9]mov qword ptr [rsp+0x50],rdxmovrdx, qword ptr [rsp+0x58]imulrdx, qword ptr [rsp+0x50]mov qword ptr [rsp+0x60],rdxmovrsi, qword ptr [rsp+0x60]movrdx, qword ptr [r8+0x08]mov qword ptr [rsp+0x40],rdxmovrdx, qword ptr [r9+0x08]mov qword ptr [rsp+0x38],rdxmovrcx, qword ptr [rsp+0x40]movrdx, qword ptr [rsp+0x38]call[System.Runtime.Intrinsics.Scalar`1[long]:Multiply(long,long):long] ;;; not inlined call!mov qword ptr [rsp+0x48],raxmovrax, qword ptr [rsp+0x48]mov qword ptr [rsp+0x20],rsimov qword ptr [rsp+0x28],rax vmovaps xmm0, xmmword ptr [rsp+0x20]vmovups xmmword ptr [rbx],xmm0movrax,rbxaddrsp,104poprbxpoprsiret; Total bytes of code: 120

New codegen:

; Method MyBench:Foovmovupsxmm0, xmmword ptr [r8]vmovupsxmm1, xmmword ptr [r9] vpmuludq xmm2,xmm1,xmm0 vpshufd xmm1,xmm1,-79 vpmulld xmm0,xmm1,xmm0 vxorps xmm1,xmm1,xmm1vphadddxmm0,xmm0,xmm1 vpshufd xmm0,xmm0,115vpaddqxmm0,xmm0,xmm2vmovups xmmword ptr [rdx],xmm0movrax,rdxret; Total bytes of code: 50

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-runtime-intrinsics
See info in area-owners.md if you want to be subscribed.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Note: results should be better if we do it in JIT, it will enable loop hoisting, cse, etc for MUL

@neon-sunset

Copy link
Copy Markdown
Contributor

Note #103539 (comment) (and https://godbolt.org/z/eqsrf341M) from xxHash128 issue.

EgorBoand others added 2 commits June 17, 2024 17:01
…sics/Vector128_1.cs
Co-authored-by: Tanner Gooding <tagoo@outlook.com>
@dotnetdotnet deleted a comment from EgorBotJun 20, 2024
@dotnetdotnet deleted a comment from EgorBotJun 20, 2024
@EgorBo

Copy link
Copy Markdown
MemberAuthor

@EgorBot -amd -intel -arm64 -profiler --envvars DOTNET_PreferredVectorBitWidth:128

usingSystem.IO.Hashing;usingBenchmarkDotNet.Attributes;publicclassBench{staticreadonlybyte[]Data=newbyte[1000000];[Benchmark]publicbyte[]BenchXxHash128(){XxHash128hash=new();hash.Append(Data);returnhash.GetHashAndReset();}}

@EgorBot

Copy link
Copy Markdown
Benchmark results on Intel
BenchmarkDotNet v0.13.12, Ubuntu 22.04.4 LTS (Jammy Jellyfish)
Intel Xeon Platinum 8370C CPU 2.80GHz, 1 CPU, 16 logical and 8 physical cores
Job-ITXSAG : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX-512F+CD+BW+DQ+VL+VBMI
Job-XSORFZ : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX-512F+CD+BW+DQ+VL+VBMI
EnvironmentVariables=DOTNET_PreferredVectorBitWidth=128
MethodToolchainMeanErrorRatio
BenchXxHash128Main43.41 μs0.087 μs1.00
BenchXxHash128PR43.33 μs0.009 μs1.00

BDN_Artifacts.zip

Flame graphs: Main vs PR 🔥
Hot asm: Main vs PR
Hot functions: Main vs PR

For clean perf results, make sure you have just one [Benchmark] in your app.

@EgorBot

Copy link
Copy Markdown
Benchmark results on Amd
BenchmarkDotNet v0.13.12, Ubuntu 22.04.4 LTS (Jammy Jellyfish)
AMD EPYC 7763, 1 CPU, 16 logical and 8 physical cores
Job-SUBLYH : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX2
Job-OPUYDY : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX2
EnvironmentVariables=DOTNET_PreferredVectorBitWidth=128
MethodToolchainMeanErrorRatio
BenchXxHash128Main71.20 μs0.022 μs1.00
BenchXxHash128PR43.84 μs0.013 μs0.62

BDN_Artifacts.zip

Flame graphs: Main vs PR 🔥
Hot asm: Main vs PR
Hot functions: Main vs PR

For clean perf results, make sure you have just one [Benchmark] in your app.

@EgorBot

Copy link
Copy Markdown
Benchmark results on Arm64
BenchmarkDotNet v0.13.12, Ubuntu 22.04.4 LTS (Jammy Jellyfish)
Unknown processor
Job-EDPWDU : .NET 9.0.0 (42.42.42.42424), Arm64 RyuJIT AdvSIMD
Job-TIALUR : .NET 9.0.0 (42.42.42.42424), Arm64 RyuJIT AdvSIMD
EnvironmentVariables=DOTNET_PreferredVectorBitWidth=128
MethodToolchainMeanErrorRatio
BenchXxHash128Main116.9 μs0.11 μs1.00
BenchXxHash128PR116.8 μs0.07 μs1.00

BDN_Artifacts.zip

Flame graphs: Main vs PR 🔥
Hot asm: Main vs PR
Hot functions: Main vs PR

For clean perf results, make sure you have just one [Benchmark] in your app.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

This comment was marked as resolved.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr jitstress-isas-x86

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@EgorBo

EgorBo commented Jun 21, 2024

Copy link
Copy Markdown
MemberAuthor

@tannergooding PTAL, I'll add arm64 separately, need to test different impls.
I've expanded it in importer similar to existing op_Multiply expansions

Benchmark improvement: #103555 (comment)

@EgorBo
EgorBo marked this pull request as ready for review June 24, 2024 14:26
Comment threadsrc/coreclr/jit/gentree.cpp Outdated
Comment on lines +21627 to +21631
// Vector256<int> tmp3 = Avx2.HorizontalAdd(tmp2.AsInt32(), Vector256<int>.Zero);
GenTreeHWIntrinsic* tmp3 =
gtNewSimdHWIntrinsicNode(type, tmp2, gtNewZeroConNode(type),
is256 ? NI_AVX2_HorizontalAdd : NI_SSSE3_HorizontalAdd,
CORINFO_TYPE_UINT, simdSize);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I know in other places we've started avoiding hadd in favor of shuffle+add, might be worth seeing if that's appropriate here too (low priority, non blocking)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I tried to benchmark different implementations for it and they all were equaly fast e.g. #99871 (comment)

if (TARGET_POINTER_SIZE == 4)
{
// TODO-XARCH-CQ: We should support long/ulong multiplication
// TODO-XARCH-CQ: 32bit support

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What's blocking 32-bit support? It doesn't look like we're using any _X64 intrinsics in the fallback logic?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure to be honest, that check was pre-existing, I only changed comment

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@EgorBo@neon-sunset@EgorBot@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Accelerate Vector128<long>::op_Multiply on x64 - #103555

Merged
EgorBo merged 21 commits into
dotnet:mainfrom
EgorBo:arm-mul-64bit
Jun 28, 2024
Merged

Accelerate Vector128<long>::op_Multiply on x64#103555
EgorBo merged 21 commits into
dotnet:mainfrom
EgorBo:arm-mul-64bit

Conversation

@EgorBo

@EgorBoEgorBo commented Jun 17, 2024

Copy link
Copy Markdown
Member

This PR optimizes Vector128 and Vector256 multiplication for long/ulong when AVX512 is not presented in the system. It makes XxHash128 faster, see #103555 (comment)

publicVector128<long>Foo(Vector128<long>a,Vector128<long>b)=>a*b;

Current codegen on x64 cpu without AVX512:

; Method MyBench:Foopushrsipushrbxsubrsp,104movrbx,rdxmovrdx, qword ptr [r8]mov qword ptr [rsp+0x58],rdxmovrdx, qword ptr [r9]mov qword ptr [rsp+0x50],rdxmovrdx, qword ptr [rsp+0x58]imulrdx, qword ptr [rsp+0x50]mov qword ptr [rsp+0x60],rdxmovrsi, qword ptr [rsp+0x60]movrdx, qword ptr [r8+0x08]mov qword ptr [rsp+0x40],rdxmovrdx, qword ptr [r9+0x08]mov qword ptr [rsp+0x38],rdxmovrcx, qword ptr [rsp+0x40]movrdx, qword ptr [rsp+0x38]call[System.Runtime.Intrinsics.Scalar`1[long]:Multiply(long,long):long] ;;; not inlined call!mov qword ptr [rsp+0x48],raxmovrax, qword ptr [rsp+0x48]mov qword ptr [rsp+0x20],rsimov qword ptr [rsp+0x28],rax vmovaps xmm0, xmmword ptr [rsp+0x20]vmovups xmmword ptr [rbx],xmm0movrax,rbxaddrsp,104poprbxpoprsiret; Total bytes of code: 120

New codegen:

; Method MyBench:Foovmovupsxmm0, xmmword ptr [r8]vmovupsxmm1, xmmword ptr [r9] vpmuludq xmm2,xmm1,xmm0 vpshufd xmm1,xmm1,-79 vpmulld xmm0,xmm1,xmm0 vxorps xmm1,xmm1,xmm1vphadddxmm0,xmm0,xmm1 vpshufd xmm0,xmm0,115vpaddqxmm0,xmm0,xmm2vmovups xmmword ptr [rdx],xmm0movrax,rdxret; Total bytes of code: 50

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-runtime-intrinsics
See info in area-owners.md if you want to be subscribed.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

Note: results should be better if we do it in JIT, it will enable loop hoisting, cse, etc for MUL

@neon-sunset

Copy link
Copy Markdown
Contributor

Note #103539 (comment) (and https://godbolt.org/z/eqsrf341M) from xxHash128 issue.

EgorBoand others added 2 commits June 17, 2024 17:01
…sics/Vector128_1.cs
Co-authored-by: Tanner Gooding <tagoo@outlook.com>
@dotnetdotnet deleted a comment from EgorBotJun 20, 2024
@dotnetdotnet deleted a comment from EgorBotJun 20, 2024
@EgorBo

Copy link
Copy Markdown
MemberAuthor

@EgorBot -amd -intel -arm64 -profiler --envvars DOTNET_PreferredVectorBitWidth:128

usingSystem.IO.Hashing;usingBenchmarkDotNet.Attributes;publicclassBench{staticreadonlybyte[]Data=newbyte[1000000];[Benchmark]publicbyte[]BenchXxHash128(){XxHash128hash=new();hash.Append(Data);returnhash.GetHashAndReset();}}

@EgorBot

Copy link
Copy Markdown
Benchmark results on Intel
BenchmarkDotNet v0.13.12, Ubuntu 22.04.4 LTS (Jammy Jellyfish)
Intel Xeon Platinum 8370C CPU 2.80GHz, 1 CPU, 16 logical and 8 physical cores
Job-ITXSAG : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX-512F+CD+BW+DQ+VL+VBMI
Job-XSORFZ : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX-512F+CD+BW+DQ+VL+VBMI
EnvironmentVariables=DOTNET_PreferredVectorBitWidth=128
MethodToolchainMeanErrorRatio
BenchXxHash128Main43.41 μs0.087 μs1.00
BenchXxHash128PR43.33 μs0.009 μs1.00

BDN_Artifacts.zip

Flame graphs: Main vs PR 🔥
Hot asm: Main vs PR
Hot functions: Main vs PR

For clean perf results, make sure you have just one [Benchmark] in your app.

@EgorBot

Copy link
Copy Markdown
Benchmark results on Amd
BenchmarkDotNet v0.13.12, Ubuntu 22.04.4 LTS (Jammy Jellyfish)
AMD EPYC 7763, 1 CPU, 16 logical and 8 physical cores
Job-SUBLYH : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX2
Job-OPUYDY : .NET 9.0.0 (42.42.42.42424), X64 RyuJIT AVX2
EnvironmentVariables=DOTNET_PreferredVectorBitWidth=128
MethodToolchainMeanErrorRatio
BenchXxHash128Main71.20 μs0.022 μs1.00
BenchXxHash128PR43.84 μs0.013 μs0.62

BDN_Artifacts.zip

Flame graphs: Main vs PR 🔥
Hot asm: Main vs PR
Hot functions: Main vs PR

For clean perf results, make sure you have just one [Benchmark] in your app.

@EgorBot

Copy link
Copy Markdown
Benchmark results on Arm64
BenchmarkDotNet v0.13.12, Ubuntu 22.04.4 LTS (Jammy Jellyfish)
Unknown processor
Job-EDPWDU : .NET 9.0.0 (42.42.42.42424), Arm64 RyuJIT AdvSIMD
Job-TIALUR : .NET 9.0.0 (42.42.42.42424), Arm64 RyuJIT AdvSIMD
EnvironmentVariables=DOTNET_PreferredVectorBitWidth=128
MethodToolchainMeanErrorRatio
BenchXxHash128Main116.9 μs0.11 μs1.00
BenchXxHash128PR116.8 μs0.07 μs1.00

BDN_Artifacts.zip

Flame graphs: Main vs PR 🔥
Hot asm: Main vs PR
Hot functions: Main vs PR

For clean perf results, make sure you have just one [Benchmark] in your app.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp list

@azure-pipelines

This comment was marked as resolved.

@EgorBo

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-coreclr jitstress-isas-x86

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@EgorBo

EgorBo commented Jun 21, 2024

Copy link
Copy Markdown
MemberAuthor

@tannergooding PTAL, I'll add arm64 separately, need to test different impls.
I've expanded it in importer similar to existing op_Multiply expansions

Benchmark improvement: #103555 (comment)

@EgorBo
EgorBo marked this pull request as ready for review June 24, 2024 14:26
Comment threadsrc/coreclr/jit/gentree.cpp Outdated
Comment on lines +21627 to +21631
// Vector256<int> tmp3 = Avx2.HorizontalAdd(tmp2.AsInt32(), Vector256<int>.Zero);
GenTreeHWIntrinsic* tmp3 =
gtNewSimdHWIntrinsicNode(type, tmp2, gtNewZeroConNode(type),
is256 ? NI_AVX2_HorizontalAdd : NI_SSSE3_HorizontalAdd,
CORINFO_TYPE_UINT, simdSize);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I know in other places we've started avoiding hadd in favor of shuffle+add, might be worth seeing if that's appropriate here too (low priority, non blocking)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I tried to benchmark different implementations for it and they all were equaly fast e.g. #99871 (comment)

if (TARGET_POINTER_SIZE == 4)
{
// TODO-XARCH-CQ: We should support long/ulong multiplication
// TODO-XARCH-CQ: 32bit support

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What's blocking 32-bit support? It doesn't look like we're using any _X64 intrinsics in the fallback logic?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure to be honest, that check was pre-existing, I only changed comment

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@EgorBo@neon-sunset@EgorBot@tannergooding