[mono] Basic SIMD support for System.Numerics.Vector2 on arm64 - #91659

Merged
matouskozak merged 17 commits into
dotnet:mainfrom
matouskozak:arm64-vector2-intrinsics
Sep 21, 2023
Merged

[mono] Basic SIMD support for System.Numerics.Vector2 on arm64#91659
matouskozak merged 17 commits into
dotnet:mainfrom
matouskozak:arm64-vector2-intrinsics

Conversation

@matouskozak

@matouskozakmatouskozak commented Sep 6, 2023

Copy link
Copy Markdown
Member

Re-created PR that adds basic SIMD support for System.Numerics.Vector2 on arm64. Equaling the current support for System.Numerics.Vector4. Rename vector2_methods table to vector_2_3_4_methods to better reflect its usage.

Current SIMD support for Vector2 with mini/llvm:

  • SN_ctor
  • SN_Abs
  • SN_Add
  • SN_Clamp
  • SN_Divide (currently disabled Vector2 / float scenario, will enable in the next PR)
  • SN_Dot
  • SN_Max
  • SN_Min
  • SN_Multiply (same as with SN_Divide)
  • SN_Negate
  • SN_SquareRoot
  • SN_Subtract
  • SN_get_Item
  • SN_get_One
  • SN_get_UnitX
  • SN_get_UnitY
  • SN_get_Zero
  • SN_op_Addition
  • SN_op_Division
  • SN_op_Equality
  • SN_op_Inequality
  • SN_op_Multiply
  • SN_op_Subtraction
  • SN_op_UnaryNegation
  • SN_set_Item

Future work on the missing intrinsic is tracked here #91394.
Contributes to: #73462


p.s. These getters currently use 128-bit code paths for emitting const values (emit_xconst_v128) even for Vector2 (64-bit vector):

  • SN_get_UnitX
  • SN_get_UnitY
  • SN_get_One

Comment from @jandupej on the original PR:
You can use a fmov to flood the lower two floats with 1.0f. This gives you the fastest SN_get_One possible (there is a 64-bit variant of this, with q=0). To make SN_get_UnitX/Y you can shift the vector left or right as doubles by 32. Zeros are shifted in, so this will give you a (0.0f, 1.0f) or reverse. This will destroy the upper 64 bits of the register, but it shouldn't be a problem as only the lower 64 bits are of importance.

@matouskozak

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

Comment threadsrc/mono/mono/mini/mini-arm64.c Outdated
Comment threadsrc/mono/mono/mini/mini-arm64.c Outdated
@matouskozak

Copy link
Copy Markdown
MemberAuthor

Perf_Vector2 microbenchmarks on osx arm64 JIT-mini:

beforeafterspeed-up
CreateFromScalar1.951.2536%
OneBenchmark1.960.8855%
UnitXBenchmark1.920.8655%
UnitYBenchmark1.910.8954%
ZeroBenchmark1.193.60-202%
AddOperatorBenchmark4.030.8080%
DivideByVector2OperatorBenchmark4.260.8580%
DivideByScalarOperatorBenchmark6.1242.321862%
EqualityOperatorBenchmark1.4110.035997%
InequalityOperatorBenchmark2.2020.022199%
MultiplyOperatorBenchmark4.0810.725882%
MultiplyByScalarOperatorBenchmark5.8291.886768%
SubtractOperatorBenchmark4.0580.713782%
NegateOperatorBenchmark4.1430.739382%
AbsBenchmark19.9890.632897%
AddFunctionBenchmark4.0550.613585%
ClampBenchmark15.8590.815795%
DivideByVector2Benchmark4.2730.772182%
DivideByScalarBenchmark6.4532.423162%
DotBenchmark2.0850.141493%
MaxBenchmark6.3540.717289%
MinBenchmark6.2140.681189%
MultiplyFunctionBenchmark4.2610.831180%
NegateBenchmark4.2870.824781%
SquareRootBenchmark10.3550.705793%
SubtractFunctionBenchmark4.2390.717483%

Vector2.Zero is reporting 202% regression even though the emitted code looks correct:

0000000000000000 stp x29, x30, [sp, #-0x50]!
0000000000000004 mov x29, sp
0000000000000008 eor.8b v0, v0, v0
000000000000000c str d0, [x29, #0x10]
0000000000000010 ldr s0, [x29, #0x10]
0000000000000014 ldr s1, [x29, #0x14]
0000000000000018 mov sp, x29
000000000000001c ldp x29, x30, [sp], #0x50
0000000000000020 ret

@jandupej

Copy link
Copy Markdown
Contributor

Perf_Vector2 microbenchmarks on osx arm64 JIT-mini:

These are some impressive speedups, nice!

As for Vector2.Zero, what you did is correct. However our register allocator likes to spill and reload every value you create, especially in FP/SIMD (see the instructions at 0x0c and 0x10). If you load a constant instead, maybe it will forgo spilling and only load from memory (?). Still, I'd keep what you did. If there are future improvements to constant folding or the reg allocator, this will likely go away.

@matouskozak

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@SamMonoRT

Copy link
Copy Markdown
Member

@LoopedBard3 - if the aot-llvm arm64 local testing script ready, please add a link to the documentation and @matouskozak you should try to get numbers for aot-llvm arm64 also if possible via that script.

@matouskozak

Copy link
Copy Markdown
MemberAuthor

The test failures are tracked/unrelated to this PR.

const int t = get_type_size_macro (ins->inst_c1);
arm_neon_fdup_e (code, VREG_FULL, t, dreg, sreg1, 0);
if (ins->opcode == OP_EXPAND_R8)
arm_neon_fdup_e (code, VREG_FULL, t, dreg, sreg1, 0);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

OP_EXPAND_R8 can be simplified to a mov dreg, sreg1 or nothing if dreg == sreg1.

Comment threadsrc/mono/mono/mini/mini-arm64.c

@fanyang-monofanyang-mono left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

@matouskozak
matouskozak merged commit 09e796a into dotnet:mainSep 21, 2023
@ghostghost locked as resolved and limited conversation to collaborators Oct 21, 2023
@matouskozak
matouskozak deleted the arm64-vector2-intrinsics branch October 3, 2024 13:15
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@matouskozak@jandupej@SamMonoRT@fanyang-mono
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

[mono] Basic SIMD support for System.Numerics.Vector2 on arm64 - #91659

Merged
matouskozak merged 17 commits into
dotnet:mainfrom
matouskozak:arm64-vector2-intrinsics
Sep 21, 2023
Merged

[mono] Basic SIMD support for System.Numerics.Vector2 on arm64#91659
matouskozak merged 17 commits into
dotnet:mainfrom
matouskozak:arm64-vector2-intrinsics

Conversation

@matouskozak

@matouskozakmatouskozak commented Sep 6, 2023

Copy link
Copy Markdown
Member

Re-created PR that adds basic SIMD support for System.Numerics.Vector2 on arm64. Equaling the current support for System.Numerics.Vector4. Rename vector2_methods table to vector_2_3_4_methods to better reflect its usage.

Current SIMD support for Vector2 with mini/llvm:

  • SN_ctor
  • SN_Abs
  • SN_Add
  • SN_Clamp
  • SN_Divide (currently disabled Vector2 / float scenario, will enable in the next PR)
  • SN_Dot
  • SN_Max
  • SN_Min
  • SN_Multiply (same as with SN_Divide)
  • SN_Negate
  • SN_SquareRoot
  • SN_Subtract
  • SN_get_Item
  • SN_get_One
  • SN_get_UnitX
  • SN_get_UnitY
  • SN_get_Zero
  • SN_op_Addition
  • SN_op_Division
  • SN_op_Equality
  • SN_op_Inequality
  • SN_op_Multiply
  • SN_op_Subtraction
  • SN_op_UnaryNegation
  • SN_set_Item

Future work on the missing intrinsic is tracked here #91394.
Contributes to: #73462


p.s. These getters currently use 128-bit code paths for emitting const values (emit_xconst_v128) even for Vector2 (64-bit vector):

  • SN_get_UnitX
  • SN_get_UnitY
  • SN_get_One

Comment from @jandupej on the original PR:
You can use a fmov to flood the lower two floats with 1.0f. This gives you the fastest SN_get_One possible (there is a 64-bit variant of this, with q=0). To make SN_get_UnitX/Y you can shift the vector left or right as doubles by 32. Zeros are shifted in, so this will give you a (0.0f, 1.0f) or reverse. This will destroy the upper 64 bits of the register, but it shouldn't be a problem as only the lower 64 bits are of importance.

@matouskozak

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

Comment threadsrc/mono/mono/mini/mini-arm64.c Outdated
Comment threadsrc/mono/mono/mini/mini-arm64.c Outdated
@matouskozak

Copy link
Copy Markdown
MemberAuthor

Perf_Vector2 microbenchmarks on osx arm64 JIT-mini:

beforeafterspeed-up
CreateFromScalar1.951.2536%
OneBenchmark1.960.8855%
UnitXBenchmark1.920.8655%
UnitYBenchmark1.910.8954%
ZeroBenchmark1.193.60-202%
AddOperatorBenchmark4.030.8080%
DivideByVector2OperatorBenchmark4.260.8580%
DivideByScalarOperatorBenchmark6.1242.321862%
EqualityOperatorBenchmark1.4110.035997%
InequalityOperatorBenchmark2.2020.022199%
MultiplyOperatorBenchmark4.0810.725882%
MultiplyByScalarOperatorBenchmark5.8291.886768%
SubtractOperatorBenchmark4.0580.713782%
NegateOperatorBenchmark4.1430.739382%
AbsBenchmark19.9890.632897%
AddFunctionBenchmark4.0550.613585%
ClampBenchmark15.8590.815795%
DivideByVector2Benchmark4.2730.772182%
DivideByScalarBenchmark6.4532.423162%
DotBenchmark2.0850.141493%
MaxBenchmark6.3540.717289%
MinBenchmark6.2140.681189%
MultiplyFunctionBenchmark4.2610.831180%
NegateBenchmark4.2870.824781%
SquareRootBenchmark10.3550.705793%
SubtractFunctionBenchmark4.2390.717483%

Vector2.Zero is reporting 202% regression even though the emitted code looks correct:

0000000000000000 stp x29, x30, [sp, #-0x50]!
0000000000000004 mov x29, sp
0000000000000008 eor.8b v0, v0, v0
000000000000000c str d0, [x29, #0x10]
0000000000000010 ldr s0, [x29, #0x10]
0000000000000014 ldr s1, [x29, #0x14]
0000000000000018 mov sp, x29
000000000000001c ldp x29, x30, [sp], #0x50
0000000000000020 ret

@jandupej

Copy link
Copy Markdown
Contributor

Perf_Vector2 microbenchmarks on osx arm64 JIT-mini:

These are some impressive speedups, nice!

As for Vector2.Zero, what you did is correct. However our register allocator likes to spill and reload every value you create, especially in FP/SIMD (see the instructions at 0x0c and 0x10). If you load a constant instead, maybe it will forgo spilling and only load from memory (?). Still, I'd keep what you did. If there are future improvements to constant folding or the reg allocator, this will likely go away.

@matouskozak

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@SamMonoRT

Copy link
Copy Markdown
Member

@LoopedBard3 - if the aot-llvm arm64 local testing script ready, please add a link to the documentation and @matouskozak you should try to get numbers for aot-llvm arm64 also if possible via that script.

@matouskozak

Copy link
Copy Markdown
MemberAuthor

The test failures are tracked/unrelated to this PR.

const int t = get_type_size_macro (ins->inst_c1);
arm_neon_fdup_e (code, VREG_FULL, t, dreg, sreg1, 0);
if (ins->opcode == OP_EXPAND_R8)
arm_neon_fdup_e (code, VREG_FULL, t, dreg, sreg1, 0);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

OP_EXPAND_R8 can be simplified to a mov dreg, sreg1 or nothing if dreg == sreg1.

Comment threadsrc/mono/mono/mini/mini-arm64.c

@fanyang-monofanyang-mono left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

@matouskozak
matouskozak merged commit 09e796a into dotnet:mainSep 21, 2023
@ghostghost locked as resolved and limited conversation to collaborators Oct 21, 2023
@matouskozak
matouskozak deleted the arm64-vector2-intrinsics branch October 3, 2024 13:15
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@matouskozak@jandupej@SamMonoRT@fanyang-mono
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[mono] Basic SIMD support for System.Numerics.Vector2 on arm64 - #91659

Merged
matouskozak merged 17 commits into
dotnet:mainfrom
matouskozak:arm64-vector2-intrinsics
Sep 21, 2023
Merged

[mono] Basic SIMD support for System.Numerics.Vector2 on arm64#91659
matouskozak merged 17 commits into
dotnet:mainfrom
matouskozak:arm64-vector2-intrinsics

Conversation

@matouskozak

@matouskozakmatouskozak commented Sep 6, 2023

Copy link
Copy Markdown
Member

Re-created PR that adds basic SIMD support for System.Numerics.Vector2 on arm64. Equaling the current support for System.Numerics.Vector4. Rename vector2_methods table to vector_2_3_4_methods to better reflect its usage.

Current SIMD support for Vector2 with mini/llvm:

  • SN_ctor
  • SN_Abs
  • SN_Add
  • SN_Clamp
  • SN_Divide (currently disabled Vector2 / float scenario, will enable in the next PR)
  • SN_Dot
  • SN_Max
  • SN_Min
  • SN_Multiply (same as with SN_Divide)
  • SN_Negate
  • SN_SquareRoot
  • SN_Subtract
  • SN_get_Item
  • SN_get_One
  • SN_get_UnitX
  • SN_get_UnitY
  • SN_get_Zero
  • SN_op_Addition
  • SN_op_Division
  • SN_op_Equality
  • SN_op_Inequality
  • SN_op_Multiply
  • SN_op_Subtraction
  • SN_op_UnaryNegation
  • SN_set_Item

Future work on the missing intrinsic is tracked here #91394.
Contributes to: #73462


p.s. These getters currently use 128-bit code paths for emitting const values (emit_xconst_v128) even for Vector2 (64-bit vector):

  • SN_get_UnitX
  • SN_get_UnitY
  • SN_get_One

Comment from @jandupej on the original PR:
You can use a fmov to flood the lower two floats with 1.0f. This gives you the fastest SN_get_One possible (there is a 64-bit variant of this, with q=0). To make SN_get_UnitX/Y you can shift the vector left or right as doubles by 32. Zeros are shifted in, so this will give you a (0.0f, 1.0f) or reverse. This will destroy the upper 64 bits of the register, but it shouldn't be a problem as only the lower 64 bits are of importance.

@matouskozak

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

Comment threadsrc/mono/mono/mini/mini-arm64.c Outdated
Comment threadsrc/mono/mono/mini/mini-arm64.c Outdated
@matouskozak

Copy link
Copy Markdown
MemberAuthor

Perf_Vector2 microbenchmarks on osx arm64 JIT-mini:

beforeafterspeed-up
CreateFromScalar1.951.2536%
OneBenchmark1.960.8855%
UnitXBenchmark1.920.8655%
UnitYBenchmark1.910.8954%
ZeroBenchmark1.193.60-202%
AddOperatorBenchmark4.030.8080%
DivideByVector2OperatorBenchmark4.260.8580%
DivideByScalarOperatorBenchmark6.1242.321862%
EqualityOperatorBenchmark1.4110.035997%
InequalityOperatorBenchmark2.2020.022199%
MultiplyOperatorBenchmark4.0810.725882%
MultiplyByScalarOperatorBenchmark5.8291.886768%
SubtractOperatorBenchmark4.0580.713782%
NegateOperatorBenchmark4.1430.739382%
AbsBenchmark19.9890.632897%
AddFunctionBenchmark4.0550.613585%
ClampBenchmark15.8590.815795%
DivideByVector2Benchmark4.2730.772182%
DivideByScalarBenchmark6.4532.423162%
DotBenchmark2.0850.141493%
MaxBenchmark6.3540.717289%
MinBenchmark6.2140.681189%
MultiplyFunctionBenchmark4.2610.831180%
NegateBenchmark4.2870.824781%
SquareRootBenchmark10.3550.705793%
SubtractFunctionBenchmark4.2390.717483%

Vector2.Zero is reporting 202% regression even though the emitted code looks correct:

0000000000000000 stp x29, x30, [sp, #-0x50]!
0000000000000004 mov x29, sp
0000000000000008 eor.8b v0, v0, v0
000000000000000c str d0, [x29, #0x10]
0000000000000010 ldr s0, [x29, #0x10]
0000000000000014 ldr s1, [x29, #0x14]
0000000000000018 mov sp, x29
000000000000001c ldp x29, x30, [sp], #0x50
0000000000000020 ret

@jandupej

Copy link
Copy Markdown
Contributor

Perf_Vector2 microbenchmarks on osx arm64 JIT-mini:

These are some impressive speedups, nice!

As for Vector2.Zero, what you did is correct. However our register allocator likes to spill and reload every value you create, especially in FP/SIMD (see the instructions at 0x0c and 0x10). If you load a constant instead, maybe it will forgo spilling and only load from memory (?). Still, I'd keep what you did. If there are future improvements to constant folding or the reg allocator, this will likely go away.

@matouskozak

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@SamMonoRT

Copy link
Copy Markdown
Member

@LoopedBard3 - if the aot-llvm arm64 local testing script ready, please add a link to the documentation and @matouskozak you should try to get numbers for aot-llvm arm64 also if possible via that script.

@matouskozak

Copy link
Copy Markdown
MemberAuthor

The test failures are tracked/unrelated to this PR.

const int t = get_type_size_macro (ins->inst_c1);
arm_neon_fdup_e (code, VREG_FULL, t, dreg, sreg1, 0);
if (ins->opcode == OP_EXPAND_R8)
arm_neon_fdup_e (code, VREG_FULL, t, dreg, sreg1, 0);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

OP_EXPAND_R8 can be simplified to a mov dreg, sreg1 or nothing if dreg == sreg1.

Comment threadsrc/mono/mono/mini/mini-arm64.c

@fanyang-monofanyang-mono left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

@matouskozak
matouskozak merged commit 09e796a into dotnet:mainSep 21, 2023
@ghostghost locked as resolved and limited conversation to collaborators Oct 21, 2023
@matouskozak
matouskozak deleted the arm64-vector2-intrinsics branch October 3, 2024 13:15
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@matouskozak@jandupej@SamMonoRT@fanyang-mono
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[mono] Basic SIMD support for System.Numerics.Vector2 on arm64 - #91659

Merged
matouskozak merged 17 commits into
dotnet:mainfrom
matouskozak:arm64-vector2-intrinsics
Sep 21, 2023
Merged

[mono] Basic SIMD support for System.Numerics.Vector2 on arm64#91659
matouskozak merged 17 commits into
dotnet:mainfrom
matouskozak:arm64-vector2-intrinsics

Conversation

@matouskozak

@matouskozakmatouskozak commented Sep 6, 2023

Copy link
Copy Markdown
Member

Re-created PR that adds basic SIMD support for System.Numerics.Vector2 on arm64. Equaling the current support for System.Numerics.Vector4. Rename vector2_methods table to vector_2_3_4_methods to better reflect its usage.

Current SIMD support for Vector2 with mini/llvm:

  • SN_ctor
  • SN_Abs
  • SN_Add
  • SN_Clamp
  • SN_Divide (currently disabled Vector2 / float scenario, will enable in the next PR)
  • SN_Dot
  • SN_Max
  • SN_Min
  • SN_Multiply (same as with SN_Divide)
  • SN_Negate
  • SN_SquareRoot
  • SN_Subtract
  • SN_get_Item
  • SN_get_One
  • SN_get_UnitX
  • SN_get_UnitY
  • SN_get_Zero
  • SN_op_Addition
  • SN_op_Division
  • SN_op_Equality
  • SN_op_Inequality
  • SN_op_Multiply
  • SN_op_Subtraction
  • SN_op_UnaryNegation
  • SN_set_Item

Future work on the missing intrinsic is tracked here #91394.
Contributes to: #73462


p.s. These getters currently use 128-bit code paths for emitting const values (emit_xconst_v128) even for Vector2 (64-bit vector):

  • SN_get_UnitX
  • SN_get_UnitY
  • SN_get_One

Comment from @jandupej on the original PR:
You can use a fmov to flood the lower two floats with 1.0f. This gives you the fastest SN_get_One possible (there is a 64-bit variant of this, with q=0). To make SN_get_UnitX/Y you can shift the vector left or right as doubles by 32. Zeros are shifted in, so this will give you a (0.0f, 1.0f) or reverse. This will destroy the upper 64 bits of the register, but it shouldn't be a problem as only the lower 64 bits are of importance.

@matouskozak

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

Comment threadsrc/mono/mono/mini/mini-arm64.c Outdated
Comment threadsrc/mono/mono/mini/mini-arm64.c Outdated
@matouskozak

Copy link
Copy Markdown
MemberAuthor

Perf_Vector2 microbenchmarks on osx arm64 JIT-mini:

beforeafterspeed-up
CreateFromScalar1.951.2536%
OneBenchmark1.960.8855%
UnitXBenchmark1.920.8655%
UnitYBenchmark1.910.8954%
ZeroBenchmark1.193.60-202%
AddOperatorBenchmark4.030.8080%
DivideByVector2OperatorBenchmark4.260.8580%
DivideByScalarOperatorBenchmark6.1242.321862%
EqualityOperatorBenchmark1.4110.035997%
InequalityOperatorBenchmark2.2020.022199%
MultiplyOperatorBenchmark4.0810.725882%
MultiplyByScalarOperatorBenchmark5.8291.886768%
SubtractOperatorBenchmark4.0580.713782%
NegateOperatorBenchmark4.1430.739382%
AbsBenchmark19.9890.632897%
AddFunctionBenchmark4.0550.613585%
ClampBenchmark15.8590.815795%
DivideByVector2Benchmark4.2730.772182%
DivideByScalarBenchmark6.4532.423162%
DotBenchmark2.0850.141493%
MaxBenchmark6.3540.717289%
MinBenchmark6.2140.681189%
MultiplyFunctionBenchmark4.2610.831180%
NegateBenchmark4.2870.824781%
SquareRootBenchmark10.3550.705793%
SubtractFunctionBenchmark4.2390.717483%

Vector2.Zero is reporting 202% regression even though the emitted code looks correct:

0000000000000000 stp x29, x30, [sp, #-0x50]!
0000000000000004 mov x29, sp
0000000000000008 eor.8b v0, v0, v0
000000000000000c str d0, [x29, #0x10]
0000000000000010 ldr s0, [x29, #0x10]
0000000000000014 ldr s1, [x29, #0x14]
0000000000000018 mov sp, x29
000000000000001c ldp x29, x30, [sp], #0x50
0000000000000020 ret

@jandupej

Copy link
Copy Markdown
Contributor

Perf_Vector2 microbenchmarks on osx arm64 JIT-mini:

These are some impressive speedups, nice!

As for Vector2.Zero, what you did is correct. However our register allocator likes to spill and reload every value you create, especially in FP/SIMD (see the instructions at 0x0c and 0x10). If you load a constant instead, maybe it will forgo spilling and only load from memory (?). Still, I'd keep what you did. If there are future improvements to constant folding or the reg allocator, this will likely go away.

@matouskozak

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@SamMonoRT

Copy link
Copy Markdown
Member

@LoopedBard3 - if the aot-llvm arm64 local testing script ready, please add a link to the documentation and @matouskozak you should try to get numbers for aot-llvm arm64 also if possible via that script.

@matouskozak

Copy link
Copy Markdown
MemberAuthor

The test failures are tracked/unrelated to this PR.

const int t = get_type_size_macro (ins->inst_c1);
arm_neon_fdup_e (code, VREG_FULL, t, dreg, sreg1, 0);
if (ins->opcode == OP_EXPAND_R8)
arm_neon_fdup_e (code, VREG_FULL, t, dreg, sreg1, 0);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

OP_EXPAND_R8 can be simplified to a mov dreg, sreg1 or nothing if dreg == sreg1.

Comment threadsrc/mono/mono/mini/mini-arm64.c

@fanyang-monofanyang-mono left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

@matouskozak
matouskozak merged commit 09e796a into dotnet:mainSep 21, 2023
@ghostghost locked as resolved and limited conversation to collaborators Oct 21, 2023
@matouskozak
matouskozak deleted the arm64-vector2-intrinsics branch October 3, 2024 13:15
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@matouskozak@jandupej@SamMonoRT@fanyang-mono
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

[mono] Basic SIMD support for System.Numerics.Vector2 on arm64 - #91659

Merged
matouskozak merged 17 commits into
dotnet:mainfrom
matouskozak:arm64-vector2-intrinsics
Sep 21, 2023
Merged

[mono] Basic SIMD support for System.Numerics.Vector2 on arm64#91659
matouskozak merged 17 commits into
dotnet:mainfrom
matouskozak:arm64-vector2-intrinsics

Conversation

@matouskozak

@matouskozakmatouskozak commented Sep 6, 2023

Copy link
Copy Markdown
Member

Re-created PR that adds basic SIMD support for System.Numerics.Vector2 on arm64. Equaling the current support for System.Numerics.Vector4. Rename vector2_methods table to vector_2_3_4_methods to better reflect its usage.

Current SIMD support for Vector2 with mini/llvm:

  • SN_ctor
  • SN_Abs
  • SN_Add
  • SN_Clamp
  • SN_Divide (currently disabled Vector2 / float scenario, will enable in the next PR)
  • SN_Dot
  • SN_Max
  • SN_Min
  • SN_Multiply (same as with SN_Divide)
  • SN_Negate
  • SN_SquareRoot
  • SN_Subtract
  • SN_get_Item
  • SN_get_One
  • SN_get_UnitX
  • SN_get_UnitY
  • SN_get_Zero
  • SN_op_Addition
  • SN_op_Division
  • SN_op_Equality
  • SN_op_Inequality
  • SN_op_Multiply
  • SN_op_Subtraction
  • SN_op_UnaryNegation
  • SN_set_Item

Future work on the missing intrinsic is tracked here #91394.
Contributes to: #73462


p.s. These getters currently use 128-bit code paths for emitting const values (emit_xconst_v128) even for Vector2 (64-bit vector):

  • SN_get_UnitX
  • SN_get_UnitY
  • SN_get_One

Comment from @jandupej on the original PR:
You can use a fmov to flood the lower two floats with 1.0f. This gives you the fastest SN_get_One possible (there is a 64-bit variant of this, with q=0). To make SN_get_UnitX/Y you can shift the vector left or right as doubles by 32. Zeros are shifted in, so this will give you a (0.0f, 1.0f) or reverse. This will destroy the upper 64 bits of the register, but it shouldn't be a problem as only the lower 64 bits are of importance.

@matouskozak

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

Comment threadsrc/mono/mono/mini/mini-arm64.c Outdated
Comment threadsrc/mono/mono/mini/mini-arm64.c Outdated
@matouskozak

Copy link
Copy Markdown
MemberAuthor

Perf_Vector2 microbenchmarks on osx arm64 JIT-mini:

beforeafterspeed-up
CreateFromScalar1.951.2536%
OneBenchmark1.960.8855%
UnitXBenchmark1.920.8655%
UnitYBenchmark1.910.8954%
ZeroBenchmark1.193.60-202%
AddOperatorBenchmark4.030.8080%
DivideByVector2OperatorBenchmark4.260.8580%
DivideByScalarOperatorBenchmark6.1242.321862%
EqualityOperatorBenchmark1.4110.035997%
InequalityOperatorBenchmark2.2020.022199%
MultiplyOperatorBenchmark4.0810.725882%
MultiplyByScalarOperatorBenchmark5.8291.886768%
SubtractOperatorBenchmark4.0580.713782%
NegateOperatorBenchmark4.1430.739382%
AbsBenchmark19.9890.632897%
AddFunctionBenchmark4.0550.613585%
ClampBenchmark15.8590.815795%
DivideByVector2Benchmark4.2730.772182%
DivideByScalarBenchmark6.4532.423162%
DotBenchmark2.0850.141493%
MaxBenchmark6.3540.717289%
MinBenchmark6.2140.681189%
MultiplyFunctionBenchmark4.2610.831180%
NegateBenchmark4.2870.824781%
SquareRootBenchmark10.3550.705793%
SubtractFunctionBenchmark4.2390.717483%

Vector2.Zero is reporting 202% regression even though the emitted code looks correct:

0000000000000000 stp x29, x30, [sp, #-0x50]!
0000000000000004 mov x29, sp
0000000000000008 eor.8b v0, v0, v0
000000000000000c str d0, [x29, #0x10]
0000000000000010 ldr s0, [x29, #0x10]
0000000000000014 ldr s1, [x29, #0x14]
0000000000000018 mov sp, x29
000000000000001c ldp x29, x30, [sp], #0x50
0000000000000020 ret

@jandupej

Copy link
Copy Markdown
Contributor

Perf_Vector2 microbenchmarks on osx arm64 JIT-mini:

These are some impressive speedups, nice!

As for Vector2.Zero, what you did is correct. However our register allocator likes to spill and reload every value you create, especially in FP/SIMD (see the instructions at 0x0c and 0x10). If you load a constant instead, maybe it will forgo spilling and only load from memory (?). Still, I'd keep what you did. If there are future improvements to constant folding or the reg allocator, this will likely go away.

@matouskozak

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@SamMonoRT

Copy link
Copy Markdown
Member

@LoopedBard3 - if the aot-llvm arm64 local testing script ready, please add a link to the documentation and @matouskozak you should try to get numbers for aot-llvm arm64 also if possible via that script.

@matouskozak

Copy link
Copy Markdown
MemberAuthor

The test failures are tracked/unrelated to this PR.

const int t = get_type_size_macro (ins->inst_c1);
arm_neon_fdup_e (code, VREG_FULL, t, dreg, sreg1, 0);
if (ins->opcode == OP_EXPAND_R8)
arm_neon_fdup_e (code, VREG_FULL, t, dreg, sreg1, 0);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

OP_EXPAND_R8 can be simplified to a mov dreg, sreg1 or nothing if dreg == sreg1.

Comment threadsrc/mono/mono/mini/mini-arm64.c

@fanyang-monofanyang-mono left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

@matouskozak
matouskozak merged commit 09e796a into dotnet:mainSep 21, 2023
@ghostghost locked as resolved and limited conversation to collaborators Oct 21, 2023
@matouskozak
matouskozak deleted the arm64-vector2-intrinsics branch October 3, 2024 13:15
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@matouskozak@jandupej@SamMonoRT@fanyang-mono
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[mono] Basic SIMD support for System.Numerics.Vector2 on arm64 - #91659

Merged
matouskozak merged 17 commits into
dotnet:mainfrom
matouskozak:arm64-vector2-intrinsics
Sep 21, 2023
Merged

[mono] Basic SIMD support for System.Numerics.Vector2 on arm64#91659
matouskozak merged 17 commits into
dotnet:mainfrom
matouskozak:arm64-vector2-intrinsics

Conversation

@matouskozak

@matouskozakmatouskozak commented Sep 6, 2023

Copy link
Copy Markdown
Member

Re-created PR that adds basic SIMD support for System.Numerics.Vector2 on arm64. Equaling the current support for System.Numerics.Vector4. Rename vector2_methods table to vector_2_3_4_methods to better reflect its usage.

Current SIMD support for Vector2 with mini/llvm:

  • SN_ctor
  • SN_Abs
  • SN_Add
  • SN_Clamp
  • SN_Divide (currently disabled Vector2 / float scenario, will enable in the next PR)
  • SN_Dot
  • SN_Max
  • SN_Min
  • SN_Multiply (same as with SN_Divide)
  • SN_Negate
  • SN_SquareRoot
  • SN_Subtract
  • SN_get_Item
  • SN_get_One
  • SN_get_UnitX
  • SN_get_UnitY
  • SN_get_Zero
  • SN_op_Addition
  • SN_op_Division
  • SN_op_Equality
  • SN_op_Inequality
  • SN_op_Multiply
  • SN_op_Subtraction
  • SN_op_UnaryNegation
  • SN_set_Item

Future work on the missing intrinsic is tracked here #91394.
Contributes to: #73462


p.s. These getters currently use 128-bit code paths for emitting const values (emit_xconst_v128) even for Vector2 (64-bit vector):

  • SN_get_UnitX
  • SN_get_UnitY
  • SN_get_One

Comment from @jandupej on the original PR:
You can use a fmov to flood the lower two floats with 1.0f. This gives you the fastest SN_get_One possible (there is a 64-bit variant of this, with q=0). To make SN_get_UnitX/Y you can shift the vector left or right as doubles by 32. Zeros are shifted in, so this will give you a (0.0f, 1.0f) or reverse. This will destroy the upper 64 bits of the register, but it shouldn't be a problem as only the lower 64 bits are of importance.

@matouskozak

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

Comment threadsrc/mono/mono/mini/mini-arm64.c Outdated
Comment threadsrc/mono/mono/mini/mini-arm64.c Outdated
@matouskozak

Copy link
Copy Markdown
MemberAuthor

Perf_Vector2 microbenchmarks on osx arm64 JIT-mini:

beforeafterspeed-up
CreateFromScalar1.951.2536%
OneBenchmark1.960.8855%
UnitXBenchmark1.920.8655%
UnitYBenchmark1.910.8954%
ZeroBenchmark1.193.60-202%
AddOperatorBenchmark4.030.8080%
DivideByVector2OperatorBenchmark4.260.8580%
DivideByScalarOperatorBenchmark6.1242.321862%
EqualityOperatorBenchmark1.4110.035997%
InequalityOperatorBenchmark2.2020.022199%
MultiplyOperatorBenchmark4.0810.725882%
MultiplyByScalarOperatorBenchmark5.8291.886768%
SubtractOperatorBenchmark4.0580.713782%
NegateOperatorBenchmark4.1430.739382%
AbsBenchmark19.9890.632897%
AddFunctionBenchmark4.0550.613585%
ClampBenchmark15.8590.815795%
DivideByVector2Benchmark4.2730.772182%
DivideByScalarBenchmark6.4532.423162%
DotBenchmark2.0850.141493%
MaxBenchmark6.3540.717289%
MinBenchmark6.2140.681189%
MultiplyFunctionBenchmark4.2610.831180%
NegateBenchmark4.2870.824781%
SquareRootBenchmark10.3550.705793%
SubtractFunctionBenchmark4.2390.717483%

Vector2.Zero is reporting 202% regression even though the emitted code looks correct:

0000000000000000 stp x29, x30, [sp, #-0x50]!
0000000000000004 mov x29, sp
0000000000000008 eor.8b v0, v0, v0
000000000000000c str d0, [x29, #0x10]
0000000000000010 ldr s0, [x29, #0x10]
0000000000000014 ldr s1, [x29, #0x14]
0000000000000018 mov sp, x29
000000000000001c ldp x29, x30, [sp], #0x50
0000000000000020 ret

@jandupej

Copy link
Copy Markdown
Contributor

Perf_Vector2 microbenchmarks on osx arm64 JIT-mini:

These are some impressive speedups, nice!

As for Vector2.Zero, what you did is correct. However our register allocator likes to spill and reload every value you create, especially in FP/SIMD (see the instructions at 0x0c and 0x10). If you load a constant instead, maybe it will forgo spilling and only load from memory (?). Still, I'd keep what you did. If there are future improvements to constant folding or the reg allocator, this will likely go away.

@matouskozak

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@SamMonoRT

Copy link
Copy Markdown
Member

@LoopedBard3 - if the aot-llvm arm64 local testing script ready, please add a link to the documentation and @matouskozak you should try to get numbers for aot-llvm arm64 also if possible via that script.

@matouskozak

Copy link
Copy Markdown
MemberAuthor

The test failures are tracked/unrelated to this PR.

const int t = get_type_size_macro (ins->inst_c1);
arm_neon_fdup_e (code, VREG_FULL, t, dreg, sreg1, 0);
if (ins->opcode == OP_EXPAND_R8)
arm_neon_fdup_e (code, VREG_FULL, t, dreg, sreg1, 0);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

OP_EXPAND_R8 can be simplified to a mov dreg, sreg1 or nothing if dreg == sreg1.

Comment threadsrc/mono/mono/mini/mini-arm64.c

@fanyang-monofanyang-mono left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

@matouskozak
matouskozak merged commit 09e796a into dotnet:mainSep 21, 2023
@ghostghost locked as resolved and limited conversation to collaborators Oct 21, 2023
@matouskozak
matouskozak deleted the arm64-vector2-intrinsics branch October 3, 2024 13:15
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@matouskozak@jandupej@SamMonoRT@fanyang-mono
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[mono] Basic SIMD support for System.Numerics.Vector2 on arm64 - #91659

Merged
matouskozak merged 17 commits into
dotnet:mainfrom
matouskozak:arm64-vector2-intrinsics
Sep 21, 2023
Merged

[mono] Basic SIMD support for System.Numerics.Vector2 on arm64#91659
matouskozak merged 17 commits into
dotnet:mainfrom
matouskozak:arm64-vector2-intrinsics

Conversation

@matouskozak

@matouskozakmatouskozak commented Sep 6, 2023

Copy link
Copy Markdown
Member

Re-created PR that adds basic SIMD support for System.Numerics.Vector2 on arm64. Equaling the current support for System.Numerics.Vector4. Rename vector2_methods table to vector_2_3_4_methods to better reflect its usage.

Current SIMD support for Vector2 with mini/llvm:

  • SN_ctor
  • SN_Abs
  • SN_Add
  • SN_Clamp
  • SN_Divide (currently disabled Vector2 / float scenario, will enable in the next PR)
  • SN_Dot
  • SN_Max
  • SN_Min
  • SN_Multiply (same as with SN_Divide)
  • SN_Negate
  • SN_SquareRoot
  • SN_Subtract
  • SN_get_Item
  • SN_get_One
  • SN_get_UnitX
  • SN_get_UnitY
  • SN_get_Zero
  • SN_op_Addition
  • SN_op_Division
  • SN_op_Equality
  • SN_op_Inequality
  • SN_op_Multiply
  • SN_op_Subtraction
  • SN_op_UnaryNegation
  • SN_set_Item

Future work on the missing intrinsic is tracked here #91394.
Contributes to: #73462


p.s. These getters currently use 128-bit code paths for emitting const values (emit_xconst_v128) even for Vector2 (64-bit vector):

  • SN_get_UnitX
  • SN_get_UnitY
  • SN_get_One

Comment from @jandupej on the original PR:
You can use a fmov to flood the lower two floats with 1.0f. This gives you the fastest SN_get_One possible (there is a 64-bit variant of this, with q=0). To make SN_get_UnitX/Y you can shift the vector left or right as doubles by 32. Zeros are shifted in, so this will give you a (0.0f, 1.0f) or reverse. This will destroy the upper 64 bits of the register, but it shouldn't be a problem as only the lower 64 bits are of importance.

@matouskozak

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

Comment threadsrc/mono/mono/mini/mini-arm64.c Outdated
Comment threadsrc/mono/mono/mini/mini-arm64.c Outdated
@matouskozak

Copy link
Copy Markdown
MemberAuthor

Perf_Vector2 microbenchmarks on osx arm64 JIT-mini:

beforeafterspeed-up
CreateFromScalar1.951.2536%
OneBenchmark1.960.8855%
UnitXBenchmark1.920.8655%
UnitYBenchmark1.910.8954%
ZeroBenchmark1.193.60-202%
AddOperatorBenchmark4.030.8080%
DivideByVector2OperatorBenchmark4.260.8580%
DivideByScalarOperatorBenchmark6.1242.321862%
EqualityOperatorBenchmark1.4110.035997%
InequalityOperatorBenchmark2.2020.022199%
MultiplyOperatorBenchmark4.0810.725882%
MultiplyByScalarOperatorBenchmark5.8291.886768%
SubtractOperatorBenchmark4.0580.713782%
NegateOperatorBenchmark4.1430.739382%
AbsBenchmark19.9890.632897%
AddFunctionBenchmark4.0550.613585%
ClampBenchmark15.8590.815795%
DivideByVector2Benchmark4.2730.772182%
DivideByScalarBenchmark6.4532.423162%
DotBenchmark2.0850.141493%
MaxBenchmark6.3540.717289%
MinBenchmark6.2140.681189%
MultiplyFunctionBenchmark4.2610.831180%
NegateBenchmark4.2870.824781%
SquareRootBenchmark10.3550.705793%
SubtractFunctionBenchmark4.2390.717483%

Vector2.Zero is reporting 202% regression even though the emitted code looks correct:

0000000000000000 stp x29, x30, [sp, #-0x50]!
0000000000000004 mov x29, sp
0000000000000008 eor.8b v0, v0, v0
000000000000000c str d0, [x29, #0x10]
0000000000000010 ldr s0, [x29, #0x10]
0000000000000014 ldr s1, [x29, #0x14]
0000000000000018 mov sp, x29
000000000000001c ldp x29, x30, [sp], #0x50
0000000000000020 ret

@jandupej

Copy link
Copy Markdown
Contributor

Perf_Vector2 microbenchmarks on osx arm64 JIT-mini:

These are some impressive speedups, nice!

As for Vector2.Zero, what you did is correct. However our register allocator likes to spill and reload every value you create, especially in FP/SIMD (see the instructions at 0x0c and 0x10). If you load a constant instead, maybe it will forgo spilling and only load from memory (?). Still, I'd keep what you did. If there are future improvements to constant folding or the reg allocator, this will likely go away.

@matouskozak

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@SamMonoRT

Copy link
Copy Markdown
Member

@LoopedBard3 - if the aot-llvm arm64 local testing script ready, please add a link to the documentation and @matouskozak you should try to get numbers for aot-llvm arm64 also if possible via that script.

@matouskozak

Copy link
Copy Markdown
MemberAuthor

The test failures are tracked/unrelated to this PR.

const int t = get_type_size_macro (ins->inst_c1);
arm_neon_fdup_e (code, VREG_FULL, t, dreg, sreg1, 0);
if (ins->opcode == OP_EXPAND_R8)
arm_neon_fdup_e (code, VREG_FULL, t, dreg, sreg1, 0);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

OP_EXPAND_R8 can be simplified to a mov dreg, sreg1 or nothing if dreg == sreg1.

Comment threadsrc/mono/mono/mini/mini-arm64.c

@fanyang-monofanyang-mono left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

@matouskozak
matouskozak merged commit 09e796a into dotnet:mainSep 21, 2023
@ghostghost locked as resolved and limited conversation to collaborators Oct 21, 2023
@matouskozak
matouskozak deleted the arm64-vector2-intrinsics branch October 3, 2024 13:15
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@matouskozak@jandupej@SamMonoRT@fanyang-mono
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

[mono] Basic SIMD support for System.Numerics.Vector2 on arm64 - #91659

Merged
matouskozak merged 17 commits into
dotnet:mainfrom
matouskozak:arm64-vector2-intrinsics
Sep 21, 2023
Merged

[mono] Basic SIMD support for System.Numerics.Vector2 on arm64#91659
matouskozak merged 17 commits into
dotnet:mainfrom
matouskozak:arm64-vector2-intrinsics

Conversation

@matouskozak

@matouskozakmatouskozak commented Sep 6, 2023

Copy link
Copy Markdown
Member

Re-created PR that adds basic SIMD support for System.Numerics.Vector2 on arm64. Equaling the current support for System.Numerics.Vector4. Rename vector2_methods table to vector_2_3_4_methods to better reflect its usage.

Current SIMD support for Vector2 with mini/llvm:

  • SN_ctor
  • SN_Abs
  • SN_Add
  • SN_Clamp
  • SN_Divide (currently disabled Vector2 / float scenario, will enable in the next PR)
  • SN_Dot
  • SN_Max
  • SN_Min
  • SN_Multiply (same as with SN_Divide)
  • SN_Negate
  • SN_SquareRoot
  • SN_Subtract
  • SN_get_Item
  • SN_get_One
  • SN_get_UnitX
  • SN_get_UnitY
  • SN_get_Zero
  • SN_op_Addition
  • SN_op_Division
  • SN_op_Equality
  • SN_op_Inequality
  • SN_op_Multiply
  • SN_op_Subtraction
  • SN_op_UnaryNegation
  • SN_set_Item

Future work on the missing intrinsic is tracked here #91394.
Contributes to: #73462


p.s. These getters currently use 128-bit code paths for emitting const values (emit_xconst_v128) even for Vector2 (64-bit vector):

  • SN_get_UnitX
  • SN_get_UnitY
  • SN_get_One

Comment from @jandupej on the original PR:
You can use a fmov to flood the lower two floats with 1.0f. This gives you the fastest SN_get_One possible (there is a 64-bit variant of this, with q=0). To make SN_get_UnitX/Y you can shift the vector left or right as doubles by 32. Zeros are shifted in, so this will give you a (0.0f, 1.0f) or reverse. This will destroy the upper 64 bits of the register, but it shouldn't be a problem as only the lower 64 bits are of importance.

@matouskozak

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

Comment threadsrc/mono/mono/mini/mini-arm64.c Outdated
Comment threadsrc/mono/mono/mini/mini-arm64.c Outdated
@matouskozak

Copy link
Copy Markdown
MemberAuthor

Perf_Vector2 microbenchmarks on osx arm64 JIT-mini:

beforeafterspeed-up
CreateFromScalar1.951.2536%
OneBenchmark1.960.8855%
UnitXBenchmark1.920.8655%
UnitYBenchmark1.910.8954%
ZeroBenchmark1.193.60-202%
AddOperatorBenchmark4.030.8080%
DivideByVector2OperatorBenchmark4.260.8580%
DivideByScalarOperatorBenchmark6.1242.321862%
EqualityOperatorBenchmark1.4110.035997%
InequalityOperatorBenchmark2.2020.022199%
MultiplyOperatorBenchmark4.0810.725882%
MultiplyByScalarOperatorBenchmark5.8291.886768%
SubtractOperatorBenchmark4.0580.713782%
NegateOperatorBenchmark4.1430.739382%
AbsBenchmark19.9890.632897%
AddFunctionBenchmark4.0550.613585%
ClampBenchmark15.8590.815795%
DivideByVector2Benchmark4.2730.772182%
DivideByScalarBenchmark6.4532.423162%
DotBenchmark2.0850.141493%
MaxBenchmark6.3540.717289%
MinBenchmark6.2140.681189%
MultiplyFunctionBenchmark4.2610.831180%
NegateBenchmark4.2870.824781%
SquareRootBenchmark10.3550.705793%
SubtractFunctionBenchmark4.2390.717483%

Vector2.Zero is reporting 202% regression even though the emitted code looks correct:

0000000000000000 stp x29, x30, [sp, #-0x50]!
0000000000000004 mov x29, sp
0000000000000008 eor.8b v0, v0, v0
000000000000000c str d0, [x29, #0x10]
0000000000000010 ldr s0, [x29, #0x10]
0000000000000014 ldr s1, [x29, #0x14]
0000000000000018 mov sp, x29
000000000000001c ldp x29, x30, [sp], #0x50
0000000000000020 ret

@jandupej

Copy link
Copy Markdown
Contributor

Perf_Vector2 microbenchmarks on osx arm64 JIT-mini:

These are some impressive speedups, nice!

As for Vector2.Zero, what you did is correct. However our register allocator likes to spill and reload every value you create, especially in FP/SIMD (see the instructions at 0x0c and 0x10). If you load a constant instead, maybe it will forgo spilling and only load from memory (?). Still, I'd keep what you did. If there are future improvements to constant folding or the reg allocator, this will likely go away.

@matouskozak

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@SamMonoRT

Copy link
Copy Markdown
Member

@LoopedBard3 - if the aot-llvm arm64 local testing script ready, please add a link to the documentation and @matouskozak you should try to get numbers for aot-llvm arm64 also if possible via that script.

@matouskozak

Copy link
Copy Markdown
MemberAuthor

The test failures are tracked/unrelated to this PR.

const int t = get_type_size_macro (ins->inst_c1);
arm_neon_fdup_e (code, VREG_FULL, t, dreg, sreg1, 0);
if (ins->opcode == OP_EXPAND_R8)
arm_neon_fdup_e (code, VREG_FULL, t, dreg, sreg1, 0);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

OP_EXPAND_R8 can be simplified to a mov dreg, sreg1 or nothing if dreg == sreg1.

Comment threadsrc/mono/mono/mini/mini-arm64.c

@fanyang-monofanyang-mono left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

@matouskozak
matouskozak merged commit 09e796a into dotnet:mainSep 21, 2023
@ghostghost locked as resolved and limited conversation to collaborators Oct 21, 2023
@matouskozak
matouskozak deleted the arm64-vector2-intrinsics branch October 3, 2024 13:15
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@matouskozak@jandupej@SamMonoRT@fanyang-mono