ARM64 intrinsic support for Vector64.Create() and Vector128.Create() - #35590

Merged
kunalspathak merged 15 commits into
dotnet:masterfrom
kunalspathak:create-scalar-multiple
May 5, 2020
Merged

ARM64 intrinsic support for Vector64.Create() and Vector128.Create()#35590
kunalspathak merged 15 commits into
dotnet:masterfrom
kunalspathak:create-scalar-multiple

Conversation

@kunalspathak

@kunalspathakkunalspathak commented Apr 28, 2020

Copy link
Copy Markdown
Contributor

Added hardware intrinsic for various overloads of Vector64.Create() and Vector128.Create():

  • Multiple arguments - The APIs that takes multiple parameters to be set in respective lanes are implemented in C# using AdvSimd.Insert.
  • Single arguments - The APIs that takes single argument and should be copied in all lanes are implemented in JIT by generating dup/mov/fmov instructions.

While I was there, I noticed an edge case where we hit assert if trying to emit an immediate int.MaxValue. Fixed it as well.

Contributes to #33308 and #33496.
Fixes: #35821

@Dotnet-GitSync-BotDotnet-GitSync-Bot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Apr 28, 2020
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib , @tannergooding

@BruceForstall

Copy link
Copy Markdown
Contributor

cc @TamarChristinaArm

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

not for this PR, but since we only have 8 argument registers the rest will be passed on the stack.
In which case would be easier to load the remaining 64 bits into a register directly from the stack into the top part of the 128-bit vector.

i.e. do something like this:

 and w0, w0, 255
fmov s0, w0
ins v0.b[1], w1
ins v0.b[2], w2
ins v0.b[3], w3
ins v0.b[4], w4
ins v0.b[5], w5
ins v0.b[6], w6
ins v0.b[7], w7
ld1 {v0.d}[1], [sp]

And you can do this is something like this for a lot of the cases with more than 8 arguments. Do you know what the code generates here at the moment?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's a good point which I didn't realize. I will explore your suggestion.

Generated code
 53001C00 uxtb w0, w0 4E011C10 ins v16.b[0], w0 53001C20 uxtb w0, w1 4E031C10 ins v16.b[1], w0 53001C40 uxtb w0, w2 4E051C10 ins v16.b[2], w0 53001C60 uxtb w0, w3 4E071C10 ins v16.b[3], w0 53001C80 uxtb w0, w4 4E091C10 ins v16.b[4], w0 53001CA0 uxtb w0, w5 4E0B1C10 ins v16.b[5], w0 53001CC0 uxtb w0, w6 4E0D1C10 ins v16.b[6], w0 53001CE0 uxtb w0, w7 4E0F1C10 ins v16.b[7], w0 B94023A0 ldr w0,[fp,#32] // [V08 arg8] 53001C00 uxtb w0, w0 4E111C10 ins v16.b[8], w0 B9402BA0 ldr w0,[fp,#40] // [V09 arg9] 53001C00 uxtb w0, w0 4E131C10 ins v16.b[9], w0 B94033A0 ldr w0,[fp,#48] // [V10 arg10] 53001C00 uxtb w0, w0 4E151C10 ins v16.b[10], w0 B9403BA0 ldr w0,[fp,#56] // [V11 arg11] 53001C00 uxtb w0, w0 4E171C10 ins v16.b[11], w0 B94043A0 ldr w0,[fp,#64] // [V12 arg12] 53001C00 uxtb w0, w0 4E191C10 ins v16.b[12], w0 B9404BA0 ldr w0,[fp,#72] // [V13 arg13] 53001C00 uxtb w0, w0 4E1B1C10 ins v16.b[13], w0 B94053A0 ldr w0,[fp,#80] // [V14 arg14] 53001C00 uxtb w0, w0 4E1D1C10 ins v16.b[14], w0 B9405BA0 ldr w0,[fp,#88] // [V15 arg15] 53001C00 uxtb w0, w0 4EB01E08 mov v8.16b, v16.16b 4E1F1C08 ins v8.b[15], w0 D28D0500 movz x0, #0x6828 F2A6EF60 movk x0, #0x377bLSL #16 F2CFFFA0 movk x0, #0x7ffdLSL #32 6E084509 mov v9.d[0], v8.d[1] 97FF53EA bl CORINFO_HELP_NEWSFAST 6E180528 mov v8.d[1], v9.d[0] 3C808008 str q8,[x0,#8]

And you can do this is something like this for a lot of the cases with more than 8 arguments.

Actually, this is the only API that has more than 8 arguments.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I wonder how hard is goint to be to get rid of the uxtb-s

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah, I am not much familiar with this, but would love to get rid of them. Just to call out, they show up for byte, sbyte, short and ushort parameters.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah, I am not much familiar with this, but would love to get rid of them. Just to call out, they show up for byte, sbyte, short and ushort parameters.

Correct, because these types are usually "normalized" (i.e. sign- or zero-extended) to a 32 bit value

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah, those aren't needed indeed, would be good to get rid of them if possible since they double the number of instructions.

Also

 B94053A0 ldr w0, [fp,#80] // [V14 arg14]
53001C00 uxtb w0, w0
4E1D1C10 ins v16.b[14], w0

Without the optimization I talked about above should ideally be

ld1 {v16.b[14]}, [x1]

where x1 gets incremented, or you could use a which prevents from having to move between register files. But also don't know how easy this is to do..

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I briefly discussed various options with @echesakovMSFT and concluded that it might not be straight forward. I would like to revisit it after we implement other APIs. Opened #35688 to track it.

Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/lowerarmarch.cpp Outdated
Comment threadsrc/coreclr/src/jit/emitarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated

@echesakovechesakov left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks Good to me with some questions/suggestions.

Comment threadsrc/coreclr/src/jit/emitarm64.cpp Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@kunalspathak You might want to add comment here in the same fashion as it's done for Sse2 case

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Spoke offline and this is not needed.

@kunalspathak
kunalspathakforce-pushed the create-scalar-multiple branch from 1a86bc1 to 76cc65eCompareMay 1, 2020 21:55
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@tannergooding and @echesakovMSFT - Just FYI, while working on something else I found 2 issues for which we don't have test coverage today.

  1. Vector64.Create((double)10) (any integer value casted to double)
  2. Passing Vector64<double> or Vector64<long> as parameter to a function.

I will fix and add test coverage for it before merging.

@kunalspathak
kunalspathak merged commit d23f1a2 into dotnet:masterMay 5, 2020
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Assertion failed '!"Didn't find a class handle for simdType"' when trying to pass Vector64<ulong> to function

6 participants

@kunalspathak@BruceForstall@echesakov@tannergooding@TamarChristinaArm@Dotnet-GitSync-Bot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

ARM64 intrinsic support for Vector64.Create() and Vector128.Create() - #35590

Merged
kunalspathak merged 15 commits into
dotnet:masterfrom
kunalspathak:create-scalar-multiple
May 5, 2020
Merged

ARM64 intrinsic support for Vector64.Create() and Vector128.Create()#35590
kunalspathak merged 15 commits into
dotnet:masterfrom
kunalspathak:create-scalar-multiple

Conversation

@kunalspathak

@kunalspathakkunalspathak commented Apr 28, 2020

Copy link
Copy Markdown
Contributor

Added hardware intrinsic for various overloads of Vector64.Create() and Vector128.Create():

  • Multiple arguments - The APIs that takes multiple parameters to be set in respective lanes are implemented in C# using AdvSimd.Insert.
  • Single arguments - The APIs that takes single argument and should be copied in all lanes are implemented in JIT by generating dup/mov/fmov instructions.

While I was there, I noticed an edge case where we hit assert if trying to emit an immediate int.MaxValue. Fixed it as well.

Contributes to #33308 and #33496.
Fixes: #35821

@Dotnet-GitSync-BotDotnet-GitSync-Bot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Apr 28, 2020
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib , @tannergooding

@BruceForstall

Copy link
Copy Markdown
Contributor

cc @TamarChristinaArm

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

not for this PR, but since we only have 8 argument registers the rest will be passed on the stack.
In which case would be easier to load the remaining 64 bits into a register directly from the stack into the top part of the 128-bit vector.

i.e. do something like this:

 and w0, w0, 255
fmov s0, w0
ins v0.b[1], w1
ins v0.b[2], w2
ins v0.b[3], w3
ins v0.b[4], w4
ins v0.b[5], w5
ins v0.b[6], w6
ins v0.b[7], w7
ld1 {v0.d}[1], [sp]

And you can do this is something like this for a lot of the cases with more than 8 arguments. Do you know what the code generates here at the moment?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's a good point which I didn't realize. I will explore your suggestion.

Generated code
 53001C00 uxtb w0, w0 4E011C10 ins v16.b[0], w0 53001C20 uxtb w0, w1 4E031C10 ins v16.b[1], w0 53001C40 uxtb w0, w2 4E051C10 ins v16.b[2], w0 53001C60 uxtb w0, w3 4E071C10 ins v16.b[3], w0 53001C80 uxtb w0, w4 4E091C10 ins v16.b[4], w0 53001CA0 uxtb w0, w5 4E0B1C10 ins v16.b[5], w0 53001CC0 uxtb w0, w6 4E0D1C10 ins v16.b[6], w0 53001CE0 uxtb w0, w7 4E0F1C10 ins v16.b[7], w0 B94023A0 ldr w0,[fp,#32] // [V08 arg8] 53001C00 uxtb w0, w0 4E111C10 ins v16.b[8], w0 B9402BA0 ldr w0,[fp,#40] // [V09 arg9] 53001C00 uxtb w0, w0 4E131C10 ins v16.b[9], w0 B94033A0 ldr w0,[fp,#48] // [V10 arg10] 53001C00 uxtb w0, w0 4E151C10 ins v16.b[10], w0 B9403BA0 ldr w0,[fp,#56] // [V11 arg11] 53001C00 uxtb w0, w0 4E171C10 ins v16.b[11], w0 B94043A0 ldr w0,[fp,#64] // [V12 arg12] 53001C00 uxtb w0, w0 4E191C10 ins v16.b[12], w0 B9404BA0 ldr w0,[fp,#72] // [V13 arg13] 53001C00 uxtb w0, w0 4E1B1C10 ins v16.b[13], w0 B94053A0 ldr w0,[fp,#80] // [V14 arg14] 53001C00 uxtb w0, w0 4E1D1C10 ins v16.b[14], w0 B9405BA0 ldr w0,[fp,#88] // [V15 arg15] 53001C00 uxtb w0, w0 4EB01E08 mov v8.16b, v16.16b 4E1F1C08 ins v8.b[15], w0 D28D0500 movz x0, #0x6828 F2A6EF60 movk x0, #0x377bLSL #16 F2CFFFA0 movk x0, #0x7ffdLSL #32 6E084509 mov v9.d[0], v8.d[1] 97FF53EA bl CORINFO_HELP_NEWSFAST 6E180528 mov v8.d[1], v9.d[0] 3C808008 str q8,[x0,#8]

And you can do this is something like this for a lot of the cases with more than 8 arguments.

Actually, this is the only API that has more than 8 arguments.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I wonder how hard is goint to be to get rid of the uxtb-s

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah, I am not much familiar with this, but would love to get rid of them. Just to call out, they show up for byte, sbyte, short and ushort parameters.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah, I am not much familiar with this, but would love to get rid of them. Just to call out, they show up for byte, sbyte, short and ushort parameters.

Correct, because these types are usually "normalized" (i.e. sign- or zero-extended) to a 32 bit value

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah, those aren't needed indeed, would be good to get rid of them if possible since they double the number of instructions.

Also

 B94053A0 ldr w0, [fp,#80] // [V14 arg14]
53001C00 uxtb w0, w0
4E1D1C10 ins v16.b[14], w0

Without the optimization I talked about above should ideally be

ld1 {v16.b[14]}, [x1]

where x1 gets incremented, or you could use a which prevents from having to move between register files. But also don't know how easy this is to do..

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I briefly discussed various options with @echesakovMSFT and concluded that it might not be straight forward. I would like to revisit it after we implement other APIs. Opened #35688 to track it.

Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/lowerarmarch.cpp Outdated
Comment threadsrc/coreclr/src/jit/emitarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated

@echesakovechesakov left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks Good to me with some questions/suggestions.

Comment threadsrc/coreclr/src/jit/emitarm64.cpp Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@kunalspathak You might want to add comment here in the same fashion as it's done for Sse2 case

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Spoke offline and this is not needed.

@kunalspathak
kunalspathakforce-pushed the create-scalar-multiple branch from 1a86bc1 to 76cc65eCompareMay 1, 2020 21:55
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@tannergooding and @echesakovMSFT - Just FYI, while working on something else I found 2 issues for which we don't have test coverage today.

  1. Vector64.Create((double)10) (any integer value casted to double)
  2. Passing Vector64<double> or Vector64<long> as parameter to a function.

I will fix and add test coverage for it before merging.

@kunalspathak
kunalspathak merged commit d23f1a2 into dotnet:masterMay 5, 2020
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Assertion failed '!"Didn't find a class handle for simdType"' when trying to pass Vector64<ulong> to function

6 participants

@kunalspathak@BruceForstall@echesakov@tannergooding@TamarChristinaArm@Dotnet-GitSync-Bot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARM64 intrinsic support for Vector64.Create() and Vector128.Create() - #35590

Merged
kunalspathak merged 15 commits into
dotnet:masterfrom
kunalspathak:create-scalar-multiple
May 5, 2020
Merged

ARM64 intrinsic support for Vector64.Create() and Vector128.Create()#35590
kunalspathak merged 15 commits into
dotnet:masterfrom
kunalspathak:create-scalar-multiple

Conversation

@kunalspathak

@kunalspathakkunalspathak commented Apr 28, 2020

Copy link
Copy Markdown
Contributor

Added hardware intrinsic for various overloads of Vector64.Create() and Vector128.Create():

  • Multiple arguments - The APIs that takes multiple parameters to be set in respective lanes are implemented in C# using AdvSimd.Insert.
  • Single arguments - The APIs that takes single argument and should be copied in all lanes are implemented in JIT by generating dup/mov/fmov instructions.

While I was there, I noticed an edge case where we hit assert if trying to emit an immediate int.MaxValue. Fixed it as well.

Contributes to #33308 and #33496.
Fixes: #35821

@Dotnet-GitSync-BotDotnet-GitSync-Bot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Apr 28, 2020
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib , @tannergooding

@BruceForstall

Copy link
Copy Markdown
Contributor

cc @TamarChristinaArm

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

not for this PR, but since we only have 8 argument registers the rest will be passed on the stack.
In which case would be easier to load the remaining 64 bits into a register directly from the stack into the top part of the 128-bit vector.

i.e. do something like this:

 and w0, w0, 255
fmov s0, w0
ins v0.b[1], w1
ins v0.b[2], w2
ins v0.b[3], w3
ins v0.b[4], w4
ins v0.b[5], w5
ins v0.b[6], w6
ins v0.b[7], w7
ld1 {v0.d}[1], [sp]

And you can do this is something like this for a lot of the cases with more than 8 arguments. Do you know what the code generates here at the moment?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's a good point which I didn't realize. I will explore your suggestion.

Generated code
 53001C00 uxtb w0, w0 4E011C10 ins v16.b[0], w0 53001C20 uxtb w0, w1 4E031C10 ins v16.b[1], w0 53001C40 uxtb w0, w2 4E051C10 ins v16.b[2], w0 53001C60 uxtb w0, w3 4E071C10 ins v16.b[3], w0 53001C80 uxtb w0, w4 4E091C10 ins v16.b[4], w0 53001CA0 uxtb w0, w5 4E0B1C10 ins v16.b[5], w0 53001CC0 uxtb w0, w6 4E0D1C10 ins v16.b[6], w0 53001CE0 uxtb w0, w7 4E0F1C10 ins v16.b[7], w0 B94023A0 ldr w0,[fp,#32] // [V08 arg8] 53001C00 uxtb w0, w0 4E111C10 ins v16.b[8], w0 B9402BA0 ldr w0,[fp,#40] // [V09 arg9] 53001C00 uxtb w0, w0 4E131C10 ins v16.b[9], w0 B94033A0 ldr w0,[fp,#48] // [V10 arg10] 53001C00 uxtb w0, w0 4E151C10 ins v16.b[10], w0 B9403BA0 ldr w0,[fp,#56] // [V11 arg11] 53001C00 uxtb w0, w0 4E171C10 ins v16.b[11], w0 B94043A0 ldr w0,[fp,#64] // [V12 arg12] 53001C00 uxtb w0, w0 4E191C10 ins v16.b[12], w0 B9404BA0 ldr w0,[fp,#72] // [V13 arg13] 53001C00 uxtb w0, w0 4E1B1C10 ins v16.b[13], w0 B94053A0 ldr w0,[fp,#80] // [V14 arg14] 53001C00 uxtb w0, w0 4E1D1C10 ins v16.b[14], w0 B9405BA0 ldr w0,[fp,#88] // [V15 arg15] 53001C00 uxtb w0, w0 4EB01E08 mov v8.16b, v16.16b 4E1F1C08 ins v8.b[15], w0 D28D0500 movz x0, #0x6828 F2A6EF60 movk x0, #0x377bLSL #16 F2CFFFA0 movk x0, #0x7ffdLSL #32 6E084509 mov v9.d[0], v8.d[1] 97FF53EA bl CORINFO_HELP_NEWSFAST 6E180528 mov v8.d[1], v9.d[0] 3C808008 str q8,[x0,#8]

And you can do this is something like this for a lot of the cases with more than 8 arguments.

Actually, this is the only API that has more than 8 arguments.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I wonder how hard is goint to be to get rid of the uxtb-s

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah, I am not much familiar with this, but would love to get rid of them. Just to call out, they show up for byte, sbyte, short and ushort parameters.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah, I am not much familiar with this, but would love to get rid of them. Just to call out, they show up for byte, sbyte, short and ushort parameters.

Correct, because these types are usually "normalized" (i.e. sign- or zero-extended) to a 32 bit value

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah, those aren't needed indeed, would be good to get rid of them if possible since they double the number of instructions.

Also

 B94053A0 ldr w0, [fp,#80] // [V14 arg14]
53001C00 uxtb w0, w0
4E1D1C10 ins v16.b[14], w0

Without the optimization I talked about above should ideally be

ld1 {v16.b[14]}, [x1]

where x1 gets incremented, or you could use a which prevents from having to move between register files. But also don't know how easy this is to do..

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I briefly discussed various options with @echesakovMSFT and concluded that it might not be straight forward. I would like to revisit it after we implement other APIs. Opened #35688 to track it.

Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/lowerarmarch.cpp Outdated
Comment threadsrc/coreclr/src/jit/emitarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated

@echesakovechesakov left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks Good to me with some questions/suggestions.

Comment threadsrc/coreclr/src/jit/emitarm64.cpp Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@kunalspathak You might want to add comment here in the same fashion as it's done for Sse2 case

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Spoke offline and this is not needed.

@kunalspathak
kunalspathakforce-pushed the create-scalar-multiple branch from 1a86bc1 to 76cc65eCompareMay 1, 2020 21:55
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@tannergooding and @echesakovMSFT - Just FYI, while working on something else I found 2 issues for which we don't have test coverage today.

  1. Vector64.Create((double)10) (any integer value casted to double)
  2. Passing Vector64<double> or Vector64<long> as parameter to a function.

I will fix and add test coverage for it before merging.

@kunalspathak
kunalspathak merged commit d23f1a2 into dotnet:masterMay 5, 2020
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Assertion failed '!"Didn't find a class handle for simdType"' when trying to pass Vector64<ulong> to function

6 participants

@kunalspathak@BruceForstall@echesakov@tannergooding@TamarChristinaArm@Dotnet-GitSync-Bot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARM64 intrinsic support for Vector64.Create() and Vector128.Create() - #35590

Merged
kunalspathak merged 15 commits into
dotnet:masterfrom
kunalspathak:create-scalar-multiple
May 5, 2020
Merged

ARM64 intrinsic support for Vector64.Create() and Vector128.Create()#35590
kunalspathak merged 15 commits into
dotnet:masterfrom
kunalspathak:create-scalar-multiple

Conversation

@kunalspathak

@kunalspathakkunalspathak commented Apr 28, 2020

Copy link
Copy Markdown
Contributor

Added hardware intrinsic for various overloads of Vector64.Create() and Vector128.Create():

  • Multiple arguments - The APIs that takes multiple parameters to be set in respective lanes are implemented in C# using AdvSimd.Insert.
  • Single arguments - The APIs that takes single argument and should be copied in all lanes are implemented in JIT by generating dup/mov/fmov instructions.

While I was there, I noticed an edge case where we hit assert if trying to emit an immediate int.MaxValue. Fixed it as well.

Contributes to #33308 and #33496.
Fixes: #35821

@Dotnet-GitSync-BotDotnet-GitSync-Bot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Apr 28, 2020
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib , @tannergooding

@BruceForstall

Copy link
Copy Markdown
Contributor

cc @TamarChristinaArm

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

not for this PR, but since we only have 8 argument registers the rest will be passed on the stack.
In which case would be easier to load the remaining 64 bits into a register directly from the stack into the top part of the 128-bit vector.

i.e. do something like this:

 and w0, w0, 255
fmov s0, w0
ins v0.b[1], w1
ins v0.b[2], w2
ins v0.b[3], w3
ins v0.b[4], w4
ins v0.b[5], w5
ins v0.b[6], w6
ins v0.b[7], w7
ld1 {v0.d}[1], [sp]

And you can do this is something like this for a lot of the cases with more than 8 arguments. Do you know what the code generates here at the moment?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's a good point which I didn't realize. I will explore your suggestion.

Generated code
 53001C00 uxtb w0, w0 4E011C10 ins v16.b[0], w0 53001C20 uxtb w0, w1 4E031C10 ins v16.b[1], w0 53001C40 uxtb w0, w2 4E051C10 ins v16.b[2], w0 53001C60 uxtb w0, w3 4E071C10 ins v16.b[3], w0 53001C80 uxtb w0, w4 4E091C10 ins v16.b[4], w0 53001CA0 uxtb w0, w5 4E0B1C10 ins v16.b[5], w0 53001CC0 uxtb w0, w6 4E0D1C10 ins v16.b[6], w0 53001CE0 uxtb w0, w7 4E0F1C10 ins v16.b[7], w0 B94023A0 ldr w0,[fp,#32] // [V08 arg8] 53001C00 uxtb w0, w0 4E111C10 ins v16.b[8], w0 B9402BA0 ldr w0,[fp,#40] // [V09 arg9] 53001C00 uxtb w0, w0 4E131C10 ins v16.b[9], w0 B94033A0 ldr w0,[fp,#48] // [V10 arg10] 53001C00 uxtb w0, w0 4E151C10 ins v16.b[10], w0 B9403BA0 ldr w0,[fp,#56] // [V11 arg11] 53001C00 uxtb w0, w0 4E171C10 ins v16.b[11], w0 B94043A0 ldr w0,[fp,#64] // [V12 arg12] 53001C00 uxtb w0, w0 4E191C10 ins v16.b[12], w0 B9404BA0 ldr w0,[fp,#72] // [V13 arg13] 53001C00 uxtb w0, w0 4E1B1C10 ins v16.b[13], w0 B94053A0 ldr w0,[fp,#80] // [V14 arg14] 53001C00 uxtb w0, w0 4E1D1C10 ins v16.b[14], w0 B9405BA0 ldr w0,[fp,#88] // [V15 arg15] 53001C00 uxtb w0, w0 4EB01E08 mov v8.16b, v16.16b 4E1F1C08 ins v8.b[15], w0 D28D0500 movz x0, #0x6828 F2A6EF60 movk x0, #0x377bLSL #16 F2CFFFA0 movk x0, #0x7ffdLSL #32 6E084509 mov v9.d[0], v8.d[1] 97FF53EA bl CORINFO_HELP_NEWSFAST 6E180528 mov v8.d[1], v9.d[0] 3C808008 str q8,[x0,#8]

And you can do this is something like this for a lot of the cases with more than 8 arguments.

Actually, this is the only API that has more than 8 arguments.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I wonder how hard is goint to be to get rid of the uxtb-s

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah, I am not much familiar with this, but would love to get rid of them. Just to call out, they show up for byte, sbyte, short and ushort parameters.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah, I am not much familiar with this, but would love to get rid of them. Just to call out, they show up for byte, sbyte, short and ushort parameters.

Correct, because these types are usually "normalized" (i.e. sign- or zero-extended) to a 32 bit value

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah, those aren't needed indeed, would be good to get rid of them if possible since they double the number of instructions.

Also

 B94053A0 ldr w0, [fp,#80] // [V14 arg14]
53001C00 uxtb w0, w0
4E1D1C10 ins v16.b[14], w0

Without the optimization I talked about above should ideally be

ld1 {v16.b[14]}, [x1]

where x1 gets incremented, or you could use a which prevents from having to move between register files. But also don't know how easy this is to do..

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I briefly discussed various options with @echesakovMSFT and concluded that it might not be straight forward. I would like to revisit it after we implement other APIs. Opened #35688 to track it.

Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/lowerarmarch.cpp Outdated
Comment threadsrc/coreclr/src/jit/emitarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated

@echesakovechesakov left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks Good to me with some questions/suggestions.

Comment threadsrc/coreclr/src/jit/emitarm64.cpp Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@kunalspathak You might want to add comment here in the same fashion as it's done for Sse2 case

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Spoke offline and this is not needed.

@kunalspathak
kunalspathakforce-pushed the create-scalar-multiple branch from 1a86bc1 to 76cc65eCompareMay 1, 2020 21:55
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@tannergooding and @echesakovMSFT - Just FYI, while working on something else I found 2 issues for which we don't have test coverage today.

  1. Vector64.Create((double)10) (any integer value casted to double)
  2. Passing Vector64<double> or Vector64<long> as parameter to a function.

I will fix and add test coverage for it before merging.

@kunalspathak
kunalspathak merged commit d23f1a2 into dotnet:masterMay 5, 2020
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Assertion failed '!"Didn't find a class handle for simdType"' when trying to pass Vector64<ulong> to function

6 participants

@kunalspathak@BruceForstall@echesakov@tannergooding@TamarChristinaArm@Dotnet-GitSync-Bot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

ARM64 intrinsic support for Vector64.Create() and Vector128.Create() - #35590

Merged
kunalspathak merged 15 commits into
dotnet:masterfrom
kunalspathak:create-scalar-multiple
May 5, 2020
Merged

ARM64 intrinsic support for Vector64.Create() and Vector128.Create()#35590
kunalspathak merged 15 commits into
dotnet:masterfrom
kunalspathak:create-scalar-multiple

Conversation

@kunalspathak

@kunalspathakkunalspathak commented Apr 28, 2020

Copy link
Copy Markdown
Contributor

Added hardware intrinsic for various overloads of Vector64.Create() and Vector128.Create():

  • Multiple arguments - The APIs that takes multiple parameters to be set in respective lanes are implemented in C# using AdvSimd.Insert.
  • Single arguments - The APIs that takes single argument and should be copied in all lanes are implemented in JIT by generating dup/mov/fmov instructions.

While I was there, I noticed an edge case where we hit assert if trying to emit an immediate int.MaxValue. Fixed it as well.

Contributes to #33308 and #33496.
Fixes: #35821

@Dotnet-GitSync-BotDotnet-GitSync-Bot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Apr 28, 2020
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib , @tannergooding

@BruceForstall

Copy link
Copy Markdown
Contributor

cc @TamarChristinaArm

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

not for this PR, but since we only have 8 argument registers the rest will be passed on the stack.
In which case would be easier to load the remaining 64 bits into a register directly from the stack into the top part of the 128-bit vector.

i.e. do something like this:

 and w0, w0, 255
fmov s0, w0
ins v0.b[1], w1
ins v0.b[2], w2
ins v0.b[3], w3
ins v0.b[4], w4
ins v0.b[5], w5
ins v0.b[6], w6
ins v0.b[7], w7
ld1 {v0.d}[1], [sp]

And you can do this is something like this for a lot of the cases with more than 8 arguments. Do you know what the code generates here at the moment?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's a good point which I didn't realize. I will explore your suggestion.

Generated code
 53001C00 uxtb w0, w0 4E011C10 ins v16.b[0], w0 53001C20 uxtb w0, w1 4E031C10 ins v16.b[1], w0 53001C40 uxtb w0, w2 4E051C10 ins v16.b[2], w0 53001C60 uxtb w0, w3 4E071C10 ins v16.b[3], w0 53001C80 uxtb w0, w4 4E091C10 ins v16.b[4], w0 53001CA0 uxtb w0, w5 4E0B1C10 ins v16.b[5], w0 53001CC0 uxtb w0, w6 4E0D1C10 ins v16.b[6], w0 53001CE0 uxtb w0, w7 4E0F1C10 ins v16.b[7], w0 B94023A0 ldr w0,[fp,#32] // [V08 arg8] 53001C00 uxtb w0, w0 4E111C10 ins v16.b[8], w0 B9402BA0 ldr w0,[fp,#40] // [V09 arg9] 53001C00 uxtb w0, w0 4E131C10 ins v16.b[9], w0 B94033A0 ldr w0,[fp,#48] // [V10 arg10] 53001C00 uxtb w0, w0 4E151C10 ins v16.b[10], w0 B9403BA0 ldr w0,[fp,#56] // [V11 arg11] 53001C00 uxtb w0, w0 4E171C10 ins v16.b[11], w0 B94043A0 ldr w0,[fp,#64] // [V12 arg12] 53001C00 uxtb w0, w0 4E191C10 ins v16.b[12], w0 B9404BA0 ldr w0,[fp,#72] // [V13 arg13] 53001C00 uxtb w0, w0 4E1B1C10 ins v16.b[13], w0 B94053A0 ldr w0,[fp,#80] // [V14 arg14] 53001C00 uxtb w0, w0 4E1D1C10 ins v16.b[14], w0 B9405BA0 ldr w0,[fp,#88] // [V15 arg15] 53001C00 uxtb w0, w0 4EB01E08 mov v8.16b, v16.16b 4E1F1C08 ins v8.b[15], w0 D28D0500 movz x0, #0x6828 F2A6EF60 movk x0, #0x377bLSL #16 F2CFFFA0 movk x0, #0x7ffdLSL #32 6E084509 mov v9.d[0], v8.d[1] 97FF53EA bl CORINFO_HELP_NEWSFAST 6E180528 mov v8.d[1], v9.d[0] 3C808008 str q8,[x0,#8]

And you can do this is something like this for a lot of the cases with more than 8 arguments.

Actually, this is the only API that has more than 8 arguments.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I wonder how hard is goint to be to get rid of the uxtb-s

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah, I am not much familiar with this, but would love to get rid of them. Just to call out, they show up for byte, sbyte, short and ushort parameters.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah, I am not much familiar with this, but would love to get rid of them. Just to call out, they show up for byte, sbyte, short and ushort parameters.

Correct, because these types are usually "normalized" (i.e. sign- or zero-extended) to a 32 bit value

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah, those aren't needed indeed, would be good to get rid of them if possible since they double the number of instructions.

Also

 B94053A0 ldr w0, [fp,#80] // [V14 arg14]
53001C00 uxtb w0, w0
4E1D1C10 ins v16.b[14], w0

Without the optimization I talked about above should ideally be

ld1 {v16.b[14]}, [x1]

where x1 gets incremented, or you could use a which prevents from having to move between register files. But also don't know how easy this is to do..

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I briefly discussed various options with @echesakovMSFT and concluded that it might not be straight forward. I would like to revisit it after we implement other APIs. Opened #35688 to track it.

Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/lowerarmarch.cpp Outdated
Comment threadsrc/coreclr/src/jit/emitarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated

@echesakovechesakov left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks Good to me with some questions/suggestions.

Comment threadsrc/coreclr/src/jit/emitarm64.cpp Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@kunalspathak You might want to add comment here in the same fashion as it's done for Sse2 case

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Spoke offline and this is not needed.

@kunalspathak
kunalspathakforce-pushed the create-scalar-multiple branch from 1a86bc1 to 76cc65eCompareMay 1, 2020 21:55
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@tannergooding and @echesakovMSFT - Just FYI, while working on something else I found 2 issues for which we don't have test coverage today.

  1. Vector64.Create((double)10) (any integer value casted to double)
  2. Passing Vector64<double> or Vector64<long> as parameter to a function.

I will fix and add test coverage for it before merging.

@kunalspathak
kunalspathak merged commit d23f1a2 into dotnet:masterMay 5, 2020
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Assertion failed '!"Didn't find a class handle for simdType"' when trying to pass Vector64<ulong> to function

6 participants

@kunalspathak@BruceForstall@echesakov@tannergooding@TamarChristinaArm@Dotnet-GitSync-Bot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARM64 intrinsic support for Vector64.Create() and Vector128.Create() - #35590

Merged
kunalspathak merged 15 commits into
dotnet:masterfrom
kunalspathak:create-scalar-multiple
May 5, 2020
Merged

ARM64 intrinsic support for Vector64.Create() and Vector128.Create()#35590
kunalspathak merged 15 commits into
dotnet:masterfrom
kunalspathak:create-scalar-multiple

Conversation

@kunalspathak

@kunalspathakkunalspathak commented Apr 28, 2020

Copy link
Copy Markdown
Contributor

Added hardware intrinsic for various overloads of Vector64.Create() and Vector128.Create():

  • Multiple arguments - The APIs that takes multiple parameters to be set in respective lanes are implemented in C# using AdvSimd.Insert.
  • Single arguments - The APIs that takes single argument and should be copied in all lanes are implemented in JIT by generating dup/mov/fmov instructions.

While I was there, I noticed an edge case where we hit assert if trying to emit an immediate int.MaxValue. Fixed it as well.

Contributes to #33308 and #33496.
Fixes: #35821

@Dotnet-GitSync-BotDotnet-GitSync-Bot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Apr 28, 2020
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib , @tannergooding

@BruceForstall

Copy link
Copy Markdown
Contributor

cc @TamarChristinaArm

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

not for this PR, but since we only have 8 argument registers the rest will be passed on the stack.
In which case would be easier to load the remaining 64 bits into a register directly from the stack into the top part of the 128-bit vector.

i.e. do something like this:

 and w0, w0, 255
fmov s0, w0
ins v0.b[1], w1
ins v0.b[2], w2
ins v0.b[3], w3
ins v0.b[4], w4
ins v0.b[5], w5
ins v0.b[6], w6
ins v0.b[7], w7
ld1 {v0.d}[1], [sp]

And you can do this is something like this for a lot of the cases with more than 8 arguments. Do you know what the code generates here at the moment?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's a good point which I didn't realize. I will explore your suggestion.

Generated code
 53001C00 uxtb w0, w0 4E011C10 ins v16.b[0], w0 53001C20 uxtb w0, w1 4E031C10 ins v16.b[1], w0 53001C40 uxtb w0, w2 4E051C10 ins v16.b[2], w0 53001C60 uxtb w0, w3 4E071C10 ins v16.b[3], w0 53001C80 uxtb w0, w4 4E091C10 ins v16.b[4], w0 53001CA0 uxtb w0, w5 4E0B1C10 ins v16.b[5], w0 53001CC0 uxtb w0, w6 4E0D1C10 ins v16.b[6], w0 53001CE0 uxtb w0, w7 4E0F1C10 ins v16.b[7], w0 B94023A0 ldr w0,[fp,#32] // [V08 arg8] 53001C00 uxtb w0, w0 4E111C10 ins v16.b[8], w0 B9402BA0 ldr w0,[fp,#40] // [V09 arg9] 53001C00 uxtb w0, w0 4E131C10 ins v16.b[9], w0 B94033A0 ldr w0,[fp,#48] // [V10 arg10] 53001C00 uxtb w0, w0 4E151C10 ins v16.b[10], w0 B9403BA0 ldr w0,[fp,#56] // [V11 arg11] 53001C00 uxtb w0, w0 4E171C10 ins v16.b[11], w0 B94043A0 ldr w0,[fp,#64] // [V12 arg12] 53001C00 uxtb w0, w0 4E191C10 ins v16.b[12], w0 B9404BA0 ldr w0,[fp,#72] // [V13 arg13] 53001C00 uxtb w0, w0 4E1B1C10 ins v16.b[13], w0 B94053A0 ldr w0,[fp,#80] // [V14 arg14] 53001C00 uxtb w0, w0 4E1D1C10 ins v16.b[14], w0 B9405BA0 ldr w0,[fp,#88] // [V15 arg15] 53001C00 uxtb w0, w0 4EB01E08 mov v8.16b, v16.16b 4E1F1C08 ins v8.b[15], w0 D28D0500 movz x0, #0x6828 F2A6EF60 movk x0, #0x377bLSL #16 F2CFFFA0 movk x0, #0x7ffdLSL #32 6E084509 mov v9.d[0], v8.d[1] 97FF53EA bl CORINFO_HELP_NEWSFAST 6E180528 mov v8.d[1], v9.d[0] 3C808008 str q8,[x0,#8]

And you can do this is something like this for a lot of the cases with more than 8 arguments.

Actually, this is the only API that has more than 8 arguments.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I wonder how hard is goint to be to get rid of the uxtb-s

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah, I am not much familiar with this, but would love to get rid of them. Just to call out, they show up for byte, sbyte, short and ushort parameters.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah, I am not much familiar with this, but would love to get rid of them. Just to call out, they show up for byte, sbyte, short and ushort parameters.

Correct, because these types are usually "normalized" (i.e. sign- or zero-extended) to a 32 bit value

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah, those aren't needed indeed, would be good to get rid of them if possible since they double the number of instructions.

Also

 B94053A0 ldr w0, [fp,#80] // [V14 arg14]
53001C00 uxtb w0, w0
4E1D1C10 ins v16.b[14], w0

Without the optimization I talked about above should ideally be

ld1 {v16.b[14]}, [x1]

where x1 gets incremented, or you could use a which prevents from having to move between register files. But also don't know how easy this is to do..

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I briefly discussed various options with @echesakovMSFT and concluded that it might not be straight forward. I would like to revisit it after we implement other APIs. Opened #35688 to track it.

Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/lowerarmarch.cpp Outdated
Comment threadsrc/coreclr/src/jit/emitarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated

@echesakovechesakov left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks Good to me with some questions/suggestions.

Comment threadsrc/coreclr/src/jit/emitarm64.cpp Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@kunalspathak You might want to add comment here in the same fashion as it's done for Sse2 case

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Spoke offline and this is not needed.

@kunalspathak
kunalspathakforce-pushed the create-scalar-multiple branch from 1a86bc1 to 76cc65eCompareMay 1, 2020 21:55
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@tannergooding and @echesakovMSFT - Just FYI, while working on something else I found 2 issues for which we don't have test coverage today.

  1. Vector64.Create((double)10) (any integer value casted to double)
  2. Passing Vector64<double> or Vector64<long> as parameter to a function.

I will fix and add test coverage for it before merging.

@kunalspathak
kunalspathak merged commit d23f1a2 into dotnet:masterMay 5, 2020
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Assertion failed '!"Didn't find a class handle for simdType"' when trying to pass Vector64<ulong> to function

6 participants

@kunalspathak@BruceForstall@echesakov@tannergooding@TamarChristinaArm@Dotnet-GitSync-Bot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ARM64 intrinsic support for Vector64.Create() and Vector128.Create() - #35590

Merged
kunalspathak merged 15 commits into
dotnet:masterfrom
kunalspathak:create-scalar-multiple
May 5, 2020
Merged

ARM64 intrinsic support for Vector64.Create() and Vector128.Create()#35590
kunalspathak merged 15 commits into
dotnet:masterfrom
kunalspathak:create-scalar-multiple

Conversation

@kunalspathak

@kunalspathakkunalspathak commented Apr 28, 2020

Copy link
Copy Markdown
Contributor

Added hardware intrinsic for various overloads of Vector64.Create() and Vector128.Create():

  • Multiple arguments - The APIs that takes multiple parameters to be set in respective lanes are implemented in C# using AdvSimd.Insert.
  • Single arguments - The APIs that takes single argument and should be copied in all lanes are implemented in JIT by generating dup/mov/fmov instructions.

While I was there, I noticed an edge case where we hit assert if trying to emit an immediate int.MaxValue. Fixed it as well.

Contributes to #33308 and #33496.
Fixes: #35821

@Dotnet-GitSync-BotDotnet-GitSync-Bot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Apr 28, 2020
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib , @tannergooding

@BruceForstall

Copy link
Copy Markdown
Contributor

cc @TamarChristinaArm

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

not for this PR, but since we only have 8 argument registers the rest will be passed on the stack.
In which case would be easier to load the remaining 64 bits into a register directly from the stack into the top part of the 128-bit vector.

i.e. do something like this:

 and w0, w0, 255
fmov s0, w0
ins v0.b[1], w1
ins v0.b[2], w2
ins v0.b[3], w3
ins v0.b[4], w4
ins v0.b[5], w5
ins v0.b[6], w6
ins v0.b[7], w7
ld1 {v0.d}[1], [sp]

And you can do this is something like this for a lot of the cases with more than 8 arguments. Do you know what the code generates here at the moment?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's a good point which I didn't realize. I will explore your suggestion.

Generated code
 53001C00 uxtb w0, w0 4E011C10 ins v16.b[0], w0 53001C20 uxtb w0, w1 4E031C10 ins v16.b[1], w0 53001C40 uxtb w0, w2 4E051C10 ins v16.b[2], w0 53001C60 uxtb w0, w3 4E071C10 ins v16.b[3], w0 53001C80 uxtb w0, w4 4E091C10 ins v16.b[4], w0 53001CA0 uxtb w0, w5 4E0B1C10 ins v16.b[5], w0 53001CC0 uxtb w0, w6 4E0D1C10 ins v16.b[6], w0 53001CE0 uxtb w0, w7 4E0F1C10 ins v16.b[7], w0 B94023A0 ldr w0,[fp,#32] // [V08 arg8] 53001C00 uxtb w0, w0 4E111C10 ins v16.b[8], w0 B9402BA0 ldr w0,[fp,#40] // [V09 arg9] 53001C00 uxtb w0, w0 4E131C10 ins v16.b[9], w0 B94033A0 ldr w0,[fp,#48] // [V10 arg10] 53001C00 uxtb w0, w0 4E151C10 ins v16.b[10], w0 B9403BA0 ldr w0,[fp,#56] // [V11 arg11] 53001C00 uxtb w0, w0 4E171C10 ins v16.b[11], w0 B94043A0 ldr w0,[fp,#64] // [V12 arg12] 53001C00 uxtb w0, w0 4E191C10 ins v16.b[12], w0 B9404BA0 ldr w0,[fp,#72] // [V13 arg13] 53001C00 uxtb w0, w0 4E1B1C10 ins v16.b[13], w0 B94053A0 ldr w0,[fp,#80] // [V14 arg14] 53001C00 uxtb w0, w0 4E1D1C10 ins v16.b[14], w0 B9405BA0 ldr w0,[fp,#88] // [V15 arg15] 53001C00 uxtb w0, w0 4EB01E08 mov v8.16b, v16.16b 4E1F1C08 ins v8.b[15], w0 D28D0500 movz x0, #0x6828 F2A6EF60 movk x0, #0x377bLSL #16 F2CFFFA0 movk x0, #0x7ffdLSL #32 6E084509 mov v9.d[0], v8.d[1] 97FF53EA bl CORINFO_HELP_NEWSFAST 6E180528 mov v8.d[1], v9.d[0] 3C808008 str q8,[x0,#8]

And you can do this is something like this for a lot of the cases with more than 8 arguments.

Actually, this is the only API that has more than 8 arguments.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I wonder how hard is goint to be to get rid of the uxtb-s

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah, I am not much familiar with this, but would love to get rid of them. Just to call out, they show up for byte, sbyte, short and ushort parameters.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah, I am not much familiar with this, but would love to get rid of them. Just to call out, they show up for byte, sbyte, short and ushort parameters.

Correct, because these types are usually "normalized" (i.e. sign- or zero-extended) to a 32 bit value

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah, those aren't needed indeed, would be good to get rid of them if possible since they double the number of instructions.

Also

 B94053A0 ldr w0, [fp,#80] // [V14 arg14]
53001C00 uxtb w0, w0
4E1D1C10 ins v16.b[14], w0

Without the optimization I talked about above should ideally be

ld1 {v16.b[14]}, [x1]

where x1 gets incremented, or you could use a which prevents from having to move between register files. But also don't know how easy this is to do..

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I briefly discussed various options with @echesakovMSFT and concluded that it might not be straight forward. I would like to revisit it after we implement other APIs. Opened #35688 to track it.

Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/lowerarmarch.cpp Outdated
Comment threadsrc/coreclr/src/jit/emitarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated

@echesakovechesakov left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks Good to me with some questions/suggestions.

Comment threadsrc/coreclr/src/jit/emitarm64.cpp Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@kunalspathak You might want to add comment here in the same fashion as it's done for Sse2 case

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Spoke offline and this is not needed.

@kunalspathak
kunalspathakforce-pushed the create-scalar-multiple branch from 1a86bc1 to 76cc65eCompareMay 1, 2020 21:55
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@tannergooding and @echesakovMSFT - Just FYI, while working on something else I found 2 issues for which we don't have test coverage today.

  1. Vector64.Create((double)10) (any integer value casted to double)
  2. Passing Vector64<double> or Vector64<long> as parameter to a function.

I will fix and add test coverage for it before merging.

@kunalspathak
kunalspathak merged commit d23f1a2 into dotnet:masterMay 5, 2020
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Assertion failed '!"Didn't find a class handle for simdType"' when trying to pass Vector64<ulong> to function

6 participants

@kunalspathak@BruceForstall@echesakov@tannergooding@TamarChristinaArm@Dotnet-GitSync-Bot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

ARM64 intrinsic support for Vector64.Create() and Vector128.Create() - #35590

Merged
kunalspathak merged 15 commits into
dotnet:masterfrom
kunalspathak:create-scalar-multiple
May 5, 2020
Merged

ARM64 intrinsic support for Vector64.Create() and Vector128.Create()#35590
kunalspathak merged 15 commits into
dotnet:masterfrom
kunalspathak:create-scalar-multiple

Conversation

@kunalspathak

@kunalspathakkunalspathak commented Apr 28, 2020

Copy link
Copy Markdown
Contributor

Added hardware intrinsic for various overloads of Vector64.Create() and Vector128.Create():

  • Multiple arguments - The APIs that takes multiple parameters to be set in respective lanes are implemented in C# using AdvSimd.Insert.
  • Single arguments - The APIs that takes single argument and should be copied in all lanes are implemented in JIT by generating dup/mov/fmov instructions.

While I was there, I noticed an edge case where we hit assert if trying to emit an immediate int.MaxValue. Fixed it as well.

Contributes to #33308 and #33496.
Fixes: #35821

@Dotnet-GitSync-BotDotnet-GitSync-Bot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Apr 28, 2020
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib , @tannergooding

@BruceForstall

Copy link
Copy Markdown
Contributor

cc @TamarChristinaArm

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

not for this PR, but since we only have 8 argument registers the rest will be passed on the stack.
In which case would be easier to load the remaining 64 bits into a register directly from the stack into the top part of the 128-bit vector.

i.e. do something like this:

 and w0, w0, 255
fmov s0, w0
ins v0.b[1], w1
ins v0.b[2], w2
ins v0.b[3], w3
ins v0.b[4], w4
ins v0.b[5], w5
ins v0.b[6], w6
ins v0.b[7], w7
ld1 {v0.d}[1], [sp]

And you can do this is something like this for a lot of the cases with more than 8 arguments. Do you know what the code generates here at the moment?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's a good point which I didn't realize. I will explore your suggestion.

Generated code
 53001C00 uxtb w0, w0 4E011C10 ins v16.b[0], w0 53001C20 uxtb w0, w1 4E031C10 ins v16.b[1], w0 53001C40 uxtb w0, w2 4E051C10 ins v16.b[2], w0 53001C60 uxtb w0, w3 4E071C10 ins v16.b[3], w0 53001C80 uxtb w0, w4 4E091C10 ins v16.b[4], w0 53001CA0 uxtb w0, w5 4E0B1C10 ins v16.b[5], w0 53001CC0 uxtb w0, w6 4E0D1C10 ins v16.b[6], w0 53001CE0 uxtb w0, w7 4E0F1C10 ins v16.b[7], w0 B94023A0 ldr w0,[fp,#32] // [V08 arg8] 53001C00 uxtb w0, w0 4E111C10 ins v16.b[8], w0 B9402BA0 ldr w0,[fp,#40] // [V09 arg9] 53001C00 uxtb w0, w0 4E131C10 ins v16.b[9], w0 B94033A0 ldr w0,[fp,#48] // [V10 arg10] 53001C00 uxtb w0, w0 4E151C10 ins v16.b[10], w0 B9403BA0 ldr w0,[fp,#56] // [V11 arg11] 53001C00 uxtb w0, w0 4E171C10 ins v16.b[11], w0 B94043A0 ldr w0,[fp,#64] // [V12 arg12] 53001C00 uxtb w0, w0 4E191C10 ins v16.b[12], w0 B9404BA0 ldr w0,[fp,#72] // [V13 arg13] 53001C00 uxtb w0, w0 4E1B1C10 ins v16.b[13], w0 B94053A0 ldr w0,[fp,#80] // [V14 arg14] 53001C00 uxtb w0, w0 4E1D1C10 ins v16.b[14], w0 B9405BA0 ldr w0,[fp,#88] // [V15 arg15] 53001C00 uxtb w0, w0 4EB01E08 mov v8.16b, v16.16b 4E1F1C08 ins v8.b[15], w0 D28D0500 movz x0, #0x6828 F2A6EF60 movk x0, #0x377bLSL #16 F2CFFFA0 movk x0, #0x7ffdLSL #32 6E084509 mov v9.d[0], v8.d[1] 97FF53EA bl CORINFO_HELP_NEWSFAST 6E180528 mov v8.d[1], v9.d[0] 3C808008 str q8,[x0,#8]

And you can do this is something like this for a lot of the cases with more than 8 arguments.

Actually, this is the only API that has more than 8 arguments.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I wonder how hard is goint to be to get rid of the uxtb-s

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah, I am not much familiar with this, but would love to get rid of them. Just to call out, they show up for byte, sbyte, short and ushort parameters.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah, I am not much familiar with this, but would love to get rid of them. Just to call out, they show up for byte, sbyte, short and ushort parameters.

Correct, because these types are usually "normalized" (i.e. sign- or zero-extended) to a 32 bit value

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah, those aren't needed indeed, would be good to get rid of them if possible since they double the number of instructions.

Also

 B94053A0 ldr w0, [fp,#80] // [V14 arg14]
53001C00 uxtb w0, w0
4E1D1C10 ins v16.b[14], w0

Without the optimization I talked about above should ideally be

ld1 {v16.b[14]}, [x1]

where x1 gets incremented, or you could use a which prevents from having to move between register files. But also don't know how easy this is to do..

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I briefly discussed various options with @echesakovMSFT and concluded that it might not be straight forward. I would like to revisit it after we implement other APIs. Opened #35688 to track it.

Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/lowerarmarch.cpp Outdated
Comment threadsrc/coreclr/src/jit/emitarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated
Comment threadsrc/coreclr/src/jit/hwintrinsiccodegenarm64.cpp Outdated

@echesakovechesakov left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks Good to me with some questions/suggestions.

Comment threadsrc/coreclr/src/jit/emitarm64.cpp Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@kunalspathak You might want to add comment here in the same fashion as it's done for Sse2 case

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Spoke offline and this is not needed.

@kunalspathak
kunalspathakforce-pushed the create-scalar-multiple branch from 1a86bc1 to 76cc65eCompareMay 1, 2020 21:55
@kunalspathak

Copy link
Copy Markdown
ContributorAuthor

@tannergooding and @echesakovMSFT - Just FYI, while working on something else I found 2 issues for which we don't have test coverage today.

  1. Vector64.Create((double)10) (any integer value casted to double)
  2. Passing Vector64<double> or Vector64<long> as parameter to a function.

I will fix and add test coverage for it before merging.

@kunalspathak
kunalspathak merged commit d23f1a2 into dotnet:masterMay 5, 2020
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Assertion failed '!"Didn't find a class handle for simdType"' when trying to pass Vector64<ulong> to function

6 participants

@kunalspathak@BruceForstall@echesakov@tannergooding@TamarChristinaArm@Dotnet-GitSync-Bot