[Arm64] Implement Vector64/128.CreateScalar() using AdvSimd.Insert - #35300

Merged
echesakov merged 2 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Vector64-Vector128-CreateScalar-Use-AdvSimd-Insert
Apr 30, 2020
Merged

[Arm64] Implement Vector64/128.CreateScalar() using AdvSimd.Insert#35300
echesakov merged 2 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Vector64-Vector128-CreateScalar-Use-AdvSimd-Insert

Conversation

@echesakov

Copy link
Copy Markdown
Contributor

No description provided.

@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @tannergooding
Notify danmosemsft if you want to be subscribed.

{
if (AdvSimd.IsSupported)
{
return AdvSimd.Insert(Vector128<byte>.Zero, 0, value);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We'll need to special-case CreateScalarUnsafe since the upper bits don't have to be zeroed for it.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep, this work is tracked by #34485

@echesakov

Copy link
Copy Markdown
ContributorAuthor

I still would like to collect the jisDisasm-s for this - I remember seeing something weird for Vector64.CreateScalar() - will do it later.

@echesakov

echesakov commented Apr 30, 2020

Copy link
Copy Markdown
ContributorAuthor
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector128`1[Byte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ubyte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M19699_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M19699_IG02: 4E010FF0 dup v16.16b, wzr 53001C00 uxtb w0, w0 4E011C10 ins v16.b[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M19699_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=a381b30c) for method System.Runtime.Intrinsics.Vector128:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector128`1[Byte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(double):System.Runtime.Intrinsics.Vector128`1[Double]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) double -> d0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M6886_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M6886_IG02: 4E080FF0 dup v16.2d, xzr 6E080410 ins v16.d[0], v0.d[0] 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M6886_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=17a3e519) for method System.Runtime.Intrinsics.Vector128:CreateScalar(double):System.Runtime.Intrinsics.Vector128`1[Double]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(short):System.Runtime.Intrinsics.Vector128`1[Int16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) short -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M37120_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M37120_IG02: 4E020FF0 dup v16.8h, wzr 13003C00 sxth w0, w0 4E021C10 ins v16.h[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M37120_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=87f66eff) for method System.Runtime.Intrinsics.Vector128:CreateScalar(short):System.Runtime.Intrinsics.Vector128`1[Int16]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[Int32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M42503_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M42503_IG02: 4E040FF0 dup v16.4s, wzr 4E041C10 ins v16.s[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M42503_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=56ed59f8) for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[Int32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[Int64]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) long -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M9853_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M9853_IG02: 4E080FF0 dup v16.2d, xzr 4E081C10 ins v16.d[0], x0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M9853_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=99c7d982) for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[Int64]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(byte):System.Runtime.Intrinsics.Vector128`1[SByte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) byte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M64149_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M64149_IG02: 4E010FF0 dup v16.16b, wzr 13001C00 sxtb w0, w0 4E011C10 ins v16.b[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M64149_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=3757056a) for method System.Runtime.Intrinsics.Vector128:CreateScalar(byte):System.Runtime.Intrinsics.Vector128`1[SByte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(float):System.Runtime.Intrinsics.Vector128`1[Single]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) float -> d0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M28940_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M28940_IG02: 4E040FF0 dup v16.4s, wzr 6E040410 ins v16.s[0], v0.s[0] 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M28940_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=910d8ef3) for method System.Runtime.Intrinsics.Vector128:CreateScalar(float):System.Runtime.Intrinsics.Vector128`1[Single]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(ushort):System.Runtime.Intrinsics.Vector128`1[UInt16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ushort -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M480_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M480_IG02: 4E020FF0 dup v16.8h, wzr 53003C00 uxth w0, w0 4E021C10 ins v16.h[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M480_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=5a77fe1f) for method System.Runtime.Intrinsics.Vector128:CreateScalar(ushort):System.Runtime.Intrinsics.Vector128`1[UInt16]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[UInt32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M21746_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M21746_IG02: 4E040FF0 dup v16.4s, wzr 4E041C10 ins v16.s[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M21746_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=4a35ab0d) for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[UInt32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[UInt64]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) long -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M2664_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M2664_IG02: 4E080FF0 dup v16.2d, xzr 4E081C10 ins v16.d[0], x0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M2664_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=714bf597) for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[UInt64]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[Int32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M25863_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M25863_IG02: 0E040FE0 dup v0.2s, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 4E041C00 ins v0.s[0], w0 ;; bbWeight=1 PerfScore 6.00G_M25863_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 12.70, (MethodHash=80a89af8) for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[Int32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(byte):System.Runtime.Intrinsics.Vector64`1[SByte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) byte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M12309_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M12309_IG02: 0E010FE0 dup v0.8b, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 13001C00 sxtb w0, w0 4E011C00 ins v0.b[0], w0 ;; bbWeight=1 PerfScore 6.50G_M12309_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=1802cfea) for method System.Runtime.Intrinsics.Vector64:CreateScalar(byte):System.Runtime.Intrinsics.Vector64`1[SByte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(float):System.Runtime.Intrinsics.Vector64`1[Single]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) float -> d0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d16 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M44268_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M44268_IG02: 0E040FF0 dup v16.2s, wzr FD000FB0 str d16,[fp,#24] FD400FB0 ldr d16,[fp,#24] 6E040410 ins v16.s[0], v0.s[0] 1E604200 fmov d0, d16 ;; bbWeight=1 PerfScore 6.50G_M44268_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=b5c65313) for method System.Runtime.Intrinsics.Vector64:CreateScalar(float):System.Runtime.Intrinsics.Vector64`1[Single]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(ushort):System.Runtime.Intrinsics.Vector64`1[UInt16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ushort -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M37504_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M37504_IG02: 0E020FE0 dup v0.4h, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 53003C00 uxth w0, w0 4E021C00 ins v0.h[0], w0 ;; bbWeight=1 PerfScore 6.50G_M37504_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=68536d7f) for method System.Runtime.Intrinsics.Vector64:CreateScalar(ushort):System.Runtime.Intrinsics.Vector64`1[UInt16]; ============================================================

Collected JIT disassemblies with the changes rebased on top of latest master

; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[UInt32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M62450_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M62450_IG02: 0E040FE0 dup v0.2s, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 4E041C00 ins v0.s[0], w0 ;; bbWeight=1 PerfScore 6.00G_M62450_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 12.70, (MethodHash=ab590c0d) for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[UInt32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector64`1[Byte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ubyte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M20083_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M20083_IG02: 0E010FE0 dup v0.8b, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 53001C00 uxtb w0, w0 4E011C00 ins v0.b[0], w0 ;; bbWeight=1 PerfScore 6.50G_M20083_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=cedeb18c) for method System.Runtime.Intrinsics.Vector64:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector64`1[Byte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(short):System.Runtime.Intrinsics.Vector64`1[Int16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) short -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M58336_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M58336_IG02: 0E020FE0 dup v0.4h, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 13003C00 sxth w0, w0 4E021C00 ins v0.h[0], w0 ;; bbWeight=1 PerfScore 6.50G_M58336_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=e95c1c1f) for method System.Runtime.Intrinsics.Vector64:CreateScalar(short):System.Runtime.Intrinsics.Vector64`1[Int16]; ============================================================

There are multiple issues here:

  1. Redundant str/ldr-s with a SIMD register - this appears only in Vector64.CreateScalar():
str d0,[fp,#24]ldr d0,[fp,#24]

The code is the worst for Vector64<float>.CreateScalar()

 0E040FF0 dup v16.2s, wzr FD000FB0 str d16,[fp,#24] FD400FB0 ldr d16,[fp,#24] 6E040410 ins v16.s[0], v0.s[0] 1E604200 fmov d0, d16

or Vector64<ushort>.CreateScalar()

 0E020FE0 dup v0.4h, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 53003C00 uxth w0, w0 4E021C00 ins v0.h[0], w0
  1. Unnecessary sign-/zero-extensions with byte,ubyte,short,ushort (the same as seen in ARM64 intrinsic support for Vector64.Create() and Vector128.Create() #35590):
uxtb w0, w0uxth w0, w0sxtb w0, w0sxth w0, w0
  1. dup Vd.T, wzr seems to be used for code generation of Vector64/128.Zero which I thought was fixed with Implement Vector{Size}<T>.AllBitsSet #33924 (cc @Gnbrkm41). I will follow up on this

cc @kunalspathak@BruceForstall

@echesakov
echesakov marked this pull request as ready for review April 30, 2020 19:35
return AdvSimd.Insert(Vector64<byte>.Zero, 0, value);
}

return SoftwareFallback(value);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Curious about wrapping the SoftwareFallback() in a static method. What is the advantage of doing it?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure, actually. I did this to be consistent with existing Vector128/256 implementations.
@tannergooding Do you know why?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With the logic for all the various paths the method is too large for the normal inlining heuristics to work so we need to mark it AggressiveInlining (since the accelerated paths will generally be pretty small).
However, we don't necessarily want the SoftwareFallback to be inlined as that may not be beneficial.
Putting it in its own method prevents it from being inlined in the normal case and allows the JIT to decide if it is "too large or not" by itself.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That make sense.

@kunalspathakkunalspathak left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@echesakov
echesakov merged commit 670bf21 into dotnet:masterApr 30, 2020
@echesakov
echesakov deleted the Arm64-ASIMD-Vector64-Vector128-CreateScalar-Use-AdvSimd-Insert branch April 30, 2020 23:35
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@echesakov@tannergooding@kunalspathak
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

[Arm64] Implement Vector64/128.CreateScalar() using AdvSimd.Insert - #35300

Merged
echesakov merged 2 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Vector64-Vector128-CreateScalar-Use-AdvSimd-Insert
Apr 30, 2020
Merged

[Arm64] Implement Vector64/128.CreateScalar() using AdvSimd.Insert#35300
echesakov merged 2 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Vector64-Vector128-CreateScalar-Use-AdvSimd-Insert

Conversation

@echesakov

Copy link
Copy Markdown
Contributor

No description provided.

@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @tannergooding
Notify danmosemsft if you want to be subscribed.

{
if (AdvSimd.IsSupported)
{
return AdvSimd.Insert(Vector128<byte>.Zero, 0, value);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We'll need to special-case CreateScalarUnsafe since the upper bits don't have to be zeroed for it.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep, this work is tracked by #34485

@echesakov

Copy link
Copy Markdown
ContributorAuthor

I still would like to collect the jisDisasm-s for this - I remember seeing something weird for Vector64.CreateScalar() - will do it later.

@echesakov

echesakov commented Apr 30, 2020

Copy link
Copy Markdown
ContributorAuthor
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector128`1[Byte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ubyte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M19699_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M19699_IG02: 4E010FF0 dup v16.16b, wzr 53001C00 uxtb w0, w0 4E011C10 ins v16.b[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M19699_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=a381b30c) for method System.Runtime.Intrinsics.Vector128:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector128`1[Byte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(double):System.Runtime.Intrinsics.Vector128`1[Double]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) double -> d0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M6886_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M6886_IG02: 4E080FF0 dup v16.2d, xzr 6E080410 ins v16.d[0], v0.d[0] 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M6886_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=17a3e519) for method System.Runtime.Intrinsics.Vector128:CreateScalar(double):System.Runtime.Intrinsics.Vector128`1[Double]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(short):System.Runtime.Intrinsics.Vector128`1[Int16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) short -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M37120_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M37120_IG02: 4E020FF0 dup v16.8h, wzr 13003C00 sxth w0, w0 4E021C10 ins v16.h[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M37120_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=87f66eff) for method System.Runtime.Intrinsics.Vector128:CreateScalar(short):System.Runtime.Intrinsics.Vector128`1[Int16]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[Int32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M42503_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M42503_IG02: 4E040FF0 dup v16.4s, wzr 4E041C10 ins v16.s[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M42503_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=56ed59f8) for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[Int32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[Int64]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) long -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M9853_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M9853_IG02: 4E080FF0 dup v16.2d, xzr 4E081C10 ins v16.d[0], x0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M9853_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=99c7d982) for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[Int64]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(byte):System.Runtime.Intrinsics.Vector128`1[SByte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) byte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M64149_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M64149_IG02: 4E010FF0 dup v16.16b, wzr 13001C00 sxtb w0, w0 4E011C10 ins v16.b[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M64149_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=3757056a) for method System.Runtime.Intrinsics.Vector128:CreateScalar(byte):System.Runtime.Intrinsics.Vector128`1[SByte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(float):System.Runtime.Intrinsics.Vector128`1[Single]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) float -> d0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M28940_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M28940_IG02: 4E040FF0 dup v16.4s, wzr 6E040410 ins v16.s[0], v0.s[0] 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M28940_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=910d8ef3) for method System.Runtime.Intrinsics.Vector128:CreateScalar(float):System.Runtime.Intrinsics.Vector128`1[Single]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(ushort):System.Runtime.Intrinsics.Vector128`1[UInt16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ushort -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M480_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M480_IG02: 4E020FF0 dup v16.8h, wzr 53003C00 uxth w0, w0 4E021C10 ins v16.h[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M480_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=5a77fe1f) for method System.Runtime.Intrinsics.Vector128:CreateScalar(ushort):System.Runtime.Intrinsics.Vector128`1[UInt16]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[UInt32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M21746_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M21746_IG02: 4E040FF0 dup v16.4s, wzr 4E041C10 ins v16.s[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M21746_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=4a35ab0d) for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[UInt32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[UInt64]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) long -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M2664_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M2664_IG02: 4E080FF0 dup v16.2d, xzr 4E081C10 ins v16.d[0], x0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M2664_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=714bf597) for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[UInt64]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[Int32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M25863_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M25863_IG02: 0E040FE0 dup v0.2s, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 4E041C00 ins v0.s[0], w0 ;; bbWeight=1 PerfScore 6.00G_M25863_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 12.70, (MethodHash=80a89af8) for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[Int32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(byte):System.Runtime.Intrinsics.Vector64`1[SByte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) byte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M12309_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M12309_IG02: 0E010FE0 dup v0.8b, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 13001C00 sxtb w0, w0 4E011C00 ins v0.b[0], w0 ;; bbWeight=1 PerfScore 6.50G_M12309_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=1802cfea) for method System.Runtime.Intrinsics.Vector64:CreateScalar(byte):System.Runtime.Intrinsics.Vector64`1[SByte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(float):System.Runtime.Intrinsics.Vector64`1[Single]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) float -> d0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d16 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M44268_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M44268_IG02: 0E040FF0 dup v16.2s, wzr FD000FB0 str d16,[fp,#24] FD400FB0 ldr d16,[fp,#24] 6E040410 ins v16.s[0], v0.s[0] 1E604200 fmov d0, d16 ;; bbWeight=1 PerfScore 6.50G_M44268_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=b5c65313) for method System.Runtime.Intrinsics.Vector64:CreateScalar(float):System.Runtime.Intrinsics.Vector64`1[Single]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(ushort):System.Runtime.Intrinsics.Vector64`1[UInt16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ushort -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M37504_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M37504_IG02: 0E020FE0 dup v0.4h, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 53003C00 uxth w0, w0 4E021C00 ins v0.h[0], w0 ;; bbWeight=1 PerfScore 6.50G_M37504_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=68536d7f) for method System.Runtime.Intrinsics.Vector64:CreateScalar(ushort):System.Runtime.Intrinsics.Vector64`1[UInt16]; ============================================================

Collected JIT disassemblies with the changes rebased on top of latest master

; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[UInt32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M62450_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M62450_IG02: 0E040FE0 dup v0.2s, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 4E041C00 ins v0.s[0], w0 ;; bbWeight=1 PerfScore 6.00G_M62450_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 12.70, (MethodHash=ab590c0d) for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[UInt32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector64`1[Byte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ubyte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M20083_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M20083_IG02: 0E010FE0 dup v0.8b, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 53001C00 uxtb w0, w0 4E011C00 ins v0.b[0], w0 ;; bbWeight=1 PerfScore 6.50G_M20083_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=cedeb18c) for method System.Runtime.Intrinsics.Vector64:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector64`1[Byte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(short):System.Runtime.Intrinsics.Vector64`1[Int16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) short -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M58336_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M58336_IG02: 0E020FE0 dup v0.4h, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 13003C00 sxth w0, w0 4E021C00 ins v0.h[0], w0 ;; bbWeight=1 PerfScore 6.50G_M58336_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=e95c1c1f) for method System.Runtime.Intrinsics.Vector64:CreateScalar(short):System.Runtime.Intrinsics.Vector64`1[Int16]; ============================================================

There are multiple issues here:

  1. Redundant str/ldr-s with a SIMD register - this appears only in Vector64.CreateScalar():
str d0,[fp,#24]ldr d0,[fp,#24]

The code is the worst for Vector64<float>.CreateScalar()

 0E040FF0 dup v16.2s, wzr FD000FB0 str d16,[fp,#24] FD400FB0 ldr d16,[fp,#24] 6E040410 ins v16.s[0], v0.s[0] 1E604200 fmov d0, d16

or Vector64<ushort>.CreateScalar()

 0E020FE0 dup v0.4h, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 53003C00 uxth w0, w0 4E021C00 ins v0.h[0], w0
  1. Unnecessary sign-/zero-extensions with byte,ubyte,short,ushort (the same as seen in ARM64 intrinsic support for Vector64.Create() and Vector128.Create() #35590):
uxtb w0, w0uxth w0, w0sxtb w0, w0sxth w0, w0
  1. dup Vd.T, wzr seems to be used for code generation of Vector64/128.Zero which I thought was fixed with Implement Vector{Size}<T>.AllBitsSet #33924 (cc @Gnbrkm41). I will follow up on this

cc @kunalspathak@BruceForstall

@echesakov
echesakov marked this pull request as ready for review April 30, 2020 19:35
return AdvSimd.Insert(Vector64<byte>.Zero, 0, value);
}

return SoftwareFallback(value);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Curious about wrapping the SoftwareFallback() in a static method. What is the advantage of doing it?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure, actually. I did this to be consistent with existing Vector128/256 implementations.
@tannergooding Do you know why?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With the logic for all the various paths the method is too large for the normal inlining heuristics to work so we need to mark it AggressiveInlining (since the accelerated paths will generally be pretty small).
However, we don't necessarily want the SoftwareFallback to be inlined as that may not be beneficial.
Putting it in its own method prevents it from being inlined in the normal case and allows the JIT to decide if it is "too large or not" by itself.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That make sense.

@kunalspathakkunalspathak left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@echesakov
echesakov merged commit 670bf21 into dotnet:masterApr 30, 2020
@echesakov
echesakov deleted the Arm64-ASIMD-Vector64-Vector128-CreateScalar-Use-AdvSimd-Insert branch April 30, 2020 23:35
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@echesakov@tannergooding@kunalspathak
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[Arm64] Implement Vector64/128.CreateScalar() using AdvSimd.Insert - #35300

Merged
echesakov merged 2 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Vector64-Vector128-CreateScalar-Use-AdvSimd-Insert
Apr 30, 2020
Merged

[Arm64] Implement Vector64/128.CreateScalar() using AdvSimd.Insert#35300
echesakov merged 2 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Vector64-Vector128-CreateScalar-Use-AdvSimd-Insert

Conversation

@echesakov

Copy link
Copy Markdown
Contributor

No description provided.

@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @tannergooding
Notify danmosemsft if you want to be subscribed.

{
if (AdvSimd.IsSupported)
{
return AdvSimd.Insert(Vector128<byte>.Zero, 0, value);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We'll need to special-case CreateScalarUnsafe since the upper bits don't have to be zeroed for it.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep, this work is tracked by #34485

@echesakov

Copy link
Copy Markdown
ContributorAuthor

I still would like to collect the jisDisasm-s for this - I remember seeing something weird for Vector64.CreateScalar() - will do it later.

@echesakov

echesakov commented Apr 30, 2020

Copy link
Copy Markdown
ContributorAuthor
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector128`1[Byte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ubyte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M19699_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M19699_IG02: 4E010FF0 dup v16.16b, wzr 53001C00 uxtb w0, w0 4E011C10 ins v16.b[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M19699_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=a381b30c) for method System.Runtime.Intrinsics.Vector128:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector128`1[Byte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(double):System.Runtime.Intrinsics.Vector128`1[Double]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) double -> d0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M6886_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M6886_IG02: 4E080FF0 dup v16.2d, xzr 6E080410 ins v16.d[0], v0.d[0] 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M6886_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=17a3e519) for method System.Runtime.Intrinsics.Vector128:CreateScalar(double):System.Runtime.Intrinsics.Vector128`1[Double]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(short):System.Runtime.Intrinsics.Vector128`1[Int16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) short -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M37120_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M37120_IG02: 4E020FF0 dup v16.8h, wzr 13003C00 sxth w0, w0 4E021C10 ins v16.h[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M37120_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=87f66eff) for method System.Runtime.Intrinsics.Vector128:CreateScalar(short):System.Runtime.Intrinsics.Vector128`1[Int16]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[Int32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M42503_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M42503_IG02: 4E040FF0 dup v16.4s, wzr 4E041C10 ins v16.s[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M42503_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=56ed59f8) for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[Int32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[Int64]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) long -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M9853_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M9853_IG02: 4E080FF0 dup v16.2d, xzr 4E081C10 ins v16.d[0], x0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M9853_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=99c7d982) for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[Int64]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(byte):System.Runtime.Intrinsics.Vector128`1[SByte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) byte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M64149_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M64149_IG02: 4E010FF0 dup v16.16b, wzr 13001C00 sxtb w0, w0 4E011C10 ins v16.b[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M64149_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=3757056a) for method System.Runtime.Intrinsics.Vector128:CreateScalar(byte):System.Runtime.Intrinsics.Vector128`1[SByte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(float):System.Runtime.Intrinsics.Vector128`1[Single]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) float -> d0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M28940_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M28940_IG02: 4E040FF0 dup v16.4s, wzr 6E040410 ins v16.s[0], v0.s[0] 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M28940_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=910d8ef3) for method System.Runtime.Intrinsics.Vector128:CreateScalar(float):System.Runtime.Intrinsics.Vector128`1[Single]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(ushort):System.Runtime.Intrinsics.Vector128`1[UInt16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ushort -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M480_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M480_IG02: 4E020FF0 dup v16.8h, wzr 53003C00 uxth w0, w0 4E021C10 ins v16.h[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M480_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=5a77fe1f) for method System.Runtime.Intrinsics.Vector128:CreateScalar(ushort):System.Runtime.Intrinsics.Vector128`1[UInt16]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[UInt32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M21746_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M21746_IG02: 4E040FF0 dup v16.4s, wzr 4E041C10 ins v16.s[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M21746_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=4a35ab0d) for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[UInt32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[UInt64]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) long -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M2664_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M2664_IG02: 4E080FF0 dup v16.2d, xzr 4E081C10 ins v16.d[0], x0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M2664_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=714bf597) for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[UInt64]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[Int32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M25863_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M25863_IG02: 0E040FE0 dup v0.2s, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 4E041C00 ins v0.s[0], w0 ;; bbWeight=1 PerfScore 6.00G_M25863_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 12.70, (MethodHash=80a89af8) for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[Int32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(byte):System.Runtime.Intrinsics.Vector64`1[SByte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) byte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M12309_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M12309_IG02: 0E010FE0 dup v0.8b, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 13001C00 sxtb w0, w0 4E011C00 ins v0.b[0], w0 ;; bbWeight=1 PerfScore 6.50G_M12309_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=1802cfea) for method System.Runtime.Intrinsics.Vector64:CreateScalar(byte):System.Runtime.Intrinsics.Vector64`1[SByte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(float):System.Runtime.Intrinsics.Vector64`1[Single]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) float -> d0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d16 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M44268_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M44268_IG02: 0E040FF0 dup v16.2s, wzr FD000FB0 str d16,[fp,#24] FD400FB0 ldr d16,[fp,#24] 6E040410 ins v16.s[0], v0.s[0] 1E604200 fmov d0, d16 ;; bbWeight=1 PerfScore 6.50G_M44268_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=b5c65313) for method System.Runtime.Intrinsics.Vector64:CreateScalar(float):System.Runtime.Intrinsics.Vector64`1[Single]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(ushort):System.Runtime.Intrinsics.Vector64`1[UInt16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ushort -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M37504_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M37504_IG02: 0E020FE0 dup v0.4h, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 53003C00 uxth w0, w0 4E021C00 ins v0.h[0], w0 ;; bbWeight=1 PerfScore 6.50G_M37504_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=68536d7f) for method System.Runtime.Intrinsics.Vector64:CreateScalar(ushort):System.Runtime.Intrinsics.Vector64`1[UInt16]; ============================================================

Collected JIT disassemblies with the changes rebased on top of latest master

; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[UInt32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M62450_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M62450_IG02: 0E040FE0 dup v0.2s, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 4E041C00 ins v0.s[0], w0 ;; bbWeight=1 PerfScore 6.00G_M62450_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 12.70, (MethodHash=ab590c0d) for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[UInt32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector64`1[Byte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ubyte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M20083_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M20083_IG02: 0E010FE0 dup v0.8b, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 53001C00 uxtb w0, w0 4E011C00 ins v0.b[0], w0 ;; bbWeight=1 PerfScore 6.50G_M20083_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=cedeb18c) for method System.Runtime.Intrinsics.Vector64:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector64`1[Byte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(short):System.Runtime.Intrinsics.Vector64`1[Int16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) short -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M58336_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M58336_IG02: 0E020FE0 dup v0.4h, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 13003C00 sxth w0, w0 4E021C00 ins v0.h[0], w0 ;; bbWeight=1 PerfScore 6.50G_M58336_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=e95c1c1f) for method System.Runtime.Intrinsics.Vector64:CreateScalar(short):System.Runtime.Intrinsics.Vector64`1[Int16]; ============================================================

There are multiple issues here:

  1. Redundant str/ldr-s with a SIMD register - this appears only in Vector64.CreateScalar():
str d0,[fp,#24]ldr d0,[fp,#24]

The code is the worst for Vector64<float>.CreateScalar()

 0E040FF0 dup v16.2s, wzr FD000FB0 str d16,[fp,#24] FD400FB0 ldr d16,[fp,#24] 6E040410 ins v16.s[0], v0.s[0] 1E604200 fmov d0, d16

or Vector64<ushort>.CreateScalar()

 0E020FE0 dup v0.4h, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 53003C00 uxth w0, w0 4E021C00 ins v0.h[0], w0
  1. Unnecessary sign-/zero-extensions with byte,ubyte,short,ushort (the same as seen in ARM64 intrinsic support for Vector64.Create() and Vector128.Create() #35590):
uxtb w0, w0uxth w0, w0sxtb w0, w0sxth w0, w0
  1. dup Vd.T, wzr seems to be used for code generation of Vector64/128.Zero which I thought was fixed with Implement Vector{Size}<T>.AllBitsSet #33924 (cc @Gnbrkm41). I will follow up on this

cc @kunalspathak@BruceForstall

@echesakov
echesakov marked this pull request as ready for review April 30, 2020 19:35
return AdvSimd.Insert(Vector64<byte>.Zero, 0, value);
}

return SoftwareFallback(value);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Curious about wrapping the SoftwareFallback() in a static method. What is the advantage of doing it?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure, actually. I did this to be consistent with existing Vector128/256 implementations.
@tannergooding Do you know why?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With the logic for all the various paths the method is too large for the normal inlining heuristics to work so we need to mark it AggressiveInlining (since the accelerated paths will generally be pretty small).
However, we don't necessarily want the SoftwareFallback to be inlined as that may not be beneficial.
Putting it in its own method prevents it from being inlined in the normal case and allows the JIT to decide if it is "too large or not" by itself.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That make sense.

@kunalspathakkunalspathak left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@echesakov
echesakov merged commit 670bf21 into dotnet:masterApr 30, 2020
@echesakov
echesakov deleted the Arm64-ASIMD-Vector64-Vector128-CreateScalar-Use-AdvSimd-Insert branch April 30, 2020 23:35
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@echesakov@tannergooding@kunalspathak
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[Arm64] Implement Vector64/128.CreateScalar() using AdvSimd.Insert - #35300

Merged
echesakov merged 2 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Vector64-Vector128-CreateScalar-Use-AdvSimd-Insert
Apr 30, 2020
Merged

[Arm64] Implement Vector64/128.CreateScalar() using AdvSimd.Insert#35300
echesakov merged 2 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Vector64-Vector128-CreateScalar-Use-AdvSimd-Insert

Conversation

@echesakov

Copy link
Copy Markdown
Contributor

No description provided.

@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @tannergooding
Notify danmosemsft if you want to be subscribed.

{
if (AdvSimd.IsSupported)
{
return AdvSimd.Insert(Vector128<byte>.Zero, 0, value);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We'll need to special-case CreateScalarUnsafe since the upper bits don't have to be zeroed for it.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep, this work is tracked by #34485

@echesakov

Copy link
Copy Markdown
ContributorAuthor

I still would like to collect the jisDisasm-s for this - I remember seeing something weird for Vector64.CreateScalar() - will do it later.

@echesakov

echesakov commented Apr 30, 2020

Copy link
Copy Markdown
ContributorAuthor
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector128`1[Byte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ubyte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M19699_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M19699_IG02: 4E010FF0 dup v16.16b, wzr 53001C00 uxtb w0, w0 4E011C10 ins v16.b[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M19699_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=a381b30c) for method System.Runtime.Intrinsics.Vector128:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector128`1[Byte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(double):System.Runtime.Intrinsics.Vector128`1[Double]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) double -> d0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M6886_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M6886_IG02: 4E080FF0 dup v16.2d, xzr 6E080410 ins v16.d[0], v0.d[0] 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M6886_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=17a3e519) for method System.Runtime.Intrinsics.Vector128:CreateScalar(double):System.Runtime.Intrinsics.Vector128`1[Double]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(short):System.Runtime.Intrinsics.Vector128`1[Int16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) short -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M37120_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M37120_IG02: 4E020FF0 dup v16.8h, wzr 13003C00 sxth w0, w0 4E021C10 ins v16.h[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M37120_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=87f66eff) for method System.Runtime.Intrinsics.Vector128:CreateScalar(short):System.Runtime.Intrinsics.Vector128`1[Int16]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[Int32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M42503_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M42503_IG02: 4E040FF0 dup v16.4s, wzr 4E041C10 ins v16.s[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M42503_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=56ed59f8) for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[Int32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[Int64]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) long -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M9853_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M9853_IG02: 4E080FF0 dup v16.2d, xzr 4E081C10 ins v16.d[0], x0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M9853_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=99c7d982) for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[Int64]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(byte):System.Runtime.Intrinsics.Vector128`1[SByte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) byte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M64149_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M64149_IG02: 4E010FF0 dup v16.16b, wzr 13001C00 sxtb w0, w0 4E011C10 ins v16.b[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M64149_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=3757056a) for method System.Runtime.Intrinsics.Vector128:CreateScalar(byte):System.Runtime.Intrinsics.Vector128`1[SByte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(float):System.Runtime.Intrinsics.Vector128`1[Single]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) float -> d0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M28940_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M28940_IG02: 4E040FF0 dup v16.4s, wzr 6E040410 ins v16.s[0], v0.s[0] 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M28940_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=910d8ef3) for method System.Runtime.Intrinsics.Vector128:CreateScalar(float):System.Runtime.Intrinsics.Vector128`1[Single]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(ushort):System.Runtime.Intrinsics.Vector128`1[UInt16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ushort -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M480_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M480_IG02: 4E020FF0 dup v16.8h, wzr 53003C00 uxth w0, w0 4E021C10 ins v16.h[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M480_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=5a77fe1f) for method System.Runtime.Intrinsics.Vector128:CreateScalar(ushort):System.Runtime.Intrinsics.Vector128`1[UInt16]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[UInt32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M21746_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M21746_IG02: 4E040FF0 dup v16.4s, wzr 4E041C10 ins v16.s[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M21746_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=4a35ab0d) for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[UInt32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[UInt64]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) long -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M2664_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M2664_IG02: 4E080FF0 dup v16.2d, xzr 4E081C10 ins v16.d[0], x0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M2664_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=714bf597) for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[UInt64]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[Int32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M25863_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M25863_IG02: 0E040FE0 dup v0.2s, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 4E041C00 ins v0.s[0], w0 ;; bbWeight=1 PerfScore 6.00G_M25863_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 12.70, (MethodHash=80a89af8) for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[Int32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(byte):System.Runtime.Intrinsics.Vector64`1[SByte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) byte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M12309_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M12309_IG02: 0E010FE0 dup v0.8b, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 13001C00 sxtb w0, w0 4E011C00 ins v0.b[0], w0 ;; bbWeight=1 PerfScore 6.50G_M12309_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=1802cfea) for method System.Runtime.Intrinsics.Vector64:CreateScalar(byte):System.Runtime.Intrinsics.Vector64`1[SByte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(float):System.Runtime.Intrinsics.Vector64`1[Single]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) float -> d0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d16 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M44268_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M44268_IG02: 0E040FF0 dup v16.2s, wzr FD000FB0 str d16,[fp,#24] FD400FB0 ldr d16,[fp,#24] 6E040410 ins v16.s[0], v0.s[0] 1E604200 fmov d0, d16 ;; bbWeight=1 PerfScore 6.50G_M44268_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=b5c65313) for method System.Runtime.Intrinsics.Vector64:CreateScalar(float):System.Runtime.Intrinsics.Vector64`1[Single]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(ushort):System.Runtime.Intrinsics.Vector64`1[UInt16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ushort -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M37504_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M37504_IG02: 0E020FE0 dup v0.4h, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 53003C00 uxth w0, w0 4E021C00 ins v0.h[0], w0 ;; bbWeight=1 PerfScore 6.50G_M37504_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=68536d7f) for method System.Runtime.Intrinsics.Vector64:CreateScalar(ushort):System.Runtime.Intrinsics.Vector64`1[UInt16]; ============================================================

Collected JIT disassemblies with the changes rebased on top of latest master

; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[UInt32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M62450_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M62450_IG02: 0E040FE0 dup v0.2s, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 4E041C00 ins v0.s[0], w0 ;; bbWeight=1 PerfScore 6.00G_M62450_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 12.70, (MethodHash=ab590c0d) for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[UInt32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector64`1[Byte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ubyte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M20083_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M20083_IG02: 0E010FE0 dup v0.8b, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 53001C00 uxtb w0, w0 4E011C00 ins v0.b[0], w0 ;; bbWeight=1 PerfScore 6.50G_M20083_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=cedeb18c) for method System.Runtime.Intrinsics.Vector64:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector64`1[Byte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(short):System.Runtime.Intrinsics.Vector64`1[Int16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) short -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M58336_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M58336_IG02: 0E020FE0 dup v0.4h, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 13003C00 sxth w0, w0 4E021C00 ins v0.h[0], w0 ;; bbWeight=1 PerfScore 6.50G_M58336_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=e95c1c1f) for method System.Runtime.Intrinsics.Vector64:CreateScalar(short):System.Runtime.Intrinsics.Vector64`1[Int16]; ============================================================

There are multiple issues here:

  1. Redundant str/ldr-s with a SIMD register - this appears only in Vector64.CreateScalar():
str d0,[fp,#24]ldr d0,[fp,#24]

The code is the worst for Vector64<float>.CreateScalar()

 0E040FF0 dup v16.2s, wzr FD000FB0 str d16,[fp,#24] FD400FB0 ldr d16,[fp,#24] 6E040410 ins v16.s[0], v0.s[0] 1E604200 fmov d0, d16

or Vector64<ushort>.CreateScalar()

 0E020FE0 dup v0.4h, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 53003C00 uxth w0, w0 4E021C00 ins v0.h[0], w0
  1. Unnecessary sign-/zero-extensions with byte,ubyte,short,ushort (the same as seen in ARM64 intrinsic support for Vector64.Create() and Vector128.Create() #35590):
uxtb w0, w0uxth w0, w0sxtb w0, w0sxth w0, w0
  1. dup Vd.T, wzr seems to be used for code generation of Vector64/128.Zero which I thought was fixed with Implement Vector{Size}<T>.AllBitsSet #33924 (cc @Gnbrkm41). I will follow up on this

cc @kunalspathak@BruceForstall

@echesakov
echesakov marked this pull request as ready for review April 30, 2020 19:35
return AdvSimd.Insert(Vector64<byte>.Zero, 0, value);
}

return SoftwareFallback(value);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Curious about wrapping the SoftwareFallback() in a static method. What is the advantage of doing it?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure, actually. I did this to be consistent with existing Vector128/256 implementations.
@tannergooding Do you know why?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With the logic for all the various paths the method is too large for the normal inlining heuristics to work so we need to mark it AggressiveInlining (since the accelerated paths will generally be pretty small).
However, we don't necessarily want the SoftwareFallback to be inlined as that may not be beneficial.
Putting it in its own method prevents it from being inlined in the normal case and allows the JIT to decide if it is "too large or not" by itself.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That make sense.

@kunalspathakkunalspathak left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@echesakov
echesakov merged commit 670bf21 into dotnet:masterApr 30, 2020
@echesakov
echesakov deleted the Arm64-ASIMD-Vector64-Vector128-CreateScalar-Use-AdvSimd-Insert branch April 30, 2020 23:35
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@echesakov@tannergooding@kunalspathak
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

[Arm64] Implement Vector64/128.CreateScalar() using AdvSimd.Insert - #35300

Merged
echesakov merged 2 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Vector64-Vector128-CreateScalar-Use-AdvSimd-Insert
Apr 30, 2020
Merged

[Arm64] Implement Vector64/128.CreateScalar() using AdvSimd.Insert#35300
echesakov merged 2 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Vector64-Vector128-CreateScalar-Use-AdvSimd-Insert

Conversation

@echesakov

Copy link
Copy Markdown
Contributor

No description provided.

@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @tannergooding
Notify danmosemsft if you want to be subscribed.

{
if (AdvSimd.IsSupported)
{
return AdvSimd.Insert(Vector128<byte>.Zero, 0, value);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We'll need to special-case CreateScalarUnsafe since the upper bits don't have to be zeroed for it.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep, this work is tracked by #34485

@echesakov

Copy link
Copy Markdown
ContributorAuthor

I still would like to collect the jisDisasm-s for this - I remember seeing something weird for Vector64.CreateScalar() - will do it later.

@echesakov

echesakov commented Apr 30, 2020

Copy link
Copy Markdown
ContributorAuthor
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector128`1[Byte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ubyte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M19699_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M19699_IG02: 4E010FF0 dup v16.16b, wzr 53001C00 uxtb w0, w0 4E011C10 ins v16.b[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M19699_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=a381b30c) for method System.Runtime.Intrinsics.Vector128:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector128`1[Byte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(double):System.Runtime.Intrinsics.Vector128`1[Double]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) double -> d0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M6886_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M6886_IG02: 4E080FF0 dup v16.2d, xzr 6E080410 ins v16.d[0], v0.d[0] 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M6886_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=17a3e519) for method System.Runtime.Intrinsics.Vector128:CreateScalar(double):System.Runtime.Intrinsics.Vector128`1[Double]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(short):System.Runtime.Intrinsics.Vector128`1[Int16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) short -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M37120_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M37120_IG02: 4E020FF0 dup v16.8h, wzr 13003C00 sxth w0, w0 4E021C10 ins v16.h[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M37120_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=87f66eff) for method System.Runtime.Intrinsics.Vector128:CreateScalar(short):System.Runtime.Intrinsics.Vector128`1[Int16]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[Int32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M42503_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M42503_IG02: 4E040FF0 dup v16.4s, wzr 4E041C10 ins v16.s[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M42503_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=56ed59f8) for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[Int32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[Int64]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) long -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M9853_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M9853_IG02: 4E080FF0 dup v16.2d, xzr 4E081C10 ins v16.d[0], x0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M9853_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=99c7d982) for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[Int64]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(byte):System.Runtime.Intrinsics.Vector128`1[SByte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) byte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M64149_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M64149_IG02: 4E010FF0 dup v16.16b, wzr 13001C00 sxtb w0, w0 4E011C10 ins v16.b[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M64149_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=3757056a) for method System.Runtime.Intrinsics.Vector128:CreateScalar(byte):System.Runtime.Intrinsics.Vector128`1[SByte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(float):System.Runtime.Intrinsics.Vector128`1[Single]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) float -> d0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M28940_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M28940_IG02: 4E040FF0 dup v16.4s, wzr 6E040410 ins v16.s[0], v0.s[0] 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M28940_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=910d8ef3) for method System.Runtime.Intrinsics.Vector128:CreateScalar(float):System.Runtime.Intrinsics.Vector128`1[Single]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(ushort):System.Runtime.Intrinsics.Vector128`1[UInt16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ushort -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M480_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M480_IG02: 4E020FF0 dup v16.8h, wzr 53003C00 uxth w0, w0 4E021C10 ins v16.h[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M480_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=5a77fe1f) for method System.Runtime.Intrinsics.Vector128:CreateScalar(ushort):System.Runtime.Intrinsics.Vector128`1[UInt16]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[UInt32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M21746_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M21746_IG02: 4E040FF0 dup v16.4s, wzr 4E041C10 ins v16.s[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M21746_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=4a35ab0d) for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[UInt32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[UInt64]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) long -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M2664_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M2664_IG02: 4E080FF0 dup v16.2d, xzr 4E081C10 ins v16.d[0], x0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M2664_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=714bf597) for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[UInt64]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[Int32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M25863_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M25863_IG02: 0E040FE0 dup v0.2s, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 4E041C00 ins v0.s[0], w0 ;; bbWeight=1 PerfScore 6.00G_M25863_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 12.70, (MethodHash=80a89af8) for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[Int32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(byte):System.Runtime.Intrinsics.Vector64`1[SByte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) byte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M12309_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M12309_IG02: 0E010FE0 dup v0.8b, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 13001C00 sxtb w0, w0 4E011C00 ins v0.b[0], w0 ;; bbWeight=1 PerfScore 6.50G_M12309_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=1802cfea) for method System.Runtime.Intrinsics.Vector64:CreateScalar(byte):System.Runtime.Intrinsics.Vector64`1[SByte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(float):System.Runtime.Intrinsics.Vector64`1[Single]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) float -> d0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d16 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M44268_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M44268_IG02: 0E040FF0 dup v16.2s, wzr FD000FB0 str d16,[fp,#24] FD400FB0 ldr d16,[fp,#24] 6E040410 ins v16.s[0], v0.s[0] 1E604200 fmov d0, d16 ;; bbWeight=1 PerfScore 6.50G_M44268_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=b5c65313) for method System.Runtime.Intrinsics.Vector64:CreateScalar(float):System.Runtime.Intrinsics.Vector64`1[Single]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(ushort):System.Runtime.Intrinsics.Vector64`1[UInt16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ushort -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M37504_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M37504_IG02: 0E020FE0 dup v0.4h, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 53003C00 uxth w0, w0 4E021C00 ins v0.h[0], w0 ;; bbWeight=1 PerfScore 6.50G_M37504_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=68536d7f) for method System.Runtime.Intrinsics.Vector64:CreateScalar(ushort):System.Runtime.Intrinsics.Vector64`1[UInt16]; ============================================================

Collected JIT disassemblies with the changes rebased on top of latest master

; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[UInt32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M62450_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M62450_IG02: 0E040FE0 dup v0.2s, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 4E041C00 ins v0.s[0], w0 ;; bbWeight=1 PerfScore 6.00G_M62450_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 12.70, (MethodHash=ab590c0d) for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[UInt32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector64`1[Byte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ubyte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M20083_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M20083_IG02: 0E010FE0 dup v0.8b, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 53001C00 uxtb w0, w0 4E011C00 ins v0.b[0], w0 ;; bbWeight=1 PerfScore 6.50G_M20083_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=cedeb18c) for method System.Runtime.Intrinsics.Vector64:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector64`1[Byte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(short):System.Runtime.Intrinsics.Vector64`1[Int16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) short -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M58336_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M58336_IG02: 0E020FE0 dup v0.4h, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 13003C00 sxth w0, w0 4E021C00 ins v0.h[0], w0 ;; bbWeight=1 PerfScore 6.50G_M58336_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=e95c1c1f) for method System.Runtime.Intrinsics.Vector64:CreateScalar(short):System.Runtime.Intrinsics.Vector64`1[Int16]; ============================================================

There are multiple issues here:

  1. Redundant str/ldr-s with a SIMD register - this appears only in Vector64.CreateScalar():
str d0,[fp,#24]ldr d0,[fp,#24]

The code is the worst for Vector64<float>.CreateScalar()

 0E040FF0 dup v16.2s, wzr FD000FB0 str d16,[fp,#24] FD400FB0 ldr d16,[fp,#24] 6E040410 ins v16.s[0], v0.s[0] 1E604200 fmov d0, d16

or Vector64<ushort>.CreateScalar()

 0E020FE0 dup v0.4h, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 53003C00 uxth w0, w0 4E021C00 ins v0.h[0], w0
  1. Unnecessary sign-/zero-extensions with byte,ubyte,short,ushort (the same as seen in ARM64 intrinsic support for Vector64.Create() and Vector128.Create() #35590):
uxtb w0, w0uxth w0, w0sxtb w0, w0sxth w0, w0
  1. dup Vd.T, wzr seems to be used for code generation of Vector64/128.Zero which I thought was fixed with Implement Vector{Size}<T>.AllBitsSet #33924 (cc @Gnbrkm41). I will follow up on this

cc @kunalspathak@BruceForstall

@echesakov
echesakov marked this pull request as ready for review April 30, 2020 19:35
return AdvSimd.Insert(Vector64<byte>.Zero, 0, value);
}

return SoftwareFallback(value);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Curious about wrapping the SoftwareFallback() in a static method. What is the advantage of doing it?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure, actually. I did this to be consistent with existing Vector128/256 implementations.
@tannergooding Do you know why?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With the logic for all the various paths the method is too large for the normal inlining heuristics to work so we need to mark it AggressiveInlining (since the accelerated paths will generally be pretty small).
However, we don't necessarily want the SoftwareFallback to be inlined as that may not be beneficial.
Putting it in its own method prevents it from being inlined in the normal case and allows the JIT to decide if it is "too large or not" by itself.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That make sense.

@kunalspathakkunalspathak left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@echesakov
echesakov merged commit 670bf21 into dotnet:masterApr 30, 2020
@echesakov
echesakov deleted the Arm64-ASIMD-Vector64-Vector128-CreateScalar-Use-AdvSimd-Insert branch April 30, 2020 23:35
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@echesakov@tannergooding@kunalspathak
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[Arm64] Implement Vector64/128.CreateScalar() using AdvSimd.Insert - #35300

Merged
echesakov merged 2 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Vector64-Vector128-CreateScalar-Use-AdvSimd-Insert
Apr 30, 2020
Merged

[Arm64] Implement Vector64/128.CreateScalar() using AdvSimd.Insert#35300
echesakov merged 2 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Vector64-Vector128-CreateScalar-Use-AdvSimd-Insert

Conversation

@echesakov

Copy link
Copy Markdown
Contributor

No description provided.

@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @tannergooding
Notify danmosemsft if you want to be subscribed.

{
if (AdvSimd.IsSupported)
{
return AdvSimd.Insert(Vector128<byte>.Zero, 0, value);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We'll need to special-case CreateScalarUnsafe since the upper bits don't have to be zeroed for it.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep, this work is tracked by #34485

@echesakov

Copy link
Copy Markdown
ContributorAuthor

I still would like to collect the jisDisasm-s for this - I remember seeing something weird for Vector64.CreateScalar() - will do it later.

@echesakov

echesakov commented Apr 30, 2020

Copy link
Copy Markdown
ContributorAuthor
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector128`1[Byte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ubyte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M19699_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M19699_IG02: 4E010FF0 dup v16.16b, wzr 53001C00 uxtb w0, w0 4E011C10 ins v16.b[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M19699_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=a381b30c) for method System.Runtime.Intrinsics.Vector128:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector128`1[Byte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(double):System.Runtime.Intrinsics.Vector128`1[Double]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) double -> d0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M6886_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M6886_IG02: 4E080FF0 dup v16.2d, xzr 6E080410 ins v16.d[0], v0.d[0] 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M6886_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=17a3e519) for method System.Runtime.Intrinsics.Vector128:CreateScalar(double):System.Runtime.Intrinsics.Vector128`1[Double]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(short):System.Runtime.Intrinsics.Vector128`1[Int16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) short -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M37120_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M37120_IG02: 4E020FF0 dup v16.8h, wzr 13003C00 sxth w0, w0 4E021C10 ins v16.h[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M37120_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=87f66eff) for method System.Runtime.Intrinsics.Vector128:CreateScalar(short):System.Runtime.Intrinsics.Vector128`1[Int16]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[Int32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M42503_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M42503_IG02: 4E040FF0 dup v16.4s, wzr 4E041C10 ins v16.s[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M42503_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=56ed59f8) for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[Int32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[Int64]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) long -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M9853_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M9853_IG02: 4E080FF0 dup v16.2d, xzr 4E081C10 ins v16.d[0], x0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M9853_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=99c7d982) for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[Int64]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(byte):System.Runtime.Intrinsics.Vector128`1[SByte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) byte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M64149_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M64149_IG02: 4E010FF0 dup v16.16b, wzr 13001C00 sxtb w0, w0 4E011C10 ins v16.b[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M64149_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=3757056a) for method System.Runtime.Intrinsics.Vector128:CreateScalar(byte):System.Runtime.Intrinsics.Vector128`1[SByte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(float):System.Runtime.Intrinsics.Vector128`1[Single]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) float -> d0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M28940_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M28940_IG02: 4E040FF0 dup v16.4s, wzr 6E040410 ins v16.s[0], v0.s[0] 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M28940_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=910d8ef3) for method System.Runtime.Intrinsics.Vector128:CreateScalar(float):System.Runtime.Intrinsics.Vector128`1[Single]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(ushort):System.Runtime.Intrinsics.Vector128`1[UInt16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ushort -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M480_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M480_IG02: 4E020FF0 dup v16.8h, wzr 53003C00 uxth w0, w0 4E021C10 ins v16.h[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M480_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=5a77fe1f) for method System.Runtime.Intrinsics.Vector128:CreateScalar(ushort):System.Runtime.Intrinsics.Vector128`1[UInt16]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[UInt32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M21746_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M21746_IG02: 4E040FF0 dup v16.4s, wzr 4E041C10 ins v16.s[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M21746_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=4a35ab0d) for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[UInt32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[UInt64]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) long -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M2664_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M2664_IG02: 4E080FF0 dup v16.2d, xzr 4E081C10 ins v16.d[0], x0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M2664_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=714bf597) for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[UInt64]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[Int32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M25863_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M25863_IG02: 0E040FE0 dup v0.2s, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 4E041C00 ins v0.s[0], w0 ;; bbWeight=1 PerfScore 6.00G_M25863_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 12.70, (MethodHash=80a89af8) for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[Int32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(byte):System.Runtime.Intrinsics.Vector64`1[SByte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) byte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M12309_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M12309_IG02: 0E010FE0 dup v0.8b, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 13001C00 sxtb w0, w0 4E011C00 ins v0.b[0], w0 ;; bbWeight=1 PerfScore 6.50G_M12309_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=1802cfea) for method System.Runtime.Intrinsics.Vector64:CreateScalar(byte):System.Runtime.Intrinsics.Vector64`1[SByte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(float):System.Runtime.Intrinsics.Vector64`1[Single]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) float -> d0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d16 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M44268_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M44268_IG02: 0E040FF0 dup v16.2s, wzr FD000FB0 str d16,[fp,#24] FD400FB0 ldr d16,[fp,#24] 6E040410 ins v16.s[0], v0.s[0] 1E604200 fmov d0, d16 ;; bbWeight=1 PerfScore 6.50G_M44268_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=b5c65313) for method System.Runtime.Intrinsics.Vector64:CreateScalar(float):System.Runtime.Intrinsics.Vector64`1[Single]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(ushort):System.Runtime.Intrinsics.Vector64`1[UInt16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ushort -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M37504_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M37504_IG02: 0E020FE0 dup v0.4h, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 53003C00 uxth w0, w0 4E021C00 ins v0.h[0], w0 ;; bbWeight=1 PerfScore 6.50G_M37504_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=68536d7f) for method System.Runtime.Intrinsics.Vector64:CreateScalar(ushort):System.Runtime.Intrinsics.Vector64`1[UInt16]; ============================================================

Collected JIT disassemblies with the changes rebased on top of latest master

; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[UInt32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M62450_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M62450_IG02: 0E040FE0 dup v0.2s, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 4E041C00 ins v0.s[0], w0 ;; bbWeight=1 PerfScore 6.00G_M62450_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 12.70, (MethodHash=ab590c0d) for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[UInt32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector64`1[Byte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ubyte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M20083_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M20083_IG02: 0E010FE0 dup v0.8b, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 53001C00 uxtb w0, w0 4E011C00 ins v0.b[0], w0 ;; bbWeight=1 PerfScore 6.50G_M20083_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=cedeb18c) for method System.Runtime.Intrinsics.Vector64:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector64`1[Byte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(short):System.Runtime.Intrinsics.Vector64`1[Int16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) short -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M58336_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M58336_IG02: 0E020FE0 dup v0.4h, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 13003C00 sxth w0, w0 4E021C00 ins v0.h[0], w0 ;; bbWeight=1 PerfScore 6.50G_M58336_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=e95c1c1f) for method System.Runtime.Intrinsics.Vector64:CreateScalar(short):System.Runtime.Intrinsics.Vector64`1[Int16]; ============================================================

There are multiple issues here:

  1. Redundant str/ldr-s with a SIMD register - this appears only in Vector64.CreateScalar():
str d0,[fp,#24]ldr d0,[fp,#24]

The code is the worst for Vector64<float>.CreateScalar()

 0E040FF0 dup v16.2s, wzr FD000FB0 str d16,[fp,#24] FD400FB0 ldr d16,[fp,#24] 6E040410 ins v16.s[0], v0.s[0] 1E604200 fmov d0, d16

or Vector64<ushort>.CreateScalar()

 0E020FE0 dup v0.4h, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 53003C00 uxth w0, w0 4E021C00 ins v0.h[0], w0
  1. Unnecessary sign-/zero-extensions with byte,ubyte,short,ushort (the same as seen in ARM64 intrinsic support for Vector64.Create() and Vector128.Create() #35590):
uxtb w0, w0uxth w0, w0sxtb w0, w0sxth w0, w0
  1. dup Vd.T, wzr seems to be used for code generation of Vector64/128.Zero which I thought was fixed with Implement Vector{Size}<T>.AllBitsSet #33924 (cc @Gnbrkm41). I will follow up on this

cc @kunalspathak@BruceForstall

@echesakov
echesakov marked this pull request as ready for review April 30, 2020 19:35
return AdvSimd.Insert(Vector64<byte>.Zero, 0, value);
}

return SoftwareFallback(value);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Curious about wrapping the SoftwareFallback() in a static method. What is the advantage of doing it?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure, actually. I did this to be consistent with existing Vector128/256 implementations.
@tannergooding Do you know why?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With the logic for all the various paths the method is too large for the normal inlining heuristics to work so we need to mark it AggressiveInlining (since the accelerated paths will generally be pretty small).
However, we don't necessarily want the SoftwareFallback to be inlined as that may not be beneficial.
Putting it in its own method prevents it from being inlined in the normal case and allows the JIT to decide if it is "too large or not" by itself.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That make sense.

@kunalspathakkunalspathak left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@echesakov
echesakov merged commit 670bf21 into dotnet:masterApr 30, 2020
@echesakov
echesakov deleted the Arm64-ASIMD-Vector64-Vector128-CreateScalar-Use-AdvSimd-Insert branch April 30, 2020 23:35
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@echesakov@tannergooding@kunalspathak
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[Arm64] Implement Vector64/128.CreateScalar() using AdvSimd.Insert - #35300

Merged
echesakov merged 2 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Vector64-Vector128-CreateScalar-Use-AdvSimd-Insert
Apr 30, 2020
Merged

[Arm64] Implement Vector64/128.CreateScalar() using AdvSimd.Insert#35300
echesakov merged 2 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Vector64-Vector128-CreateScalar-Use-AdvSimd-Insert

Conversation

@echesakov

Copy link
Copy Markdown
Contributor

No description provided.

@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @tannergooding
Notify danmosemsft if you want to be subscribed.

{
if (AdvSimd.IsSupported)
{
return AdvSimd.Insert(Vector128<byte>.Zero, 0, value);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We'll need to special-case CreateScalarUnsafe since the upper bits don't have to be zeroed for it.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep, this work is tracked by #34485

@echesakov

Copy link
Copy Markdown
ContributorAuthor

I still would like to collect the jisDisasm-s for this - I remember seeing something weird for Vector64.CreateScalar() - will do it later.

@echesakov

echesakov commented Apr 30, 2020

Copy link
Copy Markdown
ContributorAuthor
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector128`1[Byte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ubyte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M19699_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M19699_IG02: 4E010FF0 dup v16.16b, wzr 53001C00 uxtb w0, w0 4E011C10 ins v16.b[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M19699_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=a381b30c) for method System.Runtime.Intrinsics.Vector128:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector128`1[Byte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(double):System.Runtime.Intrinsics.Vector128`1[Double]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) double -> d0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M6886_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M6886_IG02: 4E080FF0 dup v16.2d, xzr 6E080410 ins v16.d[0], v0.d[0] 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M6886_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=17a3e519) for method System.Runtime.Intrinsics.Vector128:CreateScalar(double):System.Runtime.Intrinsics.Vector128`1[Double]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(short):System.Runtime.Intrinsics.Vector128`1[Int16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) short -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M37120_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M37120_IG02: 4E020FF0 dup v16.8h, wzr 13003C00 sxth w0, w0 4E021C10 ins v16.h[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M37120_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=87f66eff) for method System.Runtime.Intrinsics.Vector128:CreateScalar(short):System.Runtime.Intrinsics.Vector128`1[Int16]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[Int32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M42503_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M42503_IG02: 4E040FF0 dup v16.4s, wzr 4E041C10 ins v16.s[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M42503_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=56ed59f8) for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[Int32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[Int64]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) long -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M9853_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M9853_IG02: 4E080FF0 dup v16.2d, xzr 4E081C10 ins v16.d[0], x0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M9853_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=99c7d982) for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[Int64]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(byte):System.Runtime.Intrinsics.Vector128`1[SByte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) byte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M64149_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M64149_IG02: 4E010FF0 dup v16.16b, wzr 13001C00 sxtb w0, w0 4E011C10 ins v16.b[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M64149_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=3757056a) for method System.Runtime.Intrinsics.Vector128:CreateScalar(byte):System.Runtime.Intrinsics.Vector128`1[SByte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(float):System.Runtime.Intrinsics.Vector128`1[Single]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) float -> d0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M28940_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M28940_IG02: 4E040FF0 dup v16.4s, wzr 6E040410 ins v16.s[0], v0.s[0] 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M28940_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=910d8ef3) for method System.Runtime.Intrinsics.Vector128:CreateScalar(float):System.Runtime.Intrinsics.Vector128`1[Single]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(ushort):System.Runtime.Intrinsics.Vector128`1[UInt16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ushort -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M480_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M480_IG02: 4E020FF0 dup v16.8h, wzr 53003C00 uxth w0, w0 4E021C10 ins v16.h[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M480_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=5a77fe1f) for method System.Runtime.Intrinsics.Vector128:CreateScalar(ushort):System.Runtime.Intrinsics.Vector128`1[UInt16]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[UInt32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M21746_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M21746_IG02: 4E040FF0 dup v16.4s, wzr 4E041C10 ins v16.s[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M21746_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=4a35ab0d) for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[UInt32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[UInt64]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) long -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M2664_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M2664_IG02: 4E080FF0 dup v16.2d, xzr 4E081C10 ins v16.d[0], x0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M2664_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=714bf597) for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[UInt64]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[Int32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M25863_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M25863_IG02: 0E040FE0 dup v0.2s, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 4E041C00 ins v0.s[0], w0 ;; bbWeight=1 PerfScore 6.00G_M25863_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 12.70, (MethodHash=80a89af8) for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[Int32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(byte):System.Runtime.Intrinsics.Vector64`1[SByte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) byte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M12309_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M12309_IG02: 0E010FE0 dup v0.8b, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 13001C00 sxtb w0, w0 4E011C00 ins v0.b[0], w0 ;; bbWeight=1 PerfScore 6.50G_M12309_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=1802cfea) for method System.Runtime.Intrinsics.Vector64:CreateScalar(byte):System.Runtime.Intrinsics.Vector64`1[SByte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(float):System.Runtime.Intrinsics.Vector64`1[Single]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) float -> d0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d16 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M44268_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M44268_IG02: 0E040FF0 dup v16.2s, wzr FD000FB0 str d16,[fp,#24] FD400FB0 ldr d16,[fp,#24] 6E040410 ins v16.s[0], v0.s[0] 1E604200 fmov d0, d16 ;; bbWeight=1 PerfScore 6.50G_M44268_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=b5c65313) for method System.Runtime.Intrinsics.Vector64:CreateScalar(float):System.Runtime.Intrinsics.Vector64`1[Single]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(ushort):System.Runtime.Intrinsics.Vector64`1[UInt16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ushort -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M37504_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M37504_IG02: 0E020FE0 dup v0.4h, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 53003C00 uxth w0, w0 4E021C00 ins v0.h[0], w0 ;; bbWeight=1 PerfScore 6.50G_M37504_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=68536d7f) for method System.Runtime.Intrinsics.Vector64:CreateScalar(ushort):System.Runtime.Intrinsics.Vector64`1[UInt16]; ============================================================

Collected JIT disassemblies with the changes rebased on top of latest master

; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[UInt32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M62450_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M62450_IG02: 0E040FE0 dup v0.2s, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 4E041C00 ins v0.s[0], w0 ;; bbWeight=1 PerfScore 6.00G_M62450_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 12.70, (MethodHash=ab590c0d) for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[UInt32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector64`1[Byte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ubyte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M20083_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M20083_IG02: 0E010FE0 dup v0.8b, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 53001C00 uxtb w0, w0 4E011C00 ins v0.b[0], w0 ;; bbWeight=1 PerfScore 6.50G_M20083_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=cedeb18c) for method System.Runtime.Intrinsics.Vector64:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector64`1[Byte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(short):System.Runtime.Intrinsics.Vector64`1[Int16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) short -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M58336_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M58336_IG02: 0E020FE0 dup v0.4h, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 13003C00 sxth w0, w0 4E021C00 ins v0.h[0], w0 ;; bbWeight=1 PerfScore 6.50G_M58336_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=e95c1c1f) for method System.Runtime.Intrinsics.Vector64:CreateScalar(short):System.Runtime.Intrinsics.Vector64`1[Int16]; ============================================================

There are multiple issues here:

  1. Redundant str/ldr-s with a SIMD register - this appears only in Vector64.CreateScalar():
str d0,[fp,#24]ldr d0,[fp,#24]

The code is the worst for Vector64<float>.CreateScalar()

 0E040FF0 dup v16.2s, wzr FD000FB0 str d16,[fp,#24] FD400FB0 ldr d16,[fp,#24] 6E040410 ins v16.s[0], v0.s[0] 1E604200 fmov d0, d16

or Vector64<ushort>.CreateScalar()

 0E020FE0 dup v0.4h, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 53003C00 uxth w0, w0 4E021C00 ins v0.h[0], w0
  1. Unnecessary sign-/zero-extensions with byte,ubyte,short,ushort (the same as seen in ARM64 intrinsic support for Vector64.Create() and Vector128.Create() #35590):
uxtb w0, w0uxth w0, w0sxtb w0, w0sxth w0, w0
  1. dup Vd.T, wzr seems to be used for code generation of Vector64/128.Zero which I thought was fixed with Implement Vector{Size}<T>.AllBitsSet #33924 (cc @Gnbrkm41). I will follow up on this

cc @kunalspathak@BruceForstall

@echesakov
echesakov marked this pull request as ready for review April 30, 2020 19:35
return AdvSimd.Insert(Vector64<byte>.Zero, 0, value);
}

return SoftwareFallback(value);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Curious about wrapping the SoftwareFallback() in a static method. What is the advantage of doing it?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure, actually. I did this to be consistent with existing Vector128/256 implementations.
@tannergooding Do you know why?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With the logic for all the various paths the method is too large for the normal inlining heuristics to work so we need to mark it AggressiveInlining (since the accelerated paths will generally be pretty small).
However, we don't necessarily want the SoftwareFallback to be inlined as that may not be beneficial.
Putting it in its own method prevents it from being inlined in the normal case and allows the JIT to decide if it is "too large or not" by itself.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That make sense.

@kunalspathakkunalspathak left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@echesakov
echesakov merged commit 670bf21 into dotnet:masterApr 30, 2020
@echesakov
echesakov deleted the Arm64-ASIMD-Vector64-Vector128-CreateScalar-Use-AdvSimd-Insert branch April 30, 2020 23:35
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@echesakov@tannergooding@kunalspathak
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

[Arm64] Implement Vector64/128.CreateScalar() using AdvSimd.Insert - #35300

Merged
echesakov merged 2 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Vector64-Vector128-CreateScalar-Use-AdvSimd-Insert
Apr 30, 2020
Merged

[Arm64] Implement Vector64/128.CreateScalar() using AdvSimd.Insert#35300
echesakov merged 2 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Vector64-Vector128-CreateScalar-Use-AdvSimd-Insert

Conversation

@echesakov

Copy link
Copy Markdown
Contributor

No description provided.

@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @tannergooding
Notify danmosemsft if you want to be subscribed.

{
if (AdvSimd.IsSupported)
{
return AdvSimd.Insert(Vector128<byte>.Zero, 0, value);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We'll need to special-case CreateScalarUnsafe since the upper bits don't have to be zeroed for it.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep, this work is tracked by #34485

@echesakov

Copy link
Copy Markdown
ContributorAuthor

I still would like to collect the jisDisasm-s for this - I remember seeing something weird for Vector64.CreateScalar() - will do it later.

@echesakov

echesakov commented Apr 30, 2020

Copy link
Copy Markdown
ContributorAuthor
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector128`1[Byte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ubyte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M19699_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M19699_IG02: 4E010FF0 dup v16.16b, wzr 53001C00 uxtb w0, w0 4E011C10 ins v16.b[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M19699_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=a381b30c) for method System.Runtime.Intrinsics.Vector128:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector128`1[Byte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(double):System.Runtime.Intrinsics.Vector128`1[Double]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) double -> d0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M6886_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M6886_IG02: 4E080FF0 dup v16.2d, xzr 6E080410 ins v16.d[0], v0.d[0] 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M6886_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=17a3e519) for method System.Runtime.Intrinsics.Vector128:CreateScalar(double):System.Runtime.Intrinsics.Vector128`1[Double]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(short):System.Runtime.Intrinsics.Vector128`1[Int16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) short -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M37120_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M37120_IG02: 4E020FF0 dup v16.8h, wzr 13003C00 sxth w0, w0 4E021C10 ins v16.h[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M37120_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=87f66eff) for method System.Runtime.Intrinsics.Vector128:CreateScalar(short):System.Runtime.Intrinsics.Vector128`1[Int16]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[Int32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M42503_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M42503_IG02: 4E040FF0 dup v16.4s, wzr 4E041C10 ins v16.s[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M42503_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=56ed59f8) for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[Int32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[Int64]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) long -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M9853_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M9853_IG02: 4E080FF0 dup v16.2d, xzr 4E081C10 ins v16.d[0], x0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M9853_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=99c7d982) for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[Int64]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(byte):System.Runtime.Intrinsics.Vector128`1[SByte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) byte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M64149_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M64149_IG02: 4E010FF0 dup v16.16b, wzr 13001C00 sxtb w0, w0 4E011C10 ins v16.b[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M64149_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=3757056a) for method System.Runtime.Intrinsics.Vector128:CreateScalar(byte):System.Runtime.Intrinsics.Vector128`1[SByte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(float):System.Runtime.Intrinsics.Vector128`1[Single]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) float -> d0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M28940_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M28940_IG02: 4E040FF0 dup v16.4s, wzr 6E040410 ins v16.s[0], v0.s[0] 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M28940_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=910d8ef3) for method System.Runtime.Intrinsics.Vector128:CreateScalar(float):System.Runtime.Intrinsics.Vector128`1[Single]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(ushort):System.Runtime.Intrinsics.Vector128`1[UInt16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ushort -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M480_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M480_IG02: 4E020FF0 dup v16.8h, wzr 53003C00 uxth w0, w0 4E021C10 ins v16.h[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 4.00G_M480_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 10.70, (MethodHash=5a77fe1f) for method System.Runtime.Intrinsics.Vector128:CreateScalar(ushort):System.Runtime.Intrinsics.Vector128`1[UInt16]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[UInt32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M21746_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M21746_IG02: 4E040FF0 dup v16.4s, wzr 4E041C10 ins v16.s[0], w0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M21746_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=4a35ab0d) for method System.Runtime.Intrinsics.Vector128:CreateScalar(int):System.Runtime.Intrinsics.Vector128`1[UInt32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[UInt64]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) long -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) simd16 -> zero-ref HFA(simd16) "struct address for call/obj"; V03 tmp2 [V03,T01] ( 2, 2 ) simd16 -> d16 HFA(simd16) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 0G_M2664_IG01: A9BF7BFD stp fp, lr,[sp,#-16]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M2664_IG02: 4E080FF0 dup v16.2d, xzr 4E081C10 ins v16.d[0], x0 4EB01E00 mov v0.16b, v16.16b ;; bbWeight=1 PerfScore 3.50G_M2664_IG03: A8C17BFD ldp fp, lr,[sp],#16 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 28, prolog size 8, PerfScore 9.80, (MethodHash=714bf597) for method System.Runtime.Intrinsics.Vector128:CreateScalar(long):System.Runtime.Intrinsics.Vector128`1[UInt64]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[Int32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M25863_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M25863_IG02: 0E040FE0 dup v0.2s, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 4E041C00 ins v0.s[0], w0 ;; bbWeight=1 PerfScore 6.00G_M25863_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 12.70, (MethodHash=80a89af8) for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[Int32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(byte):System.Runtime.Intrinsics.Vector64`1[SByte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) byte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M12309_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M12309_IG02: 0E010FE0 dup v0.8b, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 13001C00 sxtb w0, w0 4E011C00 ins v0.b[0], w0 ;; bbWeight=1 PerfScore 6.50G_M12309_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=1802cfea) for method System.Runtime.Intrinsics.Vector64:CreateScalar(byte):System.Runtime.Intrinsics.Vector64`1[SByte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(float):System.Runtime.Intrinsics.Vector64`1[Single]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) float -> d0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d16 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M44268_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M44268_IG02: 0E040FF0 dup v16.2s, wzr FD000FB0 str d16,[fp,#24] FD400FB0 ldr d16,[fp,#24] 6E040410 ins v16.s[0], v0.s[0] 1E604200 fmov d0, d16 ;; bbWeight=1 PerfScore 6.50G_M44268_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=b5c65313) for method System.Runtime.Intrinsics.Vector64:CreateScalar(float):System.Runtime.Intrinsics.Vector64`1[Single]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(ushort):System.Runtime.Intrinsics.Vector64`1[UInt16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ushort -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M37504_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M37504_IG02: 0E020FE0 dup v0.4h, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 53003C00 uxth w0, w0 4E021C00 ins v0.h[0], w0 ;; bbWeight=1 PerfScore 6.50G_M37504_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=68536d7f) for method System.Runtime.Intrinsics.Vector64:CreateScalar(ushort):System.Runtime.Intrinsics.Vector64`1[UInt16]; ============================================================

Collected JIT disassemblies with the changes rebased on top of latest master

; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[UInt32]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) int -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M62450_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M62450_IG02: 0E040FE0 dup v0.2s, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 4E041C00 ins v0.s[0], w0 ;; bbWeight=1 PerfScore 6.00G_M62450_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 32, prolog size 8, PerfScore 12.70, (MethodHash=ab590c0d) for method System.Runtime.Intrinsics.Vector64:CreateScalar(int):System.Runtime.Intrinsics.Vector64`1[UInt32]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector64`1[Byte]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) ubyte -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M20083_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M20083_IG02: 0E010FE0 dup v0.8b, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 53001C00 uxtb w0, w0 4E011C00 ins v0.b[0], w0 ;; bbWeight=1 PerfScore 6.50G_M20083_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=cedeb18c) for method System.Runtime.Intrinsics.Vector64:CreateScalar(ubyte):System.Runtime.Intrinsics.Vector64`1[Byte]; ============================================================
; Assembly listing for method System.Runtime.Intrinsics.Vector64:CreateScalar(short):System.Runtime.Intrinsics.Vector64`1[Int16]; Emitting BLENDED_CODE for generic ARM64 CPU - Windows; optimized code; fp based frame; partially interruptible; Final local variable assignments;; V00 arg0 [V00,T00] ( 3, 3 ) short -> x0 ;# V01 OutArgs [V01 ] ( 1, 1 ) lclBlk ( 0) [sp+0x00] "OutgoingArgSpace"; V02 tmp1 [V02,T01] ( 2, 4 ) simd8 -> [fp+0x18] HFA(double) do-not-enreg[SF] "struct address for call/obj"; V03 tmp2 [V03,T02] ( 2, 2 ) simd8 -> d0 HFA(double) ld-addr-op "Inline ldloca(s) first use temp";; Lcl frame size = 16G_M58336_IG01: A9BE7BFD stp fp, lr,[sp,#-32]! 910003FD mov fp,sp ;; bbWeight=1 PerfScore 1.50G_M58336_IG02: 0E020FE0 dup v0.4h, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 13003C00 sxth w0, w0 4E021C00 ins v0.h[0], w0 ;; bbWeight=1 PerfScore 6.50G_M58336_IG03: A8C27BFD ldp fp, lr,[sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00; Total bytes of code 36, prolog size 8, PerfScore 13.60, (MethodHash=e95c1c1f) for method System.Runtime.Intrinsics.Vector64:CreateScalar(short):System.Runtime.Intrinsics.Vector64`1[Int16]; ============================================================

There are multiple issues here:

  1. Redundant str/ldr-s with a SIMD register - this appears only in Vector64.CreateScalar():
str d0,[fp,#24]ldr d0,[fp,#24]

The code is the worst for Vector64<float>.CreateScalar()

 0E040FF0 dup v16.2s, wzr FD000FB0 str d16,[fp,#24] FD400FB0 ldr d16,[fp,#24] 6E040410 ins v16.s[0], v0.s[0] 1E604200 fmov d0, d16

or Vector64<ushort>.CreateScalar()

 0E020FE0 dup v0.4h, wzr FD000FA0 str d0,[fp,#24] FD400FA0 ldr d0,[fp,#24] 53003C00 uxth w0, w0 4E021C00 ins v0.h[0], w0
  1. Unnecessary sign-/zero-extensions with byte,ubyte,short,ushort (the same as seen in ARM64 intrinsic support for Vector64.Create() and Vector128.Create() #35590):
uxtb w0, w0uxth w0, w0sxtb w0, w0sxth w0, w0
  1. dup Vd.T, wzr seems to be used for code generation of Vector64/128.Zero which I thought was fixed with Implement Vector{Size}<T>.AllBitsSet #33924 (cc @Gnbrkm41). I will follow up on this

cc @kunalspathak@BruceForstall

@echesakov
echesakov marked this pull request as ready for review April 30, 2020 19:35
return AdvSimd.Insert(Vector64<byte>.Zero, 0, value);
}

return SoftwareFallback(value);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Curious about wrapping the SoftwareFallback() in a static method. What is the advantage of doing it?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure, actually. I did this to be consistent with existing Vector128/256 implementations.
@tannergooding Do you know why?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With the logic for all the various paths the method is too large for the normal inlining heuristics to work so we need to mark it AggressiveInlining (since the accelerated paths will generally be pretty small).
However, we don't necessarily want the SoftwareFallback to be inlined as that may not be beneficial.
Putting it in its own method prevents it from being inlined in the normal case and allows the JIT to decide if it is "too large or not" by itself.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That make sense.

@kunalspathakkunalspathak left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@echesakov
echesakov merged commit 670bf21 into dotnet:masterApr 30, 2020
@echesakov
echesakov deleted the Arm64-ASIMD-Vector64-Vector128-CreateScalar-Use-AdvSimd-Insert branch April 30, 2020 23:35
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@echesakov@tannergooding@kunalspathak