Optimize scalar conversions with AVX512 - #84384

Merged
tannergooding merged 41 commits into
dotnet:mainfrom
khushal1996:avx512-scalar-convert-rebased
Jul 16, 2023
Merged

Optimize scalar conversions with AVX512#84384
tannergooding merged 41 commits into
dotnet:mainfrom
khushal1996:avx512-scalar-convert-rebased

Conversation

@khushal1996

@khushal1996khushal1996 commented Apr 5, 2023

Copy link
Copy Markdown
Member

This PR optimize the following cases:


CasePrevious CodeOptimized Instruction
ulong -> doublevcvtsi2sdvcvtusi2sd
publicstaticdoubleUIntToDouble(UInt64val){return(double)val;}

Assembly before optimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F17C0857C0 vxorps xmm0,xmm0 62F1FF082AC1 vcvtsi2sd xmm0,rcx 4885C9 testrcx,rcx 7D0A jge SHORT G_M33997_IG03 62F1FF08580502000000 vaddsd xmm0, qword ptr [reloc @RWD00] ;; size=27 bbWeight=1 PerfScore 12.58G_M33997_IG03: ;; offset=001EH C3 ret

Assembly after optimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F1FF087BC1 vcvtusi2sdxmm0,rcx ;; size=6 bbWeight=1 PerfScore 4.00G_M33997_IG03: ;; offset=0009H C3 ret
CasePrevious CodeOptimized Instruction
ulong -> floatulong->double->floatvcvttusi2ss
[MethodImplAttribute(MethodImplOptions.NoInlining)]publicstaticfloatConvUlongToFloat(ulongval){return(float)val;}

Assembly before optimization

G_M2883_IG01: ;; offset=0000Hpushrbpsubrsp,48vzeroupperlearbp,[rsp+30H]xoreax,eaxmov dword ptr [rbp-04H],eaxmov qword ptr [rbp+10H],rcx ;; size=22 bbWeight=1 PerfScore 5.00G_M2883_IG02: ;; offset=0016Hcmp dword ptr [(reloc 0x7ff8bc0d2898)],0je SHORT G_M2883_IG04 ;; size=9 bbWeight=1 PerfScore 4.00G_M2883_IG03: ;; offset=001FHcall CORINFO_HELP_DBG_IS_JUST_MY_CODE ;; size=5 bbWeight=0.50 PerfScore 0.50G_M2883_IG04: ;; offset=0024Hnopmovrax, qword ptr [rbp+10H]vcvtusi2sdxmm0,rax vcvtsd2ss xmm0,xmm0,xmm0vmovss dword ptr [rbp-04H],xmm0nop ;; size=25 bbWeight=1 PerfScore 10.50G_M2883_IG05: ;; offset=003DHvmovssxmm0, dword ptr [rbp-04H] ;; size=7 bbWeight=1 PerfScore 3.00G_M2883_IG06: ;; offset=0044Haddrsp,48poprbpret ;; size=6 bbWeight=1 PerfScore 1.75

Assembly after optimization

G_M2883_IG01: ;; offset=0000Hpushrbpsubrsp,48vzeroupperlearbp,[rsp+30H]xoreax,eaxmov dword ptr [rbp-04H],eaxmov qword ptr [rbp+10H],rcx ;; size=22 bbWeight=1 PerfScore 5.00G_M2883_IG02: ;; offset=0016Hcmp dword ptr [(reloc 0x7ff8b54b2898)],0je SHORT G_M2883_IG04 ;; size=9 bbWeight=1 PerfScore 4.00G_M2883_IG03: ;; offset=001FHcall CORINFO_HELP_DBG_IS_JUST_MY_CODE ;; size=5 bbWeight=0.50 PerfScore 0.50G_M2883_IG04: ;; offset=0024Hnopmovrax, qword ptr [rbp+10H]vcvtusi2ssxmm0,raxvmovss dword ptr [rbp-04H],xmm0nop ;; size=19 bbWeight=1 PerfScore 8.50G_M2883_IG05: ;; offset=0037Hvmovssxmm0, dword ptr [rbp-04H] ;; size=7 bbWeight=1 PerfScore 3.00G_M2883_IG06: ;; offset=003EHaddrsp,48poprbpret ;; size=6 bbWeight=1 PerfScore 1.75

@ghostghost added area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI community-contribution Indicates that the PR has been added by a community member labels Apr 5, 2023
@ghost

ghost commented Apr 5, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch, @kunalspathak
See info in area-owners.md if you want to be subscribed.

Issue Details

Draft PR for testing purposes. No need for review at this time.

This PR optimize the following cases:


CasePrevious CodeOptimized Instruction
float -> ulongCORINFO_HELP_DBL2ULNG Helpervcvttss2usi
publicstaticUInt64FloatToULong(floatval){return(UInt64)val;}

Assembly before optimization

G_M22196_IG01: ;; offset=0000H 4883EC28 subrsp,40 C5F877 vzeroupper ;; size=7 bbWeight=1 PerfScore 1.25G_M22196_IG02: ;; offset=0007H 62F17E085AC0 vcvtss2sd xmm0,xmm0 E87E57815E call CORINFO_HELP_DBL2ULNG90nop ;; size=12 bbWeight=1 PerfScore 5.25G_M22196_IG03: ;; offset=0013H 4883C428 addrsp,40 C3 ret ;; size=5 bbWeight=1 PerfScore 1.25

Assembly afteroptimization

G_M22196_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M22196_IG02: ;; offset=0003H 62F1FE0878C0 vcvttss2usirax,xmm0 ;; size=6 bbWeight=1 PerfScore 6.00G_M22196_IG03: ;; offset=0009H C3 ret

CasePrevious CodeOptimized Instruction
double -> ulongCORINFO_HELP_DBL2ULNG Helpervcvttsd2usi
publicstaticUInt64DoubleToULong(doubleval){return(UInt64)val;}

Assembly before optimization

G_M30068_IG01: ;; offset=0000H 4883EC28 subrsp,40 C5F877 vzeroupper ;; size=7 bbWeight=1 PerfScore 1.25G_M30068_IG02: ;; offset=0007H E874577F5E call CORINFO_HELP_DBL2ULNG90nop ;; size=6 bbWeight=1 PerfScore 1.25G_M30068_IG03: ;; offset=000DH 4883C428 addrsp,40 C3 ret

Assembly afteroptimization

G_M30068_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M30068_IG02: ;; offset=0003H 62F1FF0878C0 vcvttsd2usirax,xmm0 ;; size=6 bbWeight=1 PerfScore 5.00G_M30068_IG03: ;; offset=0009H C3 ret

CasePrevious CodeOptimized Instruction
ulong -> doublevcvtsi2sdvcvtusi2sd
publicstaticdoubleUIntToDouble(UInt64val){return(double)val;}

Assembly before optimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F17C0857C0 vxorps xmm0,xmm0 62F1FF082AC1 vcvtsi2sd xmm0,rcx 4885C9 testrcx,rcx 7D0A jge SHORT G_M33997_IG03 62F1FF08580502000000 vaddsd xmm0, qword ptr [reloc @RWD00] ;; size=27 bbWeight=1 PerfScore 12.58G_M33997_IG03: ;; offset=001EH C3 ret

Assembly afteroptimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F1FF087BC1 vcvtusi2sdxmm0,rcx ;; size=6 bbWeight=1 PerfScore 4.00G_M33997_IG03: ;; offset=0009H C3 ret
Author:khushal1996
Assignees:-
Labels:

area-CodeGen-coreclr

Milestone:-

@khushal1996khushal1996 changed the title Avx512 scalar convert rebasedOptimize scalar conversions with AVX512Apr 5, 2023
@BruceForstallBruceForstall added the avx512 Related to the AVX-512 architecture label Apr 7, 2023
@khushal1996

khushal1996 commented Apr 7, 2023

Copy link
Copy Markdown
MemberAuthor

I am making Submissions in the course of work for my employer (or my employer has intellectual property rights in my Submissions by contract or applicable law). I have permission from my employer to make Submissions and enter into this Agreement on behalf of my employer. By signing below, the defined term “You” includes me and my employer.

@dotnet-policy-service agree company="Intel"

Comment threadsrc/coreclr/vm/jithelpers.cpp Outdated
Comment threadsrc/tests/JIT/Directed/Convert/out_of_range_fp_to_int_conversions.cpp Outdated
Comment threadsrc/tests/JIT/Directed/Convert/out_of_range_fp_to_int_conversions.cpp Outdated
Comment threadsrc/libraries/System.Runtime/tests/System/UIntPtrTests.GenericMath.cs Outdated
Comment threadsrc/coreclr/jit/instrsxarch.h Outdated
Comment threadsrc/coreclr/vm/jithelpers.cpp Outdated
Comment threadsrc/libraries/System.Runtime/tests/System/UIntPtrTests.GenericMath.cs Outdated
Comment threadsrc/coreclr/vm/jithelpers.cpp Outdated
@khushal1996

Copy link
Copy Markdown
MemberAuthor

We are narrowing down the scope of the PR to ulong -> float conversions. The changes will b e updated soon.

@@ -18,8 +18,46 @@ FORCEINLINE int64_t FastDbl2Lng(double val)
#endif

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The changes in this file and in jithelpers.cpp should be rolled back too.

@khushal1996

Copy link
Copy Markdown
MemberAuthor

@tannergooding@jkotas does this look good for merge? I have rolled back the changes for float/double -> I long and optimized ulong -> float/double cases.

@jkotasjkotas left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks good to me. I have commented on a few nits.

Somebody on @dotnet/jit-contrib should do final review and merge.

Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/lowerxarch.cpp
Comment threadsrc/coreclr/jit/instr.cpp
Comment threadsrc/coreclr/jit/importer.cpp Outdated
@khushal1996

Copy link
Copy Markdown
MemberAuthor

@jkotas I think we are good to go here. Would you please help with the approval and merge.

@jkotas

Copy link
Copy Markdown
Member

@kunalspathak Could you please do final review and merge?

@khushal1996

Copy link
Copy Markdown
MemberAuthor

@kunalspathak can you help to merge this changes? They have been approved and we are trying to get them in before the next release.

Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated
@tannergooding

Copy link
Copy Markdown
Member

SPMI failures are the general No Azure Storage MCH files to download from cef79bc8-29bf-4f7b-9d05-9fc06832098c/osx/arm64/ impacting other PRs.

CI was passing before the minor formatting cleanup requested

@tannergooding
tannergooding merged commit b99a279 into dotnet:mainJul 16, 2023
@ghostghost locked as resolved and limited conversation to collaborators Aug 15, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIavx512Related to the AVX-512 architecturecommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@khushal1996@tannergooding@jkotas@EgorBo@BruceForstall@kunalspathak@SingleAccretion@anthonycanino
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Optimize scalar conversions with AVX512 - #84384

Merged
tannergooding merged 41 commits into
dotnet:mainfrom
khushal1996:avx512-scalar-convert-rebased
Jul 16, 2023
Merged

Optimize scalar conversions with AVX512#84384
tannergooding merged 41 commits into
dotnet:mainfrom
khushal1996:avx512-scalar-convert-rebased

Conversation

@khushal1996

@khushal1996khushal1996 commented Apr 5, 2023

Copy link
Copy Markdown
Member

This PR optimize the following cases:


CasePrevious CodeOptimized Instruction
ulong -> doublevcvtsi2sdvcvtusi2sd
publicstaticdoubleUIntToDouble(UInt64val){return(double)val;}

Assembly before optimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F17C0857C0 vxorps xmm0,xmm0 62F1FF082AC1 vcvtsi2sd xmm0,rcx 4885C9 testrcx,rcx 7D0A jge SHORT G_M33997_IG03 62F1FF08580502000000 vaddsd xmm0, qword ptr [reloc @RWD00] ;; size=27 bbWeight=1 PerfScore 12.58G_M33997_IG03: ;; offset=001EH C3 ret

Assembly after optimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F1FF087BC1 vcvtusi2sdxmm0,rcx ;; size=6 bbWeight=1 PerfScore 4.00G_M33997_IG03: ;; offset=0009H C3 ret
CasePrevious CodeOptimized Instruction
ulong -> floatulong->double->floatvcvttusi2ss
[MethodImplAttribute(MethodImplOptions.NoInlining)]publicstaticfloatConvUlongToFloat(ulongval){return(float)val;}

Assembly before optimization

G_M2883_IG01: ;; offset=0000Hpushrbpsubrsp,48vzeroupperlearbp,[rsp+30H]xoreax,eaxmov dword ptr [rbp-04H],eaxmov qword ptr [rbp+10H],rcx ;; size=22 bbWeight=1 PerfScore 5.00G_M2883_IG02: ;; offset=0016Hcmp dword ptr [(reloc 0x7ff8bc0d2898)],0je SHORT G_M2883_IG04 ;; size=9 bbWeight=1 PerfScore 4.00G_M2883_IG03: ;; offset=001FHcall CORINFO_HELP_DBG_IS_JUST_MY_CODE ;; size=5 bbWeight=0.50 PerfScore 0.50G_M2883_IG04: ;; offset=0024Hnopmovrax, qword ptr [rbp+10H]vcvtusi2sdxmm0,rax vcvtsd2ss xmm0,xmm0,xmm0vmovss dword ptr [rbp-04H],xmm0nop ;; size=25 bbWeight=1 PerfScore 10.50G_M2883_IG05: ;; offset=003DHvmovssxmm0, dword ptr [rbp-04H] ;; size=7 bbWeight=1 PerfScore 3.00G_M2883_IG06: ;; offset=0044Haddrsp,48poprbpret ;; size=6 bbWeight=1 PerfScore 1.75

Assembly after optimization

G_M2883_IG01: ;; offset=0000Hpushrbpsubrsp,48vzeroupperlearbp,[rsp+30H]xoreax,eaxmov dword ptr [rbp-04H],eaxmov qword ptr [rbp+10H],rcx ;; size=22 bbWeight=1 PerfScore 5.00G_M2883_IG02: ;; offset=0016Hcmp dword ptr [(reloc 0x7ff8b54b2898)],0je SHORT G_M2883_IG04 ;; size=9 bbWeight=1 PerfScore 4.00G_M2883_IG03: ;; offset=001FHcall CORINFO_HELP_DBG_IS_JUST_MY_CODE ;; size=5 bbWeight=0.50 PerfScore 0.50G_M2883_IG04: ;; offset=0024Hnopmovrax, qword ptr [rbp+10H]vcvtusi2ssxmm0,raxvmovss dword ptr [rbp-04H],xmm0nop ;; size=19 bbWeight=1 PerfScore 8.50G_M2883_IG05: ;; offset=0037Hvmovssxmm0, dword ptr [rbp-04H] ;; size=7 bbWeight=1 PerfScore 3.00G_M2883_IG06: ;; offset=003EHaddrsp,48poprbpret ;; size=6 bbWeight=1 PerfScore 1.75

@ghostghost added area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI community-contribution Indicates that the PR has been added by a community member labels Apr 5, 2023
@ghost

ghost commented Apr 5, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch, @kunalspathak
See info in area-owners.md if you want to be subscribed.

Issue Details

Draft PR for testing purposes. No need for review at this time.

This PR optimize the following cases:


CasePrevious CodeOptimized Instruction
float -> ulongCORINFO_HELP_DBL2ULNG Helpervcvttss2usi
publicstaticUInt64FloatToULong(floatval){return(UInt64)val;}

Assembly before optimization

G_M22196_IG01: ;; offset=0000H 4883EC28 subrsp,40 C5F877 vzeroupper ;; size=7 bbWeight=1 PerfScore 1.25G_M22196_IG02: ;; offset=0007H 62F17E085AC0 vcvtss2sd xmm0,xmm0 E87E57815E call CORINFO_HELP_DBL2ULNG90nop ;; size=12 bbWeight=1 PerfScore 5.25G_M22196_IG03: ;; offset=0013H 4883C428 addrsp,40 C3 ret ;; size=5 bbWeight=1 PerfScore 1.25

Assembly afteroptimization

G_M22196_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M22196_IG02: ;; offset=0003H 62F1FE0878C0 vcvttss2usirax,xmm0 ;; size=6 bbWeight=1 PerfScore 6.00G_M22196_IG03: ;; offset=0009H C3 ret

CasePrevious CodeOptimized Instruction
double -> ulongCORINFO_HELP_DBL2ULNG Helpervcvttsd2usi
publicstaticUInt64DoubleToULong(doubleval){return(UInt64)val;}

Assembly before optimization

G_M30068_IG01: ;; offset=0000H 4883EC28 subrsp,40 C5F877 vzeroupper ;; size=7 bbWeight=1 PerfScore 1.25G_M30068_IG02: ;; offset=0007H E874577F5E call CORINFO_HELP_DBL2ULNG90nop ;; size=6 bbWeight=1 PerfScore 1.25G_M30068_IG03: ;; offset=000DH 4883C428 addrsp,40 C3 ret

Assembly afteroptimization

G_M30068_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M30068_IG02: ;; offset=0003H 62F1FF0878C0 vcvttsd2usirax,xmm0 ;; size=6 bbWeight=1 PerfScore 5.00G_M30068_IG03: ;; offset=0009H C3 ret

CasePrevious CodeOptimized Instruction
ulong -> doublevcvtsi2sdvcvtusi2sd
publicstaticdoubleUIntToDouble(UInt64val){return(double)val;}

Assembly before optimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F17C0857C0 vxorps xmm0,xmm0 62F1FF082AC1 vcvtsi2sd xmm0,rcx 4885C9 testrcx,rcx 7D0A jge SHORT G_M33997_IG03 62F1FF08580502000000 vaddsd xmm0, qword ptr [reloc @RWD00] ;; size=27 bbWeight=1 PerfScore 12.58G_M33997_IG03: ;; offset=001EH C3 ret

Assembly afteroptimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F1FF087BC1 vcvtusi2sdxmm0,rcx ;; size=6 bbWeight=1 PerfScore 4.00G_M33997_IG03: ;; offset=0009H C3 ret
Author:khushal1996
Assignees:-
Labels:

area-CodeGen-coreclr

Milestone:-

@khushal1996khushal1996 changed the title Avx512 scalar convert rebasedOptimize scalar conversions with AVX512Apr 5, 2023
@BruceForstallBruceForstall added the avx512 Related to the AVX-512 architecture label Apr 7, 2023
@khushal1996

khushal1996 commented Apr 7, 2023

Copy link
Copy Markdown
MemberAuthor

I am making Submissions in the course of work for my employer (or my employer has intellectual property rights in my Submissions by contract or applicable law). I have permission from my employer to make Submissions and enter into this Agreement on behalf of my employer. By signing below, the defined term “You” includes me and my employer.

@dotnet-policy-service agree company="Intel"

Comment threadsrc/coreclr/vm/jithelpers.cpp Outdated
Comment threadsrc/tests/JIT/Directed/Convert/out_of_range_fp_to_int_conversions.cpp Outdated
Comment threadsrc/tests/JIT/Directed/Convert/out_of_range_fp_to_int_conversions.cpp Outdated
Comment threadsrc/libraries/System.Runtime/tests/System/UIntPtrTests.GenericMath.cs Outdated
Comment threadsrc/coreclr/jit/instrsxarch.h Outdated
Comment threadsrc/coreclr/vm/jithelpers.cpp Outdated
Comment threadsrc/libraries/System.Runtime/tests/System/UIntPtrTests.GenericMath.cs Outdated
Comment threadsrc/coreclr/vm/jithelpers.cpp Outdated
@khushal1996

Copy link
Copy Markdown
MemberAuthor

We are narrowing down the scope of the PR to ulong -> float conversions. The changes will b e updated soon.

@@ -18,8 +18,46 @@ FORCEINLINE int64_t FastDbl2Lng(double val)
#endif

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The changes in this file and in jithelpers.cpp should be rolled back too.

@khushal1996

Copy link
Copy Markdown
MemberAuthor

@tannergooding@jkotas does this look good for merge? I have rolled back the changes for float/double -> I long and optimized ulong -> float/double cases.

@jkotasjkotas left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks good to me. I have commented on a few nits.

Somebody on @dotnet/jit-contrib should do final review and merge.

Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/lowerxarch.cpp
Comment threadsrc/coreclr/jit/instr.cpp
Comment threadsrc/coreclr/jit/importer.cpp Outdated
@khushal1996

Copy link
Copy Markdown
MemberAuthor

@jkotas I think we are good to go here. Would you please help with the approval and merge.

@jkotas

Copy link
Copy Markdown
Member

@kunalspathak Could you please do final review and merge?

@khushal1996

Copy link
Copy Markdown
MemberAuthor

@kunalspathak can you help to merge this changes? They have been approved and we are trying to get them in before the next release.

Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated
@tannergooding

Copy link
Copy Markdown
Member

SPMI failures are the general No Azure Storage MCH files to download from cef79bc8-29bf-4f7b-9d05-9fc06832098c/osx/arm64/ impacting other PRs.

CI was passing before the minor formatting cleanup requested

@tannergooding
tannergooding merged commit b99a279 into dotnet:mainJul 16, 2023
@ghostghost locked as resolved and limited conversation to collaborators Aug 15, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIavx512Related to the AVX-512 architecturecommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@khushal1996@tannergooding@jkotas@EgorBo@BruceForstall@kunalspathak@SingleAccretion@anthonycanino
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Optimize scalar conversions with AVX512 - #84384

Merged
tannergooding merged 41 commits into
dotnet:mainfrom
khushal1996:avx512-scalar-convert-rebased
Jul 16, 2023
Merged

Optimize scalar conversions with AVX512#84384
tannergooding merged 41 commits into
dotnet:mainfrom
khushal1996:avx512-scalar-convert-rebased

Conversation

@khushal1996

@khushal1996khushal1996 commented Apr 5, 2023

Copy link
Copy Markdown
Member

This PR optimize the following cases:


CasePrevious CodeOptimized Instruction
ulong -> doublevcvtsi2sdvcvtusi2sd
publicstaticdoubleUIntToDouble(UInt64val){return(double)val;}

Assembly before optimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F17C0857C0 vxorps xmm0,xmm0 62F1FF082AC1 vcvtsi2sd xmm0,rcx 4885C9 testrcx,rcx 7D0A jge SHORT G_M33997_IG03 62F1FF08580502000000 vaddsd xmm0, qword ptr [reloc @RWD00] ;; size=27 bbWeight=1 PerfScore 12.58G_M33997_IG03: ;; offset=001EH C3 ret

Assembly after optimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F1FF087BC1 vcvtusi2sdxmm0,rcx ;; size=6 bbWeight=1 PerfScore 4.00G_M33997_IG03: ;; offset=0009H C3 ret
CasePrevious CodeOptimized Instruction
ulong -> floatulong->double->floatvcvttusi2ss
[MethodImplAttribute(MethodImplOptions.NoInlining)]publicstaticfloatConvUlongToFloat(ulongval){return(float)val;}

Assembly before optimization

G_M2883_IG01: ;; offset=0000Hpushrbpsubrsp,48vzeroupperlearbp,[rsp+30H]xoreax,eaxmov dword ptr [rbp-04H],eaxmov qword ptr [rbp+10H],rcx ;; size=22 bbWeight=1 PerfScore 5.00G_M2883_IG02: ;; offset=0016Hcmp dword ptr [(reloc 0x7ff8bc0d2898)],0je SHORT G_M2883_IG04 ;; size=9 bbWeight=1 PerfScore 4.00G_M2883_IG03: ;; offset=001FHcall CORINFO_HELP_DBG_IS_JUST_MY_CODE ;; size=5 bbWeight=0.50 PerfScore 0.50G_M2883_IG04: ;; offset=0024Hnopmovrax, qword ptr [rbp+10H]vcvtusi2sdxmm0,rax vcvtsd2ss xmm0,xmm0,xmm0vmovss dword ptr [rbp-04H],xmm0nop ;; size=25 bbWeight=1 PerfScore 10.50G_M2883_IG05: ;; offset=003DHvmovssxmm0, dword ptr [rbp-04H] ;; size=7 bbWeight=1 PerfScore 3.00G_M2883_IG06: ;; offset=0044Haddrsp,48poprbpret ;; size=6 bbWeight=1 PerfScore 1.75

Assembly after optimization

G_M2883_IG01: ;; offset=0000Hpushrbpsubrsp,48vzeroupperlearbp,[rsp+30H]xoreax,eaxmov dword ptr [rbp-04H],eaxmov qword ptr [rbp+10H],rcx ;; size=22 bbWeight=1 PerfScore 5.00G_M2883_IG02: ;; offset=0016Hcmp dword ptr [(reloc 0x7ff8b54b2898)],0je SHORT G_M2883_IG04 ;; size=9 bbWeight=1 PerfScore 4.00G_M2883_IG03: ;; offset=001FHcall CORINFO_HELP_DBG_IS_JUST_MY_CODE ;; size=5 bbWeight=0.50 PerfScore 0.50G_M2883_IG04: ;; offset=0024Hnopmovrax, qword ptr [rbp+10H]vcvtusi2ssxmm0,raxvmovss dword ptr [rbp-04H],xmm0nop ;; size=19 bbWeight=1 PerfScore 8.50G_M2883_IG05: ;; offset=0037Hvmovssxmm0, dword ptr [rbp-04H] ;; size=7 bbWeight=1 PerfScore 3.00G_M2883_IG06: ;; offset=003EHaddrsp,48poprbpret ;; size=6 bbWeight=1 PerfScore 1.75

@ghostghost added area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI community-contribution Indicates that the PR has been added by a community member labels Apr 5, 2023
@ghost

ghost commented Apr 5, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch, @kunalspathak
See info in area-owners.md if you want to be subscribed.

Issue Details

Draft PR for testing purposes. No need for review at this time.

This PR optimize the following cases:


CasePrevious CodeOptimized Instruction
float -> ulongCORINFO_HELP_DBL2ULNG Helpervcvttss2usi
publicstaticUInt64FloatToULong(floatval){return(UInt64)val;}

Assembly before optimization

G_M22196_IG01: ;; offset=0000H 4883EC28 subrsp,40 C5F877 vzeroupper ;; size=7 bbWeight=1 PerfScore 1.25G_M22196_IG02: ;; offset=0007H 62F17E085AC0 vcvtss2sd xmm0,xmm0 E87E57815E call CORINFO_HELP_DBL2ULNG90nop ;; size=12 bbWeight=1 PerfScore 5.25G_M22196_IG03: ;; offset=0013H 4883C428 addrsp,40 C3 ret ;; size=5 bbWeight=1 PerfScore 1.25

Assembly afteroptimization

G_M22196_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M22196_IG02: ;; offset=0003H 62F1FE0878C0 vcvttss2usirax,xmm0 ;; size=6 bbWeight=1 PerfScore 6.00G_M22196_IG03: ;; offset=0009H C3 ret

CasePrevious CodeOptimized Instruction
double -> ulongCORINFO_HELP_DBL2ULNG Helpervcvttsd2usi
publicstaticUInt64DoubleToULong(doubleval){return(UInt64)val;}

Assembly before optimization

G_M30068_IG01: ;; offset=0000H 4883EC28 subrsp,40 C5F877 vzeroupper ;; size=7 bbWeight=1 PerfScore 1.25G_M30068_IG02: ;; offset=0007H E874577F5E call CORINFO_HELP_DBL2ULNG90nop ;; size=6 bbWeight=1 PerfScore 1.25G_M30068_IG03: ;; offset=000DH 4883C428 addrsp,40 C3 ret

Assembly afteroptimization

G_M30068_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M30068_IG02: ;; offset=0003H 62F1FF0878C0 vcvttsd2usirax,xmm0 ;; size=6 bbWeight=1 PerfScore 5.00G_M30068_IG03: ;; offset=0009H C3 ret

CasePrevious CodeOptimized Instruction
ulong -> doublevcvtsi2sdvcvtusi2sd
publicstaticdoubleUIntToDouble(UInt64val){return(double)val;}

Assembly before optimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F17C0857C0 vxorps xmm0,xmm0 62F1FF082AC1 vcvtsi2sd xmm0,rcx 4885C9 testrcx,rcx 7D0A jge SHORT G_M33997_IG03 62F1FF08580502000000 vaddsd xmm0, qword ptr [reloc @RWD00] ;; size=27 bbWeight=1 PerfScore 12.58G_M33997_IG03: ;; offset=001EH C3 ret

Assembly afteroptimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F1FF087BC1 vcvtusi2sdxmm0,rcx ;; size=6 bbWeight=1 PerfScore 4.00G_M33997_IG03: ;; offset=0009H C3 ret
Author:khushal1996
Assignees:-
Labels:

area-CodeGen-coreclr

Milestone:-

@khushal1996khushal1996 changed the title Avx512 scalar convert rebasedOptimize scalar conversions with AVX512Apr 5, 2023
@BruceForstallBruceForstall added the avx512 Related to the AVX-512 architecture label Apr 7, 2023
@khushal1996

khushal1996 commented Apr 7, 2023

Copy link
Copy Markdown
MemberAuthor

I am making Submissions in the course of work for my employer (or my employer has intellectual property rights in my Submissions by contract or applicable law). I have permission from my employer to make Submissions and enter into this Agreement on behalf of my employer. By signing below, the defined term “You” includes me and my employer.

@dotnet-policy-service agree company="Intel"

Comment threadsrc/coreclr/vm/jithelpers.cpp Outdated
Comment threadsrc/tests/JIT/Directed/Convert/out_of_range_fp_to_int_conversions.cpp Outdated
Comment threadsrc/tests/JIT/Directed/Convert/out_of_range_fp_to_int_conversions.cpp Outdated
Comment threadsrc/libraries/System.Runtime/tests/System/UIntPtrTests.GenericMath.cs Outdated
Comment threadsrc/coreclr/jit/instrsxarch.h Outdated
Comment threadsrc/coreclr/vm/jithelpers.cpp Outdated
Comment threadsrc/libraries/System.Runtime/tests/System/UIntPtrTests.GenericMath.cs Outdated
Comment threadsrc/coreclr/vm/jithelpers.cpp Outdated
@khushal1996

Copy link
Copy Markdown
MemberAuthor

We are narrowing down the scope of the PR to ulong -> float conversions. The changes will b e updated soon.

@@ -18,8 +18,46 @@ FORCEINLINE int64_t FastDbl2Lng(double val)
#endif

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The changes in this file and in jithelpers.cpp should be rolled back too.

@khushal1996

Copy link
Copy Markdown
MemberAuthor

@tannergooding@jkotas does this look good for merge? I have rolled back the changes for float/double -> I long and optimized ulong -> float/double cases.

@jkotasjkotas left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks good to me. I have commented on a few nits.

Somebody on @dotnet/jit-contrib should do final review and merge.

Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/lowerxarch.cpp
Comment threadsrc/coreclr/jit/instr.cpp
Comment threadsrc/coreclr/jit/importer.cpp Outdated
@khushal1996

Copy link
Copy Markdown
MemberAuthor

@jkotas I think we are good to go here. Would you please help with the approval and merge.

@jkotas

Copy link
Copy Markdown
Member

@kunalspathak Could you please do final review and merge?

@khushal1996

Copy link
Copy Markdown
MemberAuthor

@kunalspathak can you help to merge this changes? They have been approved and we are trying to get them in before the next release.

Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated
@tannergooding

Copy link
Copy Markdown
Member

SPMI failures are the general No Azure Storage MCH files to download from cef79bc8-29bf-4f7b-9d05-9fc06832098c/osx/arm64/ impacting other PRs.

CI was passing before the minor formatting cleanup requested

@tannergooding
tannergooding merged commit b99a279 into dotnet:mainJul 16, 2023
@ghostghost locked as resolved and limited conversation to collaborators Aug 15, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIavx512Related to the AVX-512 architecturecommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@khushal1996@tannergooding@jkotas@EgorBo@BruceForstall@kunalspathak@SingleAccretion@anthonycanino
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Optimize scalar conversions with AVX512 - #84384

Merged
tannergooding merged 41 commits into
dotnet:mainfrom
khushal1996:avx512-scalar-convert-rebased
Jul 16, 2023
Merged

Optimize scalar conversions with AVX512#84384
tannergooding merged 41 commits into
dotnet:mainfrom
khushal1996:avx512-scalar-convert-rebased

Conversation

@khushal1996

@khushal1996khushal1996 commented Apr 5, 2023

Copy link
Copy Markdown
Member

This PR optimize the following cases:


CasePrevious CodeOptimized Instruction
ulong -> doublevcvtsi2sdvcvtusi2sd
publicstaticdoubleUIntToDouble(UInt64val){return(double)val;}

Assembly before optimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F17C0857C0 vxorps xmm0,xmm0 62F1FF082AC1 vcvtsi2sd xmm0,rcx 4885C9 testrcx,rcx 7D0A jge SHORT G_M33997_IG03 62F1FF08580502000000 vaddsd xmm0, qword ptr [reloc @RWD00] ;; size=27 bbWeight=1 PerfScore 12.58G_M33997_IG03: ;; offset=001EH C3 ret

Assembly after optimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F1FF087BC1 vcvtusi2sdxmm0,rcx ;; size=6 bbWeight=1 PerfScore 4.00G_M33997_IG03: ;; offset=0009H C3 ret
CasePrevious CodeOptimized Instruction
ulong -> floatulong->double->floatvcvttusi2ss
[MethodImplAttribute(MethodImplOptions.NoInlining)]publicstaticfloatConvUlongToFloat(ulongval){return(float)val;}

Assembly before optimization

G_M2883_IG01: ;; offset=0000Hpushrbpsubrsp,48vzeroupperlearbp,[rsp+30H]xoreax,eaxmov dword ptr [rbp-04H],eaxmov qword ptr [rbp+10H],rcx ;; size=22 bbWeight=1 PerfScore 5.00G_M2883_IG02: ;; offset=0016Hcmp dword ptr [(reloc 0x7ff8bc0d2898)],0je SHORT G_M2883_IG04 ;; size=9 bbWeight=1 PerfScore 4.00G_M2883_IG03: ;; offset=001FHcall CORINFO_HELP_DBG_IS_JUST_MY_CODE ;; size=5 bbWeight=0.50 PerfScore 0.50G_M2883_IG04: ;; offset=0024Hnopmovrax, qword ptr [rbp+10H]vcvtusi2sdxmm0,rax vcvtsd2ss xmm0,xmm0,xmm0vmovss dword ptr [rbp-04H],xmm0nop ;; size=25 bbWeight=1 PerfScore 10.50G_M2883_IG05: ;; offset=003DHvmovssxmm0, dword ptr [rbp-04H] ;; size=7 bbWeight=1 PerfScore 3.00G_M2883_IG06: ;; offset=0044Haddrsp,48poprbpret ;; size=6 bbWeight=1 PerfScore 1.75

Assembly after optimization

G_M2883_IG01: ;; offset=0000Hpushrbpsubrsp,48vzeroupperlearbp,[rsp+30H]xoreax,eaxmov dword ptr [rbp-04H],eaxmov qword ptr [rbp+10H],rcx ;; size=22 bbWeight=1 PerfScore 5.00G_M2883_IG02: ;; offset=0016Hcmp dword ptr [(reloc 0x7ff8b54b2898)],0je SHORT G_M2883_IG04 ;; size=9 bbWeight=1 PerfScore 4.00G_M2883_IG03: ;; offset=001FHcall CORINFO_HELP_DBG_IS_JUST_MY_CODE ;; size=5 bbWeight=0.50 PerfScore 0.50G_M2883_IG04: ;; offset=0024Hnopmovrax, qword ptr [rbp+10H]vcvtusi2ssxmm0,raxvmovss dword ptr [rbp-04H],xmm0nop ;; size=19 bbWeight=1 PerfScore 8.50G_M2883_IG05: ;; offset=0037Hvmovssxmm0, dword ptr [rbp-04H] ;; size=7 bbWeight=1 PerfScore 3.00G_M2883_IG06: ;; offset=003EHaddrsp,48poprbpret ;; size=6 bbWeight=1 PerfScore 1.75

@ghostghost added area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI community-contribution Indicates that the PR has been added by a community member labels Apr 5, 2023
@ghost

ghost commented Apr 5, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch, @kunalspathak
See info in area-owners.md if you want to be subscribed.

Issue Details

Draft PR for testing purposes. No need for review at this time.

This PR optimize the following cases:


CasePrevious CodeOptimized Instruction
float -> ulongCORINFO_HELP_DBL2ULNG Helpervcvttss2usi
publicstaticUInt64FloatToULong(floatval){return(UInt64)val;}

Assembly before optimization

G_M22196_IG01: ;; offset=0000H 4883EC28 subrsp,40 C5F877 vzeroupper ;; size=7 bbWeight=1 PerfScore 1.25G_M22196_IG02: ;; offset=0007H 62F17E085AC0 vcvtss2sd xmm0,xmm0 E87E57815E call CORINFO_HELP_DBL2ULNG90nop ;; size=12 bbWeight=1 PerfScore 5.25G_M22196_IG03: ;; offset=0013H 4883C428 addrsp,40 C3 ret ;; size=5 bbWeight=1 PerfScore 1.25

Assembly afteroptimization

G_M22196_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M22196_IG02: ;; offset=0003H 62F1FE0878C0 vcvttss2usirax,xmm0 ;; size=6 bbWeight=1 PerfScore 6.00G_M22196_IG03: ;; offset=0009H C3 ret

CasePrevious CodeOptimized Instruction
double -> ulongCORINFO_HELP_DBL2ULNG Helpervcvttsd2usi
publicstaticUInt64DoubleToULong(doubleval){return(UInt64)val;}

Assembly before optimization

G_M30068_IG01: ;; offset=0000H 4883EC28 subrsp,40 C5F877 vzeroupper ;; size=7 bbWeight=1 PerfScore 1.25G_M30068_IG02: ;; offset=0007H E874577F5E call CORINFO_HELP_DBL2ULNG90nop ;; size=6 bbWeight=1 PerfScore 1.25G_M30068_IG03: ;; offset=000DH 4883C428 addrsp,40 C3 ret

Assembly afteroptimization

G_M30068_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M30068_IG02: ;; offset=0003H 62F1FF0878C0 vcvttsd2usirax,xmm0 ;; size=6 bbWeight=1 PerfScore 5.00G_M30068_IG03: ;; offset=0009H C3 ret

CasePrevious CodeOptimized Instruction
ulong -> doublevcvtsi2sdvcvtusi2sd
publicstaticdoubleUIntToDouble(UInt64val){return(double)val;}

Assembly before optimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F17C0857C0 vxorps xmm0,xmm0 62F1FF082AC1 vcvtsi2sd xmm0,rcx 4885C9 testrcx,rcx 7D0A jge SHORT G_M33997_IG03 62F1FF08580502000000 vaddsd xmm0, qword ptr [reloc @RWD00] ;; size=27 bbWeight=1 PerfScore 12.58G_M33997_IG03: ;; offset=001EH C3 ret

Assembly afteroptimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F1FF087BC1 vcvtusi2sdxmm0,rcx ;; size=6 bbWeight=1 PerfScore 4.00G_M33997_IG03: ;; offset=0009H C3 ret
Author:khushal1996
Assignees:-
Labels:

area-CodeGen-coreclr

Milestone:-

@khushal1996khushal1996 changed the title Avx512 scalar convert rebasedOptimize scalar conversions with AVX512Apr 5, 2023
@BruceForstallBruceForstall added the avx512 Related to the AVX-512 architecture label Apr 7, 2023
@khushal1996

khushal1996 commented Apr 7, 2023

Copy link
Copy Markdown
MemberAuthor

I am making Submissions in the course of work for my employer (or my employer has intellectual property rights in my Submissions by contract or applicable law). I have permission from my employer to make Submissions and enter into this Agreement on behalf of my employer. By signing below, the defined term “You” includes me and my employer.

@dotnet-policy-service agree company="Intel"

Comment threadsrc/coreclr/vm/jithelpers.cpp Outdated
Comment threadsrc/tests/JIT/Directed/Convert/out_of_range_fp_to_int_conversions.cpp Outdated
Comment threadsrc/tests/JIT/Directed/Convert/out_of_range_fp_to_int_conversions.cpp Outdated
Comment threadsrc/libraries/System.Runtime/tests/System/UIntPtrTests.GenericMath.cs Outdated
Comment threadsrc/coreclr/jit/instrsxarch.h Outdated
Comment threadsrc/coreclr/vm/jithelpers.cpp Outdated
Comment threadsrc/libraries/System.Runtime/tests/System/UIntPtrTests.GenericMath.cs Outdated
Comment threadsrc/coreclr/vm/jithelpers.cpp Outdated
@khushal1996

Copy link
Copy Markdown
MemberAuthor

We are narrowing down the scope of the PR to ulong -> float conversions. The changes will b e updated soon.

@@ -18,8 +18,46 @@ FORCEINLINE int64_t FastDbl2Lng(double val)
#endif

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The changes in this file and in jithelpers.cpp should be rolled back too.

@khushal1996

Copy link
Copy Markdown
MemberAuthor

@tannergooding@jkotas does this look good for merge? I have rolled back the changes for float/double -> I long and optimized ulong -> float/double cases.

@jkotasjkotas left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks good to me. I have commented on a few nits.

Somebody on @dotnet/jit-contrib should do final review and merge.

Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/lowerxarch.cpp
Comment threadsrc/coreclr/jit/instr.cpp
Comment threadsrc/coreclr/jit/importer.cpp Outdated
@khushal1996

Copy link
Copy Markdown
MemberAuthor

@jkotas I think we are good to go here. Would you please help with the approval and merge.

@jkotas

Copy link
Copy Markdown
Member

@kunalspathak Could you please do final review and merge?

@khushal1996

Copy link
Copy Markdown
MemberAuthor

@kunalspathak can you help to merge this changes? They have been approved and we are trying to get them in before the next release.

Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated
@tannergooding

Copy link
Copy Markdown
Member

SPMI failures are the general No Azure Storage MCH files to download from cef79bc8-29bf-4f7b-9d05-9fc06832098c/osx/arm64/ impacting other PRs.

CI was passing before the minor formatting cleanup requested

@tannergooding
tannergooding merged commit b99a279 into dotnet:mainJul 16, 2023
@ghostghost locked as resolved and limited conversation to collaborators Aug 15, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIavx512Related to the AVX-512 architecturecommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@khushal1996@tannergooding@jkotas@EgorBo@BruceForstall@kunalspathak@SingleAccretion@anthonycanino
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Optimize scalar conversions with AVX512 - #84384

Merged
tannergooding merged 41 commits into
dotnet:mainfrom
khushal1996:avx512-scalar-convert-rebased
Jul 16, 2023
Merged

Optimize scalar conversions with AVX512#84384
tannergooding merged 41 commits into
dotnet:mainfrom
khushal1996:avx512-scalar-convert-rebased

Conversation

@khushal1996

@khushal1996khushal1996 commented Apr 5, 2023

Copy link
Copy Markdown
Member

This PR optimize the following cases:


CasePrevious CodeOptimized Instruction
ulong -> doublevcvtsi2sdvcvtusi2sd
publicstaticdoubleUIntToDouble(UInt64val){return(double)val;}

Assembly before optimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F17C0857C0 vxorps xmm0,xmm0 62F1FF082AC1 vcvtsi2sd xmm0,rcx 4885C9 testrcx,rcx 7D0A jge SHORT G_M33997_IG03 62F1FF08580502000000 vaddsd xmm0, qword ptr [reloc @RWD00] ;; size=27 bbWeight=1 PerfScore 12.58G_M33997_IG03: ;; offset=001EH C3 ret

Assembly after optimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F1FF087BC1 vcvtusi2sdxmm0,rcx ;; size=6 bbWeight=1 PerfScore 4.00G_M33997_IG03: ;; offset=0009H C3 ret
CasePrevious CodeOptimized Instruction
ulong -> floatulong->double->floatvcvttusi2ss
[MethodImplAttribute(MethodImplOptions.NoInlining)]publicstaticfloatConvUlongToFloat(ulongval){return(float)val;}

Assembly before optimization

G_M2883_IG01: ;; offset=0000Hpushrbpsubrsp,48vzeroupperlearbp,[rsp+30H]xoreax,eaxmov dword ptr [rbp-04H],eaxmov qword ptr [rbp+10H],rcx ;; size=22 bbWeight=1 PerfScore 5.00G_M2883_IG02: ;; offset=0016Hcmp dword ptr [(reloc 0x7ff8bc0d2898)],0je SHORT G_M2883_IG04 ;; size=9 bbWeight=1 PerfScore 4.00G_M2883_IG03: ;; offset=001FHcall CORINFO_HELP_DBG_IS_JUST_MY_CODE ;; size=5 bbWeight=0.50 PerfScore 0.50G_M2883_IG04: ;; offset=0024Hnopmovrax, qword ptr [rbp+10H]vcvtusi2sdxmm0,rax vcvtsd2ss xmm0,xmm0,xmm0vmovss dword ptr [rbp-04H],xmm0nop ;; size=25 bbWeight=1 PerfScore 10.50G_M2883_IG05: ;; offset=003DHvmovssxmm0, dword ptr [rbp-04H] ;; size=7 bbWeight=1 PerfScore 3.00G_M2883_IG06: ;; offset=0044Haddrsp,48poprbpret ;; size=6 bbWeight=1 PerfScore 1.75

Assembly after optimization

G_M2883_IG01: ;; offset=0000Hpushrbpsubrsp,48vzeroupperlearbp,[rsp+30H]xoreax,eaxmov dword ptr [rbp-04H],eaxmov qword ptr [rbp+10H],rcx ;; size=22 bbWeight=1 PerfScore 5.00G_M2883_IG02: ;; offset=0016Hcmp dword ptr [(reloc 0x7ff8b54b2898)],0je SHORT G_M2883_IG04 ;; size=9 bbWeight=1 PerfScore 4.00G_M2883_IG03: ;; offset=001FHcall CORINFO_HELP_DBG_IS_JUST_MY_CODE ;; size=5 bbWeight=0.50 PerfScore 0.50G_M2883_IG04: ;; offset=0024Hnopmovrax, qword ptr [rbp+10H]vcvtusi2ssxmm0,raxvmovss dword ptr [rbp-04H],xmm0nop ;; size=19 bbWeight=1 PerfScore 8.50G_M2883_IG05: ;; offset=0037Hvmovssxmm0, dword ptr [rbp-04H] ;; size=7 bbWeight=1 PerfScore 3.00G_M2883_IG06: ;; offset=003EHaddrsp,48poprbpret ;; size=6 bbWeight=1 PerfScore 1.75

@ghostghost added area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI community-contribution Indicates that the PR has been added by a community member labels Apr 5, 2023
@ghost

ghost commented Apr 5, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch, @kunalspathak
See info in area-owners.md if you want to be subscribed.

Issue Details

Draft PR for testing purposes. No need for review at this time.

This PR optimize the following cases:


CasePrevious CodeOptimized Instruction
float -> ulongCORINFO_HELP_DBL2ULNG Helpervcvttss2usi
publicstaticUInt64FloatToULong(floatval){return(UInt64)val;}

Assembly before optimization

G_M22196_IG01: ;; offset=0000H 4883EC28 subrsp,40 C5F877 vzeroupper ;; size=7 bbWeight=1 PerfScore 1.25G_M22196_IG02: ;; offset=0007H 62F17E085AC0 vcvtss2sd xmm0,xmm0 E87E57815E call CORINFO_HELP_DBL2ULNG90nop ;; size=12 bbWeight=1 PerfScore 5.25G_M22196_IG03: ;; offset=0013H 4883C428 addrsp,40 C3 ret ;; size=5 bbWeight=1 PerfScore 1.25

Assembly afteroptimization

G_M22196_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M22196_IG02: ;; offset=0003H 62F1FE0878C0 vcvttss2usirax,xmm0 ;; size=6 bbWeight=1 PerfScore 6.00G_M22196_IG03: ;; offset=0009H C3 ret

CasePrevious CodeOptimized Instruction
double -> ulongCORINFO_HELP_DBL2ULNG Helpervcvttsd2usi
publicstaticUInt64DoubleToULong(doubleval){return(UInt64)val;}

Assembly before optimization

G_M30068_IG01: ;; offset=0000H 4883EC28 subrsp,40 C5F877 vzeroupper ;; size=7 bbWeight=1 PerfScore 1.25G_M30068_IG02: ;; offset=0007H E874577F5E call CORINFO_HELP_DBL2ULNG90nop ;; size=6 bbWeight=1 PerfScore 1.25G_M30068_IG03: ;; offset=000DH 4883C428 addrsp,40 C3 ret

Assembly afteroptimization

G_M30068_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M30068_IG02: ;; offset=0003H 62F1FF0878C0 vcvttsd2usirax,xmm0 ;; size=6 bbWeight=1 PerfScore 5.00G_M30068_IG03: ;; offset=0009H C3 ret

CasePrevious CodeOptimized Instruction
ulong -> doublevcvtsi2sdvcvtusi2sd
publicstaticdoubleUIntToDouble(UInt64val){return(double)val;}

Assembly before optimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F17C0857C0 vxorps xmm0,xmm0 62F1FF082AC1 vcvtsi2sd xmm0,rcx 4885C9 testrcx,rcx 7D0A jge SHORT G_M33997_IG03 62F1FF08580502000000 vaddsd xmm0, qword ptr [reloc @RWD00] ;; size=27 bbWeight=1 PerfScore 12.58G_M33997_IG03: ;; offset=001EH C3 ret

Assembly afteroptimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F1FF087BC1 vcvtusi2sdxmm0,rcx ;; size=6 bbWeight=1 PerfScore 4.00G_M33997_IG03: ;; offset=0009H C3 ret
Author:khushal1996
Assignees:-
Labels:

area-CodeGen-coreclr

Milestone:-

@khushal1996khushal1996 changed the title Avx512 scalar convert rebasedOptimize scalar conversions with AVX512Apr 5, 2023
@BruceForstallBruceForstall added the avx512 Related to the AVX-512 architecture label Apr 7, 2023
@khushal1996

khushal1996 commented Apr 7, 2023

Copy link
Copy Markdown
MemberAuthor

I am making Submissions in the course of work for my employer (or my employer has intellectual property rights in my Submissions by contract or applicable law). I have permission from my employer to make Submissions and enter into this Agreement on behalf of my employer. By signing below, the defined term “You” includes me and my employer.

@dotnet-policy-service agree company="Intel"

Comment threadsrc/coreclr/vm/jithelpers.cpp Outdated
Comment threadsrc/tests/JIT/Directed/Convert/out_of_range_fp_to_int_conversions.cpp Outdated
Comment threadsrc/tests/JIT/Directed/Convert/out_of_range_fp_to_int_conversions.cpp Outdated
Comment threadsrc/libraries/System.Runtime/tests/System/UIntPtrTests.GenericMath.cs Outdated
Comment threadsrc/coreclr/jit/instrsxarch.h Outdated
Comment threadsrc/coreclr/vm/jithelpers.cpp Outdated
Comment threadsrc/libraries/System.Runtime/tests/System/UIntPtrTests.GenericMath.cs Outdated
Comment threadsrc/coreclr/vm/jithelpers.cpp Outdated
@khushal1996

Copy link
Copy Markdown
MemberAuthor

We are narrowing down the scope of the PR to ulong -> float conversions. The changes will b e updated soon.

@@ -18,8 +18,46 @@ FORCEINLINE int64_t FastDbl2Lng(double val)
#endif

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The changes in this file and in jithelpers.cpp should be rolled back too.

@khushal1996

Copy link
Copy Markdown
MemberAuthor

@tannergooding@jkotas does this look good for merge? I have rolled back the changes for float/double -> I long and optimized ulong -> float/double cases.

@jkotasjkotas left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks good to me. I have commented on a few nits.

Somebody on @dotnet/jit-contrib should do final review and merge.

Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/lowerxarch.cpp
Comment threadsrc/coreclr/jit/instr.cpp
Comment threadsrc/coreclr/jit/importer.cpp Outdated
@khushal1996

Copy link
Copy Markdown
MemberAuthor

@jkotas I think we are good to go here. Would you please help with the approval and merge.

@jkotas

Copy link
Copy Markdown
Member

@kunalspathak Could you please do final review and merge?

@khushal1996

Copy link
Copy Markdown
MemberAuthor

@kunalspathak can you help to merge this changes? They have been approved and we are trying to get them in before the next release.

Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated
@tannergooding

Copy link
Copy Markdown
Member

SPMI failures are the general No Azure Storage MCH files to download from cef79bc8-29bf-4f7b-9d05-9fc06832098c/osx/arm64/ impacting other PRs.

CI was passing before the minor formatting cleanup requested

@tannergooding
tannergooding merged commit b99a279 into dotnet:mainJul 16, 2023
@ghostghost locked as resolved and limited conversation to collaborators Aug 15, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIavx512Related to the AVX-512 architecturecommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@khushal1996@tannergooding@jkotas@EgorBo@BruceForstall@kunalspathak@SingleAccretion@anthonycanino
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Optimize scalar conversions with AVX512 - #84384

Merged
tannergooding merged 41 commits into
dotnet:mainfrom
khushal1996:avx512-scalar-convert-rebased
Jul 16, 2023
Merged

Optimize scalar conversions with AVX512#84384
tannergooding merged 41 commits into
dotnet:mainfrom
khushal1996:avx512-scalar-convert-rebased

Conversation

@khushal1996

@khushal1996khushal1996 commented Apr 5, 2023

Copy link
Copy Markdown
Member

This PR optimize the following cases:


CasePrevious CodeOptimized Instruction
ulong -> doublevcvtsi2sdvcvtusi2sd
publicstaticdoubleUIntToDouble(UInt64val){return(double)val;}

Assembly before optimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F17C0857C0 vxorps xmm0,xmm0 62F1FF082AC1 vcvtsi2sd xmm0,rcx 4885C9 testrcx,rcx 7D0A jge SHORT G_M33997_IG03 62F1FF08580502000000 vaddsd xmm0, qword ptr [reloc @RWD00] ;; size=27 bbWeight=1 PerfScore 12.58G_M33997_IG03: ;; offset=001EH C3 ret

Assembly after optimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F1FF087BC1 vcvtusi2sdxmm0,rcx ;; size=6 bbWeight=1 PerfScore 4.00G_M33997_IG03: ;; offset=0009H C3 ret
CasePrevious CodeOptimized Instruction
ulong -> floatulong->double->floatvcvttusi2ss
[MethodImplAttribute(MethodImplOptions.NoInlining)]publicstaticfloatConvUlongToFloat(ulongval){return(float)val;}

Assembly before optimization

G_M2883_IG01: ;; offset=0000Hpushrbpsubrsp,48vzeroupperlearbp,[rsp+30H]xoreax,eaxmov dword ptr [rbp-04H],eaxmov qword ptr [rbp+10H],rcx ;; size=22 bbWeight=1 PerfScore 5.00G_M2883_IG02: ;; offset=0016Hcmp dword ptr [(reloc 0x7ff8bc0d2898)],0je SHORT G_M2883_IG04 ;; size=9 bbWeight=1 PerfScore 4.00G_M2883_IG03: ;; offset=001FHcall CORINFO_HELP_DBG_IS_JUST_MY_CODE ;; size=5 bbWeight=0.50 PerfScore 0.50G_M2883_IG04: ;; offset=0024Hnopmovrax, qword ptr [rbp+10H]vcvtusi2sdxmm0,rax vcvtsd2ss xmm0,xmm0,xmm0vmovss dword ptr [rbp-04H],xmm0nop ;; size=25 bbWeight=1 PerfScore 10.50G_M2883_IG05: ;; offset=003DHvmovssxmm0, dword ptr [rbp-04H] ;; size=7 bbWeight=1 PerfScore 3.00G_M2883_IG06: ;; offset=0044Haddrsp,48poprbpret ;; size=6 bbWeight=1 PerfScore 1.75

Assembly after optimization

G_M2883_IG01: ;; offset=0000Hpushrbpsubrsp,48vzeroupperlearbp,[rsp+30H]xoreax,eaxmov dword ptr [rbp-04H],eaxmov qword ptr [rbp+10H],rcx ;; size=22 bbWeight=1 PerfScore 5.00G_M2883_IG02: ;; offset=0016Hcmp dword ptr [(reloc 0x7ff8b54b2898)],0je SHORT G_M2883_IG04 ;; size=9 bbWeight=1 PerfScore 4.00G_M2883_IG03: ;; offset=001FHcall CORINFO_HELP_DBG_IS_JUST_MY_CODE ;; size=5 bbWeight=0.50 PerfScore 0.50G_M2883_IG04: ;; offset=0024Hnopmovrax, qword ptr [rbp+10H]vcvtusi2ssxmm0,raxvmovss dword ptr [rbp-04H],xmm0nop ;; size=19 bbWeight=1 PerfScore 8.50G_M2883_IG05: ;; offset=0037Hvmovssxmm0, dword ptr [rbp-04H] ;; size=7 bbWeight=1 PerfScore 3.00G_M2883_IG06: ;; offset=003EHaddrsp,48poprbpret ;; size=6 bbWeight=1 PerfScore 1.75

@ghostghost added area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI community-contribution Indicates that the PR has been added by a community member labels Apr 5, 2023
@ghost

ghost commented Apr 5, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch, @kunalspathak
See info in area-owners.md if you want to be subscribed.

Issue Details

Draft PR for testing purposes. No need for review at this time.

This PR optimize the following cases:


CasePrevious CodeOptimized Instruction
float -> ulongCORINFO_HELP_DBL2ULNG Helpervcvttss2usi
publicstaticUInt64FloatToULong(floatval){return(UInt64)val;}

Assembly before optimization

G_M22196_IG01: ;; offset=0000H 4883EC28 subrsp,40 C5F877 vzeroupper ;; size=7 bbWeight=1 PerfScore 1.25G_M22196_IG02: ;; offset=0007H 62F17E085AC0 vcvtss2sd xmm0,xmm0 E87E57815E call CORINFO_HELP_DBL2ULNG90nop ;; size=12 bbWeight=1 PerfScore 5.25G_M22196_IG03: ;; offset=0013H 4883C428 addrsp,40 C3 ret ;; size=5 bbWeight=1 PerfScore 1.25

Assembly afteroptimization

G_M22196_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M22196_IG02: ;; offset=0003H 62F1FE0878C0 vcvttss2usirax,xmm0 ;; size=6 bbWeight=1 PerfScore 6.00G_M22196_IG03: ;; offset=0009H C3 ret

CasePrevious CodeOptimized Instruction
double -> ulongCORINFO_HELP_DBL2ULNG Helpervcvttsd2usi
publicstaticUInt64DoubleToULong(doubleval){return(UInt64)val;}

Assembly before optimization

G_M30068_IG01: ;; offset=0000H 4883EC28 subrsp,40 C5F877 vzeroupper ;; size=7 bbWeight=1 PerfScore 1.25G_M30068_IG02: ;; offset=0007H E874577F5E call CORINFO_HELP_DBL2ULNG90nop ;; size=6 bbWeight=1 PerfScore 1.25G_M30068_IG03: ;; offset=000DH 4883C428 addrsp,40 C3 ret

Assembly afteroptimization

G_M30068_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M30068_IG02: ;; offset=0003H 62F1FF0878C0 vcvttsd2usirax,xmm0 ;; size=6 bbWeight=1 PerfScore 5.00G_M30068_IG03: ;; offset=0009H C3 ret

CasePrevious CodeOptimized Instruction
ulong -> doublevcvtsi2sdvcvtusi2sd
publicstaticdoubleUIntToDouble(UInt64val){return(double)val;}

Assembly before optimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F17C0857C0 vxorps xmm0,xmm0 62F1FF082AC1 vcvtsi2sd xmm0,rcx 4885C9 testrcx,rcx 7D0A jge SHORT G_M33997_IG03 62F1FF08580502000000 vaddsd xmm0, qword ptr [reloc @RWD00] ;; size=27 bbWeight=1 PerfScore 12.58G_M33997_IG03: ;; offset=001EH C3 ret

Assembly afteroptimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F1FF087BC1 vcvtusi2sdxmm0,rcx ;; size=6 bbWeight=1 PerfScore 4.00G_M33997_IG03: ;; offset=0009H C3 ret
Author:khushal1996
Assignees:-
Labels:

area-CodeGen-coreclr

Milestone:-

@khushal1996khushal1996 changed the title Avx512 scalar convert rebasedOptimize scalar conversions with AVX512Apr 5, 2023
@BruceForstallBruceForstall added the avx512 Related to the AVX-512 architecture label Apr 7, 2023
@khushal1996

khushal1996 commented Apr 7, 2023

Copy link
Copy Markdown
MemberAuthor

I am making Submissions in the course of work for my employer (or my employer has intellectual property rights in my Submissions by contract or applicable law). I have permission from my employer to make Submissions and enter into this Agreement on behalf of my employer. By signing below, the defined term “You” includes me and my employer.

@dotnet-policy-service agree company="Intel"

Comment threadsrc/coreclr/vm/jithelpers.cpp Outdated
Comment threadsrc/tests/JIT/Directed/Convert/out_of_range_fp_to_int_conversions.cpp Outdated
Comment threadsrc/tests/JIT/Directed/Convert/out_of_range_fp_to_int_conversions.cpp Outdated
Comment threadsrc/libraries/System.Runtime/tests/System/UIntPtrTests.GenericMath.cs Outdated
Comment threadsrc/coreclr/jit/instrsxarch.h Outdated
Comment threadsrc/coreclr/vm/jithelpers.cpp Outdated
Comment threadsrc/libraries/System.Runtime/tests/System/UIntPtrTests.GenericMath.cs Outdated
Comment threadsrc/coreclr/vm/jithelpers.cpp Outdated
@khushal1996

Copy link
Copy Markdown
MemberAuthor

We are narrowing down the scope of the PR to ulong -> float conversions. The changes will b e updated soon.

@@ -18,8 +18,46 @@ FORCEINLINE int64_t FastDbl2Lng(double val)
#endif

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The changes in this file and in jithelpers.cpp should be rolled back too.

@khushal1996

Copy link
Copy Markdown
MemberAuthor

@tannergooding@jkotas does this look good for merge? I have rolled back the changes for float/double -> I long and optimized ulong -> float/double cases.

@jkotasjkotas left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks good to me. I have commented on a few nits.

Somebody on @dotnet/jit-contrib should do final review and merge.

Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/lowerxarch.cpp
Comment threadsrc/coreclr/jit/instr.cpp
Comment threadsrc/coreclr/jit/importer.cpp Outdated
@khushal1996

Copy link
Copy Markdown
MemberAuthor

@jkotas I think we are good to go here. Would you please help with the approval and merge.

@jkotas

Copy link
Copy Markdown
Member

@kunalspathak Could you please do final review and merge?

@khushal1996

Copy link
Copy Markdown
MemberAuthor

@kunalspathak can you help to merge this changes? They have been approved and we are trying to get them in before the next release.

Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated
@tannergooding

Copy link
Copy Markdown
Member

SPMI failures are the general No Azure Storage MCH files to download from cef79bc8-29bf-4f7b-9d05-9fc06832098c/osx/arm64/ impacting other PRs.

CI was passing before the minor formatting cleanup requested

@tannergooding
tannergooding merged commit b99a279 into dotnet:mainJul 16, 2023
@ghostghost locked as resolved and limited conversation to collaborators Aug 15, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIavx512Related to the AVX-512 architecturecommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@khushal1996@tannergooding@jkotas@EgorBo@BruceForstall@kunalspathak@SingleAccretion@anthonycanino
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Optimize scalar conversions with AVX512 - #84384

Merged
tannergooding merged 41 commits into
dotnet:mainfrom
khushal1996:avx512-scalar-convert-rebased
Jul 16, 2023
Merged

Optimize scalar conversions with AVX512#84384
tannergooding merged 41 commits into
dotnet:mainfrom
khushal1996:avx512-scalar-convert-rebased

Conversation

@khushal1996

@khushal1996khushal1996 commented Apr 5, 2023

Copy link
Copy Markdown
Member

This PR optimize the following cases:


CasePrevious CodeOptimized Instruction
ulong -> doublevcvtsi2sdvcvtusi2sd
publicstaticdoubleUIntToDouble(UInt64val){return(double)val;}

Assembly before optimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F17C0857C0 vxorps xmm0,xmm0 62F1FF082AC1 vcvtsi2sd xmm0,rcx 4885C9 testrcx,rcx 7D0A jge SHORT G_M33997_IG03 62F1FF08580502000000 vaddsd xmm0, qword ptr [reloc @RWD00] ;; size=27 bbWeight=1 PerfScore 12.58G_M33997_IG03: ;; offset=001EH C3 ret

Assembly after optimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F1FF087BC1 vcvtusi2sdxmm0,rcx ;; size=6 bbWeight=1 PerfScore 4.00G_M33997_IG03: ;; offset=0009H C3 ret
CasePrevious CodeOptimized Instruction
ulong -> floatulong->double->floatvcvttusi2ss
[MethodImplAttribute(MethodImplOptions.NoInlining)]publicstaticfloatConvUlongToFloat(ulongval){return(float)val;}

Assembly before optimization

G_M2883_IG01: ;; offset=0000Hpushrbpsubrsp,48vzeroupperlearbp,[rsp+30H]xoreax,eaxmov dword ptr [rbp-04H],eaxmov qword ptr [rbp+10H],rcx ;; size=22 bbWeight=1 PerfScore 5.00G_M2883_IG02: ;; offset=0016Hcmp dword ptr [(reloc 0x7ff8bc0d2898)],0je SHORT G_M2883_IG04 ;; size=9 bbWeight=1 PerfScore 4.00G_M2883_IG03: ;; offset=001FHcall CORINFO_HELP_DBG_IS_JUST_MY_CODE ;; size=5 bbWeight=0.50 PerfScore 0.50G_M2883_IG04: ;; offset=0024Hnopmovrax, qword ptr [rbp+10H]vcvtusi2sdxmm0,rax vcvtsd2ss xmm0,xmm0,xmm0vmovss dword ptr [rbp-04H],xmm0nop ;; size=25 bbWeight=1 PerfScore 10.50G_M2883_IG05: ;; offset=003DHvmovssxmm0, dword ptr [rbp-04H] ;; size=7 bbWeight=1 PerfScore 3.00G_M2883_IG06: ;; offset=0044Haddrsp,48poprbpret ;; size=6 bbWeight=1 PerfScore 1.75

Assembly after optimization

G_M2883_IG01: ;; offset=0000Hpushrbpsubrsp,48vzeroupperlearbp,[rsp+30H]xoreax,eaxmov dword ptr [rbp-04H],eaxmov qword ptr [rbp+10H],rcx ;; size=22 bbWeight=1 PerfScore 5.00G_M2883_IG02: ;; offset=0016Hcmp dword ptr [(reloc 0x7ff8b54b2898)],0je SHORT G_M2883_IG04 ;; size=9 bbWeight=1 PerfScore 4.00G_M2883_IG03: ;; offset=001FHcall CORINFO_HELP_DBG_IS_JUST_MY_CODE ;; size=5 bbWeight=0.50 PerfScore 0.50G_M2883_IG04: ;; offset=0024Hnopmovrax, qword ptr [rbp+10H]vcvtusi2ssxmm0,raxvmovss dword ptr [rbp-04H],xmm0nop ;; size=19 bbWeight=1 PerfScore 8.50G_M2883_IG05: ;; offset=0037Hvmovssxmm0, dword ptr [rbp-04H] ;; size=7 bbWeight=1 PerfScore 3.00G_M2883_IG06: ;; offset=003EHaddrsp,48poprbpret ;; size=6 bbWeight=1 PerfScore 1.75

@ghostghost added area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI community-contribution Indicates that the PR has been added by a community member labels Apr 5, 2023
@ghost

ghost commented Apr 5, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch, @kunalspathak
See info in area-owners.md if you want to be subscribed.

Issue Details

Draft PR for testing purposes. No need for review at this time.

This PR optimize the following cases:


CasePrevious CodeOptimized Instruction
float -> ulongCORINFO_HELP_DBL2ULNG Helpervcvttss2usi
publicstaticUInt64FloatToULong(floatval){return(UInt64)val;}

Assembly before optimization

G_M22196_IG01: ;; offset=0000H 4883EC28 subrsp,40 C5F877 vzeroupper ;; size=7 bbWeight=1 PerfScore 1.25G_M22196_IG02: ;; offset=0007H 62F17E085AC0 vcvtss2sd xmm0,xmm0 E87E57815E call CORINFO_HELP_DBL2ULNG90nop ;; size=12 bbWeight=1 PerfScore 5.25G_M22196_IG03: ;; offset=0013H 4883C428 addrsp,40 C3 ret ;; size=5 bbWeight=1 PerfScore 1.25

Assembly afteroptimization

G_M22196_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M22196_IG02: ;; offset=0003H 62F1FE0878C0 vcvttss2usirax,xmm0 ;; size=6 bbWeight=1 PerfScore 6.00G_M22196_IG03: ;; offset=0009H C3 ret

CasePrevious CodeOptimized Instruction
double -> ulongCORINFO_HELP_DBL2ULNG Helpervcvttsd2usi
publicstaticUInt64DoubleToULong(doubleval){return(UInt64)val;}

Assembly before optimization

G_M30068_IG01: ;; offset=0000H 4883EC28 subrsp,40 C5F877 vzeroupper ;; size=7 bbWeight=1 PerfScore 1.25G_M30068_IG02: ;; offset=0007H E874577F5E call CORINFO_HELP_DBL2ULNG90nop ;; size=6 bbWeight=1 PerfScore 1.25G_M30068_IG03: ;; offset=000DH 4883C428 addrsp,40 C3 ret

Assembly afteroptimization

G_M30068_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M30068_IG02: ;; offset=0003H 62F1FF0878C0 vcvttsd2usirax,xmm0 ;; size=6 bbWeight=1 PerfScore 5.00G_M30068_IG03: ;; offset=0009H C3 ret

CasePrevious CodeOptimized Instruction
ulong -> doublevcvtsi2sdvcvtusi2sd
publicstaticdoubleUIntToDouble(UInt64val){return(double)val;}

Assembly before optimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F17C0857C0 vxorps xmm0,xmm0 62F1FF082AC1 vcvtsi2sd xmm0,rcx 4885C9 testrcx,rcx 7D0A jge SHORT G_M33997_IG03 62F1FF08580502000000 vaddsd xmm0, qword ptr [reloc @RWD00] ;; size=27 bbWeight=1 PerfScore 12.58G_M33997_IG03: ;; offset=001EH C3 ret

Assembly afteroptimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F1FF087BC1 vcvtusi2sdxmm0,rcx ;; size=6 bbWeight=1 PerfScore 4.00G_M33997_IG03: ;; offset=0009H C3 ret
Author:khushal1996
Assignees:-
Labels:

area-CodeGen-coreclr

Milestone:-

@khushal1996khushal1996 changed the title Avx512 scalar convert rebasedOptimize scalar conversions with AVX512Apr 5, 2023
@BruceForstallBruceForstall added the avx512 Related to the AVX-512 architecture label Apr 7, 2023
@khushal1996

khushal1996 commented Apr 7, 2023

Copy link
Copy Markdown
MemberAuthor

I am making Submissions in the course of work for my employer (or my employer has intellectual property rights in my Submissions by contract or applicable law). I have permission from my employer to make Submissions and enter into this Agreement on behalf of my employer. By signing below, the defined term “You” includes me and my employer.

@dotnet-policy-service agree company="Intel"

Comment threadsrc/coreclr/vm/jithelpers.cpp Outdated
Comment threadsrc/tests/JIT/Directed/Convert/out_of_range_fp_to_int_conversions.cpp Outdated
Comment threadsrc/tests/JIT/Directed/Convert/out_of_range_fp_to_int_conversions.cpp Outdated
Comment threadsrc/libraries/System.Runtime/tests/System/UIntPtrTests.GenericMath.cs Outdated
Comment threadsrc/coreclr/jit/instrsxarch.h Outdated
Comment threadsrc/coreclr/vm/jithelpers.cpp Outdated
Comment threadsrc/libraries/System.Runtime/tests/System/UIntPtrTests.GenericMath.cs Outdated
Comment threadsrc/coreclr/vm/jithelpers.cpp Outdated
@khushal1996

Copy link
Copy Markdown
MemberAuthor

We are narrowing down the scope of the PR to ulong -> float conversions. The changes will b e updated soon.

@@ -18,8 +18,46 @@ FORCEINLINE int64_t FastDbl2Lng(double val)
#endif

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The changes in this file and in jithelpers.cpp should be rolled back too.

@khushal1996

Copy link
Copy Markdown
MemberAuthor

@tannergooding@jkotas does this look good for merge? I have rolled back the changes for float/double -> I long and optimized ulong -> float/double cases.

@jkotasjkotas left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks good to me. I have commented on a few nits.

Somebody on @dotnet/jit-contrib should do final review and merge.

Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/lowerxarch.cpp
Comment threadsrc/coreclr/jit/instr.cpp
Comment threadsrc/coreclr/jit/importer.cpp Outdated
@khushal1996

Copy link
Copy Markdown
MemberAuthor

@jkotas I think we are good to go here. Would you please help with the approval and merge.

@jkotas

Copy link
Copy Markdown
Member

@kunalspathak Could you please do final review and merge?

@khushal1996

Copy link
Copy Markdown
MemberAuthor

@kunalspathak can you help to merge this changes? They have been approved and we are trying to get them in before the next release.

Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated
@tannergooding

Copy link
Copy Markdown
Member

SPMI failures are the general No Azure Storage MCH files to download from cef79bc8-29bf-4f7b-9d05-9fc06832098c/osx/arm64/ impacting other PRs.

CI was passing before the minor formatting cleanup requested

@tannergooding
tannergooding merged commit b99a279 into dotnet:mainJul 16, 2023
@ghostghost locked as resolved and limited conversation to collaborators Aug 15, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIavx512Related to the AVX-512 architecturecommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@khushal1996@tannergooding@jkotas@EgorBo@BruceForstall@kunalspathak@SingleAccretion@anthonycanino
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Optimize scalar conversions with AVX512 - #84384

Merged
tannergooding merged 41 commits into
dotnet:mainfrom
khushal1996:avx512-scalar-convert-rebased
Jul 16, 2023
Merged

Optimize scalar conversions with AVX512#84384
tannergooding merged 41 commits into
dotnet:mainfrom
khushal1996:avx512-scalar-convert-rebased

Conversation

@khushal1996

@khushal1996khushal1996 commented Apr 5, 2023

Copy link
Copy Markdown
Member

This PR optimize the following cases:


CasePrevious CodeOptimized Instruction
ulong -> doublevcvtsi2sdvcvtusi2sd
publicstaticdoubleUIntToDouble(UInt64val){return(double)val;}

Assembly before optimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F17C0857C0 vxorps xmm0,xmm0 62F1FF082AC1 vcvtsi2sd xmm0,rcx 4885C9 testrcx,rcx 7D0A jge SHORT G_M33997_IG03 62F1FF08580502000000 vaddsd xmm0, qword ptr [reloc @RWD00] ;; size=27 bbWeight=1 PerfScore 12.58G_M33997_IG03: ;; offset=001EH C3 ret

Assembly after optimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F1FF087BC1 vcvtusi2sdxmm0,rcx ;; size=6 bbWeight=1 PerfScore 4.00G_M33997_IG03: ;; offset=0009H C3 ret
CasePrevious CodeOptimized Instruction
ulong -> floatulong->double->floatvcvttusi2ss
[MethodImplAttribute(MethodImplOptions.NoInlining)]publicstaticfloatConvUlongToFloat(ulongval){return(float)val;}

Assembly before optimization

G_M2883_IG01: ;; offset=0000Hpushrbpsubrsp,48vzeroupperlearbp,[rsp+30H]xoreax,eaxmov dword ptr [rbp-04H],eaxmov qword ptr [rbp+10H],rcx ;; size=22 bbWeight=1 PerfScore 5.00G_M2883_IG02: ;; offset=0016Hcmp dword ptr [(reloc 0x7ff8bc0d2898)],0je SHORT G_M2883_IG04 ;; size=9 bbWeight=1 PerfScore 4.00G_M2883_IG03: ;; offset=001FHcall CORINFO_HELP_DBG_IS_JUST_MY_CODE ;; size=5 bbWeight=0.50 PerfScore 0.50G_M2883_IG04: ;; offset=0024Hnopmovrax, qword ptr [rbp+10H]vcvtusi2sdxmm0,rax vcvtsd2ss xmm0,xmm0,xmm0vmovss dword ptr [rbp-04H],xmm0nop ;; size=25 bbWeight=1 PerfScore 10.50G_M2883_IG05: ;; offset=003DHvmovssxmm0, dword ptr [rbp-04H] ;; size=7 bbWeight=1 PerfScore 3.00G_M2883_IG06: ;; offset=0044Haddrsp,48poprbpret ;; size=6 bbWeight=1 PerfScore 1.75

Assembly after optimization

G_M2883_IG01: ;; offset=0000Hpushrbpsubrsp,48vzeroupperlearbp,[rsp+30H]xoreax,eaxmov dword ptr [rbp-04H],eaxmov qword ptr [rbp+10H],rcx ;; size=22 bbWeight=1 PerfScore 5.00G_M2883_IG02: ;; offset=0016Hcmp dword ptr [(reloc 0x7ff8b54b2898)],0je SHORT G_M2883_IG04 ;; size=9 bbWeight=1 PerfScore 4.00G_M2883_IG03: ;; offset=001FHcall CORINFO_HELP_DBG_IS_JUST_MY_CODE ;; size=5 bbWeight=0.50 PerfScore 0.50G_M2883_IG04: ;; offset=0024Hnopmovrax, qword ptr [rbp+10H]vcvtusi2ssxmm0,raxvmovss dword ptr [rbp-04H],xmm0nop ;; size=19 bbWeight=1 PerfScore 8.50G_M2883_IG05: ;; offset=0037Hvmovssxmm0, dword ptr [rbp-04H] ;; size=7 bbWeight=1 PerfScore 3.00G_M2883_IG06: ;; offset=003EHaddrsp,48poprbpret ;; size=6 bbWeight=1 PerfScore 1.75

@ghostghost added area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI community-contribution Indicates that the PR has been added by a community member labels Apr 5, 2023
@ghost

ghost commented Apr 5, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch, @kunalspathak
See info in area-owners.md if you want to be subscribed.

Issue Details

Draft PR for testing purposes. No need for review at this time.

This PR optimize the following cases:


CasePrevious CodeOptimized Instruction
float -> ulongCORINFO_HELP_DBL2ULNG Helpervcvttss2usi
publicstaticUInt64FloatToULong(floatval){return(UInt64)val;}

Assembly before optimization

G_M22196_IG01: ;; offset=0000H 4883EC28 subrsp,40 C5F877 vzeroupper ;; size=7 bbWeight=1 PerfScore 1.25G_M22196_IG02: ;; offset=0007H 62F17E085AC0 vcvtss2sd xmm0,xmm0 E87E57815E call CORINFO_HELP_DBL2ULNG90nop ;; size=12 bbWeight=1 PerfScore 5.25G_M22196_IG03: ;; offset=0013H 4883C428 addrsp,40 C3 ret ;; size=5 bbWeight=1 PerfScore 1.25

Assembly afteroptimization

G_M22196_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M22196_IG02: ;; offset=0003H 62F1FE0878C0 vcvttss2usirax,xmm0 ;; size=6 bbWeight=1 PerfScore 6.00G_M22196_IG03: ;; offset=0009H C3 ret

CasePrevious CodeOptimized Instruction
double -> ulongCORINFO_HELP_DBL2ULNG Helpervcvttsd2usi
publicstaticUInt64DoubleToULong(doubleval){return(UInt64)val;}

Assembly before optimization

G_M30068_IG01: ;; offset=0000H 4883EC28 subrsp,40 C5F877 vzeroupper ;; size=7 bbWeight=1 PerfScore 1.25G_M30068_IG02: ;; offset=0007H E874577F5E call CORINFO_HELP_DBL2ULNG90nop ;; size=6 bbWeight=1 PerfScore 1.25G_M30068_IG03: ;; offset=000DH 4883C428 addrsp,40 C3 ret

Assembly afteroptimization

G_M30068_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M30068_IG02: ;; offset=0003H 62F1FF0878C0 vcvttsd2usirax,xmm0 ;; size=6 bbWeight=1 PerfScore 5.00G_M30068_IG03: ;; offset=0009H C3 ret

CasePrevious CodeOptimized Instruction
ulong -> doublevcvtsi2sdvcvtusi2sd
publicstaticdoubleUIntToDouble(UInt64val){return(double)val;}

Assembly before optimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F17C0857C0 vxorps xmm0,xmm0 62F1FF082AC1 vcvtsi2sd xmm0,rcx 4885C9 testrcx,rcx 7D0A jge SHORT G_M33997_IG03 62F1FF08580502000000 vaddsd xmm0, qword ptr [reloc @RWD00] ;; size=27 bbWeight=1 PerfScore 12.58G_M33997_IG03: ;; offset=001EH C3 ret

Assembly afteroptimization

G_M33997_IG01: ;; offset=0000H C5F877 vzeroupper ;; size=3 bbWeight=1 PerfScore 1.00G_M33997_IG02: ;; offset=0003H 62F1FF087BC1 vcvtusi2sdxmm0,rcx ;; size=6 bbWeight=1 PerfScore 4.00G_M33997_IG03: ;; offset=0009H C3 ret
Author:khushal1996
Assignees:-
Labels:

area-CodeGen-coreclr

Milestone:-

@khushal1996khushal1996 changed the title Avx512 scalar convert rebasedOptimize scalar conversions with AVX512Apr 5, 2023
@BruceForstallBruceForstall added the avx512 Related to the AVX-512 architecture label Apr 7, 2023
@khushal1996

khushal1996 commented Apr 7, 2023

Copy link
Copy Markdown
MemberAuthor

I am making Submissions in the course of work for my employer (or my employer has intellectual property rights in my Submissions by contract or applicable law). I have permission from my employer to make Submissions and enter into this Agreement on behalf of my employer. By signing below, the defined term “You” includes me and my employer.

@dotnet-policy-service agree company="Intel"

Comment threadsrc/coreclr/vm/jithelpers.cpp Outdated
Comment threadsrc/tests/JIT/Directed/Convert/out_of_range_fp_to_int_conversions.cpp Outdated
Comment threadsrc/tests/JIT/Directed/Convert/out_of_range_fp_to_int_conversions.cpp Outdated
Comment threadsrc/libraries/System.Runtime/tests/System/UIntPtrTests.GenericMath.cs Outdated
Comment threadsrc/coreclr/jit/instrsxarch.h Outdated
Comment threadsrc/coreclr/vm/jithelpers.cpp Outdated
Comment threadsrc/libraries/System.Runtime/tests/System/UIntPtrTests.GenericMath.cs Outdated
Comment threadsrc/coreclr/vm/jithelpers.cpp Outdated
@khushal1996

Copy link
Copy Markdown
MemberAuthor

We are narrowing down the scope of the PR to ulong -> float conversions. The changes will b e updated soon.

@@ -18,8 +18,46 @@ FORCEINLINE int64_t FastDbl2Lng(double val)
#endif

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The changes in this file and in jithelpers.cpp should be rolled back too.

@khushal1996

Copy link
Copy Markdown
MemberAuthor

@tannergooding@jkotas does this look good for merge? I have rolled back the changes for float/double -> I long and optimized ulong -> float/double cases.

@jkotasjkotas left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks good to me. I have commented on a few nits.

Somebody on @dotnet/jit-contrib should do final review and merge.

Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/lowerxarch.cpp
Comment threadsrc/coreclr/jit/instr.cpp
Comment threadsrc/coreclr/jit/importer.cpp Outdated
@khushal1996

Copy link
Copy Markdown
MemberAuthor

@jkotas I think we are good to go here. Would you please help with the approval and merge.

@jkotas

Copy link
Copy Markdown
Member

@kunalspathak Could you please do final review and merge?

@khushal1996

Copy link
Copy Markdown
MemberAuthor

@kunalspathak can you help to merge this changes? They have been approved and we are trying to get them in before the next release.

Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated
@tannergooding

Copy link
Copy Markdown
Member

SPMI failures are the general No Azure Storage MCH files to download from cef79bc8-29bf-4f7b-9d05-9fc06832098c/osx/arm64/ impacting other PRs.

CI was passing before the minor formatting cleanup requested

@tannergooding
tannergooding merged commit b99a279 into dotnet:mainJul 16, 2023
@ghostghost locked as resolved and limited conversation to collaborators Aug 15, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIavx512Related to the AVX-512 architecturecommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@khushal1996@tannergooding@jkotas@EgorBo@BruceForstall@kunalspathak@SingleAccretion@anthonycanino