[Arm64] Implement ASIMD widening, narrowing, saturating intrinsics - #35612

Merged
echesakov merged 82 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics
May 1, 2020
Merged

[Arm64] Implement ASIMD widening, narrowing, saturating intrinsics#35612
echesakov merged 82 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics

Conversation

@echesakov

@echesakovechesakov commented Apr 29, 2020

Copy link
Copy Markdown
Contributor

Implement API as approved in #32512 (with corrections regarding their ISA class - see my comment here)

Note that I am using the following temporary names as per comment:

  • AddReturningHighNarrowUpper and AddReturningHighNarrowLower
  • AddReturningRoundedHighNarrowUpper and AddReturningRoundedHighNarrowLower
  • SubtractReturningHighNarrowUpper and SubtractReturningHighNarrowLower.

I will update the methods names after the next round of API review.

The change is quite straightforward - the special handling is required only for AddWideningLower, AddWideningUpper, SubractWideningLower and SubtractWideningUpper since each of those are mapped to four different instructions (saddl{2}, saddw{2}, uaddl{2} and uaddw{2}; ssubl{2}, ssubw{2}, usubl{2} and usubw{2}) and do not fit into our table-driven model.

The change depends on removing HW_Flag_UnfixedSIMDSize in #35594 - I didn't want to keep adding the flag for the added intrinsics for no particular reason.

Fixes#32512

@echesakovechesakov added arch-arm64 area-System.Runtime.Intrinsics area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI labels Apr 29, 2020
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @tannergooding
Notify danmosemsft if you want to be subscribed.

@Dotnet-GitSync-Bot

Copy link
Copy Markdown
Collaborator

Note regarding the new-api-needs-documentation label:

This serves as a reminder for when your PR is modifying a ref *.cs file and adding/modifying public APIs, to please make sure the API implementation in the src *.cs file is documented with triple slash comments, so the PR reviewers can sign off that change.

@echesakov

Copy link
Copy Markdown
ContributorAuthor

@CarolEidt@tannergooding PTAL

@echesakov

Copy link
Copy Markdown
ContributorAuthor

@TamarChristinaArm PTAL

@echesakov
echesakovforce-pushed the Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics branch from f5004a2 to 97e7d77CompareApril 30, 2020 02:46
@echesakov

Copy link
Copy Markdown
ContributorAuthor

Rebased on top of 1f436e0

Comment threadsrc/coreclr/src/jit/hwintrinsic.cpp Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why do widening and narrowing need an "other" base type? The return type is always twice the size and of the same signedness of the inputs for these operations.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why do widening and narrowing need an "other" base type?

Narrowing intrinsics do not need - only AddWideningUpper and SubractWideningUpper do.

They don't need the operand 1 type itself (which is always going to be TYP_SIMD16) but the operand 1 element type so I can distinguish between each pair of these three

Vector128<int>AddWideningUpper(Vector128<short>left,Vector128<short>right);Vector128<short>AddWideningUpper(Vector128<short>left,Vector128<sbyte>right);Vector128<int>AddWideningUpper(Vector128<int>left,Vector128<short>right);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why can't we just use the second argument for the base type? That is what HW_Flag_BaseTypeFromSecondArg is for.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh nevermind, I missed that int = short + short and int = int + short need to be differentiated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Interesting - I renamed the gtIndexBaseType to gtOtherBaseType when I restructured the IR, because I figured there would be cases aside from gather where we'd need an additional base type - I didn't think it would be used that quickly :-)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should be able to get rid of this logic once we change the tables to take simdSize into account during lookup, correct?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can if we make simdSize and baseType of an intrinsic "independent" on each other - i.e. if we in addition to HW_Flag_BaseTypeFromFirstArg, HW_Flag_BaseTypeFromSecondArg we have HW_Flag_SimdSizeFromFirstArg, HW_Flag_SimdSizeFromSecondArg which would allow, for example, to set simdSize based on op1 and baseType from op2.

Then, the corresponding rows in hwintrinsiclistarm64.h could be updated as

HARDWARE_INTRINSIC(AdvSimd, AddWideningLower, 8, 2, {INS_saddl, INS_uaddl, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)
HARDWARE_INTRINSIC(AdvSimd, AddWideningLower, 16, 2, {INS_saddw, INS_uaddw, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)
HARDWARE_INTRINSIC(AdvSimd, AddWideningUpper, 8, 2, {INS_saddl2, INS_uaddl2, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)
HARDWARE_INTRINSIC(AdvSimd, AddWideningUpper, 16, 2, {INS_saddw2, INS_uaddw2, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)

In other words, if we allow x and y coordinates (where x is baseType and y is a tuple of (intrinsicId, simdSize)) of hwintrinsiclistarm64.h table to be orthogonal then the answer to your question is yes.

…nd SubtractWideningUpper in hwintrinsiccodegenarm64.cpp
…per and SubtractWideningUpper in hwintrinsic.cpp
@echesakov
echesakovforce-pushed the Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics branch from 97e7d77 to 417c253CompareApril 30, 2020 17:51
@echesakov

Copy link
Copy Markdown
ContributorAuthor

Fixed merge conflicts and rebased on top of aa81328

@CarolEidtCarolEidt left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM with some non-blocking comments & questions.
I only lightly reviewed the test helper functions and related changes.

Comment threadsrc/coreclr/src/jit/hwintrinsic.cpp Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Interesting - I renamed the gtIndexBaseType to gtOtherBaseType when I restructured the IR, because I figured there would be cases aside from gather where we'd need an additional base type - I didn't think it would be used that quickly :-)

insOpts opt = INS_OPTS_NONE;

if ((intrin.category == HW_Category_SIMDScalar) || (intrin.category == HW_Category_Scalar))
if (intrin.category == HW_Category_SIMDScalar)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The changes here highlight the fact that there are many subtleties with regard to emit sizes - the "actual type" of the baseType, the "raw" (emitTypeSize) of the baseType, and the emitSize of the node itself. I think it would be worth some additional comments both here and in the cases where we use emitTypeSize(node) for moves.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agree, I will follow up and add the comments

case NI_AdvSimd_AddWideningUpper:
case NI_AdvSimd_SubtractWideningLower:
case NI_AdvSimd_SubtractWideningUpper:
GetEmitter()->emitIns_R_R_R(ins, emitSize, targetReg, op1Reg, op2Reg, opt);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I believe this is identical to case NI_Crc32_ComputeCrc32: et al. Is there a reason you didn't combine it? Is the idea to keep them in order, and rely on the C++ compiler to de-duplicate the code? (I see there's some duplication already).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No, I just missed this. Thanks for spotting this.

I will follow up and de-duplicate this in a separate PR and I also need to address one of your another suggestions here

@echesakov

Copy link
Copy Markdown
ContributorAuthor

Both runtime (Mono Product Build Android x86 debug) and runtime (Mono Product Build OSX x64 debug) has succeeded on the second attempt - https://dev.azure.com/dnceng/public/_build/results?buildId=625278 - merging

@echesakov
echesakov merged commit a156293 into dotnet:masterMay 1, 2020
@echesakov
echesakov deleted the Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics branch May 1, 2020 02:47
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

arch-arm64area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIarea-System.Runtime.Intrinsicsnew-api-needs-documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

ARM additional arithmetic intrinsics

4 participants

@echesakov@Dotnet-GitSync-Bot@CarolEidt@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

[Arm64] Implement ASIMD widening, narrowing, saturating intrinsics - #35612

Merged
echesakov merged 82 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics
May 1, 2020
Merged

[Arm64] Implement ASIMD widening, narrowing, saturating intrinsics#35612
echesakov merged 82 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics

Conversation

@echesakov

@echesakovechesakov commented Apr 29, 2020

Copy link
Copy Markdown
Contributor

Implement API as approved in #32512 (with corrections regarding their ISA class - see my comment here)

Note that I am using the following temporary names as per comment:

  • AddReturningHighNarrowUpper and AddReturningHighNarrowLower
  • AddReturningRoundedHighNarrowUpper and AddReturningRoundedHighNarrowLower
  • SubtractReturningHighNarrowUpper and SubtractReturningHighNarrowLower.

I will update the methods names after the next round of API review.

The change is quite straightforward - the special handling is required only for AddWideningLower, AddWideningUpper, SubractWideningLower and SubtractWideningUpper since each of those are mapped to four different instructions (saddl{2}, saddw{2}, uaddl{2} and uaddw{2}; ssubl{2}, ssubw{2}, usubl{2} and usubw{2}) and do not fit into our table-driven model.

The change depends on removing HW_Flag_UnfixedSIMDSize in #35594 - I didn't want to keep adding the flag for the added intrinsics for no particular reason.

Fixes#32512

@echesakovechesakov added arch-arm64 area-System.Runtime.Intrinsics area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI labels Apr 29, 2020
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @tannergooding
Notify danmosemsft if you want to be subscribed.

@Dotnet-GitSync-Bot

Copy link
Copy Markdown
Collaborator

Note regarding the new-api-needs-documentation label:

This serves as a reminder for when your PR is modifying a ref *.cs file and adding/modifying public APIs, to please make sure the API implementation in the src *.cs file is documented with triple slash comments, so the PR reviewers can sign off that change.

@echesakov

Copy link
Copy Markdown
ContributorAuthor

@CarolEidt@tannergooding PTAL

@echesakov

Copy link
Copy Markdown
ContributorAuthor

@TamarChristinaArm PTAL

@echesakov
echesakovforce-pushed the Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics branch from f5004a2 to 97e7d77CompareApril 30, 2020 02:46
@echesakov

Copy link
Copy Markdown
ContributorAuthor

Rebased on top of 1f436e0

Comment threadsrc/coreclr/src/jit/hwintrinsic.cpp Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why do widening and narrowing need an "other" base type? The return type is always twice the size and of the same signedness of the inputs for these operations.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why do widening and narrowing need an "other" base type?

Narrowing intrinsics do not need - only AddWideningUpper and SubractWideningUpper do.

They don't need the operand 1 type itself (which is always going to be TYP_SIMD16) but the operand 1 element type so I can distinguish between each pair of these three

Vector128<int>AddWideningUpper(Vector128<short>left,Vector128<short>right);Vector128<short>AddWideningUpper(Vector128<short>left,Vector128<sbyte>right);Vector128<int>AddWideningUpper(Vector128<int>left,Vector128<short>right);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why can't we just use the second argument for the base type? That is what HW_Flag_BaseTypeFromSecondArg is for.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh nevermind, I missed that int = short + short and int = int + short need to be differentiated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Interesting - I renamed the gtIndexBaseType to gtOtherBaseType when I restructured the IR, because I figured there would be cases aside from gather where we'd need an additional base type - I didn't think it would be used that quickly :-)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should be able to get rid of this logic once we change the tables to take simdSize into account during lookup, correct?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can if we make simdSize and baseType of an intrinsic "independent" on each other - i.e. if we in addition to HW_Flag_BaseTypeFromFirstArg, HW_Flag_BaseTypeFromSecondArg we have HW_Flag_SimdSizeFromFirstArg, HW_Flag_SimdSizeFromSecondArg which would allow, for example, to set simdSize based on op1 and baseType from op2.

Then, the corresponding rows in hwintrinsiclistarm64.h could be updated as

HARDWARE_INTRINSIC(AdvSimd, AddWideningLower, 8, 2, {INS_saddl, INS_uaddl, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)
HARDWARE_INTRINSIC(AdvSimd, AddWideningLower, 16, 2, {INS_saddw, INS_uaddw, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)
HARDWARE_INTRINSIC(AdvSimd, AddWideningUpper, 8, 2, {INS_saddl2, INS_uaddl2, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)
HARDWARE_INTRINSIC(AdvSimd, AddWideningUpper, 16, 2, {INS_saddw2, INS_uaddw2, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)

In other words, if we allow x and y coordinates (where x is baseType and y is a tuple of (intrinsicId, simdSize)) of hwintrinsiclistarm64.h table to be orthogonal then the answer to your question is yes.

…nd SubtractWideningUpper in hwintrinsiccodegenarm64.cpp
…per and SubtractWideningUpper in hwintrinsic.cpp
@echesakov
echesakovforce-pushed the Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics branch from 97e7d77 to 417c253CompareApril 30, 2020 17:51
@echesakov

Copy link
Copy Markdown
ContributorAuthor

Fixed merge conflicts and rebased on top of aa81328

@CarolEidtCarolEidt left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM with some non-blocking comments & questions.
I only lightly reviewed the test helper functions and related changes.

Comment threadsrc/coreclr/src/jit/hwintrinsic.cpp Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Interesting - I renamed the gtIndexBaseType to gtOtherBaseType when I restructured the IR, because I figured there would be cases aside from gather where we'd need an additional base type - I didn't think it would be used that quickly :-)

insOpts opt = INS_OPTS_NONE;

if ((intrin.category == HW_Category_SIMDScalar) || (intrin.category == HW_Category_Scalar))
if (intrin.category == HW_Category_SIMDScalar)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The changes here highlight the fact that there are many subtleties with regard to emit sizes - the "actual type" of the baseType, the "raw" (emitTypeSize) of the baseType, and the emitSize of the node itself. I think it would be worth some additional comments both here and in the cases where we use emitTypeSize(node) for moves.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agree, I will follow up and add the comments

case NI_AdvSimd_AddWideningUpper:
case NI_AdvSimd_SubtractWideningLower:
case NI_AdvSimd_SubtractWideningUpper:
GetEmitter()->emitIns_R_R_R(ins, emitSize, targetReg, op1Reg, op2Reg, opt);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I believe this is identical to case NI_Crc32_ComputeCrc32: et al. Is there a reason you didn't combine it? Is the idea to keep them in order, and rely on the C++ compiler to de-duplicate the code? (I see there's some duplication already).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No, I just missed this. Thanks for spotting this.

I will follow up and de-duplicate this in a separate PR and I also need to address one of your another suggestions here

@echesakov

Copy link
Copy Markdown
ContributorAuthor

Both runtime (Mono Product Build Android x86 debug) and runtime (Mono Product Build OSX x64 debug) has succeeded on the second attempt - https://dev.azure.com/dnceng/public/_build/results?buildId=625278 - merging

@echesakov
echesakov merged commit a156293 into dotnet:masterMay 1, 2020
@echesakov
echesakov deleted the Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics branch May 1, 2020 02:47
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

arch-arm64area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIarea-System.Runtime.Intrinsicsnew-api-needs-documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

ARM additional arithmetic intrinsics

4 participants

@echesakov@Dotnet-GitSync-Bot@CarolEidt@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[Arm64] Implement ASIMD widening, narrowing, saturating intrinsics - #35612

Merged
echesakov merged 82 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics
May 1, 2020
Merged

[Arm64] Implement ASIMD widening, narrowing, saturating intrinsics#35612
echesakov merged 82 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics

Conversation

@echesakov

@echesakovechesakov commented Apr 29, 2020

Copy link
Copy Markdown
Contributor

Implement API as approved in #32512 (with corrections regarding their ISA class - see my comment here)

Note that I am using the following temporary names as per comment:

  • AddReturningHighNarrowUpper and AddReturningHighNarrowLower
  • AddReturningRoundedHighNarrowUpper and AddReturningRoundedHighNarrowLower
  • SubtractReturningHighNarrowUpper and SubtractReturningHighNarrowLower.

I will update the methods names after the next round of API review.

The change is quite straightforward - the special handling is required only for AddWideningLower, AddWideningUpper, SubractWideningLower and SubtractWideningUpper since each of those are mapped to four different instructions (saddl{2}, saddw{2}, uaddl{2} and uaddw{2}; ssubl{2}, ssubw{2}, usubl{2} and usubw{2}) and do not fit into our table-driven model.

The change depends on removing HW_Flag_UnfixedSIMDSize in #35594 - I didn't want to keep adding the flag for the added intrinsics for no particular reason.

Fixes#32512

@echesakovechesakov added arch-arm64 area-System.Runtime.Intrinsics area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI labels Apr 29, 2020
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @tannergooding
Notify danmosemsft if you want to be subscribed.

@Dotnet-GitSync-Bot

Copy link
Copy Markdown
Collaborator

Note regarding the new-api-needs-documentation label:

This serves as a reminder for when your PR is modifying a ref *.cs file and adding/modifying public APIs, to please make sure the API implementation in the src *.cs file is documented with triple slash comments, so the PR reviewers can sign off that change.

@echesakov

Copy link
Copy Markdown
ContributorAuthor

@CarolEidt@tannergooding PTAL

@echesakov

Copy link
Copy Markdown
ContributorAuthor

@TamarChristinaArm PTAL

@echesakov
echesakovforce-pushed the Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics branch from f5004a2 to 97e7d77CompareApril 30, 2020 02:46
@echesakov

Copy link
Copy Markdown
ContributorAuthor

Rebased on top of 1f436e0

Comment threadsrc/coreclr/src/jit/hwintrinsic.cpp Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why do widening and narrowing need an "other" base type? The return type is always twice the size and of the same signedness of the inputs for these operations.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why do widening and narrowing need an "other" base type?

Narrowing intrinsics do not need - only AddWideningUpper and SubractWideningUpper do.

They don't need the operand 1 type itself (which is always going to be TYP_SIMD16) but the operand 1 element type so I can distinguish between each pair of these three

Vector128<int>AddWideningUpper(Vector128<short>left,Vector128<short>right);Vector128<short>AddWideningUpper(Vector128<short>left,Vector128<sbyte>right);Vector128<int>AddWideningUpper(Vector128<int>left,Vector128<short>right);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why can't we just use the second argument for the base type? That is what HW_Flag_BaseTypeFromSecondArg is for.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh nevermind, I missed that int = short + short and int = int + short need to be differentiated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Interesting - I renamed the gtIndexBaseType to gtOtherBaseType when I restructured the IR, because I figured there would be cases aside from gather where we'd need an additional base type - I didn't think it would be used that quickly :-)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should be able to get rid of this logic once we change the tables to take simdSize into account during lookup, correct?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can if we make simdSize and baseType of an intrinsic "independent" on each other - i.e. if we in addition to HW_Flag_BaseTypeFromFirstArg, HW_Flag_BaseTypeFromSecondArg we have HW_Flag_SimdSizeFromFirstArg, HW_Flag_SimdSizeFromSecondArg which would allow, for example, to set simdSize based on op1 and baseType from op2.

Then, the corresponding rows in hwintrinsiclistarm64.h could be updated as

HARDWARE_INTRINSIC(AdvSimd, AddWideningLower, 8, 2, {INS_saddl, INS_uaddl, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)
HARDWARE_INTRINSIC(AdvSimd, AddWideningLower, 16, 2, {INS_saddw, INS_uaddw, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)
HARDWARE_INTRINSIC(AdvSimd, AddWideningUpper, 8, 2, {INS_saddl2, INS_uaddl2, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)
HARDWARE_INTRINSIC(AdvSimd, AddWideningUpper, 16, 2, {INS_saddw2, INS_uaddw2, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)

In other words, if we allow x and y coordinates (where x is baseType and y is a tuple of (intrinsicId, simdSize)) of hwintrinsiclistarm64.h table to be orthogonal then the answer to your question is yes.

…nd SubtractWideningUpper in hwintrinsiccodegenarm64.cpp
…per and SubtractWideningUpper in hwintrinsic.cpp
@echesakov
echesakovforce-pushed the Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics branch from 97e7d77 to 417c253CompareApril 30, 2020 17:51
@echesakov

Copy link
Copy Markdown
ContributorAuthor

Fixed merge conflicts and rebased on top of aa81328

@CarolEidtCarolEidt left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM with some non-blocking comments & questions.
I only lightly reviewed the test helper functions and related changes.

Comment threadsrc/coreclr/src/jit/hwintrinsic.cpp Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Interesting - I renamed the gtIndexBaseType to gtOtherBaseType when I restructured the IR, because I figured there would be cases aside from gather where we'd need an additional base type - I didn't think it would be used that quickly :-)

insOpts opt = INS_OPTS_NONE;

if ((intrin.category == HW_Category_SIMDScalar) || (intrin.category == HW_Category_Scalar))
if (intrin.category == HW_Category_SIMDScalar)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The changes here highlight the fact that there are many subtleties with regard to emit sizes - the "actual type" of the baseType, the "raw" (emitTypeSize) of the baseType, and the emitSize of the node itself. I think it would be worth some additional comments both here and in the cases where we use emitTypeSize(node) for moves.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agree, I will follow up and add the comments

case NI_AdvSimd_AddWideningUpper:
case NI_AdvSimd_SubtractWideningLower:
case NI_AdvSimd_SubtractWideningUpper:
GetEmitter()->emitIns_R_R_R(ins, emitSize, targetReg, op1Reg, op2Reg, opt);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I believe this is identical to case NI_Crc32_ComputeCrc32: et al. Is there a reason you didn't combine it? Is the idea to keep them in order, and rely on the C++ compiler to de-duplicate the code? (I see there's some duplication already).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No, I just missed this. Thanks for spotting this.

I will follow up and de-duplicate this in a separate PR and I also need to address one of your another suggestions here

@echesakov

Copy link
Copy Markdown
ContributorAuthor

Both runtime (Mono Product Build Android x86 debug) and runtime (Mono Product Build OSX x64 debug) has succeeded on the second attempt - https://dev.azure.com/dnceng/public/_build/results?buildId=625278 - merging

@echesakov
echesakov merged commit a156293 into dotnet:masterMay 1, 2020
@echesakov
echesakov deleted the Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics branch May 1, 2020 02:47
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

arch-arm64area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIarea-System.Runtime.Intrinsicsnew-api-needs-documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

ARM additional arithmetic intrinsics

4 participants

@echesakov@Dotnet-GitSync-Bot@CarolEidt@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[Arm64] Implement ASIMD widening, narrowing, saturating intrinsics - #35612

Merged
echesakov merged 82 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics
May 1, 2020
Merged

[Arm64] Implement ASIMD widening, narrowing, saturating intrinsics#35612
echesakov merged 82 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics

Conversation

@echesakov

@echesakovechesakov commented Apr 29, 2020

Copy link
Copy Markdown
Contributor

Implement API as approved in #32512 (with corrections regarding their ISA class - see my comment here)

Note that I am using the following temporary names as per comment:

  • AddReturningHighNarrowUpper and AddReturningHighNarrowLower
  • AddReturningRoundedHighNarrowUpper and AddReturningRoundedHighNarrowLower
  • SubtractReturningHighNarrowUpper and SubtractReturningHighNarrowLower.

I will update the methods names after the next round of API review.

The change is quite straightforward - the special handling is required only for AddWideningLower, AddWideningUpper, SubractWideningLower and SubtractWideningUpper since each of those are mapped to four different instructions (saddl{2}, saddw{2}, uaddl{2} and uaddw{2}; ssubl{2}, ssubw{2}, usubl{2} and usubw{2}) and do not fit into our table-driven model.

The change depends on removing HW_Flag_UnfixedSIMDSize in #35594 - I didn't want to keep adding the flag for the added intrinsics for no particular reason.

Fixes#32512

@echesakovechesakov added arch-arm64 area-System.Runtime.Intrinsics area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI labels Apr 29, 2020
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @tannergooding
Notify danmosemsft if you want to be subscribed.

@Dotnet-GitSync-Bot

Copy link
Copy Markdown
Collaborator

Note regarding the new-api-needs-documentation label:

This serves as a reminder for when your PR is modifying a ref *.cs file and adding/modifying public APIs, to please make sure the API implementation in the src *.cs file is documented with triple slash comments, so the PR reviewers can sign off that change.

@echesakov

Copy link
Copy Markdown
ContributorAuthor

@CarolEidt@tannergooding PTAL

@echesakov

Copy link
Copy Markdown
ContributorAuthor

@TamarChristinaArm PTAL

@echesakov
echesakovforce-pushed the Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics branch from f5004a2 to 97e7d77CompareApril 30, 2020 02:46
@echesakov

Copy link
Copy Markdown
ContributorAuthor

Rebased on top of 1f436e0

Comment threadsrc/coreclr/src/jit/hwintrinsic.cpp Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why do widening and narrowing need an "other" base type? The return type is always twice the size and of the same signedness of the inputs for these operations.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why do widening and narrowing need an "other" base type?

Narrowing intrinsics do not need - only AddWideningUpper and SubractWideningUpper do.

They don't need the operand 1 type itself (which is always going to be TYP_SIMD16) but the operand 1 element type so I can distinguish between each pair of these three

Vector128<int>AddWideningUpper(Vector128<short>left,Vector128<short>right);Vector128<short>AddWideningUpper(Vector128<short>left,Vector128<sbyte>right);Vector128<int>AddWideningUpper(Vector128<int>left,Vector128<short>right);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why can't we just use the second argument for the base type? That is what HW_Flag_BaseTypeFromSecondArg is for.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh nevermind, I missed that int = short + short and int = int + short need to be differentiated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Interesting - I renamed the gtIndexBaseType to gtOtherBaseType when I restructured the IR, because I figured there would be cases aside from gather where we'd need an additional base type - I didn't think it would be used that quickly :-)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should be able to get rid of this logic once we change the tables to take simdSize into account during lookup, correct?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can if we make simdSize and baseType of an intrinsic "independent" on each other - i.e. if we in addition to HW_Flag_BaseTypeFromFirstArg, HW_Flag_BaseTypeFromSecondArg we have HW_Flag_SimdSizeFromFirstArg, HW_Flag_SimdSizeFromSecondArg which would allow, for example, to set simdSize based on op1 and baseType from op2.

Then, the corresponding rows in hwintrinsiclistarm64.h could be updated as

HARDWARE_INTRINSIC(AdvSimd, AddWideningLower, 8, 2, {INS_saddl, INS_uaddl, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)
HARDWARE_INTRINSIC(AdvSimd, AddWideningLower, 16, 2, {INS_saddw, INS_uaddw, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)
HARDWARE_INTRINSIC(AdvSimd, AddWideningUpper, 8, 2, {INS_saddl2, INS_uaddl2, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)
HARDWARE_INTRINSIC(AdvSimd, AddWideningUpper, 16, 2, {INS_saddw2, INS_uaddw2, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)

In other words, if we allow x and y coordinates (where x is baseType and y is a tuple of (intrinsicId, simdSize)) of hwintrinsiclistarm64.h table to be orthogonal then the answer to your question is yes.

…nd SubtractWideningUpper in hwintrinsiccodegenarm64.cpp
…per and SubtractWideningUpper in hwintrinsic.cpp
@echesakov
echesakovforce-pushed the Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics branch from 97e7d77 to 417c253CompareApril 30, 2020 17:51
@echesakov

Copy link
Copy Markdown
ContributorAuthor

Fixed merge conflicts and rebased on top of aa81328

@CarolEidtCarolEidt left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM with some non-blocking comments & questions.
I only lightly reviewed the test helper functions and related changes.

Comment threadsrc/coreclr/src/jit/hwintrinsic.cpp Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Interesting - I renamed the gtIndexBaseType to gtOtherBaseType when I restructured the IR, because I figured there would be cases aside from gather where we'd need an additional base type - I didn't think it would be used that quickly :-)

insOpts opt = INS_OPTS_NONE;

if ((intrin.category == HW_Category_SIMDScalar) || (intrin.category == HW_Category_Scalar))
if (intrin.category == HW_Category_SIMDScalar)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The changes here highlight the fact that there are many subtleties with regard to emit sizes - the "actual type" of the baseType, the "raw" (emitTypeSize) of the baseType, and the emitSize of the node itself. I think it would be worth some additional comments both here and in the cases where we use emitTypeSize(node) for moves.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agree, I will follow up and add the comments

case NI_AdvSimd_AddWideningUpper:
case NI_AdvSimd_SubtractWideningLower:
case NI_AdvSimd_SubtractWideningUpper:
GetEmitter()->emitIns_R_R_R(ins, emitSize, targetReg, op1Reg, op2Reg, opt);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I believe this is identical to case NI_Crc32_ComputeCrc32: et al. Is there a reason you didn't combine it? Is the idea to keep them in order, and rely on the C++ compiler to de-duplicate the code? (I see there's some duplication already).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No, I just missed this. Thanks for spotting this.

I will follow up and de-duplicate this in a separate PR and I also need to address one of your another suggestions here

@echesakov

Copy link
Copy Markdown
ContributorAuthor

Both runtime (Mono Product Build Android x86 debug) and runtime (Mono Product Build OSX x64 debug) has succeeded on the second attempt - https://dev.azure.com/dnceng/public/_build/results?buildId=625278 - merging

@echesakov
echesakov merged commit a156293 into dotnet:masterMay 1, 2020
@echesakov
echesakov deleted the Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics branch May 1, 2020 02:47
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

arch-arm64area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIarea-System.Runtime.Intrinsicsnew-api-needs-documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

ARM additional arithmetic intrinsics

4 participants

@echesakov@Dotnet-GitSync-Bot@CarolEidt@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

[Arm64] Implement ASIMD widening, narrowing, saturating intrinsics - #35612

Merged
echesakov merged 82 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics
May 1, 2020
Merged

[Arm64] Implement ASIMD widening, narrowing, saturating intrinsics#35612
echesakov merged 82 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics

Conversation

@echesakov

@echesakovechesakov commented Apr 29, 2020

Copy link
Copy Markdown
Contributor

Implement API as approved in #32512 (with corrections regarding their ISA class - see my comment here)

Note that I am using the following temporary names as per comment:

  • AddReturningHighNarrowUpper and AddReturningHighNarrowLower
  • AddReturningRoundedHighNarrowUpper and AddReturningRoundedHighNarrowLower
  • SubtractReturningHighNarrowUpper and SubtractReturningHighNarrowLower.

I will update the methods names after the next round of API review.

The change is quite straightforward - the special handling is required only for AddWideningLower, AddWideningUpper, SubractWideningLower and SubtractWideningUpper since each of those are mapped to four different instructions (saddl{2}, saddw{2}, uaddl{2} and uaddw{2}; ssubl{2}, ssubw{2}, usubl{2} and usubw{2}) and do not fit into our table-driven model.

The change depends on removing HW_Flag_UnfixedSIMDSize in #35594 - I didn't want to keep adding the flag for the added intrinsics for no particular reason.

Fixes#32512

@echesakovechesakov added arch-arm64 area-System.Runtime.Intrinsics area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI labels Apr 29, 2020
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @tannergooding
Notify danmosemsft if you want to be subscribed.

@Dotnet-GitSync-Bot

Copy link
Copy Markdown
Collaborator

Note regarding the new-api-needs-documentation label:

This serves as a reminder for when your PR is modifying a ref *.cs file and adding/modifying public APIs, to please make sure the API implementation in the src *.cs file is documented with triple slash comments, so the PR reviewers can sign off that change.

@echesakov

Copy link
Copy Markdown
ContributorAuthor

@CarolEidt@tannergooding PTAL

@echesakov

Copy link
Copy Markdown
ContributorAuthor

@TamarChristinaArm PTAL

@echesakov
echesakovforce-pushed the Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics branch from f5004a2 to 97e7d77CompareApril 30, 2020 02:46
@echesakov

Copy link
Copy Markdown
ContributorAuthor

Rebased on top of 1f436e0

Comment threadsrc/coreclr/src/jit/hwintrinsic.cpp Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why do widening and narrowing need an "other" base type? The return type is always twice the size and of the same signedness of the inputs for these operations.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why do widening and narrowing need an "other" base type?

Narrowing intrinsics do not need - only AddWideningUpper and SubractWideningUpper do.

They don't need the operand 1 type itself (which is always going to be TYP_SIMD16) but the operand 1 element type so I can distinguish between each pair of these three

Vector128<int>AddWideningUpper(Vector128<short>left,Vector128<short>right);Vector128<short>AddWideningUpper(Vector128<short>left,Vector128<sbyte>right);Vector128<int>AddWideningUpper(Vector128<int>left,Vector128<short>right);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why can't we just use the second argument for the base type? That is what HW_Flag_BaseTypeFromSecondArg is for.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh nevermind, I missed that int = short + short and int = int + short need to be differentiated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Interesting - I renamed the gtIndexBaseType to gtOtherBaseType when I restructured the IR, because I figured there would be cases aside from gather where we'd need an additional base type - I didn't think it would be used that quickly :-)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should be able to get rid of this logic once we change the tables to take simdSize into account during lookup, correct?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can if we make simdSize and baseType of an intrinsic "independent" on each other - i.e. if we in addition to HW_Flag_BaseTypeFromFirstArg, HW_Flag_BaseTypeFromSecondArg we have HW_Flag_SimdSizeFromFirstArg, HW_Flag_SimdSizeFromSecondArg which would allow, for example, to set simdSize based on op1 and baseType from op2.

Then, the corresponding rows in hwintrinsiclistarm64.h could be updated as

HARDWARE_INTRINSIC(AdvSimd, AddWideningLower, 8, 2, {INS_saddl, INS_uaddl, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)
HARDWARE_INTRINSIC(AdvSimd, AddWideningLower, 16, 2, {INS_saddw, INS_uaddw, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)
HARDWARE_INTRINSIC(AdvSimd, AddWideningUpper, 8, 2, {INS_saddl2, INS_uaddl2, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)
HARDWARE_INTRINSIC(AdvSimd, AddWideningUpper, 16, 2, {INS_saddw2, INS_uaddw2, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)

In other words, if we allow x and y coordinates (where x is baseType and y is a tuple of (intrinsicId, simdSize)) of hwintrinsiclistarm64.h table to be orthogonal then the answer to your question is yes.

…nd SubtractWideningUpper in hwintrinsiccodegenarm64.cpp
…per and SubtractWideningUpper in hwintrinsic.cpp
@echesakov
echesakovforce-pushed the Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics branch from 97e7d77 to 417c253CompareApril 30, 2020 17:51
@echesakov

Copy link
Copy Markdown
ContributorAuthor

Fixed merge conflicts and rebased on top of aa81328

@CarolEidtCarolEidt left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM with some non-blocking comments & questions.
I only lightly reviewed the test helper functions and related changes.

Comment threadsrc/coreclr/src/jit/hwintrinsic.cpp Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Interesting - I renamed the gtIndexBaseType to gtOtherBaseType when I restructured the IR, because I figured there would be cases aside from gather where we'd need an additional base type - I didn't think it would be used that quickly :-)

insOpts opt = INS_OPTS_NONE;

if ((intrin.category == HW_Category_SIMDScalar) || (intrin.category == HW_Category_Scalar))
if (intrin.category == HW_Category_SIMDScalar)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The changes here highlight the fact that there are many subtleties with regard to emit sizes - the "actual type" of the baseType, the "raw" (emitTypeSize) of the baseType, and the emitSize of the node itself. I think it would be worth some additional comments both here and in the cases where we use emitTypeSize(node) for moves.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agree, I will follow up and add the comments

case NI_AdvSimd_AddWideningUpper:
case NI_AdvSimd_SubtractWideningLower:
case NI_AdvSimd_SubtractWideningUpper:
GetEmitter()->emitIns_R_R_R(ins, emitSize, targetReg, op1Reg, op2Reg, opt);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I believe this is identical to case NI_Crc32_ComputeCrc32: et al. Is there a reason you didn't combine it? Is the idea to keep them in order, and rely on the C++ compiler to de-duplicate the code? (I see there's some duplication already).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No, I just missed this. Thanks for spotting this.

I will follow up and de-duplicate this in a separate PR and I also need to address one of your another suggestions here

@echesakov

Copy link
Copy Markdown
ContributorAuthor

Both runtime (Mono Product Build Android x86 debug) and runtime (Mono Product Build OSX x64 debug) has succeeded on the second attempt - https://dev.azure.com/dnceng/public/_build/results?buildId=625278 - merging

@echesakov
echesakov merged commit a156293 into dotnet:masterMay 1, 2020
@echesakov
echesakov deleted the Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics branch May 1, 2020 02:47
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

arch-arm64area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIarea-System.Runtime.Intrinsicsnew-api-needs-documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

ARM additional arithmetic intrinsics

4 participants

@echesakov@Dotnet-GitSync-Bot@CarolEidt@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[Arm64] Implement ASIMD widening, narrowing, saturating intrinsics - #35612

Merged
echesakov merged 82 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics
May 1, 2020
Merged

[Arm64] Implement ASIMD widening, narrowing, saturating intrinsics#35612
echesakov merged 82 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics

Conversation

@echesakov

@echesakovechesakov commented Apr 29, 2020

Copy link
Copy Markdown
Contributor

Implement API as approved in #32512 (with corrections regarding their ISA class - see my comment here)

Note that I am using the following temporary names as per comment:

  • AddReturningHighNarrowUpper and AddReturningHighNarrowLower
  • AddReturningRoundedHighNarrowUpper and AddReturningRoundedHighNarrowLower
  • SubtractReturningHighNarrowUpper and SubtractReturningHighNarrowLower.

I will update the methods names after the next round of API review.

The change is quite straightforward - the special handling is required only for AddWideningLower, AddWideningUpper, SubractWideningLower and SubtractWideningUpper since each of those are mapped to four different instructions (saddl{2}, saddw{2}, uaddl{2} and uaddw{2}; ssubl{2}, ssubw{2}, usubl{2} and usubw{2}) and do not fit into our table-driven model.

The change depends on removing HW_Flag_UnfixedSIMDSize in #35594 - I didn't want to keep adding the flag for the added intrinsics for no particular reason.

Fixes#32512

@echesakovechesakov added arch-arm64 area-System.Runtime.Intrinsics area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI labels Apr 29, 2020
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @tannergooding
Notify danmosemsft if you want to be subscribed.

@Dotnet-GitSync-Bot

Copy link
Copy Markdown
Collaborator

Note regarding the new-api-needs-documentation label:

This serves as a reminder for when your PR is modifying a ref *.cs file and adding/modifying public APIs, to please make sure the API implementation in the src *.cs file is documented with triple slash comments, so the PR reviewers can sign off that change.

@echesakov

Copy link
Copy Markdown
ContributorAuthor

@CarolEidt@tannergooding PTAL

@echesakov

Copy link
Copy Markdown
ContributorAuthor

@TamarChristinaArm PTAL

@echesakov
echesakovforce-pushed the Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics branch from f5004a2 to 97e7d77CompareApril 30, 2020 02:46
@echesakov

Copy link
Copy Markdown
ContributorAuthor

Rebased on top of 1f436e0

Comment threadsrc/coreclr/src/jit/hwintrinsic.cpp Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why do widening and narrowing need an "other" base type? The return type is always twice the size and of the same signedness of the inputs for these operations.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why do widening and narrowing need an "other" base type?

Narrowing intrinsics do not need - only AddWideningUpper and SubractWideningUpper do.

They don't need the operand 1 type itself (which is always going to be TYP_SIMD16) but the operand 1 element type so I can distinguish between each pair of these three

Vector128<int>AddWideningUpper(Vector128<short>left,Vector128<short>right);Vector128<short>AddWideningUpper(Vector128<short>left,Vector128<sbyte>right);Vector128<int>AddWideningUpper(Vector128<int>left,Vector128<short>right);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why can't we just use the second argument for the base type? That is what HW_Flag_BaseTypeFromSecondArg is for.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh nevermind, I missed that int = short + short and int = int + short need to be differentiated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Interesting - I renamed the gtIndexBaseType to gtOtherBaseType when I restructured the IR, because I figured there would be cases aside from gather where we'd need an additional base type - I didn't think it would be used that quickly :-)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should be able to get rid of this logic once we change the tables to take simdSize into account during lookup, correct?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can if we make simdSize and baseType of an intrinsic "independent" on each other - i.e. if we in addition to HW_Flag_BaseTypeFromFirstArg, HW_Flag_BaseTypeFromSecondArg we have HW_Flag_SimdSizeFromFirstArg, HW_Flag_SimdSizeFromSecondArg which would allow, for example, to set simdSize based on op1 and baseType from op2.

Then, the corresponding rows in hwintrinsiclistarm64.h could be updated as

HARDWARE_INTRINSIC(AdvSimd, AddWideningLower, 8, 2, {INS_saddl, INS_uaddl, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)
HARDWARE_INTRINSIC(AdvSimd, AddWideningLower, 16, 2, {INS_saddw, INS_uaddw, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)
HARDWARE_INTRINSIC(AdvSimd, AddWideningUpper, 8, 2, {INS_saddl2, INS_uaddl2, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)
HARDWARE_INTRINSIC(AdvSimd, AddWideningUpper, 16, 2, {INS_saddw2, INS_uaddw2, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)

In other words, if we allow x and y coordinates (where x is baseType and y is a tuple of (intrinsicId, simdSize)) of hwintrinsiclistarm64.h table to be orthogonal then the answer to your question is yes.

…nd SubtractWideningUpper in hwintrinsiccodegenarm64.cpp
…per and SubtractWideningUpper in hwintrinsic.cpp
@echesakov
echesakovforce-pushed the Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics branch from 97e7d77 to 417c253CompareApril 30, 2020 17:51
@echesakov

Copy link
Copy Markdown
ContributorAuthor

Fixed merge conflicts and rebased on top of aa81328

@CarolEidtCarolEidt left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM with some non-blocking comments & questions.
I only lightly reviewed the test helper functions and related changes.

Comment threadsrc/coreclr/src/jit/hwintrinsic.cpp Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Interesting - I renamed the gtIndexBaseType to gtOtherBaseType when I restructured the IR, because I figured there would be cases aside from gather where we'd need an additional base type - I didn't think it would be used that quickly :-)

insOpts opt = INS_OPTS_NONE;

if ((intrin.category == HW_Category_SIMDScalar) || (intrin.category == HW_Category_Scalar))
if (intrin.category == HW_Category_SIMDScalar)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The changes here highlight the fact that there are many subtleties with regard to emit sizes - the "actual type" of the baseType, the "raw" (emitTypeSize) of the baseType, and the emitSize of the node itself. I think it would be worth some additional comments both here and in the cases where we use emitTypeSize(node) for moves.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agree, I will follow up and add the comments

case NI_AdvSimd_AddWideningUpper:
case NI_AdvSimd_SubtractWideningLower:
case NI_AdvSimd_SubtractWideningUpper:
GetEmitter()->emitIns_R_R_R(ins, emitSize, targetReg, op1Reg, op2Reg, opt);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I believe this is identical to case NI_Crc32_ComputeCrc32: et al. Is there a reason you didn't combine it? Is the idea to keep them in order, and rely on the C++ compiler to de-duplicate the code? (I see there's some duplication already).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No, I just missed this. Thanks for spotting this.

I will follow up and de-duplicate this in a separate PR and I also need to address one of your another suggestions here

@echesakov

Copy link
Copy Markdown
ContributorAuthor

Both runtime (Mono Product Build Android x86 debug) and runtime (Mono Product Build OSX x64 debug) has succeeded on the second attempt - https://dev.azure.com/dnceng/public/_build/results?buildId=625278 - merging

@echesakov
echesakov merged commit a156293 into dotnet:masterMay 1, 2020
@echesakov
echesakov deleted the Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics branch May 1, 2020 02:47
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

arch-arm64area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIarea-System.Runtime.Intrinsicsnew-api-needs-documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

ARM additional arithmetic intrinsics

4 participants

@echesakov@Dotnet-GitSync-Bot@CarolEidt@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[Arm64] Implement ASIMD widening, narrowing, saturating intrinsics - #35612

Merged
echesakov merged 82 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics
May 1, 2020
Merged

[Arm64] Implement ASIMD widening, narrowing, saturating intrinsics#35612
echesakov merged 82 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics

Conversation

@echesakov

@echesakovechesakov commented Apr 29, 2020

Copy link
Copy Markdown
Contributor

Implement API as approved in #32512 (with corrections regarding their ISA class - see my comment here)

Note that I am using the following temporary names as per comment:

  • AddReturningHighNarrowUpper and AddReturningHighNarrowLower
  • AddReturningRoundedHighNarrowUpper and AddReturningRoundedHighNarrowLower
  • SubtractReturningHighNarrowUpper and SubtractReturningHighNarrowLower.

I will update the methods names after the next round of API review.

The change is quite straightforward - the special handling is required only for AddWideningLower, AddWideningUpper, SubractWideningLower and SubtractWideningUpper since each of those are mapped to four different instructions (saddl{2}, saddw{2}, uaddl{2} and uaddw{2}; ssubl{2}, ssubw{2}, usubl{2} and usubw{2}) and do not fit into our table-driven model.

The change depends on removing HW_Flag_UnfixedSIMDSize in #35594 - I didn't want to keep adding the flag for the added intrinsics for no particular reason.

Fixes#32512

@echesakovechesakov added arch-arm64 area-System.Runtime.Intrinsics area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI labels Apr 29, 2020
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @tannergooding
Notify danmosemsft if you want to be subscribed.

@Dotnet-GitSync-Bot

Copy link
Copy Markdown
Collaborator

Note regarding the new-api-needs-documentation label:

This serves as a reminder for when your PR is modifying a ref *.cs file and adding/modifying public APIs, to please make sure the API implementation in the src *.cs file is documented with triple slash comments, so the PR reviewers can sign off that change.

@echesakov

Copy link
Copy Markdown
ContributorAuthor

@CarolEidt@tannergooding PTAL

@echesakov

Copy link
Copy Markdown
ContributorAuthor

@TamarChristinaArm PTAL

@echesakov
echesakovforce-pushed the Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics branch from f5004a2 to 97e7d77CompareApril 30, 2020 02:46
@echesakov

Copy link
Copy Markdown
ContributorAuthor

Rebased on top of 1f436e0

Comment threadsrc/coreclr/src/jit/hwintrinsic.cpp Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why do widening and narrowing need an "other" base type? The return type is always twice the size and of the same signedness of the inputs for these operations.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why do widening and narrowing need an "other" base type?

Narrowing intrinsics do not need - only AddWideningUpper and SubractWideningUpper do.

They don't need the operand 1 type itself (which is always going to be TYP_SIMD16) but the operand 1 element type so I can distinguish between each pair of these three

Vector128<int>AddWideningUpper(Vector128<short>left,Vector128<short>right);Vector128<short>AddWideningUpper(Vector128<short>left,Vector128<sbyte>right);Vector128<int>AddWideningUpper(Vector128<int>left,Vector128<short>right);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why can't we just use the second argument for the base type? That is what HW_Flag_BaseTypeFromSecondArg is for.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh nevermind, I missed that int = short + short and int = int + short need to be differentiated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Interesting - I renamed the gtIndexBaseType to gtOtherBaseType when I restructured the IR, because I figured there would be cases aside from gather where we'd need an additional base type - I didn't think it would be used that quickly :-)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should be able to get rid of this logic once we change the tables to take simdSize into account during lookup, correct?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can if we make simdSize and baseType of an intrinsic "independent" on each other - i.e. if we in addition to HW_Flag_BaseTypeFromFirstArg, HW_Flag_BaseTypeFromSecondArg we have HW_Flag_SimdSizeFromFirstArg, HW_Flag_SimdSizeFromSecondArg which would allow, for example, to set simdSize based on op1 and baseType from op2.

Then, the corresponding rows in hwintrinsiclistarm64.h could be updated as

HARDWARE_INTRINSIC(AdvSimd, AddWideningLower, 8, 2, {INS_saddl, INS_uaddl, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)
HARDWARE_INTRINSIC(AdvSimd, AddWideningLower, 16, 2, {INS_saddw, INS_uaddw, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)
HARDWARE_INTRINSIC(AdvSimd, AddWideningUpper, 8, 2, {INS_saddl2, INS_uaddl2, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)
HARDWARE_INTRINSIC(AdvSimd, AddWideningUpper, 16, 2, {INS_saddw2, INS_uaddw2, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)

In other words, if we allow x and y coordinates (where x is baseType and y is a tuple of (intrinsicId, simdSize)) of hwintrinsiclistarm64.h table to be orthogonal then the answer to your question is yes.

…nd SubtractWideningUpper in hwintrinsiccodegenarm64.cpp
…per and SubtractWideningUpper in hwintrinsic.cpp
@echesakov
echesakovforce-pushed the Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics branch from 97e7d77 to 417c253CompareApril 30, 2020 17:51
@echesakov

Copy link
Copy Markdown
ContributorAuthor

Fixed merge conflicts and rebased on top of aa81328

@CarolEidtCarolEidt left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM with some non-blocking comments & questions.
I only lightly reviewed the test helper functions and related changes.

Comment threadsrc/coreclr/src/jit/hwintrinsic.cpp Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Interesting - I renamed the gtIndexBaseType to gtOtherBaseType when I restructured the IR, because I figured there would be cases aside from gather where we'd need an additional base type - I didn't think it would be used that quickly :-)

insOpts opt = INS_OPTS_NONE;

if ((intrin.category == HW_Category_SIMDScalar) || (intrin.category == HW_Category_Scalar))
if (intrin.category == HW_Category_SIMDScalar)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The changes here highlight the fact that there are many subtleties with regard to emit sizes - the "actual type" of the baseType, the "raw" (emitTypeSize) of the baseType, and the emitSize of the node itself. I think it would be worth some additional comments both here and in the cases where we use emitTypeSize(node) for moves.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agree, I will follow up and add the comments

case NI_AdvSimd_AddWideningUpper:
case NI_AdvSimd_SubtractWideningLower:
case NI_AdvSimd_SubtractWideningUpper:
GetEmitter()->emitIns_R_R_R(ins, emitSize, targetReg, op1Reg, op2Reg, opt);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I believe this is identical to case NI_Crc32_ComputeCrc32: et al. Is there a reason you didn't combine it? Is the idea to keep them in order, and rely on the C++ compiler to de-duplicate the code? (I see there's some duplication already).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No, I just missed this. Thanks for spotting this.

I will follow up and de-duplicate this in a separate PR and I also need to address one of your another suggestions here

@echesakov

Copy link
Copy Markdown
ContributorAuthor

Both runtime (Mono Product Build Android x86 debug) and runtime (Mono Product Build OSX x64 debug) has succeeded on the second attempt - https://dev.azure.com/dnceng/public/_build/results?buildId=625278 - merging

@echesakov
echesakov merged commit a156293 into dotnet:masterMay 1, 2020
@echesakov
echesakov deleted the Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics branch May 1, 2020 02:47
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

arch-arm64area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIarea-System.Runtime.Intrinsicsnew-api-needs-documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

ARM additional arithmetic intrinsics

4 participants

@echesakov@Dotnet-GitSync-Bot@CarolEidt@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

[Arm64] Implement ASIMD widening, narrowing, saturating intrinsics - #35612

Merged
echesakov merged 82 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics
May 1, 2020
Merged

[Arm64] Implement ASIMD widening, narrowing, saturating intrinsics#35612
echesakov merged 82 commits into
dotnet:masterfrom
echesakov:Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics

Conversation

@echesakov

@echesakovechesakov commented Apr 29, 2020

Copy link
Copy Markdown
Contributor

Implement API as approved in #32512 (with corrections regarding their ISA class - see my comment here)

Note that I am using the following temporary names as per comment:

  • AddReturningHighNarrowUpper and AddReturningHighNarrowLower
  • AddReturningRoundedHighNarrowUpper and AddReturningRoundedHighNarrowLower
  • SubtractReturningHighNarrowUpper and SubtractReturningHighNarrowLower.

I will update the methods names after the next round of API review.

The change is quite straightforward - the special handling is required only for AddWideningLower, AddWideningUpper, SubractWideningLower and SubtractWideningUpper since each of those are mapped to four different instructions (saddl{2}, saddw{2}, uaddl{2} and uaddw{2}; ssubl{2}, ssubw{2}, usubl{2} and usubw{2}) and do not fit into our table-driven model.

The change depends on removing HW_Flag_UnfixedSIMDSize in #35594 - I didn't want to keep adding the flag for the added intrinsics for no particular reason.

Fixes#32512

@echesakovechesakov added arch-arm64 area-System.Runtime.Intrinsics area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI labels Apr 29, 2020
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @tannergooding
Notify danmosemsft if you want to be subscribed.

@Dotnet-GitSync-Bot

Copy link
Copy Markdown
Collaborator

Note regarding the new-api-needs-documentation label:

This serves as a reminder for when your PR is modifying a ref *.cs file and adding/modifying public APIs, to please make sure the API implementation in the src *.cs file is documented with triple slash comments, so the PR reviewers can sign off that change.

@echesakov

Copy link
Copy Markdown
ContributorAuthor

@CarolEidt@tannergooding PTAL

@echesakov

Copy link
Copy Markdown
ContributorAuthor

@TamarChristinaArm PTAL

@echesakov
echesakovforce-pushed the Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics branch from f5004a2 to 97e7d77CompareApril 30, 2020 02:46
@echesakov

Copy link
Copy Markdown
ContributorAuthor

Rebased on top of 1f436e0

Comment threadsrc/coreclr/src/jit/hwintrinsic.cpp Outdated

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why do widening and narrowing need an "other" base type? The return type is always twice the size and of the same signedness of the inputs for these operations.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why do widening and narrowing need an "other" base type?

Narrowing intrinsics do not need - only AddWideningUpper and SubractWideningUpper do.

They don't need the operand 1 type itself (which is always going to be TYP_SIMD16) but the operand 1 element type so I can distinguish between each pair of these three

Vector128<int>AddWideningUpper(Vector128<short>left,Vector128<short>right);Vector128<short>AddWideningUpper(Vector128<short>left,Vector128<sbyte>right);Vector128<int>AddWideningUpper(Vector128<int>left,Vector128<short>right);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why can't we just use the second argument for the base type? That is what HW_Flag_BaseTypeFromSecondArg is for.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh nevermind, I missed that int = short + short and int = int + short need to be differentiated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Interesting - I renamed the gtIndexBaseType to gtOtherBaseType when I restructured the IR, because I figured there would be cases aside from gather where we'd need an additional base type - I didn't think it would be used that quickly :-)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should be able to get rid of this logic once we change the tables to take simdSize into account during lookup, correct?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can if we make simdSize and baseType of an intrinsic "independent" on each other - i.e. if we in addition to HW_Flag_BaseTypeFromFirstArg, HW_Flag_BaseTypeFromSecondArg we have HW_Flag_SimdSizeFromFirstArg, HW_Flag_SimdSizeFromSecondArg which would allow, for example, to set simdSize based on op1 and baseType from op2.

Then, the corresponding rows in hwintrinsiclistarm64.h could be updated as

HARDWARE_INTRINSIC(AdvSimd, AddWideningLower, 8, 2, {INS_saddl, INS_uaddl, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)
HARDWARE_INTRINSIC(AdvSimd, AddWideningLower, 16, 2, {INS_saddw, INS_uaddw, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)
HARDWARE_INTRINSIC(AdvSimd, AddWideningUpper, 8, 2, {INS_saddl2, INS_uaddl2, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)
HARDWARE_INTRINSIC(AdvSimd, AddWideningUpper, 16, 2, {INS_saddw2, INS_uaddw2, … HW_Flag_SimdSizeFromFirstArg|HW_Flag_BaseTypeFromSecondArg)

In other words, if we allow x and y coordinates (where x is baseType and y is a tuple of (intrinsicId, simdSize)) of hwintrinsiclistarm64.h table to be orthogonal then the answer to your question is yes.

…nd SubtractWideningUpper in hwintrinsiccodegenarm64.cpp
…per and SubtractWideningUpper in hwintrinsic.cpp
@echesakov
echesakovforce-pushed the Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics branch from 97e7d77 to 417c253CompareApril 30, 2020 17:51
@echesakov

Copy link
Copy Markdown
ContributorAuthor

Fixed merge conflicts and rebased on top of aa81328

@CarolEidtCarolEidt left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM with some non-blocking comments & questions.
I only lightly reviewed the test helper functions and related changes.

Comment threadsrc/coreclr/src/jit/hwintrinsic.cpp Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Interesting - I renamed the gtIndexBaseType to gtOtherBaseType when I restructured the IR, because I figured there would be cases aside from gather where we'd need an additional base type - I didn't think it would be used that quickly :-)

insOpts opt = INS_OPTS_NONE;

if ((intrin.category == HW_Category_SIMDScalar) || (intrin.category == HW_Category_Scalar))
if (intrin.category == HW_Category_SIMDScalar)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The changes here highlight the fact that there are many subtleties with regard to emit sizes - the "actual type" of the baseType, the "raw" (emitTypeSize) of the baseType, and the emitSize of the node itself. I think it would be worth some additional comments both here and in the cases where we use emitTypeSize(node) for moves.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agree, I will follow up and add the comments

case NI_AdvSimd_AddWideningUpper:
case NI_AdvSimd_SubtractWideningLower:
case NI_AdvSimd_SubtractWideningUpper:
GetEmitter()->emitIns_R_R_R(ins, emitSize, targetReg, op1Reg, op2Reg, opt);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I believe this is identical to case NI_Crc32_ComputeCrc32: et al. Is there a reason you didn't combine it? Is the idea to keep them in order, and rely on the C++ compiler to de-duplicate the code? (I see there's some duplication already).

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No, I just missed this. Thanks for spotting this.

I will follow up and de-duplicate this in a separate PR and I also need to address one of your another suggestions here

@echesakov

Copy link
Copy Markdown
ContributorAuthor

Both runtime (Mono Product Build Android x86 debug) and runtime (Mono Product Build OSX x64 debug) has succeeded on the second attempt - https://dev.azure.com/dnceng/public/_build/results?buildId=625278 - merging

@echesakov
echesakov merged commit a156293 into dotnet:masterMay 1, 2020
@echesakov
echesakov deleted the Arm64-ASIMD-Widening-Narrowing-Saturating-Intrinsics branch May 1, 2020 02:47
@ghostghost locked as resolved and limited conversation to collaborators Dec 9, 2020
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

arch-arm64area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIarea-System.Runtime.Intrinsicsnew-api-needs-documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

ARM additional arithmetic intrinsics

4 participants

@echesakov@Dotnet-GitSync-Bot@CarolEidt@tannergooding