Constant folding for SIMD comparisons - #85584

Closed
EgorBo wants to merge 9 commits into
dotnet:mainfrom
EgorBo:simd-constant-fold-eq-ne
Closed

Constant folding for SIMD comparisons#85584
EgorBo wants to merge 9 commits into
dotnet:mainfrom
EgorBo:simd-constant-fold-eq-ne

Conversation

@EgorBo

@EgorBoEgorBo commented May 1, 2023

Copy link
Copy Markdown
Member

This PR extends constant folding for SIMD to support EQ/NE. Can also be extended for GT_GT/GE/LT/LE and, perhaps, floating point, but I was mainly interested in EQ/NE for integral types as I found a few real uses cases. Also, improved constant folding for RVA slightly.

boolTest(){returnIsHelloWorldString("Hello World SIMD constant folding!"u8);}boolIsHelloWorldString(ReadOnlySpan<byte>data){Debug.Assert(data.Lengthis>=16 and <=32);// Load data into two Vector128, the 2nd vector may overlap if neededrefbyteptr=refMemoryMarshal.GetReference(data);Vector128<byte>v1=Vector128.LoadUnsafe(refptr);Vector128<byte>v2=Vector128.LoadUnsafe(refptr,(nuint)(data.Length-Vector128<byte>.Count));// The constant utf8 string we compare against:varcns="Hello World SIMD constant folding!"u8;Vector128<byte>v1Cns=Vector128.Create(cns.Slice(0,Vector128<byte>.Count));Vector128<byte>v2Cns=Vector128.Create(cns.Slice(cns.Length-Vector128<byte>.Count));// Compare pairwisereturn((v1^v1Cns)|(v2^v2Cns))==Vector128<byte>.Zero;}

Codegen for Test():

; Method Test():bool:thismoveax,1ret; Total bytes of code: 6

the whole thing is collapsed into just return true after inlining and constant folding.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 1, 2023
@ghostghost assigned EgorBoMay 1, 2023
@ghost

ghost commented May 1, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

Can be easily extended for GT_GT/GE/LT/LE and, perhaps, floating point, but I was mainly interested in EQ/NE for integeres as I found a few real uses cases. E.g.

boolTest(){vardata="hello world simd folding demo :)"u8;returnIsHelloWorldString(refMemoryMarshal.GetReference(data));}[MethodImpl(MethodImplOptions.AggressiveInlining)]boolIsHelloWorldString(refbytedata){varv1=Vector128.LoadUnsafe(refdata);varv2=Vector128.LoadUnsafe(refdata,16);varv1Cns=Vector128.Create("hello world simd"u8);varv2Cns=Vector128.Create(" folding demo :)"u8);return((v1^v1Cns)|(v2^v2Cns))==Vector128<byte>.Zero;}

Codegen for Test():

; Method MyClass:Foo():bool:thismoveax,1ret; Total bytes of code: 6

the whole thing is collapsed into just return true after inlining and constant folding.

Author:EgorBo
Assignees:EgorBo
Labels:

area-CodeGen-coreclr

Milestone:-

Comment on lines +443 to +448
template <typename TBase>
TBase GetAllBitsSetScalar()
{
uint8_t bitWidth = (sizeof(TBase) * 8);
return static_cast<TBase>((1ULL << bitWidth) - 1);
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Won't this fail for int64_t and uint64_t?

Why not simply return ~static_cast<TBase>(0);? Then for float/double you could specialize if needed.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, addressed

Comment on lines +530 to +538
case GT_EQ:
{
return arg0 == arg1 ? GetAllBitsSetScalar<TBase>() : 0;
}

case GT_NE:
{
return arg0 != arg1 ? GetAllBitsSetScalar<TBase>() : 0;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is going to be incorrect for float/double since it will do a bitwise comparison.

You may need to special-case in EvaluateBinaryScalarSpecialized<float> before it calls this main method.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You might also be able to do it in EvaluateBinaryScalar instead but special-case GetAllBitsSetScalar<float>(), which is probably simpler overall.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This path is never taken for floating point, I added asserts just in case. I was mostly interested in non-fp cases and these can be added if needed

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Any reason we aren’t covering floating-point as well? I imagine this might get hits in ImageSharp or PDN.

Potentially in other apps, games, or graphics processing, etc

assert((vn1Type == vn2Type) && varTypeIsSIMD(vn1Type));
assert(!varTypeIsFloating(baseType));

ValueNum packed = EvaluateBinarySimd(vns, GT_EQ, scalar, vn1Type, baseType, arg0VN, arg1VN);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

shouldn't this be EvaluateBinarySimd(vns, oper, ...)?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I wonder if it would be simpler to implement this as EvaluateVector and then simplify check result IsAllBitsSet or !IsZero

Which would also simplify the other relational comparisons

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

shouldn't this be EvaluateBinarySimd(vns, oper, ...)?

It makes it harder to check output, doesn't it?
Currently I just pass EQ and then return AllBitsSet or !AllBitsSet depending on source oper

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does it? We have a simple property for any simd_T and it makes it easier to cover the other cases.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That is, if you just evaluate per element you get a result simd_T result and can then just check IsAllBitsSet or IsZero

return EvaluateBinarySimd(this, GT_NE, /* scalar */ false, type, baseType, arg0VN, arg1VN);
}
break;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we want to handle GT, LT, GE, and LE simultaneously?

Likewise given you've added the support for computing the vector version in order to make bool work, should we just handle the intrinsics that produce a vector as well?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Those had no hits so I just didn't want to add more code and tests 🙂 Although, even EQ/NE don't have hits, I just found a use case outside.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is the type of scenario where it’s a core comparison on SIMD and so we want to generally handle it even if our own code doesn’t have any hits today.

it’s odd to cover some of the relational comparisons and not the others

@EgorBo

EgorBo commented May 18, 2023

Copy link
Copy Markdown
MemberAuthor

No diffs and I won't have spare time to cover all cases with tests so closing for now

@EgorBoEgorBo closed this May 18, 2023
@ghostghost locked as resolved and limited conversation to collaborators Jun 17, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@EgorBo@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Constant folding for SIMD comparisons - #85584

Closed
EgorBo wants to merge 9 commits into
dotnet:mainfrom
EgorBo:simd-constant-fold-eq-ne
Closed

Constant folding for SIMD comparisons#85584
EgorBo wants to merge 9 commits into
dotnet:mainfrom
EgorBo:simd-constant-fold-eq-ne

Conversation

@EgorBo

@EgorBoEgorBo commented May 1, 2023

Copy link
Copy Markdown
Member

This PR extends constant folding for SIMD to support EQ/NE. Can also be extended for GT_GT/GE/LT/LE and, perhaps, floating point, but I was mainly interested in EQ/NE for integral types as I found a few real uses cases. Also, improved constant folding for RVA slightly.

boolTest(){returnIsHelloWorldString("Hello World SIMD constant folding!"u8);}boolIsHelloWorldString(ReadOnlySpan<byte>data){Debug.Assert(data.Lengthis>=16 and <=32);// Load data into two Vector128, the 2nd vector may overlap if neededrefbyteptr=refMemoryMarshal.GetReference(data);Vector128<byte>v1=Vector128.LoadUnsafe(refptr);Vector128<byte>v2=Vector128.LoadUnsafe(refptr,(nuint)(data.Length-Vector128<byte>.Count));// The constant utf8 string we compare against:varcns="Hello World SIMD constant folding!"u8;Vector128<byte>v1Cns=Vector128.Create(cns.Slice(0,Vector128<byte>.Count));Vector128<byte>v2Cns=Vector128.Create(cns.Slice(cns.Length-Vector128<byte>.Count));// Compare pairwisereturn((v1^v1Cns)|(v2^v2Cns))==Vector128<byte>.Zero;}

Codegen for Test():

; Method Test():bool:thismoveax,1ret; Total bytes of code: 6

the whole thing is collapsed into just return true after inlining and constant folding.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 1, 2023
@ghostghost assigned EgorBoMay 1, 2023
@ghost

ghost commented May 1, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

Can be easily extended for GT_GT/GE/LT/LE and, perhaps, floating point, but I was mainly interested in EQ/NE for integeres as I found a few real uses cases. E.g.

boolTest(){vardata="hello world simd folding demo :)"u8;returnIsHelloWorldString(refMemoryMarshal.GetReference(data));}[MethodImpl(MethodImplOptions.AggressiveInlining)]boolIsHelloWorldString(refbytedata){varv1=Vector128.LoadUnsafe(refdata);varv2=Vector128.LoadUnsafe(refdata,16);varv1Cns=Vector128.Create("hello world simd"u8);varv2Cns=Vector128.Create(" folding demo :)"u8);return((v1^v1Cns)|(v2^v2Cns))==Vector128<byte>.Zero;}

Codegen for Test():

; Method MyClass:Foo():bool:thismoveax,1ret; Total bytes of code: 6

the whole thing is collapsed into just return true after inlining and constant folding.

Author:EgorBo
Assignees:EgorBo
Labels:

area-CodeGen-coreclr

Milestone:-

Comment on lines +443 to +448
template <typename TBase>
TBase GetAllBitsSetScalar()
{
uint8_t bitWidth = (sizeof(TBase) * 8);
return static_cast<TBase>((1ULL << bitWidth) - 1);
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Won't this fail for int64_t and uint64_t?

Why not simply return ~static_cast<TBase>(0);? Then for float/double you could specialize if needed.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, addressed

Comment on lines +530 to +538
case GT_EQ:
{
return arg0 == arg1 ? GetAllBitsSetScalar<TBase>() : 0;
}

case GT_NE:
{
return arg0 != arg1 ? GetAllBitsSetScalar<TBase>() : 0;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is going to be incorrect for float/double since it will do a bitwise comparison.

You may need to special-case in EvaluateBinaryScalarSpecialized<float> before it calls this main method.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You might also be able to do it in EvaluateBinaryScalar instead but special-case GetAllBitsSetScalar<float>(), which is probably simpler overall.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This path is never taken for floating point, I added asserts just in case. I was mostly interested in non-fp cases and these can be added if needed

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Any reason we aren’t covering floating-point as well? I imagine this might get hits in ImageSharp or PDN.

Potentially in other apps, games, or graphics processing, etc

assert((vn1Type == vn2Type) && varTypeIsSIMD(vn1Type));
assert(!varTypeIsFloating(baseType));

ValueNum packed = EvaluateBinarySimd(vns, GT_EQ, scalar, vn1Type, baseType, arg0VN, arg1VN);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

shouldn't this be EvaluateBinarySimd(vns, oper, ...)?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I wonder if it would be simpler to implement this as EvaluateVector and then simplify check result IsAllBitsSet or !IsZero

Which would also simplify the other relational comparisons

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

shouldn't this be EvaluateBinarySimd(vns, oper, ...)?

It makes it harder to check output, doesn't it?
Currently I just pass EQ and then return AllBitsSet or !AllBitsSet depending on source oper

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does it? We have a simple property for any simd_T and it makes it easier to cover the other cases.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That is, if you just evaluate per element you get a result simd_T result and can then just check IsAllBitsSet or IsZero

return EvaluateBinarySimd(this, GT_NE, /* scalar */ false, type, baseType, arg0VN, arg1VN);
}
break;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we want to handle GT, LT, GE, and LE simultaneously?

Likewise given you've added the support for computing the vector version in order to make bool work, should we just handle the intrinsics that produce a vector as well?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Those had no hits so I just didn't want to add more code and tests 🙂 Although, even EQ/NE don't have hits, I just found a use case outside.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is the type of scenario where it’s a core comparison on SIMD and so we want to generally handle it even if our own code doesn’t have any hits today.

it’s odd to cover some of the relational comparisons and not the others

@EgorBo

EgorBo commented May 18, 2023

Copy link
Copy Markdown
MemberAuthor

No diffs and I won't have spare time to cover all cases with tests so closing for now

@EgorBoEgorBo closed this May 18, 2023
@ghostghost locked as resolved and limited conversation to collaborators Jun 17, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@EgorBo@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Constant folding for SIMD comparisons - #85584

Closed
EgorBo wants to merge 9 commits into
dotnet:mainfrom
EgorBo:simd-constant-fold-eq-ne
Closed

Constant folding for SIMD comparisons#85584
EgorBo wants to merge 9 commits into
dotnet:mainfrom
EgorBo:simd-constant-fold-eq-ne

Conversation

@EgorBo

@EgorBoEgorBo commented May 1, 2023

Copy link
Copy Markdown
Member

This PR extends constant folding for SIMD to support EQ/NE. Can also be extended for GT_GT/GE/LT/LE and, perhaps, floating point, but I was mainly interested in EQ/NE for integral types as I found a few real uses cases. Also, improved constant folding for RVA slightly.

boolTest(){returnIsHelloWorldString("Hello World SIMD constant folding!"u8);}boolIsHelloWorldString(ReadOnlySpan<byte>data){Debug.Assert(data.Lengthis>=16 and <=32);// Load data into two Vector128, the 2nd vector may overlap if neededrefbyteptr=refMemoryMarshal.GetReference(data);Vector128<byte>v1=Vector128.LoadUnsafe(refptr);Vector128<byte>v2=Vector128.LoadUnsafe(refptr,(nuint)(data.Length-Vector128<byte>.Count));// The constant utf8 string we compare against:varcns="Hello World SIMD constant folding!"u8;Vector128<byte>v1Cns=Vector128.Create(cns.Slice(0,Vector128<byte>.Count));Vector128<byte>v2Cns=Vector128.Create(cns.Slice(cns.Length-Vector128<byte>.Count));// Compare pairwisereturn((v1^v1Cns)|(v2^v2Cns))==Vector128<byte>.Zero;}

Codegen for Test():

; Method Test():bool:thismoveax,1ret; Total bytes of code: 6

the whole thing is collapsed into just return true after inlining and constant folding.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 1, 2023
@ghostghost assigned EgorBoMay 1, 2023
@ghost

ghost commented May 1, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

Can be easily extended for GT_GT/GE/LT/LE and, perhaps, floating point, but I was mainly interested in EQ/NE for integeres as I found a few real uses cases. E.g.

boolTest(){vardata="hello world simd folding demo :)"u8;returnIsHelloWorldString(refMemoryMarshal.GetReference(data));}[MethodImpl(MethodImplOptions.AggressiveInlining)]boolIsHelloWorldString(refbytedata){varv1=Vector128.LoadUnsafe(refdata);varv2=Vector128.LoadUnsafe(refdata,16);varv1Cns=Vector128.Create("hello world simd"u8);varv2Cns=Vector128.Create(" folding demo :)"u8);return((v1^v1Cns)|(v2^v2Cns))==Vector128<byte>.Zero;}

Codegen for Test():

; Method MyClass:Foo():bool:thismoveax,1ret; Total bytes of code: 6

the whole thing is collapsed into just return true after inlining and constant folding.

Author:EgorBo
Assignees:EgorBo
Labels:

area-CodeGen-coreclr

Milestone:-

Comment on lines +443 to +448
template <typename TBase>
TBase GetAllBitsSetScalar()
{
uint8_t bitWidth = (sizeof(TBase) * 8);
return static_cast<TBase>((1ULL << bitWidth) - 1);
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Won't this fail for int64_t and uint64_t?

Why not simply return ~static_cast<TBase>(0);? Then for float/double you could specialize if needed.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, addressed

Comment on lines +530 to +538
case GT_EQ:
{
return arg0 == arg1 ? GetAllBitsSetScalar<TBase>() : 0;
}

case GT_NE:
{
return arg0 != arg1 ? GetAllBitsSetScalar<TBase>() : 0;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is going to be incorrect for float/double since it will do a bitwise comparison.

You may need to special-case in EvaluateBinaryScalarSpecialized<float> before it calls this main method.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You might also be able to do it in EvaluateBinaryScalar instead but special-case GetAllBitsSetScalar<float>(), which is probably simpler overall.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This path is never taken for floating point, I added asserts just in case. I was mostly interested in non-fp cases and these can be added if needed

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Any reason we aren’t covering floating-point as well? I imagine this might get hits in ImageSharp or PDN.

Potentially in other apps, games, or graphics processing, etc

assert((vn1Type == vn2Type) && varTypeIsSIMD(vn1Type));
assert(!varTypeIsFloating(baseType));

ValueNum packed = EvaluateBinarySimd(vns, GT_EQ, scalar, vn1Type, baseType, arg0VN, arg1VN);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

shouldn't this be EvaluateBinarySimd(vns, oper, ...)?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I wonder if it would be simpler to implement this as EvaluateVector and then simplify check result IsAllBitsSet or !IsZero

Which would also simplify the other relational comparisons

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

shouldn't this be EvaluateBinarySimd(vns, oper, ...)?

It makes it harder to check output, doesn't it?
Currently I just pass EQ and then return AllBitsSet or !AllBitsSet depending on source oper

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does it? We have a simple property for any simd_T and it makes it easier to cover the other cases.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That is, if you just evaluate per element you get a result simd_T result and can then just check IsAllBitsSet or IsZero

return EvaluateBinarySimd(this, GT_NE, /* scalar */ false, type, baseType, arg0VN, arg1VN);
}
break;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we want to handle GT, LT, GE, and LE simultaneously?

Likewise given you've added the support for computing the vector version in order to make bool work, should we just handle the intrinsics that produce a vector as well?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Those had no hits so I just didn't want to add more code and tests 🙂 Although, even EQ/NE don't have hits, I just found a use case outside.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is the type of scenario where it’s a core comparison on SIMD and so we want to generally handle it even if our own code doesn’t have any hits today.

it’s odd to cover some of the relational comparisons and not the others

@EgorBo

EgorBo commented May 18, 2023

Copy link
Copy Markdown
MemberAuthor

No diffs and I won't have spare time to cover all cases with tests so closing for now

@EgorBoEgorBo closed this May 18, 2023
@ghostghost locked as resolved and limited conversation to collaborators Jun 17, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@EgorBo@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Constant folding for SIMD comparisons - #85584

Closed
EgorBo wants to merge 9 commits into
dotnet:mainfrom
EgorBo:simd-constant-fold-eq-ne
Closed

Constant folding for SIMD comparisons#85584
EgorBo wants to merge 9 commits into
dotnet:mainfrom
EgorBo:simd-constant-fold-eq-ne

Conversation

@EgorBo

@EgorBoEgorBo commented May 1, 2023

Copy link
Copy Markdown
Member

This PR extends constant folding for SIMD to support EQ/NE. Can also be extended for GT_GT/GE/LT/LE and, perhaps, floating point, but I was mainly interested in EQ/NE for integral types as I found a few real uses cases. Also, improved constant folding for RVA slightly.

boolTest(){returnIsHelloWorldString("Hello World SIMD constant folding!"u8);}boolIsHelloWorldString(ReadOnlySpan<byte>data){Debug.Assert(data.Lengthis>=16 and <=32);// Load data into two Vector128, the 2nd vector may overlap if neededrefbyteptr=refMemoryMarshal.GetReference(data);Vector128<byte>v1=Vector128.LoadUnsafe(refptr);Vector128<byte>v2=Vector128.LoadUnsafe(refptr,(nuint)(data.Length-Vector128<byte>.Count));// The constant utf8 string we compare against:varcns="Hello World SIMD constant folding!"u8;Vector128<byte>v1Cns=Vector128.Create(cns.Slice(0,Vector128<byte>.Count));Vector128<byte>v2Cns=Vector128.Create(cns.Slice(cns.Length-Vector128<byte>.Count));// Compare pairwisereturn((v1^v1Cns)|(v2^v2Cns))==Vector128<byte>.Zero;}

Codegen for Test():

; Method Test():bool:thismoveax,1ret; Total bytes of code: 6

the whole thing is collapsed into just return true after inlining and constant folding.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 1, 2023
@ghostghost assigned EgorBoMay 1, 2023
@ghost

ghost commented May 1, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

Can be easily extended for GT_GT/GE/LT/LE and, perhaps, floating point, but I was mainly interested in EQ/NE for integeres as I found a few real uses cases. E.g.

boolTest(){vardata="hello world simd folding demo :)"u8;returnIsHelloWorldString(refMemoryMarshal.GetReference(data));}[MethodImpl(MethodImplOptions.AggressiveInlining)]boolIsHelloWorldString(refbytedata){varv1=Vector128.LoadUnsafe(refdata);varv2=Vector128.LoadUnsafe(refdata,16);varv1Cns=Vector128.Create("hello world simd"u8);varv2Cns=Vector128.Create(" folding demo :)"u8);return((v1^v1Cns)|(v2^v2Cns))==Vector128<byte>.Zero;}

Codegen for Test():

; Method MyClass:Foo():bool:thismoveax,1ret; Total bytes of code: 6

the whole thing is collapsed into just return true after inlining and constant folding.

Author:EgorBo
Assignees:EgorBo
Labels:

area-CodeGen-coreclr

Milestone:-

Comment on lines +443 to +448
template <typename TBase>
TBase GetAllBitsSetScalar()
{
uint8_t bitWidth = (sizeof(TBase) * 8);
return static_cast<TBase>((1ULL << bitWidth) - 1);
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Won't this fail for int64_t and uint64_t?

Why not simply return ~static_cast<TBase>(0);? Then for float/double you could specialize if needed.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, addressed

Comment on lines +530 to +538
case GT_EQ:
{
return arg0 == arg1 ? GetAllBitsSetScalar<TBase>() : 0;
}

case GT_NE:
{
return arg0 != arg1 ? GetAllBitsSetScalar<TBase>() : 0;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is going to be incorrect for float/double since it will do a bitwise comparison.

You may need to special-case in EvaluateBinaryScalarSpecialized<float> before it calls this main method.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You might also be able to do it in EvaluateBinaryScalar instead but special-case GetAllBitsSetScalar<float>(), which is probably simpler overall.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This path is never taken for floating point, I added asserts just in case. I was mostly interested in non-fp cases and these can be added if needed

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Any reason we aren’t covering floating-point as well? I imagine this might get hits in ImageSharp or PDN.

Potentially in other apps, games, or graphics processing, etc

assert((vn1Type == vn2Type) && varTypeIsSIMD(vn1Type));
assert(!varTypeIsFloating(baseType));

ValueNum packed = EvaluateBinarySimd(vns, GT_EQ, scalar, vn1Type, baseType, arg0VN, arg1VN);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

shouldn't this be EvaluateBinarySimd(vns, oper, ...)?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I wonder if it would be simpler to implement this as EvaluateVector and then simplify check result IsAllBitsSet or !IsZero

Which would also simplify the other relational comparisons

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

shouldn't this be EvaluateBinarySimd(vns, oper, ...)?

It makes it harder to check output, doesn't it?
Currently I just pass EQ and then return AllBitsSet or !AllBitsSet depending on source oper

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does it? We have a simple property for any simd_T and it makes it easier to cover the other cases.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That is, if you just evaluate per element you get a result simd_T result and can then just check IsAllBitsSet or IsZero

return EvaluateBinarySimd(this, GT_NE, /* scalar */ false, type, baseType, arg0VN, arg1VN);
}
break;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we want to handle GT, LT, GE, and LE simultaneously?

Likewise given you've added the support for computing the vector version in order to make bool work, should we just handle the intrinsics that produce a vector as well?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Those had no hits so I just didn't want to add more code and tests 🙂 Although, even EQ/NE don't have hits, I just found a use case outside.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is the type of scenario where it’s a core comparison on SIMD and so we want to generally handle it even if our own code doesn’t have any hits today.

it’s odd to cover some of the relational comparisons and not the others

@EgorBo

EgorBo commented May 18, 2023

Copy link
Copy Markdown
MemberAuthor

No diffs and I won't have spare time to cover all cases with tests so closing for now

@EgorBoEgorBo closed this May 18, 2023
@ghostghost locked as resolved and limited conversation to collaborators Jun 17, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@EgorBo@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Constant folding for SIMD comparisons - #85584

Closed
EgorBo wants to merge 9 commits into
dotnet:mainfrom
EgorBo:simd-constant-fold-eq-ne
Closed

Constant folding for SIMD comparisons#85584
EgorBo wants to merge 9 commits into
dotnet:mainfrom
EgorBo:simd-constant-fold-eq-ne

Conversation

@EgorBo

@EgorBoEgorBo commented May 1, 2023

Copy link
Copy Markdown
Member

This PR extends constant folding for SIMD to support EQ/NE. Can also be extended for GT_GT/GE/LT/LE and, perhaps, floating point, but I was mainly interested in EQ/NE for integral types as I found a few real uses cases. Also, improved constant folding for RVA slightly.

boolTest(){returnIsHelloWorldString("Hello World SIMD constant folding!"u8);}boolIsHelloWorldString(ReadOnlySpan<byte>data){Debug.Assert(data.Lengthis>=16 and <=32);// Load data into two Vector128, the 2nd vector may overlap if neededrefbyteptr=refMemoryMarshal.GetReference(data);Vector128<byte>v1=Vector128.LoadUnsafe(refptr);Vector128<byte>v2=Vector128.LoadUnsafe(refptr,(nuint)(data.Length-Vector128<byte>.Count));// The constant utf8 string we compare against:varcns="Hello World SIMD constant folding!"u8;Vector128<byte>v1Cns=Vector128.Create(cns.Slice(0,Vector128<byte>.Count));Vector128<byte>v2Cns=Vector128.Create(cns.Slice(cns.Length-Vector128<byte>.Count));// Compare pairwisereturn((v1^v1Cns)|(v2^v2Cns))==Vector128<byte>.Zero;}

Codegen for Test():

; Method Test():bool:thismoveax,1ret; Total bytes of code: 6

the whole thing is collapsed into just return true after inlining and constant folding.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 1, 2023
@ghostghost assigned EgorBoMay 1, 2023
@ghost

ghost commented May 1, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

Can be easily extended for GT_GT/GE/LT/LE and, perhaps, floating point, but I was mainly interested in EQ/NE for integeres as I found a few real uses cases. E.g.

boolTest(){vardata="hello world simd folding demo :)"u8;returnIsHelloWorldString(refMemoryMarshal.GetReference(data));}[MethodImpl(MethodImplOptions.AggressiveInlining)]boolIsHelloWorldString(refbytedata){varv1=Vector128.LoadUnsafe(refdata);varv2=Vector128.LoadUnsafe(refdata,16);varv1Cns=Vector128.Create("hello world simd"u8);varv2Cns=Vector128.Create(" folding demo :)"u8);return((v1^v1Cns)|(v2^v2Cns))==Vector128<byte>.Zero;}

Codegen for Test():

; Method MyClass:Foo():bool:thismoveax,1ret; Total bytes of code: 6

the whole thing is collapsed into just return true after inlining and constant folding.

Author:EgorBo
Assignees:EgorBo
Labels:

area-CodeGen-coreclr

Milestone:-

Comment on lines +443 to +448
template <typename TBase>
TBase GetAllBitsSetScalar()
{
uint8_t bitWidth = (sizeof(TBase) * 8);
return static_cast<TBase>((1ULL << bitWidth) - 1);
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Won't this fail for int64_t and uint64_t?

Why not simply return ~static_cast<TBase>(0);? Then for float/double you could specialize if needed.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, addressed

Comment on lines +530 to +538
case GT_EQ:
{
return arg0 == arg1 ? GetAllBitsSetScalar<TBase>() : 0;
}

case GT_NE:
{
return arg0 != arg1 ? GetAllBitsSetScalar<TBase>() : 0;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is going to be incorrect for float/double since it will do a bitwise comparison.

You may need to special-case in EvaluateBinaryScalarSpecialized<float> before it calls this main method.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You might also be able to do it in EvaluateBinaryScalar instead but special-case GetAllBitsSetScalar<float>(), which is probably simpler overall.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This path is never taken for floating point, I added asserts just in case. I was mostly interested in non-fp cases and these can be added if needed

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Any reason we aren’t covering floating-point as well? I imagine this might get hits in ImageSharp or PDN.

Potentially in other apps, games, or graphics processing, etc

assert((vn1Type == vn2Type) && varTypeIsSIMD(vn1Type));
assert(!varTypeIsFloating(baseType));

ValueNum packed = EvaluateBinarySimd(vns, GT_EQ, scalar, vn1Type, baseType, arg0VN, arg1VN);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

shouldn't this be EvaluateBinarySimd(vns, oper, ...)?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I wonder if it would be simpler to implement this as EvaluateVector and then simplify check result IsAllBitsSet or !IsZero

Which would also simplify the other relational comparisons

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

shouldn't this be EvaluateBinarySimd(vns, oper, ...)?

It makes it harder to check output, doesn't it?
Currently I just pass EQ and then return AllBitsSet or !AllBitsSet depending on source oper

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does it? We have a simple property for any simd_T and it makes it easier to cover the other cases.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That is, if you just evaluate per element you get a result simd_T result and can then just check IsAllBitsSet or IsZero

return EvaluateBinarySimd(this, GT_NE, /* scalar */ false, type, baseType, arg0VN, arg1VN);
}
break;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we want to handle GT, LT, GE, and LE simultaneously?

Likewise given you've added the support for computing the vector version in order to make bool work, should we just handle the intrinsics that produce a vector as well?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Those had no hits so I just didn't want to add more code and tests 🙂 Although, even EQ/NE don't have hits, I just found a use case outside.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is the type of scenario where it’s a core comparison on SIMD and so we want to generally handle it even if our own code doesn’t have any hits today.

it’s odd to cover some of the relational comparisons and not the others

@EgorBo

EgorBo commented May 18, 2023

Copy link
Copy Markdown
MemberAuthor

No diffs and I won't have spare time to cover all cases with tests so closing for now

@EgorBoEgorBo closed this May 18, 2023
@ghostghost locked as resolved and limited conversation to collaborators Jun 17, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@EgorBo@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Constant folding for SIMD comparisons - #85584

Closed
EgorBo wants to merge 9 commits into
dotnet:mainfrom
EgorBo:simd-constant-fold-eq-ne
Closed

Constant folding for SIMD comparisons#85584
EgorBo wants to merge 9 commits into
dotnet:mainfrom
EgorBo:simd-constant-fold-eq-ne

Conversation

@EgorBo

@EgorBoEgorBo commented May 1, 2023

Copy link
Copy Markdown
Member

This PR extends constant folding for SIMD to support EQ/NE. Can also be extended for GT_GT/GE/LT/LE and, perhaps, floating point, but I was mainly interested in EQ/NE for integral types as I found a few real uses cases. Also, improved constant folding for RVA slightly.

boolTest(){returnIsHelloWorldString("Hello World SIMD constant folding!"u8);}boolIsHelloWorldString(ReadOnlySpan<byte>data){Debug.Assert(data.Lengthis>=16 and <=32);// Load data into two Vector128, the 2nd vector may overlap if neededrefbyteptr=refMemoryMarshal.GetReference(data);Vector128<byte>v1=Vector128.LoadUnsafe(refptr);Vector128<byte>v2=Vector128.LoadUnsafe(refptr,(nuint)(data.Length-Vector128<byte>.Count));// The constant utf8 string we compare against:varcns="Hello World SIMD constant folding!"u8;Vector128<byte>v1Cns=Vector128.Create(cns.Slice(0,Vector128<byte>.Count));Vector128<byte>v2Cns=Vector128.Create(cns.Slice(cns.Length-Vector128<byte>.Count));// Compare pairwisereturn((v1^v1Cns)|(v2^v2Cns))==Vector128<byte>.Zero;}

Codegen for Test():

; Method Test():bool:thismoveax,1ret; Total bytes of code: 6

the whole thing is collapsed into just return true after inlining and constant folding.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 1, 2023
@ghostghost assigned EgorBoMay 1, 2023
@ghost

ghost commented May 1, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

Can be easily extended for GT_GT/GE/LT/LE and, perhaps, floating point, but I was mainly interested in EQ/NE for integeres as I found a few real uses cases. E.g.

boolTest(){vardata="hello world simd folding demo :)"u8;returnIsHelloWorldString(refMemoryMarshal.GetReference(data));}[MethodImpl(MethodImplOptions.AggressiveInlining)]boolIsHelloWorldString(refbytedata){varv1=Vector128.LoadUnsafe(refdata);varv2=Vector128.LoadUnsafe(refdata,16);varv1Cns=Vector128.Create("hello world simd"u8);varv2Cns=Vector128.Create(" folding demo :)"u8);return((v1^v1Cns)|(v2^v2Cns))==Vector128<byte>.Zero;}

Codegen for Test():

; Method MyClass:Foo():bool:thismoveax,1ret; Total bytes of code: 6

the whole thing is collapsed into just return true after inlining and constant folding.

Author:EgorBo
Assignees:EgorBo
Labels:

area-CodeGen-coreclr

Milestone:-

Comment on lines +443 to +448
template <typename TBase>
TBase GetAllBitsSetScalar()
{
uint8_t bitWidth = (sizeof(TBase) * 8);
return static_cast<TBase>((1ULL << bitWidth) - 1);
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Won't this fail for int64_t and uint64_t?

Why not simply return ~static_cast<TBase>(0);? Then for float/double you could specialize if needed.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, addressed

Comment on lines +530 to +538
case GT_EQ:
{
return arg0 == arg1 ? GetAllBitsSetScalar<TBase>() : 0;
}

case GT_NE:
{
return arg0 != arg1 ? GetAllBitsSetScalar<TBase>() : 0;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is going to be incorrect for float/double since it will do a bitwise comparison.

You may need to special-case in EvaluateBinaryScalarSpecialized<float> before it calls this main method.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You might also be able to do it in EvaluateBinaryScalar instead but special-case GetAllBitsSetScalar<float>(), which is probably simpler overall.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This path is never taken for floating point, I added asserts just in case. I was mostly interested in non-fp cases and these can be added if needed

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Any reason we aren’t covering floating-point as well? I imagine this might get hits in ImageSharp or PDN.

Potentially in other apps, games, or graphics processing, etc

assert((vn1Type == vn2Type) && varTypeIsSIMD(vn1Type));
assert(!varTypeIsFloating(baseType));

ValueNum packed = EvaluateBinarySimd(vns, GT_EQ, scalar, vn1Type, baseType, arg0VN, arg1VN);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

shouldn't this be EvaluateBinarySimd(vns, oper, ...)?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I wonder if it would be simpler to implement this as EvaluateVector and then simplify check result IsAllBitsSet or !IsZero

Which would also simplify the other relational comparisons

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

shouldn't this be EvaluateBinarySimd(vns, oper, ...)?

It makes it harder to check output, doesn't it?
Currently I just pass EQ and then return AllBitsSet or !AllBitsSet depending on source oper

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does it? We have a simple property for any simd_T and it makes it easier to cover the other cases.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That is, if you just evaluate per element you get a result simd_T result and can then just check IsAllBitsSet or IsZero

return EvaluateBinarySimd(this, GT_NE, /* scalar */ false, type, baseType, arg0VN, arg1VN);
}
break;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we want to handle GT, LT, GE, and LE simultaneously?

Likewise given you've added the support for computing the vector version in order to make bool work, should we just handle the intrinsics that produce a vector as well?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Those had no hits so I just didn't want to add more code and tests 🙂 Although, even EQ/NE don't have hits, I just found a use case outside.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is the type of scenario where it’s a core comparison on SIMD and so we want to generally handle it even if our own code doesn’t have any hits today.

it’s odd to cover some of the relational comparisons and not the others

@EgorBo

EgorBo commented May 18, 2023

Copy link
Copy Markdown
MemberAuthor

No diffs and I won't have spare time to cover all cases with tests so closing for now

@EgorBoEgorBo closed this May 18, 2023
@ghostghost locked as resolved and limited conversation to collaborators Jun 17, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@EgorBo@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Constant folding for SIMD comparisons - #85584

Closed
EgorBo wants to merge 9 commits into
dotnet:mainfrom
EgorBo:simd-constant-fold-eq-ne
Closed

Constant folding for SIMD comparisons#85584
EgorBo wants to merge 9 commits into
dotnet:mainfrom
EgorBo:simd-constant-fold-eq-ne

Conversation

@EgorBo

@EgorBoEgorBo commented May 1, 2023

Copy link
Copy Markdown
Member

This PR extends constant folding for SIMD to support EQ/NE. Can also be extended for GT_GT/GE/LT/LE and, perhaps, floating point, but I was mainly interested in EQ/NE for integral types as I found a few real uses cases. Also, improved constant folding for RVA slightly.

boolTest(){returnIsHelloWorldString("Hello World SIMD constant folding!"u8);}boolIsHelloWorldString(ReadOnlySpan<byte>data){Debug.Assert(data.Lengthis>=16 and <=32);// Load data into two Vector128, the 2nd vector may overlap if neededrefbyteptr=refMemoryMarshal.GetReference(data);Vector128<byte>v1=Vector128.LoadUnsafe(refptr);Vector128<byte>v2=Vector128.LoadUnsafe(refptr,(nuint)(data.Length-Vector128<byte>.Count));// The constant utf8 string we compare against:varcns="Hello World SIMD constant folding!"u8;Vector128<byte>v1Cns=Vector128.Create(cns.Slice(0,Vector128<byte>.Count));Vector128<byte>v2Cns=Vector128.Create(cns.Slice(cns.Length-Vector128<byte>.Count));// Compare pairwisereturn((v1^v1Cns)|(v2^v2Cns))==Vector128<byte>.Zero;}

Codegen for Test():

; Method Test():bool:thismoveax,1ret; Total bytes of code: 6

the whole thing is collapsed into just return true after inlining and constant folding.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 1, 2023
@ghostghost assigned EgorBoMay 1, 2023
@ghost

ghost commented May 1, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

Can be easily extended for GT_GT/GE/LT/LE and, perhaps, floating point, but I was mainly interested in EQ/NE for integeres as I found a few real uses cases. E.g.

boolTest(){vardata="hello world simd folding demo :)"u8;returnIsHelloWorldString(refMemoryMarshal.GetReference(data));}[MethodImpl(MethodImplOptions.AggressiveInlining)]boolIsHelloWorldString(refbytedata){varv1=Vector128.LoadUnsafe(refdata);varv2=Vector128.LoadUnsafe(refdata,16);varv1Cns=Vector128.Create("hello world simd"u8);varv2Cns=Vector128.Create(" folding demo :)"u8);return((v1^v1Cns)|(v2^v2Cns))==Vector128<byte>.Zero;}

Codegen for Test():

; Method MyClass:Foo():bool:thismoveax,1ret; Total bytes of code: 6

the whole thing is collapsed into just return true after inlining and constant folding.

Author:EgorBo
Assignees:EgorBo
Labels:

area-CodeGen-coreclr

Milestone:-

Comment on lines +443 to +448
template <typename TBase>
TBase GetAllBitsSetScalar()
{
uint8_t bitWidth = (sizeof(TBase) * 8);
return static_cast<TBase>((1ULL << bitWidth) - 1);
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Won't this fail for int64_t and uint64_t?

Why not simply return ~static_cast<TBase>(0);? Then for float/double you could specialize if needed.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, addressed

Comment on lines +530 to +538
case GT_EQ:
{
return arg0 == arg1 ? GetAllBitsSetScalar<TBase>() : 0;
}

case GT_NE:
{
return arg0 != arg1 ? GetAllBitsSetScalar<TBase>() : 0;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is going to be incorrect for float/double since it will do a bitwise comparison.

You may need to special-case in EvaluateBinaryScalarSpecialized<float> before it calls this main method.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You might also be able to do it in EvaluateBinaryScalar instead but special-case GetAllBitsSetScalar<float>(), which is probably simpler overall.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This path is never taken for floating point, I added asserts just in case. I was mostly interested in non-fp cases and these can be added if needed

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Any reason we aren’t covering floating-point as well? I imagine this might get hits in ImageSharp or PDN.

Potentially in other apps, games, or graphics processing, etc

assert((vn1Type == vn2Type) && varTypeIsSIMD(vn1Type));
assert(!varTypeIsFloating(baseType));

ValueNum packed = EvaluateBinarySimd(vns, GT_EQ, scalar, vn1Type, baseType, arg0VN, arg1VN);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

shouldn't this be EvaluateBinarySimd(vns, oper, ...)?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I wonder if it would be simpler to implement this as EvaluateVector and then simplify check result IsAllBitsSet or !IsZero

Which would also simplify the other relational comparisons

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

shouldn't this be EvaluateBinarySimd(vns, oper, ...)?

It makes it harder to check output, doesn't it?
Currently I just pass EQ and then return AllBitsSet or !AllBitsSet depending on source oper

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does it? We have a simple property for any simd_T and it makes it easier to cover the other cases.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That is, if you just evaluate per element you get a result simd_T result and can then just check IsAllBitsSet or IsZero

return EvaluateBinarySimd(this, GT_NE, /* scalar */ false, type, baseType, arg0VN, arg1VN);
}
break;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we want to handle GT, LT, GE, and LE simultaneously?

Likewise given you've added the support for computing the vector version in order to make bool work, should we just handle the intrinsics that produce a vector as well?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Those had no hits so I just didn't want to add more code and tests 🙂 Although, even EQ/NE don't have hits, I just found a use case outside.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is the type of scenario where it’s a core comparison on SIMD and so we want to generally handle it even if our own code doesn’t have any hits today.

it’s odd to cover some of the relational comparisons and not the others

@EgorBo

EgorBo commented May 18, 2023

Copy link
Copy Markdown
MemberAuthor

No diffs and I won't have spare time to cover all cases with tests so closing for now

@EgorBoEgorBo closed this May 18, 2023
@ghostghost locked as resolved and limited conversation to collaborators Jun 17, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@EgorBo@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Constant folding for SIMD comparisons - #85584

Closed
EgorBo wants to merge 9 commits into
dotnet:mainfrom
EgorBo:simd-constant-fold-eq-ne
Closed

Constant folding for SIMD comparisons#85584
EgorBo wants to merge 9 commits into
dotnet:mainfrom
EgorBo:simd-constant-fold-eq-ne

Conversation

@EgorBo

@EgorBoEgorBo commented May 1, 2023

Copy link
Copy Markdown
Member

This PR extends constant folding for SIMD to support EQ/NE. Can also be extended for GT_GT/GE/LT/LE and, perhaps, floating point, but I was mainly interested in EQ/NE for integral types as I found a few real uses cases. Also, improved constant folding for RVA slightly.

boolTest(){returnIsHelloWorldString("Hello World SIMD constant folding!"u8);}boolIsHelloWorldString(ReadOnlySpan<byte>data){Debug.Assert(data.Lengthis>=16 and <=32);// Load data into two Vector128, the 2nd vector may overlap if neededrefbyteptr=refMemoryMarshal.GetReference(data);Vector128<byte>v1=Vector128.LoadUnsafe(refptr);Vector128<byte>v2=Vector128.LoadUnsafe(refptr,(nuint)(data.Length-Vector128<byte>.Count));// The constant utf8 string we compare against:varcns="Hello World SIMD constant folding!"u8;Vector128<byte>v1Cns=Vector128.Create(cns.Slice(0,Vector128<byte>.Count));Vector128<byte>v2Cns=Vector128.Create(cns.Slice(cns.Length-Vector128<byte>.Count));// Compare pairwisereturn((v1^v1Cns)|(v2^v2Cns))==Vector128<byte>.Zero;}

Codegen for Test():

; Method Test():bool:thismoveax,1ret; Total bytes of code: 6

the whole thing is collapsed into just return true after inlining and constant folding.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 1, 2023
@ghostghost assigned EgorBoMay 1, 2023
@ghost

ghost commented May 1, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

Can be easily extended for GT_GT/GE/LT/LE and, perhaps, floating point, but I was mainly interested in EQ/NE for integeres as I found a few real uses cases. E.g.

boolTest(){vardata="hello world simd folding demo :)"u8;returnIsHelloWorldString(refMemoryMarshal.GetReference(data));}[MethodImpl(MethodImplOptions.AggressiveInlining)]boolIsHelloWorldString(refbytedata){varv1=Vector128.LoadUnsafe(refdata);varv2=Vector128.LoadUnsafe(refdata,16);varv1Cns=Vector128.Create("hello world simd"u8);varv2Cns=Vector128.Create(" folding demo :)"u8);return((v1^v1Cns)|(v2^v2Cns))==Vector128<byte>.Zero;}

Codegen for Test():

; Method MyClass:Foo():bool:thismoveax,1ret; Total bytes of code: 6

the whole thing is collapsed into just return true after inlining and constant folding.

Author:EgorBo
Assignees:EgorBo
Labels:

area-CodeGen-coreclr

Milestone:-

Comment on lines +443 to +448
template <typename TBase>
TBase GetAllBitsSetScalar()
{
uint8_t bitWidth = (sizeof(TBase) * 8);
return static_cast<TBase>((1ULL << bitWidth) - 1);
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Won't this fail for int64_t and uint64_t?

Why not simply return ~static_cast<TBase>(0);? Then for float/double you could specialize if needed.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, addressed

Comment on lines +530 to +538
case GT_EQ:
{
return arg0 == arg1 ? GetAllBitsSetScalar<TBase>() : 0;
}

case GT_NE:
{
return arg0 != arg1 ? GetAllBitsSetScalar<TBase>() : 0;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is going to be incorrect for float/double since it will do a bitwise comparison.

You may need to special-case in EvaluateBinaryScalarSpecialized<float> before it calls this main method.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You might also be able to do it in EvaluateBinaryScalar instead but special-case GetAllBitsSetScalar<float>(), which is probably simpler overall.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This path is never taken for floating point, I added asserts just in case. I was mostly interested in non-fp cases and these can be added if needed

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Any reason we aren’t covering floating-point as well? I imagine this might get hits in ImageSharp or PDN.

Potentially in other apps, games, or graphics processing, etc

assert((vn1Type == vn2Type) && varTypeIsSIMD(vn1Type));
assert(!varTypeIsFloating(baseType));

ValueNum packed = EvaluateBinarySimd(vns, GT_EQ, scalar, vn1Type, baseType, arg0VN, arg1VN);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

shouldn't this be EvaluateBinarySimd(vns, oper, ...)?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I wonder if it would be simpler to implement this as EvaluateVector and then simplify check result IsAllBitsSet or !IsZero

Which would also simplify the other relational comparisons

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

shouldn't this be EvaluateBinarySimd(vns, oper, ...)?

It makes it harder to check output, doesn't it?
Currently I just pass EQ and then return AllBitsSet or !AllBitsSet depending on source oper

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does it? We have a simple property for any simd_T and it makes it easier to cover the other cases.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That is, if you just evaluate per element you get a result simd_T result and can then just check IsAllBitsSet or IsZero

return EvaluateBinarySimd(this, GT_NE, /* scalar */ false, type, baseType, arg0VN, arg1VN);
}
break;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we want to handle GT, LT, GE, and LE simultaneously?

Likewise given you've added the support for computing the vector version in order to make bool work, should we just handle the intrinsics that produce a vector as well?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Those had no hits so I just didn't want to add more code and tests 🙂 Although, even EQ/NE don't have hits, I just found a use case outside.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is the type of scenario where it’s a core comparison on SIMD and so we want to generally handle it even if our own code doesn’t have any hits today.

it’s odd to cover some of the relational comparisons and not the others

@EgorBo

EgorBo commented May 18, 2023

Copy link
Copy Markdown
MemberAuthor

No diffs and I won't have spare time to cover all cases with tests so closing for now

@EgorBoEgorBo closed this May 18, 2023
@ghostghost locked as resolved and limited conversation to collaborators Jun 17, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@EgorBo@tannergooding