Switch to iSimdVector and Align WidenAsciiToUtf16 - #99982

Merged
tannergooding merged 5 commits into
dotnet:mainfrom
DeepakRajendrakumaran:align
Apr 8, 2024
Merged

Switch to iSimdVector and Align WidenAsciiToUtf16#99982
tannergooding merged 5 commits into
dotnet:mainfrom
DeepakRajendrakumaran:align

Conversation

@DeepakRajendrakumaran

@DeepakRajendrakumaranDeepakRajendrakumaran commented Mar 19, 2024

Copy link
Copy Markdown
Contributor

This is an updated version of this PR(#89892). It does the following

  1. Add ' AnyMatches' support for iSimdVector
  2. Use iSimdVector to clean up 'WidenAsciiToUtf16' implementation
  3. Align memory stores

Perf Results

Ran the following tests(sizes : 16, 512, 1024, 5120, 10240) on EMR: (base = main branch, diff = with change)https://github.com/dotnet/performance/blob/47d21ee9571164a8e3f8088d8709ca4061d96827/src/benchmarks/micro/libraries/System.Text.Encoding/Perf.Encoding.cs

On EMR
image

On ICX - Not much diff
image

@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Mar 19, 2024
@DeepakRajendrakumaran
DeepakRajendrakumaran marked this pull request as ready for review March 20, 2024 22:25
@DeepakRajendrakumaran

DeepakRajendrakumaran commented Mar 20, 2024

Copy link
Copy Markdown
ContributorAuthor

@tannergooding @dotnet/avx512-contrib Can you please review this?

private static unsafe bool HasMatch<TVectorByte>(TVectorByte vector)
where TVectorByte : unmanaged, ISimdVector<TVectorByte, byte>
{
return !(vector & TVectorByte.Create((byte)0x80)).Equals(TVectorByte.Zero);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why not (vector & TVectorByte.Create((byte)0x80)) != TVectorByte.Zero?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I basically ran into a weird issue where perf degraded significantly for specific cases on my ICX when I had '(vector & TVectorByte.Create((byte)0x80)) != TVectorByte.Zero' and the issue went away with what I have now. It was pretty consistent. I had some trouble narrowing down the exact why with VTune though.

I decided to go with the performant version for now

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Was there a codegen difference between them? I'd expect them to generate the same code

@DeepakRajendrakumaranDeepakRajendrakumaranMar 27, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I did a quick check on godbolt and they all look the same - https://godbolt.org/z/P87dPdTGa

Now I'm curious if it was just something off when I ran it locally

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I did some more digging and something is off. Tried 3 versions with a handwritten benchmark
image
Benchmark
image

Match3 is significantly faster than Match1 and Match 2

VTune for Match1 vs Match3
image

The inlining makes it hard to narrow down

image

@DeepakRajendrakumaranDeepakRajendrakumaranMar 28, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do you have a strong preference for any of these patterns?

I can look into the why this is happening if it's important. For now, I'm just keeping the fast version

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@tannergooding I have created an issue detailing this as discussed: #100493

Please let me know if there is anything else needed for this PR

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment on lines +2164 to +2196
if (!HasMatch<TVectorByte>(asciiVector))
{
(TVectorUShort utf16LowVector, TVectorUShort utf16HighVector) = Widen<TVectorByte, TVectorUShort>(asciiVector);
utf16LowVector.Store(pCurrentWriteAddress);
utf16HighVector.Store(pCurrentWriteAddress + TVectorUShort.Count);
pCurrentWriteAddress += (nuint)(TVectorUShort.Count * 2);
if (((int)pCurrentWriteAddress & 1) == 0)
{
// Bump write buffer up to the next aligned boundary
pCurrentWriteAddress = (ushort*)((nuint)pCurrentWriteAddress & ~(nuint)(TVectorUShort.Alignment - 1));
nuint numBytesWritten = (nuint)pCurrentWriteAddress - (nuint)pUtf16Buffer;
currentOffset += (nuint)numBytesWritten / 2;
}
else
{
// If input isn't char aligned, we won't be able to align it to a Vector
currentOffset += (nuint)TVectorByte.Count;
}
while (currentOffset <= finalOffsetWhereCanRunLoop)
{
asciiVector = TVectorByte.Load(pAsciiBuffer + currentOffset);
if (HasMatch<TVectorByte>(asciiVector))
{
break;
}
(utf16LowVector, utf16HighVector) = Widen<TVectorByte, TVectorUShort>(asciiVector);
utf16LowVector.StoreAligned(pCurrentWriteAddress);
utf16HighVector.StoreAligned(pCurrentWriteAddress + TVectorUShort.Count);
currentOffset += (nuint)TVectorByte.Count;
pCurrentWriteAddress += (nuint)(TVectorUShort.Count * 2);
}
}

@tannergoodingtannergoodingMar 27, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The code here looks generally good, and I don't expect anything to be changed for this PR based on what I'm about to comment.

However, I would like to refer to how we set up everything for TensorPrimitives as it works very well and allows a lot of code sharing (noting it isn't using ISimdVector yet since it's out of band, but could easily do so in the future): https://source.dot.net/#System.Numerics.Tensors/System/Numerics/Tensors/netcore/Common/TensorPrimitives.IUnaryOperator.cs,fdab74764af40a1e

In general it tries to do the pre-checks up front and only ever execute 1 vector path (so it doesn't have to fallthrough from Vector512->Vector256->Vector128 as remainders exist). To achieve this, it has the Vectorized### helpers (which are all identical, except for the size they operate on, this is what would eventually use ISimdVector) and then a shared VectorizedSmall which is simply a jump table designed to handle any data that is less than a full vector using a single branch.

The core logic for the vectorized algorithm (https://source.dot.net/#System.Numerics.Tensors/System/Numerics/Tensors/netcore/Common/TensorPrimitives.IUnaryOperator.cs,ef9adce4e9561b04) then basically has a path to handle the main loop (which is currently unrolled by a factor of 8) and otherwise hits a jump table to handle remaining blocks (so they can likewise be handled with a single branch).

To help optimize, it preloads the beginning and ending vectors. In the worst case this will result in double processing of some inputs for very small sizes, but its ultimately only 2 main operations which is fine.

The main loop then attempts to align and has an optimized path for extremely large inputs for non-temporal data if alignment could be achieved. Smaller inputs just do regular unaligned stores since the actual address will have been aligned if that was feasible.

This general approach is done because it allows all paths, but particularly the smallest inputs, to minimize the total number of branches done (no more than 2 branches for non-vectorized data and no more 3 to hit the main loop for vectorized code). It also allows us to separate the "algorithm logic" from the "vectorization logic" and share that vectorization logic between multiple vectorized algorithms.

This general setup has worked so well and provided very stable perf numbers for all sizes, such that we opened #93217 as a means of investigating if we could make it more general purpose and public. Long term, it'd probably be desirable to move algorithms like this WidenAsciiToUtf16 to follow the same approach so that we get the best perf, with the least overhead.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is really useful. And something I can add on in future. Thanks for the detailed comment

// New Surface Area
//

static bool ISimdVector<Vector128<T>, T>.AnyMatches(Vector128<T> vector)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We did review/approve this (#98055 (comment)) and settled on bool Any(Vector128<T> vector, T value) and bool AnyWhereAllBitsSet(Vector<T> vector)

It would be nice to fix this to follow that. The PR otherwise looks good and should be mergeable.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So, the way I understand is we do not need AnyMatches() anymore

And Any and AnyWhereAllBitsSet would look something like follows


static bool ISimdVector<Vector512<T>, T>.AnyWhereAllBitsSet(Vector512<T> vector)
{
return (vector.EqualsAny(Vector512<T>.AllBitsSet));
}
static bool ISimdVector<Vector512<T>, T>.Any(Vector512<T> vector, T value)
{
return (vector.EqualsAny(Vector512.Create((T)value)));
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, essentially, with the ability for the JIT to recognize these as intrinsic and optimize them more in appropriate scenarios (but that's not necessary for this PR)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done - took out AnyMatches() and added AnyWhereAllBitsSet() and Any()

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@tannergooding Does this look correct? Anything else you'd like me to fix?


static bool ISimdVector<Vector128<T>, T>.AnyWhereAllBitsSet(Vector128<T> vector)
{
return (Vector128.EqualsAny(vector, Vector128<T>.AllBitsSet));

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: unnecessary parens here and in Any

If you could fix that in a follow up PR, that'd be great (going to merge this)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good! Will put it up later today.

@DrewScoggins

DrewScoggins commented Apr 11, 2024

Copy link
Copy Markdown
Member

@matouskozak

matouskozak commented Apr 16, 2024

Copy link
Copy Markdown
Member

@lewing

Copy link
Copy Markdown
Member

mono wasm aot regression dotnet/perf-autofiling-issues#32601

matouskozak pushed a commit to matouskozak/runtime that referenced this pull request Apr 30, 2024
* Add AnyMatches() to iSimdVector interface
* Switch to iSimdVector and Align WidenAsciiToUtf16.
* Fixing perf
* Addressing Review Comments.
* Mirroring API change : dotnet#98055 (comment)
@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 21, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-System.Runtime.Intrinsicscommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@DeepakRajendrakumaran@DrewScoggins@matouskozak@lewing@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Switch to iSimdVector and Align WidenAsciiToUtf16 - #99982

Merged
tannergooding merged 5 commits into
dotnet:mainfrom
DeepakRajendrakumaran:align
Apr 8, 2024
Merged

Switch to iSimdVector and Align WidenAsciiToUtf16#99982
tannergooding merged 5 commits into
dotnet:mainfrom
DeepakRajendrakumaran:align

Conversation

@DeepakRajendrakumaran

@DeepakRajendrakumaranDeepakRajendrakumaran commented Mar 19, 2024

Copy link
Copy Markdown
Contributor

This is an updated version of this PR(#89892). It does the following

  1. Add ' AnyMatches' support for iSimdVector
  2. Use iSimdVector to clean up 'WidenAsciiToUtf16' implementation
  3. Align memory stores

Perf Results

Ran the following tests(sizes : 16, 512, 1024, 5120, 10240) on EMR: (base = main branch, diff = with change)https://github.com/dotnet/performance/blob/47d21ee9571164a8e3f8088d8709ca4061d96827/src/benchmarks/micro/libraries/System.Text.Encoding/Perf.Encoding.cs

On EMR
image

On ICX - Not much diff
image

@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Mar 19, 2024
@DeepakRajendrakumaran
DeepakRajendrakumaran marked this pull request as ready for review March 20, 2024 22:25
@DeepakRajendrakumaran

DeepakRajendrakumaran commented Mar 20, 2024

Copy link
Copy Markdown
ContributorAuthor

@tannergooding @dotnet/avx512-contrib Can you please review this?

private static unsafe bool HasMatch<TVectorByte>(TVectorByte vector)
where TVectorByte : unmanaged, ISimdVector<TVectorByte, byte>
{
return !(vector & TVectorByte.Create((byte)0x80)).Equals(TVectorByte.Zero);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why not (vector & TVectorByte.Create((byte)0x80)) != TVectorByte.Zero?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I basically ran into a weird issue where perf degraded significantly for specific cases on my ICX when I had '(vector & TVectorByte.Create((byte)0x80)) != TVectorByte.Zero' and the issue went away with what I have now. It was pretty consistent. I had some trouble narrowing down the exact why with VTune though.

I decided to go with the performant version for now

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Was there a codegen difference between them? I'd expect them to generate the same code

@DeepakRajendrakumaranDeepakRajendrakumaranMar 27, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I did a quick check on godbolt and they all look the same - https://godbolt.org/z/P87dPdTGa

Now I'm curious if it was just something off when I ran it locally

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I did some more digging and something is off. Tried 3 versions with a handwritten benchmark
image
Benchmark
image

Match3 is significantly faster than Match1 and Match 2

VTune for Match1 vs Match3
image

The inlining makes it hard to narrow down

image

@DeepakRajendrakumaranDeepakRajendrakumaranMar 28, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do you have a strong preference for any of these patterns?

I can look into the why this is happening if it's important. For now, I'm just keeping the fast version

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@tannergooding I have created an issue detailing this as discussed: #100493

Please let me know if there is anything else needed for this PR

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment on lines +2164 to +2196
if (!HasMatch<TVectorByte>(asciiVector))
{
(TVectorUShort utf16LowVector, TVectorUShort utf16HighVector) = Widen<TVectorByte, TVectorUShort>(asciiVector);
utf16LowVector.Store(pCurrentWriteAddress);
utf16HighVector.Store(pCurrentWriteAddress + TVectorUShort.Count);
pCurrentWriteAddress += (nuint)(TVectorUShort.Count * 2);
if (((int)pCurrentWriteAddress & 1) == 0)
{
// Bump write buffer up to the next aligned boundary
pCurrentWriteAddress = (ushort*)((nuint)pCurrentWriteAddress & ~(nuint)(TVectorUShort.Alignment - 1));
nuint numBytesWritten = (nuint)pCurrentWriteAddress - (nuint)pUtf16Buffer;
currentOffset += (nuint)numBytesWritten / 2;
}
else
{
// If input isn't char aligned, we won't be able to align it to a Vector
currentOffset += (nuint)TVectorByte.Count;
}
while (currentOffset <= finalOffsetWhereCanRunLoop)
{
asciiVector = TVectorByte.Load(pAsciiBuffer + currentOffset);
if (HasMatch<TVectorByte>(asciiVector))
{
break;
}
(utf16LowVector, utf16HighVector) = Widen<TVectorByte, TVectorUShort>(asciiVector);
utf16LowVector.StoreAligned(pCurrentWriteAddress);
utf16HighVector.StoreAligned(pCurrentWriteAddress + TVectorUShort.Count);
currentOffset += (nuint)TVectorByte.Count;
pCurrentWriteAddress += (nuint)(TVectorUShort.Count * 2);
}
}

@tannergoodingtannergoodingMar 27, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The code here looks generally good, and I don't expect anything to be changed for this PR based on what I'm about to comment.

However, I would like to refer to how we set up everything for TensorPrimitives as it works very well and allows a lot of code sharing (noting it isn't using ISimdVector yet since it's out of band, but could easily do so in the future): https://source.dot.net/#System.Numerics.Tensors/System/Numerics/Tensors/netcore/Common/TensorPrimitives.IUnaryOperator.cs,fdab74764af40a1e

In general it tries to do the pre-checks up front and only ever execute 1 vector path (so it doesn't have to fallthrough from Vector512->Vector256->Vector128 as remainders exist). To achieve this, it has the Vectorized### helpers (which are all identical, except for the size they operate on, this is what would eventually use ISimdVector) and then a shared VectorizedSmall which is simply a jump table designed to handle any data that is less than a full vector using a single branch.

The core logic for the vectorized algorithm (https://source.dot.net/#System.Numerics.Tensors/System/Numerics/Tensors/netcore/Common/TensorPrimitives.IUnaryOperator.cs,ef9adce4e9561b04) then basically has a path to handle the main loop (which is currently unrolled by a factor of 8) and otherwise hits a jump table to handle remaining blocks (so they can likewise be handled with a single branch).

To help optimize, it preloads the beginning and ending vectors. In the worst case this will result in double processing of some inputs for very small sizes, but its ultimately only 2 main operations which is fine.

The main loop then attempts to align and has an optimized path for extremely large inputs for non-temporal data if alignment could be achieved. Smaller inputs just do regular unaligned stores since the actual address will have been aligned if that was feasible.

This general approach is done because it allows all paths, but particularly the smallest inputs, to minimize the total number of branches done (no more than 2 branches for non-vectorized data and no more 3 to hit the main loop for vectorized code). It also allows us to separate the "algorithm logic" from the "vectorization logic" and share that vectorization logic between multiple vectorized algorithms.

This general setup has worked so well and provided very stable perf numbers for all sizes, such that we opened #93217 as a means of investigating if we could make it more general purpose and public. Long term, it'd probably be desirable to move algorithms like this WidenAsciiToUtf16 to follow the same approach so that we get the best perf, with the least overhead.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is really useful. And something I can add on in future. Thanks for the detailed comment

// New Surface Area
//

static bool ISimdVector<Vector128<T>, T>.AnyMatches(Vector128<T> vector)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We did review/approve this (#98055 (comment)) and settled on bool Any(Vector128<T> vector, T value) and bool AnyWhereAllBitsSet(Vector<T> vector)

It would be nice to fix this to follow that. The PR otherwise looks good and should be mergeable.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So, the way I understand is we do not need AnyMatches() anymore

And Any and AnyWhereAllBitsSet would look something like follows


static bool ISimdVector<Vector512<T>, T>.AnyWhereAllBitsSet(Vector512<T> vector)
{
return (vector.EqualsAny(Vector512<T>.AllBitsSet));
}
static bool ISimdVector<Vector512<T>, T>.Any(Vector512<T> vector, T value)
{
return (vector.EqualsAny(Vector512.Create((T)value)));
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, essentially, with the ability for the JIT to recognize these as intrinsic and optimize them more in appropriate scenarios (but that's not necessary for this PR)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done - took out AnyMatches() and added AnyWhereAllBitsSet() and Any()

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@tannergooding Does this look correct? Anything else you'd like me to fix?


static bool ISimdVector<Vector128<T>, T>.AnyWhereAllBitsSet(Vector128<T> vector)
{
return (Vector128.EqualsAny(vector, Vector128<T>.AllBitsSet));

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: unnecessary parens here and in Any

If you could fix that in a follow up PR, that'd be great (going to merge this)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good! Will put it up later today.

@DrewScoggins

DrewScoggins commented Apr 11, 2024

Copy link
Copy Markdown
Member

@matouskozak

matouskozak commented Apr 16, 2024

Copy link
Copy Markdown
Member

@lewing

Copy link
Copy Markdown
Member

mono wasm aot regression dotnet/perf-autofiling-issues#32601

matouskozak pushed a commit to matouskozak/runtime that referenced this pull request Apr 30, 2024
* Add AnyMatches() to iSimdVector interface
* Switch to iSimdVector and Align WidenAsciiToUtf16.
* Fixing perf
* Addressing Review Comments.
* Mirroring API change : dotnet#98055 (comment)
@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 21, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-System.Runtime.Intrinsicscommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@DeepakRajendrakumaran@DrewScoggins@matouskozak@lewing@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Switch to iSimdVector and Align WidenAsciiToUtf16 - #99982

Merged
tannergooding merged 5 commits into
dotnet:mainfrom
DeepakRajendrakumaran:align
Apr 8, 2024
Merged

Switch to iSimdVector and Align WidenAsciiToUtf16#99982
tannergooding merged 5 commits into
dotnet:mainfrom
DeepakRajendrakumaran:align

Conversation

@DeepakRajendrakumaran

@DeepakRajendrakumaranDeepakRajendrakumaran commented Mar 19, 2024

Copy link
Copy Markdown
Contributor

This is an updated version of this PR(#89892). It does the following

  1. Add ' AnyMatches' support for iSimdVector
  2. Use iSimdVector to clean up 'WidenAsciiToUtf16' implementation
  3. Align memory stores

Perf Results

Ran the following tests(sizes : 16, 512, 1024, 5120, 10240) on EMR: (base = main branch, diff = with change)https://github.com/dotnet/performance/blob/47d21ee9571164a8e3f8088d8709ca4061d96827/src/benchmarks/micro/libraries/System.Text.Encoding/Perf.Encoding.cs

On EMR
image

On ICX - Not much diff
image

@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Mar 19, 2024
@DeepakRajendrakumaran
DeepakRajendrakumaran marked this pull request as ready for review March 20, 2024 22:25
@DeepakRajendrakumaran

DeepakRajendrakumaran commented Mar 20, 2024

Copy link
Copy Markdown
ContributorAuthor

@tannergooding @dotnet/avx512-contrib Can you please review this?

private static unsafe bool HasMatch<TVectorByte>(TVectorByte vector)
where TVectorByte : unmanaged, ISimdVector<TVectorByte, byte>
{
return !(vector & TVectorByte.Create((byte)0x80)).Equals(TVectorByte.Zero);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why not (vector & TVectorByte.Create((byte)0x80)) != TVectorByte.Zero?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I basically ran into a weird issue where perf degraded significantly for specific cases on my ICX when I had '(vector & TVectorByte.Create((byte)0x80)) != TVectorByte.Zero' and the issue went away with what I have now. It was pretty consistent. I had some trouble narrowing down the exact why with VTune though.

I decided to go with the performant version for now

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Was there a codegen difference between them? I'd expect them to generate the same code

@DeepakRajendrakumaranDeepakRajendrakumaranMar 27, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I did a quick check on godbolt and they all look the same - https://godbolt.org/z/P87dPdTGa

Now I'm curious if it was just something off when I ran it locally

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I did some more digging and something is off. Tried 3 versions with a handwritten benchmark
image
Benchmark
image

Match3 is significantly faster than Match1 and Match 2

VTune for Match1 vs Match3
image

The inlining makes it hard to narrow down

image

@DeepakRajendrakumaranDeepakRajendrakumaranMar 28, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do you have a strong preference for any of these patterns?

I can look into the why this is happening if it's important. For now, I'm just keeping the fast version

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@tannergooding I have created an issue detailing this as discussed: #100493

Please let me know if there is anything else needed for this PR

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment on lines +2164 to +2196
if (!HasMatch<TVectorByte>(asciiVector))
{
(TVectorUShort utf16LowVector, TVectorUShort utf16HighVector) = Widen<TVectorByte, TVectorUShort>(asciiVector);
utf16LowVector.Store(pCurrentWriteAddress);
utf16HighVector.Store(pCurrentWriteAddress + TVectorUShort.Count);
pCurrentWriteAddress += (nuint)(TVectorUShort.Count * 2);
if (((int)pCurrentWriteAddress & 1) == 0)
{
// Bump write buffer up to the next aligned boundary
pCurrentWriteAddress = (ushort*)((nuint)pCurrentWriteAddress & ~(nuint)(TVectorUShort.Alignment - 1));
nuint numBytesWritten = (nuint)pCurrentWriteAddress - (nuint)pUtf16Buffer;
currentOffset += (nuint)numBytesWritten / 2;
}
else
{
// If input isn't char aligned, we won't be able to align it to a Vector
currentOffset += (nuint)TVectorByte.Count;
}
while (currentOffset <= finalOffsetWhereCanRunLoop)
{
asciiVector = TVectorByte.Load(pAsciiBuffer + currentOffset);
if (HasMatch<TVectorByte>(asciiVector))
{
break;
}
(utf16LowVector, utf16HighVector) = Widen<TVectorByte, TVectorUShort>(asciiVector);
utf16LowVector.StoreAligned(pCurrentWriteAddress);
utf16HighVector.StoreAligned(pCurrentWriteAddress + TVectorUShort.Count);
currentOffset += (nuint)TVectorByte.Count;
pCurrentWriteAddress += (nuint)(TVectorUShort.Count * 2);
}
}

@tannergoodingtannergoodingMar 27, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The code here looks generally good, and I don't expect anything to be changed for this PR based on what I'm about to comment.

However, I would like to refer to how we set up everything for TensorPrimitives as it works very well and allows a lot of code sharing (noting it isn't using ISimdVector yet since it's out of band, but could easily do so in the future): https://source.dot.net/#System.Numerics.Tensors/System/Numerics/Tensors/netcore/Common/TensorPrimitives.IUnaryOperator.cs,fdab74764af40a1e

In general it tries to do the pre-checks up front and only ever execute 1 vector path (so it doesn't have to fallthrough from Vector512->Vector256->Vector128 as remainders exist). To achieve this, it has the Vectorized### helpers (which are all identical, except for the size they operate on, this is what would eventually use ISimdVector) and then a shared VectorizedSmall which is simply a jump table designed to handle any data that is less than a full vector using a single branch.

The core logic for the vectorized algorithm (https://source.dot.net/#System.Numerics.Tensors/System/Numerics/Tensors/netcore/Common/TensorPrimitives.IUnaryOperator.cs,ef9adce4e9561b04) then basically has a path to handle the main loop (which is currently unrolled by a factor of 8) and otherwise hits a jump table to handle remaining blocks (so they can likewise be handled with a single branch).

To help optimize, it preloads the beginning and ending vectors. In the worst case this will result in double processing of some inputs for very small sizes, but its ultimately only 2 main operations which is fine.

The main loop then attempts to align and has an optimized path for extremely large inputs for non-temporal data if alignment could be achieved. Smaller inputs just do regular unaligned stores since the actual address will have been aligned if that was feasible.

This general approach is done because it allows all paths, but particularly the smallest inputs, to minimize the total number of branches done (no more than 2 branches for non-vectorized data and no more 3 to hit the main loop for vectorized code). It also allows us to separate the "algorithm logic" from the "vectorization logic" and share that vectorization logic between multiple vectorized algorithms.

This general setup has worked so well and provided very stable perf numbers for all sizes, such that we opened #93217 as a means of investigating if we could make it more general purpose and public. Long term, it'd probably be desirable to move algorithms like this WidenAsciiToUtf16 to follow the same approach so that we get the best perf, with the least overhead.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is really useful. And something I can add on in future. Thanks for the detailed comment

// New Surface Area
//

static bool ISimdVector<Vector128<T>, T>.AnyMatches(Vector128<T> vector)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We did review/approve this (#98055 (comment)) and settled on bool Any(Vector128<T> vector, T value) and bool AnyWhereAllBitsSet(Vector<T> vector)

It would be nice to fix this to follow that. The PR otherwise looks good and should be mergeable.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So, the way I understand is we do not need AnyMatches() anymore

And Any and AnyWhereAllBitsSet would look something like follows


static bool ISimdVector<Vector512<T>, T>.AnyWhereAllBitsSet(Vector512<T> vector)
{
return (vector.EqualsAny(Vector512<T>.AllBitsSet));
}
static bool ISimdVector<Vector512<T>, T>.Any(Vector512<T> vector, T value)
{
return (vector.EqualsAny(Vector512.Create((T)value)));
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, essentially, with the ability for the JIT to recognize these as intrinsic and optimize them more in appropriate scenarios (but that's not necessary for this PR)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done - took out AnyMatches() and added AnyWhereAllBitsSet() and Any()

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@tannergooding Does this look correct? Anything else you'd like me to fix?


static bool ISimdVector<Vector128<T>, T>.AnyWhereAllBitsSet(Vector128<T> vector)
{
return (Vector128.EqualsAny(vector, Vector128<T>.AllBitsSet));

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: unnecessary parens here and in Any

If you could fix that in a follow up PR, that'd be great (going to merge this)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good! Will put it up later today.

@DrewScoggins

DrewScoggins commented Apr 11, 2024

Copy link
Copy Markdown
Member

@matouskozak

matouskozak commented Apr 16, 2024

Copy link
Copy Markdown
Member

@lewing

Copy link
Copy Markdown
Member

mono wasm aot regression dotnet/perf-autofiling-issues#32601

matouskozak pushed a commit to matouskozak/runtime that referenced this pull request Apr 30, 2024
* Add AnyMatches() to iSimdVector interface
* Switch to iSimdVector and Align WidenAsciiToUtf16.
* Fixing perf
* Addressing Review Comments.
* Mirroring API change : dotnet#98055 (comment)
@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 21, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-System.Runtime.Intrinsicscommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@DeepakRajendrakumaran@DrewScoggins@matouskozak@lewing@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Switch to iSimdVector and Align WidenAsciiToUtf16 - #99982

Merged
tannergooding merged 5 commits into
dotnet:mainfrom
DeepakRajendrakumaran:align
Apr 8, 2024
Merged

Switch to iSimdVector and Align WidenAsciiToUtf16#99982
tannergooding merged 5 commits into
dotnet:mainfrom
DeepakRajendrakumaran:align

Conversation

@DeepakRajendrakumaran

@DeepakRajendrakumaranDeepakRajendrakumaran commented Mar 19, 2024

Copy link
Copy Markdown
Contributor

This is an updated version of this PR(#89892). It does the following

  1. Add ' AnyMatches' support for iSimdVector
  2. Use iSimdVector to clean up 'WidenAsciiToUtf16' implementation
  3. Align memory stores

Perf Results

Ran the following tests(sizes : 16, 512, 1024, 5120, 10240) on EMR: (base = main branch, diff = with change)https://github.com/dotnet/performance/blob/47d21ee9571164a8e3f8088d8709ca4061d96827/src/benchmarks/micro/libraries/System.Text.Encoding/Perf.Encoding.cs

On EMR
image

On ICX - Not much diff
image

@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Mar 19, 2024
@DeepakRajendrakumaran
DeepakRajendrakumaran marked this pull request as ready for review March 20, 2024 22:25
@DeepakRajendrakumaran

DeepakRajendrakumaran commented Mar 20, 2024

Copy link
Copy Markdown
ContributorAuthor

@tannergooding @dotnet/avx512-contrib Can you please review this?

private static unsafe bool HasMatch<TVectorByte>(TVectorByte vector)
where TVectorByte : unmanaged, ISimdVector<TVectorByte, byte>
{
return !(vector & TVectorByte.Create((byte)0x80)).Equals(TVectorByte.Zero);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why not (vector & TVectorByte.Create((byte)0x80)) != TVectorByte.Zero?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I basically ran into a weird issue where perf degraded significantly for specific cases on my ICX when I had '(vector & TVectorByte.Create((byte)0x80)) != TVectorByte.Zero' and the issue went away with what I have now. It was pretty consistent. I had some trouble narrowing down the exact why with VTune though.

I decided to go with the performant version for now

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Was there a codegen difference between them? I'd expect them to generate the same code

@DeepakRajendrakumaranDeepakRajendrakumaranMar 27, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I did a quick check on godbolt and they all look the same - https://godbolt.org/z/P87dPdTGa

Now I'm curious if it was just something off when I ran it locally

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I did some more digging and something is off. Tried 3 versions with a handwritten benchmark
image
Benchmark
image

Match3 is significantly faster than Match1 and Match 2

VTune for Match1 vs Match3
image

The inlining makes it hard to narrow down

image

@DeepakRajendrakumaranDeepakRajendrakumaranMar 28, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do you have a strong preference for any of these patterns?

I can look into the why this is happening if it's important. For now, I'm just keeping the fast version

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@tannergooding I have created an issue detailing this as discussed: #100493

Please let me know if there is anything else needed for this PR

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment on lines +2164 to +2196
if (!HasMatch<TVectorByte>(asciiVector))
{
(TVectorUShort utf16LowVector, TVectorUShort utf16HighVector) = Widen<TVectorByte, TVectorUShort>(asciiVector);
utf16LowVector.Store(pCurrentWriteAddress);
utf16HighVector.Store(pCurrentWriteAddress + TVectorUShort.Count);
pCurrentWriteAddress += (nuint)(TVectorUShort.Count * 2);
if (((int)pCurrentWriteAddress & 1) == 0)
{
// Bump write buffer up to the next aligned boundary
pCurrentWriteAddress = (ushort*)((nuint)pCurrentWriteAddress & ~(nuint)(TVectorUShort.Alignment - 1));
nuint numBytesWritten = (nuint)pCurrentWriteAddress - (nuint)pUtf16Buffer;
currentOffset += (nuint)numBytesWritten / 2;
}
else
{
// If input isn't char aligned, we won't be able to align it to a Vector
currentOffset += (nuint)TVectorByte.Count;
}
while (currentOffset <= finalOffsetWhereCanRunLoop)
{
asciiVector = TVectorByte.Load(pAsciiBuffer + currentOffset);
if (HasMatch<TVectorByte>(asciiVector))
{
break;
}
(utf16LowVector, utf16HighVector) = Widen<TVectorByte, TVectorUShort>(asciiVector);
utf16LowVector.StoreAligned(pCurrentWriteAddress);
utf16HighVector.StoreAligned(pCurrentWriteAddress + TVectorUShort.Count);
currentOffset += (nuint)TVectorByte.Count;
pCurrentWriteAddress += (nuint)(TVectorUShort.Count * 2);
}
}

@tannergoodingtannergoodingMar 27, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The code here looks generally good, and I don't expect anything to be changed for this PR based on what I'm about to comment.

However, I would like to refer to how we set up everything for TensorPrimitives as it works very well and allows a lot of code sharing (noting it isn't using ISimdVector yet since it's out of band, but could easily do so in the future): https://source.dot.net/#System.Numerics.Tensors/System/Numerics/Tensors/netcore/Common/TensorPrimitives.IUnaryOperator.cs,fdab74764af40a1e

In general it tries to do the pre-checks up front and only ever execute 1 vector path (so it doesn't have to fallthrough from Vector512->Vector256->Vector128 as remainders exist). To achieve this, it has the Vectorized### helpers (which are all identical, except for the size they operate on, this is what would eventually use ISimdVector) and then a shared VectorizedSmall which is simply a jump table designed to handle any data that is less than a full vector using a single branch.

The core logic for the vectorized algorithm (https://source.dot.net/#System.Numerics.Tensors/System/Numerics/Tensors/netcore/Common/TensorPrimitives.IUnaryOperator.cs,ef9adce4e9561b04) then basically has a path to handle the main loop (which is currently unrolled by a factor of 8) and otherwise hits a jump table to handle remaining blocks (so they can likewise be handled with a single branch).

To help optimize, it preloads the beginning and ending vectors. In the worst case this will result in double processing of some inputs for very small sizes, but its ultimately only 2 main operations which is fine.

The main loop then attempts to align and has an optimized path for extremely large inputs for non-temporal data if alignment could be achieved. Smaller inputs just do regular unaligned stores since the actual address will have been aligned if that was feasible.

This general approach is done because it allows all paths, but particularly the smallest inputs, to minimize the total number of branches done (no more than 2 branches for non-vectorized data and no more 3 to hit the main loop for vectorized code). It also allows us to separate the "algorithm logic" from the "vectorization logic" and share that vectorization logic between multiple vectorized algorithms.

This general setup has worked so well and provided very stable perf numbers for all sizes, such that we opened #93217 as a means of investigating if we could make it more general purpose and public. Long term, it'd probably be desirable to move algorithms like this WidenAsciiToUtf16 to follow the same approach so that we get the best perf, with the least overhead.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is really useful. And something I can add on in future. Thanks for the detailed comment

// New Surface Area
//

static bool ISimdVector<Vector128<T>, T>.AnyMatches(Vector128<T> vector)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We did review/approve this (#98055 (comment)) and settled on bool Any(Vector128<T> vector, T value) and bool AnyWhereAllBitsSet(Vector<T> vector)

It would be nice to fix this to follow that. The PR otherwise looks good and should be mergeable.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So, the way I understand is we do not need AnyMatches() anymore

And Any and AnyWhereAllBitsSet would look something like follows


static bool ISimdVector<Vector512<T>, T>.AnyWhereAllBitsSet(Vector512<T> vector)
{
return (vector.EqualsAny(Vector512<T>.AllBitsSet));
}
static bool ISimdVector<Vector512<T>, T>.Any(Vector512<T> vector, T value)
{
return (vector.EqualsAny(Vector512.Create((T)value)));
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, essentially, with the ability for the JIT to recognize these as intrinsic and optimize them more in appropriate scenarios (but that's not necessary for this PR)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done - took out AnyMatches() and added AnyWhereAllBitsSet() and Any()

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@tannergooding Does this look correct? Anything else you'd like me to fix?


static bool ISimdVector<Vector128<T>, T>.AnyWhereAllBitsSet(Vector128<T> vector)
{
return (Vector128.EqualsAny(vector, Vector128<T>.AllBitsSet));

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: unnecessary parens here and in Any

If you could fix that in a follow up PR, that'd be great (going to merge this)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good! Will put it up later today.

@DrewScoggins

DrewScoggins commented Apr 11, 2024

Copy link
Copy Markdown
Member

@matouskozak

matouskozak commented Apr 16, 2024

Copy link
Copy Markdown
Member

@lewing

Copy link
Copy Markdown
Member

mono wasm aot regression dotnet/perf-autofiling-issues#32601

matouskozak pushed a commit to matouskozak/runtime that referenced this pull request Apr 30, 2024
* Add AnyMatches() to iSimdVector interface
* Switch to iSimdVector and Align WidenAsciiToUtf16.
* Fixing perf
* Addressing Review Comments.
* Mirroring API change : dotnet#98055 (comment)
@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 21, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-System.Runtime.Intrinsicscommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@DeepakRajendrakumaran@DrewScoggins@matouskozak@lewing@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Switch to iSimdVector and Align WidenAsciiToUtf16 - #99982

Merged
tannergooding merged 5 commits into
dotnet:mainfrom
DeepakRajendrakumaran:align
Apr 8, 2024
Merged

Switch to iSimdVector and Align WidenAsciiToUtf16#99982
tannergooding merged 5 commits into
dotnet:mainfrom
DeepakRajendrakumaran:align

Conversation

@DeepakRajendrakumaran

@DeepakRajendrakumaranDeepakRajendrakumaran commented Mar 19, 2024

Copy link
Copy Markdown
Contributor

This is an updated version of this PR(#89892). It does the following

  1. Add ' AnyMatches' support for iSimdVector
  2. Use iSimdVector to clean up 'WidenAsciiToUtf16' implementation
  3. Align memory stores

Perf Results

Ran the following tests(sizes : 16, 512, 1024, 5120, 10240) on EMR: (base = main branch, diff = with change)https://github.com/dotnet/performance/blob/47d21ee9571164a8e3f8088d8709ca4061d96827/src/benchmarks/micro/libraries/System.Text.Encoding/Perf.Encoding.cs

On EMR
image

On ICX - Not much diff
image

@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Mar 19, 2024
@DeepakRajendrakumaran
DeepakRajendrakumaran marked this pull request as ready for review March 20, 2024 22:25
@DeepakRajendrakumaran

DeepakRajendrakumaran commented Mar 20, 2024

Copy link
Copy Markdown
ContributorAuthor

@tannergooding @dotnet/avx512-contrib Can you please review this?

private static unsafe bool HasMatch<TVectorByte>(TVectorByte vector)
where TVectorByte : unmanaged, ISimdVector<TVectorByte, byte>
{
return !(vector & TVectorByte.Create((byte)0x80)).Equals(TVectorByte.Zero);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why not (vector & TVectorByte.Create((byte)0x80)) != TVectorByte.Zero?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I basically ran into a weird issue where perf degraded significantly for specific cases on my ICX when I had '(vector & TVectorByte.Create((byte)0x80)) != TVectorByte.Zero' and the issue went away with what I have now. It was pretty consistent. I had some trouble narrowing down the exact why with VTune though.

I decided to go with the performant version for now

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Was there a codegen difference between them? I'd expect them to generate the same code

@DeepakRajendrakumaranDeepakRajendrakumaranMar 27, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I did a quick check on godbolt and they all look the same - https://godbolt.org/z/P87dPdTGa

Now I'm curious if it was just something off when I ran it locally

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I did some more digging and something is off. Tried 3 versions with a handwritten benchmark
image
Benchmark
image

Match3 is significantly faster than Match1 and Match 2

VTune for Match1 vs Match3
image

The inlining makes it hard to narrow down

image

@DeepakRajendrakumaranDeepakRajendrakumaranMar 28, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do you have a strong preference for any of these patterns?

I can look into the why this is happening if it's important. For now, I'm just keeping the fast version

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@tannergooding I have created an issue detailing this as discussed: #100493

Please let me know if there is anything else needed for this PR

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment on lines +2164 to +2196
if (!HasMatch<TVectorByte>(asciiVector))
{
(TVectorUShort utf16LowVector, TVectorUShort utf16HighVector) = Widen<TVectorByte, TVectorUShort>(asciiVector);
utf16LowVector.Store(pCurrentWriteAddress);
utf16HighVector.Store(pCurrentWriteAddress + TVectorUShort.Count);
pCurrentWriteAddress += (nuint)(TVectorUShort.Count * 2);
if (((int)pCurrentWriteAddress & 1) == 0)
{
// Bump write buffer up to the next aligned boundary
pCurrentWriteAddress = (ushort*)((nuint)pCurrentWriteAddress & ~(nuint)(TVectorUShort.Alignment - 1));
nuint numBytesWritten = (nuint)pCurrentWriteAddress - (nuint)pUtf16Buffer;
currentOffset += (nuint)numBytesWritten / 2;
}
else
{
// If input isn't char aligned, we won't be able to align it to a Vector
currentOffset += (nuint)TVectorByte.Count;
}
while (currentOffset <= finalOffsetWhereCanRunLoop)
{
asciiVector = TVectorByte.Load(pAsciiBuffer + currentOffset);
if (HasMatch<TVectorByte>(asciiVector))
{
break;
}
(utf16LowVector, utf16HighVector) = Widen<TVectorByte, TVectorUShort>(asciiVector);
utf16LowVector.StoreAligned(pCurrentWriteAddress);
utf16HighVector.StoreAligned(pCurrentWriteAddress + TVectorUShort.Count);
currentOffset += (nuint)TVectorByte.Count;
pCurrentWriteAddress += (nuint)(TVectorUShort.Count * 2);
}
}

@tannergoodingtannergoodingMar 27, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The code here looks generally good, and I don't expect anything to be changed for this PR based on what I'm about to comment.

However, I would like to refer to how we set up everything for TensorPrimitives as it works very well and allows a lot of code sharing (noting it isn't using ISimdVector yet since it's out of band, but could easily do so in the future): https://source.dot.net/#System.Numerics.Tensors/System/Numerics/Tensors/netcore/Common/TensorPrimitives.IUnaryOperator.cs,fdab74764af40a1e

In general it tries to do the pre-checks up front and only ever execute 1 vector path (so it doesn't have to fallthrough from Vector512->Vector256->Vector128 as remainders exist). To achieve this, it has the Vectorized### helpers (which are all identical, except for the size they operate on, this is what would eventually use ISimdVector) and then a shared VectorizedSmall which is simply a jump table designed to handle any data that is less than a full vector using a single branch.

The core logic for the vectorized algorithm (https://source.dot.net/#System.Numerics.Tensors/System/Numerics/Tensors/netcore/Common/TensorPrimitives.IUnaryOperator.cs,ef9adce4e9561b04) then basically has a path to handle the main loop (which is currently unrolled by a factor of 8) and otherwise hits a jump table to handle remaining blocks (so they can likewise be handled with a single branch).

To help optimize, it preloads the beginning and ending vectors. In the worst case this will result in double processing of some inputs for very small sizes, but its ultimately only 2 main operations which is fine.

The main loop then attempts to align and has an optimized path for extremely large inputs for non-temporal data if alignment could be achieved. Smaller inputs just do regular unaligned stores since the actual address will have been aligned if that was feasible.

This general approach is done because it allows all paths, but particularly the smallest inputs, to minimize the total number of branches done (no more than 2 branches for non-vectorized data and no more 3 to hit the main loop for vectorized code). It also allows us to separate the "algorithm logic" from the "vectorization logic" and share that vectorization logic between multiple vectorized algorithms.

This general setup has worked so well and provided very stable perf numbers for all sizes, such that we opened #93217 as a means of investigating if we could make it more general purpose and public. Long term, it'd probably be desirable to move algorithms like this WidenAsciiToUtf16 to follow the same approach so that we get the best perf, with the least overhead.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is really useful. And something I can add on in future. Thanks for the detailed comment

// New Surface Area
//

static bool ISimdVector<Vector128<T>, T>.AnyMatches(Vector128<T> vector)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We did review/approve this (#98055 (comment)) and settled on bool Any(Vector128<T> vector, T value) and bool AnyWhereAllBitsSet(Vector<T> vector)

It would be nice to fix this to follow that. The PR otherwise looks good and should be mergeable.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So, the way I understand is we do not need AnyMatches() anymore

And Any and AnyWhereAllBitsSet would look something like follows


static bool ISimdVector<Vector512<T>, T>.AnyWhereAllBitsSet(Vector512<T> vector)
{
return (vector.EqualsAny(Vector512<T>.AllBitsSet));
}
static bool ISimdVector<Vector512<T>, T>.Any(Vector512<T> vector, T value)
{
return (vector.EqualsAny(Vector512.Create((T)value)));
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, essentially, with the ability for the JIT to recognize these as intrinsic and optimize them more in appropriate scenarios (but that's not necessary for this PR)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done - took out AnyMatches() and added AnyWhereAllBitsSet() and Any()

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@tannergooding Does this look correct? Anything else you'd like me to fix?


static bool ISimdVector<Vector128<T>, T>.AnyWhereAllBitsSet(Vector128<T> vector)
{
return (Vector128.EqualsAny(vector, Vector128<T>.AllBitsSet));

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: unnecessary parens here and in Any

If you could fix that in a follow up PR, that'd be great (going to merge this)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good! Will put it up later today.

@DrewScoggins

DrewScoggins commented Apr 11, 2024

Copy link
Copy Markdown
Member

@matouskozak

matouskozak commented Apr 16, 2024

Copy link
Copy Markdown
Member

@lewing

Copy link
Copy Markdown
Member

mono wasm aot regression dotnet/perf-autofiling-issues#32601

matouskozak pushed a commit to matouskozak/runtime that referenced this pull request Apr 30, 2024
* Add AnyMatches() to iSimdVector interface
* Switch to iSimdVector and Align WidenAsciiToUtf16.
* Fixing perf
* Addressing Review Comments.
* Mirroring API change : dotnet#98055 (comment)
@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 21, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-System.Runtime.Intrinsicscommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@DeepakRajendrakumaran@DrewScoggins@matouskozak@lewing@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Switch to iSimdVector and Align WidenAsciiToUtf16 - #99982

Merged
tannergooding merged 5 commits into
dotnet:mainfrom
DeepakRajendrakumaran:align
Apr 8, 2024
Merged

Switch to iSimdVector and Align WidenAsciiToUtf16#99982
tannergooding merged 5 commits into
dotnet:mainfrom
DeepakRajendrakumaran:align

Conversation

@DeepakRajendrakumaran

@DeepakRajendrakumaranDeepakRajendrakumaran commented Mar 19, 2024

Copy link
Copy Markdown
Contributor

This is an updated version of this PR(#89892). It does the following

  1. Add ' AnyMatches' support for iSimdVector
  2. Use iSimdVector to clean up 'WidenAsciiToUtf16' implementation
  3. Align memory stores

Perf Results

Ran the following tests(sizes : 16, 512, 1024, 5120, 10240) on EMR: (base = main branch, diff = with change)https://github.com/dotnet/performance/blob/47d21ee9571164a8e3f8088d8709ca4061d96827/src/benchmarks/micro/libraries/System.Text.Encoding/Perf.Encoding.cs

On EMR
image

On ICX - Not much diff
image

@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Mar 19, 2024
@DeepakRajendrakumaran
DeepakRajendrakumaran marked this pull request as ready for review March 20, 2024 22:25
@DeepakRajendrakumaran

DeepakRajendrakumaran commented Mar 20, 2024

Copy link
Copy Markdown
ContributorAuthor

@tannergooding @dotnet/avx512-contrib Can you please review this?

private static unsafe bool HasMatch<TVectorByte>(TVectorByte vector)
where TVectorByte : unmanaged, ISimdVector<TVectorByte, byte>
{
return !(vector & TVectorByte.Create((byte)0x80)).Equals(TVectorByte.Zero);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why not (vector & TVectorByte.Create((byte)0x80)) != TVectorByte.Zero?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I basically ran into a weird issue where perf degraded significantly for specific cases on my ICX when I had '(vector & TVectorByte.Create((byte)0x80)) != TVectorByte.Zero' and the issue went away with what I have now. It was pretty consistent. I had some trouble narrowing down the exact why with VTune though.

I decided to go with the performant version for now

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Was there a codegen difference between them? I'd expect them to generate the same code

@DeepakRajendrakumaranDeepakRajendrakumaranMar 27, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I did a quick check on godbolt and they all look the same - https://godbolt.org/z/P87dPdTGa

Now I'm curious if it was just something off when I ran it locally

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I did some more digging and something is off. Tried 3 versions with a handwritten benchmark
image
Benchmark
image

Match3 is significantly faster than Match1 and Match 2

VTune for Match1 vs Match3
image

The inlining makes it hard to narrow down

image

@DeepakRajendrakumaranDeepakRajendrakumaranMar 28, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do you have a strong preference for any of these patterns?

I can look into the why this is happening if it's important. For now, I'm just keeping the fast version

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@tannergooding I have created an issue detailing this as discussed: #100493

Please let me know if there is anything else needed for this PR

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment on lines +2164 to +2196
if (!HasMatch<TVectorByte>(asciiVector))
{
(TVectorUShort utf16LowVector, TVectorUShort utf16HighVector) = Widen<TVectorByte, TVectorUShort>(asciiVector);
utf16LowVector.Store(pCurrentWriteAddress);
utf16HighVector.Store(pCurrentWriteAddress + TVectorUShort.Count);
pCurrentWriteAddress += (nuint)(TVectorUShort.Count * 2);
if (((int)pCurrentWriteAddress & 1) == 0)
{
// Bump write buffer up to the next aligned boundary
pCurrentWriteAddress = (ushort*)((nuint)pCurrentWriteAddress & ~(nuint)(TVectorUShort.Alignment - 1));
nuint numBytesWritten = (nuint)pCurrentWriteAddress - (nuint)pUtf16Buffer;
currentOffset += (nuint)numBytesWritten / 2;
}
else
{
// If input isn't char aligned, we won't be able to align it to a Vector
currentOffset += (nuint)TVectorByte.Count;
}
while (currentOffset <= finalOffsetWhereCanRunLoop)
{
asciiVector = TVectorByte.Load(pAsciiBuffer + currentOffset);
if (HasMatch<TVectorByte>(asciiVector))
{
break;
}
(utf16LowVector, utf16HighVector) = Widen<TVectorByte, TVectorUShort>(asciiVector);
utf16LowVector.StoreAligned(pCurrentWriteAddress);
utf16HighVector.StoreAligned(pCurrentWriteAddress + TVectorUShort.Count);
currentOffset += (nuint)TVectorByte.Count;
pCurrentWriteAddress += (nuint)(TVectorUShort.Count * 2);
}
}

@tannergoodingtannergoodingMar 27, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The code here looks generally good, and I don't expect anything to be changed for this PR based on what I'm about to comment.

However, I would like to refer to how we set up everything for TensorPrimitives as it works very well and allows a lot of code sharing (noting it isn't using ISimdVector yet since it's out of band, but could easily do so in the future): https://source.dot.net/#System.Numerics.Tensors/System/Numerics/Tensors/netcore/Common/TensorPrimitives.IUnaryOperator.cs,fdab74764af40a1e

In general it tries to do the pre-checks up front and only ever execute 1 vector path (so it doesn't have to fallthrough from Vector512->Vector256->Vector128 as remainders exist). To achieve this, it has the Vectorized### helpers (which are all identical, except for the size they operate on, this is what would eventually use ISimdVector) and then a shared VectorizedSmall which is simply a jump table designed to handle any data that is less than a full vector using a single branch.

The core logic for the vectorized algorithm (https://source.dot.net/#System.Numerics.Tensors/System/Numerics/Tensors/netcore/Common/TensorPrimitives.IUnaryOperator.cs,ef9adce4e9561b04) then basically has a path to handle the main loop (which is currently unrolled by a factor of 8) and otherwise hits a jump table to handle remaining blocks (so they can likewise be handled with a single branch).

To help optimize, it preloads the beginning and ending vectors. In the worst case this will result in double processing of some inputs for very small sizes, but its ultimately only 2 main operations which is fine.

The main loop then attempts to align and has an optimized path for extremely large inputs for non-temporal data if alignment could be achieved. Smaller inputs just do regular unaligned stores since the actual address will have been aligned if that was feasible.

This general approach is done because it allows all paths, but particularly the smallest inputs, to minimize the total number of branches done (no more than 2 branches for non-vectorized data and no more 3 to hit the main loop for vectorized code). It also allows us to separate the "algorithm logic" from the "vectorization logic" and share that vectorization logic between multiple vectorized algorithms.

This general setup has worked so well and provided very stable perf numbers for all sizes, such that we opened #93217 as a means of investigating if we could make it more general purpose and public. Long term, it'd probably be desirable to move algorithms like this WidenAsciiToUtf16 to follow the same approach so that we get the best perf, with the least overhead.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is really useful. And something I can add on in future. Thanks for the detailed comment

// New Surface Area
//

static bool ISimdVector<Vector128<T>, T>.AnyMatches(Vector128<T> vector)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We did review/approve this (#98055 (comment)) and settled on bool Any(Vector128<T> vector, T value) and bool AnyWhereAllBitsSet(Vector<T> vector)

It would be nice to fix this to follow that. The PR otherwise looks good and should be mergeable.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So, the way I understand is we do not need AnyMatches() anymore

And Any and AnyWhereAllBitsSet would look something like follows


static bool ISimdVector<Vector512<T>, T>.AnyWhereAllBitsSet(Vector512<T> vector)
{
return (vector.EqualsAny(Vector512<T>.AllBitsSet));
}
static bool ISimdVector<Vector512<T>, T>.Any(Vector512<T> vector, T value)
{
return (vector.EqualsAny(Vector512.Create((T)value)));
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, essentially, with the ability for the JIT to recognize these as intrinsic and optimize them more in appropriate scenarios (but that's not necessary for this PR)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done - took out AnyMatches() and added AnyWhereAllBitsSet() and Any()

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@tannergooding Does this look correct? Anything else you'd like me to fix?


static bool ISimdVector<Vector128<T>, T>.AnyWhereAllBitsSet(Vector128<T> vector)
{
return (Vector128.EqualsAny(vector, Vector128<T>.AllBitsSet));

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: unnecessary parens here and in Any

If you could fix that in a follow up PR, that'd be great (going to merge this)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good! Will put it up later today.

@DrewScoggins

DrewScoggins commented Apr 11, 2024

Copy link
Copy Markdown
Member

@matouskozak

matouskozak commented Apr 16, 2024

Copy link
Copy Markdown
Member

@lewing

Copy link
Copy Markdown
Member

mono wasm aot regression dotnet/perf-autofiling-issues#32601

matouskozak pushed a commit to matouskozak/runtime that referenced this pull request Apr 30, 2024
* Add AnyMatches() to iSimdVector interface
* Switch to iSimdVector and Align WidenAsciiToUtf16.
* Fixing perf
* Addressing Review Comments.
* Mirroring API change : dotnet#98055 (comment)
@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 21, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-System.Runtime.Intrinsicscommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@DeepakRajendrakumaran@DrewScoggins@matouskozak@lewing@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Switch to iSimdVector and Align WidenAsciiToUtf16 - #99982

Merged
tannergooding merged 5 commits into
dotnet:mainfrom
DeepakRajendrakumaran:align
Apr 8, 2024
Merged

Switch to iSimdVector and Align WidenAsciiToUtf16#99982
tannergooding merged 5 commits into
dotnet:mainfrom
DeepakRajendrakumaran:align

Conversation

@DeepakRajendrakumaran

@DeepakRajendrakumaranDeepakRajendrakumaran commented Mar 19, 2024

Copy link
Copy Markdown
Contributor

This is an updated version of this PR(#89892). It does the following

  1. Add ' AnyMatches' support for iSimdVector
  2. Use iSimdVector to clean up 'WidenAsciiToUtf16' implementation
  3. Align memory stores

Perf Results

Ran the following tests(sizes : 16, 512, 1024, 5120, 10240) on EMR: (base = main branch, diff = with change)https://github.com/dotnet/performance/blob/47d21ee9571164a8e3f8088d8709ca4061d96827/src/benchmarks/micro/libraries/System.Text.Encoding/Perf.Encoding.cs

On EMR
image

On ICX - Not much diff
image

@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Mar 19, 2024
@DeepakRajendrakumaran
DeepakRajendrakumaran marked this pull request as ready for review March 20, 2024 22:25
@DeepakRajendrakumaran

DeepakRajendrakumaran commented Mar 20, 2024

Copy link
Copy Markdown
ContributorAuthor

@tannergooding @dotnet/avx512-contrib Can you please review this?

private static unsafe bool HasMatch<TVectorByte>(TVectorByte vector)
where TVectorByte : unmanaged, ISimdVector<TVectorByte, byte>
{
return !(vector & TVectorByte.Create((byte)0x80)).Equals(TVectorByte.Zero);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why not (vector & TVectorByte.Create((byte)0x80)) != TVectorByte.Zero?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I basically ran into a weird issue where perf degraded significantly for specific cases on my ICX when I had '(vector & TVectorByte.Create((byte)0x80)) != TVectorByte.Zero' and the issue went away with what I have now. It was pretty consistent. I had some trouble narrowing down the exact why with VTune though.

I decided to go with the performant version for now

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Was there a codegen difference between them? I'd expect them to generate the same code

@DeepakRajendrakumaranDeepakRajendrakumaranMar 27, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I did a quick check on godbolt and they all look the same - https://godbolt.org/z/P87dPdTGa

Now I'm curious if it was just something off when I ran it locally

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I did some more digging and something is off. Tried 3 versions with a handwritten benchmark
image
Benchmark
image

Match3 is significantly faster than Match1 and Match 2

VTune for Match1 vs Match3
image

The inlining makes it hard to narrow down

image

@DeepakRajendrakumaranDeepakRajendrakumaranMar 28, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do you have a strong preference for any of these patterns?

I can look into the why this is happening if it's important. For now, I'm just keeping the fast version

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@tannergooding I have created an issue detailing this as discussed: #100493

Please let me know if there is anything else needed for this PR

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment on lines +2164 to +2196
if (!HasMatch<TVectorByte>(asciiVector))
{
(TVectorUShort utf16LowVector, TVectorUShort utf16HighVector) = Widen<TVectorByte, TVectorUShort>(asciiVector);
utf16LowVector.Store(pCurrentWriteAddress);
utf16HighVector.Store(pCurrentWriteAddress + TVectorUShort.Count);
pCurrentWriteAddress += (nuint)(TVectorUShort.Count * 2);
if (((int)pCurrentWriteAddress & 1) == 0)
{
// Bump write buffer up to the next aligned boundary
pCurrentWriteAddress = (ushort*)((nuint)pCurrentWriteAddress & ~(nuint)(TVectorUShort.Alignment - 1));
nuint numBytesWritten = (nuint)pCurrentWriteAddress - (nuint)pUtf16Buffer;
currentOffset += (nuint)numBytesWritten / 2;
}
else
{
// If input isn't char aligned, we won't be able to align it to a Vector
currentOffset += (nuint)TVectorByte.Count;
}
while (currentOffset <= finalOffsetWhereCanRunLoop)
{
asciiVector = TVectorByte.Load(pAsciiBuffer + currentOffset);
if (HasMatch<TVectorByte>(asciiVector))
{
break;
}
(utf16LowVector, utf16HighVector) = Widen<TVectorByte, TVectorUShort>(asciiVector);
utf16LowVector.StoreAligned(pCurrentWriteAddress);
utf16HighVector.StoreAligned(pCurrentWriteAddress + TVectorUShort.Count);
currentOffset += (nuint)TVectorByte.Count;
pCurrentWriteAddress += (nuint)(TVectorUShort.Count * 2);
}
}

@tannergoodingtannergoodingMar 27, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The code here looks generally good, and I don't expect anything to be changed for this PR based on what I'm about to comment.

However, I would like to refer to how we set up everything for TensorPrimitives as it works very well and allows a lot of code sharing (noting it isn't using ISimdVector yet since it's out of band, but could easily do so in the future): https://source.dot.net/#System.Numerics.Tensors/System/Numerics/Tensors/netcore/Common/TensorPrimitives.IUnaryOperator.cs,fdab74764af40a1e

In general it tries to do the pre-checks up front and only ever execute 1 vector path (so it doesn't have to fallthrough from Vector512->Vector256->Vector128 as remainders exist). To achieve this, it has the Vectorized### helpers (which are all identical, except for the size they operate on, this is what would eventually use ISimdVector) and then a shared VectorizedSmall which is simply a jump table designed to handle any data that is less than a full vector using a single branch.

The core logic for the vectorized algorithm (https://source.dot.net/#System.Numerics.Tensors/System/Numerics/Tensors/netcore/Common/TensorPrimitives.IUnaryOperator.cs,ef9adce4e9561b04) then basically has a path to handle the main loop (which is currently unrolled by a factor of 8) and otherwise hits a jump table to handle remaining blocks (so they can likewise be handled with a single branch).

To help optimize, it preloads the beginning and ending vectors. In the worst case this will result in double processing of some inputs for very small sizes, but its ultimately only 2 main operations which is fine.

The main loop then attempts to align and has an optimized path for extremely large inputs for non-temporal data if alignment could be achieved. Smaller inputs just do regular unaligned stores since the actual address will have been aligned if that was feasible.

This general approach is done because it allows all paths, but particularly the smallest inputs, to minimize the total number of branches done (no more than 2 branches for non-vectorized data and no more 3 to hit the main loop for vectorized code). It also allows us to separate the "algorithm logic" from the "vectorization logic" and share that vectorization logic between multiple vectorized algorithms.

This general setup has worked so well and provided very stable perf numbers for all sizes, such that we opened #93217 as a means of investigating if we could make it more general purpose and public. Long term, it'd probably be desirable to move algorithms like this WidenAsciiToUtf16 to follow the same approach so that we get the best perf, with the least overhead.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is really useful. And something I can add on in future. Thanks for the detailed comment

// New Surface Area
//

static bool ISimdVector<Vector128<T>, T>.AnyMatches(Vector128<T> vector)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We did review/approve this (#98055 (comment)) and settled on bool Any(Vector128<T> vector, T value) and bool AnyWhereAllBitsSet(Vector<T> vector)

It would be nice to fix this to follow that. The PR otherwise looks good and should be mergeable.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So, the way I understand is we do not need AnyMatches() anymore

And Any and AnyWhereAllBitsSet would look something like follows


static bool ISimdVector<Vector512<T>, T>.AnyWhereAllBitsSet(Vector512<T> vector)
{
return (vector.EqualsAny(Vector512<T>.AllBitsSet));
}
static bool ISimdVector<Vector512<T>, T>.Any(Vector512<T> vector, T value)
{
return (vector.EqualsAny(Vector512.Create((T)value)));
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, essentially, with the ability for the JIT to recognize these as intrinsic and optimize them more in appropriate scenarios (but that's not necessary for this PR)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done - took out AnyMatches() and added AnyWhereAllBitsSet() and Any()

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@tannergooding Does this look correct? Anything else you'd like me to fix?


static bool ISimdVector<Vector128<T>, T>.AnyWhereAllBitsSet(Vector128<T> vector)
{
return (Vector128.EqualsAny(vector, Vector128<T>.AllBitsSet));

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: unnecessary parens here and in Any

If you could fix that in a follow up PR, that'd be great (going to merge this)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good! Will put it up later today.

@DrewScoggins

DrewScoggins commented Apr 11, 2024

Copy link
Copy Markdown
Member

@matouskozak

matouskozak commented Apr 16, 2024

Copy link
Copy Markdown
Member

@lewing

Copy link
Copy Markdown
Member

mono wasm aot regression dotnet/perf-autofiling-issues#32601

matouskozak pushed a commit to matouskozak/runtime that referenced this pull request Apr 30, 2024
* Add AnyMatches() to iSimdVector interface
* Switch to iSimdVector and Align WidenAsciiToUtf16.
* Fixing perf
* Addressing Review Comments.
* Mirroring API change : dotnet#98055 (comment)
@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 21, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-System.Runtime.Intrinsicscommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@DeepakRajendrakumaran@DrewScoggins@matouskozak@lewing@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Switch to iSimdVector and Align WidenAsciiToUtf16 - #99982

Merged
tannergooding merged 5 commits into
dotnet:mainfrom
DeepakRajendrakumaran:align
Apr 8, 2024
Merged

Switch to iSimdVector and Align WidenAsciiToUtf16#99982
tannergooding merged 5 commits into
dotnet:mainfrom
DeepakRajendrakumaran:align

Conversation

@DeepakRajendrakumaran

@DeepakRajendrakumaranDeepakRajendrakumaran commented Mar 19, 2024

Copy link
Copy Markdown
Contributor

This is an updated version of this PR(#89892). It does the following

  1. Add ' AnyMatches' support for iSimdVector
  2. Use iSimdVector to clean up 'WidenAsciiToUtf16' implementation
  3. Align memory stores

Perf Results

Ran the following tests(sizes : 16, 512, 1024, 5120, 10240) on EMR: (base = main branch, diff = with change)https://github.com/dotnet/performance/blob/47d21ee9571164a8e3f8088d8709ca4061d96827/src/benchmarks/micro/libraries/System.Text.Encoding/Perf.Encoding.cs

On EMR
image

On ICX - Not much diff
image

@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Mar 19, 2024
@DeepakRajendrakumaran
DeepakRajendrakumaran marked this pull request as ready for review March 20, 2024 22:25
@DeepakRajendrakumaran

DeepakRajendrakumaran commented Mar 20, 2024

Copy link
Copy Markdown
ContributorAuthor

@tannergooding @dotnet/avx512-contrib Can you please review this?

private static unsafe bool HasMatch<TVectorByte>(TVectorByte vector)
where TVectorByte : unmanaged, ISimdVector<TVectorByte, byte>
{
return !(vector & TVectorByte.Create((byte)0x80)).Equals(TVectorByte.Zero);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why not (vector & TVectorByte.Create((byte)0x80)) != TVectorByte.Zero?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I basically ran into a weird issue where perf degraded significantly for specific cases on my ICX when I had '(vector & TVectorByte.Create((byte)0x80)) != TVectorByte.Zero' and the issue went away with what I have now. It was pretty consistent. I had some trouble narrowing down the exact why with VTune though.

I decided to go with the performant version for now

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Was there a codegen difference between them? I'd expect them to generate the same code

@DeepakRajendrakumaranDeepakRajendrakumaranMar 27, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I did a quick check on godbolt and they all look the same - https://godbolt.org/z/P87dPdTGa

Now I'm curious if it was just something off when I ran it locally

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I did some more digging and something is off. Tried 3 versions with a handwritten benchmark
image
Benchmark
image

Match3 is significantly faster than Match1 and Match 2

VTune for Match1 vs Match3
image

The inlining makes it hard to narrow down

image

@DeepakRajendrakumaranDeepakRajendrakumaranMar 28, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do you have a strong preference for any of these patterns?

I can look into the why this is happening if it's important. For now, I'm just keeping the fast version

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@tannergooding I have created an issue detailing this as discussed: #100493

Please let me know if there is anything else needed for this PR

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
Comment on lines +2164 to +2196
if (!HasMatch<TVectorByte>(asciiVector))
{
(TVectorUShort utf16LowVector, TVectorUShort utf16HighVector) = Widen<TVectorByte, TVectorUShort>(asciiVector);
utf16LowVector.Store(pCurrentWriteAddress);
utf16HighVector.Store(pCurrentWriteAddress + TVectorUShort.Count);
pCurrentWriteAddress += (nuint)(TVectorUShort.Count * 2);
if (((int)pCurrentWriteAddress & 1) == 0)
{
// Bump write buffer up to the next aligned boundary
pCurrentWriteAddress = (ushort*)((nuint)pCurrentWriteAddress & ~(nuint)(TVectorUShort.Alignment - 1));
nuint numBytesWritten = (nuint)pCurrentWriteAddress - (nuint)pUtf16Buffer;
currentOffset += (nuint)numBytesWritten / 2;
}
else
{
// If input isn't char aligned, we won't be able to align it to a Vector
currentOffset += (nuint)TVectorByte.Count;
}
while (currentOffset <= finalOffsetWhereCanRunLoop)
{
asciiVector = TVectorByte.Load(pAsciiBuffer + currentOffset);
if (HasMatch<TVectorByte>(asciiVector))
{
break;
}
(utf16LowVector, utf16HighVector) = Widen<TVectorByte, TVectorUShort>(asciiVector);
utf16LowVector.StoreAligned(pCurrentWriteAddress);
utf16HighVector.StoreAligned(pCurrentWriteAddress + TVectorUShort.Count);
currentOffset += (nuint)TVectorByte.Count;
pCurrentWriteAddress += (nuint)(TVectorUShort.Count * 2);
}
}

@tannergoodingtannergoodingMar 27, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The code here looks generally good, and I don't expect anything to be changed for this PR based on what I'm about to comment.

However, I would like to refer to how we set up everything for TensorPrimitives as it works very well and allows a lot of code sharing (noting it isn't using ISimdVector yet since it's out of band, but could easily do so in the future): https://source.dot.net/#System.Numerics.Tensors/System/Numerics/Tensors/netcore/Common/TensorPrimitives.IUnaryOperator.cs,fdab74764af40a1e

In general it tries to do the pre-checks up front and only ever execute 1 vector path (so it doesn't have to fallthrough from Vector512->Vector256->Vector128 as remainders exist). To achieve this, it has the Vectorized### helpers (which are all identical, except for the size they operate on, this is what would eventually use ISimdVector) and then a shared VectorizedSmall which is simply a jump table designed to handle any data that is less than a full vector using a single branch.

The core logic for the vectorized algorithm (https://source.dot.net/#System.Numerics.Tensors/System/Numerics/Tensors/netcore/Common/TensorPrimitives.IUnaryOperator.cs,ef9adce4e9561b04) then basically has a path to handle the main loop (which is currently unrolled by a factor of 8) and otherwise hits a jump table to handle remaining blocks (so they can likewise be handled with a single branch).

To help optimize, it preloads the beginning and ending vectors. In the worst case this will result in double processing of some inputs for very small sizes, but its ultimately only 2 main operations which is fine.

The main loop then attempts to align and has an optimized path for extremely large inputs for non-temporal data if alignment could be achieved. Smaller inputs just do regular unaligned stores since the actual address will have been aligned if that was feasible.

This general approach is done because it allows all paths, but particularly the smallest inputs, to minimize the total number of branches done (no more than 2 branches for non-vectorized data and no more 3 to hit the main loop for vectorized code). It also allows us to separate the "algorithm logic" from the "vectorization logic" and share that vectorization logic between multiple vectorized algorithms.

This general setup has worked so well and provided very stable perf numbers for all sizes, such that we opened #93217 as a means of investigating if we could make it more general purpose and public. Long term, it'd probably be desirable to move algorithms like this WidenAsciiToUtf16 to follow the same approach so that we get the best perf, with the least overhead.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is really useful. And something I can add on in future. Thanks for the detailed comment

// New Surface Area
//

static bool ISimdVector<Vector128<T>, T>.AnyMatches(Vector128<T> vector)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We did review/approve this (#98055 (comment)) and settled on bool Any(Vector128<T> vector, T value) and bool AnyWhereAllBitsSet(Vector<T> vector)

It would be nice to fix this to follow that. The PR otherwise looks good and should be mergeable.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So, the way I understand is we do not need AnyMatches() anymore

And Any and AnyWhereAllBitsSet would look something like follows


static bool ISimdVector<Vector512<T>, T>.AnyWhereAllBitsSet(Vector512<T> vector)
{
return (vector.EqualsAny(Vector512<T>.AllBitsSet));
}
static bool ISimdVector<Vector512<T>, T>.Any(Vector512<T> vector, T value)
{
return (vector.EqualsAny(Vector512.Create((T)value)));
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, essentially, with the ability for the JIT to recognize these as intrinsic and optimize them more in appropriate scenarios (but that's not necessary for this PR)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done - took out AnyMatches() and added AnyWhereAllBitsSet() and Any()

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@tannergooding Does this look correct? Anything else you'd like me to fix?


static bool ISimdVector<Vector128<T>, T>.AnyWhereAllBitsSet(Vector128<T> vector)
{
return (Vector128.EqualsAny(vector, Vector128<T>.AllBitsSet));

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: unnecessary parens here and in Any

If you could fix that in a follow up PR, that'd be great (going to merge this)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good! Will put it up later today.

@DrewScoggins

DrewScoggins commented Apr 11, 2024

Copy link
Copy Markdown
Member

@matouskozak

matouskozak commented Apr 16, 2024

Copy link
Copy Markdown
Member

@lewing

Copy link
Copy Markdown
Member

mono wasm aot regression dotnet/perf-autofiling-issues#32601

matouskozak pushed a commit to matouskozak/runtime that referenced this pull request Apr 30, 2024
* Add AnyMatches() to iSimdVector interface
* Switch to iSimdVector and Align WidenAsciiToUtf16.
* Fixing perf
* Addressing Review Comments.
* Mirroring API change : dotnet#98055 (comment)
@github-actionsgithub-actionsBot locked and limited conversation to collaborators May 21, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-System.Runtime.Intrinsicscommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@DeepakRajendrakumaran@DrewScoggins@matouskozak@lewing@tannergooding