Use StoreAligned not Store in WidenAsciiToUtf16 - #89892

Closed
Ruihan-Yin wants to merge 3 commits into
dotnet:mainfrom
Ruihan-Yin:WriteAlign
Closed

Use StoreAligned not Store in WidenAsciiToUtf16#89892
Ruihan-Yin wants to merge 3 commits into
dotnet:mainfrom
Ruihan-Yin:WriteAlign

Conversation

@Ruihan-Yin

Copy link
Copy Markdown
Member

Description

This PR is to improve the in-loop write logic in WidenAsciiToUtf16, the major change is to replace StoreAligned with Store inside the loop to reduce the penalty caused by split loads.

We are open to adjusting the implementation style for either path. Perf number attached in the comments.

@ghostghost added community-contribution Indicates that the PR has been added by a community member area-System.Text.Encoding labels Aug 2, 2023
@ghost

ghost commented Aug 2, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/area-system-text-encoding
See info in area-owners.md if you want to be subscribed.

Issue Details

Description

This PR is to improve the in-loop write logic in WidenAsciiToUtf16, the major change is to replace StoreAligned with Store inside the loop to reduce the penalty caused by split loads.

We are open to adjusting the implementation style for either path. Perf number attached in the comments.

Author:Ruihan-Yin
Assignees:-
Labels:

area-System.Text.Encoding, community-contribution

Milestone:-

@Ruihan-Yin

Copy link
Copy Markdown
MemberAuthor

Perf numbers

Base: main (531ad95)
Diff: main + changes on WidenAsciiToUtf16

Avx512

summary:
better: 5, geomean: 1.099
worse: 1, geomean: 1.022
total diff: 6

Slowerdiff/baseBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "utf-8")1.0285.2587.11
Fasterbase/diffBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "utf-8")1.18176.81150.47
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "ascii")1.14152.16133.17
System.Text.Tests.Perf_Encoding.GetChars(size: 16, encName: "ascii")1.0917.9716.54
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "ascii")1.0679.4774.72
System.Text.Tests.Perf_Encoding.GetChars(size: 16, encName: "utf-8")1.0325.7924.94

AVX

summary:
better: 3, geomean: 1.049
total diff: 3

No Slower results for the provided threshold = 1% and noise filter = 0.5 ns.

Fasterbase/diffBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "ascii")1.0781.6876.09
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "utf-8")1.05174.54165.92
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "ascii")1.02148.66145.48

SSE

summary:
better: 5, geomean: 1.065
total diff: 5

No Slower results for the provided threshold = 1% and noise filter = 0.5 ns.

Fasterbase/diffBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "ascii")1.11108.0797.21
System.Text.Tests.Perf_Encoding.GetChars(size: 16, encName: "ascii")1.1017.7416.14
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "ascii")1.05183.11173.62
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "utf-8")1.04219.66211.76
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "utf-8")1.02119.36116.59

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
@neon-sunset

neon-sunset commented Aug 3, 2023

Copy link
Copy Markdown
Contributor

How does this change affect *-arm64 targets? The PR introduces unconditional .ExtractMostSignificantBits() which is expected to regress its performance.

UPD: @Ruihan-Yin thank you

@Ruihan-Yin

Ruihan-Yin commented Aug 4, 2023

Copy link
Copy Markdown
MemberAuthor

How does this change affect *-arm64 targets? The PR introduces unconditional .ExtractMostSignificantBits() which is expected to regress its performance.

Thanks for pointing out. That was a mistake, changed to VectorContainsNonAsciiChar to make sure Arm64 won't be affected by the use of ExtractMSB.

@Ruihan-Yin

Copy link
Copy Markdown
MemberAuthor

Hi @xtqqczze@neon-sunset, kindly asking if there is any further change needed?

@xtqqczze

xtqqczze commented Aug 9, 2023

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte), as System.Text.Encoding.GetChars is a public API.

I'm not sure how to handle this, but System.SpanHelpers.IndexOfNullCharacter(System.Char*) does the following check:

if(((int)searchSpace&1)!=0)
{
// Input isn't char aligned, we won't be able to align it to a Vector
}

@MichalPetryka

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte), as System.Text.Encoding.GetChars is a public API.

AFAIR the guideline is that if the platform the code is running on works with unaligned memory, dotnet APIs should work too.

@xtqqczze

Copy link
Copy Markdown
Contributor

AFAIR the guideline is that if the platform the code is running on works with unaligned memory, dotnet APIs should work too.

@MichalPetryka We have existing code that does not account for this, e.g. WidenLatin1ToUtf16, see #90319.

@anthonycanino

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte), as System.Text.Encoding.GetChars is a public API.

I'm not sure how to handle this, but System.SpanHelpers.IndexOfNullCharacter(System.Char*) does the following check:

if(((int)searchSpace&1)!=0)
{
// Input isn't char aligned, we won't be able to align it to a Vector
}

@tannergooding based on some of the conversation we had, I am not sure how to proceed.

Perhaps we want to implement a fallback method that does not pin, and does not explicitly use StoreAligned and instead uses Store but skips the initial alignment if the char pointer is not naturally aligned?

@xtqqczze

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte)

See also comments at #90319.

@eiriktsarpaliseiriktsarpalis added this to the 9.0.0 milestone Aug 14, 2023
@tannergooding

tannergooding commented Aug 14, 2023

Copy link
Copy Markdown
Member

Perhaps we want to implement a fallback method that does not pin, and does not explicitly use StoreAligned and instead uses Store but skips the initial alignment if the char pointer is not naturally aligned?

For right now we should keep the pin and try to align, but if its unalignable then we should just continue as-is. It's basically the same code just with a check for "is this alignable at all" and using Store rather than StoreAligned (this is ultimately the same codegen and perf on modern hardware since the underlying data will actually be aligned in most cases).

For the future, we probably want to come to an agreement about how to universally handle this. My guess/vote is that's probably going to involve not pinning and using StoreUnsafe with optimistic alignment of the underlying data (basically optimistically presuming it won't be moved by the GC, which will be the common case but still using ref so that if it is moved, everything still works as expected). There's ultimately a balance between writing safe/readable code and performant code, so we need to find the right spot to land. Ideally we'd just be using Vector128.Create(ROSpan<T>) and the JIT would elide the bounds checks, so we don't have any unsafeness; but that's not possible today.

@eiriktsarpalis

Copy link
Copy Markdown
Member

Just checking up on the status of this PR, @anthonycanino have you had the chance to address @tannergooding's feedback?

@anthonycanino

Copy link
Copy Markdown
Contributor

Just checking up on the status of this PR, @anthonycanino have you had the chance to address @tannergooding's feedback?

@tannergooding how do we feel about this change now that we are moving to ISimdVector? Is it better to address the alignment with that change in one PR?

@tannergooding

Copy link
Copy Markdown
Member

I don't think it's changed from my last feedback.

We really need to account for unalignable data and the easiest way to do that is to pin, check the alignment, align if possible, and then continue processing using Store which works regardless of whether the data is aligned or unaligned.

In general the code pattern for efficiently handling alignment efficiently, for idempotent data, looks something like https://source.dot.net/#System.Numerics.Tensors/System/Numerics/Tensors/TensorPrimitives.netcore.cs,2930

Using ISimdVector then lets you merge the 3 different vectorized code paths down to 1 shared code path. The amount of unrolling and other factors can depend on the exact algorithm, the number of inputs, etc. But this is the general basic shape that works well for both large and small inputs.

@tannergoodingtannergooding added the needs-author-action An issue or pull request that requires more info or actions from the author. label Nov 3, 2023
@eiriktsarpaliseiriktsarpalis self-assigned this Nov 15, 2023
@ghost

Copy link
Copy Markdown

This pull request has been automatically marked no-recent-activity because it has not had any activity for 14 days. It will be closed if no further activity occurs within 14 more days. Any new comment (by anyone, not necessarily the author) will remove no-recent-activity.

@ghost

Copy link
Copy Markdown

This pull request will now be closed since it had been marked no-recent-activity but received no further activity in the past 14 days. It is still possible to reopen or comment on the pull request, but please note that it will be locked if it remains inactive for another 30 days.

@ghostghost closed this Dec 13, 2023
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jan 13, 2024
This pull request was closed.
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-System.Text.Encodingcommunity-contributionIndicates that the PR has been added by a community memberneeds-author-actionAn issue or pull request that requires more info or actions from the author.no-recent-activity

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@Ruihan-Yin@neon-sunset@xtqqczze@MichalPetryka@anthonycanino@tannergooding@eiriktsarpalis
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Use StoreAligned not Store in WidenAsciiToUtf16 - #89892

Closed
Ruihan-Yin wants to merge 3 commits into
dotnet:mainfrom
Ruihan-Yin:WriteAlign
Closed

Use StoreAligned not Store in WidenAsciiToUtf16#89892
Ruihan-Yin wants to merge 3 commits into
dotnet:mainfrom
Ruihan-Yin:WriteAlign

Conversation

@Ruihan-Yin

Copy link
Copy Markdown
Member

Description

This PR is to improve the in-loop write logic in WidenAsciiToUtf16, the major change is to replace StoreAligned with Store inside the loop to reduce the penalty caused by split loads.

We are open to adjusting the implementation style for either path. Perf number attached in the comments.

@ghostghost added community-contribution Indicates that the PR has been added by a community member area-System.Text.Encoding labels Aug 2, 2023
@ghost

ghost commented Aug 2, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/area-system-text-encoding
See info in area-owners.md if you want to be subscribed.

Issue Details

Description

This PR is to improve the in-loop write logic in WidenAsciiToUtf16, the major change is to replace StoreAligned with Store inside the loop to reduce the penalty caused by split loads.

We are open to adjusting the implementation style for either path. Perf number attached in the comments.

Author:Ruihan-Yin
Assignees:-
Labels:

area-System.Text.Encoding, community-contribution

Milestone:-

@Ruihan-Yin

Copy link
Copy Markdown
MemberAuthor

Perf numbers

Base: main (531ad95)
Diff: main + changes on WidenAsciiToUtf16

Avx512

summary:
better: 5, geomean: 1.099
worse: 1, geomean: 1.022
total diff: 6

Slowerdiff/baseBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "utf-8")1.0285.2587.11
Fasterbase/diffBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "utf-8")1.18176.81150.47
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "ascii")1.14152.16133.17
System.Text.Tests.Perf_Encoding.GetChars(size: 16, encName: "ascii")1.0917.9716.54
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "ascii")1.0679.4774.72
System.Text.Tests.Perf_Encoding.GetChars(size: 16, encName: "utf-8")1.0325.7924.94

AVX

summary:
better: 3, geomean: 1.049
total diff: 3

No Slower results for the provided threshold = 1% and noise filter = 0.5 ns.

Fasterbase/diffBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "ascii")1.0781.6876.09
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "utf-8")1.05174.54165.92
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "ascii")1.02148.66145.48

SSE

summary:
better: 5, geomean: 1.065
total diff: 5

No Slower results for the provided threshold = 1% and noise filter = 0.5 ns.

Fasterbase/diffBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "ascii")1.11108.0797.21
System.Text.Tests.Perf_Encoding.GetChars(size: 16, encName: "ascii")1.1017.7416.14
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "ascii")1.05183.11173.62
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "utf-8")1.04219.66211.76
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "utf-8")1.02119.36116.59

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
@neon-sunset

neon-sunset commented Aug 3, 2023

Copy link
Copy Markdown
Contributor

How does this change affect *-arm64 targets? The PR introduces unconditional .ExtractMostSignificantBits() which is expected to regress its performance.

UPD: @Ruihan-Yin thank you

@Ruihan-Yin

Ruihan-Yin commented Aug 4, 2023

Copy link
Copy Markdown
MemberAuthor

How does this change affect *-arm64 targets? The PR introduces unconditional .ExtractMostSignificantBits() which is expected to regress its performance.

Thanks for pointing out. That was a mistake, changed to VectorContainsNonAsciiChar to make sure Arm64 won't be affected by the use of ExtractMSB.

@Ruihan-Yin

Copy link
Copy Markdown
MemberAuthor

Hi @xtqqczze@neon-sunset, kindly asking if there is any further change needed?

@xtqqczze

xtqqczze commented Aug 9, 2023

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte), as System.Text.Encoding.GetChars is a public API.

I'm not sure how to handle this, but System.SpanHelpers.IndexOfNullCharacter(System.Char*) does the following check:

if(((int)searchSpace&1)!=0)
{
// Input isn't char aligned, we won't be able to align it to a Vector
}

@MichalPetryka

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte), as System.Text.Encoding.GetChars is a public API.

AFAIR the guideline is that if the platform the code is running on works with unaligned memory, dotnet APIs should work too.

@xtqqczze

Copy link
Copy Markdown
Contributor

AFAIR the guideline is that if the platform the code is running on works with unaligned memory, dotnet APIs should work too.

@MichalPetryka We have existing code that does not account for this, e.g. WidenLatin1ToUtf16, see #90319.

@anthonycanino

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte), as System.Text.Encoding.GetChars is a public API.

I'm not sure how to handle this, but System.SpanHelpers.IndexOfNullCharacter(System.Char*) does the following check:

if(((int)searchSpace&1)!=0)
{
// Input isn't char aligned, we won't be able to align it to a Vector
}

@tannergooding based on some of the conversation we had, I am not sure how to proceed.

Perhaps we want to implement a fallback method that does not pin, and does not explicitly use StoreAligned and instead uses Store but skips the initial alignment if the char pointer is not naturally aligned?

@xtqqczze

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte)

See also comments at #90319.

@eiriktsarpaliseiriktsarpalis added this to the 9.0.0 milestone Aug 14, 2023
@tannergooding

tannergooding commented Aug 14, 2023

Copy link
Copy Markdown
Member

Perhaps we want to implement a fallback method that does not pin, and does not explicitly use StoreAligned and instead uses Store but skips the initial alignment if the char pointer is not naturally aligned?

For right now we should keep the pin and try to align, but if its unalignable then we should just continue as-is. It's basically the same code just with a check for "is this alignable at all" and using Store rather than StoreAligned (this is ultimately the same codegen and perf on modern hardware since the underlying data will actually be aligned in most cases).

For the future, we probably want to come to an agreement about how to universally handle this. My guess/vote is that's probably going to involve not pinning and using StoreUnsafe with optimistic alignment of the underlying data (basically optimistically presuming it won't be moved by the GC, which will be the common case but still using ref so that if it is moved, everything still works as expected). There's ultimately a balance between writing safe/readable code and performant code, so we need to find the right spot to land. Ideally we'd just be using Vector128.Create(ROSpan<T>) and the JIT would elide the bounds checks, so we don't have any unsafeness; but that's not possible today.

@eiriktsarpalis

Copy link
Copy Markdown
Member

Just checking up on the status of this PR, @anthonycanino have you had the chance to address @tannergooding's feedback?

@anthonycanino

Copy link
Copy Markdown
Contributor

Just checking up on the status of this PR, @anthonycanino have you had the chance to address @tannergooding's feedback?

@tannergooding how do we feel about this change now that we are moving to ISimdVector? Is it better to address the alignment with that change in one PR?

@tannergooding

Copy link
Copy Markdown
Member

I don't think it's changed from my last feedback.

We really need to account for unalignable data and the easiest way to do that is to pin, check the alignment, align if possible, and then continue processing using Store which works regardless of whether the data is aligned or unaligned.

In general the code pattern for efficiently handling alignment efficiently, for idempotent data, looks something like https://source.dot.net/#System.Numerics.Tensors/System/Numerics/Tensors/TensorPrimitives.netcore.cs,2930

Using ISimdVector then lets you merge the 3 different vectorized code paths down to 1 shared code path. The amount of unrolling and other factors can depend on the exact algorithm, the number of inputs, etc. But this is the general basic shape that works well for both large and small inputs.

@tannergoodingtannergooding added the needs-author-action An issue or pull request that requires more info or actions from the author. label Nov 3, 2023
@eiriktsarpaliseiriktsarpalis self-assigned this Nov 15, 2023
@ghost

Copy link
Copy Markdown

This pull request has been automatically marked no-recent-activity because it has not had any activity for 14 days. It will be closed if no further activity occurs within 14 more days. Any new comment (by anyone, not necessarily the author) will remove no-recent-activity.

@ghost

Copy link
Copy Markdown

This pull request will now be closed since it had been marked no-recent-activity but received no further activity in the past 14 days. It is still possible to reopen or comment on the pull request, but please note that it will be locked if it remains inactive for another 30 days.

@ghostghost closed this Dec 13, 2023
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jan 13, 2024
This pull request was closed.
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-System.Text.Encodingcommunity-contributionIndicates that the PR has been added by a community memberneeds-author-actionAn issue or pull request that requires more info or actions from the author.no-recent-activity

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@Ruihan-Yin@neon-sunset@xtqqczze@MichalPetryka@anthonycanino@tannergooding@eiriktsarpalis
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use StoreAligned not Store in WidenAsciiToUtf16 - #89892

Closed
Ruihan-Yin wants to merge 3 commits into
dotnet:mainfrom
Ruihan-Yin:WriteAlign
Closed

Use StoreAligned not Store in WidenAsciiToUtf16#89892
Ruihan-Yin wants to merge 3 commits into
dotnet:mainfrom
Ruihan-Yin:WriteAlign

Conversation

@Ruihan-Yin

Copy link
Copy Markdown
Member

Description

This PR is to improve the in-loop write logic in WidenAsciiToUtf16, the major change is to replace StoreAligned with Store inside the loop to reduce the penalty caused by split loads.

We are open to adjusting the implementation style for either path. Perf number attached in the comments.

@ghostghost added community-contribution Indicates that the PR has been added by a community member area-System.Text.Encoding labels Aug 2, 2023
@ghost

ghost commented Aug 2, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/area-system-text-encoding
See info in area-owners.md if you want to be subscribed.

Issue Details

Description

This PR is to improve the in-loop write logic in WidenAsciiToUtf16, the major change is to replace StoreAligned with Store inside the loop to reduce the penalty caused by split loads.

We are open to adjusting the implementation style for either path. Perf number attached in the comments.

Author:Ruihan-Yin
Assignees:-
Labels:

area-System.Text.Encoding, community-contribution

Milestone:-

@Ruihan-Yin

Copy link
Copy Markdown
MemberAuthor

Perf numbers

Base: main (531ad95)
Diff: main + changes on WidenAsciiToUtf16

Avx512

summary:
better: 5, geomean: 1.099
worse: 1, geomean: 1.022
total diff: 6

Slowerdiff/baseBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "utf-8")1.0285.2587.11
Fasterbase/diffBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "utf-8")1.18176.81150.47
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "ascii")1.14152.16133.17
System.Text.Tests.Perf_Encoding.GetChars(size: 16, encName: "ascii")1.0917.9716.54
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "ascii")1.0679.4774.72
System.Text.Tests.Perf_Encoding.GetChars(size: 16, encName: "utf-8")1.0325.7924.94

AVX

summary:
better: 3, geomean: 1.049
total diff: 3

No Slower results for the provided threshold = 1% and noise filter = 0.5 ns.

Fasterbase/diffBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "ascii")1.0781.6876.09
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "utf-8")1.05174.54165.92
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "ascii")1.02148.66145.48

SSE

summary:
better: 5, geomean: 1.065
total diff: 5

No Slower results for the provided threshold = 1% and noise filter = 0.5 ns.

Fasterbase/diffBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "ascii")1.11108.0797.21
System.Text.Tests.Perf_Encoding.GetChars(size: 16, encName: "ascii")1.1017.7416.14
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "ascii")1.05183.11173.62
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "utf-8")1.04219.66211.76
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "utf-8")1.02119.36116.59

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
@neon-sunset

neon-sunset commented Aug 3, 2023

Copy link
Copy Markdown
Contributor

How does this change affect *-arm64 targets? The PR introduces unconditional .ExtractMostSignificantBits() which is expected to regress its performance.

UPD: @Ruihan-Yin thank you

@Ruihan-Yin

Ruihan-Yin commented Aug 4, 2023

Copy link
Copy Markdown
MemberAuthor

How does this change affect *-arm64 targets? The PR introduces unconditional .ExtractMostSignificantBits() which is expected to regress its performance.

Thanks for pointing out. That was a mistake, changed to VectorContainsNonAsciiChar to make sure Arm64 won't be affected by the use of ExtractMSB.

@Ruihan-Yin

Copy link
Copy Markdown
MemberAuthor

Hi @xtqqczze@neon-sunset, kindly asking if there is any further change needed?

@xtqqczze

xtqqczze commented Aug 9, 2023

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte), as System.Text.Encoding.GetChars is a public API.

I'm not sure how to handle this, but System.SpanHelpers.IndexOfNullCharacter(System.Char*) does the following check:

if(((int)searchSpace&1)!=0)
{
// Input isn't char aligned, we won't be able to align it to a Vector
}

@MichalPetryka

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte), as System.Text.Encoding.GetChars is a public API.

AFAIR the guideline is that if the platform the code is running on works with unaligned memory, dotnet APIs should work too.

@xtqqczze

Copy link
Copy Markdown
Contributor

AFAIR the guideline is that if the platform the code is running on works with unaligned memory, dotnet APIs should work too.

@MichalPetryka We have existing code that does not account for this, e.g. WidenLatin1ToUtf16, see #90319.

@anthonycanino

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte), as System.Text.Encoding.GetChars is a public API.

I'm not sure how to handle this, but System.SpanHelpers.IndexOfNullCharacter(System.Char*) does the following check:

if(((int)searchSpace&1)!=0)
{
// Input isn't char aligned, we won't be able to align it to a Vector
}

@tannergooding based on some of the conversation we had, I am not sure how to proceed.

Perhaps we want to implement a fallback method that does not pin, and does not explicitly use StoreAligned and instead uses Store but skips the initial alignment if the char pointer is not naturally aligned?

@xtqqczze

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte)

See also comments at #90319.

@eiriktsarpaliseiriktsarpalis added this to the 9.0.0 milestone Aug 14, 2023
@tannergooding

tannergooding commented Aug 14, 2023

Copy link
Copy Markdown
Member

Perhaps we want to implement a fallback method that does not pin, and does not explicitly use StoreAligned and instead uses Store but skips the initial alignment if the char pointer is not naturally aligned?

For right now we should keep the pin and try to align, but if its unalignable then we should just continue as-is. It's basically the same code just with a check for "is this alignable at all" and using Store rather than StoreAligned (this is ultimately the same codegen and perf on modern hardware since the underlying data will actually be aligned in most cases).

For the future, we probably want to come to an agreement about how to universally handle this. My guess/vote is that's probably going to involve not pinning and using StoreUnsafe with optimistic alignment of the underlying data (basically optimistically presuming it won't be moved by the GC, which will be the common case but still using ref so that if it is moved, everything still works as expected). There's ultimately a balance between writing safe/readable code and performant code, so we need to find the right spot to land. Ideally we'd just be using Vector128.Create(ROSpan<T>) and the JIT would elide the bounds checks, so we don't have any unsafeness; but that's not possible today.

@eiriktsarpalis

Copy link
Copy Markdown
Member

Just checking up on the status of this PR, @anthonycanino have you had the chance to address @tannergooding's feedback?

@anthonycanino

Copy link
Copy Markdown
Contributor

Just checking up on the status of this PR, @anthonycanino have you had the chance to address @tannergooding's feedback?

@tannergooding how do we feel about this change now that we are moving to ISimdVector? Is it better to address the alignment with that change in one PR?

@tannergooding

Copy link
Copy Markdown
Member

I don't think it's changed from my last feedback.

We really need to account for unalignable data and the easiest way to do that is to pin, check the alignment, align if possible, and then continue processing using Store which works regardless of whether the data is aligned or unaligned.

In general the code pattern for efficiently handling alignment efficiently, for idempotent data, looks something like https://source.dot.net/#System.Numerics.Tensors/System/Numerics/Tensors/TensorPrimitives.netcore.cs,2930

Using ISimdVector then lets you merge the 3 different vectorized code paths down to 1 shared code path. The amount of unrolling and other factors can depend on the exact algorithm, the number of inputs, etc. But this is the general basic shape that works well for both large and small inputs.

@tannergoodingtannergooding added the needs-author-action An issue or pull request that requires more info or actions from the author. label Nov 3, 2023
@eiriktsarpaliseiriktsarpalis self-assigned this Nov 15, 2023
@ghost

Copy link
Copy Markdown

This pull request has been automatically marked no-recent-activity because it has not had any activity for 14 days. It will be closed if no further activity occurs within 14 more days. Any new comment (by anyone, not necessarily the author) will remove no-recent-activity.

@ghost

Copy link
Copy Markdown

This pull request will now be closed since it had been marked no-recent-activity but received no further activity in the past 14 days. It is still possible to reopen or comment on the pull request, but please note that it will be locked if it remains inactive for another 30 days.

@ghostghost closed this Dec 13, 2023
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jan 13, 2024
This pull request was closed.
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-System.Text.Encodingcommunity-contributionIndicates that the PR has been added by a community memberneeds-author-actionAn issue or pull request that requires more info or actions from the author.no-recent-activity

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@Ruihan-Yin@neon-sunset@xtqqczze@MichalPetryka@anthonycanino@tannergooding@eiriktsarpalis
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use StoreAligned not Store in WidenAsciiToUtf16 - #89892

Closed
Ruihan-Yin wants to merge 3 commits into
dotnet:mainfrom
Ruihan-Yin:WriteAlign
Closed

Use StoreAligned not Store in WidenAsciiToUtf16#89892
Ruihan-Yin wants to merge 3 commits into
dotnet:mainfrom
Ruihan-Yin:WriteAlign

Conversation

@Ruihan-Yin

Copy link
Copy Markdown
Member

Description

This PR is to improve the in-loop write logic in WidenAsciiToUtf16, the major change is to replace StoreAligned with Store inside the loop to reduce the penalty caused by split loads.

We are open to adjusting the implementation style for either path. Perf number attached in the comments.

@ghostghost added community-contribution Indicates that the PR has been added by a community member area-System.Text.Encoding labels Aug 2, 2023
@ghost

ghost commented Aug 2, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/area-system-text-encoding
See info in area-owners.md if you want to be subscribed.

Issue Details

Description

This PR is to improve the in-loop write logic in WidenAsciiToUtf16, the major change is to replace StoreAligned with Store inside the loop to reduce the penalty caused by split loads.

We are open to adjusting the implementation style for either path. Perf number attached in the comments.

Author:Ruihan-Yin
Assignees:-
Labels:

area-System.Text.Encoding, community-contribution

Milestone:-

@Ruihan-Yin

Copy link
Copy Markdown
MemberAuthor

Perf numbers

Base: main (531ad95)
Diff: main + changes on WidenAsciiToUtf16

Avx512

summary:
better: 5, geomean: 1.099
worse: 1, geomean: 1.022
total diff: 6

Slowerdiff/baseBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "utf-8")1.0285.2587.11
Fasterbase/diffBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "utf-8")1.18176.81150.47
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "ascii")1.14152.16133.17
System.Text.Tests.Perf_Encoding.GetChars(size: 16, encName: "ascii")1.0917.9716.54
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "ascii")1.0679.4774.72
System.Text.Tests.Perf_Encoding.GetChars(size: 16, encName: "utf-8")1.0325.7924.94

AVX

summary:
better: 3, geomean: 1.049
total diff: 3

No Slower results for the provided threshold = 1% and noise filter = 0.5 ns.

Fasterbase/diffBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "ascii")1.0781.6876.09
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "utf-8")1.05174.54165.92
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "ascii")1.02148.66145.48

SSE

summary:
better: 5, geomean: 1.065
total diff: 5

No Slower results for the provided threshold = 1% and noise filter = 0.5 ns.

Fasterbase/diffBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "ascii")1.11108.0797.21
System.Text.Tests.Perf_Encoding.GetChars(size: 16, encName: "ascii")1.1017.7416.14
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "ascii")1.05183.11173.62
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "utf-8")1.04219.66211.76
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "utf-8")1.02119.36116.59

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
@neon-sunset

neon-sunset commented Aug 3, 2023

Copy link
Copy Markdown
Contributor

How does this change affect *-arm64 targets? The PR introduces unconditional .ExtractMostSignificantBits() which is expected to regress its performance.

UPD: @Ruihan-Yin thank you

@Ruihan-Yin

Ruihan-Yin commented Aug 4, 2023

Copy link
Copy Markdown
MemberAuthor

How does this change affect *-arm64 targets? The PR introduces unconditional .ExtractMostSignificantBits() which is expected to regress its performance.

Thanks for pointing out. That was a mistake, changed to VectorContainsNonAsciiChar to make sure Arm64 won't be affected by the use of ExtractMSB.

@Ruihan-Yin

Copy link
Copy Markdown
MemberAuthor

Hi @xtqqczze@neon-sunset, kindly asking if there is any further change needed?

@xtqqczze

xtqqczze commented Aug 9, 2023

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte), as System.Text.Encoding.GetChars is a public API.

I'm not sure how to handle this, but System.SpanHelpers.IndexOfNullCharacter(System.Char*) does the following check:

if(((int)searchSpace&1)!=0)
{
// Input isn't char aligned, we won't be able to align it to a Vector
}

@MichalPetryka

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte), as System.Text.Encoding.GetChars is a public API.

AFAIR the guideline is that if the platform the code is running on works with unaligned memory, dotnet APIs should work too.

@xtqqczze

Copy link
Copy Markdown
Contributor

AFAIR the guideline is that if the platform the code is running on works with unaligned memory, dotnet APIs should work too.

@MichalPetryka We have existing code that does not account for this, e.g. WidenLatin1ToUtf16, see #90319.

@anthonycanino

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte), as System.Text.Encoding.GetChars is a public API.

I'm not sure how to handle this, but System.SpanHelpers.IndexOfNullCharacter(System.Char*) does the following check:

if(((int)searchSpace&1)!=0)
{
// Input isn't char aligned, we won't be able to align it to a Vector
}

@tannergooding based on some of the conversation we had, I am not sure how to proceed.

Perhaps we want to implement a fallback method that does not pin, and does not explicitly use StoreAligned and instead uses Store but skips the initial alignment if the char pointer is not naturally aligned?

@xtqqczze

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte)

See also comments at #90319.

@eiriktsarpaliseiriktsarpalis added this to the 9.0.0 milestone Aug 14, 2023
@tannergooding

tannergooding commented Aug 14, 2023

Copy link
Copy Markdown
Member

Perhaps we want to implement a fallback method that does not pin, and does not explicitly use StoreAligned and instead uses Store but skips the initial alignment if the char pointer is not naturally aligned?

For right now we should keep the pin and try to align, but if its unalignable then we should just continue as-is. It's basically the same code just with a check for "is this alignable at all" and using Store rather than StoreAligned (this is ultimately the same codegen and perf on modern hardware since the underlying data will actually be aligned in most cases).

For the future, we probably want to come to an agreement about how to universally handle this. My guess/vote is that's probably going to involve not pinning and using StoreUnsafe with optimistic alignment of the underlying data (basically optimistically presuming it won't be moved by the GC, which will be the common case but still using ref so that if it is moved, everything still works as expected). There's ultimately a balance between writing safe/readable code and performant code, so we need to find the right spot to land. Ideally we'd just be using Vector128.Create(ROSpan<T>) and the JIT would elide the bounds checks, so we don't have any unsafeness; but that's not possible today.

@eiriktsarpalis

Copy link
Copy Markdown
Member

Just checking up on the status of this PR, @anthonycanino have you had the chance to address @tannergooding's feedback?

@anthonycanino

Copy link
Copy Markdown
Contributor

Just checking up on the status of this PR, @anthonycanino have you had the chance to address @tannergooding's feedback?

@tannergooding how do we feel about this change now that we are moving to ISimdVector? Is it better to address the alignment with that change in one PR?

@tannergooding

Copy link
Copy Markdown
Member

I don't think it's changed from my last feedback.

We really need to account for unalignable data and the easiest way to do that is to pin, check the alignment, align if possible, and then continue processing using Store which works regardless of whether the data is aligned or unaligned.

In general the code pattern for efficiently handling alignment efficiently, for idempotent data, looks something like https://source.dot.net/#System.Numerics.Tensors/System/Numerics/Tensors/TensorPrimitives.netcore.cs,2930

Using ISimdVector then lets you merge the 3 different vectorized code paths down to 1 shared code path. The amount of unrolling and other factors can depend on the exact algorithm, the number of inputs, etc. But this is the general basic shape that works well for both large and small inputs.

@tannergoodingtannergooding added the needs-author-action An issue or pull request that requires more info or actions from the author. label Nov 3, 2023
@eiriktsarpaliseiriktsarpalis self-assigned this Nov 15, 2023
@ghost

Copy link
Copy Markdown

This pull request has been automatically marked no-recent-activity because it has not had any activity for 14 days. It will be closed if no further activity occurs within 14 more days. Any new comment (by anyone, not necessarily the author) will remove no-recent-activity.

@ghost

Copy link
Copy Markdown

This pull request will now be closed since it had been marked no-recent-activity but received no further activity in the past 14 days. It is still possible to reopen or comment on the pull request, but please note that it will be locked if it remains inactive for another 30 days.

@ghostghost closed this Dec 13, 2023
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jan 13, 2024
This pull request was closed.
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-System.Text.Encodingcommunity-contributionIndicates that the PR has been added by a community memberneeds-author-actionAn issue or pull request that requires more info or actions from the author.no-recent-activity

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@Ruihan-Yin@neon-sunset@xtqqczze@MichalPetryka@anthonycanino@tannergooding@eiriktsarpalis
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Use StoreAligned not Store in WidenAsciiToUtf16 - #89892

Closed
Ruihan-Yin wants to merge 3 commits into
dotnet:mainfrom
Ruihan-Yin:WriteAlign
Closed

Use StoreAligned not Store in WidenAsciiToUtf16#89892
Ruihan-Yin wants to merge 3 commits into
dotnet:mainfrom
Ruihan-Yin:WriteAlign

Conversation

@Ruihan-Yin

Copy link
Copy Markdown
Member

Description

This PR is to improve the in-loop write logic in WidenAsciiToUtf16, the major change is to replace StoreAligned with Store inside the loop to reduce the penalty caused by split loads.

We are open to adjusting the implementation style for either path. Perf number attached in the comments.

@ghostghost added community-contribution Indicates that the PR has been added by a community member area-System.Text.Encoding labels Aug 2, 2023
@ghost

ghost commented Aug 2, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/area-system-text-encoding
See info in area-owners.md if you want to be subscribed.

Issue Details

Description

This PR is to improve the in-loop write logic in WidenAsciiToUtf16, the major change is to replace StoreAligned with Store inside the loop to reduce the penalty caused by split loads.

We are open to adjusting the implementation style for either path. Perf number attached in the comments.

Author:Ruihan-Yin
Assignees:-
Labels:

area-System.Text.Encoding, community-contribution

Milestone:-

@Ruihan-Yin

Copy link
Copy Markdown
MemberAuthor

Perf numbers

Base: main (531ad95)
Diff: main + changes on WidenAsciiToUtf16

Avx512

summary:
better: 5, geomean: 1.099
worse: 1, geomean: 1.022
total diff: 6

Slowerdiff/baseBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "utf-8")1.0285.2587.11
Fasterbase/diffBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "utf-8")1.18176.81150.47
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "ascii")1.14152.16133.17
System.Text.Tests.Perf_Encoding.GetChars(size: 16, encName: "ascii")1.0917.9716.54
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "ascii")1.0679.4774.72
System.Text.Tests.Perf_Encoding.GetChars(size: 16, encName: "utf-8")1.0325.7924.94

AVX

summary:
better: 3, geomean: 1.049
total diff: 3

No Slower results for the provided threshold = 1% and noise filter = 0.5 ns.

Fasterbase/diffBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "ascii")1.0781.6876.09
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "utf-8")1.05174.54165.92
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "ascii")1.02148.66145.48

SSE

summary:
better: 5, geomean: 1.065
total diff: 5

No Slower results for the provided threshold = 1% and noise filter = 0.5 ns.

Fasterbase/diffBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "ascii")1.11108.0797.21
System.Text.Tests.Perf_Encoding.GetChars(size: 16, encName: "ascii")1.1017.7416.14
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "ascii")1.05183.11173.62
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "utf-8")1.04219.66211.76
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "utf-8")1.02119.36116.59

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
@neon-sunset

neon-sunset commented Aug 3, 2023

Copy link
Copy Markdown
Contributor

How does this change affect *-arm64 targets? The PR introduces unconditional .ExtractMostSignificantBits() which is expected to regress its performance.

UPD: @Ruihan-Yin thank you

@Ruihan-Yin

Ruihan-Yin commented Aug 4, 2023

Copy link
Copy Markdown
MemberAuthor

How does this change affect *-arm64 targets? The PR introduces unconditional .ExtractMostSignificantBits() which is expected to regress its performance.

Thanks for pointing out. That was a mistake, changed to VectorContainsNonAsciiChar to make sure Arm64 won't be affected by the use of ExtractMSB.

@Ruihan-Yin

Copy link
Copy Markdown
MemberAuthor

Hi @xtqqczze@neon-sunset, kindly asking if there is any further change needed?

@xtqqczze

xtqqczze commented Aug 9, 2023

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte), as System.Text.Encoding.GetChars is a public API.

I'm not sure how to handle this, but System.SpanHelpers.IndexOfNullCharacter(System.Char*) does the following check:

if(((int)searchSpace&1)!=0)
{
// Input isn't char aligned, we won't be able to align it to a Vector
}

@MichalPetryka

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte), as System.Text.Encoding.GetChars is a public API.

AFAIR the guideline is that if the platform the code is running on works with unaligned memory, dotnet APIs should work too.

@xtqqczze

Copy link
Copy Markdown
Contributor

AFAIR the guideline is that if the platform the code is running on works with unaligned memory, dotnet APIs should work too.

@MichalPetryka We have existing code that does not account for this, e.g. WidenLatin1ToUtf16, see #90319.

@anthonycanino

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte), as System.Text.Encoding.GetChars is a public API.

I'm not sure how to handle this, but System.SpanHelpers.IndexOfNullCharacter(System.Char*) does the following check:

if(((int)searchSpace&1)!=0)
{
// Input isn't char aligned, we won't be able to align it to a Vector
}

@tannergooding based on some of the conversation we had, I am not sure how to proceed.

Perhaps we want to implement a fallback method that does not pin, and does not explicitly use StoreAligned and instead uses Store but skips the initial alignment if the char pointer is not naturally aligned?

@xtqqczze

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte)

See also comments at #90319.

@eiriktsarpaliseiriktsarpalis added this to the 9.0.0 milestone Aug 14, 2023
@tannergooding

tannergooding commented Aug 14, 2023

Copy link
Copy Markdown
Member

Perhaps we want to implement a fallback method that does not pin, and does not explicitly use StoreAligned and instead uses Store but skips the initial alignment if the char pointer is not naturally aligned?

For right now we should keep the pin and try to align, but if its unalignable then we should just continue as-is. It's basically the same code just with a check for "is this alignable at all" and using Store rather than StoreAligned (this is ultimately the same codegen and perf on modern hardware since the underlying data will actually be aligned in most cases).

For the future, we probably want to come to an agreement about how to universally handle this. My guess/vote is that's probably going to involve not pinning and using StoreUnsafe with optimistic alignment of the underlying data (basically optimistically presuming it won't be moved by the GC, which will be the common case but still using ref so that if it is moved, everything still works as expected). There's ultimately a balance between writing safe/readable code and performant code, so we need to find the right spot to land. Ideally we'd just be using Vector128.Create(ROSpan<T>) and the JIT would elide the bounds checks, so we don't have any unsafeness; but that's not possible today.

@eiriktsarpalis

Copy link
Copy Markdown
Member

Just checking up on the status of this PR, @anthonycanino have you had the chance to address @tannergooding's feedback?

@anthonycanino

Copy link
Copy Markdown
Contributor

Just checking up on the status of this PR, @anthonycanino have you had the chance to address @tannergooding's feedback?

@tannergooding how do we feel about this change now that we are moving to ISimdVector? Is it better to address the alignment with that change in one PR?

@tannergooding

Copy link
Copy Markdown
Member

I don't think it's changed from my last feedback.

We really need to account for unalignable data and the easiest way to do that is to pin, check the alignment, align if possible, and then continue processing using Store which works regardless of whether the data is aligned or unaligned.

In general the code pattern for efficiently handling alignment efficiently, for idempotent data, looks something like https://source.dot.net/#System.Numerics.Tensors/System/Numerics/Tensors/TensorPrimitives.netcore.cs,2930

Using ISimdVector then lets you merge the 3 different vectorized code paths down to 1 shared code path. The amount of unrolling and other factors can depend on the exact algorithm, the number of inputs, etc. But this is the general basic shape that works well for both large and small inputs.

@tannergoodingtannergooding added the needs-author-action An issue or pull request that requires more info or actions from the author. label Nov 3, 2023
@eiriktsarpaliseiriktsarpalis self-assigned this Nov 15, 2023
@ghost

Copy link
Copy Markdown

This pull request has been automatically marked no-recent-activity because it has not had any activity for 14 days. It will be closed if no further activity occurs within 14 more days. Any new comment (by anyone, not necessarily the author) will remove no-recent-activity.

@ghost

Copy link
Copy Markdown

This pull request will now be closed since it had been marked no-recent-activity but received no further activity in the past 14 days. It is still possible to reopen or comment on the pull request, but please note that it will be locked if it remains inactive for another 30 days.

@ghostghost closed this Dec 13, 2023
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jan 13, 2024
This pull request was closed.
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-System.Text.Encodingcommunity-contributionIndicates that the PR has been added by a community memberneeds-author-actionAn issue or pull request that requires more info or actions from the author.no-recent-activity

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@Ruihan-Yin@neon-sunset@xtqqczze@MichalPetryka@anthonycanino@tannergooding@eiriktsarpalis
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use StoreAligned not Store in WidenAsciiToUtf16 - #89892

Closed
Ruihan-Yin wants to merge 3 commits into
dotnet:mainfrom
Ruihan-Yin:WriteAlign
Closed

Use StoreAligned not Store in WidenAsciiToUtf16#89892
Ruihan-Yin wants to merge 3 commits into
dotnet:mainfrom
Ruihan-Yin:WriteAlign

Conversation

@Ruihan-Yin

Copy link
Copy Markdown
Member

Description

This PR is to improve the in-loop write logic in WidenAsciiToUtf16, the major change is to replace StoreAligned with Store inside the loop to reduce the penalty caused by split loads.

We are open to adjusting the implementation style for either path. Perf number attached in the comments.

@ghostghost added community-contribution Indicates that the PR has been added by a community member area-System.Text.Encoding labels Aug 2, 2023
@ghost

ghost commented Aug 2, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/area-system-text-encoding
See info in area-owners.md if you want to be subscribed.

Issue Details

Description

This PR is to improve the in-loop write logic in WidenAsciiToUtf16, the major change is to replace StoreAligned with Store inside the loop to reduce the penalty caused by split loads.

We are open to adjusting the implementation style for either path. Perf number attached in the comments.

Author:Ruihan-Yin
Assignees:-
Labels:

area-System.Text.Encoding, community-contribution

Milestone:-

@Ruihan-Yin

Copy link
Copy Markdown
MemberAuthor

Perf numbers

Base: main (531ad95)
Diff: main + changes on WidenAsciiToUtf16

Avx512

summary:
better: 5, geomean: 1.099
worse: 1, geomean: 1.022
total diff: 6

Slowerdiff/baseBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "utf-8")1.0285.2587.11
Fasterbase/diffBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "utf-8")1.18176.81150.47
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "ascii")1.14152.16133.17
System.Text.Tests.Perf_Encoding.GetChars(size: 16, encName: "ascii")1.0917.9716.54
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "ascii")1.0679.4774.72
System.Text.Tests.Perf_Encoding.GetChars(size: 16, encName: "utf-8")1.0325.7924.94

AVX

summary:
better: 3, geomean: 1.049
total diff: 3

No Slower results for the provided threshold = 1% and noise filter = 0.5 ns.

Fasterbase/diffBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "ascii")1.0781.6876.09
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "utf-8")1.05174.54165.92
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "ascii")1.02148.66145.48

SSE

summary:
better: 5, geomean: 1.065
total diff: 5

No Slower results for the provided threshold = 1% and noise filter = 0.5 ns.

Fasterbase/diffBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "ascii")1.11108.0797.21
System.Text.Tests.Perf_Encoding.GetChars(size: 16, encName: "ascii")1.1017.7416.14
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "ascii")1.05183.11173.62
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "utf-8")1.04219.66211.76
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "utf-8")1.02119.36116.59

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
@neon-sunset

neon-sunset commented Aug 3, 2023

Copy link
Copy Markdown
Contributor

How does this change affect *-arm64 targets? The PR introduces unconditional .ExtractMostSignificantBits() which is expected to regress its performance.

UPD: @Ruihan-Yin thank you

@Ruihan-Yin

Ruihan-Yin commented Aug 4, 2023

Copy link
Copy Markdown
MemberAuthor

How does this change affect *-arm64 targets? The PR introduces unconditional .ExtractMostSignificantBits() which is expected to regress its performance.

Thanks for pointing out. That was a mistake, changed to VectorContainsNonAsciiChar to make sure Arm64 won't be affected by the use of ExtractMSB.

@Ruihan-Yin

Copy link
Copy Markdown
MemberAuthor

Hi @xtqqczze@neon-sunset, kindly asking if there is any further change needed?

@xtqqczze

xtqqczze commented Aug 9, 2023

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte), as System.Text.Encoding.GetChars is a public API.

I'm not sure how to handle this, but System.SpanHelpers.IndexOfNullCharacter(System.Char*) does the following check:

if(((int)searchSpace&1)!=0)
{
// Input isn't char aligned, we won't be able to align it to a Vector
}

@MichalPetryka

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte), as System.Text.Encoding.GetChars is a public API.

AFAIR the guideline is that if the platform the code is running on works with unaligned memory, dotnet APIs should work too.

@xtqqczze

Copy link
Copy Markdown
Contributor

AFAIR the guideline is that if the platform the code is running on works with unaligned memory, dotnet APIs should work too.

@MichalPetryka We have existing code that does not account for this, e.g. WidenLatin1ToUtf16, see #90319.

@anthonycanino

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte), as System.Text.Encoding.GetChars is a public API.

I'm not sure how to handle this, but System.SpanHelpers.IndexOfNullCharacter(System.Char*) does the following check:

if(((int)searchSpace&1)!=0)
{
// Input isn't char aligned, we won't be able to align it to a Vector
}

@tannergooding based on some of the conversation we had, I am not sure how to proceed.

Perhaps we want to implement a fallback method that does not pin, and does not explicitly use StoreAligned and instead uses Store but skips the initial alignment if the char pointer is not naturally aligned?

@xtqqczze

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte)

See also comments at #90319.

@eiriktsarpaliseiriktsarpalis added this to the 9.0.0 milestone Aug 14, 2023
@tannergooding

tannergooding commented Aug 14, 2023

Copy link
Copy Markdown
Member

Perhaps we want to implement a fallback method that does not pin, and does not explicitly use StoreAligned and instead uses Store but skips the initial alignment if the char pointer is not naturally aligned?

For right now we should keep the pin and try to align, but if its unalignable then we should just continue as-is. It's basically the same code just with a check for "is this alignable at all" and using Store rather than StoreAligned (this is ultimately the same codegen and perf on modern hardware since the underlying data will actually be aligned in most cases).

For the future, we probably want to come to an agreement about how to universally handle this. My guess/vote is that's probably going to involve not pinning and using StoreUnsafe with optimistic alignment of the underlying data (basically optimistically presuming it won't be moved by the GC, which will be the common case but still using ref so that if it is moved, everything still works as expected). There's ultimately a balance between writing safe/readable code and performant code, so we need to find the right spot to land. Ideally we'd just be using Vector128.Create(ROSpan<T>) and the JIT would elide the bounds checks, so we don't have any unsafeness; but that's not possible today.

@eiriktsarpalis

Copy link
Copy Markdown
Member

Just checking up on the status of this PR, @anthonycanino have you had the chance to address @tannergooding's feedback?

@anthonycanino

Copy link
Copy Markdown
Contributor

Just checking up on the status of this PR, @anthonycanino have you had the chance to address @tannergooding's feedback?

@tannergooding how do we feel about this change now that we are moving to ISimdVector? Is it better to address the alignment with that change in one PR?

@tannergooding

Copy link
Copy Markdown
Member

I don't think it's changed from my last feedback.

We really need to account for unalignable data and the easiest way to do that is to pin, check the alignment, align if possible, and then continue processing using Store which works regardless of whether the data is aligned or unaligned.

In general the code pattern for efficiently handling alignment efficiently, for idempotent data, looks something like https://source.dot.net/#System.Numerics.Tensors/System/Numerics/Tensors/TensorPrimitives.netcore.cs,2930

Using ISimdVector then lets you merge the 3 different vectorized code paths down to 1 shared code path. The amount of unrolling and other factors can depend on the exact algorithm, the number of inputs, etc. But this is the general basic shape that works well for both large and small inputs.

@tannergoodingtannergooding added the needs-author-action An issue or pull request that requires more info or actions from the author. label Nov 3, 2023
@eiriktsarpaliseiriktsarpalis self-assigned this Nov 15, 2023
@ghost

Copy link
Copy Markdown

This pull request has been automatically marked no-recent-activity because it has not had any activity for 14 days. It will be closed if no further activity occurs within 14 more days. Any new comment (by anyone, not necessarily the author) will remove no-recent-activity.

@ghost

Copy link
Copy Markdown

This pull request will now be closed since it had been marked no-recent-activity but received no further activity in the past 14 days. It is still possible to reopen or comment on the pull request, but please note that it will be locked if it remains inactive for another 30 days.

@ghostghost closed this Dec 13, 2023
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jan 13, 2024
This pull request was closed.
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-System.Text.Encodingcommunity-contributionIndicates that the PR has been added by a community memberneeds-author-actionAn issue or pull request that requires more info or actions from the author.no-recent-activity

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@Ruihan-Yin@neon-sunset@xtqqczze@MichalPetryka@anthonycanino@tannergooding@eiriktsarpalis
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use StoreAligned not Store in WidenAsciiToUtf16 - #89892

Closed
Ruihan-Yin wants to merge 3 commits into
dotnet:mainfrom
Ruihan-Yin:WriteAlign
Closed

Use StoreAligned not Store in WidenAsciiToUtf16#89892
Ruihan-Yin wants to merge 3 commits into
dotnet:mainfrom
Ruihan-Yin:WriteAlign

Conversation

@Ruihan-Yin

Copy link
Copy Markdown
Member

Description

This PR is to improve the in-loop write logic in WidenAsciiToUtf16, the major change is to replace StoreAligned with Store inside the loop to reduce the penalty caused by split loads.

We are open to adjusting the implementation style for either path. Perf number attached in the comments.

@ghostghost added community-contribution Indicates that the PR has been added by a community member area-System.Text.Encoding labels Aug 2, 2023
@ghost

ghost commented Aug 2, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/area-system-text-encoding
See info in area-owners.md if you want to be subscribed.

Issue Details

Description

This PR is to improve the in-loop write logic in WidenAsciiToUtf16, the major change is to replace StoreAligned with Store inside the loop to reduce the penalty caused by split loads.

We are open to adjusting the implementation style for either path. Perf number attached in the comments.

Author:Ruihan-Yin
Assignees:-
Labels:

area-System.Text.Encoding, community-contribution

Milestone:-

@Ruihan-Yin

Copy link
Copy Markdown
MemberAuthor

Perf numbers

Base: main (531ad95)
Diff: main + changes on WidenAsciiToUtf16

Avx512

summary:
better: 5, geomean: 1.099
worse: 1, geomean: 1.022
total diff: 6

Slowerdiff/baseBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "utf-8")1.0285.2587.11
Fasterbase/diffBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "utf-8")1.18176.81150.47
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "ascii")1.14152.16133.17
System.Text.Tests.Perf_Encoding.GetChars(size: 16, encName: "ascii")1.0917.9716.54
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "ascii")1.0679.4774.72
System.Text.Tests.Perf_Encoding.GetChars(size: 16, encName: "utf-8")1.0325.7924.94

AVX

summary:
better: 3, geomean: 1.049
total diff: 3

No Slower results for the provided threshold = 1% and noise filter = 0.5 ns.

Fasterbase/diffBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "ascii")1.0781.6876.09
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "utf-8")1.05174.54165.92
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "ascii")1.02148.66145.48

SSE

summary:
better: 5, geomean: 1.065
total diff: 5

No Slower results for the provided threshold = 1% and noise filter = 0.5 ns.

Fasterbase/diffBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "ascii")1.11108.0797.21
System.Text.Tests.Perf_Encoding.GetChars(size: 16, encName: "ascii")1.1017.7416.14
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "ascii")1.05183.11173.62
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "utf-8")1.04219.66211.76
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "utf-8")1.02119.36116.59

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
@neon-sunset

neon-sunset commented Aug 3, 2023

Copy link
Copy Markdown
Contributor

How does this change affect *-arm64 targets? The PR introduces unconditional .ExtractMostSignificantBits() which is expected to regress its performance.

UPD: @Ruihan-Yin thank you

@Ruihan-Yin

Ruihan-Yin commented Aug 4, 2023

Copy link
Copy Markdown
MemberAuthor

How does this change affect *-arm64 targets? The PR introduces unconditional .ExtractMostSignificantBits() which is expected to regress its performance.

Thanks for pointing out. That was a mistake, changed to VectorContainsNonAsciiChar to make sure Arm64 won't be affected by the use of ExtractMSB.

@Ruihan-Yin

Copy link
Copy Markdown
MemberAuthor

Hi @xtqqczze@neon-sunset, kindly asking if there is any further change needed?

@xtqqczze

xtqqczze commented Aug 9, 2023

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte), as System.Text.Encoding.GetChars is a public API.

I'm not sure how to handle this, but System.SpanHelpers.IndexOfNullCharacter(System.Char*) does the following check:

if(((int)searchSpace&1)!=0)
{
// Input isn't char aligned, we won't be able to align it to a Vector
}

@MichalPetryka

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte), as System.Text.Encoding.GetChars is a public API.

AFAIR the guideline is that if the platform the code is running on works with unaligned memory, dotnet APIs should work too.

@xtqqczze

Copy link
Copy Markdown
Contributor

AFAIR the guideline is that if the platform the code is running on works with unaligned memory, dotnet APIs should work too.

@MichalPetryka We have existing code that does not account for this, e.g. WidenLatin1ToUtf16, see #90319.

@anthonycanino

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte), as System.Text.Encoding.GetChars is a public API.

I'm not sure how to handle this, but System.SpanHelpers.IndexOfNullCharacter(System.Char*) does the following check:

if(((int)searchSpace&1)!=0)
{
// Input isn't char aligned, we won't be able to align it to a Vector
}

@tannergooding based on some of the conversation we had, I am not sure how to proceed.

Perhaps we want to implement a fallback method that does not pin, and does not explicitly use StoreAligned and instead uses Store but skips the initial alignment if the char pointer is not naturally aligned?

@xtqqczze

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte)

See also comments at #90319.

@eiriktsarpaliseiriktsarpalis added this to the 9.0.0 milestone Aug 14, 2023
@tannergooding

tannergooding commented Aug 14, 2023

Copy link
Copy Markdown
Member

Perhaps we want to implement a fallback method that does not pin, and does not explicitly use StoreAligned and instead uses Store but skips the initial alignment if the char pointer is not naturally aligned?

For right now we should keep the pin and try to align, but if its unalignable then we should just continue as-is. It's basically the same code just with a check for "is this alignable at all" and using Store rather than StoreAligned (this is ultimately the same codegen and perf on modern hardware since the underlying data will actually be aligned in most cases).

For the future, we probably want to come to an agreement about how to universally handle this. My guess/vote is that's probably going to involve not pinning and using StoreUnsafe with optimistic alignment of the underlying data (basically optimistically presuming it won't be moved by the GC, which will be the common case but still using ref so that if it is moved, everything still works as expected). There's ultimately a balance between writing safe/readable code and performant code, so we need to find the right spot to land. Ideally we'd just be using Vector128.Create(ROSpan<T>) and the JIT would elide the bounds checks, so we don't have any unsafeness; but that's not possible today.

@eiriktsarpalis

Copy link
Copy Markdown
Member

Just checking up on the status of this PR, @anthonycanino have you had the chance to address @tannergooding's feedback?

@anthonycanino

Copy link
Copy Markdown
Contributor

Just checking up on the status of this PR, @anthonycanino have you had the chance to address @tannergooding's feedback?

@tannergooding how do we feel about this change now that we are moving to ISimdVector? Is it better to address the alignment with that change in one PR?

@tannergooding

Copy link
Copy Markdown
Member

I don't think it's changed from my last feedback.

We really need to account for unalignable data and the easiest way to do that is to pin, check the alignment, align if possible, and then continue processing using Store which works regardless of whether the data is aligned or unaligned.

In general the code pattern for efficiently handling alignment efficiently, for idempotent data, looks something like https://source.dot.net/#System.Numerics.Tensors/System/Numerics/Tensors/TensorPrimitives.netcore.cs,2930

Using ISimdVector then lets you merge the 3 different vectorized code paths down to 1 shared code path. The amount of unrolling and other factors can depend on the exact algorithm, the number of inputs, etc. But this is the general basic shape that works well for both large and small inputs.

@tannergoodingtannergooding added the needs-author-action An issue or pull request that requires more info or actions from the author. label Nov 3, 2023
@eiriktsarpaliseiriktsarpalis self-assigned this Nov 15, 2023
@ghost

Copy link
Copy Markdown

This pull request has been automatically marked no-recent-activity because it has not had any activity for 14 days. It will be closed if no further activity occurs within 14 more days. Any new comment (by anyone, not necessarily the author) will remove no-recent-activity.

@ghost

Copy link
Copy Markdown

This pull request will now be closed since it had been marked no-recent-activity but received no further activity in the past 14 days. It is still possible to reopen or comment on the pull request, but please note that it will be locked if it remains inactive for another 30 days.

@ghostghost closed this Dec 13, 2023
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jan 13, 2024
This pull request was closed.
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-System.Text.Encodingcommunity-contributionIndicates that the PR has been added by a community memberneeds-author-actionAn issue or pull request that requires more info or actions from the author.no-recent-activity

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@Ruihan-Yin@neon-sunset@xtqqczze@MichalPetryka@anthonycanino@tannergooding@eiriktsarpalis
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Use StoreAligned not Store in WidenAsciiToUtf16 - #89892

Closed
Ruihan-Yin wants to merge 3 commits into
dotnet:mainfrom
Ruihan-Yin:WriteAlign
Closed

Use StoreAligned not Store in WidenAsciiToUtf16#89892
Ruihan-Yin wants to merge 3 commits into
dotnet:mainfrom
Ruihan-Yin:WriteAlign

Conversation

@Ruihan-Yin

Copy link
Copy Markdown
Member

Description

This PR is to improve the in-loop write logic in WidenAsciiToUtf16, the major change is to replace StoreAligned with Store inside the loop to reduce the penalty caused by split loads.

We are open to adjusting the implementation style for either path. Perf number attached in the comments.

@ghostghost added community-contribution Indicates that the PR has been added by a community member area-System.Text.Encoding labels Aug 2, 2023
@ghost

ghost commented Aug 2, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @dotnet/area-system-text-encoding
See info in area-owners.md if you want to be subscribed.

Issue Details

Description

This PR is to improve the in-loop write logic in WidenAsciiToUtf16, the major change is to replace StoreAligned with Store inside the loop to reduce the penalty caused by split loads.

We are open to adjusting the implementation style for either path. Perf number attached in the comments.

Author:Ruihan-Yin
Assignees:-
Labels:

area-System.Text.Encoding, community-contribution

Milestone:-

@Ruihan-Yin

Copy link
Copy Markdown
MemberAuthor

Perf numbers

Base: main (531ad95)
Diff: main + changes on WidenAsciiToUtf16

Avx512

summary:
better: 5, geomean: 1.099
worse: 1, geomean: 1.022
total diff: 6

Slowerdiff/baseBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "utf-8")1.0285.2587.11
Fasterbase/diffBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "utf-8")1.18176.81150.47
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "ascii")1.14152.16133.17
System.Text.Tests.Perf_Encoding.GetChars(size: 16, encName: "ascii")1.0917.9716.54
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "ascii")1.0679.4774.72
System.Text.Tests.Perf_Encoding.GetChars(size: 16, encName: "utf-8")1.0325.7924.94

AVX

summary:
better: 3, geomean: 1.049
total diff: 3

No Slower results for the provided threshold = 1% and noise filter = 0.5 ns.

Fasterbase/diffBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "ascii")1.0781.6876.09
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "utf-8")1.05174.54165.92
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "ascii")1.02148.66145.48

SSE

summary:
better: 5, geomean: 1.065
total diff: 5

No Slower results for the provided threshold = 1% and noise filter = 0.5 ns.

Fasterbase/diffBase Median (ns)Diff Median (ns)Modality
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "ascii")1.11108.0797.21
System.Text.Tests.Perf_Encoding.GetChars(size: 16, encName: "ascii")1.1017.7416.14
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "ascii")1.05183.11173.62
System.Text.Tests.Perf_Encoding.GetChars(size: 1024, encName: "utf-8")1.04219.66211.76
System.Text.Tests.Perf_Encoding.GetChars(size: 512, encName: "utf-8")1.02119.36116.59

Comment threadsrc/libraries/System.Private.CoreLib/src/System/Text/Ascii.Utility.cs Outdated
@neon-sunset

neon-sunset commented Aug 3, 2023

Copy link
Copy Markdown
Contributor

How does this change affect *-arm64 targets? The PR introduces unconditional .ExtractMostSignificantBits() which is expected to regress its performance.

UPD: @Ruihan-Yin thank you

@Ruihan-Yin

Ruihan-Yin commented Aug 4, 2023

Copy link
Copy Markdown
MemberAuthor

How does this change affect *-arm64 targets? The PR introduces unconditional .ExtractMostSignificantBits() which is expected to regress its performance.

Thanks for pointing out. That was a mistake, changed to VectorContainsNonAsciiChar to make sure Arm64 won't be affected by the use of ExtractMSB.

@Ruihan-Yin

Copy link
Copy Markdown
MemberAuthor

Hi @xtqqczze@neon-sunset, kindly asking if there is any further change needed?

@xtqqczze

xtqqczze commented Aug 9, 2023

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte), as System.Text.Encoding.GetChars is a public API.

I'm not sure how to handle this, but System.SpanHelpers.IndexOfNullCharacter(System.Char*) does the following check:

if(((int)searchSpace&1)!=0)
{
// Input isn't char aligned, we won't be able to align it to a Vector
}

@MichalPetryka

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte), as System.Text.Encoding.GetChars is a public API.

AFAIR the guideline is that if the platform the code is running on works with unaligned memory, dotnet APIs should work too.

@xtqqczze

Copy link
Copy Markdown
Contributor

AFAIR the guideline is that if the platform the code is running on works with unaligned memory, dotnet APIs should work too.

@MichalPetryka We have existing code that does not account for this, e.g. WidenLatin1ToUtf16, see #90319.

@anthonycanino

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte), as System.Text.Encoding.GetChars is a public API.

I'm not sure how to handle this, but System.SpanHelpers.IndexOfNullCharacter(System.Char*) does the following check:

if(((int)searchSpace&1)!=0)
{
// Input isn't char aligned, we won't be able to align it to a Vector
}

@tannergooding based on some of the conversation we had, I am not sure how to proceed.

Perhaps we want to implement a fallback method that does not pin, and does not explicitly use StoreAligned and instead uses Store but skips the initial alignment if the char pointer is not naturally aligned?

@xtqqczze

Copy link
Copy Markdown
Contributor

Perhaps we shouldn't assume the argument pUtf16Buffer is naturally aligned (i.e. 2 byte)

See also comments at #90319.

@eiriktsarpaliseiriktsarpalis added this to the 9.0.0 milestone Aug 14, 2023
@tannergooding

tannergooding commented Aug 14, 2023

Copy link
Copy Markdown
Member

Perhaps we want to implement a fallback method that does not pin, and does not explicitly use StoreAligned and instead uses Store but skips the initial alignment if the char pointer is not naturally aligned?

For right now we should keep the pin and try to align, but if its unalignable then we should just continue as-is. It's basically the same code just with a check for "is this alignable at all" and using Store rather than StoreAligned (this is ultimately the same codegen and perf on modern hardware since the underlying data will actually be aligned in most cases).

For the future, we probably want to come to an agreement about how to universally handle this. My guess/vote is that's probably going to involve not pinning and using StoreUnsafe with optimistic alignment of the underlying data (basically optimistically presuming it won't be moved by the GC, which will be the common case but still using ref so that if it is moved, everything still works as expected). There's ultimately a balance between writing safe/readable code and performant code, so we need to find the right spot to land. Ideally we'd just be using Vector128.Create(ROSpan<T>) and the JIT would elide the bounds checks, so we don't have any unsafeness; but that's not possible today.

@eiriktsarpalis

Copy link
Copy Markdown
Member

Just checking up on the status of this PR, @anthonycanino have you had the chance to address @tannergooding's feedback?

@anthonycanino

Copy link
Copy Markdown
Contributor

Just checking up on the status of this PR, @anthonycanino have you had the chance to address @tannergooding's feedback?

@tannergooding how do we feel about this change now that we are moving to ISimdVector? Is it better to address the alignment with that change in one PR?

@tannergooding

Copy link
Copy Markdown
Member

I don't think it's changed from my last feedback.

We really need to account for unalignable data and the easiest way to do that is to pin, check the alignment, align if possible, and then continue processing using Store which works regardless of whether the data is aligned or unaligned.

In general the code pattern for efficiently handling alignment efficiently, for idempotent data, looks something like https://source.dot.net/#System.Numerics.Tensors/System/Numerics/Tensors/TensorPrimitives.netcore.cs,2930

Using ISimdVector then lets you merge the 3 different vectorized code paths down to 1 shared code path. The amount of unrolling and other factors can depend on the exact algorithm, the number of inputs, etc. But this is the general basic shape that works well for both large and small inputs.

@tannergoodingtannergooding added the needs-author-action An issue or pull request that requires more info or actions from the author. label Nov 3, 2023
@eiriktsarpaliseiriktsarpalis self-assigned this Nov 15, 2023
@ghost

Copy link
Copy Markdown

This pull request has been automatically marked no-recent-activity because it has not had any activity for 14 days. It will be closed if no further activity occurs within 14 more days. Any new comment (by anyone, not necessarily the author) will remove no-recent-activity.

@ghost

Copy link
Copy Markdown

This pull request will now be closed since it had been marked no-recent-activity but received no further activity in the past 14 days. It is still possible to reopen or comment on the pull request, but please note that it will be locked if it remains inactive for another 30 days.

@ghostghost closed this Dec 13, 2023
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jan 13, 2024
This pull request was closed.
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-System.Text.Encodingcommunity-contributionIndicates that the PR has been added by a community memberneeds-author-actionAn issue or pull request that requires more info or actions from the author.no-recent-activity

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@Ruihan-Yin@neon-sunset@xtqqczze@MichalPetryka@anthonycanino@tannergooding@eiriktsarpalis