Skip to content

Fix "Cleanup some string operating functions" PR - #100451

Closed
AaronRobinsonMSFT wants to merge 3 commits into
dotnet:mainfrom
AaronRobinsonMSFT:unrevert_string_convert_pr
Closed

Fix "Cleanup some string operating functions" PR#100451
AaronRobinsonMSFT wants to merge 3 commits into
dotnet:mainfrom
AaronRobinsonMSFT:unrevert_string_convert_pr

Conversation

@AaronRobinsonMSFT

@AaronRobinsonMSFTAaronRobinsonMSFT commented Mar 29, 2024

Copy link
Copy Markdown
Member

Unrevert #96099

The revert of #96099 was done in #97264. Commits after the first fix troublesome locations.

See #97264 for details on what the original PR impacted.

/cc @tommcdon@hoyosjs@huoyaoyuan

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/inc/corhlprpriv.h Outdated
Comment threadsrc/coreclr/inc/corhlprpriv.h Outdated
Comment threadsrc/coreclr/inc/corhlprpriv.h Outdated
Use correct error code mechanism.
Comment threadsrc/coreclr/vm/methodtablebuilder.cpp
Comment threadsrc/coreclr/vm/methodtablebuilder.cpp
Comment threadsrc/coreclr/vm/classcompat.cpp
if (SUCCEEDED(hr))
{
s.Resize(length, REPRESENTATION_UTF8);
COUNT_T length = WszWideCharToMultiByte(CP_UTF8, 0, GetRawUnicode(), GetRawCount()+1, NULL, 0, NULL, NULL);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I assume that this used FString::Unicode_Utf8_Length to avoid some perf problems in WszWideCharToMultiByte. Is switching to WszWideCharToMultiByte going to introduce performance regression? (I am particularly worried about Windows.)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I will check here.

Regardless, in principle I would prefer to defer to system APIs that we can/should assume are optimized. Having these narrowly defined "optimizations" is a non-trivial tax for an already covoluted space like SString.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree that we should get rid of FString::*. We may want to replace it with a call to the minipal UTF8 methods for perf reasons like what the original change tried to do. I would expect that minipal UTF8 convertors are going to have significantly better perf than Windows OS WideCharToMultiByte.

system APIs that we can/should assume are optimized

Historically, Windows OS WideCharToMultiByte has been significantly slower for UTF8 than one would expect from a reasonably optimized implementation.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would static linking simdutf to the VM for such conversions make sense here?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We may want to replace it with a call to the minipal UTF8 methods for perf reasons like what the original change tried to do.

So that I can get behind. There is a small performance regression here. Let me replace this with the minipal APIs. I chose this direction to match the other API. I'd prefer symmetrical API usage in this case as it avoids confusion.

Would static linking simdutf to the VM for such conversions make sense here?

That is something we can discuss, but it not for this PR.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we can push the hotpaths to managed

That seems unlikely in this case. The use of SString is litered throughout the system and there are many places where it is going to be unnatural and difficult to call into managed code. We can try jumping out in some cases, but that isn't something I am inclined to do in this PR.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

My bar for these types of changes is to avoid perf regressions. (Of course, perf improvements are nice - but it is better to evaluate them separately if they come with tradeoffs.)

Neither the existing FString fast path code nor the WideCharToMultiByte implementation built into Windows use any vectorization tricks. It suggests that we should not need vectorization tricks to avoid the perf regression here.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The shape of the code in WideCharToMultiByte built into Windows is actually very similar to the corefx implementation and the minipal. It seems that the minipal implementation picked up some cruft alone the way that makes it quite a bit slower.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is interesting that we picked up minipal implementation for Windows + Unix in mono (and Unix in coreclr as before) in .NET 8 and haven't seen any report of regression. Maybe there is some low-hanging alignment issue which can be tweaked via cl switch?

@jkotasjkotasApr 12, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

picked up minipal implementation for Windows + Unix in mono

Mono does not get as much perf scrutiny as CoreCLR.

low-hanging alignment issue

I do not think it is that.

From a cursory look, I see two problems:

  • For smaller lengths, there is a fixed overhead. Adding a simple FString-like loop that deals with small ASCII-only strings as the first thing in the minipal_* methods should fix that.
  • For longer lengths, the code of the core look does not look great (e.g. it uses more registers than necessary). For example, it may help to change
    *pTarget= (unsigned char)ch;
    *(pTarget+1) = (unsigned char)(ch >> 16);
    pSrc+=4;
    *(pTarget+2) = (unsigned char)chc;
    *(pTarget+3) = (unsigned char)(chc >> 16);
    pTarget+=4;
    to
 *pTarget = (unsigned char)ch;
*(pTarget + 2) = (unsigned char)(chc);
pSrc += 4;
*(pTarget + 1) = (unsigned char)(ch >> 16);
*(pTarget + 3) = (unsigned char)(chc >> 16);
pTarget += 4;

@AaronRobinsonMSFT
AaronRobinsonMSFT marked this pull request as draft May 3, 2024 04:10
@am11am11 mentioned this pull request May 21, 2024
@dotnet-policy-servicedotnet-policy-serviceBot removed this from the 9.0.0 milestone Jun 2, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Draft Pull Request was automatically closed for 30 days of inactivity. Please let us know if you'd like to reopen it.

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jul 2, 2024
@AaronRobinsonMSFT
AaronRobinsonMSFT deleted the unrevert_string_convert_pr branch November 10, 2025 18:14
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@AaronRobinsonMSFT@am11@jkotas@MichalPetryka
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Fix "Cleanup some string operating functions" PR by AaronRobinsonMSFT · Pull Request #100451 · dotnet/runtime · GitHub
Skip to content

Fix "Cleanup some string operating functions" PR - #100451

Closed
AaronRobinsonMSFT wants to merge 3 commits into
dotnet:mainfrom
AaronRobinsonMSFT:unrevert_string_convert_pr
Closed

Fix "Cleanup some string operating functions" PR#100451
AaronRobinsonMSFT wants to merge 3 commits into
dotnet:mainfrom
AaronRobinsonMSFT:unrevert_string_convert_pr

Conversation

@AaronRobinsonMSFT

@AaronRobinsonMSFTAaronRobinsonMSFT commented Mar 29, 2024

Copy link
Copy Markdown
Member

Unrevert #96099

The revert of #96099 was done in #97264. Commits after the first fix troublesome locations.

See #97264 for details on what the original PR impacted.

/cc @tommcdon@hoyosjs@huoyaoyuan

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/inc/corhlprpriv.h Outdated
Comment threadsrc/coreclr/inc/corhlprpriv.h Outdated
Comment threadsrc/coreclr/inc/corhlprpriv.h Outdated
Use correct error code mechanism.
Comment threadsrc/coreclr/vm/methodtablebuilder.cpp
Comment threadsrc/coreclr/vm/methodtablebuilder.cpp
Comment threadsrc/coreclr/vm/classcompat.cpp
if (SUCCEEDED(hr))
{
s.Resize(length, REPRESENTATION_UTF8);
COUNT_T length = WszWideCharToMultiByte(CP_UTF8, 0, GetRawUnicode(), GetRawCount()+1, NULL, 0, NULL, NULL);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I assume that this used FString::Unicode_Utf8_Length to avoid some perf problems in WszWideCharToMultiByte. Is switching to WszWideCharToMultiByte going to introduce performance regression? (I am particularly worried about Windows.)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I will check here.

Regardless, in principle I would prefer to defer to system APIs that we can/should assume are optimized. Having these narrowly defined "optimizations" is a non-trivial tax for an already covoluted space like SString.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree that we should get rid of FString::*. We may want to replace it with a call to the minipal UTF8 methods for perf reasons like what the original change tried to do. I would expect that minipal UTF8 convertors are going to have significantly better perf than Windows OS WideCharToMultiByte.

system APIs that we can/should assume are optimized

Historically, Windows OS WideCharToMultiByte has been significantly slower for UTF8 than one would expect from a reasonably optimized implementation.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would static linking simdutf to the VM for such conversions make sense here?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We may want to replace it with a call to the minipal UTF8 methods for perf reasons like what the original change tried to do.

So that I can get behind. There is a small performance regression here. Let me replace this with the minipal APIs. I chose this direction to match the other API. I'd prefer symmetrical API usage in this case as it avoids confusion.

Would static linking simdutf to the VM for such conversions make sense here?

That is something we can discuss, but it not for this PR.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we can push the hotpaths to managed

That seems unlikely in this case. The use of SString is litered throughout the system and there are many places where it is going to be unnatural and difficult to call into managed code. We can try jumping out in some cases, but that isn't something I am inclined to do in this PR.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

My bar for these types of changes is to avoid perf regressions. (Of course, perf improvements are nice - but it is better to evaluate them separately if they come with tradeoffs.)

Neither the existing FString fast path code nor the WideCharToMultiByte implementation built into Windows use any vectorization tricks. It suggests that we should not need vectorization tricks to avoid the perf regression here.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The shape of the code in WideCharToMultiByte built into Windows is actually very similar to the corefx implementation and the minipal. It seems that the minipal implementation picked up some cruft alone the way that makes it quite a bit slower.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is interesting that we picked up minipal implementation for Windows + Unix in mono (and Unix in coreclr as before) in .NET 8 and haven't seen any report of regression. Maybe there is some low-hanging alignment issue which can be tweaked via cl switch?

@jkotasjkotasApr 12, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

picked up minipal implementation for Windows + Unix in mono

Mono does not get as much perf scrutiny as CoreCLR.

low-hanging alignment issue

I do not think it is that.

From a cursory look, I see two problems:

  • For smaller lengths, there is a fixed overhead. Adding a simple FString-like loop that deals with small ASCII-only strings as the first thing in the minipal_* methods should fix that.
  • For longer lengths, the code of the core look does not look great (e.g. it uses more registers than necessary). For example, it may help to change
    *pTarget= (unsigned char)ch;
    *(pTarget+1) = (unsigned char)(ch >> 16);
    pSrc+=4;
    *(pTarget+2) = (unsigned char)chc;
    *(pTarget+3) = (unsigned char)(chc >> 16);
    pTarget+=4;
    to
 *pTarget = (unsigned char)ch;
*(pTarget + 2) = (unsigned char)(chc);
pSrc += 4;
*(pTarget + 1) = (unsigned char)(ch >> 16);
*(pTarget + 3) = (unsigned char)(chc >> 16);
pTarget += 4;

@AaronRobinsonMSFT
AaronRobinsonMSFT marked this pull request as draft May 3, 2024 04:10
@am11am11 mentioned this pull request May 21, 2024
@dotnet-policy-servicedotnet-policy-serviceBot removed this from the 9.0.0 milestone Jun 2, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Draft Pull Request was automatically closed for 30 days of inactivity. Please let us know if you'd like to reopen it.

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jul 2, 2024
@AaronRobinsonMSFT
AaronRobinsonMSFT deleted the unrevert_string_convert_pr branch November 10, 2025 18:14
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@AaronRobinsonMSFT@am11@jkotas@MichalPetryka
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Fix "Cleanup some string operating functions" PR by AaronRobinsonMSFT · Pull Request #100451 · dotnet/runtime · GitHub
Skip to content

Fix "Cleanup some string operating functions" PR - #100451

Closed
AaronRobinsonMSFT wants to merge 3 commits into
dotnet:mainfrom
AaronRobinsonMSFT:unrevert_string_convert_pr
Closed

Fix "Cleanup some string operating functions" PR#100451
AaronRobinsonMSFT wants to merge 3 commits into
dotnet:mainfrom
AaronRobinsonMSFT:unrevert_string_convert_pr

Conversation

@AaronRobinsonMSFT

@AaronRobinsonMSFTAaronRobinsonMSFT commented Mar 29, 2024

Copy link
Copy Markdown
Member

Unrevert #96099

The revert of #96099 was done in #97264. Commits after the first fix troublesome locations.

See #97264 for details on what the original PR impacted.

/cc @tommcdon@hoyosjs@huoyaoyuan

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/inc/corhlprpriv.h Outdated
Comment threadsrc/coreclr/inc/corhlprpriv.h Outdated
Comment threadsrc/coreclr/inc/corhlprpriv.h Outdated
Use correct error code mechanism.
Comment threadsrc/coreclr/vm/methodtablebuilder.cpp
Comment threadsrc/coreclr/vm/methodtablebuilder.cpp
Comment threadsrc/coreclr/vm/classcompat.cpp
if (SUCCEEDED(hr))
{
s.Resize(length, REPRESENTATION_UTF8);
COUNT_T length = WszWideCharToMultiByte(CP_UTF8, 0, GetRawUnicode(), GetRawCount()+1, NULL, 0, NULL, NULL);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I assume that this used FString::Unicode_Utf8_Length to avoid some perf problems in WszWideCharToMultiByte. Is switching to WszWideCharToMultiByte going to introduce performance regression? (I am particularly worried about Windows.)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I will check here.

Regardless, in principle I would prefer to defer to system APIs that we can/should assume are optimized. Having these narrowly defined "optimizations" is a non-trivial tax for an already covoluted space like SString.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree that we should get rid of FString::*. We may want to replace it with a call to the minipal UTF8 methods for perf reasons like what the original change tried to do. I would expect that minipal UTF8 convertors are going to have significantly better perf than Windows OS WideCharToMultiByte.

system APIs that we can/should assume are optimized

Historically, Windows OS WideCharToMultiByte has been significantly slower for UTF8 than one would expect from a reasonably optimized implementation.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would static linking simdutf to the VM for such conversions make sense here?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We may want to replace it with a call to the minipal UTF8 methods for perf reasons like what the original change tried to do.

So that I can get behind. There is a small performance regression here. Let me replace this with the minipal APIs. I chose this direction to match the other API. I'd prefer symmetrical API usage in this case as it avoids confusion.

Would static linking simdutf to the VM for such conversions make sense here?

That is something we can discuss, but it not for this PR.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we can push the hotpaths to managed

That seems unlikely in this case. The use of SString is litered throughout the system and there are many places where it is going to be unnatural and difficult to call into managed code. We can try jumping out in some cases, but that isn't something I am inclined to do in this PR.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

My bar for these types of changes is to avoid perf regressions. (Of course, perf improvements are nice - but it is better to evaluate them separately if they come with tradeoffs.)

Neither the existing FString fast path code nor the WideCharToMultiByte implementation built into Windows use any vectorization tricks. It suggests that we should not need vectorization tricks to avoid the perf regression here.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The shape of the code in WideCharToMultiByte built into Windows is actually very similar to the corefx implementation and the minipal. It seems that the minipal implementation picked up some cruft alone the way that makes it quite a bit slower.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is interesting that we picked up minipal implementation for Windows + Unix in mono (and Unix in coreclr as before) in .NET 8 and haven't seen any report of regression. Maybe there is some low-hanging alignment issue which can be tweaked via cl switch?

@jkotasjkotasApr 12, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

picked up minipal implementation for Windows + Unix in mono

Mono does not get as much perf scrutiny as CoreCLR.

low-hanging alignment issue

I do not think it is that.

From a cursory look, I see two problems:

  • For smaller lengths, there is a fixed overhead. Adding a simple FString-like loop that deals with small ASCII-only strings as the first thing in the minipal_* methods should fix that.
  • For longer lengths, the code of the core look does not look great (e.g. it uses more registers than necessary). For example, it may help to change
    *pTarget= (unsigned char)ch;
    *(pTarget+1) = (unsigned char)(ch >> 16);
    pSrc+=4;
    *(pTarget+2) = (unsigned char)chc;
    *(pTarget+3) = (unsigned char)(chc >> 16);
    pTarget+=4;
    to
 *pTarget = (unsigned char)ch;
*(pTarget + 2) = (unsigned char)(chc);
pSrc += 4;
*(pTarget + 1) = (unsigned char)(ch >> 16);
*(pTarget + 3) = (unsigned char)(chc >> 16);
pTarget += 4;

@AaronRobinsonMSFT
AaronRobinsonMSFT marked this pull request as draft May 3, 2024 04:10
@am11am11 mentioned this pull request May 21, 2024
@dotnet-policy-servicedotnet-policy-serviceBot removed this from the 9.0.0 milestone Jun 2, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Draft Pull Request was automatically closed for 30 days of inactivity. Please let us know if you'd like to reopen it.

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jul 2, 2024
@AaronRobinsonMSFT
AaronRobinsonMSFT deleted the unrevert_string_convert_pr branch November 10, 2025 18:14
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@AaronRobinsonMSFT@am11@jkotas@MichalPetryka
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Fix "Cleanup some string operating functions" PR by AaronRobinsonMSFT · Pull Request #100451 · dotnet/runtime · GitHub
Skip to content

Fix "Cleanup some string operating functions" PR - #100451

Closed
AaronRobinsonMSFT wants to merge 3 commits into
dotnet:mainfrom
AaronRobinsonMSFT:unrevert_string_convert_pr
Closed

Fix "Cleanup some string operating functions" PR#100451
AaronRobinsonMSFT wants to merge 3 commits into
dotnet:mainfrom
AaronRobinsonMSFT:unrevert_string_convert_pr

Conversation

@AaronRobinsonMSFT

@AaronRobinsonMSFTAaronRobinsonMSFT commented Mar 29, 2024

Copy link
Copy Markdown
Member

Unrevert #96099

The revert of #96099 was done in #97264. Commits after the first fix troublesome locations.

See #97264 for details on what the original PR impacted.

/cc @tommcdon@hoyosjs@huoyaoyuan

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/inc/corhlprpriv.h Outdated
Comment threadsrc/coreclr/inc/corhlprpriv.h Outdated
Comment threadsrc/coreclr/inc/corhlprpriv.h Outdated
Use correct error code mechanism.
Comment threadsrc/coreclr/vm/methodtablebuilder.cpp
Comment threadsrc/coreclr/vm/methodtablebuilder.cpp
Comment threadsrc/coreclr/vm/classcompat.cpp
if (SUCCEEDED(hr))
{
s.Resize(length, REPRESENTATION_UTF8);
COUNT_T length = WszWideCharToMultiByte(CP_UTF8, 0, GetRawUnicode(), GetRawCount()+1, NULL, 0, NULL, NULL);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I assume that this used FString::Unicode_Utf8_Length to avoid some perf problems in WszWideCharToMultiByte. Is switching to WszWideCharToMultiByte going to introduce performance regression? (I am particularly worried about Windows.)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I will check here.

Regardless, in principle I would prefer to defer to system APIs that we can/should assume are optimized. Having these narrowly defined "optimizations" is a non-trivial tax for an already covoluted space like SString.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree that we should get rid of FString::*. We may want to replace it with a call to the minipal UTF8 methods for perf reasons like what the original change tried to do. I would expect that minipal UTF8 convertors are going to have significantly better perf than Windows OS WideCharToMultiByte.

system APIs that we can/should assume are optimized

Historically, Windows OS WideCharToMultiByte has been significantly slower for UTF8 than one would expect from a reasonably optimized implementation.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would static linking simdutf to the VM for such conversions make sense here?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We may want to replace it with a call to the minipal UTF8 methods for perf reasons like what the original change tried to do.

So that I can get behind. There is a small performance regression here. Let me replace this with the minipal APIs. I chose this direction to match the other API. I'd prefer symmetrical API usage in this case as it avoids confusion.

Would static linking simdutf to the VM for such conversions make sense here?

That is something we can discuss, but it not for this PR.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we can push the hotpaths to managed

That seems unlikely in this case. The use of SString is litered throughout the system and there are many places where it is going to be unnatural and difficult to call into managed code. We can try jumping out in some cases, but that isn't something I am inclined to do in this PR.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

My bar for these types of changes is to avoid perf regressions. (Of course, perf improvements are nice - but it is better to evaluate them separately if they come with tradeoffs.)

Neither the existing FString fast path code nor the WideCharToMultiByte implementation built into Windows use any vectorization tricks. It suggests that we should not need vectorization tricks to avoid the perf regression here.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The shape of the code in WideCharToMultiByte built into Windows is actually very similar to the corefx implementation and the minipal. It seems that the minipal implementation picked up some cruft alone the way that makes it quite a bit slower.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is interesting that we picked up minipal implementation for Windows + Unix in mono (and Unix in coreclr as before) in .NET 8 and haven't seen any report of regression. Maybe there is some low-hanging alignment issue which can be tweaked via cl switch?

@jkotasjkotasApr 12, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

picked up minipal implementation for Windows + Unix in mono

Mono does not get as much perf scrutiny as CoreCLR.

low-hanging alignment issue

I do not think it is that.

From a cursory look, I see two problems:

  • For smaller lengths, there is a fixed overhead. Adding a simple FString-like loop that deals with small ASCII-only strings as the first thing in the minipal_* methods should fix that.
  • For longer lengths, the code of the core look does not look great (e.g. it uses more registers than necessary). For example, it may help to change
    *pTarget= (unsigned char)ch;
    *(pTarget+1) = (unsigned char)(ch >> 16);
    pSrc+=4;
    *(pTarget+2) = (unsigned char)chc;
    *(pTarget+3) = (unsigned char)(chc >> 16);
    pTarget+=4;
    to
 *pTarget = (unsigned char)ch;
*(pTarget + 2) = (unsigned char)(chc);
pSrc += 4;
*(pTarget + 1) = (unsigned char)(ch >> 16);
*(pTarget + 3) = (unsigned char)(chc >> 16);
pTarget += 4;

@AaronRobinsonMSFT
AaronRobinsonMSFT marked this pull request as draft May 3, 2024 04:10
@am11am11 mentioned this pull request May 21, 2024
@dotnet-policy-servicedotnet-policy-serviceBot removed this from the 9.0.0 milestone Jun 2, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Draft Pull Request was automatically closed for 30 days of inactivity. Please let us know if you'd like to reopen it.

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jul 2, 2024
@AaronRobinsonMSFT
AaronRobinsonMSFT deleted the unrevert_string_convert_pr branch November 10, 2025 18:14
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@AaronRobinsonMSFT@am11@jkotas@MichalPetryka
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' Fix "Cleanup some string operating functions" PR by AaronRobinsonMSFT · Pull Request #100451 · dotnet/runtime · GitHub
Skip to content

Fix "Cleanup some string operating functions" PR - #100451

Closed
AaronRobinsonMSFT wants to merge 3 commits into
dotnet:mainfrom
AaronRobinsonMSFT:unrevert_string_convert_pr
Closed

Fix "Cleanup some string operating functions" PR#100451
AaronRobinsonMSFT wants to merge 3 commits into
dotnet:mainfrom
AaronRobinsonMSFT:unrevert_string_convert_pr

Conversation

@AaronRobinsonMSFT

@AaronRobinsonMSFTAaronRobinsonMSFT commented Mar 29, 2024

Copy link
Copy Markdown
Member

Unrevert #96099

The revert of #96099 was done in #97264. Commits after the first fix troublesome locations.

See #97264 for details on what the original PR impacted.

/cc @tommcdon@hoyosjs@huoyaoyuan

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/inc/corhlprpriv.h Outdated
Comment threadsrc/coreclr/inc/corhlprpriv.h Outdated
Comment threadsrc/coreclr/inc/corhlprpriv.h Outdated
Use correct error code mechanism.
Comment threadsrc/coreclr/vm/methodtablebuilder.cpp
Comment threadsrc/coreclr/vm/methodtablebuilder.cpp
Comment threadsrc/coreclr/vm/classcompat.cpp
if (SUCCEEDED(hr))
{
s.Resize(length, REPRESENTATION_UTF8);
COUNT_T length = WszWideCharToMultiByte(CP_UTF8, 0, GetRawUnicode(), GetRawCount()+1, NULL, 0, NULL, NULL);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I assume that this used FString::Unicode_Utf8_Length to avoid some perf problems in WszWideCharToMultiByte. Is switching to WszWideCharToMultiByte going to introduce performance regression? (I am particularly worried about Windows.)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I will check here.

Regardless, in principle I would prefer to defer to system APIs that we can/should assume are optimized. Having these narrowly defined "optimizations" is a non-trivial tax for an already covoluted space like SString.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree that we should get rid of FString::*. We may want to replace it with a call to the minipal UTF8 methods for perf reasons like what the original change tried to do. I would expect that minipal UTF8 convertors are going to have significantly better perf than Windows OS WideCharToMultiByte.

system APIs that we can/should assume are optimized

Historically, Windows OS WideCharToMultiByte has been significantly slower for UTF8 than one would expect from a reasonably optimized implementation.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would static linking simdutf to the VM for such conversions make sense here?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We may want to replace it with a call to the minipal UTF8 methods for perf reasons like what the original change tried to do.

So that I can get behind. There is a small performance regression here. Let me replace this with the minipal APIs. I chose this direction to match the other API. I'd prefer symmetrical API usage in this case as it avoids confusion.

Would static linking simdutf to the VM for such conversions make sense here?

That is something we can discuss, but it not for this PR.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we can push the hotpaths to managed

That seems unlikely in this case. The use of SString is litered throughout the system and there are many places where it is going to be unnatural and difficult to call into managed code. We can try jumping out in some cases, but that isn't something I am inclined to do in this PR.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

My bar for these types of changes is to avoid perf regressions. (Of course, perf improvements are nice - but it is better to evaluate them separately if they come with tradeoffs.)

Neither the existing FString fast path code nor the WideCharToMultiByte implementation built into Windows use any vectorization tricks. It suggests that we should not need vectorization tricks to avoid the perf regression here.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The shape of the code in WideCharToMultiByte built into Windows is actually very similar to the corefx implementation and the minipal. It seems that the minipal implementation picked up some cruft alone the way that makes it quite a bit slower.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is interesting that we picked up minipal implementation for Windows + Unix in mono (and Unix in coreclr as before) in .NET 8 and haven't seen any report of regression. Maybe there is some low-hanging alignment issue which can be tweaked via cl switch?

@jkotasjkotasApr 12, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

picked up minipal implementation for Windows + Unix in mono

Mono does not get as much perf scrutiny as CoreCLR.

low-hanging alignment issue

I do not think it is that.

From a cursory look, I see two problems:

  • For smaller lengths, there is a fixed overhead. Adding a simple FString-like loop that deals with small ASCII-only strings as the first thing in the minipal_* methods should fix that.
  • For longer lengths, the code of the core look does not look great (e.g. it uses more registers than necessary). For example, it may help to change
    *pTarget= (unsigned char)ch;
    *(pTarget+1) = (unsigned char)(ch >> 16);
    pSrc+=4;
    *(pTarget+2) = (unsigned char)chc;
    *(pTarget+3) = (unsigned char)(chc >> 16);
    pTarget+=4;
    to
 *pTarget = (unsigned char)ch;
*(pTarget + 2) = (unsigned char)(chc);
pSrc += 4;
*(pTarget + 1) = (unsigned char)(ch >> 16);
*(pTarget + 3) = (unsigned char)(chc >> 16);
pTarget += 4;

@AaronRobinsonMSFT
AaronRobinsonMSFT marked this pull request as draft May 3, 2024 04:10
@am11am11 mentioned this pull request May 21, 2024
@dotnet-policy-servicedotnet-policy-serviceBot removed this from the 9.0.0 milestone Jun 2, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Draft Pull Request was automatically closed for 30 days of inactivity. Please let us know if you'd like to reopen it.

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jul 2, 2024
@AaronRobinsonMSFT
AaronRobinsonMSFT deleted the unrevert_string_convert_pr branch November 10, 2025 18:14
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@AaronRobinsonMSFT@am11@jkotas@MichalPetryka
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Fix "Cleanup some string operating functions" PR by AaronRobinsonMSFT · Pull Request #100451 · dotnet/runtime · GitHub
Skip to content

Fix "Cleanup some string operating functions" PR - #100451

Closed
AaronRobinsonMSFT wants to merge 3 commits into
dotnet:mainfrom
AaronRobinsonMSFT:unrevert_string_convert_pr
Closed

Fix "Cleanup some string operating functions" PR#100451
AaronRobinsonMSFT wants to merge 3 commits into
dotnet:mainfrom
AaronRobinsonMSFT:unrevert_string_convert_pr

Conversation

@AaronRobinsonMSFT

@AaronRobinsonMSFTAaronRobinsonMSFT commented Mar 29, 2024

Copy link
Copy Markdown
Member

Unrevert #96099

The revert of #96099 was done in #97264. Commits after the first fix troublesome locations.

See #97264 for details on what the original PR impacted.

/cc @tommcdon@hoyosjs@huoyaoyuan

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/inc/corhlprpriv.h Outdated
Comment threadsrc/coreclr/inc/corhlprpriv.h Outdated
Comment threadsrc/coreclr/inc/corhlprpriv.h Outdated
Use correct error code mechanism.
Comment threadsrc/coreclr/vm/methodtablebuilder.cpp
Comment threadsrc/coreclr/vm/methodtablebuilder.cpp
Comment threadsrc/coreclr/vm/classcompat.cpp
if (SUCCEEDED(hr))
{
s.Resize(length, REPRESENTATION_UTF8);
COUNT_T length = WszWideCharToMultiByte(CP_UTF8, 0, GetRawUnicode(), GetRawCount()+1, NULL, 0, NULL, NULL);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I assume that this used FString::Unicode_Utf8_Length to avoid some perf problems in WszWideCharToMultiByte. Is switching to WszWideCharToMultiByte going to introduce performance regression? (I am particularly worried about Windows.)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I will check here.

Regardless, in principle I would prefer to defer to system APIs that we can/should assume are optimized. Having these narrowly defined "optimizations" is a non-trivial tax for an already covoluted space like SString.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree that we should get rid of FString::*. We may want to replace it with a call to the minipal UTF8 methods for perf reasons like what the original change tried to do. I would expect that minipal UTF8 convertors are going to have significantly better perf than Windows OS WideCharToMultiByte.

system APIs that we can/should assume are optimized

Historically, Windows OS WideCharToMultiByte has been significantly slower for UTF8 than one would expect from a reasonably optimized implementation.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would static linking simdutf to the VM for such conversions make sense here?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We may want to replace it with a call to the minipal UTF8 methods for perf reasons like what the original change tried to do.

So that I can get behind. There is a small performance regression here. Let me replace this with the minipal APIs. I chose this direction to match the other API. I'd prefer symmetrical API usage in this case as it avoids confusion.

Would static linking simdutf to the VM for such conversions make sense here?

That is something we can discuss, but it not for this PR.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we can push the hotpaths to managed

That seems unlikely in this case. The use of SString is litered throughout the system and there are many places where it is going to be unnatural and difficult to call into managed code. We can try jumping out in some cases, but that isn't something I am inclined to do in this PR.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

My bar for these types of changes is to avoid perf regressions. (Of course, perf improvements are nice - but it is better to evaluate them separately if they come with tradeoffs.)

Neither the existing FString fast path code nor the WideCharToMultiByte implementation built into Windows use any vectorization tricks. It suggests that we should not need vectorization tricks to avoid the perf regression here.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The shape of the code in WideCharToMultiByte built into Windows is actually very similar to the corefx implementation and the minipal. It seems that the minipal implementation picked up some cruft alone the way that makes it quite a bit slower.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is interesting that we picked up minipal implementation for Windows + Unix in mono (and Unix in coreclr as before) in .NET 8 and haven't seen any report of regression. Maybe there is some low-hanging alignment issue which can be tweaked via cl switch?

@jkotasjkotasApr 12, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

picked up minipal implementation for Windows + Unix in mono

Mono does not get as much perf scrutiny as CoreCLR.

low-hanging alignment issue

I do not think it is that.

From a cursory look, I see two problems:

  • For smaller lengths, there is a fixed overhead. Adding a simple FString-like loop that deals with small ASCII-only strings as the first thing in the minipal_* methods should fix that.
  • For longer lengths, the code of the core look does not look great (e.g. it uses more registers than necessary). For example, it may help to change
    *pTarget= (unsigned char)ch;
    *(pTarget+1) = (unsigned char)(ch >> 16);
    pSrc+=4;
    *(pTarget+2) = (unsigned char)chc;
    *(pTarget+3) = (unsigned char)(chc >> 16);
    pTarget+=4;
    to
 *pTarget = (unsigned char)ch;
*(pTarget + 2) = (unsigned char)(chc);
pSrc += 4;
*(pTarget + 1) = (unsigned char)(ch >> 16);
*(pTarget + 3) = (unsigned char)(chc >> 16);
pTarget += 4;

@AaronRobinsonMSFT
AaronRobinsonMSFT marked this pull request as draft May 3, 2024 04:10
@am11am11 mentioned this pull request May 21, 2024
@dotnet-policy-servicedotnet-policy-serviceBot removed this from the 9.0.0 milestone Jun 2, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Draft Pull Request was automatically closed for 30 days of inactivity. Please let us know if you'd like to reopen it.

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jul 2, 2024
@AaronRobinsonMSFT
AaronRobinsonMSFT deleted the unrevert_string_convert_pr branch November 10, 2025 18:14
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@AaronRobinsonMSFT@am11@jkotas@MichalPetryka
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Fix "Cleanup some string operating functions" PR by AaronRobinsonMSFT · Pull Request #100451 · dotnet/runtime · GitHub
Skip to content

Fix "Cleanup some string operating functions" PR - #100451

Closed
AaronRobinsonMSFT wants to merge 3 commits into
dotnet:mainfrom
AaronRobinsonMSFT:unrevert_string_convert_pr
Closed

Fix "Cleanup some string operating functions" PR#100451
AaronRobinsonMSFT wants to merge 3 commits into
dotnet:mainfrom
AaronRobinsonMSFT:unrevert_string_convert_pr

Conversation

@AaronRobinsonMSFT

@AaronRobinsonMSFTAaronRobinsonMSFT commented Mar 29, 2024

Copy link
Copy Markdown
Member

Unrevert #96099

The revert of #96099 was done in #97264. Commits after the first fix troublesome locations.

See #97264 for details on what the original PR impacted.

/cc @tommcdon@hoyosjs@huoyaoyuan

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/inc/corhlprpriv.h Outdated
Comment threadsrc/coreclr/inc/corhlprpriv.h Outdated
Comment threadsrc/coreclr/inc/corhlprpriv.h Outdated
Use correct error code mechanism.
Comment threadsrc/coreclr/vm/methodtablebuilder.cpp
Comment threadsrc/coreclr/vm/methodtablebuilder.cpp
Comment threadsrc/coreclr/vm/classcompat.cpp
if (SUCCEEDED(hr))
{
s.Resize(length, REPRESENTATION_UTF8);
COUNT_T length = WszWideCharToMultiByte(CP_UTF8, 0, GetRawUnicode(), GetRawCount()+1, NULL, 0, NULL, NULL);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I assume that this used FString::Unicode_Utf8_Length to avoid some perf problems in WszWideCharToMultiByte. Is switching to WszWideCharToMultiByte going to introduce performance regression? (I am particularly worried about Windows.)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I will check here.

Regardless, in principle I would prefer to defer to system APIs that we can/should assume are optimized. Having these narrowly defined "optimizations" is a non-trivial tax for an already covoluted space like SString.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree that we should get rid of FString::*. We may want to replace it with a call to the minipal UTF8 methods for perf reasons like what the original change tried to do. I would expect that minipal UTF8 convertors are going to have significantly better perf than Windows OS WideCharToMultiByte.

system APIs that we can/should assume are optimized

Historically, Windows OS WideCharToMultiByte has been significantly slower for UTF8 than one would expect from a reasonably optimized implementation.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would static linking simdutf to the VM for such conversions make sense here?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We may want to replace it with a call to the minipal UTF8 methods for perf reasons like what the original change tried to do.

So that I can get behind. There is a small performance regression here. Let me replace this with the minipal APIs. I chose this direction to match the other API. I'd prefer symmetrical API usage in this case as it avoids confusion.

Would static linking simdutf to the VM for such conversions make sense here?

That is something we can discuss, but it not for this PR.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we can push the hotpaths to managed

That seems unlikely in this case. The use of SString is litered throughout the system and there are many places where it is going to be unnatural and difficult to call into managed code. We can try jumping out in some cases, but that isn't something I am inclined to do in this PR.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

My bar for these types of changes is to avoid perf regressions. (Of course, perf improvements are nice - but it is better to evaluate them separately if they come with tradeoffs.)

Neither the existing FString fast path code nor the WideCharToMultiByte implementation built into Windows use any vectorization tricks. It suggests that we should not need vectorization tricks to avoid the perf regression here.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The shape of the code in WideCharToMultiByte built into Windows is actually very similar to the corefx implementation and the minipal. It seems that the minipal implementation picked up some cruft alone the way that makes it quite a bit slower.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is interesting that we picked up minipal implementation for Windows + Unix in mono (and Unix in coreclr as before) in .NET 8 and haven't seen any report of regression. Maybe there is some low-hanging alignment issue which can be tweaked via cl switch?

@jkotasjkotasApr 12, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

picked up minipal implementation for Windows + Unix in mono

Mono does not get as much perf scrutiny as CoreCLR.

low-hanging alignment issue

I do not think it is that.

From a cursory look, I see two problems:

  • For smaller lengths, there is a fixed overhead. Adding a simple FString-like loop that deals with small ASCII-only strings as the first thing in the minipal_* methods should fix that.
  • For longer lengths, the code of the core look does not look great (e.g. it uses more registers than necessary). For example, it may help to change
    *pTarget= (unsigned char)ch;
    *(pTarget+1) = (unsigned char)(ch >> 16);
    pSrc+=4;
    *(pTarget+2) = (unsigned char)chc;
    *(pTarget+3) = (unsigned char)(chc >> 16);
    pTarget+=4;
    to
 *pTarget = (unsigned char)ch;
*(pTarget + 2) = (unsigned char)(chc);
pSrc += 4;
*(pTarget + 1) = (unsigned char)(ch >> 16);
*(pTarget + 3) = (unsigned char)(chc >> 16);
pTarget += 4;

@AaronRobinsonMSFT
AaronRobinsonMSFT marked this pull request as draft May 3, 2024 04:10
@am11am11 mentioned this pull request May 21, 2024
@dotnet-policy-servicedotnet-policy-serviceBot removed this from the 9.0.0 milestone Jun 2, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Draft Pull Request was automatically closed for 30 days of inactivity. Please let us know if you'd like to reopen it.

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jul 2, 2024
@AaronRobinsonMSFT
AaronRobinsonMSFT deleted the unrevert_string_convert_pr branch November 10, 2025 18:14
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@AaronRobinsonMSFT@am11@jkotas@MichalPetryka
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); Fix "Cleanup some string operating functions" PR by AaronRobinsonMSFT · Pull Request #100451 · dotnet/runtime · GitHub
Skip to content

Fix "Cleanup some string operating functions" PR - #100451

Closed
AaronRobinsonMSFT wants to merge 3 commits into
dotnet:mainfrom
AaronRobinsonMSFT:unrevert_string_convert_pr
Closed

Fix "Cleanup some string operating functions" PR#100451
AaronRobinsonMSFT wants to merge 3 commits into
dotnet:mainfrom
AaronRobinsonMSFT:unrevert_string_convert_pr

Conversation

@AaronRobinsonMSFT

@AaronRobinsonMSFTAaronRobinsonMSFT commented Mar 29, 2024

Copy link
Copy Markdown
Member

Unrevert #96099

The revert of #96099 was done in #97264. Commits after the first fix troublesome locations.

See #97264 for details on what the original PR impacted.

/cc @tommcdon@hoyosjs@huoyaoyuan

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @mangod9
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/inc/corhlprpriv.h Outdated
Comment threadsrc/coreclr/inc/corhlprpriv.h Outdated
Comment threadsrc/coreclr/inc/corhlprpriv.h Outdated
Use correct error code mechanism.
Comment threadsrc/coreclr/vm/methodtablebuilder.cpp
Comment threadsrc/coreclr/vm/methodtablebuilder.cpp
Comment threadsrc/coreclr/vm/classcompat.cpp
if (SUCCEEDED(hr))
{
s.Resize(length, REPRESENTATION_UTF8);
COUNT_T length = WszWideCharToMultiByte(CP_UTF8, 0, GetRawUnicode(), GetRawCount()+1, NULL, 0, NULL, NULL);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I assume that this used FString::Unicode_Utf8_Length to avoid some perf problems in WszWideCharToMultiByte. Is switching to WszWideCharToMultiByte going to introduce performance regression? (I am particularly worried about Windows.)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I will check here.

Regardless, in principle I would prefer to defer to system APIs that we can/should assume are optimized. Having these narrowly defined "optimizations" is a non-trivial tax for an already covoluted space like SString.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree that we should get rid of FString::*. We may want to replace it with a call to the minipal UTF8 methods for perf reasons like what the original change tried to do. I would expect that minipal UTF8 convertors are going to have significantly better perf than Windows OS WideCharToMultiByte.

system APIs that we can/should assume are optimized

Historically, Windows OS WideCharToMultiByte has been significantly slower for UTF8 than one would expect from a reasonably optimized implementation.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would static linking simdutf to the VM for such conversions make sense here?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We may want to replace it with a call to the minipal UTF8 methods for perf reasons like what the original change tried to do.

So that I can get behind. There is a small performance regression here. Let me replace this with the minipal APIs. I chose this direction to match the other API. I'd prefer symmetrical API usage in this case as it avoids confusion.

Would static linking simdutf to the VM for such conversions make sense here?

That is something we can discuss, but it not for this PR.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we can push the hotpaths to managed

That seems unlikely in this case. The use of SString is litered throughout the system and there are many places where it is going to be unnatural and difficult to call into managed code. We can try jumping out in some cases, but that isn't something I am inclined to do in this PR.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

My bar for these types of changes is to avoid perf regressions. (Of course, perf improvements are nice - but it is better to evaluate them separately if they come with tradeoffs.)

Neither the existing FString fast path code nor the WideCharToMultiByte implementation built into Windows use any vectorization tricks. It suggests that we should not need vectorization tricks to avoid the perf regression here.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The shape of the code in WideCharToMultiByte built into Windows is actually very similar to the corefx implementation and the minipal. It seems that the minipal implementation picked up some cruft alone the way that makes it quite a bit slower.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is interesting that we picked up minipal implementation for Windows + Unix in mono (and Unix in coreclr as before) in .NET 8 and haven't seen any report of regression. Maybe there is some low-hanging alignment issue which can be tweaked via cl switch?

@jkotasjkotasApr 12, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

picked up minipal implementation for Windows + Unix in mono

Mono does not get as much perf scrutiny as CoreCLR.

low-hanging alignment issue

I do not think it is that.

From a cursory look, I see two problems:

  • For smaller lengths, there is a fixed overhead. Adding a simple FString-like loop that deals with small ASCII-only strings as the first thing in the minipal_* methods should fix that.
  • For longer lengths, the code of the core look does not look great (e.g. it uses more registers than necessary). For example, it may help to change
    *pTarget= (unsigned char)ch;
    *(pTarget+1) = (unsigned char)(ch >> 16);
    pSrc+=4;
    *(pTarget+2) = (unsigned char)chc;
    *(pTarget+3) = (unsigned char)(chc >> 16);
    pTarget+=4;
    to
 *pTarget = (unsigned char)ch;
*(pTarget + 2) = (unsigned char)(chc);
pSrc += 4;
*(pTarget + 1) = (unsigned char)(ch >> 16);
*(pTarget + 3) = (unsigned char)(chc >> 16);
pTarget += 4;

@AaronRobinsonMSFT
AaronRobinsonMSFT marked this pull request as draft May 3, 2024 04:10
@am11am11 mentioned this pull request May 21, 2024
@dotnet-policy-servicedotnet-policy-serviceBot removed this from the 9.0.0 milestone Jun 2, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Draft Pull Request was automatically closed for 30 days of inactivity. Please let us know if you'd like to reopen it.

@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jul 2, 2024
@AaronRobinsonMSFT
AaronRobinsonMSFT deleted the unrevert_string_convert_pr branch November 10, 2025 18:14
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@AaronRobinsonMSFT@am11@jkotas@MichalPetryka