Mostly switch object writer to UTF-8 - #122160

Merged
MichalStrehovsky merged 5 commits into
dotnet:mainfrom
MichalStrehovsky:utf8objwriter
Dec 8, 2025
Merged

Mostly switch object writer to UTF-8#122160
MichalStrehovsky merged 5 commits into
dotnet:mainfrom
MichalStrehovsky:utf8objwriter

Conversation

@MichalStrehovsky

@MichalStrehovskyMichalStrehovsky commented Dec 3, 2025

Copy link
Copy Markdown
Member

When compiling hello world:

byte[] allocations before: 564000. string allocations before: 302000
byte[] allocations after: 625000. string allocations after: 241000

So this is mostly a wash allocation-wise, however, not allocating string means we're also avoiding the UTF-8 -> UTF-16 -> UTF-8 conversions.

Cc @dotnet/ilc-contrib

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/tools/Common/Internal/Text/Utf8String.cs Outdated
@MichalStrehovskyMichalStrehovsky changed the title Switch object writer to UTF-8Mostly switch object writer to UTF-8Dec 5, 2025
@MichalStrehovsky
MichalStrehovsky marked this pull request as ready for review December 5, 2025 05:49
CopilotAI review requested due to automatic review settings December 5, 2025 05:49

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR converts the object writer subsystem from using string to Utf8String to avoid UTF-8 ↔ UTF-16 ↔ UTF-8 conversions during compilation. The change primarily affects symbol naming, mangling, and object file generation code.

Key Changes

  • Introduced IsNull property to Utf8String for null checking
  • Added multiple Concat overload methods for efficient UTF-8 string concatenation
  • Updated all object writer interfaces and implementations to accept/return Utf8String instead of string
  • Changed dictionaries and collections to use Utf8String keys

Reviewed changes

Copilot reviewed 44 out of 44 changed files in this pull request and generated 11 comments.

Show a summary per file
FileDescription
Internal/Text/Utf8String.csAdded IsNull property and new Concat overloads for UTF-8 string operations
ObjectWriter/ObjectWriter.csCore object writer changed to use Utf8String for symbol names and relocations
ObjectWriter/StringTableBuilder.csUpdated string table to work directly with UTF-8 bytes
ObjectWriter/CoffObjectWriter.csCOFF format writer updated to use Utf8String
ObjectWriter/ElfObjectWriter.csELF format writer updated to use Utf8String
ObjectWriter/MachObjectWriter.csMach-O format writer updated to use Utf8String
ObjectWriter/UnixObjectWriter.csUnix-specific object writer base updated
Compiler/NodeMangler.csName mangling infrastructure changed to return Utf8String
Compiler/WindowsNodeMangler.csWindows-specific name mangling updated with UTF-8 concatenation
Compiler/UnixNodeMangler.csUnix-specific name mangling updated with UTF-8 concatenation
Compiler/UserDefinedTypeDescriptor.csDebug info type descriptors updated to use Utf8String
DependencyAnalysis/NodeFactory.csFactory methods updated for Utf8String symbol names
DependencyAnalysis/*Node.csVarious node types updated to use Utf8String for mangled names

Comment threadsrc/coreclr/tools/Common/Compiler/ObjectWriter/StringTableBuilder.cs Outdated
@MichalStrehovsky

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-nativeaot-outerloop

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@MichalStrehovsky
MichalStrehovsky merged commit 8c8bfb2 into dotnet:mainDec 8, 2025
111 of 120 checks passed
@MichalStrehovsky
MichalStrehovsky deleted the utf8objwriter branch December 8, 2025 12:59

private Utf8String SanitizeNameWithHash(Utf8String literal)
{
Utf8String mangledName = SanitizeName(literal);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That code seems to be dangerous.
The result is not really an Utf8-String but ASCII - and it needs to be because the next few lines would be incorrect if it would be UTF8 and would contain any multi-byte chars (you can't simply cut off arbitrary Utf8 at byte position 30).
So if anybody ever changes SanitizeName to actually be Utf8 this will create hard-to-spot errors.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So if anybody ever changes SanitizeName to actually be Utf8 this will create hard-to-spot errors.

It's not likely we'd ever allow SanitizeName to return characters outside the basic ASCII set.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would a Debug.Assert(Ascii.IsValid(mangledName)) make sense here?

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@MichalStrehovsky@ANahr@filipnavara@am11@jkotas@PaulusParssinen@MichalPetryka
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Mostly switch object writer to UTF-8 - #122160

Merged
MichalStrehovsky merged 5 commits into
dotnet:mainfrom
MichalStrehovsky:utf8objwriter
Dec 8, 2025
Merged

Mostly switch object writer to UTF-8#122160
MichalStrehovsky merged 5 commits into
dotnet:mainfrom
MichalStrehovsky:utf8objwriter

Conversation

@MichalStrehovsky

@MichalStrehovskyMichalStrehovsky commented Dec 3, 2025

Copy link
Copy Markdown
Member

When compiling hello world:

byte[] allocations before: 564000. string allocations before: 302000
byte[] allocations after: 625000. string allocations after: 241000

So this is mostly a wash allocation-wise, however, not allocating string means we're also avoiding the UTF-8 -> UTF-16 -> UTF-8 conversions.

Cc @dotnet/ilc-contrib

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/tools/Common/Internal/Text/Utf8String.cs Outdated
@MichalStrehovskyMichalStrehovsky changed the title Switch object writer to UTF-8Mostly switch object writer to UTF-8Dec 5, 2025
@MichalStrehovsky
MichalStrehovsky marked this pull request as ready for review December 5, 2025 05:49
CopilotAI review requested due to automatic review settings December 5, 2025 05:49

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR converts the object writer subsystem from using string to Utf8String to avoid UTF-8 ↔ UTF-16 ↔ UTF-8 conversions during compilation. The change primarily affects symbol naming, mangling, and object file generation code.

Key Changes

  • Introduced IsNull property to Utf8String for null checking
  • Added multiple Concat overload methods for efficient UTF-8 string concatenation
  • Updated all object writer interfaces and implementations to accept/return Utf8String instead of string
  • Changed dictionaries and collections to use Utf8String keys

Reviewed changes

Copilot reviewed 44 out of 44 changed files in this pull request and generated 11 comments.

Show a summary per file
FileDescription
Internal/Text/Utf8String.csAdded IsNull property and new Concat overloads for UTF-8 string operations
ObjectWriter/ObjectWriter.csCore object writer changed to use Utf8String for symbol names and relocations
ObjectWriter/StringTableBuilder.csUpdated string table to work directly with UTF-8 bytes
ObjectWriter/CoffObjectWriter.csCOFF format writer updated to use Utf8String
ObjectWriter/ElfObjectWriter.csELF format writer updated to use Utf8String
ObjectWriter/MachObjectWriter.csMach-O format writer updated to use Utf8String
ObjectWriter/UnixObjectWriter.csUnix-specific object writer base updated
Compiler/NodeMangler.csName mangling infrastructure changed to return Utf8String
Compiler/WindowsNodeMangler.csWindows-specific name mangling updated with UTF-8 concatenation
Compiler/UnixNodeMangler.csUnix-specific name mangling updated with UTF-8 concatenation
Compiler/UserDefinedTypeDescriptor.csDebug info type descriptors updated to use Utf8String
DependencyAnalysis/NodeFactory.csFactory methods updated for Utf8String symbol names
DependencyAnalysis/*Node.csVarious node types updated to use Utf8String for mangled names

Comment threadsrc/coreclr/tools/Common/Compiler/ObjectWriter/StringTableBuilder.cs Outdated
@MichalStrehovsky

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-nativeaot-outerloop

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@MichalStrehovsky
MichalStrehovsky merged commit 8c8bfb2 into dotnet:mainDec 8, 2025
111 of 120 checks passed
@MichalStrehovsky
MichalStrehovsky deleted the utf8objwriter branch December 8, 2025 12:59

private Utf8String SanitizeNameWithHash(Utf8String literal)
{
Utf8String mangledName = SanitizeName(literal);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That code seems to be dangerous.
The result is not really an Utf8-String but ASCII - and it needs to be because the next few lines would be incorrect if it would be UTF8 and would contain any multi-byte chars (you can't simply cut off arbitrary Utf8 at byte position 30).
So if anybody ever changes SanitizeName to actually be Utf8 this will create hard-to-spot errors.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So if anybody ever changes SanitizeName to actually be Utf8 this will create hard-to-spot errors.

It's not likely we'd ever allow SanitizeName to return characters outside the basic ASCII set.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would a Debug.Assert(Ascii.IsValid(mangledName)) make sense here?

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@MichalStrehovsky@ANahr@filipnavara@am11@jkotas@PaulusParssinen@MichalPetryka
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Mostly switch object writer to UTF-8 - #122160

Merged
MichalStrehovsky merged 5 commits into
dotnet:mainfrom
MichalStrehovsky:utf8objwriter
Dec 8, 2025
Merged

Mostly switch object writer to UTF-8#122160
MichalStrehovsky merged 5 commits into
dotnet:mainfrom
MichalStrehovsky:utf8objwriter

Conversation

@MichalStrehovsky

@MichalStrehovskyMichalStrehovsky commented Dec 3, 2025

Copy link
Copy Markdown
Member

When compiling hello world:

byte[] allocations before: 564000. string allocations before: 302000
byte[] allocations after: 625000. string allocations after: 241000

So this is mostly a wash allocation-wise, however, not allocating string means we're also avoiding the UTF-8 -> UTF-16 -> UTF-8 conversions.

Cc @dotnet/ilc-contrib

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/tools/Common/Internal/Text/Utf8String.cs Outdated
@MichalStrehovskyMichalStrehovsky changed the title Switch object writer to UTF-8Mostly switch object writer to UTF-8Dec 5, 2025
@MichalStrehovsky
MichalStrehovsky marked this pull request as ready for review December 5, 2025 05:49
CopilotAI review requested due to automatic review settings December 5, 2025 05:49

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR converts the object writer subsystem from using string to Utf8String to avoid UTF-8 ↔ UTF-16 ↔ UTF-8 conversions during compilation. The change primarily affects symbol naming, mangling, and object file generation code.

Key Changes

  • Introduced IsNull property to Utf8String for null checking
  • Added multiple Concat overload methods for efficient UTF-8 string concatenation
  • Updated all object writer interfaces and implementations to accept/return Utf8String instead of string
  • Changed dictionaries and collections to use Utf8String keys

Reviewed changes

Copilot reviewed 44 out of 44 changed files in this pull request and generated 11 comments.

Show a summary per file
FileDescription
Internal/Text/Utf8String.csAdded IsNull property and new Concat overloads for UTF-8 string operations
ObjectWriter/ObjectWriter.csCore object writer changed to use Utf8String for symbol names and relocations
ObjectWriter/StringTableBuilder.csUpdated string table to work directly with UTF-8 bytes
ObjectWriter/CoffObjectWriter.csCOFF format writer updated to use Utf8String
ObjectWriter/ElfObjectWriter.csELF format writer updated to use Utf8String
ObjectWriter/MachObjectWriter.csMach-O format writer updated to use Utf8String
ObjectWriter/UnixObjectWriter.csUnix-specific object writer base updated
Compiler/NodeMangler.csName mangling infrastructure changed to return Utf8String
Compiler/WindowsNodeMangler.csWindows-specific name mangling updated with UTF-8 concatenation
Compiler/UnixNodeMangler.csUnix-specific name mangling updated with UTF-8 concatenation
Compiler/UserDefinedTypeDescriptor.csDebug info type descriptors updated to use Utf8String
DependencyAnalysis/NodeFactory.csFactory methods updated for Utf8String symbol names
DependencyAnalysis/*Node.csVarious node types updated to use Utf8String for mangled names

Comment threadsrc/coreclr/tools/Common/Compiler/ObjectWriter/StringTableBuilder.cs Outdated
@MichalStrehovsky

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-nativeaot-outerloop

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@MichalStrehovsky
MichalStrehovsky merged commit 8c8bfb2 into dotnet:mainDec 8, 2025
111 of 120 checks passed
@MichalStrehovsky
MichalStrehovsky deleted the utf8objwriter branch December 8, 2025 12:59

private Utf8String SanitizeNameWithHash(Utf8String literal)
{
Utf8String mangledName = SanitizeName(literal);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That code seems to be dangerous.
The result is not really an Utf8-String but ASCII - and it needs to be because the next few lines would be incorrect if it would be UTF8 and would contain any multi-byte chars (you can't simply cut off arbitrary Utf8 at byte position 30).
So if anybody ever changes SanitizeName to actually be Utf8 this will create hard-to-spot errors.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So if anybody ever changes SanitizeName to actually be Utf8 this will create hard-to-spot errors.

It's not likely we'd ever allow SanitizeName to return characters outside the basic ASCII set.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would a Debug.Assert(Ascii.IsValid(mangledName)) make sense here?

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@MichalStrehovsky@ANahr@filipnavara@am11@jkotas@PaulusParssinen@MichalPetryka
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Mostly switch object writer to UTF-8 - #122160

Merged
MichalStrehovsky merged 5 commits into
dotnet:mainfrom
MichalStrehovsky:utf8objwriter
Dec 8, 2025
Merged

Mostly switch object writer to UTF-8#122160
MichalStrehovsky merged 5 commits into
dotnet:mainfrom
MichalStrehovsky:utf8objwriter

Conversation

@MichalStrehovsky

@MichalStrehovskyMichalStrehovsky commented Dec 3, 2025

Copy link
Copy Markdown
Member

When compiling hello world:

byte[] allocations before: 564000. string allocations before: 302000
byte[] allocations after: 625000. string allocations after: 241000

So this is mostly a wash allocation-wise, however, not allocating string means we're also avoiding the UTF-8 -> UTF-16 -> UTF-8 conversions.

Cc @dotnet/ilc-contrib

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/tools/Common/Internal/Text/Utf8String.cs Outdated
@MichalStrehovskyMichalStrehovsky changed the title Switch object writer to UTF-8Mostly switch object writer to UTF-8Dec 5, 2025
@MichalStrehovsky
MichalStrehovsky marked this pull request as ready for review December 5, 2025 05:49
CopilotAI review requested due to automatic review settings December 5, 2025 05:49

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR converts the object writer subsystem from using string to Utf8String to avoid UTF-8 ↔ UTF-16 ↔ UTF-8 conversions during compilation. The change primarily affects symbol naming, mangling, and object file generation code.

Key Changes

  • Introduced IsNull property to Utf8String for null checking
  • Added multiple Concat overload methods for efficient UTF-8 string concatenation
  • Updated all object writer interfaces and implementations to accept/return Utf8String instead of string
  • Changed dictionaries and collections to use Utf8String keys

Reviewed changes

Copilot reviewed 44 out of 44 changed files in this pull request and generated 11 comments.

Show a summary per file
FileDescription
Internal/Text/Utf8String.csAdded IsNull property and new Concat overloads for UTF-8 string operations
ObjectWriter/ObjectWriter.csCore object writer changed to use Utf8String for symbol names and relocations
ObjectWriter/StringTableBuilder.csUpdated string table to work directly with UTF-8 bytes
ObjectWriter/CoffObjectWriter.csCOFF format writer updated to use Utf8String
ObjectWriter/ElfObjectWriter.csELF format writer updated to use Utf8String
ObjectWriter/MachObjectWriter.csMach-O format writer updated to use Utf8String
ObjectWriter/UnixObjectWriter.csUnix-specific object writer base updated
Compiler/NodeMangler.csName mangling infrastructure changed to return Utf8String
Compiler/WindowsNodeMangler.csWindows-specific name mangling updated with UTF-8 concatenation
Compiler/UnixNodeMangler.csUnix-specific name mangling updated with UTF-8 concatenation
Compiler/UserDefinedTypeDescriptor.csDebug info type descriptors updated to use Utf8String
DependencyAnalysis/NodeFactory.csFactory methods updated for Utf8String symbol names
DependencyAnalysis/*Node.csVarious node types updated to use Utf8String for mangled names

Comment threadsrc/coreclr/tools/Common/Compiler/ObjectWriter/StringTableBuilder.cs Outdated
@MichalStrehovsky

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-nativeaot-outerloop

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@MichalStrehovsky
MichalStrehovsky merged commit 8c8bfb2 into dotnet:mainDec 8, 2025
111 of 120 checks passed
@MichalStrehovsky
MichalStrehovsky deleted the utf8objwriter branch December 8, 2025 12:59

private Utf8String SanitizeNameWithHash(Utf8String literal)
{
Utf8String mangledName = SanitizeName(literal);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That code seems to be dangerous.
The result is not really an Utf8-String but ASCII - and it needs to be because the next few lines would be incorrect if it would be UTF8 and would contain any multi-byte chars (you can't simply cut off arbitrary Utf8 at byte position 30).
So if anybody ever changes SanitizeName to actually be Utf8 this will create hard-to-spot errors.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So if anybody ever changes SanitizeName to actually be Utf8 this will create hard-to-spot errors.

It's not likely we'd ever allow SanitizeName to return characters outside the basic ASCII set.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would a Debug.Assert(Ascii.IsValid(mangledName)) make sense here?

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@MichalStrehovsky@ANahr@filipnavara@am11@jkotas@PaulusParssinen@MichalPetryka
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Mostly switch object writer to UTF-8 - #122160

Merged
MichalStrehovsky merged 5 commits into
dotnet:mainfrom
MichalStrehovsky:utf8objwriter
Dec 8, 2025
Merged

Mostly switch object writer to UTF-8#122160
MichalStrehovsky merged 5 commits into
dotnet:mainfrom
MichalStrehovsky:utf8objwriter

Conversation

@MichalStrehovsky

@MichalStrehovskyMichalStrehovsky commented Dec 3, 2025

Copy link
Copy Markdown
Member

When compiling hello world:

byte[] allocations before: 564000. string allocations before: 302000
byte[] allocations after: 625000. string allocations after: 241000

So this is mostly a wash allocation-wise, however, not allocating string means we're also avoiding the UTF-8 -> UTF-16 -> UTF-8 conversions.

Cc @dotnet/ilc-contrib

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/tools/Common/Internal/Text/Utf8String.cs Outdated
@MichalStrehovskyMichalStrehovsky changed the title Switch object writer to UTF-8Mostly switch object writer to UTF-8Dec 5, 2025
@MichalStrehovsky
MichalStrehovsky marked this pull request as ready for review December 5, 2025 05:49
CopilotAI review requested due to automatic review settings December 5, 2025 05:49

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR converts the object writer subsystem from using string to Utf8String to avoid UTF-8 ↔ UTF-16 ↔ UTF-8 conversions during compilation. The change primarily affects symbol naming, mangling, and object file generation code.

Key Changes

  • Introduced IsNull property to Utf8String for null checking
  • Added multiple Concat overload methods for efficient UTF-8 string concatenation
  • Updated all object writer interfaces and implementations to accept/return Utf8String instead of string
  • Changed dictionaries and collections to use Utf8String keys

Reviewed changes

Copilot reviewed 44 out of 44 changed files in this pull request and generated 11 comments.

Show a summary per file
FileDescription
Internal/Text/Utf8String.csAdded IsNull property and new Concat overloads for UTF-8 string operations
ObjectWriter/ObjectWriter.csCore object writer changed to use Utf8String for symbol names and relocations
ObjectWriter/StringTableBuilder.csUpdated string table to work directly with UTF-8 bytes
ObjectWriter/CoffObjectWriter.csCOFF format writer updated to use Utf8String
ObjectWriter/ElfObjectWriter.csELF format writer updated to use Utf8String
ObjectWriter/MachObjectWriter.csMach-O format writer updated to use Utf8String
ObjectWriter/UnixObjectWriter.csUnix-specific object writer base updated
Compiler/NodeMangler.csName mangling infrastructure changed to return Utf8String
Compiler/WindowsNodeMangler.csWindows-specific name mangling updated with UTF-8 concatenation
Compiler/UnixNodeMangler.csUnix-specific name mangling updated with UTF-8 concatenation
Compiler/UserDefinedTypeDescriptor.csDebug info type descriptors updated to use Utf8String
DependencyAnalysis/NodeFactory.csFactory methods updated for Utf8String symbol names
DependencyAnalysis/*Node.csVarious node types updated to use Utf8String for mangled names

Comment threadsrc/coreclr/tools/Common/Compiler/ObjectWriter/StringTableBuilder.cs Outdated
@MichalStrehovsky

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-nativeaot-outerloop

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@MichalStrehovsky
MichalStrehovsky merged commit 8c8bfb2 into dotnet:mainDec 8, 2025
111 of 120 checks passed
@MichalStrehovsky
MichalStrehovsky deleted the utf8objwriter branch December 8, 2025 12:59

private Utf8String SanitizeNameWithHash(Utf8String literal)
{
Utf8String mangledName = SanitizeName(literal);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That code seems to be dangerous.
The result is not really an Utf8-String but ASCII - and it needs to be because the next few lines would be incorrect if it would be UTF8 and would contain any multi-byte chars (you can't simply cut off arbitrary Utf8 at byte position 30).
So if anybody ever changes SanitizeName to actually be Utf8 this will create hard-to-spot errors.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So if anybody ever changes SanitizeName to actually be Utf8 this will create hard-to-spot errors.

It's not likely we'd ever allow SanitizeName to return characters outside the basic ASCII set.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would a Debug.Assert(Ascii.IsValid(mangledName)) make sense here?

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@MichalStrehovsky@ANahr@filipnavara@am11@jkotas@PaulusParssinen@MichalPetryka
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Mostly switch object writer to UTF-8 - #122160

Merged
MichalStrehovsky merged 5 commits into
dotnet:mainfrom
MichalStrehovsky:utf8objwriter
Dec 8, 2025
Merged

Mostly switch object writer to UTF-8#122160
MichalStrehovsky merged 5 commits into
dotnet:mainfrom
MichalStrehovsky:utf8objwriter

Conversation

@MichalStrehovsky

@MichalStrehovskyMichalStrehovsky commented Dec 3, 2025

Copy link
Copy Markdown
Member

When compiling hello world:

byte[] allocations before: 564000. string allocations before: 302000
byte[] allocations after: 625000. string allocations after: 241000

So this is mostly a wash allocation-wise, however, not allocating string means we're also avoiding the UTF-8 -> UTF-16 -> UTF-8 conversions.

Cc @dotnet/ilc-contrib

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/tools/Common/Internal/Text/Utf8String.cs Outdated
@MichalStrehovskyMichalStrehovsky changed the title Switch object writer to UTF-8Mostly switch object writer to UTF-8Dec 5, 2025
@MichalStrehovsky
MichalStrehovsky marked this pull request as ready for review December 5, 2025 05:49
CopilotAI review requested due to automatic review settings December 5, 2025 05:49

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR converts the object writer subsystem from using string to Utf8String to avoid UTF-8 ↔ UTF-16 ↔ UTF-8 conversions during compilation. The change primarily affects symbol naming, mangling, and object file generation code.

Key Changes

  • Introduced IsNull property to Utf8String for null checking
  • Added multiple Concat overload methods for efficient UTF-8 string concatenation
  • Updated all object writer interfaces and implementations to accept/return Utf8String instead of string
  • Changed dictionaries and collections to use Utf8String keys

Reviewed changes

Copilot reviewed 44 out of 44 changed files in this pull request and generated 11 comments.

Show a summary per file
FileDescription
Internal/Text/Utf8String.csAdded IsNull property and new Concat overloads for UTF-8 string operations
ObjectWriter/ObjectWriter.csCore object writer changed to use Utf8String for symbol names and relocations
ObjectWriter/StringTableBuilder.csUpdated string table to work directly with UTF-8 bytes
ObjectWriter/CoffObjectWriter.csCOFF format writer updated to use Utf8String
ObjectWriter/ElfObjectWriter.csELF format writer updated to use Utf8String
ObjectWriter/MachObjectWriter.csMach-O format writer updated to use Utf8String
ObjectWriter/UnixObjectWriter.csUnix-specific object writer base updated
Compiler/NodeMangler.csName mangling infrastructure changed to return Utf8String
Compiler/WindowsNodeMangler.csWindows-specific name mangling updated with UTF-8 concatenation
Compiler/UnixNodeMangler.csUnix-specific name mangling updated with UTF-8 concatenation
Compiler/UserDefinedTypeDescriptor.csDebug info type descriptors updated to use Utf8String
DependencyAnalysis/NodeFactory.csFactory methods updated for Utf8String symbol names
DependencyAnalysis/*Node.csVarious node types updated to use Utf8String for mangled names

Comment threadsrc/coreclr/tools/Common/Compiler/ObjectWriter/StringTableBuilder.cs Outdated
@MichalStrehovsky

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-nativeaot-outerloop

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@MichalStrehovsky
MichalStrehovsky merged commit 8c8bfb2 into dotnet:mainDec 8, 2025
111 of 120 checks passed
@MichalStrehovsky
MichalStrehovsky deleted the utf8objwriter branch December 8, 2025 12:59

private Utf8String SanitizeNameWithHash(Utf8String literal)
{
Utf8String mangledName = SanitizeName(literal);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That code seems to be dangerous.
The result is not really an Utf8-String but ASCII - and it needs to be because the next few lines would be incorrect if it would be UTF8 and would contain any multi-byte chars (you can't simply cut off arbitrary Utf8 at byte position 30).
So if anybody ever changes SanitizeName to actually be Utf8 this will create hard-to-spot errors.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So if anybody ever changes SanitizeName to actually be Utf8 this will create hard-to-spot errors.

It's not likely we'd ever allow SanitizeName to return characters outside the basic ASCII set.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would a Debug.Assert(Ascii.IsValid(mangledName)) make sense here?

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@MichalStrehovsky@ANahr@filipnavara@am11@jkotas@PaulusParssinen@MichalPetryka
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Mostly switch object writer to UTF-8 - #122160

Merged
MichalStrehovsky merged 5 commits into
dotnet:mainfrom
MichalStrehovsky:utf8objwriter
Dec 8, 2025
Merged

Mostly switch object writer to UTF-8#122160
MichalStrehovsky merged 5 commits into
dotnet:mainfrom
MichalStrehovsky:utf8objwriter

Conversation

@MichalStrehovsky

@MichalStrehovskyMichalStrehovsky commented Dec 3, 2025

Copy link
Copy Markdown
Member

When compiling hello world:

byte[] allocations before: 564000. string allocations before: 302000
byte[] allocations after: 625000. string allocations after: 241000

So this is mostly a wash allocation-wise, however, not allocating string means we're also avoiding the UTF-8 -> UTF-16 -> UTF-8 conversions.

Cc @dotnet/ilc-contrib

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/tools/Common/Internal/Text/Utf8String.cs Outdated
@MichalStrehovskyMichalStrehovsky changed the title Switch object writer to UTF-8Mostly switch object writer to UTF-8Dec 5, 2025
@MichalStrehovsky
MichalStrehovsky marked this pull request as ready for review December 5, 2025 05:49
CopilotAI review requested due to automatic review settings December 5, 2025 05:49

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR converts the object writer subsystem from using string to Utf8String to avoid UTF-8 ↔ UTF-16 ↔ UTF-8 conversions during compilation. The change primarily affects symbol naming, mangling, and object file generation code.

Key Changes

  • Introduced IsNull property to Utf8String for null checking
  • Added multiple Concat overload methods for efficient UTF-8 string concatenation
  • Updated all object writer interfaces and implementations to accept/return Utf8String instead of string
  • Changed dictionaries and collections to use Utf8String keys

Reviewed changes

Copilot reviewed 44 out of 44 changed files in this pull request and generated 11 comments.

Show a summary per file
FileDescription
Internal/Text/Utf8String.csAdded IsNull property and new Concat overloads for UTF-8 string operations
ObjectWriter/ObjectWriter.csCore object writer changed to use Utf8String for symbol names and relocations
ObjectWriter/StringTableBuilder.csUpdated string table to work directly with UTF-8 bytes
ObjectWriter/CoffObjectWriter.csCOFF format writer updated to use Utf8String
ObjectWriter/ElfObjectWriter.csELF format writer updated to use Utf8String
ObjectWriter/MachObjectWriter.csMach-O format writer updated to use Utf8String
ObjectWriter/UnixObjectWriter.csUnix-specific object writer base updated
Compiler/NodeMangler.csName mangling infrastructure changed to return Utf8String
Compiler/WindowsNodeMangler.csWindows-specific name mangling updated with UTF-8 concatenation
Compiler/UnixNodeMangler.csUnix-specific name mangling updated with UTF-8 concatenation
Compiler/UserDefinedTypeDescriptor.csDebug info type descriptors updated to use Utf8String
DependencyAnalysis/NodeFactory.csFactory methods updated for Utf8String symbol names
DependencyAnalysis/*Node.csVarious node types updated to use Utf8String for mangled names

Comment threadsrc/coreclr/tools/Common/Compiler/ObjectWriter/StringTableBuilder.cs Outdated
@MichalStrehovsky

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-nativeaot-outerloop

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@MichalStrehovsky
MichalStrehovsky merged commit 8c8bfb2 into dotnet:mainDec 8, 2025
111 of 120 checks passed
@MichalStrehovsky
MichalStrehovsky deleted the utf8objwriter branch December 8, 2025 12:59

private Utf8String SanitizeNameWithHash(Utf8String literal)
{
Utf8String mangledName = SanitizeName(literal);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That code seems to be dangerous.
The result is not really an Utf8-String but ASCII - and it needs to be because the next few lines would be incorrect if it would be UTF8 and would contain any multi-byte chars (you can't simply cut off arbitrary Utf8 at byte position 30).
So if anybody ever changes SanitizeName to actually be Utf8 this will create hard-to-spot errors.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So if anybody ever changes SanitizeName to actually be Utf8 this will create hard-to-spot errors.

It's not likely we'd ever allow SanitizeName to return characters outside the basic ASCII set.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would a Debug.Assert(Ascii.IsValid(mangledName)) make sense here?

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@MichalStrehovsky@ANahr@filipnavara@am11@jkotas@PaulusParssinen@MichalPetryka
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Mostly switch object writer to UTF-8 - #122160

Merged
MichalStrehovsky merged 5 commits into
dotnet:mainfrom
MichalStrehovsky:utf8objwriter
Dec 8, 2025
Merged

Mostly switch object writer to UTF-8#122160
MichalStrehovsky merged 5 commits into
dotnet:mainfrom
MichalStrehovsky:utf8objwriter

Conversation

@MichalStrehovsky

@MichalStrehovskyMichalStrehovsky commented Dec 3, 2025

Copy link
Copy Markdown
Member

When compiling hello world:

byte[] allocations before: 564000. string allocations before: 302000
byte[] allocations after: 625000. string allocations after: 241000

So this is mostly a wash allocation-wise, however, not allocating string means we're also avoiding the UTF-8 -> UTF-16 -> UTF-8 conversions.

Cc @dotnet/ilc-contrib

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/tools/Common/Internal/Text/Utf8String.cs Outdated
@MichalStrehovskyMichalStrehovsky changed the title Switch object writer to UTF-8Mostly switch object writer to UTF-8Dec 5, 2025
@MichalStrehovsky
MichalStrehovsky marked this pull request as ready for review December 5, 2025 05:49
CopilotAI review requested due to automatic review settings December 5, 2025 05:49

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR converts the object writer subsystem from using string to Utf8String to avoid UTF-8 ↔ UTF-16 ↔ UTF-8 conversions during compilation. The change primarily affects symbol naming, mangling, and object file generation code.

Key Changes

  • Introduced IsNull property to Utf8String for null checking
  • Added multiple Concat overload methods for efficient UTF-8 string concatenation
  • Updated all object writer interfaces and implementations to accept/return Utf8String instead of string
  • Changed dictionaries and collections to use Utf8String keys

Reviewed changes

Copilot reviewed 44 out of 44 changed files in this pull request and generated 11 comments.

Show a summary per file
FileDescription
Internal/Text/Utf8String.csAdded IsNull property and new Concat overloads for UTF-8 string operations
ObjectWriter/ObjectWriter.csCore object writer changed to use Utf8String for symbol names and relocations
ObjectWriter/StringTableBuilder.csUpdated string table to work directly with UTF-8 bytes
ObjectWriter/CoffObjectWriter.csCOFF format writer updated to use Utf8String
ObjectWriter/ElfObjectWriter.csELF format writer updated to use Utf8String
ObjectWriter/MachObjectWriter.csMach-O format writer updated to use Utf8String
ObjectWriter/UnixObjectWriter.csUnix-specific object writer base updated
Compiler/NodeMangler.csName mangling infrastructure changed to return Utf8String
Compiler/WindowsNodeMangler.csWindows-specific name mangling updated with UTF-8 concatenation
Compiler/UnixNodeMangler.csUnix-specific name mangling updated with UTF-8 concatenation
Compiler/UserDefinedTypeDescriptor.csDebug info type descriptors updated to use Utf8String
DependencyAnalysis/NodeFactory.csFactory methods updated for Utf8String symbol names
DependencyAnalysis/*Node.csVarious node types updated to use Utf8String for mangled names

Comment threadsrc/coreclr/tools/Common/Compiler/ObjectWriter/StringTableBuilder.cs Outdated
@MichalStrehovsky

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-nativeaot-outerloop

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@MichalStrehovsky
MichalStrehovsky merged commit 8c8bfb2 into dotnet:mainDec 8, 2025
111 of 120 checks passed
@MichalStrehovsky
MichalStrehovsky deleted the utf8objwriter branch December 8, 2025 12:59

private Utf8String SanitizeNameWithHash(Utf8String literal)
{
Utf8String mangledName = SanitizeName(literal);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That code seems to be dangerous.
The result is not really an Utf8-String but ASCII - and it needs to be because the next few lines would be incorrect if it would be UTF8 and would contain any multi-byte chars (you can't simply cut off arbitrary Utf8 at byte position 30).
So if anybody ever changes SanitizeName to actually be Utf8 this will create hard-to-spot errors.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So if anybody ever changes SanitizeName to actually be Utf8 this will create hard-to-spot errors.

It's not likely we'd ever allow SanitizeName to return characters outside the basic ASCII set.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would a Debug.Assert(Ascii.IsValid(mangledName)) make sense here?

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants

@MichalStrehovsky@ANahr@filipnavara@am11@jkotas@PaulusParssinen@MichalPetryka