Use Unsafe.BitCast for Int128UInt128 operators - #104506

Merged
stephentoub merged 1 commit into
dotnet:mainfrom
xtqqczze:Int128BitCast
Jul 9, 2024
Merged

Use Unsafe.BitCast for Int128UInt128 operators#104506
stephentoub merged 1 commit into
dotnet:mainfrom
xtqqczze:Int128BitCast

Conversation

@xtqqczze

Copy link
Copy Markdown
Contributor

Diffs show increased inlining and tail calls.

MihuBot/runtime-utils#478
MihuBot/runtime-utils#479

@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jul 6, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-numerics
See info in area-owners.md if you want to be subscribed.

@tannergooding

Copy link
Copy Markdown
Member

This in general LGTM, but it might be interesting to understand why promotion didn't handle this given it's a simple struct containing 2x ulong fields.

CC. @jakobbotsch

@jakobbotsch

Copy link
Copy Markdown
Member

This in general LGTM, but it might be interesting to understand why promotion didn't handle this given it's a simple struct containing 2x ulong fields.

CC. @jakobbotsch

I would need some concrete cases to look at. As it is this looks like a size wise regression. @xtqqczze do you have concrete benchmarks showing improvements from the change?

@stephentoub

Copy link
Copy Markdown
Member

As it is this looks like a size wise regression

Is that not just because the system chose to inline more things as a result of this? If BitCast had been available when this code was first written, presumably we'd have chosen to use it then.

@jakobbotsch

Copy link
Copy Markdown
Member

Is that not just because the system chose to inline more things as a result of this? If BitCast had been available when this code was first written, presumably we'd have chosen to use it then.

Yeah, definitely looks like there are different inlining decisions here (both new inlines we perform and cases where we no longer inline, it looks like).

Also looks like the second diff is a larger size-wise improvement than the first diff is a regression, so in that sense this isn't actually a size-wise regression (but as you said, it's hard to compare in the face of different inlining decisions).

Either way I'd be happy to look at concrete cases if there are any, but I wasn't immediately able to identify anything that looks related to deficiencies in promotion in the diffs. If you and @tannergooding prefer BitCast over the explicit field accesses then certainly that's fine with me. I don't think the JIT should have any problem handling either pattern.

@tannergooding

tannergooding commented Jul 9, 2024

Copy link
Copy Markdown
Member

Is that not just because the system chose to inline more things as a result of this?

That's what it looks like to me. Size diffs are always a little wonky for xarch due to the variable sized encoding and differing encoding cost for some sizes or registers. LSRA choosing to use R8 instead of RAX can lead to an additional byte, for example.

In this case, it looks like we eliminate code, which allows inlining to kick in and that changes register preferences and some operation sizes causes the size to increase even though the number of instructions often decreases. For example in System.Int128:TryFormat(System.Span1[ushort],byref,System.ReadOnlySpan1[ushort],System.IFormatProvider):ubyte:this,

- mov rsi, rax- or rsi, 1- lzcnt rsi, rsi- xor esi, 63- movsxd rsi, esi+ mov rdi, rax+ or rdi, 1+ lzcnt rdi, rdi+ xor edi, 63+ movsxd rdi, edi+ mov rsi, 0xD1FFAB1E ; static handle+ movzx rdi, byte ptr [rdi+rsi]+ mov esi, edi
mov rdx, 0xD1FFAB1E ; static handle
- movzx rsi, byte ptr [rsi+rdx]- mov edx, esi- mov rcx, 0xD1FFAB1E ; static handle- cmp rax, qword ptr [rcx+8*rdx]- setb dl- movzx rdx, dl- sub esi, edx- lea eax, [rsi+0x14]+ cmp rax, qword ptr [rdx+8*rsi]+ setb sil+ movzx rsi, sil+ sub edi, esi+ lea eax, [rdi+0x14]
jmp SHORT G_M10567_IG08
- ;; size=106 bbWeight=0.50 PerfScore 9.25+ ;; size=108 bbWeight=0.50 PerfScore 9.25

This code is 2 bytes bigger, but it actually hasn't fundamentally changed and if you were to replace the registers used with abstract names like reg1 you'd find they're identical. The reason the latter is bigger is because LSRA decided to use RSI for the "second register" rather than RDX, this requires an extra byte to encode the setb sil and subsequent movzx rsi, sil.

We see quite a lot of diffs that are in this realm and it'd probably be beneficial to track the number of instructions in addition to the size so we can get a better view over what's actually changed.

Later on in the method we get 96 extra bytes that are "real" due to the call to UInt128.DivRem being inlined and thus us having 3 calls to op_Division, op_Multiply, and op_Subtraction instead
-- Notably there's also a call to the System.ValueTuple constructor which seems non-ideal given its just setting two fields
-- Ideally we'd also optimize DivRem to avoid needing the multiply/subtract. The main division algorithm can just return the remainder directly and avoid needing the extra work

@tannergooding

Copy link
Copy Markdown
Member

If you and @tannergooding prefer BitCast over the explicit field accesses then certainly that's fine with me. I don't think the JIT should have any problem handling either pattern.

I think I prefer BitCast just because it's less IL and gives a small inlining profitability boost due to being intrinsic. Had just seemed a little odd that there were other improvements beyond just inlining.

@stephentoub

Copy link
Copy Markdown
Member

Thanks.

@stephentoub
stephentoub merged commit 01dbf51 into dotnet:mainJul 9, 2024
matouskozak added a commit to matouskozak/runtime that referenced this pull request Jul 11, 2024
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 9, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-System.Numericscommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@xtqqczze@tannergooding@jakobbotsch@stephentoub
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Use Unsafe.BitCast for Int128UInt128 operators - #104506

Merged
stephentoub merged 1 commit into
dotnet:mainfrom
xtqqczze:Int128BitCast
Jul 9, 2024
Merged

Use Unsafe.BitCast for Int128UInt128 operators#104506
stephentoub merged 1 commit into
dotnet:mainfrom
xtqqczze:Int128BitCast

Conversation

@xtqqczze

Copy link
Copy Markdown
Contributor

Diffs show increased inlining and tail calls.

MihuBot/runtime-utils#478
MihuBot/runtime-utils#479

@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jul 6, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-numerics
See info in area-owners.md if you want to be subscribed.

@tannergooding

Copy link
Copy Markdown
Member

This in general LGTM, but it might be interesting to understand why promotion didn't handle this given it's a simple struct containing 2x ulong fields.

CC. @jakobbotsch

@jakobbotsch

Copy link
Copy Markdown
Member

This in general LGTM, but it might be interesting to understand why promotion didn't handle this given it's a simple struct containing 2x ulong fields.

CC. @jakobbotsch

I would need some concrete cases to look at. As it is this looks like a size wise regression. @xtqqczze do you have concrete benchmarks showing improvements from the change?

@stephentoub

Copy link
Copy Markdown
Member

As it is this looks like a size wise regression

Is that not just because the system chose to inline more things as a result of this? If BitCast had been available when this code was first written, presumably we'd have chosen to use it then.

@jakobbotsch

Copy link
Copy Markdown
Member

Is that not just because the system chose to inline more things as a result of this? If BitCast had been available when this code was first written, presumably we'd have chosen to use it then.

Yeah, definitely looks like there are different inlining decisions here (both new inlines we perform and cases where we no longer inline, it looks like).

Also looks like the second diff is a larger size-wise improvement than the first diff is a regression, so in that sense this isn't actually a size-wise regression (but as you said, it's hard to compare in the face of different inlining decisions).

Either way I'd be happy to look at concrete cases if there are any, but I wasn't immediately able to identify anything that looks related to deficiencies in promotion in the diffs. If you and @tannergooding prefer BitCast over the explicit field accesses then certainly that's fine with me. I don't think the JIT should have any problem handling either pattern.

@tannergooding

tannergooding commented Jul 9, 2024

Copy link
Copy Markdown
Member

Is that not just because the system chose to inline more things as a result of this?

That's what it looks like to me. Size diffs are always a little wonky for xarch due to the variable sized encoding and differing encoding cost for some sizes or registers. LSRA choosing to use R8 instead of RAX can lead to an additional byte, for example.

In this case, it looks like we eliminate code, which allows inlining to kick in and that changes register preferences and some operation sizes causes the size to increase even though the number of instructions often decreases. For example in System.Int128:TryFormat(System.Span1[ushort],byref,System.ReadOnlySpan1[ushort],System.IFormatProvider):ubyte:this,

- mov rsi, rax- or rsi, 1- lzcnt rsi, rsi- xor esi, 63- movsxd rsi, esi+ mov rdi, rax+ or rdi, 1+ lzcnt rdi, rdi+ xor edi, 63+ movsxd rdi, edi+ mov rsi, 0xD1FFAB1E ; static handle+ movzx rdi, byte ptr [rdi+rsi]+ mov esi, edi
mov rdx, 0xD1FFAB1E ; static handle
- movzx rsi, byte ptr [rsi+rdx]- mov edx, esi- mov rcx, 0xD1FFAB1E ; static handle- cmp rax, qword ptr [rcx+8*rdx]- setb dl- movzx rdx, dl- sub esi, edx- lea eax, [rsi+0x14]+ cmp rax, qword ptr [rdx+8*rsi]+ setb sil+ movzx rsi, sil+ sub edi, esi+ lea eax, [rdi+0x14]
jmp SHORT G_M10567_IG08
- ;; size=106 bbWeight=0.50 PerfScore 9.25+ ;; size=108 bbWeight=0.50 PerfScore 9.25

This code is 2 bytes bigger, but it actually hasn't fundamentally changed and if you were to replace the registers used with abstract names like reg1 you'd find they're identical. The reason the latter is bigger is because LSRA decided to use RSI for the "second register" rather than RDX, this requires an extra byte to encode the setb sil and subsequent movzx rsi, sil.

We see quite a lot of diffs that are in this realm and it'd probably be beneficial to track the number of instructions in addition to the size so we can get a better view over what's actually changed.

Later on in the method we get 96 extra bytes that are "real" due to the call to UInt128.DivRem being inlined and thus us having 3 calls to op_Division, op_Multiply, and op_Subtraction instead
-- Notably there's also a call to the System.ValueTuple constructor which seems non-ideal given its just setting two fields
-- Ideally we'd also optimize DivRem to avoid needing the multiply/subtract. The main division algorithm can just return the remainder directly and avoid needing the extra work

@tannergooding

Copy link
Copy Markdown
Member

If you and @tannergooding prefer BitCast over the explicit field accesses then certainly that's fine with me. I don't think the JIT should have any problem handling either pattern.

I think I prefer BitCast just because it's less IL and gives a small inlining profitability boost due to being intrinsic. Had just seemed a little odd that there were other improvements beyond just inlining.

@stephentoub

Copy link
Copy Markdown
Member

Thanks.

@stephentoub
stephentoub merged commit 01dbf51 into dotnet:mainJul 9, 2024
matouskozak added a commit to matouskozak/runtime that referenced this pull request Jul 11, 2024
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 9, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-System.Numericscommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@xtqqczze@tannergooding@jakobbotsch@stephentoub
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use Unsafe.BitCast for Int128UInt128 operators - #104506

Merged
stephentoub merged 1 commit into
dotnet:mainfrom
xtqqczze:Int128BitCast
Jul 9, 2024
Merged

Use Unsafe.BitCast for Int128UInt128 operators#104506
stephentoub merged 1 commit into
dotnet:mainfrom
xtqqczze:Int128BitCast

Conversation

@xtqqczze

Copy link
Copy Markdown
Contributor

Diffs show increased inlining and tail calls.

MihuBot/runtime-utils#478
MihuBot/runtime-utils#479

@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jul 6, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-numerics
See info in area-owners.md if you want to be subscribed.

@tannergooding

Copy link
Copy Markdown
Member

This in general LGTM, but it might be interesting to understand why promotion didn't handle this given it's a simple struct containing 2x ulong fields.

CC. @jakobbotsch

@jakobbotsch

Copy link
Copy Markdown
Member

This in general LGTM, but it might be interesting to understand why promotion didn't handle this given it's a simple struct containing 2x ulong fields.

CC. @jakobbotsch

I would need some concrete cases to look at. As it is this looks like a size wise regression. @xtqqczze do you have concrete benchmarks showing improvements from the change?

@stephentoub

Copy link
Copy Markdown
Member

As it is this looks like a size wise regression

Is that not just because the system chose to inline more things as a result of this? If BitCast had been available when this code was first written, presumably we'd have chosen to use it then.

@jakobbotsch

Copy link
Copy Markdown
Member

Is that not just because the system chose to inline more things as a result of this? If BitCast had been available when this code was first written, presumably we'd have chosen to use it then.

Yeah, definitely looks like there are different inlining decisions here (both new inlines we perform and cases where we no longer inline, it looks like).

Also looks like the second diff is a larger size-wise improvement than the first diff is a regression, so in that sense this isn't actually a size-wise regression (but as you said, it's hard to compare in the face of different inlining decisions).

Either way I'd be happy to look at concrete cases if there are any, but I wasn't immediately able to identify anything that looks related to deficiencies in promotion in the diffs. If you and @tannergooding prefer BitCast over the explicit field accesses then certainly that's fine with me. I don't think the JIT should have any problem handling either pattern.

@tannergooding

tannergooding commented Jul 9, 2024

Copy link
Copy Markdown
Member

Is that not just because the system chose to inline more things as a result of this?

That's what it looks like to me. Size diffs are always a little wonky for xarch due to the variable sized encoding and differing encoding cost for some sizes or registers. LSRA choosing to use R8 instead of RAX can lead to an additional byte, for example.

In this case, it looks like we eliminate code, which allows inlining to kick in and that changes register preferences and some operation sizes causes the size to increase even though the number of instructions often decreases. For example in System.Int128:TryFormat(System.Span1[ushort],byref,System.ReadOnlySpan1[ushort],System.IFormatProvider):ubyte:this,

- mov rsi, rax- or rsi, 1- lzcnt rsi, rsi- xor esi, 63- movsxd rsi, esi+ mov rdi, rax+ or rdi, 1+ lzcnt rdi, rdi+ xor edi, 63+ movsxd rdi, edi+ mov rsi, 0xD1FFAB1E ; static handle+ movzx rdi, byte ptr [rdi+rsi]+ mov esi, edi
mov rdx, 0xD1FFAB1E ; static handle
- movzx rsi, byte ptr [rsi+rdx]- mov edx, esi- mov rcx, 0xD1FFAB1E ; static handle- cmp rax, qword ptr [rcx+8*rdx]- setb dl- movzx rdx, dl- sub esi, edx- lea eax, [rsi+0x14]+ cmp rax, qword ptr [rdx+8*rsi]+ setb sil+ movzx rsi, sil+ sub edi, esi+ lea eax, [rdi+0x14]
jmp SHORT G_M10567_IG08
- ;; size=106 bbWeight=0.50 PerfScore 9.25+ ;; size=108 bbWeight=0.50 PerfScore 9.25

This code is 2 bytes bigger, but it actually hasn't fundamentally changed and if you were to replace the registers used with abstract names like reg1 you'd find they're identical. The reason the latter is bigger is because LSRA decided to use RSI for the "second register" rather than RDX, this requires an extra byte to encode the setb sil and subsequent movzx rsi, sil.

We see quite a lot of diffs that are in this realm and it'd probably be beneficial to track the number of instructions in addition to the size so we can get a better view over what's actually changed.

Later on in the method we get 96 extra bytes that are "real" due to the call to UInt128.DivRem being inlined and thus us having 3 calls to op_Division, op_Multiply, and op_Subtraction instead
-- Notably there's also a call to the System.ValueTuple constructor which seems non-ideal given its just setting two fields
-- Ideally we'd also optimize DivRem to avoid needing the multiply/subtract. The main division algorithm can just return the remainder directly and avoid needing the extra work

@tannergooding

Copy link
Copy Markdown
Member

If you and @tannergooding prefer BitCast over the explicit field accesses then certainly that's fine with me. I don't think the JIT should have any problem handling either pattern.

I think I prefer BitCast just because it's less IL and gives a small inlining profitability boost due to being intrinsic. Had just seemed a little odd that there were other improvements beyond just inlining.

@stephentoub

Copy link
Copy Markdown
Member

Thanks.

@stephentoub
stephentoub merged commit 01dbf51 into dotnet:mainJul 9, 2024
matouskozak added a commit to matouskozak/runtime that referenced this pull request Jul 11, 2024
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 9, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-System.Numericscommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@xtqqczze@tannergooding@jakobbotsch@stephentoub
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use Unsafe.BitCast for Int128UInt128 operators - #104506

Merged
stephentoub merged 1 commit into
dotnet:mainfrom
xtqqczze:Int128BitCast
Jul 9, 2024
Merged

Use Unsafe.BitCast for Int128UInt128 operators#104506
stephentoub merged 1 commit into
dotnet:mainfrom
xtqqczze:Int128BitCast

Conversation

@xtqqczze

Copy link
Copy Markdown
Contributor

Diffs show increased inlining and tail calls.

MihuBot/runtime-utils#478
MihuBot/runtime-utils#479

@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jul 6, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-numerics
See info in area-owners.md if you want to be subscribed.

@tannergooding

Copy link
Copy Markdown
Member

This in general LGTM, but it might be interesting to understand why promotion didn't handle this given it's a simple struct containing 2x ulong fields.

CC. @jakobbotsch

@jakobbotsch

Copy link
Copy Markdown
Member

This in general LGTM, but it might be interesting to understand why promotion didn't handle this given it's a simple struct containing 2x ulong fields.

CC. @jakobbotsch

I would need some concrete cases to look at. As it is this looks like a size wise regression. @xtqqczze do you have concrete benchmarks showing improvements from the change?

@stephentoub

Copy link
Copy Markdown
Member

As it is this looks like a size wise regression

Is that not just because the system chose to inline more things as a result of this? If BitCast had been available when this code was first written, presumably we'd have chosen to use it then.

@jakobbotsch

Copy link
Copy Markdown
Member

Is that not just because the system chose to inline more things as a result of this? If BitCast had been available when this code was first written, presumably we'd have chosen to use it then.

Yeah, definitely looks like there are different inlining decisions here (both new inlines we perform and cases where we no longer inline, it looks like).

Also looks like the second diff is a larger size-wise improvement than the first diff is a regression, so in that sense this isn't actually a size-wise regression (but as you said, it's hard to compare in the face of different inlining decisions).

Either way I'd be happy to look at concrete cases if there are any, but I wasn't immediately able to identify anything that looks related to deficiencies in promotion in the diffs. If you and @tannergooding prefer BitCast over the explicit field accesses then certainly that's fine with me. I don't think the JIT should have any problem handling either pattern.

@tannergooding

tannergooding commented Jul 9, 2024

Copy link
Copy Markdown
Member

Is that not just because the system chose to inline more things as a result of this?

That's what it looks like to me. Size diffs are always a little wonky for xarch due to the variable sized encoding and differing encoding cost for some sizes or registers. LSRA choosing to use R8 instead of RAX can lead to an additional byte, for example.

In this case, it looks like we eliminate code, which allows inlining to kick in and that changes register preferences and some operation sizes causes the size to increase even though the number of instructions often decreases. For example in System.Int128:TryFormat(System.Span1[ushort],byref,System.ReadOnlySpan1[ushort],System.IFormatProvider):ubyte:this,

- mov rsi, rax- or rsi, 1- lzcnt rsi, rsi- xor esi, 63- movsxd rsi, esi+ mov rdi, rax+ or rdi, 1+ lzcnt rdi, rdi+ xor edi, 63+ movsxd rdi, edi+ mov rsi, 0xD1FFAB1E ; static handle+ movzx rdi, byte ptr [rdi+rsi]+ mov esi, edi
mov rdx, 0xD1FFAB1E ; static handle
- movzx rsi, byte ptr [rsi+rdx]- mov edx, esi- mov rcx, 0xD1FFAB1E ; static handle- cmp rax, qword ptr [rcx+8*rdx]- setb dl- movzx rdx, dl- sub esi, edx- lea eax, [rsi+0x14]+ cmp rax, qword ptr [rdx+8*rsi]+ setb sil+ movzx rsi, sil+ sub edi, esi+ lea eax, [rdi+0x14]
jmp SHORT G_M10567_IG08
- ;; size=106 bbWeight=0.50 PerfScore 9.25+ ;; size=108 bbWeight=0.50 PerfScore 9.25

This code is 2 bytes bigger, but it actually hasn't fundamentally changed and if you were to replace the registers used with abstract names like reg1 you'd find they're identical. The reason the latter is bigger is because LSRA decided to use RSI for the "second register" rather than RDX, this requires an extra byte to encode the setb sil and subsequent movzx rsi, sil.

We see quite a lot of diffs that are in this realm and it'd probably be beneficial to track the number of instructions in addition to the size so we can get a better view over what's actually changed.

Later on in the method we get 96 extra bytes that are "real" due to the call to UInt128.DivRem being inlined and thus us having 3 calls to op_Division, op_Multiply, and op_Subtraction instead
-- Notably there's also a call to the System.ValueTuple constructor which seems non-ideal given its just setting two fields
-- Ideally we'd also optimize DivRem to avoid needing the multiply/subtract. The main division algorithm can just return the remainder directly and avoid needing the extra work

@tannergooding

Copy link
Copy Markdown
Member

If you and @tannergooding prefer BitCast over the explicit field accesses then certainly that's fine with me. I don't think the JIT should have any problem handling either pattern.

I think I prefer BitCast just because it's less IL and gives a small inlining profitability boost due to being intrinsic. Had just seemed a little odd that there were other improvements beyond just inlining.

@stephentoub

Copy link
Copy Markdown
Member

Thanks.

@stephentoub
stephentoub merged commit 01dbf51 into dotnet:mainJul 9, 2024
matouskozak added a commit to matouskozak/runtime that referenced this pull request Jul 11, 2024
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 9, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-System.Numericscommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@xtqqczze@tannergooding@jakobbotsch@stephentoub
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Use Unsafe.BitCast for Int128UInt128 operators - #104506

Merged
stephentoub merged 1 commit into
dotnet:mainfrom
xtqqczze:Int128BitCast
Jul 9, 2024
Merged

Use Unsafe.BitCast for Int128UInt128 operators#104506
stephentoub merged 1 commit into
dotnet:mainfrom
xtqqczze:Int128BitCast

Conversation

@xtqqczze

Copy link
Copy Markdown
Contributor

Diffs show increased inlining and tail calls.

MihuBot/runtime-utils#478
MihuBot/runtime-utils#479

@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jul 6, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-numerics
See info in area-owners.md if you want to be subscribed.

@tannergooding

Copy link
Copy Markdown
Member

This in general LGTM, but it might be interesting to understand why promotion didn't handle this given it's a simple struct containing 2x ulong fields.

CC. @jakobbotsch

@jakobbotsch

Copy link
Copy Markdown
Member

This in general LGTM, but it might be interesting to understand why promotion didn't handle this given it's a simple struct containing 2x ulong fields.

CC. @jakobbotsch

I would need some concrete cases to look at. As it is this looks like a size wise regression. @xtqqczze do you have concrete benchmarks showing improvements from the change?

@stephentoub

Copy link
Copy Markdown
Member

As it is this looks like a size wise regression

Is that not just because the system chose to inline more things as a result of this? If BitCast had been available when this code was first written, presumably we'd have chosen to use it then.

@jakobbotsch

Copy link
Copy Markdown
Member

Is that not just because the system chose to inline more things as a result of this? If BitCast had been available when this code was first written, presumably we'd have chosen to use it then.

Yeah, definitely looks like there are different inlining decisions here (both new inlines we perform and cases where we no longer inline, it looks like).

Also looks like the second diff is a larger size-wise improvement than the first diff is a regression, so in that sense this isn't actually a size-wise regression (but as you said, it's hard to compare in the face of different inlining decisions).

Either way I'd be happy to look at concrete cases if there are any, but I wasn't immediately able to identify anything that looks related to deficiencies in promotion in the diffs. If you and @tannergooding prefer BitCast over the explicit field accesses then certainly that's fine with me. I don't think the JIT should have any problem handling either pattern.

@tannergooding

tannergooding commented Jul 9, 2024

Copy link
Copy Markdown
Member

Is that not just because the system chose to inline more things as a result of this?

That's what it looks like to me. Size diffs are always a little wonky for xarch due to the variable sized encoding and differing encoding cost for some sizes or registers. LSRA choosing to use R8 instead of RAX can lead to an additional byte, for example.

In this case, it looks like we eliminate code, which allows inlining to kick in and that changes register preferences and some operation sizes causes the size to increase even though the number of instructions often decreases. For example in System.Int128:TryFormat(System.Span1[ushort],byref,System.ReadOnlySpan1[ushort],System.IFormatProvider):ubyte:this,

- mov rsi, rax- or rsi, 1- lzcnt rsi, rsi- xor esi, 63- movsxd rsi, esi+ mov rdi, rax+ or rdi, 1+ lzcnt rdi, rdi+ xor edi, 63+ movsxd rdi, edi+ mov rsi, 0xD1FFAB1E ; static handle+ movzx rdi, byte ptr [rdi+rsi]+ mov esi, edi
mov rdx, 0xD1FFAB1E ; static handle
- movzx rsi, byte ptr [rsi+rdx]- mov edx, esi- mov rcx, 0xD1FFAB1E ; static handle- cmp rax, qword ptr [rcx+8*rdx]- setb dl- movzx rdx, dl- sub esi, edx- lea eax, [rsi+0x14]+ cmp rax, qword ptr [rdx+8*rsi]+ setb sil+ movzx rsi, sil+ sub edi, esi+ lea eax, [rdi+0x14]
jmp SHORT G_M10567_IG08
- ;; size=106 bbWeight=0.50 PerfScore 9.25+ ;; size=108 bbWeight=0.50 PerfScore 9.25

This code is 2 bytes bigger, but it actually hasn't fundamentally changed and if you were to replace the registers used with abstract names like reg1 you'd find they're identical. The reason the latter is bigger is because LSRA decided to use RSI for the "second register" rather than RDX, this requires an extra byte to encode the setb sil and subsequent movzx rsi, sil.

We see quite a lot of diffs that are in this realm and it'd probably be beneficial to track the number of instructions in addition to the size so we can get a better view over what's actually changed.

Later on in the method we get 96 extra bytes that are "real" due to the call to UInt128.DivRem being inlined and thus us having 3 calls to op_Division, op_Multiply, and op_Subtraction instead
-- Notably there's also a call to the System.ValueTuple constructor which seems non-ideal given its just setting two fields
-- Ideally we'd also optimize DivRem to avoid needing the multiply/subtract. The main division algorithm can just return the remainder directly and avoid needing the extra work

@tannergooding

Copy link
Copy Markdown
Member

If you and @tannergooding prefer BitCast over the explicit field accesses then certainly that's fine with me. I don't think the JIT should have any problem handling either pattern.

I think I prefer BitCast just because it's less IL and gives a small inlining profitability boost due to being intrinsic. Had just seemed a little odd that there were other improvements beyond just inlining.

@stephentoub

Copy link
Copy Markdown
Member

Thanks.

@stephentoub
stephentoub merged commit 01dbf51 into dotnet:mainJul 9, 2024
matouskozak added a commit to matouskozak/runtime that referenced this pull request Jul 11, 2024
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 9, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-System.Numericscommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@xtqqczze@tannergooding@jakobbotsch@stephentoub
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use Unsafe.BitCast for Int128UInt128 operators - #104506

Merged
stephentoub merged 1 commit into
dotnet:mainfrom
xtqqczze:Int128BitCast
Jul 9, 2024
Merged

Use Unsafe.BitCast for Int128UInt128 operators#104506
stephentoub merged 1 commit into
dotnet:mainfrom
xtqqczze:Int128BitCast

Conversation

@xtqqczze

Copy link
Copy Markdown
Contributor

Diffs show increased inlining and tail calls.

MihuBot/runtime-utils#478
MihuBot/runtime-utils#479

@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jul 6, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-numerics
See info in area-owners.md if you want to be subscribed.

@tannergooding

Copy link
Copy Markdown
Member

This in general LGTM, but it might be interesting to understand why promotion didn't handle this given it's a simple struct containing 2x ulong fields.

CC. @jakobbotsch

@jakobbotsch

Copy link
Copy Markdown
Member

This in general LGTM, but it might be interesting to understand why promotion didn't handle this given it's a simple struct containing 2x ulong fields.

CC. @jakobbotsch

I would need some concrete cases to look at. As it is this looks like a size wise regression. @xtqqczze do you have concrete benchmarks showing improvements from the change?

@stephentoub

Copy link
Copy Markdown
Member

As it is this looks like a size wise regression

Is that not just because the system chose to inline more things as a result of this? If BitCast had been available when this code was first written, presumably we'd have chosen to use it then.

@jakobbotsch

Copy link
Copy Markdown
Member

Is that not just because the system chose to inline more things as a result of this? If BitCast had been available when this code was first written, presumably we'd have chosen to use it then.

Yeah, definitely looks like there are different inlining decisions here (both new inlines we perform and cases where we no longer inline, it looks like).

Also looks like the second diff is a larger size-wise improvement than the first diff is a regression, so in that sense this isn't actually a size-wise regression (but as you said, it's hard to compare in the face of different inlining decisions).

Either way I'd be happy to look at concrete cases if there are any, but I wasn't immediately able to identify anything that looks related to deficiencies in promotion in the diffs. If you and @tannergooding prefer BitCast over the explicit field accesses then certainly that's fine with me. I don't think the JIT should have any problem handling either pattern.

@tannergooding

tannergooding commented Jul 9, 2024

Copy link
Copy Markdown
Member

Is that not just because the system chose to inline more things as a result of this?

That's what it looks like to me. Size diffs are always a little wonky for xarch due to the variable sized encoding and differing encoding cost for some sizes or registers. LSRA choosing to use R8 instead of RAX can lead to an additional byte, for example.

In this case, it looks like we eliminate code, which allows inlining to kick in and that changes register preferences and some operation sizes causes the size to increase even though the number of instructions often decreases. For example in System.Int128:TryFormat(System.Span1[ushort],byref,System.ReadOnlySpan1[ushort],System.IFormatProvider):ubyte:this,

- mov rsi, rax- or rsi, 1- lzcnt rsi, rsi- xor esi, 63- movsxd rsi, esi+ mov rdi, rax+ or rdi, 1+ lzcnt rdi, rdi+ xor edi, 63+ movsxd rdi, edi+ mov rsi, 0xD1FFAB1E ; static handle+ movzx rdi, byte ptr [rdi+rsi]+ mov esi, edi
mov rdx, 0xD1FFAB1E ; static handle
- movzx rsi, byte ptr [rsi+rdx]- mov edx, esi- mov rcx, 0xD1FFAB1E ; static handle- cmp rax, qword ptr [rcx+8*rdx]- setb dl- movzx rdx, dl- sub esi, edx- lea eax, [rsi+0x14]+ cmp rax, qword ptr [rdx+8*rsi]+ setb sil+ movzx rsi, sil+ sub edi, esi+ lea eax, [rdi+0x14]
jmp SHORT G_M10567_IG08
- ;; size=106 bbWeight=0.50 PerfScore 9.25+ ;; size=108 bbWeight=0.50 PerfScore 9.25

This code is 2 bytes bigger, but it actually hasn't fundamentally changed and if you were to replace the registers used with abstract names like reg1 you'd find they're identical. The reason the latter is bigger is because LSRA decided to use RSI for the "second register" rather than RDX, this requires an extra byte to encode the setb sil and subsequent movzx rsi, sil.

We see quite a lot of diffs that are in this realm and it'd probably be beneficial to track the number of instructions in addition to the size so we can get a better view over what's actually changed.

Later on in the method we get 96 extra bytes that are "real" due to the call to UInt128.DivRem being inlined and thus us having 3 calls to op_Division, op_Multiply, and op_Subtraction instead
-- Notably there's also a call to the System.ValueTuple constructor which seems non-ideal given its just setting two fields
-- Ideally we'd also optimize DivRem to avoid needing the multiply/subtract. The main division algorithm can just return the remainder directly and avoid needing the extra work

@tannergooding

Copy link
Copy Markdown
Member

If you and @tannergooding prefer BitCast over the explicit field accesses then certainly that's fine with me. I don't think the JIT should have any problem handling either pattern.

I think I prefer BitCast just because it's less IL and gives a small inlining profitability boost due to being intrinsic. Had just seemed a little odd that there were other improvements beyond just inlining.

@stephentoub

Copy link
Copy Markdown
Member

Thanks.

@stephentoub
stephentoub merged commit 01dbf51 into dotnet:mainJul 9, 2024
matouskozak added a commit to matouskozak/runtime that referenced this pull request Jul 11, 2024
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 9, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-System.Numericscommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@xtqqczze@tannergooding@jakobbotsch@stephentoub
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use Unsafe.BitCast for Int128UInt128 operators - #104506

Merged
stephentoub merged 1 commit into
dotnet:mainfrom
xtqqczze:Int128BitCast
Jul 9, 2024
Merged

Use Unsafe.BitCast for Int128UInt128 operators#104506
stephentoub merged 1 commit into
dotnet:mainfrom
xtqqczze:Int128BitCast

Conversation

@xtqqczze

Copy link
Copy Markdown
Contributor

Diffs show increased inlining and tail calls.

MihuBot/runtime-utils#478
MihuBot/runtime-utils#479

@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jul 6, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-numerics
See info in area-owners.md if you want to be subscribed.

@tannergooding

Copy link
Copy Markdown
Member

This in general LGTM, but it might be interesting to understand why promotion didn't handle this given it's a simple struct containing 2x ulong fields.

CC. @jakobbotsch

@jakobbotsch

Copy link
Copy Markdown
Member

This in general LGTM, but it might be interesting to understand why promotion didn't handle this given it's a simple struct containing 2x ulong fields.

CC. @jakobbotsch

I would need some concrete cases to look at. As it is this looks like a size wise regression. @xtqqczze do you have concrete benchmarks showing improvements from the change?

@stephentoub

Copy link
Copy Markdown
Member

As it is this looks like a size wise regression

Is that not just because the system chose to inline more things as a result of this? If BitCast had been available when this code was first written, presumably we'd have chosen to use it then.

@jakobbotsch

Copy link
Copy Markdown
Member

Is that not just because the system chose to inline more things as a result of this? If BitCast had been available when this code was first written, presumably we'd have chosen to use it then.

Yeah, definitely looks like there are different inlining decisions here (both new inlines we perform and cases where we no longer inline, it looks like).

Also looks like the second diff is a larger size-wise improvement than the first diff is a regression, so in that sense this isn't actually a size-wise regression (but as you said, it's hard to compare in the face of different inlining decisions).

Either way I'd be happy to look at concrete cases if there are any, but I wasn't immediately able to identify anything that looks related to deficiencies in promotion in the diffs. If you and @tannergooding prefer BitCast over the explicit field accesses then certainly that's fine with me. I don't think the JIT should have any problem handling either pattern.

@tannergooding

tannergooding commented Jul 9, 2024

Copy link
Copy Markdown
Member

Is that not just because the system chose to inline more things as a result of this?

That's what it looks like to me. Size diffs are always a little wonky for xarch due to the variable sized encoding and differing encoding cost for some sizes or registers. LSRA choosing to use R8 instead of RAX can lead to an additional byte, for example.

In this case, it looks like we eliminate code, which allows inlining to kick in and that changes register preferences and some operation sizes causes the size to increase even though the number of instructions often decreases. For example in System.Int128:TryFormat(System.Span1[ushort],byref,System.ReadOnlySpan1[ushort],System.IFormatProvider):ubyte:this,

- mov rsi, rax- or rsi, 1- lzcnt rsi, rsi- xor esi, 63- movsxd rsi, esi+ mov rdi, rax+ or rdi, 1+ lzcnt rdi, rdi+ xor edi, 63+ movsxd rdi, edi+ mov rsi, 0xD1FFAB1E ; static handle+ movzx rdi, byte ptr [rdi+rsi]+ mov esi, edi
mov rdx, 0xD1FFAB1E ; static handle
- movzx rsi, byte ptr [rsi+rdx]- mov edx, esi- mov rcx, 0xD1FFAB1E ; static handle- cmp rax, qword ptr [rcx+8*rdx]- setb dl- movzx rdx, dl- sub esi, edx- lea eax, [rsi+0x14]+ cmp rax, qword ptr [rdx+8*rsi]+ setb sil+ movzx rsi, sil+ sub edi, esi+ lea eax, [rdi+0x14]
jmp SHORT G_M10567_IG08
- ;; size=106 bbWeight=0.50 PerfScore 9.25+ ;; size=108 bbWeight=0.50 PerfScore 9.25

This code is 2 bytes bigger, but it actually hasn't fundamentally changed and if you were to replace the registers used with abstract names like reg1 you'd find they're identical. The reason the latter is bigger is because LSRA decided to use RSI for the "second register" rather than RDX, this requires an extra byte to encode the setb sil and subsequent movzx rsi, sil.

We see quite a lot of diffs that are in this realm and it'd probably be beneficial to track the number of instructions in addition to the size so we can get a better view over what's actually changed.

Later on in the method we get 96 extra bytes that are "real" due to the call to UInt128.DivRem being inlined and thus us having 3 calls to op_Division, op_Multiply, and op_Subtraction instead
-- Notably there's also a call to the System.ValueTuple constructor which seems non-ideal given its just setting two fields
-- Ideally we'd also optimize DivRem to avoid needing the multiply/subtract. The main division algorithm can just return the remainder directly and avoid needing the extra work

@tannergooding

Copy link
Copy Markdown
Member

If you and @tannergooding prefer BitCast over the explicit field accesses then certainly that's fine with me. I don't think the JIT should have any problem handling either pattern.

I think I prefer BitCast just because it's less IL and gives a small inlining profitability boost due to being intrinsic. Had just seemed a little odd that there were other improvements beyond just inlining.

@stephentoub

Copy link
Copy Markdown
Member

Thanks.

@stephentoub
stephentoub merged commit 01dbf51 into dotnet:mainJul 9, 2024
matouskozak added a commit to matouskozak/runtime that referenced this pull request Jul 11, 2024
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 9, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-System.Numericscommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@xtqqczze@tannergooding@jakobbotsch@stephentoub
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Use Unsafe.BitCast for Int128UInt128 operators - #104506

Merged
stephentoub merged 1 commit into
dotnet:mainfrom
xtqqczze:Int128BitCast
Jul 9, 2024
Merged

Use Unsafe.BitCast for Int128UInt128 operators#104506
stephentoub merged 1 commit into
dotnet:mainfrom
xtqqczze:Int128BitCast

Conversation

@xtqqczze

Copy link
Copy Markdown
Contributor

Diffs show increased inlining and tail calls.

MihuBot/runtime-utils#478
MihuBot/runtime-utils#479

@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jul 6, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-numerics
See info in area-owners.md if you want to be subscribed.

@tannergooding

Copy link
Copy Markdown
Member

This in general LGTM, but it might be interesting to understand why promotion didn't handle this given it's a simple struct containing 2x ulong fields.

CC. @jakobbotsch

@jakobbotsch

Copy link
Copy Markdown
Member

This in general LGTM, but it might be interesting to understand why promotion didn't handle this given it's a simple struct containing 2x ulong fields.

CC. @jakobbotsch

I would need some concrete cases to look at. As it is this looks like a size wise regression. @xtqqczze do you have concrete benchmarks showing improvements from the change?

@stephentoub

Copy link
Copy Markdown
Member

As it is this looks like a size wise regression

Is that not just because the system chose to inline more things as a result of this? If BitCast had been available when this code was first written, presumably we'd have chosen to use it then.

@jakobbotsch

Copy link
Copy Markdown
Member

Is that not just because the system chose to inline more things as a result of this? If BitCast had been available when this code was first written, presumably we'd have chosen to use it then.

Yeah, definitely looks like there are different inlining decisions here (both new inlines we perform and cases where we no longer inline, it looks like).

Also looks like the second diff is a larger size-wise improvement than the first diff is a regression, so in that sense this isn't actually a size-wise regression (but as you said, it's hard to compare in the face of different inlining decisions).

Either way I'd be happy to look at concrete cases if there are any, but I wasn't immediately able to identify anything that looks related to deficiencies in promotion in the diffs. If you and @tannergooding prefer BitCast over the explicit field accesses then certainly that's fine with me. I don't think the JIT should have any problem handling either pattern.

@tannergooding

tannergooding commented Jul 9, 2024

Copy link
Copy Markdown
Member

Is that not just because the system chose to inline more things as a result of this?

That's what it looks like to me. Size diffs are always a little wonky for xarch due to the variable sized encoding and differing encoding cost for some sizes or registers. LSRA choosing to use R8 instead of RAX can lead to an additional byte, for example.

In this case, it looks like we eliminate code, which allows inlining to kick in and that changes register preferences and some operation sizes causes the size to increase even though the number of instructions often decreases. For example in System.Int128:TryFormat(System.Span1[ushort],byref,System.ReadOnlySpan1[ushort],System.IFormatProvider):ubyte:this,

- mov rsi, rax- or rsi, 1- lzcnt rsi, rsi- xor esi, 63- movsxd rsi, esi+ mov rdi, rax+ or rdi, 1+ lzcnt rdi, rdi+ xor edi, 63+ movsxd rdi, edi+ mov rsi, 0xD1FFAB1E ; static handle+ movzx rdi, byte ptr [rdi+rsi]+ mov esi, edi
mov rdx, 0xD1FFAB1E ; static handle
- movzx rsi, byte ptr [rsi+rdx]- mov edx, esi- mov rcx, 0xD1FFAB1E ; static handle- cmp rax, qword ptr [rcx+8*rdx]- setb dl- movzx rdx, dl- sub esi, edx- lea eax, [rsi+0x14]+ cmp rax, qword ptr [rdx+8*rsi]+ setb sil+ movzx rsi, sil+ sub edi, esi+ lea eax, [rdi+0x14]
jmp SHORT G_M10567_IG08
- ;; size=106 bbWeight=0.50 PerfScore 9.25+ ;; size=108 bbWeight=0.50 PerfScore 9.25

This code is 2 bytes bigger, but it actually hasn't fundamentally changed and if you were to replace the registers used with abstract names like reg1 you'd find they're identical. The reason the latter is bigger is because LSRA decided to use RSI for the "second register" rather than RDX, this requires an extra byte to encode the setb sil and subsequent movzx rsi, sil.

We see quite a lot of diffs that are in this realm and it'd probably be beneficial to track the number of instructions in addition to the size so we can get a better view over what's actually changed.

Later on in the method we get 96 extra bytes that are "real" due to the call to UInt128.DivRem being inlined and thus us having 3 calls to op_Division, op_Multiply, and op_Subtraction instead
-- Notably there's also a call to the System.ValueTuple constructor which seems non-ideal given its just setting two fields
-- Ideally we'd also optimize DivRem to avoid needing the multiply/subtract. The main division algorithm can just return the remainder directly and avoid needing the extra work

@tannergooding

Copy link
Copy Markdown
Member

If you and @tannergooding prefer BitCast over the explicit field accesses then certainly that's fine with me. I don't think the JIT should have any problem handling either pattern.

I think I prefer BitCast just because it's less IL and gives a small inlining profitability boost due to being intrinsic. Had just seemed a little odd that there were other improvements beyond just inlining.

@stephentoub

Copy link
Copy Markdown
Member

Thanks.

@stephentoub
stephentoub merged commit 01dbf51 into dotnet:mainJul 9, 2024
matouskozak added a commit to matouskozak/runtime that referenced this pull request Jul 11, 2024
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Aug 9, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-System.Numericscommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@xtqqczze@tannergooding@jakobbotsch@stephentoub