Contain memory operands under casts - #72719

Merged
jakobbotsch merged 6 commits into
dotnet:mainfrom
SingleAccretion:Casts-Containment
Aug 24, 2022
Merged

Contain memory operands under casts#72719
jakobbotsch merged 6 commits into
dotnet:mainfrom
SingleAccretion:Casts-Containment

Conversation

@SingleAccretion

@SingleAccretionSingleAccretion commented Jul 23, 2022

Copy link
Copy Markdown
Contributor

Recognize legal patterns in lowering and then fold the sign/zero-extension into an appropriate load at codegen time.

The change implements full support on XARCH, partial support on ARM/64, and punts LA. A follow-up issue will be filed to track the remaining work. The ARM/64 commit contains some notes on what is required for said support on our load/store architectures.

The diffs are nice and simple:

- ldr lr, [sp+0x34] // [V07 arg7]- uxtb lr, lr+ ldrb lr, [sp+0x34] // [V07 arg7]

Regressions are RA/alignment.

@ghostghost added area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI community-contribution Indicates that the PR has been added by a community member labels Jul 23, 2022
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

Fold the sign/zero-extension into an appropriate load.

The change implements full support on XARCH, partial support on ARM/64, and punts LA. A follow-up issue will be filed to track the remaining work. The ARM/64 commit contains some notes on what is required for said support on our load/store architectures.

The diffs are nice and simple:

- ldr lr, [sp+0x34] // [V07 arg7]- uxtb lr, lr+ ldrb lr, [sp+0x34] // [V07 arg7]
Author:SingleAccretion
Assignees:-
Labels:

area-CodeGen-coreclr

Milestone:-

@SingleAccretion
SingleAccretionforce-pushed the Casts-Containment branch 3 times, most recently from d3d8af4 to 6ffddaeCompareJuly 23, 2022 23:14
@SingleAccretion
SingleAccretion marked this pull request as ready for review July 24, 2022 12:19
@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib

Comment threadsrc/coreclr/jit/codegenarmarch.cpp Outdated

@TIHanTIHan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The diffs look really good.

I tried doing something similar here: #70756 - but I didn't get to work on it longer.

Comment threadsrc/coreclr/jit/codegenxarch.cpp Outdated
@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

I think it would be good to run the fuzzers and [libraries] stress on this change.

@EgorBo

Copy link
Copy Markdown
Member

/azp run Fuzzlyn, Antigen, runtime-coreclr jitstress, runtime-coreclr libraries-jitstress, runtime-coreclr gcstress0x3-gcstress0xc, runtime-coreclr jitstressregs

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 6 pipeline(s).

@jakobbotschjakobbotsch added this to the 8.0.0 milestone Jul 29, 2022
TODO: consider using a dedicated IND_EXT oper for ARM/ARM64 instead of containment.
This would allow us to cleany handle all indirections. It would not mean we'd give
up on the casts containment, as we'd still need to handle the "reg optional" case.
IND_EXT will be much like an ordinary IND, but have a "source" and "target" types.
The "target" type would always be int/long, while "source" could be of any integral
type.
This design would be a bit more natural, and nicely separable from casts. However,
the main problem with the current state of things, apart from the fact codegen of
indirections is tied strongly to "GenTreeIndir", is the fact that changing type of
the load can invalidate LEA containment. One would think this is solvable with some
tricks, like re-running containment analysis on an indirection after processing the
cast, but ARM64 codegen doesn't support uncontained LEAs in some cases.
A possible solution to that problem is uncontaining the whole address tree. That
would be messy, but ought to work. An additional complication is that these trees
can contain a lot of contained operands as part of ADDEX and BFIZ, so what would
have to be done first is the making of these into proper EXOPs.
In any case, this is all future work.
In CAST<short>(IND<byte>(...)), "m_extendSrcSize" must be "1".
Modulo the above, stress runs looked clean.

@jakobbotschjakobbotsch left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's curious to me that there are so many x86 improvements. Do we not already removed most of these unnecessary casts there by retyping loads in morph?

@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

It's curious to me that there are so many x86 improvements. Do we not already removed most of these unnecessary casts there by retyping loads in morph?

I think most of these come from normalize-on-load stack parameters, which we do not retype.

@jakobbotsch

Copy link
Copy Markdown
Member

I think most of these come from normalize-on-load stack parameters, which we do not retype.

Doesn't morph introduce this IR shape itself in fgMorphLocalVar? So doing that is a large regression?

@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

Doesn't morph introduce this IR shape itself in fgMorphLocalVar? So doing that is a large regression?

It of course does, and it is, but it is necessary for in-register parameters (and on-stack ones as well, though that should be relatively easy to fix by using extending loads in genEnregisterIncomingStackArgs).

Morph's logic there is questionable in more than one way. If we know the local will be DNER at that point, we could retype it to its small type instead of introducing a cast. However, this turns out to sometimes not be a win because we don't CSE locals, but do CSE casts.

@jakobbotsch
jakobbotsch merged commit 0469020 into dotnet:mainAug 24, 2022
@jakobbotsch

Copy link
Copy Markdown
Member

Makes sense about CSE. The way fgMorphLocalVar worked always seemed odd to me.

Anyway, thanks for the contribution as usual.

@DrewScoggins

DrewScoggins commented Aug 30, 2022

Copy link
Copy Markdown
Member

@AndyAyersMS

AndyAyersMS commented Sep 1, 2022

Copy link
Copy Markdown
Member

Possible regressions:
ubuntu arm64: dotnet/perf-autofiling-issues#8221
ubuntu x64: dotnet/perf-autofiling-issues#8255
windows x64: dotnet/perf-autofiling-issues#8258

@ghostghost locked as resolved and limited conversation to collaborators Oct 2, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@SingleAccretion@EgorBo@jakobbotsch@DrewScoggins@AndyAyersMS@TIHan@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Contain memory operands under casts - #72719

Merged
jakobbotsch merged 6 commits into
dotnet:mainfrom
SingleAccretion:Casts-Containment
Aug 24, 2022
Merged

Contain memory operands under casts#72719
jakobbotsch merged 6 commits into
dotnet:mainfrom
SingleAccretion:Casts-Containment

Conversation

@SingleAccretion

@SingleAccretionSingleAccretion commented Jul 23, 2022

Copy link
Copy Markdown
Contributor

Recognize legal patterns in lowering and then fold the sign/zero-extension into an appropriate load at codegen time.

The change implements full support on XARCH, partial support on ARM/64, and punts LA. A follow-up issue will be filed to track the remaining work. The ARM/64 commit contains some notes on what is required for said support on our load/store architectures.

The diffs are nice and simple:

- ldr lr, [sp+0x34] // [V07 arg7]- uxtb lr, lr+ ldrb lr, [sp+0x34] // [V07 arg7]

Regressions are RA/alignment.

@ghostghost added area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI community-contribution Indicates that the PR has been added by a community member labels Jul 23, 2022
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

Fold the sign/zero-extension into an appropriate load.

The change implements full support on XARCH, partial support on ARM/64, and punts LA. A follow-up issue will be filed to track the remaining work. The ARM/64 commit contains some notes on what is required for said support on our load/store architectures.

The diffs are nice and simple:

- ldr lr, [sp+0x34] // [V07 arg7]- uxtb lr, lr+ ldrb lr, [sp+0x34] // [V07 arg7]
Author:SingleAccretion
Assignees:-
Labels:

area-CodeGen-coreclr

Milestone:-

@SingleAccretion
SingleAccretionforce-pushed the Casts-Containment branch 3 times, most recently from d3d8af4 to 6ffddaeCompareJuly 23, 2022 23:14
@SingleAccretion
SingleAccretion marked this pull request as ready for review July 24, 2022 12:19
@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib

Comment threadsrc/coreclr/jit/codegenarmarch.cpp Outdated

@TIHanTIHan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The diffs look really good.

I tried doing something similar here: #70756 - but I didn't get to work on it longer.

Comment threadsrc/coreclr/jit/codegenxarch.cpp Outdated
@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

I think it would be good to run the fuzzers and [libraries] stress on this change.

@EgorBo

Copy link
Copy Markdown
Member

/azp run Fuzzlyn, Antigen, runtime-coreclr jitstress, runtime-coreclr libraries-jitstress, runtime-coreclr gcstress0x3-gcstress0xc, runtime-coreclr jitstressregs

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 6 pipeline(s).

@jakobbotschjakobbotsch added this to the 8.0.0 milestone Jul 29, 2022
TODO: consider using a dedicated IND_EXT oper for ARM/ARM64 instead of containment.
This would allow us to cleany handle all indirections. It would not mean we'd give
up on the casts containment, as we'd still need to handle the "reg optional" case.
IND_EXT will be much like an ordinary IND, but have a "source" and "target" types.
The "target" type would always be int/long, while "source" could be of any integral
type.
This design would be a bit more natural, and nicely separable from casts. However,
the main problem with the current state of things, apart from the fact codegen of
indirections is tied strongly to "GenTreeIndir", is the fact that changing type of
the load can invalidate LEA containment. One would think this is solvable with some
tricks, like re-running containment analysis on an indirection after processing the
cast, but ARM64 codegen doesn't support uncontained LEAs in some cases.
A possible solution to that problem is uncontaining the whole address tree. That
would be messy, but ought to work. An additional complication is that these trees
can contain a lot of contained operands as part of ADDEX and BFIZ, so what would
have to be done first is the making of these into proper EXOPs.
In any case, this is all future work.
In CAST<short>(IND<byte>(...)), "m_extendSrcSize" must be "1".
Modulo the above, stress runs looked clean.

@jakobbotschjakobbotsch left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's curious to me that there are so many x86 improvements. Do we not already removed most of these unnecessary casts there by retyping loads in morph?

@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

It's curious to me that there are so many x86 improvements. Do we not already removed most of these unnecessary casts there by retyping loads in morph?

I think most of these come from normalize-on-load stack parameters, which we do not retype.

@jakobbotsch

Copy link
Copy Markdown
Member

I think most of these come from normalize-on-load stack parameters, which we do not retype.

Doesn't morph introduce this IR shape itself in fgMorphLocalVar? So doing that is a large regression?

@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

Doesn't morph introduce this IR shape itself in fgMorphLocalVar? So doing that is a large regression?

It of course does, and it is, but it is necessary for in-register parameters (and on-stack ones as well, though that should be relatively easy to fix by using extending loads in genEnregisterIncomingStackArgs).

Morph's logic there is questionable in more than one way. If we know the local will be DNER at that point, we could retype it to its small type instead of introducing a cast. However, this turns out to sometimes not be a win because we don't CSE locals, but do CSE casts.

@jakobbotsch
jakobbotsch merged commit 0469020 into dotnet:mainAug 24, 2022
@jakobbotsch

Copy link
Copy Markdown
Member

Makes sense about CSE. The way fgMorphLocalVar worked always seemed odd to me.

Anyway, thanks for the contribution as usual.

@DrewScoggins

DrewScoggins commented Aug 30, 2022

Copy link
Copy Markdown
Member

@AndyAyersMS

AndyAyersMS commented Sep 1, 2022

Copy link
Copy Markdown
Member

Possible regressions:
ubuntu arm64: dotnet/perf-autofiling-issues#8221
ubuntu x64: dotnet/perf-autofiling-issues#8255
windows x64: dotnet/perf-autofiling-issues#8258

@ghostghost locked as resolved and limited conversation to collaborators Oct 2, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@SingleAccretion@EgorBo@jakobbotsch@DrewScoggins@AndyAyersMS@TIHan@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Contain memory operands under casts - #72719

Merged
jakobbotsch merged 6 commits into
dotnet:mainfrom
SingleAccretion:Casts-Containment
Aug 24, 2022
Merged

Contain memory operands under casts#72719
jakobbotsch merged 6 commits into
dotnet:mainfrom
SingleAccretion:Casts-Containment

Conversation

@SingleAccretion

@SingleAccretionSingleAccretion commented Jul 23, 2022

Copy link
Copy Markdown
Contributor

Recognize legal patterns in lowering and then fold the sign/zero-extension into an appropriate load at codegen time.

The change implements full support on XARCH, partial support on ARM/64, and punts LA. A follow-up issue will be filed to track the remaining work. The ARM/64 commit contains some notes on what is required for said support on our load/store architectures.

The diffs are nice and simple:

- ldr lr, [sp+0x34] // [V07 arg7]- uxtb lr, lr+ ldrb lr, [sp+0x34] // [V07 arg7]

Regressions are RA/alignment.

@ghostghost added area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI community-contribution Indicates that the PR has been added by a community member labels Jul 23, 2022
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

Fold the sign/zero-extension into an appropriate load.

The change implements full support on XARCH, partial support on ARM/64, and punts LA. A follow-up issue will be filed to track the remaining work. The ARM/64 commit contains some notes on what is required for said support on our load/store architectures.

The diffs are nice and simple:

- ldr lr, [sp+0x34] // [V07 arg7]- uxtb lr, lr+ ldrb lr, [sp+0x34] // [V07 arg7]
Author:SingleAccretion
Assignees:-
Labels:

area-CodeGen-coreclr

Milestone:-

@SingleAccretion
SingleAccretionforce-pushed the Casts-Containment branch 3 times, most recently from d3d8af4 to 6ffddaeCompareJuly 23, 2022 23:14
@SingleAccretion
SingleAccretion marked this pull request as ready for review July 24, 2022 12:19
@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib

Comment threadsrc/coreclr/jit/codegenarmarch.cpp Outdated

@TIHanTIHan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The diffs look really good.

I tried doing something similar here: #70756 - but I didn't get to work on it longer.

Comment threadsrc/coreclr/jit/codegenxarch.cpp Outdated
@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

I think it would be good to run the fuzzers and [libraries] stress on this change.

@EgorBo

Copy link
Copy Markdown
Member

/azp run Fuzzlyn, Antigen, runtime-coreclr jitstress, runtime-coreclr libraries-jitstress, runtime-coreclr gcstress0x3-gcstress0xc, runtime-coreclr jitstressregs

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 6 pipeline(s).

@jakobbotschjakobbotsch added this to the 8.0.0 milestone Jul 29, 2022
TODO: consider using a dedicated IND_EXT oper for ARM/ARM64 instead of containment.
This would allow us to cleany handle all indirections. It would not mean we'd give
up on the casts containment, as we'd still need to handle the "reg optional" case.
IND_EXT will be much like an ordinary IND, but have a "source" and "target" types.
The "target" type would always be int/long, while "source" could be of any integral
type.
This design would be a bit more natural, and nicely separable from casts. However,
the main problem with the current state of things, apart from the fact codegen of
indirections is tied strongly to "GenTreeIndir", is the fact that changing type of
the load can invalidate LEA containment. One would think this is solvable with some
tricks, like re-running containment analysis on an indirection after processing the
cast, but ARM64 codegen doesn't support uncontained LEAs in some cases.
A possible solution to that problem is uncontaining the whole address tree. That
would be messy, but ought to work. An additional complication is that these trees
can contain a lot of contained operands as part of ADDEX and BFIZ, so what would
have to be done first is the making of these into proper EXOPs.
In any case, this is all future work.
In CAST<short>(IND<byte>(...)), "m_extendSrcSize" must be "1".
Modulo the above, stress runs looked clean.

@jakobbotschjakobbotsch left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's curious to me that there are so many x86 improvements. Do we not already removed most of these unnecessary casts there by retyping loads in morph?

@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

It's curious to me that there are so many x86 improvements. Do we not already removed most of these unnecessary casts there by retyping loads in morph?

I think most of these come from normalize-on-load stack parameters, which we do not retype.

@jakobbotsch

Copy link
Copy Markdown
Member

I think most of these come from normalize-on-load stack parameters, which we do not retype.

Doesn't morph introduce this IR shape itself in fgMorphLocalVar? So doing that is a large regression?

@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

Doesn't morph introduce this IR shape itself in fgMorphLocalVar? So doing that is a large regression?

It of course does, and it is, but it is necessary for in-register parameters (and on-stack ones as well, though that should be relatively easy to fix by using extending loads in genEnregisterIncomingStackArgs).

Morph's logic there is questionable in more than one way. If we know the local will be DNER at that point, we could retype it to its small type instead of introducing a cast. However, this turns out to sometimes not be a win because we don't CSE locals, but do CSE casts.

@jakobbotsch
jakobbotsch merged commit 0469020 into dotnet:mainAug 24, 2022
@jakobbotsch

Copy link
Copy Markdown
Member

Makes sense about CSE. The way fgMorphLocalVar worked always seemed odd to me.

Anyway, thanks for the contribution as usual.

@DrewScoggins

DrewScoggins commented Aug 30, 2022

Copy link
Copy Markdown
Member

@AndyAyersMS

AndyAyersMS commented Sep 1, 2022

Copy link
Copy Markdown
Member

Possible regressions:
ubuntu arm64: dotnet/perf-autofiling-issues#8221
ubuntu x64: dotnet/perf-autofiling-issues#8255
windows x64: dotnet/perf-autofiling-issues#8258

@ghostghost locked as resolved and limited conversation to collaborators Oct 2, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@SingleAccretion@EgorBo@jakobbotsch@DrewScoggins@AndyAyersMS@TIHan@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Contain memory operands under casts - #72719

Merged
jakobbotsch merged 6 commits into
dotnet:mainfrom
SingleAccretion:Casts-Containment
Aug 24, 2022
Merged

Contain memory operands under casts#72719
jakobbotsch merged 6 commits into
dotnet:mainfrom
SingleAccretion:Casts-Containment

Conversation

@SingleAccretion

@SingleAccretionSingleAccretion commented Jul 23, 2022

Copy link
Copy Markdown
Contributor

Recognize legal patterns in lowering and then fold the sign/zero-extension into an appropriate load at codegen time.

The change implements full support on XARCH, partial support on ARM/64, and punts LA. A follow-up issue will be filed to track the remaining work. The ARM/64 commit contains some notes on what is required for said support on our load/store architectures.

The diffs are nice and simple:

- ldr lr, [sp+0x34] // [V07 arg7]- uxtb lr, lr+ ldrb lr, [sp+0x34] // [V07 arg7]

Regressions are RA/alignment.

@ghostghost added area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI community-contribution Indicates that the PR has been added by a community member labels Jul 23, 2022
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

Fold the sign/zero-extension into an appropriate load.

The change implements full support on XARCH, partial support on ARM/64, and punts LA. A follow-up issue will be filed to track the remaining work. The ARM/64 commit contains some notes on what is required for said support on our load/store architectures.

The diffs are nice and simple:

- ldr lr, [sp+0x34] // [V07 arg7]- uxtb lr, lr+ ldrb lr, [sp+0x34] // [V07 arg7]
Author:SingleAccretion
Assignees:-
Labels:

area-CodeGen-coreclr

Milestone:-

@SingleAccretion
SingleAccretionforce-pushed the Casts-Containment branch 3 times, most recently from d3d8af4 to 6ffddaeCompareJuly 23, 2022 23:14
@SingleAccretion
SingleAccretion marked this pull request as ready for review July 24, 2022 12:19
@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib

Comment threadsrc/coreclr/jit/codegenarmarch.cpp Outdated

@TIHanTIHan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The diffs look really good.

I tried doing something similar here: #70756 - but I didn't get to work on it longer.

Comment threadsrc/coreclr/jit/codegenxarch.cpp Outdated
@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

I think it would be good to run the fuzzers and [libraries] stress on this change.

@EgorBo

Copy link
Copy Markdown
Member

/azp run Fuzzlyn, Antigen, runtime-coreclr jitstress, runtime-coreclr libraries-jitstress, runtime-coreclr gcstress0x3-gcstress0xc, runtime-coreclr jitstressregs

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 6 pipeline(s).

@jakobbotschjakobbotsch added this to the 8.0.0 milestone Jul 29, 2022
TODO: consider using a dedicated IND_EXT oper for ARM/ARM64 instead of containment.
This would allow us to cleany handle all indirections. It would not mean we'd give
up on the casts containment, as we'd still need to handle the "reg optional" case.
IND_EXT will be much like an ordinary IND, but have a "source" and "target" types.
The "target" type would always be int/long, while "source" could be of any integral
type.
This design would be a bit more natural, and nicely separable from casts. However,
the main problem with the current state of things, apart from the fact codegen of
indirections is tied strongly to "GenTreeIndir", is the fact that changing type of
the load can invalidate LEA containment. One would think this is solvable with some
tricks, like re-running containment analysis on an indirection after processing the
cast, but ARM64 codegen doesn't support uncontained LEAs in some cases.
A possible solution to that problem is uncontaining the whole address tree. That
would be messy, but ought to work. An additional complication is that these trees
can contain a lot of contained operands as part of ADDEX and BFIZ, so what would
have to be done first is the making of these into proper EXOPs.
In any case, this is all future work.
In CAST<short>(IND<byte>(...)), "m_extendSrcSize" must be "1".
Modulo the above, stress runs looked clean.

@jakobbotschjakobbotsch left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's curious to me that there are so many x86 improvements. Do we not already removed most of these unnecessary casts there by retyping loads in morph?

@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

It's curious to me that there are so many x86 improvements. Do we not already removed most of these unnecessary casts there by retyping loads in morph?

I think most of these come from normalize-on-load stack parameters, which we do not retype.

@jakobbotsch

Copy link
Copy Markdown
Member

I think most of these come from normalize-on-load stack parameters, which we do not retype.

Doesn't morph introduce this IR shape itself in fgMorphLocalVar? So doing that is a large regression?

@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

Doesn't morph introduce this IR shape itself in fgMorphLocalVar? So doing that is a large regression?

It of course does, and it is, but it is necessary for in-register parameters (and on-stack ones as well, though that should be relatively easy to fix by using extending loads in genEnregisterIncomingStackArgs).

Morph's logic there is questionable in more than one way. If we know the local will be DNER at that point, we could retype it to its small type instead of introducing a cast. However, this turns out to sometimes not be a win because we don't CSE locals, but do CSE casts.

@jakobbotsch
jakobbotsch merged commit 0469020 into dotnet:mainAug 24, 2022
@jakobbotsch

Copy link
Copy Markdown
Member

Makes sense about CSE. The way fgMorphLocalVar worked always seemed odd to me.

Anyway, thanks for the contribution as usual.

@DrewScoggins

DrewScoggins commented Aug 30, 2022

Copy link
Copy Markdown
Member

@AndyAyersMS

AndyAyersMS commented Sep 1, 2022

Copy link
Copy Markdown
Member

Possible regressions:
ubuntu arm64: dotnet/perf-autofiling-issues#8221
ubuntu x64: dotnet/perf-autofiling-issues#8255
windows x64: dotnet/perf-autofiling-issues#8258

@ghostghost locked as resolved and limited conversation to collaborators Oct 2, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@SingleAccretion@EgorBo@jakobbotsch@DrewScoggins@AndyAyersMS@TIHan@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Contain memory operands under casts - #72719

Merged
jakobbotsch merged 6 commits into
dotnet:mainfrom
SingleAccretion:Casts-Containment
Aug 24, 2022
Merged

Contain memory operands under casts#72719
jakobbotsch merged 6 commits into
dotnet:mainfrom
SingleAccretion:Casts-Containment

Conversation

@SingleAccretion

@SingleAccretionSingleAccretion commented Jul 23, 2022

Copy link
Copy Markdown
Contributor

Recognize legal patterns in lowering and then fold the sign/zero-extension into an appropriate load at codegen time.

The change implements full support on XARCH, partial support on ARM/64, and punts LA. A follow-up issue will be filed to track the remaining work. The ARM/64 commit contains some notes on what is required for said support on our load/store architectures.

The diffs are nice and simple:

- ldr lr, [sp+0x34] // [V07 arg7]- uxtb lr, lr+ ldrb lr, [sp+0x34] // [V07 arg7]

Regressions are RA/alignment.

@ghostghost added area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI community-contribution Indicates that the PR has been added by a community member labels Jul 23, 2022
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

Fold the sign/zero-extension into an appropriate load.

The change implements full support on XARCH, partial support on ARM/64, and punts LA. A follow-up issue will be filed to track the remaining work. The ARM/64 commit contains some notes on what is required for said support on our load/store architectures.

The diffs are nice and simple:

- ldr lr, [sp+0x34] // [V07 arg7]- uxtb lr, lr+ ldrb lr, [sp+0x34] // [V07 arg7]
Author:SingleAccretion
Assignees:-
Labels:

area-CodeGen-coreclr

Milestone:-

@SingleAccretion
SingleAccretionforce-pushed the Casts-Containment branch 3 times, most recently from d3d8af4 to 6ffddaeCompareJuly 23, 2022 23:14
@SingleAccretion
SingleAccretion marked this pull request as ready for review July 24, 2022 12:19
@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib

Comment threadsrc/coreclr/jit/codegenarmarch.cpp Outdated

@TIHanTIHan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The diffs look really good.

I tried doing something similar here: #70756 - but I didn't get to work on it longer.

Comment threadsrc/coreclr/jit/codegenxarch.cpp Outdated
@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

I think it would be good to run the fuzzers and [libraries] stress on this change.

@EgorBo

Copy link
Copy Markdown
Member

/azp run Fuzzlyn, Antigen, runtime-coreclr jitstress, runtime-coreclr libraries-jitstress, runtime-coreclr gcstress0x3-gcstress0xc, runtime-coreclr jitstressregs

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 6 pipeline(s).

@jakobbotschjakobbotsch added this to the 8.0.0 milestone Jul 29, 2022
TODO: consider using a dedicated IND_EXT oper for ARM/ARM64 instead of containment.
This would allow us to cleany handle all indirections. It would not mean we'd give
up on the casts containment, as we'd still need to handle the "reg optional" case.
IND_EXT will be much like an ordinary IND, but have a "source" and "target" types.
The "target" type would always be int/long, while "source" could be of any integral
type.
This design would be a bit more natural, and nicely separable from casts. However,
the main problem with the current state of things, apart from the fact codegen of
indirections is tied strongly to "GenTreeIndir", is the fact that changing type of
the load can invalidate LEA containment. One would think this is solvable with some
tricks, like re-running containment analysis on an indirection after processing the
cast, but ARM64 codegen doesn't support uncontained LEAs in some cases.
A possible solution to that problem is uncontaining the whole address tree. That
would be messy, but ought to work. An additional complication is that these trees
can contain a lot of contained operands as part of ADDEX and BFIZ, so what would
have to be done first is the making of these into proper EXOPs.
In any case, this is all future work.
In CAST<short>(IND<byte>(...)), "m_extendSrcSize" must be "1".
Modulo the above, stress runs looked clean.

@jakobbotschjakobbotsch left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's curious to me that there are so many x86 improvements. Do we not already removed most of these unnecessary casts there by retyping loads in morph?

@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

It's curious to me that there are so many x86 improvements. Do we not already removed most of these unnecessary casts there by retyping loads in morph?

I think most of these come from normalize-on-load stack parameters, which we do not retype.

@jakobbotsch

Copy link
Copy Markdown
Member

I think most of these come from normalize-on-load stack parameters, which we do not retype.

Doesn't morph introduce this IR shape itself in fgMorphLocalVar? So doing that is a large regression?

@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

Doesn't morph introduce this IR shape itself in fgMorphLocalVar? So doing that is a large regression?

It of course does, and it is, but it is necessary for in-register parameters (and on-stack ones as well, though that should be relatively easy to fix by using extending loads in genEnregisterIncomingStackArgs).

Morph's logic there is questionable in more than one way. If we know the local will be DNER at that point, we could retype it to its small type instead of introducing a cast. However, this turns out to sometimes not be a win because we don't CSE locals, but do CSE casts.

@jakobbotsch
jakobbotsch merged commit 0469020 into dotnet:mainAug 24, 2022
@jakobbotsch

Copy link
Copy Markdown
Member

Makes sense about CSE. The way fgMorphLocalVar worked always seemed odd to me.

Anyway, thanks for the contribution as usual.

@DrewScoggins

DrewScoggins commented Aug 30, 2022

Copy link
Copy Markdown
Member

@AndyAyersMS

AndyAyersMS commented Sep 1, 2022

Copy link
Copy Markdown
Member

Possible regressions:
ubuntu arm64: dotnet/perf-autofiling-issues#8221
ubuntu x64: dotnet/perf-autofiling-issues#8255
windows x64: dotnet/perf-autofiling-issues#8258

@ghostghost locked as resolved and limited conversation to collaborators Oct 2, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@SingleAccretion@EgorBo@jakobbotsch@DrewScoggins@AndyAyersMS@TIHan@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Contain memory operands under casts - #72719

Merged
jakobbotsch merged 6 commits into
dotnet:mainfrom
SingleAccretion:Casts-Containment
Aug 24, 2022
Merged

Contain memory operands under casts#72719
jakobbotsch merged 6 commits into
dotnet:mainfrom
SingleAccretion:Casts-Containment

Conversation

@SingleAccretion

@SingleAccretionSingleAccretion commented Jul 23, 2022

Copy link
Copy Markdown
Contributor

Recognize legal patterns in lowering and then fold the sign/zero-extension into an appropriate load at codegen time.

The change implements full support on XARCH, partial support on ARM/64, and punts LA. A follow-up issue will be filed to track the remaining work. The ARM/64 commit contains some notes on what is required for said support on our load/store architectures.

The diffs are nice and simple:

- ldr lr, [sp+0x34] // [V07 arg7]- uxtb lr, lr+ ldrb lr, [sp+0x34] // [V07 arg7]

Regressions are RA/alignment.

@ghostghost added area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI community-contribution Indicates that the PR has been added by a community member labels Jul 23, 2022
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

Fold the sign/zero-extension into an appropriate load.

The change implements full support on XARCH, partial support on ARM/64, and punts LA. A follow-up issue will be filed to track the remaining work. The ARM/64 commit contains some notes on what is required for said support on our load/store architectures.

The diffs are nice and simple:

- ldr lr, [sp+0x34] // [V07 arg7]- uxtb lr, lr+ ldrb lr, [sp+0x34] // [V07 arg7]
Author:SingleAccretion
Assignees:-
Labels:

area-CodeGen-coreclr

Milestone:-

@SingleAccretion
SingleAccretionforce-pushed the Casts-Containment branch 3 times, most recently from d3d8af4 to 6ffddaeCompareJuly 23, 2022 23:14
@SingleAccretion
SingleAccretion marked this pull request as ready for review July 24, 2022 12:19
@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib

Comment threadsrc/coreclr/jit/codegenarmarch.cpp Outdated

@TIHanTIHan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The diffs look really good.

I tried doing something similar here: #70756 - but I didn't get to work on it longer.

Comment threadsrc/coreclr/jit/codegenxarch.cpp Outdated
@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

I think it would be good to run the fuzzers and [libraries] stress on this change.

@EgorBo

Copy link
Copy Markdown
Member

/azp run Fuzzlyn, Antigen, runtime-coreclr jitstress, runtime-coreclr libraries-jitstress, runtime-coreclr gcstress0x3-gcstress0xc, runtime-coreclr jitstressregs

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 6 pipeline(s).

@jakobbotschjakobbotsch added this to the 8.0.0 milestone Jul 29, 2022
TODO: consider using a dedicated IND_EXT oper for ARM/ARM64 instead of containment.
This would allow us to cleany handle all indirections. It would not mean we'd give
up on the casts containment, as we'd still need to handle the "reg optional" case.
IND_EXT will be much like an ordinary IND, but have a "source" and "target" types.
The "target" type would always be int/long, while "source" could be of any integral
type.
This design would be a bit more natural, and nicely separable from casts. However,
the main problem with the current state of things, apart from the fact codegen of
indirections is tied strongly to "GenTreeIndir", is the fact that changing type of
the load can invalidate LEA containment. One would think this is solvable with some
tricks, like re-running containment analysis on an indirection after processing the
cast, but ARM64 codegen doesn't support uncontained LEAs in some cases.
A possible solution to that problem is uncontaining the whole address tree. That
would be messy, but ought to work. An additional complication is that these trees
can contain a lot of contained operands as part of ADDEX and BFIZ, so what would
have to be done first is the making of these into proper EXOPs.
In any case, this is all future work.
In CAST<short>(IND<byte>(...)), "m_extendSrcSize" must be "1".
Modulo the above, stress runs looked clean.

@jakobbotschjakobbotsch left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's curious to me that there are so many x86 improvements. Do we not already removed most of these unnecessary casts there by retyping loads in morph?

@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

It's curious to me that there are so many x86 improvements. Do we not already removed most of these unnecessary casts there by retyping loads in morph?

I think most of these come from normalize-on-load stack parameters, which we do not retype.

@jakobbotsch

Copy link
Copy Markdown
Member

I think most of these come from normalize-on-load stack parameters, which we do not retype.

Doesn't morph introduce this IR shape itself in fgMorphLocalVar? So doing that is a large regression?

@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

Doesn't morph introduce this IR shape itself in fgMorphLocalVar? So doing that is a large regression?

It of course does, and it is, but it is necessary for in-register parameters (and on-stack ones as well, though that should be relatively easy to fix by using extending loads in genEnregisterIncomingStackArgs).

Morph's logic there is questionable in more than one way. If we know the local will be DNER at that point, we could retype it to its small type instead of introducing a cast. However, this turns out to sometimes not be a win because we don't CSE locals, but do CSE casts.

@jakobbotsch
jakobbotsch merged commit 0469020 into dotnet:mainAug 24, 2022
@jakobbotsch

Copy link
Copy Markdown
Member

Makes sense about CSE. The way fgMorphLocalVar worked always seemed odd to me.

Anyway, thanks for the contribution as usual.

@DrewScoggins

DrewScoggins commented Aug 30, 2022

Copy link
Copy Markdown
Member

@AndyAyersMS

AndyAyersMS commented Sep 1, 2022

Copy link
Copy Markdown
Member

Possible regressions:
ubuntu arm64: dotnet/perf-autofiling-issues#8221
ubuntu x64: dotnet/perf-autofiling-issues#8255
windows x64: dotnet/perf-autofiling-issues#8258

@ghostghost locked as resolved and limited conversation to collaborators Oct 2, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@SingleAccretion@EgorBo@jakobbotsch@DrewScoggins@AndyAyersMS@TIHan@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Contain memory operands under casts - #72719

Merged
jakobbotsch merged 6 commits into
dotnet:mainfrom
SingleAccretion:Casts-Containment
Aug 24, 2022
Merged

Contain memory operands under casts#72719
jakobbotsch merged 6 commits into
dotnet:mainfrom
SingleAccretion:Casts-Containment

Conversation

@SingleAccretion

@SingleAccretionSingleAccretion commented Jul 23, 2022

Copy link
Copy Markdown
Contributor

Recognize legal patterns in lowering and then fold the sign/zero-extension into an appropriate load at codegen time.

The change implements full support on XARCH, partial support on ARM/64, and punts LA. A follow-up issue will be filed to track the remaining work. The ARM/64 commit contains some notes on what is required for said support on our load/store architectures.

The diffs are nice and simple:

- ldr lr, [sp+0x34] // [V07 arg7]- uxtb lr, lr+ ldrb lr, [sp+0x34] // [V07 arg7]

Regressions are RA/alignment.

@ghostghost added area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI community-contribution Indicates that the PR has been added by a community member labels Jul 23, 2022
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

Fold the sign/zero-extension into an appropriate load.

The change implements full support on XARCH, partial support on ARM/64, and punts LA. A follow-up issue will be filed to track the remaining work. The ARM/64 commit contains some notes on what is required for said support on our load/store architectures.

The diffs are nice and simple:

- ldr lr, [sp+0x34] // [V07 arg7]- uxtb lr, lr+ ldrb lr, [sp+0x34] // [V07 arg7]
Author:SingleAccretion
Assignees:-
Labels:

area-CodeGen-coreclr

Milestone:-

@SingleAccretion
SingleAccretionforce-pushed the Casts-Containment branch 3 times, most recently from d3d8af4 to 6ffddaeCompareJuly 23, 2022 23:14
@SingleAccretion
SingleAccretion marked this pull request as ready for review July 24, 2022 12:19
@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib

Comment threadsrc/coreclr/jit/codegenarmarch.cpp Outdated

@TIHanTIHan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The diffs look really good.

I tried doing something similar here: #70756 - but I didn't get to work on it longer.

Comment threadsrc/coreclr/jit/codegenxarch.cpp Outdated
@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

I think it would be good to run the fuzzers and [libraries] stress on this change.

@EgorBo

Copy link
Copy Markdown
Member

/azp run Fuzzlyn, Antigen, runtime-coreclr jitstress, runtime-coreclr libraries-jitstress, runtime-coreclr gcstress0x3-gcstress0xc, runtime-coreclr jitstressregs

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 6 pipeline(s).

@jakobbotschjakobbotsch added this to the 8.0.0 milestone Jul 29, 2022
TODO: consider using a dedicated IND_EXT oper for ARM/ARM64 instead of containment.
This would allow us to cleany handle all indirections. It would not mean we'd give
up on the casts containment, as we'd still need to handle the "reg optional" case.
IND_EXT will be much like an ordinary IND, but have a "source" and "target" types.
The "target" type would always be int/long, while "source" could be of any integral
type.
This design would be a bit more natural, and nicely separable from casts. However,
the main problem with the current state of things, apart from the fact codegen of
indirections is tied strongly to "GenTreeIndir", is the fact that changing type of
the load can invalidate LEA containment. One would think this is solvable with some
tricks, like re-running containment analysis on an indirection after processing the
cast, but ARM64 codegen doesn't support uncontained LEAs in some cases.
A possible solution to that problem is uncontaining the whole address tree. That
would be messy, but ought to work. An additional complication is that these trees
can contain a lot of contained operands as part of ADDEX and BFIZ, so what would
have to be done first is the making of these into proper EXOPs.
In any case, this is all future work.
In CAST<short>(IND<byte>(...)), "m_extendSrcSize" must be "1".
Modulo the above, stress runs looked clean.

@jakobbotschjakobbotsch left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's curious to me that there are so many x86 improvements. Do we not already removed most of these unnecessary casts there by retyping loads in morph?

@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

It's curious to me that there are so many x86 improvements. Do we not already removed most of these unnecessary casts there by retyping loads in morph?

I think most of these come from normalize-on-load stack parameters, which we do not retype.

@jakobbotsch

Copy link
Copy Markdown
Member

I think most of these come from normalize-on-load stack parameters, which we do not retype.

Doesn't morph introduce this IR shape itself in fgMorphLocalVar? So doing that is a large regression?

@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

Doesn't morph introduce this IR shape itself in fgMorphLocalVar? So doing that is a large regression?

It of course does, and it is, but it is necessary for in-register parameters (and on-stack ones as well, though that should be relatively easy to fix by using extending loads in genEnregisterIncomingStackArgs).

Morph's logic there is questionable in more than one way. If we know the local will be DNER at that point, we could retype it to its small type instead of introducing a cast. However, this turns out to sometimes not be a win because we don't CSE locals, but do CSE casts.

@jakobbotsch
jakobbotsch merged commit 0469020 into dotnet:mainAug 24, 2022
@jakobbotsch

Copy link
Copy Markdown
Member

Makes sense about CSE. The way fgMorphLocalVar worked always seemed odd to me.

Anyway, thanks for the contribution as usual.

@DrewScoggins

DrewScoggins commented Aug 30, 2022

Copy link
Copy Markdown
Member

@AndyAyersMS

AndyAyersMS commented Sep 1, 2022

Copy link
Copy Markdown
Member

Possible regressions:
ubuntu arm64: dotnet/perf-autofiling-issues#8221
ubuntu x64: dotnet/perf-autofiling-issues#8255
windows x64: dotnet/perf-autofiling-issues#8258

@ghostghost locked as resolved and limited conversation to collaborators Oct 2, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@SingleAccretion@EgorBo@jakobbotsch@DrewScoggins@AndyAyersMS@TIHan@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Contain memory operands under casts - #72719

Merged
jakobbotsch merged 6 commits into
dotnet:mainfrom
SingleAccretion:Casts-Containment
Aug 24, 2022
Merged

Contain memory operands under casts#72719
jakobbotsch merged 6 commits into
dotnet:mainfrom
SingleAccretion:Casts-Containment

Conversation

@SingleAccretion

@SingleAccretionSingleAccretion commented Jul 23, 2022

Copy link
Copy Markdown
Contributor

Recognize legal patterns in lowering and then fold the sign/zero-extension into an appropriate load at codegen time.

The change implements full support on XARCH, partial support on ARM/64, and punts LA. A follow-up issue will be filed to track the remaining work. The ARM/64 commit contains some notes on what is required for said support on our load/store architectures.

The diffs are nice and simple:

- ldr lr, [sp+0x34] // [V07 arg7]- uxtb lr, lr+ ldrb lr, [sp+0x34] // [V07 arg7]

Regressions are RA/alignment.

@ghostghost added area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI community-contribution Indicates that the PR has been added by a community member labels Jul 23, 2022
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Issue Details

Fold the sign/zero-extension into an appropriate load.

The change implements full support on XARCH, partial support on ARM/64, and punts LA. A follow-up issue will be filed to track the remaining work. The ARM/64 commit contains some notes on what is required for said support on our load/store architectures.

The diffs are nice and simple:

- ldr lr, [sp+0x34] // [V07 arg7]- uxtb lr, lr+ ldrb lr, [sp+0x34] // [V07 arg7]
Author:SingleAccretion
Assignees:-
Labels:

area-CodeGen-coreclr

Milestone:-

@SingleAccretion
SingleAccretionforce-pushed the Casts-Containment branch 3 times, most recently from d3d8af4 to 6ffddaeCompareJuly 23, 2022 23:14
@SingleAccretion
SingleAccretion marked this pull request as ready for review July 24, 2022 12:19
@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

@dotnet/jit-contrib

Comment threadsrc/coreclr/jit/codegenarmarch.cpp Outdated

@TIHanTIHan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The diffs look really good.

I tried doing something similar here: #70756 - but I didn't get to work on it longer.

Comment threadsrc/coreclr/jit/codegenxarch.cpp Outdated
@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

I think it would be good to run the fuzzers and [libraries] stress on this change.

@EgorBo

Copy link
Copy Markdown
Member

/azp run Fuzzlyn, Antigen, runtime-coreclr jitstress, runtime-coreclr libraries-jitstress, runtime-coreclr gcstress0x3-gcstress0xc, runtime-coreclr jitstressregs

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 6 pipeline(s).

@jakobbotschjakobbotsch added this to the 8.0.0 milestone Jul 29, 2022
TODO: consider using a dedicated IND_EXT oper for ARM/ARM64 instead of containment.
This would allow us to cleany handle all indirections. It would not mean we'd give
up on the casts containment, as we'd still need to handle the "reg optional" case.
IND_EXT will be much like an ordinary IND, but have a "source" and "target" types.
The "target" type would always be int/long, while "source" could be of any integral
type.
This design would be a bit more natural, and nicely separable from casts. However,
the main problem with the current state of things, apart from the fact codegen of
indirections is tied strongly to "GenTreeIndir", is the fact that changing type of
the load can invalidate LEA containment. One would think this is solvable with some
tricks, like re-running containment analysis on an indirection after processing the
cast, but ARM64 codegen doesn't support uncontained LEAs in some cases.
A possible solution to that problem is uncontaining the whole address tree. That
would be messy, but ought to work. An additional complication is that these trees
can contain a lot of contained operands as part of ADDEX and BFIZ, so what would
have to be done first is the making of these into proper EXOPs.
In any case, this is all future work.
In CAST<short>(IND<byte>(...)), "m_extendSrcSize" must be "1".
Modulo the above, stress runs looked clean.

@jakobbotschjakobbotsch left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's curious to me that there are so many x86 improvements. Do we not already removed most of these unnecessary casts there by retyping loads in morph?

@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

It's curious to me that there are so many x86 improvements. Do we not already removed most of these unnecessary casts there by retyping loads in morph?

I think most of these come from normalize-on-load stack parameters, which we do not retype.

@jakobbotsch

Copy link
Copy Markdown
Member

I think most of these come from normalize-on-load stack parameters, which we do not retype.

Doesn't morph introduce this IR shape itself in fgMorphLocalVar? So doing that is a large regression?

@SingleAccretion

Copy link
Copy Markdown
ContributorAuthor

Doesn't morph introduce this IR shape itself in fgMorphLocalVar? So doing that is a large regression?

It of course does, and it is, but it is necessary for in-register parameters (and on-stack ones as well, though that should be relatively easy to fix by using extending loads in genEnregisterIncomingStackArgs).

Morph's logic there is questionable in more than one way. If we know the local will be DNER at that point, we could retype it to its small type instead of introducing a cast. However, this turns out to sometimes not be a win because we don't CSE locals, but do CSE casts.

@jakobbotsch
jakobbotsch merged commit 0469020 into dotnet:mainAug 24, 2022
@jakobbotsch

Copy link
Copy Markdown
Member

Makes sense about CSE. The way fgMorphLocalVar worked always seemed odd to me.

Anyway, thanks for the contribution as usual.

@DrewScoggins

DrewScoggins commented Aug 30, 2022

Copy link
Copy Markdown
Member

@AndyAyersMS

AndyAyersMS commented Sep 1, 2022

Copy link
Copy Markdown
Member

Possible regressions:
ubuntu arm64: dotnet/perf-autofiling-issues#8221
ubuntu x64: dotnet/perf-autofiling-issues#8255
windows x64: dotnet/perf-autofiling-issues#8258

@ghostghost locked as resolved and limited conversation to collaborators Oct 2, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants

@SingleAccretion@EgorBo@jakobbotsch@DrewScoggins@AndyAyersMS@TIHan@tannergooding