Decompose some bitwise operations in HIR to allow more overall optimizations to kick in - #104517

Merged
tannergooding merged 14 commits into
dotnet:mainfrom
tannergooding:simd-andn
Jul 13, 2024
Merged

Decompose some bitwise operations in HIR to allow more overall optimizations to kick in#104517
tannergooding merged 14 commits into
dotnet:mainfrom
tannergooding:simd-andn

Conversation

@tannergooding

@tannergoodingtannergooding commented Jul 7, 2024

Copy link
Copy Markdown
Member

Doing this allows more scenarios involving CSE, morph transforms, and other optimizations to kick in. This will also longer term allow ~Compare to be optimized to the inverse comparison as well.

When such optimizations aren't feasible, we still end up getting the good codegen in lowering anyways.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 7, 2024
@tannergooding
tannergooding marked this pull request as ready for review July 9, 2024 04:59
@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @dotnet/jit-contrib, this is ready for review.

It gives some nice improvements on both Arm64 and x64 due to additional optimization opportunities. On xarch in particular it also gives better codegen almost anywhere vpternlog can be used and handles more (but not all) of the possible vpternlog lightups.

There are a couple methods where the code size has increased, rather than decreased, due to vpternlog but its a minority overall and typically due to 2 of the inputs being containable. These should be addressable, but I'd rather do that separately since this is still a large net improvement (generally speaking this should entail checking if two operands are containable and opting to not combine them into a vpternlog in that scenario).

@tannergooding

Copy link
Copy Markdown
MemberAuthor

The SPMI replay failure is #104585

@tannergooding

Copy link
Copy Markdown
MemberAuthor

Rerunning now that the SPMI replay issue should be fixed. This should still be ready for review and notably also fixes some issues that were found by antigen

// We specially handle this here since we're only producing a
// native intrinsic node in LIR

std::swap(op1, op2);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can you remind me - is it fine to do this swap here in terms of side-effect reordering?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ah, it's LIR so I guess it is

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, in this case its fine specifically because we've validated we're in LIR already.

Otherwise, we should be applying GTF_REVERSE_OPS or have some comment about expecting the user to have already spilled side effects, etc.

@EgorBoEgorBo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Changes LGTM assuming CI failures are unrelated

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@tannergooding@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Decompose some bitwise operations in HIR to allow more overall optimizations to kick in - #104517

Merged
tannergooding merged 14 commits into
dotnet:mainfrom
tannergooding:simd-andn
Jul 13, 2024
Merged

Decompose some bitwise operations in HIR to allow more overall optimizations to kick in#104517
tannergooding merged 14 commits into
dotnet:mainfrom
tannergooding:simd-andn

Conversation

@tannergooding

@tannergoodingtannergooding commented Jul 7, 2024

Copy link
Copy Markdown
Member

Doing this allows more scenarios involving CSE, morph transforms, and other optimizations to kick in. This will also longer term allow ~Compare to be optimized to the inverse comparison as well.

When such optimizations aren't feasible, we still end up getting the good codegen in lowering anyways.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 7, 2024
@tannergooding
tannergooding marked this pull request as ready for review July 9, 2024 04:59
@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @dotnet/jit-contrib, this is ready for review.

It gives some nice improvements on both Arm64 and x64 due to additional optimization opportunities. On xarch in particular it also gives better codegen almost anywhere vpternlog can be used and handles more (but not all) of the possible vpternlog lightups.

There are a couple methods where the code size has increased, rather than decreased, due to vpternlog but its a minority overall and typically due to 2 of the inputs being containable. These should be addressable, but I'd rather do that separately since this is still a large net improvement (generally speaking this should entail checking if two operands are containable and opting to not combine them into a vpternlog in that scenario).

@tannergooding

Copy link
Copy Markdown
MemberAuthor

The SPMI replay failure is #104585

@tannergooding

Copy link
Copy Markdown
MemberAuthor

Rerunning now that the SPMI replay issue should be fixed. This should still be ready for review and notably also fixes some issues that were found by antigen

// We specially handle this here since we're only producing a
// native intrinsic node in LIR

std::swap(op1, op2);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can you remind me - is it fine to do this swap here in terms of side-effect reordering?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ah, it's LIR so I guess it is

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, in this case its fine specifically because we've validated we're in LIR already.

Otherwise, we should be applying GTF_REVERSE_OPS or have some comment about expecting the user to have already spilled side effects, etc.

@EgorBoEgorBo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Changes LGTM assuming CI failures are unrelated

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@tannergooding@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Decompose some bitwise operations in HIR to allow more overall optimizations to kick in - #104517

Merged
tannergooding merged 14 commits into
dotnet:mainfrom
tannergooding:simd-andn
Jul 13, 2024
Merged

Decompose some bitwise operations in HIR to allow more overall optimizations to kick in#104517
tannergooding merged 14 commits into
dotnet:mainfrom
tannergooding:simd-andn

Conversation

@tannergooding

@tannergoodingtannergooding commented Jul 7, 2024

Copy link
Copy Markdown
Member

Doing this allows more scenarios involving CSE, morph transforms, and other optimizations to kick in. This will also longer term allow ~Compare to be optimized to the inverse comparison as well.

When such optimizations aren't feasible, we still end up getting the good codegen in lowering anyways.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 7, 2024
@tannergooding
tannergooding marked this pull request as ready for review July 9, 2024 04:59
@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @dotnet/jit-contrib, this is ready for review.

It gives some nice improvements on both Arm64 and x64 due to additional optimization opportunities. On xarch in particular it also gives better codegen almost anywhere vpternlog can be used and handles more (but not all) of the possible vpternlog lightups.

There are a couple methods where the code size has increased, rather than decreased, due to vpternlog but its a minority overall and typically due to 2 of the inputs being containable. These should be addressable, but I'd rather do that separately since this is still a large net improvement (generally speaking this should entail checking if two operands are containable and opting to not combine them into a vpternlog in that scenario).

@tannergooding

Copy link
Copy Markdown
MemberAuthor

The SPMI replay failure is #104585

@tannergooding

Copy link
Copy Markdown
MemberAuthor

Rerunning now that the SPMI replay issue should be fixed. This should still be ready for review and notably also fixes some issues that were found by antigen

// We specially handle this here since we're only producing a
// native intrinsic node in LIR

std::swap(op1, op2);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can you remind me - is it fine to do this swap here in terms of side-effect reordering?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ah, it's LIR so I guess it is

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, in this case its fine specifically because we've validated we're in LIR already.

Otherwise, we should be applying GTF_REVERSE_OPS or have some comment about expecting the user to have already spilled side effects, etc.

@EgorBoEgorBo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Changes LGTM assuming CI failures are unrelated

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@tannergooding@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Decompose some bitwise operations in HIR to allow more overall optimizations to kick in - #104517

Merged
tannergooding merged 14 commits into
dotnet:mainfrom
tannergooding:simd-andn
Jul 13, 2024
Merged

Decompose some bitwise operations in HIR to allow more overall optimizations to kick in#104517
tannergooding merged 14 commits into
dotnet:mainfrom
tannergooding:simd-andn

Conversation

@tannergooding

@tannergoodingtannergooding commented Jul 7, 2024

Copy link
Copy Markdown
Member

Doing this allows more scenarios involving CSE, morph transforms, and other optimizations to kick in. This will also longer term allow ~Compare to be optimized to the inverse comparison as well.

When such optimizations aren't feasible, we still end up getting the good codegen in lowering anyways.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 7, 2024
@tannergooding
tannergooding marked this pull request as ready for review July 9, 2024 04:59
@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @dotnet/jit-contrib, this is ready for review.

It gives some nice improvements on both Arm64 and x64 due to additional optimization opportunities. On xarch in particular it also gives better codegen almost anywhere vpternlog can be used and handles more (but not all) of the possible vpternlog lightups.

There are a couple methods where the code size has increased, rather than decreased, due to vpternlog but its a minority overall and typically due to 2 of the inputs being containable. These should be addressable, but I'd rather do that separately since this is still a large net improvement (generally speaking this should entail checking if two operands are containable and opting to not combine them into a vpternlog in that scenario).

@tannergooding

Copy link
Copy Markdown
MemberAuthor

The SPMI replay failure is #104585

@tannergooding

Copy link
Copy Markdown
MemberAuthor

Rerunning now that the SPMI replay issue should be fixed. This should still be ready for review and notably also fixes some issues that were found by antigen

// We specially handle this here since we're only producing a
// native intrinsic node in LIR

std::swap(op1, op2);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can you remind me - is it fine to do this swap here in terms of side-effect reordering?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ah, it's LIR so I guess it is

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, in this case its fine specifically because we've validated we're in LIR already.

Otherwise, we should be applying GTF_REVERSE_OPS or have some comment about expecting the user to have already spilled side effects, etc.

@EgorBoEgorBo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Changes LGTM assuming CI failures are unrelated

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@tannergooding@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Decompose some bitwise operations in HIR to allow more overall optimizations to kick in - #104517

Merged
tannergooding merged 14 commits into
dotnet:mainfrom
tannergooding:simd-andn
Jul 13, 2024
Merged

Decompose some bitwise operations in HIR to allow more overall optimizations to kick in#104517
tannergooding merged 14 commits into
dotnet:mainfrom
tannergooding:simd-andn

Conversation

@tannergooding

@tannergoodingtannergooding commented Jul 7, 2024

Copy link
Copy Markdown
Member

Doing this allows more scenarios involving CSE, morph transforms, and other optimizations to kick in. This will also longer term allow ~Compare to be optimized to the inverse comparison as well.

When such optimizations aren't feasible, we still end up getting the good codegen in lowering anyways.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 7, 2024
@tannergooding
tannergooding marked this pull request as ready for review July 9, 2024 04:59
@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @dotnet/jit-contrib, this is ready for review.

It gives some nice improvements on both Arm64 and x64 due to additional optimization opportunities. On xarch in particular it also gives better codegen almost anywhere vpternlog can be used and handles more (but not all) of the possible vpternlog lightups.

There are a couple methods where the code size has increased, rather than decreased, due to vpternlog but its a minority overall and typically due to 2 of the inputs being containable. These should be addressable, but I'd rather do that separately since this is still a large net improvement (generally speaking this should entail checking if two operands are containable and opting to not combine them into a vpternlog in that scenario).

@tannergooding

Copy link
Copy Markdown
MemberAuthor

The SPMI replay failure is #104585

@tannergooding

Copy link
Copy Markdown
MemberAuthor

Rerunning now that the SPMI replay issue should be fixed. This should still be ready for review and notably also fixes some issues that were found by antigen

// We specially handle this here since we're only producing a
// native intrinsic node in LIR

std::swap(op1, op2);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can you remind me - is it fine to do this swap here in terms of side-effect reordering?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ah, it's LIR so I guess it is

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, in this case its fine specifically because we've validated we're in LIR already.

Otherwise, we should be applying GTF_REVERSE_OPS or have some comment about expecting the user to have already spilled side effects, etc.

@EgorBoEgorBo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Changes LGTM assuming CI failures are unrelated

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@tannergooding@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Decompose some bitwise operations in HIR to allow more overall optimizations to kick in - #104517

Merged
tannergooding merged 14 commits into
dotnet:mainfrom
tannergooding:simd-andn
Jul 13, 2024
Merged

Decompose some bitwise operations in HIR to allow more overall optimizations to kick in#104517
tannergooding merged 14 commits into
dotnet:mainfrom
tannergooding:simd-andn

Conversation

@tannergooding

@tannergoodingtannergooding commented Jul 7, 2024

Copy link
Copy Markdown
Member

Doing this allows more scenarios involving CSE, morph transforms, and other optimizations to kick in. This will also longer term allow ~Compare to be optimized to the inverse comparison as well.

When such optimizations aren't feasible, we still end up getting the good codegen in lowering anyways.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 7, 2024
@tannergooding
tannergooding marked this pull request as ready for review July 9, 2024 04:59
@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @dotnet/jit-contrib, this is ready for review.

It gives some nice improvements on both Arm64 and x64 due to additional optimization opportunities. On xarch in particular it also gives better codegen almost anywhere vpternlog can be used and handles more (but not all) of the possible vpternlog lightups.

There are a couple methods where the code size has increased, rather than decreased, due to vpternlog but its a minority overall and typically due to 2 of the inputs being containable. These should be addressable, but I'd rather do that separately since this is still a large net improvement (generally speaking this should entail checking if two operands are containable and opting to not combine them into a vpternlog in that scenario).

@tannergooding

Copy link
Copy Markdown
MemberAuthor

The SPMI replay failure is #104585

@tannergooding

Copy link
Copy Markdown
MemberAuthor

Rerunning now that the SPMI replay issue should be fixed. This should still be ready for review and notably also fixes some issues that were found by antigen

// We specially handle this here since we're only producing a
// native intrinsic node in LIR

std::swap(op1, op2);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can you remind me - is it fine to do this swap here in terms of side-effect reordering?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ah, it's LIR so I guess it is

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, in this case its fine specifically because we've validated we're in LIR already.

Otherwise, we should be applying GTF_REVERSE_OPS or have some comment about expecting the user to have already spilled side effects, etc.

@EgorBoEgorBo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Changes LGTM assuming CI failures are unrelated

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@tannergooding@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Decompose some bitwise operations in HIR to allow more overall optimizations to kick in - #104517

Merged
tannergooding merged 14 commits into
dotnet:mainfrom
tannergooding:simd-andn
Jul 13, 2024
Merged

Decompose some bitwise operations in HIR to allow more overall optimizations to kick in#104517
tannergooding merged 14 commits into
dotnet:mainfrom
tannergooding:simd-andn

Conversation

@tannergooding

@tannergoodingtannergooding commented Jul 7, 2024

Copy link
Copy Markdown
Member

Doing this allows more scenarios involving CSE, morph transforms, and other optimizations to kick in. This will also longer term allow ~Compare to be optimized to the inverse comparison as well.

When such optimizations aren't feasible, we still end up getting the good codegen in lowering anyways.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 7, 2024
@tannergooding
tannergooding marked this pull request as ready for review July 9, 2024 04:59
@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @dotnet/jit-contrib, this is ready for review.

It gives some nice improvements on both Arm64 and x64 due to additional optimization opportunities. On xarch in particular it also gives better codegen almost anywhere vpternlog can be used and handles more (but not all) of the possible vpternlog lightups.

There are a couple methods where the code size has increased, rather than decreased, due to vpternlog but its a minority overall and typically due to 2 of the inputs being containable. These should be addressable, but I'd rather do that separately since this is still a large net improvement (generally speaking this should entail checking if two operands are containable and opting to not combine them into a vpternlog in that scenario).

@tannergooding

Copy link
Copy Markdown
MemberAuthor

The SPMI replay failure is #104585

@tannergooding

Copy link
Copy Markdown
MemberAuthor

Rerunning now that the SPMI replay issue should be fixed. This should still be ready for review and notably also fixes some issues that were found by antigen

// We specially handle this here since we're only producing a
// native intrinsic node in LIR

std::swap(op1, op2);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can you remind me - is it fine to do this swap here in terms of side-effect reordering?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ah, it's LIR so I guess it is

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, in this case its fine specifically because we've validated we're in LIR already.

Otherwise, we should be applying GTF_REVERSE_OPS or have some comment about expecting the user to have already spilled side effects, etc.

@EgorBoEgorBo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Changes LGTM assuming CI failures are unrelated

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@tannergooding@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Decompose some bitwise operations in HIR to allow more overall optimizations to kick in - #104517

Merged
tannergooding merged 14 commits into
dotnet:mainfrom
tannergooding:simd-andn
Jul 13, 2024
Merged

Decompose some bitwise operations in HIR to allow more overall optimizations to kick in#104517
tannergooding merged 14 commits into
dotnet:mainfrom
tannergooding:simd-andn

Conversation

@tannergooding

@tannergoodingtannergooding commented Jul 7, 2024

Copy link
Copy Markdown
Member

Doing this allows more scenarios involving CSE, morph transforms, and other optimizations to kick in. This will also longer term allow ~Compare to be optimized to the inverse comparison as well.

When such optimizations aren't feasible, we still end up getting the good codegen in lowering anyways.

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 7, 2024
@tannergooding
tannergooding marked this pull request as ready for review July 9, 2024 04:59
@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @dotnet/jit-contrib, this is ready for review.

It gives some nice improvements on both Arm64 and x64 due to additional optimization opportunities. On xarch in particular it also gives better codegen almost anywhere vpternlog can be used and handles more (but not all) of the possible vpternlog lightups.

There are a couple methods where the code size has increased, rather than decreased, due to vpternlog but its a minority overall and typically due to 2 of the inputs being containable. These should be addressable, but I'd rather do that separately since this is still a large net improvement (generally speaking this should entail checking if two operands are containable and opting to not combine them into a vpternlog in that scenario).

@tannergooding

Copy link
Copy Markdown
MemberAuthor

The SPMI replay failure is #104585

@tannergooding

Copy link
Copy Markdown
MemberAuthor

Rerunning now that the SPMI replay issue should be fixed. This should still be ready for review and notably also fixes some issues that were found by antigen

// We specially handle this here since we're only producing a
// native intrinsic node in LIR

std::swap(op1, op2);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can you remind me - is it fine to do this swap here in terms of side-effect reordering?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ah, it's LIR so I guess it is

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, in this case its fine specifically because we've validated we're in LIR already.

Otherwise, we should be applying GTF_REVERSE_OPS or have some comment about expecting the user to have already spilled side effects, etc.

@EgorBoEgorBo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Changes LGTM assuming CI failures are unrelated

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@tannergooding@EgorBo