JIT: Consume FMA intrinsic operands in right order - #102914

Merged
jakobbotsch merged 1 commit into
dotnet:mainfrom
jakobbotsch:fix-102773
May 31, 2024
Merged

JIT: Consume FMA intrinsic operands in right order#102914
jakobbotsch merged 1 commit into
dotnet:mainfrom
jakobbotsch:fix-102773

Conversation

@jakobbotsch

@jakobbotschjakobbotsch commented May 31, 2024

Copy link
Copy Markdown
Member

The operands of the FMA intrinsics are permuted in a non-standard way during LSRA. Codegen already takes this into account, but the handling was missing when consuming the operands.

Ideally we would permute these during lowering instead to avoid these hacks.

Fix#102773

The operands of the FMA intrinsic are permuted in a non-standard way
during LSRA. Codegen already takes this into account, but the handling
was missing when consuming the operands.
Ideally we would permute these during lowering instead to avoid these
hacks.
Fixdotnet#102773
@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 31, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@jakobbotsch
jakobbotsch marked this pull request as ready for review May 31, 2024 15:08
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

cc @dotnet/jit-contrib PTAL @tannergooding@kunalspathak

No diffs

This fixes the superpmi-replay blocking issue.

Comment on lines +3144 to +3158
#ifdef DEBUG
// Use nums are assigned in LIR order but this node is special and doesn't
// actually use operands. Fix up the use nums here to avoid asserts.
unsigned useNum1 = op1->gtUseNum;
unsigned useNum2 = op2->gtUseNum;
unsigned useNum3 = op3->gtUseNum;
emitOp1->gtUseNum = useNum1;
emitOp2->gtUseNum = useNum2;
emitOp3->gtUseNum = useNum3;
#endif

genConsumeRegs(emitOp1);
genConsumeRegs(emitOp2);
genConsumeRegs(emitOp3);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ideally we would permute these during lowering instead to avoid these hacks.

It's worth noting the reason we didn't do this in lowering (unlike most of the other cases that do swap operands) is because it would require introducing a lot of new synthetic intrinsics and it was believed to be overall more costly.

Each FusedMultiplyAdd intrinsic has three forms, where op3 is always the node that can be optionally contained:

  • 132 - op1 = (op1 * op3) + op2
  • 213 - `op1 = (op2 * op1) + op3
  • 231 - op1 = (op2 * op3) + op1

The managed API we expose is the 213 form and there are 10 different FMA intrinsics, so we'd need to expose 10 more for the 132 and 10 more for the 231 form. Then we'd need to repeat this for the Avx512 specific variants, giving us at least 50 new synthetic intrinsics in lowering just to cover FMA.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We might be able to avoid new synthetic intrinsics if we had a way to track what permutation it was, but free bits are fairly sparse right now. So I think we'd need to get clever in how we tracked that.

@kunalspathakkunalspathakMay 31, 2024

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We cannot use genConsumeMultiOpOperands() here instead?
Edit: I assume because we won't be using the same order as swapped operands?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, op1, op2, op3 (the order that genConsumeMultiOpOperands consumes in) is not the same as emitOp1, emitOp2, emitOp3, which as I understand it is the order that uses were built in by LSRA. We should consume in that order.

I don't have the context necessary to completely understand why we can't build and consume the operands in the op1, op2, op3 order even if we end up emitting different instructions using the registers in different orders in the instruction we emit.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@tannergooding - do you know?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That I don't. It's an area of the register allocator I'm not well versed in.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That I don't. It's an area of the register allocator I'm not well versed in.

Hhm, I think it is to do with how we are consuming them in codegen as opposed to the LSRA ordering.

and consume the operands in the op1, op2, op3 order even if we end up emitting different instructions using the registers in different order

@jakobbotsch
jakobbotsch merged commit e9cd3f1 into dotnet:mainMay 31, 2024
@jakobbotsch
jakobbotsch deleted the fix-102773 branch May 31, 2024 18:59
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jul 1, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

SPMI Replay failing sporadically

3 participants

@jakobbotsch@tannergooding@kunalspathak
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

JIT: Consume FMA intrinsic operands in right order - #102914

Merged
jakobbotsch merged 1 commit into
dotnet:mainfrom
jakobbotsch:fix-102773
May 31, 2024
Merged

JIT: Consume FMA intrinsic operands in right order#102914
jakobbotsch merged 1 commit into
dotnet:mainfrom
jakobbotsch:fix-102773

Conversation

@jakobbotsch

@jakobbotschjakobbotsch commented May 31, 2024

Copy link
Copy Markdown
Member

The operands of the FMA intrinsics are permuted in a non-standard way during LSRA. Codegen already takes this into account, but the handling was missing when consuming the operands.

Ideally we would permute these during lowering instead to avoid these hacks.

Fix#102773

The operands of the FMA intrinsic are permuted in a non-standard way
during LSRA. Codegen already takes this into account, but the handling
was missing when consuming the operands.
Ideally we would permute these during lowering instead to avoid these
hacks.
Fixdotnet#102773
@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 31, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@jakobbotsch
jakobbotsch marked this pull request as ready for review May 31, 2024 15:08
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

cc @dotnet/jit-contrib PTAL @tannergooding@kunalspathak

No diffs

This fixes the superpmi-replay blocking issue.

Comment on lines +3144 to +3158
#ifdef DEBUG
// Use nums are assigned in LIR order but this node is special and doesn't
// actually use operands. Fix up the use nums here to avoid asserts.
unsigned useNum1 = op1->gtUseNum;
unsigned useNum2 = op2->gtUseNum;
unsigned useNum3 = op3->gtUseNum;
emitOp1->gtUseNum = useNum1;
emitOp2->gtUseNum = useNum2;
emitOp3->gtUseNum = useNum3;
#endif

genConsumeRegs(emitOp1);
genConsumeRegs(emitOp2);
genConsumeRegs(emitOp3);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ideally we would permute these during lowering instead to avoid these hacks.

It's worth noting the reason we didn't do this in lowering (unlike most of the other cases that do swap operands) is because it would require introducing a lot of new synthetic intrinsics and it was believed to be overall more costly.

Each FusedMultiplyAdd intrinsic has three forms, where op3 is always the node that can be optionally contained:

  • 132 - op1 = (op1 * op3) + op2
  • 213 - `op1 = (op2 * op1) + op3
  • 231 - op1 = (op2 * op3) + op1

The managed API we expose is the 213 form and there are 10 different FMA intrinsics, so we'd need to expose 10 more for the 132 and 10 more for the 231 form. Then we'd need to repeat this for the Avx512 specific variants, giving us at least 50 new synthetic intrinsics in lowering just to cover FMA.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We might be able to avoid new synthetic intrinsics if we had a way to track what permutation it was, but free bits are fairly sparse right now. So I think we'd need to get clever in how we tracked that.

@kunalspathakkunalspathakMay 31, 2024

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We cannot use genConsumeMultiOpOperands() here instead?
Edit: I assume because we won't be using the same order as swapped operands?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, op1, op2, op3 (the order that genConsumeMultiOpOperands consumes in) is not the same as emitOp1, emitOp2, emitOp3, which as I understand it is the order that uses were built in by LSRA. We should consume in that order.

I don't have the context necessary to completely understand why we can't build and consume the operands in the op1, op2, op3 order even if we end up emitting different instructions using the registers in different orders in the instruction we emit.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@tannergooding - do you know?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That I don't. It's an area of the register allocator I'm not well versed in.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That I don't. It's an area of the register allocator I'm not well versed in.

Hhm, I think it is to do with how we are consuming them in codegen as opposed to the LSRA ordering.

and consume the operands in the op1, op2, op3 order even if we end up emitting different instructions using the registers in different order

@jakobbotsch
jakobbotsch merged commit e9cd3f1 into dotnet:mainMay 31, 2024
@jakobbotsch
jakobbotsch deleted the fix-102773 branch May 31, 2024 18:59
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jul 1, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

SPMI Replay failing sporadically

3 participants

@jakobbotsch@tannergooding@kunalspathak
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: Consume FMA intrinsic operands in right order - #102914

Merged
jakobbotsch merged 1 commit into
dotnet:mainfrom
jakobbotsch:fix-102773
May 31, 2024
Merged

JIT: Consume FMA intrinsic operands in right order#102914
jakobbotsch merged 1 commit into
dotnet:mainfrom
jakobbotsch:fix-102773

Conversation

@jakobbotsch

@jakobbotschjakobbotsch commented May 31, 2024

Copy link
Copy Markdown
Member

The operands of the FMA intrinsics are permuted in a non-standard way during LSRA. Codegen already takes this into account, but the handling was missing when consuming the operands.

Ideally we would permute these during lowering instead to avoid these hacks.

Fix#102773

The operands of the FMA intrinsic are permuted in a non-standard way
during LSRA. Codegen already takes this into account, but the handling
was missing when consuming the operands.
Ideally we would permute these during lowering instead to avoid these
hacks.
Fixdotnet#102773
@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 31, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@jakobbotsch
jakobbotsch marked this pull request as ready for review May 31, 2024 15:08
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

cc @dotnet/jit-contrib PTAL @tannergooding@kunalspathak

No diffs

This fixes the superpmi-replay blocking issue.

Comment on lines +3144 to +3158
#ifdef DEBUG
// Use nums are assigned in LIR order but this node is special and doesn't
// actually use operands. Fix up the use nums here to avoid asserts.
unsigned useNum1 = op1->gtUseNum;
unsigned useNum2 = op2->gtUseNum;
unsigned useNum3 = op3->gtUseNum;
emitOp1->gtUseNum = useNum1;
emitOp2->gtUseNum = useNum2;
emitOp3->gtUseNum = useNum3;
#endif

genConsumeRegs(emitOp1);
genConsumeRegs(emitOp2);
genConsumeRegs(emitOp3);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ideally we would permute these during lowering instead to avoid these hacks.

It's worth noting the reason we didn't do this in lowering (unlike most of the other cases that do swap operands) is because it would require introducing a lot of new synthetic intrinsics and it was believed to be overall more costly.

Each FusedMultiplyAdd intrinsic has three forms, where op3 is always the node that can be optionally contained:

  • 132 - op1 = (op1 * op3) + op2
  • 213 - `op1 = (op2 * op1) + op3
  • 231 - op1 = (op2 * op3) + op1

The managed API we expose is the 213 form and there are 10 different FMA intrinsics, so we'd need to expose 10 more for the 132 and 10 more for the 231 form. Then we'd need to repeat this for the Avx512 specific variants, giving us at least 50 new synthetic intrinsics in lowering just to cover FMA.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We might be able to avoid new synthetic intrinsics if we had a way to track what permutation it was, but free bits are fairly sparse right now. So I think we'd need to get clever in how we tracked that.

@kunalspathakkunalspathakMay 31, 2024

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We cannot use genConsumeMultiOpOperands() here instead?
Edit: I assume because we won't be using the same order as swapped operands?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, op1, op2, op3 (the order that genConsumeMultiOpOperands consumes in) is not the same as emitOp1, emitOp2, emitOp3, which as I understand it is the order that uses were built in by LSRA. We should consume in that order.

I don't have the context necessary to completely understand why we can't build and consume the operands in the op1, op2, op3 order even if we end up emitting different instructions using the registers in different orders in the instruction we emit.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@tannergooding - do you know?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That I don't. It's an area of the register allocator I'm not well versed in.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That I don't. It's an area of the register allocator I'm not well versed in.

Hhm, I think it is to do with how we are consuming them in codegen as opposed to the LSRA ordering.

and consume the operands in the op1, op2, op3 order even if we end up emitting different instructions using the registers in different order

@jakobbotsch
jakobbotsch merged commit e9cd3f1 into dotnet:mainMay 31, 2024
@jakobbotsch
jakobbotsch deleted the fix-102773 branch May 31, 2024 18:59
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jul 1, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

SPMI Replay failing sporadically

3 participants

@jakobbotsch@tannergooding@kunalspathak
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: Consume FMA intrinsic operands in right order - #102914

Merged
jakobbotsch merged 1 commit into
dotnet:mainfrom
jakobbotsch:fix-102773
May 31, 2024
Merged

JIT: Consume FMA intrinsic operands in right order#102914
jakobbotsch merged 1 commit into
dotnet:mainfrom
jakobbotsch:fix-102773

Conversation

@jakobbotsch

@jakobbotschjakobbotsch commented May 31, 2024

Copy link
Copy Markdown
Member

The operands of the FMA intrinsics are permuted in a non-standard way during LSRA. Codegen already takes this into account, but the handling was missing when consuming the operands.

Ideally we would permute these during lowering instead to avoid these hacks.

Fix#102773

The operands of the FMA intrinsic are permuted in a non-standard way
during LSRA. Codegen already takes this into account, but the handling
was missing when consuming the operands.
Ideally we would permute these during lowering instead to avoid these
hacks.
Fixdotnet#102773
@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 31, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@jakobbotsch
jakobbotsch marked this pull request as ready for review May 31, 2024 15:08
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

cc @dotnet/jit-contrib PTAL @tannergooding@kunalspathak

No diffs

This fixes the superpmi-replay blocking issue.

Comment on lines +3144 to +3158
#ifdef DEBUG
// Use nums are assigned in LIR order but this node is special and doesn't
// actually use operands. Fix up the use nums here to avoid asserts.
unsigned useNum1 = op1->gtUseNum;
unsigned useNum2 = op2->gtUseNum;
unsigned useNum3 = op3->gtUseNum;
emitOp1->gtUseNum = useNum1;
emitOp2->gtUseNum = useNum2;
emitOp3->gtUseNum = useNum3;
#endif

genConsumeRegs(emitOp1);
genConsumeRegs(emitOp2);
genConsumeRegs(emitOp3);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ideally we would permute these during lowering instead to avoid these hacks.

It's worth noting the reason we didn't do this in lowering (unlike most of the other cases that do swap operands) is because it would require introducing a lot of new synthetic intrinsics and it was believed to be overall more costly.

Each FusedMultiplyAdd intrinsic has three forms, where op3 is always the node that can be optionally contained:

  • 132 - op1 = (op1 * op3) + op2
  • 213 - `op1 = (op2 * op1) + op3
  • 231 - op1 = (op2 * op3) + op1

The managed API we expose is the 213 form and there are 10 different FMA intrinsics, so we'd need to expose 10 more for the 132 and 10 more for the 231 form. Then we'd need to repeat this for the Avx512 specific variants, giving us at least 50 new synthetic intrinsics in lowering just to cover FMA.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We might be able to avoid new synthetic intrinsics if we had a way to track what permutation it was, but free bits are fairly sparse right now. So I think we'd need to get clever in how we tracked that.

@kunalspathakkunalspathakMay 31, 2024

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We cannot use genConsumeMultiOpOperands() here instead?
Edit: I assume because we won't be using the same order as swapped operands?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, op1, op2, op3 (the order that genConsumeMultiOpOperands consumes in) is not the same as emitOp1, emitOp2, emitOp3, which as I understand it is the order that uses were built in by LSRA. We should consume in that order.

I don't have the context necessary to completely understand why we can't build and consume the operands in the op1, op2, op3 order even if we end up emitting different instructions using the registers in different orders in the instruction we emit.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@tannergooding - do you know?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That I don't. It's an area of the register allocator I'm not well versed in.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That I don't. It's an area of the register allocator I'm not well versed in.

Hhm, I think it is to do with how we are consuming them in codegen as opposed to the LSRA ordering.

and consume the operands in the op1, op2, op3 order even if we end up emitting different instructions using the registers in different order

@jakobbotsch
jakobbotsch merged commit e9cd3f1 into dotnet:mainMay 31, 2024
@jakobbotsch
jakobbotsch deleted the fix-102773 branch May 31, 2024 18:59
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jul 1, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

SPMI Replay failing sporadically

3 participants

@jakobbotsch@tannergooding@kunalspathak
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

JIT: Consume FMA intrinsic operands in right order - #102914

Merged
jakobbotsch merged 1 commit into
dotnet:mainfrom
jakobbotsch:fix-102773
May 31, 2024
Merged

JIT: Consume FMA intrinsic operands in right order#102914
jakobbotsch merged 1 commit into
dotnet:mainfrom
jakobbotsch:fix-102773

Conversation

@jakobbotsch

@jakobbotschjakobbotsch commented May 31, 2024

Copy link
Copy Markdown
Member

The operands of the FMA intrinsics are permuted in a non-standard way during LSRA. Codegen already takes this into account, but the handling was missing when consuming the operands.

Ideally we would permute these during lowering instead to avoid these hacks.

Fix#102773

The operands of the FMA intrinsic are permuted in a non-standard way
during LSRA. Codegen already takes this into account, but the handling
was missing when consuming the operands.
Ideally we would permute these during lowering instead to avoid these
hacks.
Fixdotnet#102773
@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 31, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@jakobbotsch
jakobbotsch marked this pull request as ready for review May 31, 2024 15:08
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

cc @dotnet/jit-contrib PTAL @tannergooding@kunalspathak

No diffs

This fixes the superpmi-replay blocking issue.

Comment on lines +3144 to +3158
#ifdef DEBUG
// Use nums are assigned in LIR order but this node is special and doesn't
// actually use operands. Fix up the use nums here to avoid asserts.
unsigned useNum1 = op1->gtUseNum;
unsigned useNum2 = op2->gtUseNum;
unsigned useNum3 = op3->gtUseNum;
emitOp1->gtUseNum = useNum1;
emitOp2->gtUseNum = useNum2;
emitOp3->gtUseNum = useNum3;
#endif

genConsumeRegs(emitOp1);
genConsumeRegs(emitOp2);
genConsumeRegs(emitOp3);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ideally we would permute these during lowering instead to avoid these hacks.

It's worth noting the reason we didn't do this in lowering (unlike most of the other cases that do swap operands) is because it would require introducing a lot of new synthetic intrinsics and it was believed to be overall more costly.

Each FusedMultiplyAdd intrinsic has three forms, where op3 is always the node that can be optionally contained:

  • 132 - op1 = (op1 * op3) + op2
  • 213 - `op1 = (op2 * op1) + op3
  • 231 - op1 = (op2 * op3) + op1

The managed API we expose is the 213 form and there are 10 different FMA intrinsics, so we'd need to expose 10 more for the 132 and 10 more for the 231 form. Then we'd need to repeat this for the Avx512 specific variants, giving us at least 50 new synthetic intrinsics in lowering just to cover FMA.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We might be able to avoid new synthetic intrinsics if we had a way to track what permutation it was, but free bits are fairly sparse right now. So I think we'd need to get clever in how we tracked that.

@kunalspathakkunalspathakMay 31, 2024

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We cannot use genConsumeMultiOpOperands() here instead?
Edit: I assume because we won't be using the same order as swapped operands?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, op1, op2, op3 (the order that genConsumeMultiOpOperands consumes in) is not the same as emitOp1, emitOp2, emitOp3, which as I understand it is the order that uses were built in by LSRA. We should consume in that order.

I don't have the context necessary to completely understand why we can't build and consume the operands in the op1, op2, op3 order even if we end up emitting different instructions using the registers in different orders in the instruction we emit.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@tannergooding - do you know?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That I don't. It's an area of the register allocator I'm not well versed in.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That I don't. It's an area of the register allocator I'm not well versed in.

Hhm, I think it is to do with how we are consuming them in codegen as opposed to the LSRA ordering.

and consume the operands in the op1, op2, op3 order even if we end up emitting different instructions using the registers in different order

@jakobbotsch
jakobbotsch merged commit e9cd3f1 into dotnet:mainMay 31, 2024
@jakobbotsch
jakobbotsch deleted the fix-102773 branch May 31, 2024 18:59
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jul 1, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

SPMI Replay failing sporadically

3 participants

@jakobbotsch@tannergooding@kunalspathak
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: Consume FMA intrinsic operands in right order - #102914

Merged
jakobbotsch merged 1 commit into
dotnet:mainfrom
jakobbotsch:fix-102773
May 31, 2024
Merged

JIT: Consume FMA intrinsic operands in right order#102914
jakobbotsch merged 1 commit into
dotnet:mainfrom
jakobbotsch:fix-102773

Conversation

@jakobbotsch

@jakobbotschjakobbotsch commented May 31, 2024

Copy link
Copy Markdown
Member

The operands of the FMA intrinsics are permuted in a non-standard way during LSRA. Codegen already takes this into account, but the handling was missing when consuming the operands.

Ideally we would permute these during lowering instead to avoid these hacks.

Fix#102773

The operands of the FMA intrinsic are permuted in a non-standard way
during LSRA. Codegen already takes this into account, but the handling
was missing when consuming the operands.
Ideally we would permute these during lowering instead to avoid these
hacks.
Fixdotnet#102773
@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 31, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@jakobbotsch
jakobbotsch marked this pull request as ready for review May 31, 2024 15:08
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

cc @dotnet/jit-contrib PTAL @tannergooding@kunalspathak

No diffs

This fixes the superpmi-replay blocking issue.

Comment on lines +3144 to +3158
#ifdef DEBUG
// Use nums are assigned in LIR order but this node is special and doesn't
// actually use operands. Fix up the use nums here to avoid asserts.
unsigned useNum1 = op1->gtUseNum;
unsigned useNum2 = op2->gtUseNum;
unsigned useNum3 = op3->gtUseNum;
emitOp1->gtUseNum = useNum1;
emitOp2->gtUseNum = useNum2;
emitOp3->gtUseNum = useNum3;
#endif

genConsumeRegs(emitOp1);
genConsumeRegs(emitOp2);
genConsumeRegs(emitOp3);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ideally we would permute these during lowering instead to avoid these hacks.

It's worth noting the reason we didn't do this in lowering (unlike most of the other cases that do swap operands) is because it would require introducing a lot of new synthetic intrinsics and it was believed to be overall more costly.

Each FusedMultiplyAdd intrinsic has three forms, where op3 is always the node that can be optionally contained:

  • 132 - op1 = (op1 * op3) + op2
  • 213 - `op1 = (op2 * op1) + op3
  • 231 - op1 = (op2 * op3) + op1

The managed API we expose is the 213 form and there are 10 different FMA intrinsics, so we'd need to expose 10 more for the 132 and 10 more for the 231 form. Then we'd need to repeat this for the Avx512 specific variants, giving us at least 50 new synthetic intrinsics in lowering just to cover FMA.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We might be able to avoid new synthetic intrinsics if we had a way to track what permutation it was, but free bits are fairly sparse right now. So I think we'd need to get clever in how we tracked that.

@kunalspathakkunalspathakMay 31, 2024

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We cannot use genConsumeMultiOpOperands() here instead?
Edit: I assume because we won't be using the same order as swapped operands?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, op1, op2, op3 (the order that genConsumeMultiOpOperands consumes in) is not the same as emitOp1, emitOp2, emitOp3, which as I understand it is the order that uses were built in by LSRA. We should consume in that order.

I don't have the context necessary to completely understand why we can't build and consume the operands in the op1, op2, op3 order even if we end up emitting different instructions using the registers in different orders in the instruction we emit.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@tannergooding - do you know?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That I don't. It's an area of the register allocator I'm not well versed in.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That I don't. It's an area of the register allocator I'm not well versed in.

Hhm, I think it is to do with how we are consuming them in codegen as opposed to the LSRA ordering.

and consume the operands in the op1, op2, op3 order even if we end up emitting different instructions using the registers in different order

@jakobbotsch
jakobbotsch merged commit e9cd3f1 into dotnet:mainMay 31, 2024
@jakobbotsch
jakobbotsch deleted the fix-102773 branch May 31, 2024 18:59
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jul 1, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

SPMI Replay failing sporadically

3 participants

@jakobbotsch@tannergooding@kunalspathak
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: Consume FMA intrinsic operands in right order - #102914

Merged
jakobbotsch merged 1 commit into
dotnet:mainfrom
jakobbotsch:fix-102773
May 31, 2024
Merged

JIT: Consume FMA intrinsic operands in right order#102914
jakobbotsch merged 1 commit into
dotnet:mainfrom
jakobbotsch:fix-102773

Conversation

@jakobbotsch

@jakobbotschjakobbotsch commented May 31, 2024

Copy link
Copy Markdown
Member

The operands of the FMA intrinsics are permuted in a non-standard way during LSRA. Codegen already takes this into account, but the handling was missing when consuming the operands.

Ideally we would permute these during lowering instead to avoid these hacks.

Fix#102773

The operands of the FMA intrinsic are permuted in a non-standard way
during LSRA. Codegen already takes this into account, but the handling
was missing when consuming the operands.
Ideally we would permute these during lowering instead to avoid these
hacks.
Fixdotnet#102773
@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 31, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@jakobbotsch
jakobbotsch marked this pull request as ready for review May 31, 2024 15:08
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

cc @dotnet/jit-contrib PTAL @tannergooding@kunalspathak

No diffs

This fixes the superpmi-replay blocking issue.

Comment on lines +3144 to +3158
#ifdef DEBUG
// Use nums are assigned in LIR order but this node is special and doesn't
// actually use operands. Fix up the use nums here to avoid asserts.
unsigned useNum1 = op1->gtUseNum;
unsigned useNum2 = op2->gtUseNum;
unsigned useNum3 = op3->gtUseNum;
emitOp1->gtUseNum = useNum1;
emitOp2->gtUseNum = useNum2;
emitOp3->gtUseNum = useNum3;
#endif

genConsumeRegs(emitOp1);
genConsumeRegs(emitOp2);
genConsumeRegs(emitOp3);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ideally we would permute these during lowering instead to avoid these hacks.

It's worth noting the reason we didn't do this in lowering (unlike most of the other cases that do swap operands) is because it would require introducing a lot of new synthetic intrinsics and it was believed to be overall more costly.

Each FusedMultiplyAdd intrinsic has three forms, where op3 is always the node that can be optionally contained:

  • 132 - op1 = (op1 * op3) + op2
  • 213 - `op1 = (op2 * op1) + op3
  • 231 - op1 = (op2 * op3) + op1

The managed API we expose is the 213 form and there are 10 different FMA intrinsics, so we'd need to expose 10 more for the 132 and 10 more for the 231 form. Then we'd need to repeat this for the Avx512 specific variants, giving us at least 50 new synthetic intrinsics in lowering just to cover FMA.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We might be able to avoid new synthetic intrinsics if we had a way to track what permutation it was, but free bits are fairly sparse right now. So I think we'd need to get clever in how we tracked that.

@kunalspathakkunalspathakMay 31, 2024

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We cannot use genConsumeMultiOpOperands() here instead?
Edit: I assume because we won't be using the same order as swapped operands?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, op1, op2, op3 (the order that genConsumeMultiOpOperands consumes in) is not the same as emitOp1, emitOp2, emitOp3, which as I understand it is the order that uses were built in by LSRA. We should consume in that order.

I don't have the context necessary to completely understand why we can't build and consume the operands in the op1, op2, op3 order even if we end up emitting different instructions using the registers in different orders in the instruction we emit.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@tannergooding - do you know?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That I don't. It's an area of the register allocator I'm not well versed in.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That I don't. It's an area of the register allocator I'm not well versed in.

Hhm, I think it is to do with how we are consuming them in codegen as opposed to the LSRA ordering.

and consume the operands in the op1, op2, op3 order even if we end up emitting different instructions using the registers in different order

@jakobbotsch
jakobbotsch merged commit e9cd3f1 into dotnet:mainMay 31, 2024
@jakobbotsch
jakobbotsch deleted the fix-102773 branch May 31, 2024 18:59
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jul 1, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

SPMI Replay failing sporadically

3 participants

@jakobbotsch@tannergooding@kunalspathak
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

JIT: Consume FMA intrinsic operands in right order - #102914

Merged
jakobbotsch merged 1 commit into
dotnet:mainfrom
jakobbotsch:fix-102773
May 31, 2024
Merged

JIT: Consume FMA intrinsic operands in right order#102914
jakobbotsch merged 1 commit into
dotnet:mainfrom
jakobbotsch:fix-102773

Conversation

@jakobbotsch

@jakobbotschjakobbotsch commented May 31, 2024

Copy link
Copy Markdown
Member

The operands of the FMA intrinsics are permuted in a non-standard way during LSRA. Codegen already takes this into account, but the handling was missing when consuming the operands.

Ideally we would permute these during lowering instead to avoid these hacks.

Fix#102773

The operands of the FMA intrinsic are permuted in a non-standard way
during LSRA. Codegen already takes this into account, but the handling
was missing when consuming the operands.
Ideally we would permute these during lowering instead to avoid these
hacks.
Fixdotnet#102773
@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label May 31, 2024
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@jakobbotsch
jakobbotsch marked this pull request as ready for review May 31, 2024 15:08
@jakobbotsch

Copy link
Copy Markdown
MemberAuthor

cc @dotnet/jit-contrib PTAL @tannergooding@kunalspathak

No diffs

This fixes the superpmi-replay blocking issue.

Comment on lines +3144 to +3158
#ifdef DEBUG
// Use nums are assigned in LIR order but this node is special and doesn't
// actually use operands. Fix up the use nums here to avoid asserts.
unsigned useNum1 = op1->gtUseNum;
unsigned useNum2 = op2->gtUseNum;
unsigned useNum3 = op3->gtUseNum;
emitOp1->gtUseNum = useNum1;
emitOp2->gtUseNum = useNum2;
emitOp3->gtUseNum = useNum3;
#endif

genConsumeRegs(emitOp1);
genConsumeRegs(emitOp2);
genConsumeRegs(emitOp3);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ideally we would permute these during lowering instead to avoid these hacks.

It's worth noting the reason we didn't do this in lowering (unlike most of the other cases that do swap operands) is because it would require introducing a lot of new synthetic intrinsics and it was believed to be overall more costly.

Each FusedMultiplyAdd intrinsic has three forms, where op3 is always the node that can be optionally contained:

  • 132 - op1 = (op1 * op3) + op2
  • 213 - `op1 = (op2 * op1) + op3
  • 231 - op1 = (op2 * op3) + op1

The managed API we expose is the 213 form and there are 10 different FMA intrinsics, so we'd need to expose 10 more for the 132 and 10 more for the 231 form. Then we'd need to repeat this for the Avx512 specific variants, giving us at least 50 new synthetic intrinsics in lowering just to cover FMA.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We might be able to avoid new synthetic intrinsics if we had a way to track what permutation it was, but free bits are fairly sparse right now. So I think we'd need to get clever in how we tracked that.

@kunalspathakkunalspathakMay 31, 2024

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We cannot use genConsumeMultiOpOperands() here instead?
Edit: I assume because we won't be using the same order as swapped operands?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, op1, op2, op3 (the order that genConsumeMultiOpOperands consumes in) is not the same as emitOp1, emitOp2, emitOp3, which as I understand it is the order that uses were built in by LSRA. We should consume in that order.

I don't have the context necessary to completely understand why we can't build and consume the operands in the op1, op2, op3 order even if we end up emitting different instructions using the registers in different orders in the instruction we emit.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@tannergooding - do you know?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That I don't. It's an area of the register allocator I'm not well versed in.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That I don't. It's an area of the register allocator I'm not well versed in.

Hhm, I think it is to do with how we are consuming them in codegen as opposed to the LSRA ordering.

and consume the operands in the op1, op2, op3 order even if we end up emitting different instructions using the registers in different order

@jakobbotsch
jakobbotsch merged commit e9cd3f1 into dotnet:mainMay 31, 2024
@jakobbotsch
jakobbotsch deleted the fix-102773 branch May 31, 2024 18:59
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jul 1, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

SPMI Replay failing sporadically

3 participants

@jakobbotsch@tannergooding@kunalspathak