JIT: Transform arithmetic using distributive property - #126852

Open
BoyBaykiller wants to merge 18 commits into
dotnet:mainfrom
BoyBaykiller:transform-using-distributive-property
Open

JIT: Transform arithmetic using distributive property#126852
BoyBaykiller wants to merge 18 commits into
dotnet:mainfrom
BoyBaykiller:transform-using-distributive-property

Conversation

@BoyBaykiller

@BoyBaykillerBoyBaykiller commented Apr 13, 2026

Copy link
Copy Markdown
Contributor

Generalization of #126070
Progress towards #93954

Basically we weren't doing any simplification based on distributive property before. So transforming
((A op1 B) op2 (A op1 C)) => (A op1 (B op2 C)). And this adds some basic support. Examples:

intMulDistedOverAdd(intA,intB,intC){return(A*B)+(A*C);}
;; ------ BASEG_M000_IG02:moveax,edximuleax,r8dimulr9d,edxaddeax,r9d;; ------ DIFFG_M48043_IG02: ;; offset=0x0000leaeax,[r8+r9]imuleax,edx
boolAfterOptimizeBools(intA,intB){return(A&4)!=0||(A&8)!=0;}
;; ------ BASEmoveax,edxandeax,4andedx,8oreax,edx setne almovzxrax,al;; ------ DIFFG_M55610_IG02:testdl,12 setne almovzxrax,al

We still need something that changes order to enable this opt. So (A | B) | C becoming A | (B | C) in this case:

uintReassociate(uintfoo,uintflags){return(foo|(flags&256))|(flags&512);}

Or here, the C * A need to be reversed.

intReassociate(intA,intB,intC){return(A*B)+(C*A);}

@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Apr 13, 2026
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Apr 13, 2026
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/jit/morph.cpp Outdated
return tree;
}

if (((tree->gtFlags & GTF_PERSISTENT_SIDE_EFFECTS) != 0) || ((tree->gtFlags & GTF_ORDER_SIDEEFF) != 0))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This could use the same optimization you are trying to apply😅... Probably handled by c++ though, so this is purely documentation ;-)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Lol. A new level of self documenting code. I didn't even notice that, I copied this check from elsewhere. I guess this just goes to show how useful the optimization is. Unfortunately even with this PR it's still not handled : (

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Update: This now gets handled by calling this opt again after the "optimizeBools" phase where it simplifies e.g
(A & 4) != 0 || (A & 8) != 0 into ((A & 4) | (A & 8)) != 0.

@BoyBaykiller

BoyBaykiller commented Apr 15, 2026

Copy link
Copy Markdown
ContributorAuthor

@EgorBo PTAL. Diffs are rather small, but nonetheless I believe it's a step in the right direction. I've written down some future work (which all improve diffs). I think the most interesting case for real-world code is arround bitwise ops.
Also should this go into post-order or pre-order morphing?

@BoyBaykiller
BoyBaykiller marked this pull request as ready for review April 15, 2026 01:07
@BoyBaykiller

Copy link
Copy Markdown
ContributorAuthor

Calling fgMorphBlockStmt after optimizeBools fixes point 2 in my list and produces much better diffs oldnew

@BoyBaykiller

BoyBaykiller commented Apr 15, 2026

Copy link
Copy Markdown
ContributorAuthor

Regarding point 1. The issue is that this runs first:

elseif (mulShiftOpt && (lowestBit > 1) && jitIsScaleIndexMul(lowestBit))
{
int shift = genLog2(lowestBit);
ssize_t factor = abs_mult >> shift;
if (factor == 3 || factor == 5 || factor == 9)
{
// if negative negate (min-int does not need negation)
if (mult < 0 && mult != SSIZE_T_MIN)
{
op1 = gtNewOperNode(GT_NEG, genActualType(op1), op1);
mul->gtOp1 = op1;
fgMorphTreeDone(op1);
}
// change the multiplication into a smaller multiplication (by 3, 5 or 9) and a shift
GenTree* const factorNode = gtNewIconNodeWithVN(this, factor, mul->TypeGet());
factorNode->SetMorphed(this);
op1 = gtNewOperNode(GT_MUL, mul->TypeGet(), op1, factorNode);
mul->gtOp1 = op1;
fgMorphTreeDone(op1);
op2->AsIntConCommon()->SetIconValue(shift);
changeToShift = true;
}
}

Perhaps this should be moved to Lower. LLVM also does it in "X86 DAG->DAG Instruction Selection" and not "InstCombine" https://godbolt.org/z/jsPea1PhP

Comment threadsrc/coreclr/jit/optimizebools.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated

auto isLeftDistributive = [](genTreeOps op1, genTreeOps op2) {
// op1 is left distributive over op2 iff:
// "A op1 (B op2 C)" <==> "(A op1 B) op2 (A op1 C)"

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We already do something similar for GT_ADD somewhere, we should unify it with that

@BoyBaykillerBoyBaykillerJun 4, 2026

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I didn't find what you are refering to here

Comment threadsrc/coreclr/jit/optimizebools.cpp Outdated
Comment threadsrc/coreclr/jit/gentree.cpp Outdated
@BoyBaykiller

Copy link
Copy Markdown
ContributorAuthor

I've limited it to OperIsAnyLocal, as discussed. Ready for an other review.

EgorBo pushed a commit that referenced this pull request Jun 4, 2026
Basically so we always have 0 at the right-side:
1. '(A & pow2) == pow2' -> '(A & pow2) != 0'
2. '(A & pow2) != pow2' -> '(A & pow2) == 0'
Tthis currently helps optimizeBools, but I am also adding this with the
future in mind.
Here is one example:
```cs
static bool Example(int foo)
{
return ((foo & 4) == 4) || ((foo & 8) == 8);
}
```
```assembly
;; ------ BASE
G_M000_IG02: ;; offset=0x0000
test cl, 4
je SHORT G_M000_IG05
G_M000_IG03: ;; offset=0x0005
mov eax, 1
G_M000_IG04: ;; offset=0x000A
ret G_M000_IG05: ;; offset=0x000B
test cl, 8
setne al
movzx rax, al
G_M000_IG06: ;; offset=0x0014
ret ;; ------ DIFF
G_M33938_IG02: ;; offset=0x0000
mov eax, ecx
and eax, 4
and ecx, 8
or eax, ecx
setne al
movzx rax, al
```
Note: With #126852 this would then
get transformed into just `test 12`.
@JulieLeeMSFT

Copy link
Copy Markdown
Member

@EgorBo, PTAL.

@EgorBo

Copy link
Copy Markdown
Member

@MihuBot -nuget

@EgorBo

Copy link
Copy Markdown
Member

@MihaZupan

Copy link
Copy Markdown
Member

Just testing error reporting changes

@MihuBot -nuget

@MihuBot

Copy link
Copy Markdown

Warning

2 extra test assemblies failed to produce JIT diffs and were excluded from the diff analysis:

  • MathNet.Numerics.dll (pr)
  • Pipelines.Sockets.Unofficial.dll (pr)

All failures happened only on the PR branch, which likely indicates a bug introduced by the PR (e.g. a JIT assert or crash while compiling these assemblies).

See job details in MihuBot/runtime-utils#2058, triggered by #126852 (comment).

eiriktsarpalis pushed a commit that referenced this pull request Jul 15, 2026
Basically so we always have 0 at the right-side:
1. '(A & pow2) == pow2' -> '(A & pow2) != 0'
2. '(A & pow2) != pow2' -> '(A & pow2) == 0'
Tthis currently helps optimizeBools, but I am also adding this with the
future in mind.
Here is one example:
```cs
static bool Example(int foo)
{
return ((foo & 4) == 4) || ((foo & 8) == 8);
}
```
```assembly
;; ------ BASE
G_M000_IG02: ;; offset=0x0000
test cl, 4
je SHORT G_M000_IG05
G_M000_IG03: ;; offset=0x0005
mov eax, 1
G_M000_IG04: ;; offset=0x000A
ret G_M000_IG05: ;; offset=0x000B
test cl, 8
setne al
movzx rax, al
G_M000_IG06: ;; offset=0x0014
ret ;; ------ DIFF
G_M33938_IG02: ;; offset=0x0000
mov eax, ecx
and eax, 4
and ecx, 8
or eax, ecx
setne al
movzx rax, al
```
Note: With #126852 this would then
get transformed into just `test 12`.
@github-actions

github-actionsBot commented Jul 19, 2026

Copy link
Copy Markdown
Contributor

Workflow state for the Holistic Review Orchestrator.

{
"version": 5,
"last_dispatched_commit": "ef82e04e0d0ec0a577f66e436b30a9adf3b47b77",
"last_dispatched_base_ref": "main",
"last_dispatched_base_sha": "f32edc0223b18865778e8e4d335e31bf9a2933f0",
"last_reviewed_commit": "ef82e04e0d0ec0a577f66e436b30a9adf3b47b77",
"last_reviewed_base_ref": "main",
"last_reviewed_base_sha": "f32edc0223b18865778e8e4d335e31bf9a2933f0",
"last_recorded_worker_run_id": "29679146097",
"review_attempt_commit": "",
"review_attempt_base_ref": "",
"review_attempt_count": 0,
"max_review_attempts": 5,
"review_history_format": "holistic-review-disclosure-v1",
"review_history": [
{
"commit": "ef82e04e0d0ec0a577f66e436b30a9adf3b47b77",
"review_id": 4730527763
}
]
}

@github-actionsgithub-actionsBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Holistic Review

Motivation: Generalizes #126070 and makes progress on #93954 by teaching the JIT to simplify arithmetic via the distributive property, i.e. (A op1 B) op2 (A op1 C) => A op1 (B op2 C). This reduces instruction count in common patterns (e.g. (A*B)+(A*C) folding to A*(B+C), and bit-test patterns after optOptimizeBools), which is a worthwhile codegen improvement.

Approach: Adds gtFoldDistributiveArithmetic invoked from gtFoldExprBinary for GT_AND/GT_OR/GT_XOR/GT_ADD/GT_SUB trees. It uses an isLeftDistributive table (AND/OR/MUL over their duals) and, when both operands share the same outer op and their left children are matching locals, rebuilds the tree with the common factor extracted. It also re-runs gtFoldExpr on the newly-combined boolean fold result in optOptimizeBoolsUpdateTrees, and tidies an unrelated cmp->gtOp2 -> cmp->gtGetOp2() reference in morph.cpp.

Summary: The transform is conceptually sound for wrapping (unchecked) integer arithmetic, and the guarding on optimization level, overflow of the outer node, side effects, and integral type is a reasonable start. However there is one correctness concern worth resolving before merge: the overflow check only examines the outer tree, not the inner operands. For checked(A*B) + checked(A*C) the inner GT_MUL nodes carry GTF_OVERFLOW, yet the rebuilt A*(B+C) uses a plain unchecked GT_MUL, dropping the overflow checks and altering observable behavior (a missing OverflowException). See the inline comment on the guard. I'd recommend adding a targeted regression test covering the checked-arithmetic case (and confirming it is rejected) once the guard is tightened. A secondary, non-blocking note: the freshly-created newOp2 node does not get value numbers assigned even though gtFoldExpr may be reached post-VN via optOptimizeBools; only result receives SetVNsFromNode. Verify this is benign in that phase (constants typically fold away, but a non-constant B op C combination would leave an unnumbered node).

Detailed Findings

  • Checked-arithmetic correctness (see inline comment on gtFoldDistributiveArithmetic): outer-only overflow guard can drop inner checked-multiply/add/sub semantics. Recommend also bailing when the operands being merged have gtOverflowEx()/GTF_EXCEPT set.
  • VN coverage of newOp2 (non-blocking): SetVNsFromNode(tree) is applied to result only; the intermediate newOp2 = gtFoldExpr(B op C) node has no VN. Since this path is also reachable from optOptimizeBoolsUpdateTrees after value numbering, confirm an unnumbered non-constant intermediate cannot reach a consumer that requires valid VNs.
  • Scope observation (non-blocking): the transform requires the shared factor to be OperIsAnyLocal. This is a safe, conservative choice (no duplicated side effects / re-evaluation), consistent with the PR description noting reassociation is not yet handled.

Note

This review was generated by this repository's Holistic Review agentic workflow to complement the built-in Copilot review.

Generated by Holistic Review · 115.6 AIC · ⌖ 11 AIC · ⊞ 10K

Comment threadsrc/coreclr/jit/gentree.cpp Outdated
{
assert(tree->OperIs(GT_AND, GT_OR, GT_XOR, GT_ADD, GT_SUB));

if (opts.OptimizationDisabled() || tree->gtOverflowEx() || !varTypeIsIntegralOrI(tree))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The overflow guard only inspects the outer tree, not the inner operands op1/op2. For an inner distributive op that can overflow-throw, this can silently drop checked semantics. Concretely, checked(A * B) + checked(A * C) has an unchecked outer GT_ADD but two inner GT_MUL nodes carrying GTF_OVERFLOW. isLeftDistributive(GT_MUL, GT_ADD) is true, so the transform rebuilds the tree as A * (B + C) via gtNewOperNode(GT_MUL, ...) with no overflow flag, discarding both checked multiplications. That changes observable behavior (a program that should throw OverflowException no longer does). For unchecked/wrapping arithmetic the distribution is correct mod 2^n, but the checked case is not. Consider also rejecting when op1/op2 (or their relevant children) have gtOverflowEx() set, or more conservatively bail when GTF_EXCEPT is present on the operands being merged.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@BoyBaykiller that is a good point, this PR basically breaks this:

staticintF(inta,intb,intc)=>checked(a*b)+checked(a*c);

With:

a = 2
b = 0x40000000
c = -0x40000000

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am now bailing on GTF_EXCEPT

Comment threadsrc/coreclr/jit/gentree.cpp Outdated
{
GenTree* newOp1 = op1->gtGetOp1();
GenTree* newOp2 =
gtFoldExpr(gtNewOperNode(tree->OperGet(), tree->TypeGet(), op1->gtGetOp2(), op2->gtGetOp2()));

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

seems like this gtNewOperNode is not marked as MORPHED

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

added fgMorphTreeDone(newOp2);

@BoyBaykiller

Copy link
Copy Markdown
ContributorAuthor

Failures fixed?

@MihuBot -nuget

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants

@BoyBaykiller@JulieLeeMSFT@EgorBo@MihaZupan@MihuBot@rhuijben@am11@jakobbotsch@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

JIT: Transform arithmetic using distributive property - #126852

Open
BoyBaykiller wants to merge 18 commits into
dotnet:mainfrom
BoyBaykiller:transform-using-distributive-property
Open

JIT: Transform arithmetic using distributive property#126852
BoyBaykiller wants to merge 18 commits into
dotnet:mainfrom
BoyBaykiller:transform-using-distributive-property

Conversation

@BoyBaykiller

@BoyBaykillerBoyBaykiller commented Apr 13, 2026

Copy link
Copy Markdown
Contributor

Generalization of #126070
Progress towards #93954

Basically we weren't doing any simplification based on distributive property before. So transforming
((A op1 B) op2 (A op1 C)) => (A op1 (B op2 C)). And this adds some basic support. Examples:

intMulDistedOverAdd(intA,intB,intC){return(A*B)+(A*C);}
;; ------ BASEG_M000_IG02:moveax,edximuleax,r8dimulr9d,edxaddeax,r9d;; ------ DIFFG_M48043_IG02: ;; offset=0x0000leaeax,[r8+r9]imuleax,edx
boolAfterOptimizeBools(intA,intB){return(A&4)!=0||(A&8)!=0;}
;; ------ BASEmoveax,edxandeax,4andedx,8oreax,edx setne almovzxrax,al;; ------ DIFFG_M55610_IG02:testdl,12 setne almovzxrax,al

We still need something that changes order to enable this opt. So (A | B) | C becoming A | (B | C) in this case:

uintReassociate(uintfoo,uintflags){return(foo|(flags&256))|(flags&512);}

Or here, the C * A need to be reversed.

intReassociate(intA,intB,intC){return(A*B)+(C*A);}

@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Apr 13, 2026
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Apr 13, 2026
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/jit/morph.cpp Outdated
return tree;
}

if (((tree->gtFlags & GTF_PERSISTENT_SIDE_EFFECTS) != 0) || ((tree->gtFlags & GTF_ORDER_SIDEEFF) != 0))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This could use the same optimization you are trying to apply😅... Probably handled by c++ though, so this is purely documentation ;-)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Lol. A new level of self documenting code. I didn't even notice that, I copied this check from elsewhere. I guess this just goes to show how useful the optimization is. Unfortunately even with this PR it's still not handled : (

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Update: This now gets handled by calling this opt again after the "optimizeBools" phase where it simplifies e.g
(A & 4) != 0 || (A & 8) != 0 into ((A & 4) | (A & 8)) != 0.

@BoyBaykiller

BoyBaykiller commented Apr 15, 2026

Copy link
Copy Markdown
ContributorAuthor

@EgorBo PTAL. Diffs are rather small, but nonetheless I believe it's a step in the right direction. I've written down some future work (which all improve diffs). I think the most interesting case for real-world code is arround bitwise ops.
Also should this go into post-order or pre-order morphing?

@BoyBaykiller
BoyBaykiller marked this pull request as ready for review April 15, 2026 01:07
@BoyBaykiller

Copy link
Copy Markdown
ContributorAuthor

Calling fgMorphBlockStmt after optimizeBools fixes point 2 in my list and produces much better diffs oldnew

@BoyBaykiller

BoyBaykiller commented Apr 15, 2026

Copy link
Copy Markdown
ContributorAuthor

Regarding point 1. The issue is that this runs first:

elseif (mulShiftOpt && (lowestBit > 1) && jitIsScaleIndexMul(lowestBit))
{
int shift = genLog2(lowestBit);
ssize_t factor = abs_mult >> shift;
if (factor == 3 || factor == 5 || factor == 9)
{
// if negative negate (min-int does not need negation)
if (mult < 0 && mult != SSIZE_T_MIN)
{
op1 = gtNewOperNode(GT_NEG, genActualType(op1), op1);
mul->gtOp1 = op1;
fgMorphTreeDone(op1);
}
// change the multiplication into a smaller multiplication (by 3, 5 or 9) and a shift
GenTree* const factorNode = gtNewIconNodeWithVN(this, factor, mul->TypeGet());
factorNode->SetMorphed(this);
op1 = gtNewOperNode(GT_MUL, mul->TypeGet(), op1, factorNode);
mul->gtOp1 = op1;
fgMorphTreeDone(op1);
op2->AsIntConCommon()->SetIconValue(shift);
changeToShift = true;
}
}

Perhaps this should be moved to Lower. LLVM also does it in "X86 DAG->DAG Instruction Selection" and not "InstCombine" https://godbolt.org/z/jsPea1PhP

Comment threadsrc/coreclr/jit/optimizebools.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated

auto isLeftDistributive = [](genTreeOps op1, genTreeOps op2) {
// op1 is left distributive over op2 iff:
// "A op1 (B op2 C)" <==> "(A op1 B) op2 (A op1 C)"

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We already do something similar for GT_ADD somewhere, we should unify it with that

@BoyBaykillerBoyBaykillerJun 4, 2026

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I didn't find what you are refering to here

Comment threadsrc/coreclr/jit/optimizebools.cpp Outdated
Comment threadsrc/coreclr/jit/gentree.cpp Outdated
@BoyBaykiller

Copy link
Copy Markdown
ContributorAuthor

I've limited it to OperIsAnyLocal, as discussed. Ready for an other review.

EgorBo pushed a commit that referenced this pull request Jun 4, 2026
Basically so we always have 0 at the right-side:
1. '(A & pow2) == pow2' -> '(A & pow2) != 0'
2. '(A & pow2) != pow2' -> '(A & pow2) == 0'
Tthis currently helps optimizeBools, but I am also adding this with the
future in mind.
Here is one example:
```cs
static bool Example(int foo)
{
return ((foo & 4) == 4) || ((foo & 8) == 8);
}
```
```assembly
;; ------ BASE
G_M000_IG02: ;; offset=0x0000
test cl, 4
je SHORT G_M000_IG05
G_M000_IG03: ;; offset=0x0005
mov eax, 1
G_M000_IG04: ;; offset=0x000A
ret G_M000_IG05: ;; offset=0x000B
test cl, 8
setne al
movzx rax, al
G_M000_IG06: ;; offset=0x0014
ret ;; ------ DIFF
G_M33938_IG02: ;; offset=0x0000
mov eax, ecx
and eax, 4
and ecx, 8
or eax, ecx
setne al
movzx rax, al
```
Note: With #126852 this would then
get transformed into just `test 12`.
@JulieLeeMSFT

Copy link
Copy Markdown
Member

@EgorBo, PTAL.

@EgorBo

Copy link
Copy Markdown
Member

@MihuBot -nuget

@EgorBo

Copy link
Copy Markdown
Member

@MihaZupan

Copy link
Copy Markdown
Member

Just testing error reporting changes

@MihuBot -nuget

@MihuBot

Copy link
Copy Markdown

Warning

2 extra test assemblies failed to produce JIT diffs and were excluded from the diff analysis:

  • MathNet.Numerics.dll (pr)
  • Pipelines.Sockets.Unofficial.dll (pr)

All failures happened only on the PR branch, which likely indicates a bug introduced by the PR (e.g. a JIT assert or crash while compiling these assemblies).

See job details in MihuBot/runtime-utils#2058, triggered by #126852 (comment).

eiriktsarpalis pushed a commit that referenced this pull request Jul 15, 2026
Basically so we always have 0 at the right-side:
1. '(A & pow2) == pow2' -> '(A & pow2) != 0'
2. '(A & pow2) != pow2' -> '(A & pow2) == 0'
Tthis currently helps optimizeBools, but I am also adding this with the
future in mind.
Here is one example:
```cs
static bool Example(int foo)
{
return ((foo & 4) == 4) || ((foo & 8) == 8);
}
```
```assembly
;; ------ BASE
G_M000_IG02: ;; offset=0x0000
test cl, 4
je SHORT G_M000_IG05
G_M000_IG03: ;; offset=0x0005
mov eax, 1
G_M000_IG04: ;; offset=0x000A
ret G_M000_IG05: ;; offset=0x000B
test cl, 8
setne al
movzx rax, al
G_M000_IG06: ;; offset=0x0014
ret ;; ------ DIFF
G_M33938_IG02: ;; offset=0x0000
mov eax, ecx
and eax, 4
and ecx, 8
or eax, ecx
setne al
movzx rax, al
```
Note: With #126852 this would then
get transformed into just `test 12`.
@github-actions

github-actionsBot commented Jul 19, 2026

Copy link
Copy Markdown
Contributor

Workflow state for the Holistic Review Orchestrator.

{
"version": 5,
"last_dispatched_commit": "ef82e04e0d0ec0a577f66e436b30a9adf3b47b77",
"last_dispatched_base_ref": "main",
"last_dispatched_base_sha": "f32edc0223b18865778e8e4d335e31bf9a2933f0",
"last_reviewed_commit": "ef82e04e0d0ec0a577f66e436b30a9adf3b47b77",
"last_reviewed_base_ref": "main",
"last_reviewed_base_sha": "f32edc0223b18865778e8e4d335e31bf9a2933f0",
"last_recorded_worker_run_id": "29679146097",
"review_attempt_commit": "",
"review_attempt_base_ref": "",
"review_attempt_count": 0,
"max_review_attempts": 5,
"review_history_format": "holistic-review-disclosure-v1",
"review_history": [
{
"commit": "ef82e04e0d0ec0a577f66e436b30a9adf3b47b77",
"review_id": 4730527763
}
]
}

@github-actionsgithub-actionsBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Holistic Review

Motivation: Generalizes #126070 and makes progress on #93954 by teaching the JIT to simplify arithmetic via the distributive property, i.e. (A op1 B) op2 (A op1 C) => A op1 (B op2 C). This reduces instruction count in common patterns (e.g. (A*B)+(A*C) folding to A*(B+C), and bit-test patterns after optOptimizeBools), which is a worthwhile codegen improvement.

Approach: Adds gtFoldDistributiveArithmetic invoked from gtFoldExprBinary for GT_AND/GT_OR/GT_XOR/GT_ADD/GT_SUB trees. It uses an isLeftDistributive table (AND/OR/MUL over their duals) and, when both operands share the same outer op and their left children are matching locals, rebuilds the tree with the common factor extracted. It also re-runs gtFoldExpr on the newly-combined boolean fold result in optOptimizeBoolsUpdateTrees, and tidies an unrelated cmp->gtOp2 -> cmp->gtGetOp2() reference in morph.cpp.

Summary: The transform is conceptually sound for wrapping (unchecked) integer arithmetic, and the guarding on optimization level, overflow of the outer node, side effects, and integral type is a reasonable start. However there is one correctness concern worth resolving before merge: the overflow check only examines the outer tree, not the inner operands. For checked(A*B) + checked(A*C) the inner GT_MUL nodes carry GTF_OVERFLOW, yet the rebuilt A*(B+C) uses a plain unchecked GT_MUL, dropping the overflow checks and altering observable behavior (a missing OverflowException). See the inline comment on the guard. I'd recommend adding a targeted regression test covering the checked-arithmetic case (and confirming it is rejected) once the guard is tightened. A secondary, non-blocking note: the freshly-created newOp2 node does not get value numbers assigned even though gtFoldExpr may be reached post-VN via optOptimizeBools; only result receives SetVNsFromNode. Verify this is benign in that phase (constants typically fold away, but a non-constant B op C combination would leave an unnumbered node).

Detailed Findings

  • Checked-arithmetic correctness (see inline comment on gtFoldDistributiveArithmetic): outer-only overflow guard can drop inner checked-multiply/add/sub semantics. Recommend also bailing when the operands being merged have gtOverflowEx()/GTF_EXCEPT set.
  • VN coverage of newOp2 (non-blocking): SetVNsFromNode(tree) is applied to result only; the intermediate newOp2 = gtFoldExpr(B op C) node has no VN. Since this path is also reachable from optOptimizeBoolsUpdateTrees after value numbering, confirm an unnumbered non-constant intermediate cannot reach a consumer that requires valid VNs.
  • Scope observation (non-blocking): the transform requires the shared factor to be OperIsAnyLocal. This is a safe, conservative choice (no duplicated side effects / re-evaluation), consistent with the PR description noting reassociation is not yet handled.

Note

This review was generated by this repository's Holistic Review agentic workflow to complement the built-in Copilot review.

Generated by Holistic Review · 115.6 AIC · ⌖ 11 AIC · ⊞ 10K

Comment threadsrc/coreclr/jit/gentree.cpp Outdated
{
assert(tree->OperIs(GT_AND, GT_OR, GT_XOR, GT_ADD, GT_SUB));

if (opts.OptimizationDisabled() || tree->gtOverflowEx() || !varTypeIsIntegralOrI(tree))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The overflow guard only inspects the outer tree, not the inner operands op1/op2. For an inner distributive op that can overflow-throw, this can silently drop checked semantics. Concretely, checked(A * B) + checked(A * C) has an unchecked outer GT_ADD but two inner GT_MUL nodes carrying GTF_OVERFLOW. isLeftDistributive(GT_MUL, GT_ADD) is true, so the transform rebuilds the tree as A * (B + C) via gtNewOperNode(GT_MUL, ...) with no overflow flag, discarding both checked multiplications. That changes observable behavior (a program that should throw OverflowException no longer does). For unchecked/wrapping arithmetic the distribution is correct mod 2^n, but the checked case is not. Consider also rejecting when op1/op2 (or their relevant children) have gtOverflowEx() set, or more conservatively bail when GTF_EXCEPT is present on the operands being merged.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@BoyBaykiller that is a good point, this PR basically breaks this:

staticintF(inta,intb,intc)=>checked(a*b)+checked(a*c);

With:

a = 2
b = 0x40000000
c = -0x40000000

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am now bailing on GTF_EXCEPT

Comment threadsrc/coreclr/jit/gentree.cpp Outdated
{
GenTree* newOp1 = op1->gtGetOp1();
GenTree* newOp2 =
gtFoldExpr(gtNewOperNode(tree->OperGet(), tree->TypeGet(), op1->gtGetOp2(), op2->gtGetOp2()));

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

seems like this gtNewOperNode is not marked as MORPHED

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

added fgMorphTreeDone(newOp2);

@BoyBaykiller

Copy link
Copy Markdown
ContributorAuthor

Failures fixed?

@MihuBot -nuget

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants

@BoyBaykiller@JulieLeeMSFT@EgorBo@MihaZupan@MihuBot@rhuijben@am11@jakobbotsch@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: Transform arithmetic using distributive property - #126852

Open
BoyBaykiller wants to merge 18 commits into
dotnet:mainfrom
BoyBaykiller:transform-using-distributive-property
Open

JIT: Transform arithmetic using distributive property#126852
BoyBaykiller wants to merge 18 commits into
dotnet:mainfrom
BoyBaykiller:transform-using-distributive-property

Conversation

@BoyBaykiller

@BoyBaykillerBoyBaykiller commented Apr 13, 2026

Copy link
Copy Markdown
Contributor

Generalization of #126070
Progress towards #93954

Basically we weren't doing any simplification based on distributive property before. So transforming
((A op1 B) op2 (A op1 C)) => (A op1 (B op2 C)). And this adds some basic support. Examples:

intMulDistedOverAdd(intA,intB,intC){return(A*B)+(A*C);}
;; ------ BASEG_M000_IG02:moveax,edximuleax,r8dimulr9d,edxaddeax,r9d;; ------ DIFFG_M48043_IG02: ;; offset=0x0000leaeax,[r8+r9]imuleax,edx
boolAfterOptimizeBools(intA,intB){return(A&4)!=0||(A&8)!=0;}
;; ------ BASEmoveax,edxandeax,4andedx,8oreax,edx setne almovzxrax,al;; ------ DIFFG_M55610_IG02:testdl,12 setne almovzxrax,al

We still need something that changes order to enable this opt. So (A | B) | C becoming A | (B | C) in this case:

uintReassociate(uintfoo,uintflags){return(foo|(flags&256))|(flags&512);}

Or here, the C * A need to be reversed.

intReassociate(intA,intB,intC){return(A*B)+(C*A);}

@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Apr 13, 2026
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Apr 13, 2026
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/jit/morph.cpp Outdated
return tree;
}

if (((tree->gtFlags & GTF_PERSISTENT_SIDE_EFFECTS) != 0) || ((tree->gtFlags & GTF_ORDER_SIDEEFF) != 0))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This could use the same optimization you are trying to apply😅... Probably handled by c++ though, so this is purely documentation ;-)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Lol. A new level of self documenting code. I didn't even notice that, I copied this check from elsewhere. I guess this just goes to show how useful the optimization is. Unfortunately even with this PR it's still not handled : (

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Update: This now gets handled by calling this opt again after the "optimizeBools" phase where it simplifies e.g
(A & 4) != 0 || (A & 8) != 0 into ((A & 4) | (A & 8)) != 0.

@BoyBaykiller

BoyBaykiller commented Apr 15, 2026

Copy link
Copy Markdown
ContributorAuthor

@EgorBo PTAL. Diffs are rather small, but nonetheless I believe it's a step in the right direction. I've written down some future work (which all improve diffs). I think the most interesting case for real-world code is arround bitwise ops.
Also should this go into post-order or pre-order morphing?

@BoyBaykiller
BoyBaykiller marked this pull request as ready for review April 15, 2026 01:07
@BoyBaykiller

Copy link
Copy Markdown
ContributorAuthor

Calling fgMorphBlockStmt after optimizeBools fixes point 2 in my list and produces much better diffs oldnew

@BoyBaykiller

BoyBaykiller commented Apr 15, 2026

Copy link
Copy Markdown
ContributorAuthor

Regarding point 1. The issue is that this runs first:

elseif (mulShiftOpt && (lowestBit > 1) && jitIsScaleIndexMul(lowestBit))
{
int shift = genLog2(lowestBit);
ssize_t factor = abs_mult >> shift;
if (factor == 3 || factor == 5 || factor == 9)
{
// if negative negate (min-int does not need negation)
if (mult < 0 && mult != SSIZE_T_MIN)
{
op1 = gtNewOperNode(GT_NEG, genActualType(op1), op1);
mul->gtOp1 = op1;
fgMorphTreeDone(op1);
}
// change the multiplication into a smaller multiplication (by 3, 5 or 9) and a shift
GenTree* const factorNode = gtNewIconNodeWithVN(this, factor, mul->TypeGet());
factorNode->SetMorphed(this);
op1 = gtNewOperNode(GT_MUL, mul->TypeGet(), op1, factorNode);
mul->gtOp1 = op1;
fgMorphTreeDone(op1);
op2->AsIntConCommon()->SetIconValue(shift);
changeToShift = true;
}
}

Perhaps this should be moved to Lower. LLVM also does it in "X86 DAG->DAG Instruction Selection" and not "InstCombine" https://godbolt.org/z/jsPea1PhP

Comment threadsrc/coreclr/jit/optimizebools.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated

auto isLeftDistributive = [](genTreeOps op1, genTreeOps op2) {
// op1 is left distributive over op2 iff:
// "A op1 (B op2 C)" <==> "(A op1 B) op2 (A op1 C)"

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We already do something similar for GT_ADD somewhere, we should unify it with that

@BoyBaykillerBoyBaykillerJun 4, 2026

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I didn't find what you are refering to here

Comment threadsrc/coreclr/jit/optimizebools.cpp Outdated
Comment threadsrc/coreclr/jit/gentree.cpp Outdated
@BoyBaykiller

Copy link
Copy Markdown
ContributorAuthor

I've limited it to OperIsAnyLocal, as discussed. Ready for an other review.

EgorBo pushed a commit that referenced this pull request Jun 4, 2026
Basically so we always have 0 at the right-side:
1. '(A & pow2) == pow2' -> '(A & pow2) != 0'
2. '(A & pow2) != pow2' -> '(A & pow2) == 0'
Tthis currently helps optimizeBools, but I am also adding this with the
future in mind.
Here is one example:
```cs
static bool Example(int foo)
{
return ((foo & 4) == 4) || ((foo & 8) == 8);
}
```
```assembly
;; ------ BASE
G_M000_IG02: ;; offset=0x0000
test cl, 4
je SHORT G_M000_IG05
G_M000_IG03: ;; offset=0x0005
mov eax, 1
G_M000_IG04: ;; offset=0x000A
ret G_M000_IG05: ;; offset=0x000B
test cl, 8
setne al
movzx rax, al
G_M000_IG06: ;; offset=0x0014
ret ;; ------ DIFF
G_M33938_IG02: ;; offset=0x0000
mov eax, ecx
and eax, 4
and ecx, 8
or eax, ecx
setne al
movzx rax, al
```
Note: With #126852 this would then
get transformed into just `test 12`.
@JulieLeeMSFT

Copy link
Copy Markdown
Member

@EgorBo, PTAL.

@EgorBo

Copy link
Copy Markdown
Member

@MihuBot -nuget

@EgorBo

Copy link
Copy Markdown
Member

@MihaZupan

Copy link
Copy Markdown
Member

Just testing error reporting changes

@MihuBot -nuget

@MihuBot

Copy link
Copy Markdown

Warning

2 extra test assemblies failed to produce JIT diffs and were excluded from the diff analysis:

  • MathNet.Numerics.dll (pr)
  • Pipelines.Sockets.Unofficial.dll (pr)

All failures happened only on the PR branch, which likely indicates a bug introduced by the PR (e.g. a JIT assert or crash while compiling these assemblies).

See job details in MihuBot/runtime-utils#2058, triggered by #126852 (comment).

eiriktsarpalis pushed a commit that referenced this pull request Jul 15, 2026
Basically so we always have 0 at the right-side:
1. '(A & pow2) == pow2' -> '(A & pow2) != 0'
2. '(A & pow2) != pow2' -> '(A & pow2) == 0'
Tthis currently helps optimizeBools, but I am also adding this with the
future in mind.
Here is one example:
```cs
static bool Example(int foo)
{
return ((foo & 4) == 4) || ((foo & 8) == 8);
}
```
```assembly
;; ------ BASE
G_M000_IG02: ;; offset=0x0000
test cl, 4
je SHORT G_M000_IG05
G_M000_IG03: ;; offset=0x0005
mov eax, 1
G_M000_IG04: ;; offset=0x000A
ret G_M000_IG05: ;; offset=0x000B
test cl, 8
setne al
movzx rax, al
G_M000_IG06: ;; offset=0x0014
ret ;; ------ DIFF
G_M33938_IG02: ;; offset=0x0000
mov eax, ecx
and eax, 4
and ecx, 8
or eax, ecx
setne al
movzx rax, al
```
Note: With #126852 this would then
get transformed into just `test 12`.
@github-actions

github-actionsBot commented Jul 19, 2026

Copy link
Copy Markdown
Contributor

Workflow state for the Holistic Review Orchestrator.

{
"version": 5,
"last_dispatched_commit": "ef82e04e0d0ec0a577f66e436b30a9adf3b47b77",
"last_dispatched_base_ref": "main",
"last_dispatched_base_sha": "f32edc0223b18865778e8e4d335e31bf9a2933f0",
"last_reviewed_commit": "ef82e04e0d0ec0a577f66e436b30a9adf3b47b77",
"last_reviewed_base_ref": "main",
"last_reviewed_base_sha": "f32edc0223b18865778e8e4d335e31bf9a2933f0",
"last_recorded_worker_run_id": "29679146097",
"review_attempt_commit": "",
"review_attempt_base_ref": "",
"review_attempt_count": 0,
"max_review_attempts": 5,
"review_history_format": "holistic-review-disclosure-v1",
"review_history": [
{
"commit": "ef82e04e0d0ec0a577f66e436b30a9adf3b47b77",
"review_id": 4730527763
}
]
}

@github-actionsgithub-actionsBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Holistic Review

Motivation: Generalizes #126070 and makes progress on #93954 by teaching the JIT to simplify arithmetic via the distributive property, i.e. (A op1 B) op2 (A op1 C) => A op1 (B op2 C). This reduces instruction count in common patterns (e.g. (A*B)+(A*C) folding to A*(B+C), and bit-test patterns after optOptimizeBools), which is a worthwhile codegen improvement.

Approach: Adds gtFoldDistributiveArithmetic invoked from gtFoldExprBinary for GT_AND/GT_OR/GT_XOR/GT_ADD/GT_SUB trees. It uses an isLeftDistributive table (AND/OR/MUL over their duals) and, when both operands share the same outer op and their left children are matching locals, rebuilds the tree with the common factor extracted. It also re-runs gtFoldExpr on the newly-combined boolean fold result in optOptimizeBoolsUpdateTrees, and tidies an unrelated cmp->gtOp2 -> cmp->gtGetOp2() reference in morph.cpp.

Summary: The transform is conceptually sound for wrapping (unchecked) integer arithmetic, and the guarding on optimization level, overflow of the outer node, side effects, and integral type is a reasonable start. However there is one correctness concern worth resolving before merge: the overflow check only examines the outer tree, not the inner operands. For checked(A*B) + checked(A*C) the inner GT_MUL nodes carry GTF_OVERFLOW, yet the rebuilt A*(B+C) uses a plain unchecked GT_MUL, dropping the overflow checks and altering observable behavior (a missing OverflowException). See the inline comment on the guard. I'd recommend adding a targeted regression test covering the checked-arithmetic case (and confirming it is rejected) once the guard is tightened. A secondary, non-blocking note: the freshly-created newOp2 node does not get value numbers assigned even though gtFoldExpr may be reached post-VN via optOptimizeBools; only result receives SetVNsFromNode. Verify this is benign in that phase (constants typically fold away, but a non-constant B op C combination would leave an unnumbered node).

Detailed Findings

  • Checked-arithmetic correctness (see inline comment on gtFoldDistributiveArithmetic): outer-only overflow guard can drop inner checked-multiply/add/sub semantics. Recommend also bailing when the operands being merged have gtOverflowEx()/GTF_EXCEPT set.
  • VN coverage of newOp2 (non-blocking): SetVNsFromNode(tree) is applied to result only; the intermediate newOp2 = gtFoldExpr(B op C) node has no VN. Since this path is also reachable from optOptimizeBoolsUpdateTrees after value numbering, confirm an unnumbered non-constant intermediate cannot reach a consumer that requires valid VNs.
  • Scope observation (non-blocking): the transform requires the shared factor to be OperIsAnyLocal. This is a safe, conservative choice (no duplicated side effects / re-evaluation), consistent with the PR description noting reassociation is not yet handled.

Note

This review was generated by this repository's Holistic Review agentic workflow to complement the built-in Copilot review.

Generated by Holistic Review · 115.6 AIC · ⌖ 11 AIC · ⊞ 10K

Comment threadsrc/coreclr/jit/gentree.cpp Outdated
{
assert(tree->OperIs(GT_AND, GT_OR, GT_XOR, GT_ADD, GT_SUB));

if (opts.OptimizationDisabled() || tree->gtOverflowEx() || !varTypeIsIntegralOrI(tree))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The overflow guard only inspects the outer tree, not the inner operands op1/op2. For an inner distributive op that can overflow-throw, this can silently drop checked semantics. Concretely, checked(A * B) + checked(A * C) has an unchecked outer GT_ADD but two inner GT_MUL nodes carrying GTF_OVERFLOW. isLeftDistributive(GT_MUL, GT_ADD) is true, so the transform rebuilds the tree as A * (B + C) via gtNewOperNode(GT_MUL, ...) with no overflow flag, discarding both checked multiplications. That changes observable behavior (a program that should throw OverflowException no longer does). For unchecked/wrapping arithmetic the distribution is correct mod 2^n, but the checked case is not. Consider also rejecting when op1/op2 (or their relevant children) have gtOverflowEx() set, or more conservatively bail when GTF_EXCEPT is present on the operands being merged.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@BoyBaykiller that is a good point, this PR basically breaks this:

staticintF(inta,intb,intc)=>checked(a*b)+checked(a*c);

With:

a = 2
b = 0x40000000
c = -0x40000000

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am now bailing on GTF_EXCEPT

Comment threadsrc/coreclr/jit/gentree.cpp Outdated
{
GenTree* newOp1 = op1->gtGetOp1();
GenTree* newOp2 =
gtFoldExpr(gtNewOperNode(tree->OperGet(), tree->TypeGet(), op1->gtGetOp2(), op2->gtGetOp2()));

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

seems like this gtNewOperNode is not marked as MORPHED

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

added fgMorphTreeDone(newOp2);

@BoyBaykiller

Copy link
Copy Markdown
ContributorAuthor

Failures fixed?

@MihuBot -nuget

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants

@BoyBaykiller@JulieLeeMSFT@EgorBo@MihaZupan@MihuBot@rhuijben@am11@jakobbotsch@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: Transform arithmetic using distributive property - #126852

Open
BoyBaykiller wants to merge 18 commits into
dotnet:mainfrom
BoyBaykiller:transform-using-distributive-property
Open

JIT: Transform arithmetic using distributive property#126852
BoyBaykiller wants to merge 18 commits into
dotnet:mainfrom
BoyBaykiller:transform-using-distributive-property

Conversation

@BoyBaykiller

@BoyBaykillerBoyBaykiller commented Apr 13, 2026

Copy link
Copy Markdown
Contributor

Generalization of #126070
Progress towards #93954

Basically we weren't doing any simplification based on distributive property before. So transforming
((A op1 B) op2 (A op1 C)) => (A op1 (B op2 C)). And this adds some basic support. Examples:

intMulDistedOverAdd(intA,intB,intC){return(A*B)+(A*C);}
;; ------ BASEG_M000_IG02:moveax,edximuleax,r8dimulr9d,edxaddeax,r9d;; ------ DIFFG_M48043_IG02: ;; offset=0x0000leaeax,[r8+r9]imuleax,edx
boolAfterOptimizeBools(intA,intB){return(A&4)!=0||(A&8)!=0;}
;; ------ BASEmoveax,edxandeax,4andedx,8oreax,edx setne almovzxrax,al;; ------ DIFFG_M55610_IG02:testdl,12 setne almovzxrax,al

We still need something that changes order to enable this opt. So (A | B) | C becoming A | (B | C) in this case:

uintReassociate(uintfoo,uintflags){return(foo|(flags&256))|(flags&512);}

Or here, the C * A need to be reversed.

intReassociate(intA,intB,intC){return(A*B)+(C*A);}

@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Apr 13, 2026
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Apr 13, 2026
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/jit/morph.cpp Outdated
return tree;
}

if (((tree->gtFlags & GTF_PERSISTENT_SIDE_EFFECTS) != 0) || ((tree->gtFlags & GTF_ORDER_SIDEEFF) != 0))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This could use the same optimization you are trying to apply😅... Probably handled by c++ though, so this is purely documentation ;-)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Lol. A new level of self documenting code. I didn't even notice that, I copied this check from elsewhere. I guess this just goes to show how useful the optimization is. Unfortunately even with this PR it's still not handled : (

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Update: This now gets handled by calling this opt again after the "optimizeBools" phase where it simplifies e.g
(A & 4) != 0 || (A & 8) != 0 into ((A & 4) | (A & 8)) != 0.

@BoyBaykiller

BoyBaykiller commented Apr 15, 2026

Copy link
Copy Markdown
ContributorAuthor

@EgorBo PTAL. Diffs are rather small, but nonetheless I believe it's a step in the right direction. I've written down some future work (which all improve diffs). I think the most interesting case for real-world code is arround bitwise ops.
Also should this go into post-order or pre-order morphing?

@BoyBaykiller
BoyBaykiller marked this pull request as ready for review April 15, 2026 01:07
@BoyBaykiller

Copy link
Copy Markdown
ContributorAuthor

Calling fgMorphBlockStmt after optimizeBools fixes point 2 in my list and produces much better diffs oldnew

@BoyBaykiller

BoyBaykiller commented Apr 15, 2026

Copy link
Copy Markdown
ContributorAuthor

Regarding point 1. The issue is that this runs first:

elseif (mulShiftOpt && (lowestBit > 1) && jitIsScaleIndexMul(lowestBit))
{
int shift = genLog2(lowestBit);
ssize_t factor = abs_mult >> shift;
if (factor == 3 || factor == 5 || factor == 9)
{
// if negative negate (min-int does not need negation)
if (mult < 0 && mult != SSIZE_T_MIN)
{
op1 = gtNewOperNode(GT_NEG, genActualType(op1), op1);
mul->gtOp1 = op1;
fgMorphTreeDone(op1);
}
// change the multiplication into a smaller multiplication (by 3, 5 or 9) and a shift
GenTree* const factorNode = gtNewIconNodeWithVN(this, factor, mul->TypeGet());
factorNode->SetMorphed(this);
op1 = gtNewOperNode(GT_MUL, mul->TypeGet(), op1, factorNode);
mul->gtOp1 = op1;
fgMorphTreeDone(op1);
op2->AsIntConCommon()->SetIconValue(shift);
changeToShift = true;
}
}

Perhaps this should be moved to Lower. LLVM also does it in "X86 DAG->DAG Instruction Selection" and not "InstCombine" https://godbolt.org/z/jsPea1PhP

Comment threadsrc/coreclr/jit/optimizebools.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated

auto isLeftDistributive = [](genTreeOps op1, genTreeOps op2) {
// op1 is left distributive over op2 iff:
// "A op1 (B op2 C)" <==> "(A op1 B) op2 (A op1 C)"

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We already do something similar for GT_ADD somewhere, we should unify it with that

@BoyBaykillerBoyBaykillerJun 4, 2026

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I didn't find what you are refering to here

Comment threadsrc/coreclr/jit/optimizebools.cpp Outdated
Comment threadsrc/coreclr/jit/gentree.cpp Outdated
@BoyBaykiller

Copy link
Copy Markdown
ContributorAuthor

I've limited it to OperIsAnyLocal, as discussed. Ready for an other review.

EgorBo pushed a commit that referenced this pull request Jun 4, 2026
Basically so we always have 0 at the right-side:
1. '(A & pow2) == pow2' -> '(A & pow2) != 0'
2. '(A & pow2) != pow2' -> '(A & pow2) == 0'
Tthis currently helps optimizeBools, but I am also adding this with the
future in mind.
Here is one example:
```cs
static bool Example(int foo)
{
return ((foo & 4) == 4) || ((foo & 8) == 8);
}
```
```assembly
;; ------ BASE
G_M000_IG02: ;; offset=0x0000
test cl, 4
je SHORT G_M000_IG05
G_M000_IG03: ;; offset=0x0005
mov eax, 1
G_M000_IG04: ;; offset=0x000A
ret G_M000_IG05: ;; offset=0x000B
test cl, 8
setne al
movzx rax, al
G_M000_IG06: ;; offset=0x0014
ret ;; ------ DIFF
G_M33938_IG02: ;; offset=0x0000
mov eax, ecx
and eax, 4
and ecx, 8
or eax, ecx
setne al
movzx rax, al
```
Note: With #126852 this would then
get transformed into just `test 12`.
@JulieLeeMSFT

Copy link
Copy Markdown
Member

@EgorBo, PTAL.

@EgorBo

Copy link
Copy Markdown
Member

@MihuBot -nuget

@EgorBo

Copy link
Copy Markdown
Member

@MihaZupan

Copy link
Copy Markdown
Member

Just testing error reporting changes

@MihuBot -nuget

@MihuBot

Copy link
Copy Markdown

Warning

2 extra test assemblies failed to produce JIT diffs and were excluded from the diff analysis:

  • MathNet.Numerics.dll (pr)
  • Pipelines.Sockets.Unofficial.dll (pr)

All failures happened only on the PR branch, which likely indicates a bug introduced by the PR (e.g. a JIT assert or crash while compiling these assemblies).

See job details in MihuBot/runtime-utils#2058, triggered by #126852 (comment).

eiriktsarpalis pushed a commit that referenced this pull request Jul 15, 2026
Basically so we always have 0 at the right-side:
1. '(A & pow2) == pow2' -> '(A & pow2) != 0'
2. '(A & pow2) != pow2' -> '(A & pow2) == 0'
Tthis currently helps optimizeBools, but I am also adding this with the
future in mind.
Here is one example:
```cs
static bool Example(int foo)
{
return ((foo & 4) == 4) || ((foo & 8) == 8);
}
```
```assembly
;; ------ BASE
G_M000_IG02: ;; offset=0x0000
test cl, 4
je SHORT G_M000_IG05
G_M000_IG03: ;; offset=0x0005
mov eax, 1
G_M000_IG04: ;; offset=0x000A
ret G_M000_IG05: ;; offset=0x000B
test cl, 8
setne al
movzx rax, al
G_M000_IG06: ;; offset=0x0014
ret ;; ------ DIFF
G_M33938_IG02: ;; offset=0x0000
mov eax, ecx
and eax, 4
and ecx, 8
or eax, ecx
setne al
movzx rax, al
```
Note: With #126852 this would then
get transformed into just `test 12`.
@github-actions

github-actionsBot commented Jul 19, 2026

Copy link
Copy Markdown
Contributor

Workflow state for the Holistic Review Orchestrator.

{
"version": 5,
"last_dispatched_commit": "ef82e04e0d0ec0a577f66e436b30a9adf3b47b77",
"last_dispatched_base_ref": "main",
"last_dispatched_base_sha": "f32edc0223b18865778e8e4d335e31bf9a2933f0",
"last_reviewed_commit": "ef82e04e0d0ec0a577f66e436b30a9adf3b47b77",
"last_reviewed_base_ref": "main",
"last_reviewed_base_sha": "f32edc0223b18865778e8e4d335e31bf9a2933f0",
"last_recorded_worker_run_id": "29679146097",
"review_attempt_commit": "",
"review_attempt_base_ref": "",
"review_attempt_count": 0,
"max_review_attempts": 5,
"review_history_format": "holistic-review-disclosure-v1",
"review_history": [
{
"commit": "ef82e04e0d0ec0a577f66e436b30a9adf3b47b77",
"review_id": 4730527763
}
]
}

@github-actionsgithub-actionsBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Holistic Review

Motivation: Generalizes #126070 and makes progress on #93954 by teaching the JIT to simplify arithmetic via the distributive property, i.e. (A op1 B) op2 (A op1 C) => A op1 (B op2 C). This reduces instruction count in common patterns (e.g. (A*B)+(A*C) folding to A*(B+C), and bit-test patterns after optOptimizeBools), which is a worthwhile codegen improvement.

Approach: Adds gtFoldDistributiveArithmetic invoked from gtFoldExprBinary for GT_AND/GT_OR/GT_XOR/GT_ADD/GT_SUB trees. It uses an isLeftDistributive table (AND/OR/MUL over their duals) and, when both operands share the same outer op and their left children are matching locals, rebuilds the tree with the common factor extracted. It also re-runs gtFoldExpr on the newly-combined boolean fold result in optOptimizeBoolsUpdateTrees, and tidies an unrelated cmp->gtOp2 -> cmp->gtGetOp2() reference in morph.cpp.

Summary: The transform is conceptually sound for wrapping (unchecked) integer arithmetic, and the guarding on optimization level, overflow of the outer node, side effects, and integral type is a reasonable start. However there is one correctness concern worth resolving before merge: the overflow check only examines the outer tree, not the inner operands. For checked(A*B) + checked(A*C) the inner GT_MUL nodes carry GTF_OVERFLOW, yet the rebuilt A*(B+C) uses a plain unchecked GT_MUL, dropping the overflow checks and altering observable behavior (a missing OverflowException). See the inline comment on the guard. I'd recommend adding a targeted regression test covering the checked-arithmetic case (and confirming it is rejected) once the guard is tightened. A secondary, non-blocking note: the freshly-created newOp2 node does not get value numbers assigned even though gtFoldExpr may be reached post-VN via optOptimizeBools; only result receives SetVNsFromNode. Verify this is benign in that phase (constants typically fold away, but a non-constant B op C combination would leave an unnumbered node).

Detailed Findings

  • Checked-arithmetic correctness (see inline comment on gtFoldDistributiveArithmetic): outer-only overflow guard can drop inner checked-multiply/add/sub semantics. Recommend also bailing when the operands being merged have gtOverflowEx()/GTF_EXCEPT set.
  • VN coverage of newOp2 (non-blocking): SetVNsFromNode(tree) is applied to result only; the intermediate newOp2 = gtFoldExpr(B op C) node has no VN. Since this path is also reachable from optOptimizeBoolsUpdateTrees after value numbering, confirm an unnumbered non-constant intermediate cannot reach a consumer that requires valid VNs.
  • Scope observation (non-blocking): the transform requires the shared factor to be OperIsAnyLocal. This is a safe, conservative choice (no duplicated side effects / re-evaluation), consistent with the PR description noting reassociation is not yet handled.

Note

This review was generated by this repository's Holistic Review agentic workflow to complement the built-in Copilot review.

Generated by Holistic Review · 115.6 AIC · ⌖ 11 AIC · ⊞ 10K

Comment threadsrc/coreclr/jit/gentree.cpp Outdated
{
assert(tree->OperIs(GT_AND, GT_OR, GT_XOR, GT_ADD, GT_SUB));

if (opts.OptimizationDisabled() || tree->gtOverflowEx() || !varTypeIsIntegralOrI(tree))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The overflow guard only inspects the outer tree, not the inner operands op1/op2. For an inner distributive op that can overflow-throw, this can silently drop checked semantics. Concretely, checked(A * B) + checked(A * C) has an unchecked outer GT_ADD but two inner GT_MUL nodes carrying GTF_OVERFLOW. isLeftDistributive(GT_MUL, GT_ADD) is true, so the transform rebuilds the tree as A * (B + C) via gtNewOperNode(GT_MUL, ...) with no overflow flag, discarding both checked multiplications. That changes observable behavior (a program that should throw OverflowException no longer does). For unchecked/wrapping arithmetic the distribution is correct mod 2^n, but the checked case is not. Consider also rejecting when op1/op2 (or their relevant children) have gtOverflowEx() set, or more conservatively bail when GTF_EXCEPT is present on the operands being merged.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@BoyBaykiller that is a good point, this PR basically breaks this:

staticintF(inta,intb,intc)=>checked(a*b)+checked(a*c);

With:

a = 2
b = 0x40000000
c = -0x40000000

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am now bailing on GTF_EXCEPT

Comment threadsrc/coreclr/jit/gentree.cpp Outdated
{
GenTree* newOp1 = op1->gtGetOp1();
GenTree* newOp2 =
gtFoldExpr(gtNewOperNode(tree->OperGet(), tree->TypeGet(), op1->gtGetOp2(), op2->gtGetOp2()));

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

seems like this gtNewOperNode is not marked as MORPHED

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

added fgMorphTreeDone(newOp2);

@BoyBaykiller

Copy link
Copy Markdown
ContributorAuthor

Failures fixed?

@MihuBot -nuget

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants

@BoyBaykiller@JulieLeeMSFT@EgorBo@MihaZupan@MihuBot@rhuijben@am11@jakobbotsch@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

JIT: Transform arithmetic using distributive property - #126852

Open
BoyBaykiller wants to merge 18 commits into
dotnet:mainfrom
BoyBaykiller:transform-using-distributive-property
Open

JIT: Transform arithmetic using distributive property#126852
BoyBaykiller wants to merge 18 commits into
dotnet:mainfrom
BoyBaykiller:transform-using-distributive-property

Conversation

@BoyBaykiller

@BoyBaykillerBoyBaykiller commented Apr 13, 2026

Copy link
Copy Markdown
Contributor

Generalization of #126070
Progress towards #93954

Basically we weren't doing any simplification based on distributive property before. So transforming
((A op1 B) op2 (A op1 C)) => (A op1 (B op2 C)). And this adds some basic support. Examples:

intMulDistedOverAdd(intA,intB,intC){return(A*B)+(A*C);}
;; ------ BASEG_M000_IG02:moveax,edximuleax,r8dimulr9d,edxaddeax,r9d;; ------ DIFFG_M48043_IG02: ;; offset=0x0000leaeax,[r8+r9]imuleax,edx
boolAfterOptimizeBools(intA,intB){return(A&4)!=0||(A&8)!=0;}
;; ------ BASEmoveax,edxandeax,4andedx,8oreax,edx setne almovzxrax,al;; ------ DIFFG_M55610_IG02:testdl,12 setne almovzxrax,al

We still need something that changes order to enable this opt. So (A | B) | C becoming A | (B | C) in this case:

uintReassociate(uintfoo,uintflags){return(foo|(flags&256))|(flags&512);}

Or here, the C * A need to be reversed.

intReassociate(intA,intB,intC){return(A*B)+(C*A);}

@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Apr 13, 2026
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Apr 13, 2026
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/jit/morph.cpp Outdated
return tree;
}

if (((tree->gtFlags & GTF_PERSISTENT_SIDE_EFFECTS) != 0) || ((tree->gtFlags & GTF_ORDER_SIDEEFF) != 0))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This could use the same optimization you are trying to apply😅... Probably handled by c++ though, so this is purely documentation ;-)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Lol. A new level of self documenting code. I didn't even notice that, I copied this check from elsewhere. I guess this just goes to show how useful the optimization is. Unfortunately even with this PR it's still not handled : (

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Update: This now gets handled by calling this opt again after the "optimizeBools" phase where it simplifies e.g
(A & 4) != 0 || (A & 8) != 0 into ((A & 4) | (A & 8)) != 0.

@BoyBaykiller

BoyBaykiller commented Apr 15, 2026

Copy link
Copy Markdown
ContributorAuthor

@EgorBo PTAL. Diffs are rather small, but nonetheless I believe it's a step in the right direction. I've written down some future work (which all improve diffs). I think the most interesting case for real-world code is arround bitwise ops.
Also should this go into post-order or pre-order morphing?

@BoyBaykiller
BoyBaykiller marked this pull request as ready for review April 15, 2026 01:07
@BoyBaykiller

Copy link
Copy Markdown
ContributorAuthor

Calling fgMorphBlockStmt after optimizeBools fixes point 2 in my list and produces much better diffs oldnew

@BoyBaykiller

BoyBaykiller commented Apr 15, 2026

Copy link
Copy Markdown
ContributorAuthor

Regarding point 1. The issue is that this runs first:

elseif (mulShiftOpt && (lowestBit > 1) && jitIsScaleIndexMul(lowestBit))
{
int shift = genLog2(lowestBit);
ssize_t factor = abs_mult >> shift;
if (factor == 3 || factor == 5 || factor == 9)
{
// if negative negate (min-int does not need negation)
if (mult < 0 && mult != SSIZE_T_MIN)
{
op1 = gtNewOperNode(GT_NEG, genActualType(op1), op1);
mul->gtOp1 = op1;
fgMorphTreeDone(op1);
}
// change the multiplication into a smaller multiplication (by 3, 5 or 9) and a shift
GenTree* const factorNode = gtNewIconNodeWithVN(this, factor, mul->TypeGet());
factorNode->SetMorphed(this);
op1 = gtNewOperNode(GT_MUL, mul->TypeGet(), op1, factorNode);
mul->gtOp1 = op1;
fgMorphTreeDone(op1);
op2->AsIntConCommon()->SetIconValue(shift);
changeToShift = true;
}
}

Perhaps this should be moved to Lower. LLVM also does it in "X86 DAG->DAG Instruction Selection" and not "InstCombine" https://godbolt.org/z/jsPea1PhP

Comment threadsrc/coreclr/jit/optimizebools.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated

auto isLeftDistributive = [](genTreeOps op1, genTreeOps op2) {
// op1 is left distributive over op2 iff:
// "A op1 (B op2 C)" <==> "(A op1 B) op2 (A op1 C)"

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We already do something similar for GT_ADD somewhere, we should unify it with that

@BoyBaykillerBoyBaykillerJun 4, 2026

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I didn't find what you are refering to here

Comment threadsrc/coreclr/jit/optimizebools.cpp Outdated
Comment threadsrc/coreclr/jit/gentree.cpp Outdated
@BoyBaykiller

Copy link
Copy Markdown
ContributorAuthor

I've limited it to OperIsAnyLocal, as discussed. Ready for an other review.

EgorBo pushed a commit that referenced this pull request Jun 4, 2026
Basically so we always have 0 at the right-side:
1. '(A & pow2) == pow2' -> '(A & pow2) != 0'
2. '(A & pow2) != pow2' -> '(A & pow2) == 0'
Tthis currently helps optimizeBools, but I am also adding this with the
future in mind.
Here is one example:
```cs
static bool Example(int foo)
{
return ((foo & 4) == 4) || ((foo & 8) == 8);
}
```
```assembly
;; ------ BASE
G_M000_IG02: ;; offset=0x0000
test cl, 4
je SHORT G_M000_IG05
G_M000_IG03: ;; offset=0x0005
mov eax, 1
G_M000_IG04: ;; offset=0x000A
ret G_M000_IG05: ;; offset=0x000B
test cl, 8
setne al
movzx rax, al
G_M000_IG06: ;; offset=0x0014
ret ;; ------ DIFF
G_M33938_IG02: ;; offset=0x0000
mov eax, ecx
and eax, 4
and ecx, 8
or eax, ecx
setne al
movzx rax, al
```
Note: With #126852 this would then
get transformed into just `test 12`.
@JulieLeeMSFT

Copy link
Copy Markdown
Member

@EgorBo, PTAL.

@EgorBo

Copy link
Copy Markdown
Member

@MihuBot -nuget

@EgorBo

Copy link
Copy Markdown
Member

@MihaZupan

Copy link
Copy Markdown
Member

Just testing error reporting changes

@MihuBot -nuget

@MihuBot

Copy link
Copy Markdown

Warning

2 extra test assemblies failed to produce JIT diffs and were excluded from the diff analysis:

  • MathNet.Numerics.dll (pr)
  • Pipelines.Sockets.Unofficial.dll (pr)

All failures happened only on the PR branch, which likely indicates a bug introduced by the PR (e.g. a JIT assert or crash while compiling these assemblies).

See job details in MihuBot/runtime-utils#2058, triggered by #126852 (comment).

eiriktsarpalis pushed a commit that referenced this pull request Jul 15, 2026
Basically so we always have 0 at the right-side:
1. '(A & pow2) == pow2' -> '(A & pow2) != 0'
2. '(A & pow2) != pow2' -> '(A & pow2) == 0'
Tthis currently helps optimizeBools, but I am also adding this with the
future in mind.
Here is one example:
```cs
static bool Example(int foo)
{
return ((foo & 4) == 4) || ((foo & 8) == 8);
}
```
```assembly
;; ------ BASE
G_M000_IG02: ;; offset=0x0000
test cl, 4
je SHORT G_M000_IG05
G_M000_IG03: ;; offset=0x0005
mov eax, 1
G_M000_IG04: ;; offset=0x000A
ret G_M000_IG05: ;; offset=0x000B
test cl, 8
setne al
movzx rax, al
G_M000_IG06: ;; offset=0x0014
ret ;; ------ DIFF
G_M33938_IG02: ;; offset=0x0000
mov eax, ecx
and eax, 4
and ecx, 8
or eax, ecx
setne al
movzx rax, al
```
Note: With #126852 this would then
get transformed into just `test 12`.
@github-actions

github-actionsBot commented Jul 19, 2026

Copy link
Copy Markdown
Contributor

Workflow state for the Holistic Review Orchestrator.

{
"version": 5,
"last_dispatched_commit": "ef82e04e0d0ec0a577f66e436b30a9adf3b47b77",
"last_dispatched_base_ref": "main",
"last_dispatched_base_sha": "f32edc0223b18865778e8e4d335e31bf9a2933f0",
"last_reviewed_commit": "ef82e04e0d0ec0a577f66e436b30a9adf3b47b77",
"last_reviewed_base_ref": "main",
"last_reviewed_base_sha": "f32edc0223b18865778e8e4d335e31bf9a2933f0",
"last_recorded_worker_run_id": "29679146097",
"review_attempt_commit": "",
"review_attempt_base_ref": "",
"review_attempt_count": 0,
"max_review_attempts": 5,
"review_history_format": "holistic-review-disclosure-v1",
"review_history": [
{
"commit": "ef82e04e0d0ec0a577f66e436b30a9adf3b47b77",
"review_id": 4730527763
}
]
}

@github-actionsgithub-actionsBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Holistic Review

Motivation: Generalizes #126070 and makes progress on #93954 by teaching the JIT to simplify arithmetic via the distributive property, i.e. (A op1 B) op2 (A op1 C) => A op1 (B op2 C). This reduces instruction count in common patterns (e.g. (A*B)+(A*C) folding to A*(B+C), and bit-test patterns after optOptimizeBools), which is a worthwhile codegen improvement.

Approach: Adds gtFoldDistributiveArithmetic invoked from gtFoldExprBinary for GT_AND/GT_OR/GT_XOR/GT_ADD/GT_SUB trees. It uses an isLeftDistributive table (AND/OR/MUL over their duals) and, when both operands share the same outer op and their left children are matching locals, rebuilds the tree with the common factor extracted. It also re-runs gtFoldExpr on the newly-combined boolean fold result in optOptimizeBoolsUpdateTrees, and tidies an unrelated cmp->gtOp2 -> cmp->gtGetOp2() reference in morph.cpp.

Summary: The transform is conceptually sound for wrapping (unchecked) integer arithmetic, and the guarding on optimization level, overflow of the outer node, side effects, and integral type is a reasonable start. However there is one correctness concern worth resolving before merge: the overflow check only examines the outer tree, not the inner operands. For checked(A*B) + checked(A*C) the inner GT_MUL nodes carry GTF_OVERFLOW, yet the rebuilt A*(B+C) uses a plain unchecked GT_MUL, dropping the overflow checks and altering observable behavior (a missing OverflowException). See the inline comment on the guard. I'd recommend adding a targeted regression test covering the checked-arithmetic case (and confirming it is rejected) once the guard is tightened. A secondary, non-blocking note: the freshly-created newOp2 node does not get value numbers assigned even though gtFoldExpr may be reached post-VN via optOptimizeBools; only result receives SetVNsFromNode. Verify this is benign in that phase (constants typically fold away, but a non-constant B op C combination would leave an unnumbered node).

Detailed Findings

  • Checked-arithmetic correctness (see inline comment on gtFoldDistributiveArithmetic): outer-only overflow guard can drop inner checked-multiply/add/sub semantics. Recommend also bailing when the operands being merged have gtOverflowEx()/GTF_EXCEPT set.
  • VN coverage of newOp2 (non-blocking): SetVNsFromNode(tree) is applied to result only; the intermediate newOp2 = gtFoldExpr(B op C) node has no VN. Since this path is also reachable from optOptimizeBoolsUpdateTrees after value numbering, confirm an unnumbered non-constant intermediate cannot reach a consumer that requires valid VNs.
  • Scope observation (non-blocking): the transform requires the shared factor to be OperIsAnyLocal. This is a safe, conservative choice (no duplicated side effects / re-evaluation), consistent with the PR description noting reassociation is not yet handled.

Note

This review was generated by this repository's Holistic Review agentic workflow to complement the built-in Copilot review.

Generated by Holistic Review · 115.6 AIC · ⌖ 11 AIC · ⊞ 10K

Comment threadsrc/coreclr/jit/gentree.cpp Outdated
{
assert(tree->OperIs(GT_AND, GT_OR, GT_XOR, GT_ADD, GT_SUB));

if (opts.OptimizationDisabled() || tree->gtOverflowEx() || !varTypeIsIntegralOrI(tree))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The overflow guard only inspects the outer tree, not the inner operands op1/op2. For an inner distributive op that can overflow-throw, this can silently drop checked semantics. Concretely, checked(A * B) + checked(A * C) has an unchecked outer GT_ADD but two inner GT_MUL nodes carrying GTF_OVERFLOW. isLeftDistributive(GT_MUL, GT_ADD) is true, so the transform rebuilds the tree as A * (B + C) via gtNewOperNode(GT_MUL, ...) with no overflow flag, discarding both checked multiplications. That changes observable behavior (a program that should throw OverflowException no longer does). For unchecked/wrapping arithmetic the distribution is correct mod 2^n, but the checked case is not. Consider also rejecting when op1/op2 (or their relevant children) have gtOverflowEx() set, or more conservatively bail when GTF_EXCEPT is present on the operands being merged.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@BoyBaykiller that is a good point, this PR basically breaks this:

staticintF(inta,intb,intc)=>checked(a*b)+checked(a*c);

With:

a = 2
b = 0x40000000
c = -0x40000000

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am now bailing on GTF_EXCEPT

Comment threadsrc/coreclr/jit/gentree.cpp Outdated
{
GenTree* newOp1 = op1->gtGetOp1();
GenTree* newOp2 =
gtFoldExpr(gtNewOperNode(tree->OperGet(), tree->TypeGet(), op1->gtGetOp2(), op2->gtGetOp2()));

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

seems like this gtNewOperNode is not marked as MORPHED

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

added fgMorphTreeDone(newOp2);

@BoyBaykiller

Copy link
Copy Markdown
ContributorAuthor

Failures fixed?

@MihuBot -nuget

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants

@BoyBaykiller@JulieLeeMSFT@EgorBo@MihaZupan@MihuBot@rhuijben@am11@jakobbotsch@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: Transform arithmetic using distributive property - #126852

Open
BoyBaykiller wants to merge 18 commits into
dotnet:mainfrom
BoyBaykiller:transform-using-distributive-property
Open

JIT: Transform arithmetic using distributive property#126852
BoyBaykiller wants to merge 18 commits into
dotnet:mainfrom
BoyBaykiller:transform-using-distributive-property

Conversation

@BoyBaykiller

@BoyBaykillerBoyBaykiller commented Apr 13, 2026

Copy link
Copy Markdown
Contributor

Generalization of #126070
Progress towards #93954

Basically we weren't doing any simplification based on distributive property before. So transforming
((A op1 B) op2 (A op1 C)) => (A op1 (B op2 C)). And this adds some basic support. Examples:

intMulDistedOverAdd(intA,intB,intC){return(A*B)+(A*C);}
;; ------ BASEG_M000_IG02:moveax,edximuleax,r8dimulr9d,edxaddeax,r9d;; ------ DIFFG_M48043_IG02: ;; offset=0x0000leaeax,[r8+r9]imuleax,edx
boolAfterOptimizeBools(intA,intB){return(A&4)!=0||(A&8)!=0;}
;; ------ BASEmoveax,edxandeax,4andedx,8oreax,edx setne almovzxrax,al;; ------ DIFFG_M55610_IG02:testdl,12 setne almovzxrax,al

We still need something that changes order to enable this opt. So (A | B) | C becoming A | (B | C) in this case:

uintReassociate(uintfoo,uintflags){return(foo|(flags&256))|(flags&512);}

Or here, the C * A need to be reversed.

intReassociate(intA,intB,intC){return(A*B)+(C*A);}

@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Apr 13, 2026
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Apr 13, 2026
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/jit/morph.cpp Outdated
return tree;
}

if (((tree->gtFlags & GTF_PERSISTENT_SIDE_EFFECTS) != 0) || ((tree->gtFlags & GTF_ORDER_SIDEEFF) != 0))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This could use the same optimization you are trying to apply😅... Probably handled by c++ though, so this is purely documentation ;-)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Lol. A new level of self documenting code. I didn't even notice that, I copied this check from elsewhere. I guess this just goes to show how useful the optimization is. Unfortunately even with this PR it's still not handled : (

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Update: This now gets handled by calling this opt again after the "optimizeBools" phase where it simplifies e.g
(A & 4) != 0 || (A & 8) != 0 into ((A & 4) | (A & 8)) != 0.

@BoyBaykiller

BoyBaykiller commented Apr 15, 2026

Copy link
Copy Markdown
ContributorAuthor

@EgorBo PTAL. Diffs are rather small, but nonetheless I believe it's a step in the right direction. I've written down some future work (which all improve diffs). I think the most interesting case for real-world code is arround bitwise ops.
Also should this go into post-order or pre-order morphing?

@BoyBaykiller
BoyBaykiller marked this pull request as ready for review April 15, 2026 01:07
@BoyBaykiller

Copy link
Copy Markdown
ContributorAuthor

Calling fgMorphBlockStmt after optimizeBools fixes point 2 in my list and produces much better diffs oldnew

@BoyBaykiller

BoyBaykiller commented Apr 15, 2026

Copy link
Copy Markdown
ContributorAuthor

Regarding point 1. The issue is that this runs first:

elseif (mulShiftOpt && (lowestBit > 1) && jitIsScaleIndexMul(lowestBit))
{
int shift = genLog2(lowestBit);
ssize_t factor = abs_mult >> shift;
if (factor == 3 || factor == 5 || factor == 9)
{
// if negative negate (min-int does not need negation)
if (mult < 0 && mult != SSIZE_T_MIN)
{
op1 = gtNewOperNode(GT_NEG, genActualType(op1), op1);
mul->gtOp1 = op1;
fgMorphTreeDone(op1);
}
// change the multiplication into a smaller multiplication (by 3, 5 or 9) and a shift
GenTree* const factorNode = gtNewIconNodeWithVN(this, factor, mul->TypeGet());
factorNode->SetMorphed(this);
op1 = gtNewOperNode(GT_MUL, mul->TypeGet(), op1, factorNode);
mul->gtOp1 = op1;
fgMorphTreeDone(op1);
op2->AsIntConCommon()->SetIconValue(shift);
changeToShift = true;
}
}

Perhaps this should be moved to Lower. LLVM also does it in "X86 DAG->DAG Instruction Selection" and not "InstCombine" https://godbolt.org/z/jsPea1PhP

Comment threadsrc/coreclr/jit/optimizebools.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated

auto isLeftDistributive = [](genTreeOps op1, genTreeOps op2) {
// op1 is left distributive over op2 iff:
// "A op1 (B op2 C)" <==> "(A op1 B) op2 (A op1 C)"

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We already do something similar for GT_ADD somewhere, we should unify it with that

@BoyBaykillerBoyBaykillerJun 4, 2026

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I didn't find what you are refering to here

Comment threadsrc/coreclr/jit/optimizebools.cpp Outdated
Comment threadsrc/coreclr/jit/gentree.cpp Outdated
@BoyBaykiller

Copy link
Copy Markdown
ContributorAuthor

I've limited it to OperIsAnyLocal, as discussed. Ready for an other review.

EgorBo pushed a commit that referenced this pull request Jun 4, 2026
Basically so we always have 0 at the right-side:
1. '(A & pow2) == pow2' -> '(A & pow2) != 0'
2. '(A & pow2) != pow2' -> '(A & pow2) == 0'
Tthis currently helps optimizeBools, but I am also adding this with the
future in mind.
Here is one example:
```cs
static bool Example(int foo)
{
return ((foo & 4) == 4) || ((foo & 8) == 8);
}
```
```assembly
;; ------ BASE
G_M000_IG02: ;; offset=0x0000
test cl, 4
je SHORT G_M000_IG05
G_M000_IG03: ;; offset=0x0005
mov eax, 1
G_M000_IG04: ;; offset=0x000A
ret G_M000_IG05: ;; offset=0x000B
test cl, 8
setne al
movzx rax, al
G_M000_IG06: ;; offset=0x0014
ret ;; ------ DIFF
G_M33938_IG02: ;; offset=0x0000
mov eax, ecx
and eax, 4
and ecx, 8
or eax, ecx
setne al
movzx rax, al
```
Note: With #126852 this would then
get transformed into just `test 12`.
@JulieLeeMSFT

Copy link
Copy Markdown
Member

@EgorBo, PTAL.

@EgorBo

Copy link
Copy Markdown
Member

@MihuBot -nuget

@EgorBo

Copy link
Copy Markdown
Member

@MihaZupan

Copy link
Copy Markdown
Member

Just testing error reporting changes

@MihuBot -nuget

@MihuBot

Copy link
Copy Markdown

Warning

2 extra test assemblies failed to produce JIT diffs and were excluded from the diff analysis:

  • MathNet.Numerics.dll (pr)
  • Pipelines.Sockets.Unofficial.dll (pr)

All failures happened only on the PR branch, which likely indicates a bug introduced by the PR (e.g. a JIT assert or crash while compiling these assemblies).

See job details in MihuBot/runtime-utils#2058, triggered by #126852 (comment).

eiriktsarpalis pushed a commit that referenced this pull request Jul 15, 2026
Basically so we always have 0 at the right-side:
1. '(A & pow2) == pow2' -> '(A & pow2) != 0'
2. '(A & pow2) != pow2' -> '(A & pow2) == 0'
Tthis currently helps optimizeBools, but I am also adding this with the
future in mind.
Here is one example:
```cs
static bool Example(int foo)
{
return ((foo & 4) == 4) || ((foo & 8) == 8);
}
```
```assembly
;; ------ BASE
G_M000_IG02: ;; offset=0x0000
test cl, 4
je SHORT G_M000_IG05
G_M000_IG03: ;; offset=0x0005
mov eax, 1
G_M000_IG04: ;; offset=0x000A
ret G_M000_IG05: ;; offset=0x000B
test cl, 8
setne al
movzx rax, al
G_M000_IG06: ;; offset=0x0014
ret ;; ------ DIFF
G_M33938_IG02: ;; offset=0x0000
mov eax, ecx
and eax, 4
and ecx, 8
or eax, ecx
setne al
movzx rax, al
```
Note: With #126852 this would then
get transformed into just `test 12`.
@github-actions

github-actionsBot commented Jul 19, 2026

Copy link
Copy Markdown
Contributor

Workflow state for the Holistic Review Orchestrator.

{
"version": 5,
"last_dispatched_commit": "ef82e04e0d0ec0a577f66e436b30a9adf3b47b77",
"last_dispatched_base_ref": "main",
"last_dispatched_base_sha": "f32edc0223b18865778e8e4d335e31bf9a2933f0",
"last_reviewed_commit": "ef82e04e0d0ec0a577f66e436b30a9adf3b47b77",
"last_reviewed_base_ref": "main",
"last_reviewed_base_sha": "f32edc0223b18865778e8e4d335e31bf9a2933f0",
"last_recorded_worker_run_id": "29679146097",
"review_attempt_commit": "",
"review_attempt_base_ref": "",
"review_attempt_count": 0,
"max_review_attempts": 5,
"review_history_format": "holistic-review-disclosure-v1",
"review_history": [
{
"commit": "ef82e04e0d0ec0a577f66e436b30a9adf3b47b77",
"review_id": 4730527763
}
]
}

@github-actionsgithub-actionsBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Holistic Review

Motivation: Generalizes #126070 and makes progress on #93954 by teaching the JIT to simplify arithmetic via the distributive property, i.e. (A op1 B) op2 (A op1 C) => A op1 (B op2 C). This reduces instruction count in common patterns (e.g. (A*B)+(A*C) folding to A*(B+C), and bit-test patterns after optOptimizeBools), which is a worthwhile codegen improvement.

Approach: Adds gtFoldDistributiveArithmetic invoked from gtFoldExprBinary for GT_AND/GT_OR/GT_XOR/GT_ADD/GT_SUB trees. It uses an isLeftDistributive table (AND/OR/MUL over their duals) and, when both operands share the same outer op and their left children are matching locals, rebuilds the tree with the common factor extracted. It also re-runs gtFoldExpr on the newly-combined boolean fold result in optOptimizeBoolsUpdateTrees, and tidies an unrelated cmp->gtOp2 -> cmp->gtGetOp2() reference in morph.cpp.

Summary: The transform is conceptually sound for wrapping (unchecked) integer arithmetic, and the guarding on optimization level, overflow of the outer node, side effects, and integral type is a reasonable start. However there is one correctness concern worth resolving before merge: the overflow check only examines the outer tree, not the inner operands. For checked(A*B) + checked(A*C) the inner GT_MUL nodes carry GTF_OVERFLOW, yet the rebuilt A*(B+C) uses a plain unchecked GT_MUL, dropping the overflow checks and altering observable behavior (a missing OverflowException). See the inline comment on the guard. I'd recommend adding a targeted regression test covering the checked-arithmetic case (and confirming it is rejected) once the guard is tightened. A secondary, non-blocking note: the freshly-created newOp2 node does not get value numbers assigned even though gtFoldExpr may be reached post-VN via optOptimizeBools; only result receives SetVNsFromNode. Verify this is benign in that phase (constants typically fold away, but a non-constant B op C combination would leave an unnumbered node).

Detailed Findings

  • Checked-arithmetic correctness (see inline comment on gtFoldDistributiveArithmetic): outer-only overflow guard can drop inner checked-multiply/add/sub semantics. Recommend also bailing when the operands being merged have gtOverflowEx()/GTF_EXCEPT set.
  • VN coverage of newOp2 (non-blocking): SetVNsFromNode(tree) is applied to result only; the intermediate newOp2 = gtFoldExpr(B op C) node has no VN. Since this path is also reachable from optOptimizeBoolsUpdateTrees after value numbering, confirm an unnumbered non-constant intermediate cannot reach a consumer that requires valid VNs.
  • Scope observation (non-blocking): the transform requires the shared factor to be OperIsAnyLocal. This is a safe, conservative choice (no duplicated side effects / re-evaluation), consistent with the PR description noting reassociation is not yet handled.

Note

This review was generated by this repository's Holistic Review agentic workflow to complement the built-in Copilot review.

Generated by Holistic Review · 115.6 AIC · ⌖ 11 AIC · ⊞ 10K

Comment threadsrc/coreclr/jit/gentree.cpp Outdated
{
assert(tree->OperIs(GT_AND, GT_OR, GT_XOR, GT_ADD, GT_SUB));

if (opts.OptimizationDisabled() || tree->gtOverflowEx() || !varTypeIsIntegralOrI(tree))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The overflow guard only inspects the outer tree, not the inner operands op1/op2. For an inner distributive op that can overflow-throw, this can silently drop checked semantics. Concretely, checked(A * B) + checked(A * C) has an unchecked outer GT_ADD but two inner GT_MUL nodes carrying GTF_OVERFLOW. isLeftDistributive(GT_MUL, GT_ADD) is true, so the transform rebuilds the tree as A * (B + C) via gtNewOperNode(GT_MUL, ...) with no overflow flag, discarding both checked multiplications. That changes observable behavior (a program that should throw OverflowException no longer does). For unchecked/wrapping arithmetic the distribution is correct mod 2^n, but the checked case is not. Consider also rejecting when op1/op2 (or their relevant children) have gtOverflowEx() set, or more conservatively bail when GTF_EXCEPT is present on the operands being merged.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@BoyBaykiller that is a good point, this PR basically breaks this:

staticintF(inta,intb,intc)=>checked(a*b)+checked(a*c);

With:

a = 2
b = 0x40000000
c = -0x40000000

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am now bailing on GTF_EXCEPT

Comment threadsrc/coreclr/jit/gentree.cpp Outdated
{
GenTree* newOp1 = op1->gtGetOp1();
GenTree* newOp2 =
gtFoldExpr(gtNewOperNode(tree->OperGet(), tree->TypeGet(), op1->gtGetOp2(), op2->gtGetOp2()));

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

seems like this gtNewOperNode is not marked as MORPHED

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

added fgMorphTreeDone(newOp2);

@BoyBaykiller

Copy link
Copy Markdown
ContributorAuthor

Failures fixed?

@MihuBot -nuget

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants

@BoyBaykiller@JulieLeeMSFT@EgorBo@MihaZupan@MihuBot@rhuijben@am11@jakobbotsch@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: Transform arithmetic using distributive property - #126852

Open
BoyBaykiller wants to merge 18 commits into
dotnet:mainfrom
BoyBaykiller:transform-using-distributive-property
Open

JIT: Transform arithmetic using distributive property#126852
BoyBaykiller wants to merge 18 commits into
dotnet:mainfrom
BoyBaykiller:transform-using-distributive-property

Conversation

@BoyBaykiller

@BoyBaykillerBoyBaykiller commented Apr 13, 2026

Copy link
Copy Markdown
Contributor

Generalization of #126070
Progress towards #93954

Basically we weren't doing any simplification based on distributive property before. So transforming
((A op1 B) op2 (A op1 C)) => (A op1 (B op2 C)). And this adds some basic support. Examples:

intMulDistedOverAdd(intA,intB,intC){return(A*B)+(A*C);}
;; ------ BASEG_M000_IG02:moveax,edximuleax,r8dimulr9d,edxaddeax,r9d;; ------ DIFFG_M48043_IG02: ;; offset=0x0000leaeax,[r8+r9]imuleax,edx
boolAfterOptimizeBools(intA,intB){return(A&4)!=0||(A&8)!=0;}
;; ------ BASEmoveax,edxandeax,4andedx,8oreax,edx setne almovzxrax,al;; ------ DIFFG_M55610_IG02:testdl,12 setne almovzxrax,al

We still need something that changes order to enable this opt. So (A | B) | C becoming A | (B | C) in this case:

uintReassociate(uintfoo,uintflags){return(foo|(flags&256))|(flags&512);}

Or here, the C * A need to be reversed.

intReassociate(intA,intB,intC){return(A*B)+(C*A);}

@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Apr 13, 2026
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Apr 13, 2026
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/jit/morph.cpp Outdated
return tree;
}

if (((tree->gtFlags & GTF_PERSISTENT_SIDE_EFFECTS) != 0) || ((tree->gtFlags & GTF_ORDER_SIDEEFF) != 0))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This could use the same optimization you are trying to apply😅... Probably handled by c++ though, so this is purely documentation ;-)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Lol. A new level of self documenting code. I didn't even notice that, I copied this check from elsewhere. I guess this just goes to show how useful the optimization is. Unfortunately even with this PR it's still not handled : (

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Update: This now gets handled by calling this opt again after the "optimizeBools" phase where it simplifies e.g
(A & 4) != 0 || (A & 8) != 0 into ((A & 4) | (A & 8)) != 0.

@BoyBaykiller

BoyBaykiller commented Apr 15, 2026

Copy link
Copy Markdown
ContributorAuthor

@EgorBo PTAL. Diffs are rather small, but nonetheless I believe it's a step in the right direction. I've written down some future work (which all improve diffs). I think the most interesting case for real-world code is arround bitwise ops.
Also should this go into post-order or pre-order morphing?

@BoyBaykiller
BoyBaykiller marked this pull request as ready for review April 15, 2026 01:07
@BoyBaykiller

Copy link
Copy Markdown
ContributorAuthor

Calling fgMorphBlockStmt after optimizeBools fixes point 2 in my list and produces much better diffs oldnew

@BoyBaykiller

BoyBaykiller commented Apr 15, 2026

Copy link
Copy Markdown
ContributorAuthor

Regarding point 1. The issue is that this runs first:

elseif (mulShiftOpt && (lowestBit > 1) && jitIsScaleIndexMul(lowestBit))
{
int shift = genLog2(lowestBit);
ssize_t factor = abs_mult >> shift;
if (factor == 3 || factor == 5 || factor == 9)
{
// if negative negate (min-int does not need negation)
if (mult < 0 && mult != SSIZE_T_MIN)
{
op1 = gtNewOperNode(GT_NEG, genActualType(op1), op1);
mul->gtOp1 = op1;
fgMorphTreeDone(op1);
}
// change the multiplication into a smaller multiplication (by 3, 5 or 9) and a shift
GenTree* const factorNode = gtNewIconNodeWithVN(this, factor, mul->TypeGet());
factorNode->SetMorphed(this);
op1 = gtNewOperNode(GT_MUL, mul->TypeGet(), op1, factorNode);
mul->gtOp1 = op1;
fgMorphTreeDone(op1);
op2->AsIntConCommon()->SetIconValue(shift);
changeToShift = true;
}
}

Perhaps this should be moved to Lower. LLVM also does it in "X86 DAG->DAG Instruction Selection" and not "InstCombine" https://godbolt.org/z/jsPea1PhP

Comment threadsrc/coreclr/jit/optimizebools.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated

auto isLeftDistributive = [](genTreeOps op1, genTreeOps op2) {
// op1 is left distributive over op2 iff:
// "A op1 (B op2 C)" <==> "(A op1 B) op2 (A op1 C)"

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We already do something similar for GT_ADD somewhere, we should unify it with that

@BoyBaykillerBoyBaykillerJun 4, 2026

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I didn't find what you are refering to here

Comment threadsrc/coreclr/jit/optimizebools.cpp Outdated
Comment threadsrc/coreclr/jit/gentree.cpp Outdated
@BoyBaykiller

Copy link
Copy Markdown
ContributorAuthor

I've limited it to OperIsAnyLocal, as discussed. Ready for an other review.

EgorBo pushed a commit that referenced this pull request Jun 4, 2026
Basically so we always have 0 at the right-side:
1. '(A & pow2) == pow2' -> '(A & pow2) != 0'
2. '(A & pow2) != pow2' -> '(A & pow2) == 0'
Tthis currently helps optimizeBools, but I am also adding this with the
future in mind.
Here is one example:
```cs
static bool Example(int foo)
{
return ((foo & 4) == 4) || ((foo & 8) == 8);
}
```
```assembly
;; ------ BASE
G_M000_IG02: ;; offset=0x0000
test cl, 4
je SHORT G_M000_IG05
G_M000_IG03: ;; offset=0x0005
mov eax, 1
G_M000_IG04: ;; offset=0x000A
ret G_M000_IG05: ;; offset=0x000B
test cl, 8
setne al
movzx rax, al
G_M000_IG06: ;; offset=0x0014
ret ;; ------ DIFF
G_M33938_IG02: ;; offset=0x0000
mov eax, ecx
and eax, 4
and ecx, 8
or eax, ecx
setne al
movzx rax, al
```
Note: With #126852 this would then
get transformed into just `test 12`.
@JulieLeeMSFT

Copy link
Copy Markdown
Member

@EgorBo, PTAL.

@EgorBo

Copy link
Copy Markdown
Member

@MihuBot -nuget

@EgorBo

Copy link
Copy Markdown
Member

@MihaZupan

Copy link
Copy Markdown
Member

Just testing error reporting changes

@MihuBot -nuget

@MihuBot

Copy link
Copy Markdown

Warning

2 extra test assemblies failed to produce JIT diffs and were excluded from the diff analysis:

  • MathNet.Numerics.dll (pr)
  • Pipelines.Sockets.Unofficial.dll (pr)

All failures happened only on the PR branch, which likely indicates a bug introduced by the PR (e.g. a JIT assert or crash while compiling these assemblies).

See job details in MihuBot/runtime-utils#2058, triggered by #126852 (comment).

eiriktsarpalis pushed a commit that referenced this pull request Jul 15, 2026
Basically so we always have 0 at the right-side:
1. '(A & pow2) == pow2' -> '(A & pow2) != 0'
2. '(A & pow2) != pow2' -> '(A & pow2) == 0'
Tthis currently helps optimizeBools, but I am also adding this with the
future in mind.
Here is one example:
```cs
static bool Example(int foo)
{
return ((foo & 4) == 4) || ((foo & 8) == 8);
}
```
```assembly
;; ------ BASE
G_M000_IG02: ;; offset=0x0000
test cl, 4
je SHORT G_M000_IG05
G_M000_IG03: ;; offset=0x0005
mov eax, 1
G_M000_IG04: ;; offset=0x000A
ret G_M000_IG05: ;; offset=0x000B
test cl, 8
setne al
movzx rax, al
G_M000_IG06: ;; offset=0x0014
ret ;; ------ DIFF
G_M33938_IG02: ;; offset=0x0000
mov eax, ecx
and eax, 4
and ecx, 8
or eax, ecx
setne al
movzx rax, al
```
Note: With #126852 this would then
get transformed into just `test 12`.
@github-actions

github-actionsBot commented Jul 19, 2026

Copy link
Copy Markdown
Contributor

Workflow state for the Holistic Review Orchestrator.

{
"version": 5,
"last_dispatched_commit": "ef82e04e0d0ec0a577f66e436b30a9adf3b47b77",
"last_dispatched_base_ref": "main",
"last_dispatched_base_sha": "f32edc0223b18865778e8e4d335e31bf9a2933f0",
"last_reviewed_commit": "ef82e04e0d0ec0a577f66e436b30a9adf3b47b77",
"last_reviewed_base_ref": "main",
"last_reviewed_base_sha": "f32edc0223b18865778e8e4d335e31bf9a2933f0",
"last_recorded_worker_run_id": "29679146097",
"review_attempt_commit": "",
"review_attempt_base_ref": "",
"review_attempt_count": 0,
"max_review_attempts": 5,
"review_history_format": "holistic-review-disclosure-v1",
"review_history": [
{
"commit": "ef82e04e0d0ec0a577f66e436b30a9adf3b47b77",
"review_id": 4730527763
}
]
}

@github-actionsgithub-actionsBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Holistic Review

Motivation: Generalizes #126070 and makes progress on #93954 by teaching the JIT to simplify arithmetic via the distributive property, i.e. (A op1 B) op2 (A op1 C) => A op1 (B op2 C). This reduces instruction count in common patterns (e.g. (A*B)+(A*C) folding to A*(B+C), and bit-test patterns after optOptimizeBools), which is a worthwhile codegen improvement.

Approach: Adds gtFoldDistributiveArithmetic invoked from gtFoldExprBinary for GT_AND/GT_OR/GT_XOR/GT_ADD/GT_SUB trees. It uses an isLeftDistributive table (AND/OR/MUL over their duals) and, when both operands share the same outer op and their left children are matching locals, rebuilds the tree with the common factor extracted. It also re-runs gtFoldExpr on the newly-combined boolean fold result in optOptimizeBoolsUpdateTrees, and tidies an unrelated cmp->gtOp2 -> cmp->gtGetOp2() reference in morph.cpp.

Summary: The transform is conceptually sound for wrapping (unchecked) integer arithmetic, and the guarding on optimization level, overflow of the outer node, side effects, and integral type is a reasonable start. However there is one correctness concern worth resolving before merge: the overflow check only examines the outer tree, not the inner operands. For checked(A*B) + checked(A*C) the inner GT_MUL nodes carry GTF_OVERFLOW, yet the rebuilt A*(B+C) uses a plain unchecked GT_MUL, dropping the overflow checks and altering observable behavior (a missing OverflowException). See the inline comment on the guard. I'd recommend adding a targeted regression test covering the checked-arithmetic case (and confirming it is rejected) once the guard is tightened. A secondary, non-blocking note: the freshly-created newOp2 node does not get value numbers assigned even though gtFoldExpr may be reached post-VN via optOptimizeBools; only result receives SetVNsFromNode. Verify this is benign in that phase (constants typically fold away, but a non-constant B op C combination would leave an unnumbered node).

Detailed Findings

  • Checked-arithmetic correctness (see inline comment on gtFoldDistributiveArithmetic): outer-only overflow guard can drop inner checked-multiply/add/sub semantics. Recommend also bailing when the operands being merged have gtOverflowEx()/GTF_EXCEPT set.
  • VN coverage of newOp2 (non-blocking): SetVNsFromNode(tree) is applied to result only; the intermediate newOp2 = gtFoldExpr(B op C) node has no VN. Since this path is also reachable from optOptimizeBoolsUpdateTrees after value numbering, confirm an unnumbered non-constant intermediate cannot reach a consumer that requires valid VNs.
  • Scope observation (non-blocking): the transform requires the shared factor to be OperIsAnyLocal. This is a safe, conservative choice (no duplicated side effects / re-evaluation), consistent with the PR description noting reassociation is not yet handled.

Note

This review was generated by this repository's Holistic Review agentic workflow to complement the built-in Copilot review.

Generated by Holistic Review · 115.6 AIC · ⌖ 11 AIC · ⊞ 10K

Comment threadsrc/coreclr/jit/gentree.cpp Outdated
{
assert(tree->OperIs(GT_AND, GT_OR, GT_XOR, GT_ADD, GT_SUB));

if (opts.OptimizationDisabled() || tree->gtOverflowEx() || !varTypeIsIntegralOrI(tree))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The overflow guard only inspects the outer tree, not the inner operands op1/op2. For an inner distributive op that can overflow-throw, this can silently drop checked semantics. Concretely, checked(A * B) + checked(A * C) has an unchecked outer GT_ADD but two inner GT_MUL nodes carrying GTF_OVERFLOW. isLeftDistributive(GT_MUL, GT_ADD) is true, so the transform rebuilds the tree as A * (B + C) via gtNewOperNode(GT_MUL, ...) with no overflow flag, discarding both checked multiplications. That changes observable behavior (a program that should throw OverflowException no longer does). For unchecked/wrapping arithmetic the distribution is correct mod 2^n, but the checked case is not. Consider also rejecting when op1/op2 (or their relevant children) have gtOverflowEx() set, or more conservatively bail when GTF_EXCEPT is present on the operands being merged.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@BoyBaykiller that is a good point, this PR basically breaks this:

staticintF(inta,intb,intc)=>checked(a*b)+checked(a*c);

With:

a = 2
b = 0x40000000
c = -0x40000000

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am now bailing on GTF_EXCEPT

Comment threadsrc/coreclr/jit/gentree.cpp Outdated
{
GenTree* newOp1 = op1->gtGetOp1();
GenTree* newOp2 =
gtFoldExpr(gtNewOperNode(tree->OperGet(), tree->TypeGet(), op1->gtGetOp2(), op2->gtGetOp2()));

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

seems like this gtNewOperNode is not marked as MORPHED

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

added fgMorphTreeDone(newOp2);

@BoyBaykiller

Copy link
Copy Markdown
ContributorAuthor

Failures fixed?

@MihuBot -nuget

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants

@BoyBaykiller@JulieLeeMSFT@EgorBo@MihaZupan@MihuBot@rhuijben@am11@jakobbotsch@tannergooding
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

JIT: Transform arithmetic using distributive property - #126852

Open
BoyBaykiller wants to merge 18 commits into
dotnet:mainfrom
BoyBaykiller:transform-using-distributive-property
Open

JIT: Transform arithmetic using distributive property#126852
BoyBaykiller wants to merge 18 commits into
dotnet:mainfrom
BoyBaykiller:transform-using-distributive-property

Conversation

@BoyBaykiller

@BoyBaykillerBoyBaykiller commented Apr 13, 2026

Copy link
Copy Markdown
Contributor

Generalization of #126070
Progress towards #93954

Basically we weren't doing any simplification based on distributive property before. So transforming
((A op1 B) op2 (A op1 C)) => (A op1 (B op2 C)). And this adds some basic support. Examples:

intMulDistedOverAdd(intA,intB,intC){return(A*B)+(A*C);}
;; ------ BASEG_M000_IG02:moveax,edximuleax,r8dimulr9d,edxaddeax,r9d;; ------ DIFFG_M48043_IG02: ;; offset=0x0000leaeax,[r8+r9]imuleax,edx
boolAfterOptimizeBools(intA,intB){return(A&4)!=0||(A&8)!=0;}
;; ------ BASEmoveax,edxandeax,4andedx,8oreax,edx setne almovzxrax,al;; ------ DIFFG_M55610_IG02:testdl,12 setne almovzxrax,al

We still need something that changes order to enable this opt. So (A | B) | C becoming A | (B | C) in this case:

uintReassociate(uintfoo,uintflags){return(foo|(flags&256))|(flags&512);}

Or here, the C * A need to be reversed.

intReassociate(intA,intB,intC){return(A*B)+(C*A);}

@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Apr 13, 2026
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Apr 13, 2026
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/jit/morph.cpp Outdated
return tree;
}

if (((tree->gtFlags & GTF_PERSISTENT_SIDE_EFFECTS) != 0) || ((tree->gtFlags & GTF_ORDER_SIDEEFF) != 0))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This could use the same optimization you are trying to apply😅... Probably handled by c++ though, so this is purely documentation ;-)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Lol. A new level of self documenting code. I didn't even notice that, I copied this check from elsewhere. I guess this just goes to show how useful the optimization is. Unfortunately even with this PR it's still not handled : (

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Update: This now gets handled by calling this opt again after the "optimizeBools" phase where it simplifies e.g
(A & 4) != 0 || (A & 8) != 0 into ((A & 4) | (A & 8)) != 0.

@BoyBaykiller

BoyBaykiller commented Apr 15, 2026

Copy link
Copy Markdown
ContributorAuthor

@EgorBo PTAL. Diffs are rather small, but nonetheless I believe it's a step in the right direction. I've written down some future work (which all improve diffs). I think the most interesting case for real-world code is arround bitwise ops.
Also should this go into post-order or pre-order morphing?

@BoyBaykiller
BoyBaykiller marked this pull request as ready for review April 15, 2026 01:07
@BoyBaykiller

Copy link
Copy Markdown
ContributorAuthor

Calling fgMorphBlockStmt after optimizeBools fixes point 2 in my list and produces much better diffs oldnew

@BoyBaykiller

BoyBaykiller commented Apr 15, 2026

Copy link
Copy Markdown
ContributorAuthor

Regarding point 1. The issue is that this runs first:

elseif (mulShiftOpt && (lowestBit > 1) && jitIsScaleIndexMul(lowestBit))
{
int shift = genLog2(lowestBit);
ssize_t factor = abs_mult >> shift;
if (factor == 3 || factor == 5 || factor == 9)
{
// if negative negate (min-int does not need negation)
if (mult < 0 && mult != SSIZE_T_MIN)
{
op1 = gtNewOperNode(GT_NEG, genActualType(op1), op1);
mul->gtOp1 = op1;
fgMorphTreeDone(op1);
}
// change the multiplication into a smaller multiplication (by 3, 5 or 9) and a shift
GenTree* const factorNode = gtNewIconNodeWithVN(this, factor, mul->TypeGet());
factorNode->SetMorphed(this);
op1 = gtNewOperNode(GT_MUL, mul->TypeGet(), op1, factorNode);
mul->gtOp1 = op1;
fgMorphTreeDone(op1);
op2->AsIntConCommon()->SetIconValue(shift);
changeToShift = true;
}
}

Perhaps this should be moved to Lower. LLVM also does it in "X86 DAG->DAG Instruction Selection" and not "InstCombine" https://godbolt.org/z/jsPea1PhP

Comment threadsrc/coreclr/jit/optimizebools.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated
Comment threadsrc/coreclr/jit/morph.cpp Outdated

auto isLeftDistributive = [](genTreeOps op1, genTreeOps op2) {
// op1 is left distributive over op2 iff:
// "A op1 (B op2 C)" <==> "(A op1 B) op2 (A op1 C)"

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We already do something similar for GT_ADD somewhere, we should unify it with that

@BoyBaykillerBoyBaykillerJun 4, 2026

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I didn't find what you are refering to here

Comment threadsrc/coreclr/jit/optimizebools.cpp Outdated
Comment threadsrc/coreclr/jit/gentree.cpp Outdated
@BoyBaykiller

Copy link
Copy Markdown
ContributorAuthor

I've limited it to OperIsAnyLocal, as discussed. Ready for an other review.

EgorBo pushed a commit that referenced this pull request Jun 4, 2026
Basically so we always have 0 at the right-side:
1. '(A & pow2) == pow2' -> '(A & pow2) != 0'
2. '(A & pow2) != pow2' -> '(A & pow2) == 0'
Tthis currently helps optimizeBools, but I am also adding this with the
future in mind.
Here is one example:
```cs
static bool Example(int foo)
{
return ((foo & 4) == 4) || ((foo & 8) == 8);
}
```
```assembly
;; ------ BASE
G_M000_IG02: ;; offset=0x0000
test cl, 4
je SHORT G_M000_IG05
G_M000_IG03: ;; offset=0x0005
mov eax, 1
G_M000_IG04: ;; offset=0x000A
ret G_M000_IG05: ;; offset=0x000B
test cl, 8
setne al
movzx rax, al
G_M000_IG06: ;; offset=0x0014
ret ;; ------ DIFF
G_M33938_IG02: ;; offset=0x0000
mov eax, ecx
and eax, 4
and ecx, 8
or eax, ecx
setne al
movzx rax, al
```
Note: With #126852 this would then
get transformed into just `test 12`.
@JulieLeeMSFT

Copy link
Copy Markdown
Member

@EgorBo, PTAL.

@EgorBo

Copy link
Copy Markdown
Member

@MihuBot -nuget

@EgorBo

Copy link
Copy Markdown
Member

@MihaZupan

Copy link
Copy Markdown
Member

Just testing error reporting changes

@MihuBot -nuget

@MihuBot

Copy link
Copy Markdown

Warning

2 extra test assemblies failed to produce JIT diffs and were excluded from the diff analysis:

  • MathNet.Numerics.dll (pr)
  • Pipelines.Sockets.Unofficial.dll (pr)

All failures happened only on the PR branch, which likely indicates a bug introduced by the PR (e.g. a JIT assert or crash while compiling these assemblies).

See job details in MihuBot/runtime-utils#2058, triggered by #126852 (comment).

eiriktsarpalis pushed a commit that referenced this pull request Jul 15, 2026
Basically so we always have 0 at the right-side:
1. '(A & pow2) == pow2' -> '(A & pow2) != 0'
2. '(A & pow2) != pow2' -> '(A & pow2) == 0'
Tthis currently helps optimizeBools, but I am also adding this with the
future in mind.
Here is one example:
```cs
static bool Example(int foo)
{
return ((foo & 4) == 4) || ((foo & 8) == 8);
}
```
```assembly
;; ------ BASE
G_M000_IG02: ;; offset=0x0000
test cl, 4
je SHORT G_M000_IG05
G_M000_IG03: ;; offset=0x0005
mov eax, 1
G_M000_IG04: ;; offset=0x000A
ret G_M000_IG05: ;; offset=0x000B
test cl, 8
setne al
movzx rax, al
G_M000_IG06: ;; offset=0x0014
ret ;; ------ DIFF
G_M33938_IG02: ;; offset=0x0000
mov eax, ecx
and eax, 4
and ecx, 8
or eax, ecx
setne al
movzx rax, al
```
Note: With #126852 this would then
get transformed into just `test 12`.
@github-actions

github-actionsBot commented Jul 19, 2026

Copy link
Copy Markdown
Contributor

Workflow state for the Holistic Review Orchestrator.

{
"version": 5,
"last_dispatched_commit": "ef82e04e0d0ec0a577f66e436b30a9adf3b47b77",
"last_dispatched_base_ref": "main",
"last_dispatched_base_sha": "f32edc0223b18865778e8e4d335e31bf9a2933f0",
"last_reviewed_commit": "ef82e04e0d0ec0a577f66e436b30a9adf3b47b77",
"last_reviewed_base_ref": "main",
"last_reviewed_base_sha": "f32edc0223b18865778e8e4d335e31bf9a2933f0",
"last_recorded_worker_run_id": "29679146097",
"review_attempt_commit": "",
"review_attempt_base_ref": "",
"review_attempt_count": 0,
"max_review_attempts": 5,
"review_history_format": "holistic-review-disclosure-v1",
"review_history": [
{
"commit": "ef82e04e0d0ec0a577f66e436b30a9adf3b47b77",
"review_id": 4730527763
}
]
}

@github-actionsgithub-actionsBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Holistic Review

Motivation: Generalizes #126070 and makes progress on #93954 by teaching the JIT to simplify arithmetic via the distributive property, i.e. (A op1 B) op2 (A op1 C) => A op1 (B op2 C). This reduces instruction count in common patterns (e.g. (A*B)+(A*C) folding to A*(B+C), and bit-test patterns after optOptimizeBools), which is a worthwhile codegen improvement.

Approach: Adds gtFoldDistributiveArithmetic invoked from gtFoldExprBinary for GT_AND/GT_OR/GT_XOR/GT_ADD/GT_SUB trees. It uses an isLeftDistributive table (AND/OR/MUL over their duals) and, when both operands share the same outer op and their left children are matching locals, rebuilds the tree with the common factor extracted. It also re-runs gtFoldExpr on the newly-combined boolean fold result in optOptimizeBoolsUpdateTrees, and tidies an unrelated cmp->gtOp2 -> cmp->gtGetOp2() reference in morph.cpp.

Summary: The transform is conceptually sound for wrapping (unchecked) integer arithmetic, and the guarding on optimization level, overflow of the outer node, side effects, and integral type is a reasonable start. However there is one correctness concern worth resolving before merge: the overflow check only examines the outer tree, not the inner operands. For checked(A*B) + checked(A*C) the inner GT_MUL nodes carry GTF_OVERFLOW, yet the rebuilt A*(B+C) uses a plain unchecked GT_MUL, dropping the overflow checks and altering observable behavior (a missing OverflowException). See the inline comment on the guard. I'd recommend adding a targeted regression test covering the checked-arithmetic case (and confirming it is rejected) once the guard is tightened. A secondary, non-blocking note: the freshly-created newOp2 node does not get value numbers assigned even though gtFoldExpr may be reached post-VN via optOptimizeBools; only result receives SetVNsFromNode. Verify this is benign in that phase (constants typically fold away, but a non-constant B op C combination would leave an unnumbered node).

Detailed Findings

  • Checked-arithmetic correctness (see inline comment on gtFoldDistributiveArithmetic): outer-only overflow guard can drop inner checked-multiply/add/sub semantics. Recommend also bailing when the operands being merged have gtOverflowEx()/GTF_EXCEPT set.
  • VN coverage of newOp2 (non-blocking): SetVNsFromNode(tree) is applied to result only; the intermediate newOp2 = gtFoldExpr(B op C) node has no VN. Since this path is also reachable from optOptimizeBoolsUpdateTrees after value numbering, confirm an unnumbered non-constant intermediate cannot reach a consumer that requires valid VNs.
  • Scope observation (non-blocking): the transform requires the shared factor to be OperIsAnyLocal. This is a safe, conservative choice (no duplicated side effects / re-evaluation), consistent with the PR description noting reassociation is not yet handled.

Note

This review was generated by this repository's Holistic Review agentic workflow to complement the built-in Copilot review.

Generated by Holistic Review · 115.6 AIC · ⌖ 11 AIC · ⊞ 10K

Comment threadsrc/coreclr/jit/gentree.cpp Outdated
{
assert(tree->OperIs(GT_AND, GT_OR, GT_XOR, GT_ADD, GT_SUB));

if (opts.OptimizationDisabled() || tree->gtOverflowEx() || !varTypeIsIntegralOrI(tree))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The overflow guard only inspects the outer tree, not the inner operands op1/op2. For an inner distributive op that can overflow-throw, this can silently drop checked semantics. Concretely, checked(A * B) + checked(A * C) has an unchecked outer GT_ADD but two inner GT_MUL nodes carrying GTF_OVERFLOW. isLeftDistributive(GT_MUL, GT_ADD) is true, so the transform rebuilds the tree as A * (B + C) via gtNewOperNode(GT_MUL, ...) with no overflow flag, discarding both checked multiplications. That changes observable behavior (a program that should throw OverflowException no longer does). For unchecked/wrapping arithmetic the distribution is correct mod 2^n, but the checked case is not. Consider also rejecting when op1/op2 (or their relevant children) have gtOverflowEx() set, or more conservatively bail when GTF_EXCEPT is present on the operands being merged.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@BoyBaykiller that is a good point, this PR basically breaks this:

staticintF(inta,intb,intc)=>checked(a*b)+checked(a*c);

With:

a = 2
b = 0x40000000
c = -0x40000000

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am now bailing on GTF_EXCEPT

Comment threadsrc/coreclr/jit/gentree.cpp Outdated
{
GenTree* newOp1 = op1->gtGetOp1();
GenTree* newOp2 =
gtFoldExpr(gtNewOperNode(tree->OperGet(), tree->TypeGet(), op1->gtGetOp2(), op2->gtGetOp2()));

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

seems like this gtNewOperNode is not marked as MORPHED

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

added fgMorphTreeDone(newOp2);

@BoyBaykiller

Copy link
Copy Markdown
ContributorAuthor

Failures fixed?

@MihuBot -nuget

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants

@BoyBaykiller@JulieLeeMSFT@EgorBo@MihaZupan@MihuBot@rhuijben@am11@jakobbotsch@tannergooding