Skip to content

arm64: Add SVE ptrue reuse pass after lowering - #128844

Closed
jonathandavies-arm wants to merge 7 commits into
dotnet:mainfrom
jonathandavies-arm:upstream/sve/ptrue-reuse
Closed

arm64: Add SVE ptrue reuse pass after lowering#128844
jonathandavies-arm wants to merge 7 commits into
dotnet:mainfrom
jonathandavies-arm:upstream/sve/ptrue-reuse

Conversation

@jonathandavies-arm

Copy link
Copy Markdown
Contributor
  • Reuses equivalent ptrue producing nodes within a block via a mask temp

- Reuses equivalent ptrue-producing nodes within a block via a mask temp
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jun 1, 2026
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jun 1, 2026
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/jit/compiler.cpp Outdated
Comment threadsrc/coreclr/jit/CMakeLists.txt Outdated
Comment threadsrc/coreclr/jit/constantmaskreuse.cpp Outdated
@tannergooding

Copy link
Copy Markdown
Member

The logic here looks overall correct to me. However, diffs are not looking very good.

Rather this comes at a significant throughput hit to Arm64 (+0.54% in MinOpts) and so likely, at a minimum, needs to be skipped if optimizations are not enabled.

Then, this actually doesn't appear to be improving codegen either. Instead, it looks to significantly regress codegen (+199k bytes), particularly for MinOpts where we get a bunch of ldr, add, ldr sequences instead of ptrue or reusing an existing constant.

If you avoid running this for T0 code, then we'll still have +65k bytes of new codegen, so I imagine there is still something to be fixed/improved here. But, those diffs are generally hard to find given how much T0 code regressed and so you can't really identify what the issue is or if its the same ldr, add, ldr issue.

@tannergooding

tannergooding commented Jun 25, 2026

Copy link
Copy Markdown
Member

diffs look better now with everything but the coreclr_tests showing improvements.

Several of the tests still show regressions (+10k bytes of codegen), however, where we effectively have the following instead:

 ptrue p0.s
+ add xip1, fp, #24+ str p0, [xip1]	// [V33 rat0]
...
- ptrue p0.s+ add xip1, fp, #24+ ldr p0, [xip1]	// [V33 rat0]

Is there some particular pattern here that we're failing to account for, such as some call that is forcing a spill between the two callsites or perhaps the local being annotated in a way that forces a spill?

@tannergoodingtannergooding left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Changes LGTM and diffs are now purely positive.

There's still a TP hit, but I doubt that can be mitigated more as its essentially just the cost of adding a phase, whether that phase is executed or not.

@dotnet/jit-contrib, @jakobbotsch, @EgorBo for secondary review and input on whether such a phase is worth having.

This would notably be extendable to AllBitsSet and Zero for general SIMD or floating-point constants on all platforms (Arm64 and xarch) so isn't SVE specific and would greatly mitigate some of the issues we see in #70182 (which is a complex LSRA issue)

@jakobbotsch

Copy link
Copy Markdown
Member

@dotnet/jit-contrib, @jakobbotsch, @EgorBo for secondary review and input on whether such a phase is worth having.

I would much rather see the effort spent on improving the constant reuse that LSRA already implements. This implementation seems to come with some rather severe limitations (single block only and impoverished reasoning about register kills being some of them), in addition to being a new phase that solves a problem we already try to solve.

@jonathandavies-arm Did you look into expanding LSRA's constant reuse and investigate out why some of the cases this PR handles are not handled by it?

@github-actions

github-actionsBot commented Jul 19, 2026

Copy link
Copy Markdown
Contributor

Workflow state for the Holistic Review Orchestrator.

{
"version": 5,
"last_dispatched_commit": "8d3f5534cb1ee5dba6f1fd30e30fc8c97674f322",
"last_dispatched_base_ref": "main",
"last_dispatched_base_sha": "e137b2c018221d7babd79162f0fd151367c34c58",
"last_reviewed_commit": "8d3f5534cb1ee5dba6f1fd30e30fc8c97674f322",
"last_reviewed_base_ref": "main",
"last_reviewed_base_sha": "e137b2c018221d7babd79162f0fd151367c34c58",
"last_recorded_worker_run_id": "29684000455",
"review_attempt_commit": "",
"review_attempt_base_ref": "",
"review_attempt_count": 0,
"max_review_attempts": 5,
"review_history_format": "holistic-review-disclosure-v1",
"review_history": [
{
"commit": "8d3f5534cb1ee5dba6f1fd30e30fc8c97674f322",
"review_id": 4730635250
}
]
}

@github-actionsgithub-actionsBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Holistic Review

Motivation: The problem is real. On arm64/SVE, predicate-producing constant masks (ptrue/pfalse) are frequently rematerialized multiple times within a block, and the JIT had no post-lowering mechanism to share an equivalent predicate via a temp. Reducing redundant ptrue/pfalse materialization is a legitimate codegen win.

Approach: The approach is sound and appropriately conservative. It introduces a new post-lowering PHASE_CONSTANT_REUSE (arm64 + FEATURE_MASKED_HW_INTRINSICS, optimizations only), materializing the first equivalent constant-mask into a local temp and rewriting later equivalent uses as loads, relying on normal local-var lifetimes so LSRA models predicate clobbers. The refactor that hoists the post-lowering liveness/ref-count/dead-block cleanup out of Lowering::DoPhase into a new Compiler::fgPostLowering phase is a clean mechanical move that preserves the original ordering and semantics, and correctly runs after the new reuse phase. The guards are careful: it bails when locals are not enregistered, skips all-true masks whose live range crosses a call, respects contained masks by poisoning the whole group, and preserves the ConditionalSelect(AllTrue, embedded-op, zero) movprfx-free codegen special case. Pattern normalization (LargestPowerOf2 -> All, pfalse sentinel keyed as .b) matches how these lower.

Summary: ✅ LGTM. The change is well-structured, guarded conservatively, and already carries an approval from a JIT maintainer (tannergooding). The refactor preserves prior post-lowering semantics, and the added FileCheck-based test covers a good spread of true/false/pattern/embedded/conversion/SVE2 cases plus single-vs-multiple reuse scenarios. One minor non-blocking observation: the new IsSveBreakMask helper in hwintrinsic.h is dead code with no callers (flagged inline). Nothing here blocks merge.

Note

This review was generated by this repository's Holistic Review agentic workflow to complement the built-in Copilot review.

Generated by Holistic Review · 146.3 AIC · ⌖ 10.5 AIC · ⊞ 10K

}
}

static bool IsSveBreakMask(NamedIntrinsic id)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 IsSveBreakMask is added here but never referenced anywhere in the JIT (IsSveCreateTrueMask above it is used by constantreuse.cpp, but this helper has no callers). If it is intended for a follow-up, a brief comment noting that would help; otherwise consider dropping it to avoid dead code.

@jonathandavies-arm

Copy link
Copy Markdown
ContributorAuthor

I'm going to close this PR and the LSRA version of the work is here #131309.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jonathandavies-arm@tannergooding@jakobbotsch
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
arm64: Add SVE ptrue reuse pass after lowering by jonathandavies-arm · Pull Request #128844 · dotnet/runtime · GitHub
Skip to content

arm64: Add SVE ptrue reuse pass after lowering - #128844

Closed
jonathandavies-arm wants to merge 7 commits into
dotnet:mainfrom
jonathandavies-arm:upstream/sve/ptrue-reuse
Closed

arm64: Add SVE ptrue reuse pass after lowering#128844
jonathandavies-arm wants to merge 7 commits into
dotnet:mainfrom
jonathandavies-arm:upstream/sve/ptrue-reuse

Conversation

@jonathandavies-arm

Copy link
Copy Markdown
Contributor
  • Reuses equivalent ptrue producing nodes within a block via a mask temp

- Reuses equivalent ptrue-producing nodes within a block via a mask temp
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jun 1, 2026
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jun 1, 2026
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/jit/compiler.cpp Outdated
Comment threadsrc/coreclr/jit/CMakeLists.txt Outdated
Comment threadsrc/coreclr/jit/constantmaskreuse.cpp Outdated
@tannergooding

Copy link
Copy Markdown
Member

The logic here looks overall correct to me. However, diffs are not looking very good.

Rather this comes at a significant throughput hit to Arm64 (+0.54% in MinOpts) and so likely, at a minimum, needs to be skipped if optimizations are not enabled.

Then, this actually doesn't appear to be improving codegen either. Instead, it looks to significantly regress codegen (+199k bytes), particularly for MinOpts where we get a bunch of ldr, add, ldr sequences instead of ptrue or reusing an existing constant.

If you avoid running this for T0 code, then we'll still have +65k bytes of new codegen, so I imagine there is still something to be fixed/improved here. But, those diffs are generally hard to find given how much T0 code regressed and so you can't really identify what the issue is or if its the same ldr, add, ldr issue.

@tannergooding

tannergooding commented Jun 25, 2026

Copy link
Copy Markdown
Member

diffs look better now with everything but the coreclr_tests showing improvements.

Several of the tests still show regressions (+10k bytes of codegen), however, where we effectively have the following instead:

 ptrue p0.s
+ add xip1, fp, #24+ str p0, [xip1]	// [V33 rat0]
...
- ptrue p0.s+ add xip1, fp, #24+ ldr p0, [xip1]	// [V33 rat0]

Is there some particular pattern here that we're failing to account for, such as some call that is forcing a spill between the two callsites or perhaps the local being annotated in a way that forces a spill?

@tannergoodingtannergooding left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Changes LGTM and diffs are now purely positive.

There's still a TP hit, but I doubt that can be mitigated more as its essentially just the cost of adding a phase, whether that phase is executed or not.

@dotnet/jit-contrib, @jakobbotsch, @EgorBo for secondary review and input on whether such a phase is worth having.

This would notably be extendable to AllBitsSet and Zero for general SIMD or floating-point constants on all platforms (Arm64 and xarch) so isn't SVE specific and would greatly mitigate some of the issues we see in #70182 (which is a complex LSRA issue)

@jakobbotsch

Copy link
Copy Markdown
Member

@dotnet/jit-contrib, @jakobbotsch, @EgorBo for secondary review and input on whether such a phase is worth having.

I would much rather see the effort spent on improving the constant reuse that LSRA already implements. This implementation seems to come with some rather severe limitations (single block only and impoverished reasoning about register kills being some of them), in addition to being a new phase that solves a problem we already try to solve.

@jonathandavies-arm Did you look into expanding LSRA's constant reuse and investigate out why some of the cases this PR handles are not handled by it?

@github-actions

github-actionsBot commented Jul 19, 2026

Copy link
Copy Markdown
Contributor

Workflow state for the Holistic Review Orchestrator.

{
"version": 5,
"last_dispatched_commit": "8d3f5534cb1ee5dba6f1fd30e30fc8c97674f322",
"last_dispatched_base_ref": "main",
"last_dispatched_base_sha": "e137b2c018221d7babd79162f0fd151367c34c58",
"last_reviewed_commit": "8d3f5534cb1ee5dba6f1fd30e30fc8c97674f322",
"last_reviewed_base_ref": "main",
"last_reviewed_base_sha": "e137b2c018221d7babd79162f0fd151367c34c58",
"last_recorded_worker_run_id": "29684000455",
"review_attempt_commit": "",
"review_attempt_base_ref": "",
"review_attempt_count": 0,
"max_review_attempts": 5,
"review_history_format": "holistic-review-disclosure-v1",
"review_history": [
{
"commit": "8d3f5534cb1ee5dba6f1fd30e30fc8c97674f322",
"review_id": 4730635250
}
]
}

@github-actionsgithub-actionsBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Holistic Review

Motivation: The problem is real. On arm64/SVE, predicate-producing constant masks (ptrue/pfalse) are frequently rematerialized multiple times within a block, and the JIT had no post-lowering mechanism to share an equivalent predicate via a temp. Reducing redundant ptrue/pfalse materialization is a legitimate codegen win.

Approach: The approach is sound and appropriately conservative. It introduces a new post-lowering PHASE_CONSTANT_REUSE (arm64 + FEATURE_MASKED_HW_INTRINSICS, optimizations only), materializing the first equivalent constant-mask into a local temp and rewriting later equivalent uses as loads, relying on normal local-var lifetimes so LSRA models predicate clobbers. The refactor that hoists the post-lowering liveness/ref-count/dead-block cleanup out of Lowering::DoPhase into a new Compiler::fgPostLowering phase is a clean mechanical move that preserves the original ordering and semantics, and correctly runs after the new reuse phase. The guards are careful: it bails when locals are not enregistered, skips all-true masks whose live range crosses a call, respects contained masks by poisoning the whole group, and preserves the ConditionalSelect(AllTrue, embedded-op, zero) movprfx-free codegen special case. Pattern normalization (LargestPowerOf2 -> All, pfalse sentinel keyed as .b) matches how these lower.

Summary: ✅ LGTM. The change is well-structured, guarded conservatively, and already carries an approval from a JIT maintainer (tannergooding). The refactor preserves prior post-lowering semantics, and the added FileCheck-based test covers a good spread of true/false/pattern/embedded/conversion/SVE2 cases plus single-vs-multiple reuse scenarios. One minor non-blocking observation: the new IsSveBreakMask helper in hwintrinsic.h is dead code with no callers (flagged inline). Nothing here blocks merge.

Note

This review was generated by this repository's Holistic Review agentic workflow to complement the built-in Copilot review.

Generated by Holistic Review · 146.3 AIC · ⌖ 10.5 AIC · ⊞ 10K

}
}

static bool IsSveBreakMask(NamedIntrinsic id)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 IsSveBreakMask is added here but never referenced anywhere in the JIT (IsSveCreateTrueMask above it is used by constantreuse.cpp, but this helper has no callers). If it is intended for a follow-up, a brief comment noting that would help; otherwise consider dropping it to avoid dead code.

@jonathandavies-arm

Copy link
Copy Markdown
ContributorAuthor

I'm going to close this PR and the LSRA version of the work is here #131309.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jonathandavies-arm@tannergooding@jakobbotsch
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' arm64: Add SVE ptrue reuse pass after lowering by jonathandavies-arm · Pull Request #128844 · dotnet/runtime · GitHub
Skip to content

arm64: Add SVE ptrue reuse pass after lowering - #128844

Closed
jonathandavies-arm wants to merge 7 commits into
dotnet:mainfrom
jonathandavies-arm:upstream/sve/ptrue-reuse
Closed

arm64: Add SVE ptrue reuse pass after lowering#128844
jonathandavies-arm wants to merge 7 commits into
dotnet:mainfrom
jonathandavies-arm:upstream/sve/ptrue-reuse

Conversation

@jonathandavies-arm

Copy link
Copy Markdown
Contributor
  • Reuses equivalent ptrue producing nodes within a block via a mask temp

- Reuses equivalent ptrue-producing nodes within a block via a mask temp
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jun 1, 2026
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jun 1, 2026
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/jit/compiler.cpp Outdated
Comment threadsrc/coreclr/jit/CMakeLists.txt Outdated
Comment threadsrc/coreclr/jit/constantmaskreuse.cpp Outdated
@tannergooding

Copy link
Copy Markdown
Member

The logic here looks overall correct to me. However, diffs are not looking very good.

Rather this comes at a significant throughput hit to Arm64 (+0.54% in MinOpts) and so likely, at a minimum, needs to be skipped if optimizations are not enabled.

Then, this actually doesn't appear to be improving codegen either. Instead, it looks to significantly regress codegen (+199k bytes), particularly for MinOpts where we get a bunch of ldr, add, ldr sequences instead of ptrue or reusing an existing constant.

If you avoid running this for T0 code, then we'll still have +65k bytes of new codegen, so I imagine there is still something to be fixed/improved here. But, those diffs are generally hard to find given how much T0 code regressed and so you can't really identify what the issue is or if its the same ldr, add, ldr issue.

@tannergooding

tannergooding commented Jun 25, 2026

Copy link
Copy Markdown
Member

diffs look better now with everything but the coreclr_tests showing improvements.

Several of the tests still show regressions (+10k bytes of codegen), however, where we effectively have the following instead:

 ptrue p0.s
+ add xip1, fp, #24+ str p0, [xip1]	// [V33 rat0]
...
- ptrue p0.s+ add xip1, fp, #24+ ldr p0, [xip1]	// [V33 rat0]

Is there some particular pattern here that we're failing to account for, such as some call that is forcing a spill between the two callsites or perhaps the local being annotated in a way that forces a spill?

@tannergoodingtannergooding left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Changes LGTM and diffs are now purely positive.

There's still a TP hit, but I doubt that can be mitigated more as its essentially just the cost of adding a phase, whether that phase is executed or not.

@dotnet/jit-contrib, @jakobbotsch, @EgorBo for secondary review and input on whether such a phase is worth having.

This would notably be extendable to AllBitsSet and Zero for general SIMD or floating-point constants on all platforms (Arm64 and xarch) so isn't SVE specific and would greatly mitigate some of the issues we see in #70182 (which is a complex LSRA issue)

@jakobbotsch

Copy link
Copy Markdown
Member

@dotnet/jit-contrib, @jakobbotsch, @EgorBo for secondary review and input on whether such a phase is worth having.

I would much rather see the effort spent on improving the constant reuse that LSRA already implements. This implementation seems to come with some rather severe limitations (single block only and impoverished reasoning about register kills being some of them), in addition to being a new phase that solves a problem we already try to solve.

@jonathandavies-arm Did you look into expanding LSRA's constant reuse and investigate out why some of the cases this PR handles are not handled by it?

@github-actions

github-actionsBot commented Jul 19, 2026

Copy link
Copy Markdown
Contributor

Workflow state for the Holistic Review Orchestrator.

{
"version": 5,
"last_dispatched_commit": "8d3f5534cb1ee5dba6f1fd30e30fc8c97674f322",
"last_dispatched_base_ref": "main",
"last_dispatched_base_sha": "e137b2c018221d7babd79162f0fd151367c34c58",
"last_reviewed_commit": "8d3f5534cb1ee5dba6f1fd30e30fc8c97674f322",
"last_reviewed_base_ref": "main",
"last_reviewed_base_sha": "e137b2c018221d7babd79162f0fd151367c34c58",
"last_recorded_worker_run_id": "29684000455",
"review_attempt_commit": "",
"review_attempt_base_ref": "",
"review_attempt_count": 0,
"max_review_attempts": 5,
"review_history_format": "holistic-review-disclosure-v1",
"review_history": [
{
"commit": "8d3f5534cb1ee5dba6f1fd30e30fc8c97674f322",
"review_id": 4730635250
}
]
}

@github-actionsgithub-actionsBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Holistic Review

Motivation: The problem is real. On arm64/SVE, predicate-producing constant masks (ptrue/pfalse) are frequently rematerialized multiple times within a block, and the JIT had no post-lowering mechanism to share an equivalent predicate via a temp. Reducing redundant ptrue/pfalse materialization is a legitimate codegen win.

Approach: The approach is sound and appropriately conservative. It introduces a new post-lowering PHASE_CONSTANT_REUSE (arm64 + FEATURE_MASKED_HW_INTRINSICS, optimizations only), materializing the first equivalent constant-mask into a local temp and rewriting later equivalent uses as loads, relying on normal local-var lifetimes so LSRA models predicate clobbers. The refactor that hoists the post-lowering liveness/ref-count/dead-block cleanup out of Lowering::DoPhase into a new Compiler::fgPostLowering phase is a clean mechanical move that preserves the original ordering and semantics, and correctly runs after the new reuse phase. The guards are careful: it bails when locals are not enregistered, skips all-true masks whose live range crosses a call, respects contained masks by poisoning the whole group, and preserves the ConditionalSelect(AllTrue, embedded-op, zero) movprfx-free codegen special case. Pattern normalization (LargestPowerOf2 -> All, pfalse sentinel keyed as .b) matches how these lower.

Summary: ✅ LGTM. The change is well-structured, guarded conservatively, and already carries an approval from a JIT maintainer (tannergooding). The refactor preserves prior post-lowering semantics, and the added FileCheck-based test covers a good spread of true/false/pattern/embedded/conversion/SVE2 cases plus single-vs-multiple reuse scenarios. One minor non-blocking observation: the new IsSveBreakMask helper in hwintrinsic.h is dead code with no callers (flagged inline). Nothing here blocks merge.

Note

This review was generated by this repository's Holistic Review agentic workflow to complement the built-in Copilot review.

Generated by Holistic Review · 146.3 AIC · ⌖ 10.5 AIC · ⊞ 10K

}
}

static bool IsSveBreakMask(NamedIntrinsic id)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 IsSveBreakMask is added here but never referenced anywhere in the JIT (IsSveCreateTrueMask above it is used by constantreuse.cpp, but this helper has no callers). If it is intended for a follow-up, a brief comment noting that would help; otherwise consider dropping it to avoid dead code.

@jonathandavies-arm

Copy link
Copy Markdown
ContributorAuthor

I'm going to close this PR and the LSRA version of the work is here #131309.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jonathandavies-arm@tannergooding@jakobbotsch
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' arm64: Add SVE ptrue reuse pass after lowering by jonathandavies-arm · Pull Request #128844 · dotnet/runtime · GitHub
Skip to content

arm64: Add SVE ptrue reuse pass after lowering - #128844

Closed
jonathandavies-arm wants to merge 7 commits into
dotnet:mainfrom
jonathandavies-arm:upstream/sve/ptrue-reuse
Closed

arm64: Add SVE ptrue reuse pass after lowering#128844
jonathandavies-arm wants to merge 7 commits into
dotnet:mainfrom
jonathandavies-arm:upstream/sve/ptrue-reuse

Conversation

@jonathandavies-arm

Copy link
Copy Markdown
Contributor
  • Reuses equivalent ptrue producing nodes within a block via a mask temp

- Reuses equivalent ptrue-producing nodes within a block via a mask temp
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jun 1, 2026
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jun 1, 2026
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/jit/compiler.cpp Outdated
Comment threadsrc/coreclr/jit/CMakeLists.txt Outdated
Comment threadsrc/coreclr/jit/constantmaskreuse.cpp Outdated
@tannergooding

Copy link
Copy Markdown
Member

The logic here looks overall correct to me. However, diffs are not looking very good.

Rather this comes at a significant throughput hit to Arm64 (+0.54% in MinOpts) and so likely, at a minimum, needs to be skipped if optimizations are not enabled.

Then, this actually doesn't appear to be improving codegen either. Instead, it looks to significantly regress codegen (+199k bytes), particularly for MinOpts where we get a bunch of ldr, add, ldr sequences instead of ptrue or reusing an existing constant.

If you avoid running this for T0 code, then we'll still have +65k bytes of new codegen, so I imagine there is still something to be fixed/improved here. But, those diffs are generally hard to find given how much T0 code regressed and so you can't really identify what the issue is or if its the same ldr, add, ldr issue.

@tannergooding

tannergooding commented Jun 25, 2026

Copy link
Copy Markdown
Member

diffs look better now with everything but the coreclr_tests showing improvements.

Several of the tests still show regressions (+10k bytes of codegen), however, where we effectively have the following instead:

 ptrue p0.s
+ add xip1, fp, #24+ str p0, [xip1]	// [V33 rat0]
...
- ptrue p0.s+ add xip1, fp, #24+ ldr p0, [xip1]	// [V33 rat0]

Is there some particular pattern here that we're failing to account for, such as some call that is forcing a spill between the two callsites or perhaps the local being annotated in a way that forces a spill?

@tannergoodingtannergooding left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Changes LGTM and diffs are now purely positive.

There's still a TP hit, but I doubt that can be mitigated more as its essentially just the cost of adding a phase, whether that phase is executed or not.

@dotnet/jit-contrib, @jakobbotsch, @EgorBo for secondary review and input on whether such a phase is worth having.

This would notably be extendable to AllBitsSet and Zero for general SIMD or floating-point constants on all platforms (Arm64 and xarch) so isn't SVE specific and would greatly mitigate some of the issues we see in #70182 (which is a complex LSRA issue)

@jakobbotsch

Copy link
Copy Markdown
Member

@dotnet/jit-contrib, @jakobbotsch, @EgorBo for secondary review and input on whether such a phase is worth having.

I would much rather see the effort spent on improving the constant reuse that LSRA already implements. This implementation seems to come with some rather severe limitations (single block only and impoverished reasoning about register kills being some of them), in addition to being a new phase that solves a problem we already try to solve.

@jonathandavies-arm Did you look into expanding LSRA's constant reuse and investigate out why some of the cases this PR handles are not handled by it?

@github-actions

github-actionsBot commented Jul 19, 2026

Copy link
Copy Markdown
Contributor

Workflow state for the Holistic Review Orchestrator.

{
"version": 5,
"last_dispatched_commit": "8d3f5534cb1ee5dba6f1fd30e30fc8c97674f322",
"last_dispatched_base_ref": "main",
"last_dispatched_base_sha": "e137b2c018221d7babd79162f0fd151367c34c58",
"last_reviewed_commit": "8d3f5534cb1ee5dba6f1fd30e30fc8c97674f322",
"last_reviewed_base_ref": "main",
"last_reviewed_base_sha": "e137b2c018221d7babd79162f0fd151367c34c58",
"last_recorded_worker_run_id": "29684000455",
"review_attempt_commit": "",
"review_attempt_base_ref": "",
"review_attempt_count": 0,
"max_review_attempts": 5,
"review_history_format": "holistic-review-disclosure-v1",
"review_history": [
{
"commit": "8d3f5534cb1ee5dba6f1fd30e30fc8c97674f322",
"review_id": 4730635250
}
]
}

@github-actionsgithub-actionsBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Holistic Review

Motivation: The problem is real. On arm64/SVE, predicate-producing constant masks (ptrue/pfalse) are frequently rematerialized multiple times within a block, and the JIT had no post-lowering mechanism to share an equivalent predicate via a temp. Reducing redundant ptrue/pfalse materialization is a legitimate codegen win.

Approach: The approach is sound and appropriately conservative. It introduces a new post-lowering PHASE_CONSTANT_REUSE (arm64 + FEATURE_MASKED_HW_INTRINSICS, optimizations only), materializing the first equivalent constant-mask into a local temp and rewriting later equivalent uses as loads, relying on normal local-var lifetimes so LSRA models predicate clobbers. The refactor that hoists the post-lowering liveness/ref-count/dead-block cleanup out of Lowering::DoPhase into a new Compiler::fgPostLowering phase is a clean mechanical move that preserves the original ordering and semantics, and correctly runs after the new reuse phase. The guards are careful: it bails when locals are not enregistered, skips all-true masks whose live range crosses a call, respects contained masks by poisoning the whole group, and preserves the ConditionalSelect(AllTrue, embedded-op, zero) movprfx-free codegen special case. Pattern normalization (LargestPowerOf2 -> All, pfalse sentinel keyed as .b) matches how these lower.

Summary: ✅ LGTM. The change is well-structured, guarded conservatively, and already carries an approval from a JIT maintainer (tannergooding). The refactor preserves prior post-lowering semantics, and the added FileCheck-based test covers a good spread of true/false/pattern/embedded/conversion/SVE2 cases plus single-vs-multiple reuse scenarios. One minor non-blocking observation: the new IsSveBreakMask helper in hwintrinsic.h is dead code with no callers (flagged inline). Nothing here blocks merge.

Note

This review was generated by this repository's Holistic Review agentic workflow to complement the built-in Copilot review.

Generated by Holistic Review · 146.3 AIC · ⌖ 10.5 AIC · ⊞ 10K

}
}

static bool IsSveBreakMask(NamedIntrinsic id)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 IsSveBreakMask is added here but never referenced anywhere in the JIT (IsSveCreateTrueMask above it is used by constantreuse.cpp, but this helper has no callers). If it is intended for a follow-up, a brief comment noting that would help; otherwise consider dropping it to avoid dead code.

@jonathandavies-arm

Copy link
Copy Markdown
ContributorAuthor

I'm going to close this PR and the LSRA version of the work is here #131309.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jonathandavies-arm@tannergooding@jakobbotsch
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' arm64: Add SVE ptrue reuse pass after lowering by jonathandavies-arm · Pull Request #128844 · dotnet/runtime · GitHub
Skip to content

arm64: Add SVE ptrue reuse pass after lowering - #128844

Closed
jonathandavies-arm wants to merge 7 commits into
dotnet:mainfrom
jonathandavies-arm:upstream/sve/ptrue-reuse
Closed

arm64: Add SVE ptrue reuse pass after lowering#128844
jonathandavies-arm wants to merge 7 commits into
dotnet:mainfrom
jonathandavies-arm:upstream/sve/ptrue-reuse

Conversation

@jonathandavies-arm

Copy link
Copy Markdown
Contributor
  • Reuses equivalent ptrue producing nodes within a block via a mask temp

- Reuses equivalent ptrue-producing nodes within a block via a mask temp
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jun 1, 2026
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jun 1, 2026
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/jit/compiler.cpp Outdated
Comment threadsrc/coreclr/jit/CMakeLists.txt Outdated
Comment threadsrc/coreclr/jit/constantmaskreuse.cpp Outdated
@tannergooding

Copy link
Copy Markdown
Member

The logic here looks overall correct to me. However, diffs are not looking very good.

Rather this comes at a significant throughput hit to Arm64 (+0.54% in MinOpts) and so likely, at a minimum, needs to be skipped if optimizations are not enabled.

Then, this actually doesn't appear to be improving codegen either. Instead, it looks to significantly regress codegen (+199k bytes), particularly for MinOpts where we get a bunch of ldr, add, ldr sequences instead of ptrue or reusing an existing constant.

If you avoid running this for T0 code, then we'll still have +65k bytes of new codegen, so I imagine there is still something to be fixed/improved here. But, those diffs are generally hard to find given how much T0 code regressed and so you can't really identify what the issue is or if its the same ldr, add, ldr issue.

@tannergooding

tannergooding commented Jun 25, 2026

Copy link
Copy Markdown
Member

diffs look better now with everything but the coreclr_tests showing improvements.

Several of the tests still show regressions (+10k bytes of codegen), however, where we effectively have the following instead:

 ptrue p0.s
+ add xip1, fp, #24+ str p0, [xip1]	// [V33 rat0]
...
- ptrue p0.s+ add xip1, fp, #24+ ldr p0, [xip1]	// [V33 rat0]

Is there some particular pattern here that we're failing to account for, such as some call that is forcing a spill between the two callsites or perhaps the local being annotated in a way that forces a spill?

@tannergoodingtannergooding left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Changes LGTM and diffs are now purely positive.

There's still a TP hit, but I doubt that can be mitigated more as its essentially just the cost of adding a phase, whether that phase is executed or not.

@dotnet/jit-contrib, @jakobbotsch, @EgorBo for secondary review and input on whether such a phase is worth having.

This would notably be extendable to AllBitsSet and Zero for general SIMD or floating-point constants on all platforms (Arm64 and xarch) so isn't SVE specific and would greatly mitigate some of the issues we see in #70182 (which is a complex LSRA issue)

@jakobbotsch

Copy link
Copy Markdown
Member

@dotnet/jit-contrib, @jakobbotsch, @EgorBo for secondary review and input on whether such a phase is worth having.

I would much rather see the effort spent on improving the constant reuse that LSRA already implements. This implementation seems to come with some rather severe limitations (single block only and impoverished reasoning about register kills being some of them), in addition to being a new phase that solves a problem we already try to solve.

@jonathandavies-arm Did you look into expanding LSRA's constant reuse and investigate out why some of the cases this PR handles are not handled by it?

@github-actions

github-actionsBot commented Jul 19, 2026

Copy link
Copy Markdown
Contributor

Workflow state for the Holistic Review Orchestrator.

{
"version": 5,
"last_dispatched_commit": "8d3f5534cb1ee5dba6f1fd30e30fc8c97674f322",
"last_dispatched_base_ref": "main",
"last_dispatched_base_sha": "e137b2c018221d7babd79162f0fd151367c34c58",
"last_reviewed_commit": "8d3f5534cb1ee5dba6f1fd30e30fc8c97674f322",
"last_reviewed_base_ref": "main",
"last_reviewed_base_sha": "e137b2c018221d7babd79162f0fd151367c34c58",
"last_recorded_worker_run_id": "29684000455",
"review_attempt_commit": "",
"review_attempt_base_ref": "",
"review_attempt_count": 0,
"max_review_attempts": 5,
"review_history_format": "holistic-review-disclosure-v1",
"review_history": [
{
"commit": "8d3f5534cb1ee5dba6f1fd30e30fc8c97674f322",
"review_id": 4730635250
}
]
}

@github-actionsgithub-actionsBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Holistic Review

Motivation: The problem is real. On arm64/SVE, predicate-producing constant masks (ptrue/pfalse) are frequently rematerialized multiple times within a block, and the JIT had no post-lowering mechanism to share an equivalent predicate via a temp. Reducing redundant ptrue/pfalse materialization is a legitimate codegen win.

Approach: The approach is sound and appropriately conservative. It introduces a new post-lowering PHASE_CONSTANT_REUSE (arm64 + FEATURE_MASKED_HW_INTRINSICS, optimizations only), materializing the first equivalent constant-mask into a local temp and rewriting later equivalent uses as loads, relying on normal local-var lifetimes so LSRA models predicate clobbers. The refactor that hoists the post-lowering liveness/ref-count/dead-block cleanup out of Lowering::DoPhase into a new Compiler::fgPostLowering phase is a clean mechanical move that preserves the original ordering and semantics, and correctly runs after the new reuse phase. The guards are careful: it bails when locals are not enregistered, skips all-true masks whose live range crosses a call, respects contained masks by poisoning the whole group, and preserves the ConditionalSelect(AllTrue, embedded-op, zero) movprfx-free codegen special case. Pattern normalization (LargestPowerOf2 -> All, pfalse sentinel keyed as .b) matches how these lower.

Summary: ✅ LGTM. The change is well-structured, guarded conservatively, and already carries an approval from a JIT maintainer (tannergooding). The refactor preserves prior post-lowering semantics, and the added FileCheck-based test covers a good spread of true/false/pattern/embedded/conversion/SVE2 cases plus single-vs-multiple reuse scenarios. One minor non-blocking observation: the new IsSveBreakMask helper in hwintrinsic.h is dead code with no callers (flagged inline). Nothing here blocks merge.

Note

This review was generated by this repository's Holistic Review agentic workflow to complement the built-in Copilot review.

Generated by Holistic Review · 146.3 AIC · ⌖ 10.5 AIC · ⊞ 10K

}
}

static bool IsSveBreakMask(NamedIntrinsic id)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 IsSveBreakMask is added here but never referenced anywhere in the JIT (IsSveCreateTrueMask above it is used by constantreuse.cpp, but this helper has no callers). If it is intended for a follow-up, a brief comment noting that would help; otherwise consider dropping it to avoid dead code.

@jonathandavies-arm

Copy link
Copy Markdown
ContributorAuthor

I'm going to close this PR and the LSRA version of the work is here #131309.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jonathandavies-arm@tannergooding@jakobbotsch
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' arm64: Add SVE ptrue reuse pass after lowering by jonathandavies-arm · Pull Request #128844 · dotnet/runtime · GitHub
Skip to content

arm64: Add SVE ptrue reuse pass after lowering - #128844

Closed
jonathandavies-arm wants to merge 7 commits into
dotnet:mainfrom
jonathandavies-arm:upstream/sve/ptrue-reuse
Closed

arm64: Add SVE ptrue reuse pass after lowering#128844
jonathandavies-arm wants to merge 7 commits into
dotnet:mainfrom
jonathandavies-arm:upstream/sve/ptrue-reuse

Conversation

@jonathandavies-arm

Copy link
Copy Markdown
Contributor
  • Reuses equivalent ptrue producing nodes within a block via a mask temp

- Reuses equivalent ptrue-producing nodes within a block via a mask temp
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jun 1, 2026
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jun 1, 2026
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/jit/compiler.cpp Outdated
Comment threadsrc/coreclr/jit/CMakeLists.txt Outdated
Comment threadsrc/coreclr/jit/constantmaskreuse.cpp Outdated
@tannergooding

Copy link
Copy Markdown
Member

The logic here looks overall correct to me. However, diffs are not looking very good.

Rather this comes at a significant throughput hit to Arm64 (+0.54% in MinOpts) and so likely, at a minimum, needs to be skipped if optimizations are not enabled.

Then, this actually doesn't appear to be improving codegen either. Instead, it looks to significantly regress codegen (+199k bytes), particularly for MinOpts where we get a bunch of ldr, add, ldr sequences instead of ptrue or reusing an existing constant.

If you avoid running this for T0 code, then we'll still have +65k bytes of new codegen, so I imagine there is still something to be fixed/improved here. But, those diffs are generally hard to find given how much T0 code regressed and so you can't really identify what the issue is or if its the same ldr, add, ldr issue.

@tannergooding

tannergooding commented Jun 25, 2026

Copy link
Copy Markdown
Member

diffs look better now with everything but the coreclr_tests showing improvements.

Several of the tests still show regressions (+10k bytes of codegen), however, where we effectively have the following instead:

 ptrue p0.s
+ add xip1, fp, #24+ str p0, [xip1]	// [V33 rat0]
...
- ptrue p0.s+ add xip1, fp, #24+ ldr p0, [xip1]	// [V33 rat0]

Is there some particular pattern here that we're failing to account for, such as some call that is forcing a spill between the two callsites or perhaps the local being annotated in a way that forces a spill?

@tannergoodingtannergooding left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Changes LGTM and diffs are now purely positive.

There's still a TP hit, but I doubt that can be mitigated more as its essentially just the cost of adding a phase, whether that phase is executed or not.

@dotnet/jit-contrib, @jakobbotsch, @EgorBo for secondary review and input on whether such a phase is worth having.

This would notably be extendable to AllBitsSet and Zero for general SIMD or floating-point constants on all platforms (Arm64 and xarch) so isn't SVE specific and would greatly mitigate some of the issues we see in #70182 (which is a complex LSRA issue)

@jakobbotsch

Copy link
Copy Markdown
Member

@dotnet/jit-contrib, @jakobbotsch, @EgorBo for secondary review and input on whether such a phase is worth having.

I would much rather see the effort spent on improving the constant reuse that LSRA already implements. This implementation seems to come with some rather severe limitations (single block only and impoverished reasoning about register kills being some of them), in addition to being a new phase that solves a problem we already try to solve.

@jonathandavies-arm Did you look into expanding LSRA's constant reuse and investigate out why some of the cases this PR handles are not handled by it?

@github-actions

github-actionsBot commented Jul 19, 2026

Copy link
Copy Markdown
Contributor

Workflow state for the Holistic Review Orchestrator.

{
"version": 5,
"last_dispatched_commit": "8d3f5534cb1ee5dba6f1fd30e30fc8c97674f322",
"last_dispatched_base_ref": "main",
"last_dispatched_base_sha": "e137b2c018221d7babd79162f0fd151367c34c58",
"last_reviewed_commit": "8d3f5534cb1ee5dba6f1fd30e30fc8c97674f322",
"last_reviewed_base_ref": "main",
"last_reviewed_base_sha": "e137b2c018221d7babd79162f0fd151367c34c58",
"last_recorded_worker_run_id": "29684000455",
"review_attempt_commit": "",
"review_attempt_base_ref": "",
"review_attempt_count": 0,
"max_review_attempts": 5,
"review_history_format": "holistic-review-disclosure-v1",
"review_history": [
{
"commit": "8d3f5534cb1ee5dba6f1fd30e30fc8c97674f322",
"review_id": 4730635250
}
]
}

@github-actionsgithub-actionsBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Holistic Review

Motivation: The problem is real. On arm64/SVE, predicate-producing constant masks (ptrue/pfalse) are frequently rematerialized multiple times within a block, and the JIT had no post-lowering mechanism to share an equivalent predicate via a temp. Reducing redundant ptrue/pfalse materialization is a legitimate codegen win.

Approach: The approach is sound and appropriately conservative. It introduces a new post-lowering PHASE_CONSTANT_REUSE (arm64 + FEATURE_MASKED_HW_INTRINSICS, optimizations only), materializing the first equivalent constant-mask into a local temp and rewriting later equivalent uses as loads, relying on normal local-var lifetimes so LSRA models predicate clobbers. The refactor that hoists the post-lowering liveness/ref-count/dead-block cleanup out of Lowering::DoPhase into a new Compiler::fgPostLowering phase is a clean mechanical move that preserves the original ordering and semantics, and correctly runs after the new reuse phase. The guards are careful: it bails when locals are not enregistered, skips all-true masks whose live range crosses a call, respects contained masks by poisoning the whole group, and preserves the ConditionalSelect(AllTrue, embedded-op, zero) movprfx-free codegen special case. Pattern normalization (LargestPowerOf2 -> All, pfalse sentinel keyed as .b) matches how these lower.

Summary: ✅ LGTM. The change is well-structured, guarded conservatively, and already carries an approval from a JIT maintainer (tannergooding). The refactor preserves prior post-lowering semantics, and the added FileCheck-based test covers a good spread of true/false/pattern/embedded/conversion/SVE2 cases plus single-vs-multiple reuse scenarios. One minor non-blocking observation: the new IsSveBreakMask helper in hwintrinsic.h is dead code with no callers (flagged inline). Nothing here blocks merge.

Note

This review was generated by this repository's Holistic Review agentic workflow to complement the built-in Copilot review.

Generated by Holistic Review · 146.3 AIC · ⌖ 10.5 AIC · ⊞ 10K

}
}

static bool IsSveBreakMask(NamedIntrinsic id)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 IsSveBreakMask is added here but never referenced anywhere in the JIT (IsSveCreateTrueMask above it is used by constantreuse.cpp, but this helper has no callers). If it is intended for a follow-up, a brief comment noting that would help; otherwise consider dropping it to avoid dead code.

@jonathandavies-arm

Copy link
Copy Markdown
ContributorAuthor

I'm going to close this PR and the LSRA version of the work is here #131309.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jonathandavies-arm@tannergooding@jakobbotsch
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' arm64: Add SVE ptrue reuse pass after lowering by jonathandavies-arm · Pull Request #128844 · dotnet/runtime · GitHub
Skip to content

arm64: Add SVE ptrue reuse pass after lowering - #128844

Closed
jonathandavies-arm wants to merge 7 commits into
dotnet:mainfrom
jonathandavies-arm:upstream/sve/ptrue-reuse
Closed

arm64: Add SVE ptrue reuse pass after lowering#128844
jonathandavies-arm wants to merge 7 commits into
dotnet:mainfrom
jonathandavies-arm:upstream/sve/ptrue-reuse

Conversation

@jonathandavies-arm

Copy link
Copy Markdown
Contributor
  • Reuses equivalent ptrue producing nodes within a block via a mask temp

- Reuses equivalent ptrue-producing nodes within a block via a mask temp
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jun 1, 2026
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jun 1, 2026
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/jit/compiler.cpp Outdated
Comment threadsrc/coreclr/jit/CMakeLists.txt Outdated
Comment threadsrc/coreclr/jit/constantmaskreuse.cpp Outdated
@tannergooding

Copy link
Copy Markdown
Member

The logic here looks overall correct to me. However, diffs are not looking very good.

Rather this comes at a significant throughput hit to Arm64 (+0.54% in MinOpts) and so likely, at a minimum, needs to be skipped if optimizations are not enabled.

Then, this actually doesn't appear to be improving codegen either. Instead, it looks to significantly regress codegen (+199k bytes), particularly for MinOpts where we get a bunch of ldr, add, ldr sequences instead of ptrue or reusing an existing constant.

If you avoid running this for T0 code, then we'll still have +65k bytes of new codegen, so I imagine there is still something to be fixed/improved here. But, those diffs are generally hard to find given how much T0 code regressed and so you can't really identify what the issue is or if its the same ldr, add, ldr issue.

@tannergooding

tannergooding commented Jun 25, 2026

Copy link
Copy Markdown
Member

diffs look better now with everything but the coreclr_tests showing improvements.

Several of the tests still show regressions (+10k bytes of codegen), however, where we effectively have the following instead:

 ptrue p0.s
+ add xip1, fp, #24+ str p0, [xip1]	// [V33 rat0]
...
- ptrue p0.s+ add xip1, fp, #24+ ldr p0, [xip1]	// [V33 rat0]

Is there some particular pattern here that we're failing to account for, such as some call that is forcing a spill between the two callsites or perhaps the local being annotated in a way that forces a spill?

@tannergoodingtannergooding left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Changes LGTM and diffs are now purely positive.

There's still a TP hit, but I doubt that can be mitigated more as its essentially just the cost of adding a phase, whether that phase is executed or not.

@dotnet/jit-contrib, @jakobbotsch, @EgorBo for secondary review and input on whether such a phase is worth having.

This would notably be extendable to AllBitsSet and Zero for general SIMD or floating-point constants on all platforms (Arm64 and xarch) so isn't SVE specific and would greatly mitigate some of the issues we see in #70182 (which is a complex LSRA issue)

@jakobbotsch

Copy link
Copy Markdown
Member

@dotnet/jit-contrib, @jakobbotsch, @EgorBo for secondary review and input on whether such a phase is worth having.

I would much rather see the effort spent on improving the constant reuse that LSRA already implements. This implementation seems to come with some rather severe limitations (single block only and impoverished reasoning about register kills being some of them), in addition to being a new phase that solves a problem we already try to solve.

@jonathandavies-arm Did you look into expanding LSRA's constant reuse and investigate out why some of the cases this PR handles are not handled by it?

@github-actions

github-actionsBot commented Jul 19, 2026

Copy link
Copy Markdown
Contributor

Workflow state for the Holistic Review Orchestrator.

{
"version": 5,
"last_dispatched_commit": "8d3f5534cb1ee5dba6f1fd30e30fc8c97674f322",
"last_dispatched_base_ref": "main",
"last_dispatched_base_sha": "e137b2c018221d7babd79162f0fd151367c34c58",
"last_reviewed_commit": "8d3f5534cb1ee5dba6f1fd30e30fc8c97674f322",
"last_reviewed_base_ref": "main",
"last_reviewed_base_sha": "e137b2c018221d7babd79162f0fd151367c34c58",
"last_recorded_worker_run_id": "29684000455",
"review_attempt_commit": "",
"review_attempt_base_ref": "",
"review_attempt_count": 0,
"max_review_attempts": 5,
"review_history_format": "holistic-review-disclosure-v1",
"review_history": [
{
"commit": "8d3f5534cb1ee5dba6f1fd30e30fc8c97674f322",
"review_id": 4730635250
}
]
}

@github-actionsgithub-actionsBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Holistic Review

Motivation: The problem is real. On arm64/SVE, predicate-producing constant masks (ptrue/pfalse) are frequently rematerialized multiple times within a block, and the JIT had no post-lowering mechanism to share an equivalent predicate via a temp. Reducing redundant ptrue/pfalse materialization is a legitimate codegen win.

Approach: The approach is sound and appropriately conservative. It introduces a new post-lowering PHASE_CONSTANT_REUSE (arm64 + FEATURE_MASKED_HW_INTRINSICS, optimizations only), materializing the first equivalent constant-mask into a local temp and rewriting later equivalent uses as loads, relying on normal local-var lifetimes so LSRA models predicate clobbers. The refactor that hoists the post-lowering liveness/ref-count/dead-block cleanup out of Lowering::DoPhase into a new Compiler::fgPostLowering phase is a clean mechanical move that preserves the original ordering and semantics, and correctly runs after the new reuse phase. The guards are careful: it bails when locals are not enregistered, skips all-true masks whose live range crosses a call, respects contained masks by poisoning the whole group, and preserves the ConditionalSelect(AllTrue, embedded-op, zero) movprfx-free codegen special case. Pattern normalization (LargestPowerOf2 -> All, pfalse sentinel keyed as .b) matches how these lower.

Summary: ✅ LGTM. The change is well-structured, guarded conservatively, and already carries an approval from a JIT maintainer (tannergooding). The refactor preserves prior post-lowering semantics, and the added FileCheck-based test covers a good spread of true/false/pattern/embedded/conversion/SVE2 cases plus single-vs-multiple reuse scenarios. One minor non-blocking observation: the new IsSveBreakMask helper in hwintrinsic.h is dead code with no callers (flagged inline). Nothing here blocks merge.

Note

This review was generated by this repository's Holistic Review agentic workflow to complement the built-in Copilot review.

Generated by Holistic Review · 146.3 AIC · ⌖ 10.5 AIC · ⊞ 10K

}
}

static bool IsSveBreakMask(NamedIntrinsic id)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 IsSveBreakMask is added here but never referenced anywhere in the JIT (IsSveCreateTrueMask above it is used by constantreuse.cpp, but this helper has no callers). If it is intended for a follow-up, a brief comment noting that would help; otherwise consider dropping it to avoid dead code.

@jonathandavies-arm

Copy link
Copy Markdown
ContributorAuthor

I'm going to close this PR and the LSRA version of the work is here #131309.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jonathandavies-arm@tannergooding@jakobbotsch
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); arm64: Add SVE ptrue reuse pass after lowering by jonathandavies-arm · Pull Request #128844 · dotnet/runtime · GitHub
Skip to content

arm64: Add SVE ptrue reuse pass after lowering - #128844

Closed
jonathandavies-arm wants to merge 7 commits into
dotnet:mainfrom
jonathandavies-arm:upstream/sve/ptrue-reuse
Closed

arm64: Add SVE ptrue reuse pass after lowering#128844
jonathandavies-arm wants to merge 7 commits into
dotnet:mainfrom
jonathandavies-arm:upstream/sve/ptrue-reuse

Conversation

@jonathandavies-arm

Copy link
Copy Markdown
Contributor
  • Reuses equivalent ptrue producing nodes within a block via a mask temp

- Reuses equivalent ptrue-producing nodes within a block via a mask temp
@github-actionsgithub-actionsBot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jun 1, 2026
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jun 1, 2026
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

Comment threadsrc/coreclr/jit/compiler.cpp Outdated
Comment threadsrc/coreclr/jit/CMakeLists.txt Outdated
Comment threadsrc/coreclr/jit/constantmaskreuse.cpp Outdated
@tannergooding

Copy link
Copy Markdown
Member

The logic here looks overall correct to me. However, diffs are not looking very good.

Rather this comes at a significant throughput hit to Arm64 (+0.54% in MinOpts) and so likely, at a minimum, needs to be skipped if optimizations are not enabled.

Then, this actually doesn't appear to be improving codegen either. Instead, it looks to significantly regress codegen (+199k bytes), particularly for MinOpts where we get a bunch of ldr, add, ldr sequences instead of ptrue or reusing an existing constant.

If you avoid running this for T0 code, then we'll still have +65k bytes of new codegen, so I imagine there is still something to be fixed/improved here. But, those diffs are generally hard to find given how much T0 code regressed and so you can't really identify what the issue is or if its the same ldr, add, ldr issue.

@tannergooding

tannergooding commented Jun 25, 2026

Copy link
Copy Markdown
Member

diffs look better now with everything but the coreclr_tests showing improvements.

Several of the tests still show regressions (+10k bytes of codegen), however, where we effectively have the following instead:

 ptrue p0.s
+ add xip1, fp, #24+ str p0, [xip1]	// [V33 rat0]
...
- ptrue p0.s+ add xip1, fp, #24+ ldr p0, [xip1]	// [V33 rat0]

Is there some particular pattern here that we're failing to account for, such as some call that is forcing a spill between the two callsites or perhaps the local being annotated in a way that forces a spill?

@tannergoodingtannergooding left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Changes LGTM and diffs are now purely positive.

There's still a TP hit, but I doubt that can be mitigated more as its essentially just the cost of adding a phase, whether that phase is executed or not.

@dotnet/jit-contrib, @jakobbotsch, @EgorBo for secondary review and input on whether such a phase is worth having.

This would notably be extendable to AllBitsSet and Zero for general SIMD or floating-point constants on all platforms (Arm64 and xarch) so isn't SVE specific and would greatly mitigate some of the issues we see in #70182 (which is a complex LSRA issue)

@jakobbotsch

Copy link
Copy Markdown
Member

@dotnet/jit-contrib, @jakobbotsch, @EgorBo for secondary review and input on whether such a phase is worth having.

I would much rather see the effort spent on improving the constant reuse that LSRA already implements. This implementation seems to come with some rather severe limitations (single block only and impoverished reasoning about register kills being some of them), in addition to being a new phase that solves a problem we already try to solve.

@jonathandavies-arm Did you look into expanding LSRA's constant reuse and investigate out why some of the cases this PR handles are not handled by it?

@github-actions

github-actionsBot commented Jul 19, 2026

Copy link
Copy Markdown
Contributor

Workflow state for the Holistic Review Orchestrator.

{
"version": 5,
"last_dispatched_commit": "8d3f5534cb1ee5dba6f1fd30e30fc8c97674f322",
"last_dispatched_base_ref": "main",
"last_dispatched_base_sha": "e137b2c018221d7babd79162f0fd151367c34c58",
"last_reviewed_commit": "8d3f5534cb1ee5dba6f1fd30e30fc8c97674f322",
"last_reviewed_base_ref": "main",
"last_reviewed_base_sha": "e137b2c018221d7babd79162f0fd151367c34c58",
"last_recorded_worker_run_id": "29684000455",
"review_attempt_commit": "",
"review_attempt_base_ref": "",
"review_attempt_count": 0,
"max_review_attempts": 5,
"review_history_format": "holistic-review-disclosure-v1",
"review_history": [
{
"commit": "8d3f5534cb1ee5dba6f1fd30e30fc8c97674f322",
"review_id": 4730635250
}
]
}

@github-actionsgithub-actionsBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Holistic Review

Motivation: The problem is real. On arm64/SVE, predicate-producing constant masks (ptrue/pfalse) are frequently rematerialized multiple times within a block, and the JIT had no post-lowering mechanism to share an equivalent predicate via a temp. Reducing redundant ptrue/pfalse materialization is a legitimate codegen win.

Approach: The approach is sound and appropriately conservative. It introduces a new post-lowering PHASE_CONSTANT_REUSE (arm64 + FEATURE_MASKED_HW_INTRINSICS, optimizations only), materializing the first equivalent constant-mask into a local temp and rewriting later equivalent uses as loads, relying on normal local-var lifetimes so LSRA models predicate clobbers. The refactor that hoists the post-lowering liveness/ref-count/dead-block cleanup out of Lowering::DoPhase into a new Compiler::fgPostLowering phase is a clean mechanical move that preserves the original ordering and semantics, and correctly runs after the new reuse phase. The guards are careful: it bails when locals are not enregistered, skips all-true masks whose live range crosses a call, respects contained masks by poisoning the whole group, and preserves the ConditionalSelect(AllTrue, embedded-op, zero) movprfx-free codegen special case. Pattern normalization (LargestPowerOf2 -> All, pfalse sentinel keyed as .b) matches how these lower.

Summary: ✅ LGTM. The change is well-structured, guarded conservatively, and already carries an approval from a JIT maintainer (tannergooding). The refactor preserves prior post-lowering semantics, and the added FileCheck-based test covers a good spread of true/false/pattern/embedded/conversion/SVE2 cases plus single-vs-multiple reuse scenarios. One minor non-blocking observation: the new IsSveBreakMask helper in hwintrinsic.h is dead code with no callers (flagged inline). Nothing here blocks merge.

Note

This review was generated by this repository's Holistic Review agentic workflow to complement the built-in Copilot review.

Generated by Holistic Review · 146.3 AIC · ⌖ 10.5 AIC · ⊞ 10K

}
}

static bool IsSveBreakMask(NamedIntrinsic id)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 IsSveBreakMask is added here but never referenced anywhere in the JIT (IsSveCreateTrueMask above it is used by constantreuse.cpp, but this helper has no callers). If it is intended for a follow-up, a brief comment noting that would help; otherwise consider dropping it to avoid dead code.

@jonathandavies-arm

Copy link
Copy Markdown
ContributorAuthor

I'm going to close this PR and the LSRA version of the work is here #131309.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jonathandavies-arm@tannergooding@jakobbotsch