JIT: Extend escape analysis to account for arrays with non-gcref elements - #104906

Merged
AndyAyersMS merged 97 commits into
dotnet:mainfrom
hez2010:value-array-stack-alloc
Jan 22, 2025
Merged

JIT: Extend escape analysis to account for arrays with non-gcref elements#104906
AndyAyersMS merged 97 commits into
dotnet:mainfrom
hez2010:value-array-stack-alloc

Conversation

@hez2010

@hez2010hez2010 commented Jul 15, 2024

Copy link
Copy Markdown
Contributor

Positive case:

varchs=newchar[42];chs[1]='a';Console.WriteLine((int)chs[1]+chs.Length);

Codegen:

; Assembly listing for method ArrayAllocator.Program:Main() (FullOpts); Emitting BLENDED_CODE for X64 with AVX - Windows; FullOpts code; optimized code; rsp based frame; partially interruptible; No PGO data; Final local variable assignments;;* V00 loc0 [V00 ] ( 0, 0 ) long -> zero-ref class-hnd exact <short[]>; V01 OutArgs [V01 ] ( 1, 1 ) struct (32) [rsp+0x00] do-not-enreg[XS] addr-exposed "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) struct (104) zero-ref do-not-enreg[SF] "stack allocated array temp";* V03 tmp2 [V03 ] ( 0, 0 ) long -> zero-ref single-def "V02.[000..008)";* V04 tmp3 [V04 ] ( 0, 0 ) int -> zero-ref single-def "V02.[008..012)";* V05 tmp4 [V05 ] ( 0, 0 ) short -> zero-ref "V02.[018..020)";; Lcl frame size = 40G_M25548_IG01: ;; offset=0x0000subrsp,40 ;; size=4 bbWeight=1 PerfScore 0.25G_M25548_IG02: ;; offset=0x0004movecx,84call[System.Console:WriteLine(int)]nop ;; size=12 bbWeight=1 PerfScore 3.50G_M25548_IG03: ;; offset=0x0010addrsp,40ret ;; size=5 bbWeight=1 PerfScore 1.25; Total bytes of code 21, prolog size 4, PerfScore 5.00, instruction count 6, allocated bytes for code 21 (MethodHash=5b0b9c33) for method ArrayAllocator.Program:Main() (FullOpts)

Negative case:

varchs=newchar[42];chs[1]='a';Console.WriteLine((int)chs[42]+chs.Length);

Codegen:

; Assembly listing for method ArrayAllocator.Program:Main() (FullOpts); Emitting BLENDED_CODE for X64 with AVX - Windows; FullOpts code; optimized code; rsp based frame; partially interruptible; No PGO data; Final local variable assignments;;* V00 loc0 [V00 ] ( 0, 0 ) long -> zero-ref class-hnd exact <short[]>; V01 OutArgs [V01 ] ( 1, 1 ) struct (32) [rsp+0x00] do-not-enreg[XS] addr-exposed "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) struct (104) zero-ref do-not-enreg[SF] "stack allocated array temp"; V03 tmp2 [V03,T00] ( 1, 0 ) byref -> rbx must-init "dummy temp of must thrown exception";* V04 tmp3 [V04 ] ( 0, 0 ) long -> zero-ref single-def "V02.[000..008)";* V05 tmp4 [V05 ] ( 0, 0 ) int -> zero-ref single-def "V02.[008..012)";* V06 tmp5 [V06 ] ( 0, 0 ) short -> zero-ref single-def "V02.[018..020)";; Lcl frame size = 32G_M25548_IG01: ;; offset=0x0000pushrbxsubrsp,32xorebx,ebx ;; size=7 bbWeight=0 PerfScore 0.00G_M25548_IG02: ;; offset=0x0007call CORINFO_HELP_RNGCHKFAILmovsxrcx, word ptr [rbx]call[System.Console:WriteLine(int)]int3 ;; size=16 bbWeight=0 PerfScore 0.00; Total bytes of code 23, prolog size 5, PerfScore 0.00, instruction count 7, allocated bytes for code 23 (MethodHash=5b0b9c33) for method ArrayAllocator.Program:Main() (FullOpts); ============================================================

Benchmark on Mandelbrot:

MethodJobMeanErrorStdDevCode SizeAllocated
MandelBrotNoStackAllocationArray199.7 us1.30 us1.22 us1,996 B2.49 KB
MandelBrotStackAllocationArray195.8 us1.16 us1.08 us2,414 B1.14 KB

Diff: https://www.diffchecker.com/bNP4qHdF/

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 15, 2024
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jul 15, 2024
Comment threadsrc/coreclr/jit/objectalloc.h Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated

@AndyAyersMSAndyAyersMS left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For arrays (and also perhaps boxes and ref classes) we ought to have some kind of size limit... possibly similar to the one we use for stackallocs.

We need to be careful we don't allocate a lot of stack for an object that might not be heavily used, as we'll pay per-call prolog zeroing costs.

Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/gentree.cpp Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated
#ifdef FEATURE_READYTORUN
if (comp->opts.IsReadyToRun() && data->IsHelperCall(comp, CORINFO_HELP_READYTORUN_NEWARR_1))
{
len = data->AsCall()->gtArgs.GetArgByIndex(0)->GetNode();

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Dead code?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I removed R2R support more completely

Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
Comment on lines +2849 to +2854
GenTree* const mtStore = gtNewStoreLclFldNode(lclNum, TYP_I_IMPL, 0, mt);
Statement* const mtStmt = gtNewStmt(mtStore);

fgInsertStmtBefore(block, newStmt, mtStmt);
gtSetStmtInfo(mtStmt);
fgSetStmtSeq(mtStmt);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
GenTree* const mtStore = gtNewStoreLclFldNode(lclNum, TYP_I_IMPL, 0, mt);
Statement* const mtStmt = gtNewStmt(mtStore);
fgInsertStmtBefore(block, newStmt, mtStmt);
gtSetStmtInfo(mtStmt);
fgSetStmtSeq(mtStmt);
GenTree* const mtStore = gtNewStoreLclFldNode(lclNum, TYP_I_IMPL, 0, mt);
Statement* const mtStmt = fgNewStmtFromTree(mtStore);
fgInsertStmtBefore(block, newStmt, mtStmt);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ditto below, more things can switch to use fgNewStmtFromTree

Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
@@ -181,6 +254,11 @@ inline bool ObjectAllocator::CanAllocateLclVarOnStack(unsigned int lclNu
return false;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems like we should have a limit on the aggregate stack allocated size. That's somewhat preexisting, but probably more important now.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agree but will do that in a future PR

Comment on lines +6249 to +6260
// If this is a local array, there are no asyncronous modifications, so we can set the
// conservative VN to the liberal VN.
//
VNFuncApp arrFn;
if (vnStore->IsVNNewLocalArr(arrVN, &arrFn))
{
loadTree->gtVNPair.SetConservative(loadValueVN);
}
else
{
loadTree->gtVNPair.SetConservative(vnStore->VNForExpr(compCurBB, loadType));
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does this actually show up as benefits?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In some limited cases, yes... eg we can const propagate through a[2] = 1; y = a[2];

@AndyAyersMS

Copy link
Copy Markdown
Member

@jakobbotsch think I've addressed most of the key points. Overall size limit will come in a future PR.

@jakobbotsch

Copy link
Copy Markdown
Member

@AndyAyersMS Did you push those changes?

@AndyAyersMS

Copy link
Copy Markdown
Member

Ah, I pushed to my fork, but ... this PR is not from my fork.

@AndyAyersMS

Copy link
Copy Markdown
Member

@jakobbotsch changes are there now

GTF_CALL_M_CAST_CAN_BE_EXPANDED = 0x04000000, // this cast (helper call) can be expanded if it's profitable. To be removed.
GTF_CALL_M_CAST_OBJ_NONNULL = 0x08000000, // if we expand this specific cast we don't need to check the input object for null
// NOTE: if needed, this flag can be removed, and we can introduce new _NONNUL cast helpers
GTF_CALL_M_STACK_ARRAY = 0x10000000, // this call is a new array helper for a stack allocated array.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I suppose we weren't able to get rid of this since it's also used in VN.

Comment on lines +14375 to +14380
// modHeap = false;
}
else if (vnf == VNF_JitReadyToRunNewArr)
{
vnf = VNF_JitReadyToRunNewLclArr;
// modHeap = false;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: commented code (feel free to remove as part of a follow up)

}
}

if (isAlloc && ((call->gtCallMoreFlags & GTF_CALL_M_STACK_ARRAY) != 0))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I suppose this could be changed to FindWellKnownArg(StackArrayLocal) != nullptr to get rid of the flag

@jakobbotschjakobbotsch left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@AndyAyersMS

Copy link
Copy Markdown
Member

Thanks. I have more changes in this area so I can handle the last few bits in a subsequent PR.

@AndyAyersMS
AndyAyersMS merged commit 7d75878 into dotnet:mainJan 22, 2025
@AndyAyersMS

Copy link
Copy Markdown
Member

@hez2010 thanks for all the work you did here.

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@hez2010@AndyAyersMS@JulieLeeMSFT@jakobbotsch
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

JIT: Extend escape analysis to account for arrays with non-gcref elements - #104906

Merged
AndyAyersMS merged 97 commits into
dotnet:mainfrom
hez2010:value-array-stack-alloc
Jan 22, 2025
Merged

JIT: Extend escape analysis to account for arrays with non-gcref elements#104906
AndyAyersMS merged 97 commits into
dotnet:mainfrom
hez2010:value-array-stack-alloc

Conversation

@hez2010

@hez2010hez2010 commented Jul 15, 2024

Copy link
Copy Markdown
Contributor

Positive case:

varchs=newchar[42];chs[1]='a';Console.WriteLine((int)chs[1]+chs.Length);

Codegen:

; Assembly listing for method ArrayAllocator.Program:Main() (FullOpts); Emitting BLENDED_CODE for X64 with AVX - Windows; FullOpts code; optimized code; rsp based frame; partially interruptible; No PGO data; Final local variable assignments;;* V00 loc0 [V00 ] ( 0, 0 ) long -> zero-ref class-hnd exact <short[]>; V01 OutArgs [V01 ] ( 1, 1 ) struct (32) [rsp+0x00] do-not-enreg[XS] addr-exposed "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) struct (104) zero-ref do-not-enreg[SF] "stack allocated array temp";* V03 tmp2 [V03 ] ( 0, 0 ) long -> zero-ref single-def "V02.[000..008)";* V04 tmp3 [V04 ] ( 0, 0 ) int -> zero-ref single-def "V02.[008..012)";* V05 tmp4 [V05 ] ( 0, 0 ) short -> zero-ref "V02.[018..020)";; Lcl frame size = 40G_M25548_IG01: ;; offset=0x0000subrsp,40 ;; size=4 bbWeight=1 PerfScore 0.25G_M25548_IG02: ;; offset=0x0004movecx,84call[System.Console:WriteLine(int)]nop ;; size=12 bbWeight=1 PerfScore 3.50G_M25548_IG03: ;; offset=0x0010addrsp,40ret ;; size=5 bbWeight=1 PerfScore 1.25; Total bytes of code 21, prolog size 4, PerfScore 5.00, instruction count 6, allocated bytes for code 21 (MethodHash=5b0b9c33) for method ArrayAllocator.Program:Main() (FullOpts)

Negative case:

varchs=newchar[42];chs[1]='a';Console.WriteLine((int)chs[42]+chs.Length);

Codegen:

; Assembly listing for method ArrayAllocator.Program:Main() (FullOpts); Emitting BLENDED_CODE for X64 with AVX - Windows; FullOpts code; optimized code; rsp based frame; partially interruptible; No PGO data; Final local variable assignments;;* V00 loc0 [V00 ] ( 0, 0 ) long -> zero-ref class-hnd exact <short[]>; V01 OutArgs [V01 ] ( 1, 1 ) struct (32) [rsp+0x00] do-not-enreg[XS] addr-exposed "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) struct (104) zero-ref do-not-enreg[SF] "stack allocated array temp"; V03 tmp2 [V03,T00] ( 1, 0 ) byref -> rbx must-init "dummy temp of must thrown exception";* V04 tmp3 [V04 ] ( 0, 0 ) long -> zero-ref single-def "V02.[000..008)";* V05 tmp4 [V05 ] ( 0, 0 ) int -> zero-ref single-def "V02.[008..012)";* V06 tmp5 [V06 ] ( 0, 0 ) short -> zero-ref single-def "V02.[018..020)";; Lcl frame size = 32G_M25548_IG01: ;; offset=0x0000pushrbxsubrsp,32xorebx,ebx ;; size=7 bbWeight=0 PerfScore 0.00G_M25548_IG02: ;; offset=0x0007call CORINFO_HELP_RNGCHKFAILmovsxrcx, word ptr [rbx]call[System.Console:WriteLine(int)]int3 ;; size=16 bbWeight=0 PerfScore 0.00; Total bytes of code 23, prolog size 5, PerfScore 0.00, instruction count 7, allocated bytes for code 23 (MethodHash=5b0b9c33) for method ArrayAllocator.Program:Main() (FullOpts); ============================================================

Benchmark on Mandelbrot:

MethodJobMeanErrorStdDevCode SizeAllocated
MandelBrotNoStackAllocationArray199.7 us1.30 us1.22 us1,996 B2.49 KB
MandelBrotStackAllocationArray195.8 us1.16 us1.08 us2,414 B1.14 KB

Diff: https://www.diffchecker.com/bNP4qHdF/

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 15, 2024
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jul 15, 2024
Comment threadsrc/coreclr/jit/objectalloc.h Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated

@AndyAyersMSAndyAyersMS left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For arrays (and also perhaps boxes and ref classes) we ought to have some kind of size limit... possibly similar to the one we use for stackallocs.

We need to be careful we don't allocate a lot of stack for an object that might not be heavily used, as we'll pay per-call prolog zeroing costs.

Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/gentree.cpp Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated
#ifdef FEATURE_READYTORUN
if (comp->opts.IsReadyToRun() && data->IsHelperCall(comp, CORINFO_HELP_READYTORUN_NEWARR_1))
{
len = data->AsCall()->gtArgs.GetArgByIndex(0)->GetNode();

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Dead code?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I removed R2R support more completely

Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
Comment on lines +2849 to +2854
GenTree* const mtStore = gtNewStoreLclFldNode(lclNum, TYP_I_IMPL, 0, mt);
Statement* const mtStmt = gtNewStmt(mtStore);

fgInsertStmtBefore(block, newStmt, mtStmt);
gtSetStmtInfo(mtStmt);
fgSetStmtSeq(mtStmt);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
GenTree* const mtStore = gtNewStoreLclFldNode(lclNum, TYP_I_IMPL, 0, mt);
Statement* const mtStmt = gtNewStmt(mtStore);
fgInsertStmtBefore(block, newStmt, mtStmt);
gtSetStmtInfo(mtStmt);
fgSetStmtSeq(mtStmt);
GenTree* const mtStore = gtNewStoreLclFldNode(lclNum, TYP_I_IMPL, 0, mt);
Statement* const mtStmt = fgNewStmtFromTree(mtStore);
fgInsertStmtBefore(block, newStmt, mtStmt);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ditto below, more things can switch to use fgNewStmtFromTree

Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
@@ -181,6 +254,11 @@ inline bool ObjectAllocator::CanAllocateLclVarOnStack(unsigned int lclNu
return false;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems like we should have a limit on the aggregate stack allocated size. That's somewhat preexisting, but probably more important now.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agree but will do that in a future PR

Comment on lines +6249 to +6260
// If this is a local array, there are no asyncronous modifications, so we can set the
// conservative VN to the liberal VN.
//
VNFuncApp arrFn;
if (vnStore->IsVNNewLocalArr(arrVN, &arrFn))
{
loadTree->gtVNPair.SetConservative(loadValueVN);
}
else
{
loadTree->gtVNPair.SetConservative(vnStore->VNForExpr(compCurBB, loadType));
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does this actually show up as benefits?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In some limited cases, yes... eg we can const propagate through a[2] = 1; y = a[2];

@AndyAyersMS

Copy link
Copy Markdown
Member

@jakobbotsch think I've addressed most of the key points. Overall size limit will come in a future PR.

@jakobbotsch

Copy link
Copy Markdown
Member

@AndyAyersMS Did you push those changes?

@AndyAyersMS

Copy link
Copy Markdown
Member

Ah, I pushed to my fork, but ... this PR is not from my fork.

@AndyAyersMS

Copy link
Copy Markdown
Member

@jakobbotsch changes are there now

GTF_CALL_M_CAST_CAN_BE_EXPANDED = 0x04000000, // this cast (helper call) can be expanded if it's profitable. To be removed.
GTF_CALL_M_CAST_OBJ_NONNULL = 0x08000000, // if we expand this specific cast we don't need to check the input object for null
// NOTE: if needed, this flag can be removed, and we can introduce new _NONNUL cast helpers
GTF_CALL_M_STACK_ARRAY = 0x10000000, // this call is a new array helper for a stack allocated array.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I suppose we weren't able to get rid of this since it's also used in VN.

Comment on lines +14375 to +14380
// modHeap = false;
}
else if (vnf == VNF_JitReadyToRunNewArr)
{
vnf = VNF_JitReadyToRunNewLclArr;
// modHeap = false;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: commented code (feel free to remove as part of a follow up)

}
}

if (isAlloc && ((call->gtCallMoreFlags & GTF_CALL_M_STACK_ARRAY) != 0))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I suppose this could be changed to FindWellKnownArg(StackArrayLocal) != nullptr to get rid of the flag

@jakobbotschjakobbotsch left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@AndyAyersMS

Copy link
Copy Markdown
Member

Thanks. I have more changes in this area so I can handle the last few bits in a subsequent PR.

@AndyAyersMS
AndyAyersMS merged commit 7d75878 into dotnet:mainJan 22, 2025
@AndyAyersMS

Copy link
Copy Markdown
Member

@hez2010 thanks for all the work you did here.

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@hez2010@AndyAyersMS@JulieLeeMSFT@jakobbotsch
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: Extend escape analysis to account for arrays with non-gcref elements - #104906

Merged
AndyAyersMS merged 97 commits into
dotnet:mainfrom
hez2010:value-array-stack-alloc
Jan 22, 2025
Merged

JIT: Extend escape analysis to account for arrays with non-gcref elements#104906
AndyAyersMS merged 97 commits into
dotnet:mainfrom
hez2010:value-array-stack-alloc

Conversation

@hez2010

@hez2010hez2010 commented Jul 15, 2024

Copy link
Copy Markdown
Contributor

Positive case:

varchs=newchar[42];chs[1]='a';Console.WriteLine((int)chs[1]+chs.Length);

Codegen:

; Assembly listing for method ArrayAllocator.Program:Main() (FullOpts); Emitting BLENDED_CODE for X64 with AVX - Windows; FullOpts code; optimized code; rsp based frame; partially interruptible; No PGO data; Final local variable assignments;;* V00 loc0 [V00 ] ( 0, 0 ) long -> zero-ref class-hnd exact <short[]>; V01 OutArgs [V01 ] ( 1, 1 ) struct (32) [rsp+0x00] do-not-enreg[XS] addr-exposed "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) struct (104) zero-ref do-not-enreg[SF] "stack allocated array temp";* V03 tmp2 [V03 ] ( 0, 0 ) long -> zero-ref single-def "V02.[000..008)";* V04 tmp3 [V04 ] ( 0, 0 ) int -> zero-ref single-def "V02.[008..012)";* V05 tmp4 [V05 ] ( 0, 0 ) short -> zero-ref "V02.[018..020)";; Lcl frame size = 40G_M25548_IG01: ;; offset=0x0000subrsp,40 ;; size=4 bbWeight=1 PerfScore 0.25G_M25548_IG02: ;; offset=0x0004movecx,84call[System.Console:WriteLine(int)]nop ;; size=12 bbWeight=1 PerfScore 3.50G_M25548_IG03: ;; offset=0x0010addrsp,40ret ;; size=5 bbWeight=1 PerfScore 1.25; Total bytes of code 21, prolog size 4, PerfScore 5.00, instruction count 6, allocated bytes for code 21 (MethodHash=5b0b9c33) for method ArrayAllocator.Program:Main() (FullOpts)

Negative case:

varchs=newchar[42];chs[1]='a';Console.WriteLine((int)chs[42]+chs.Length);

Codegen:

; Assembly listing for method ArrayAllocator.Program:Main() (FullOpts); Emitting BLENDED_CODE for X64 with AVX - Windows; FullOpts code; optimized code; rsp based frame; partially interruptible; No PGO data; Final local variable assignments;;* V00 loc0 [V00 ] ( 0, 0 ) long -> zero-ref class-hnd exact <short[]>; V01 OutArgs [V01 ] ( 1, 1 ) struct (32) [rsp+0x00] do-not-enreg[XS] addr-exposed "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) struct (104) zero-ref do-not-enreg[SF] "stack allocated array temp"; V03 tmp2 [V03,T00] ( 1, 0 ) byref -> rbx must-init "dummy temp of must thrown exception";* V04 tmp3 [V04 ] ( 0, 0 ) long -> zero-ref single-def "V02.[000..008)";* V05 tmp4 [V05 ] ( 0, 0 ) int -> zero-ref single-def "V02.[008..012)";* V06 tmp5 [V06 ] ( 0, 0 ) short -> zero-ref single-def "V02.[018..020)";; Lcl frame size = 32G_M25548_IG01: ;; offset=0x0000pushrbxsubrsp,32xorebx,ebx ;; size=7 bbWeight=0 PerfScore 0.00G_M25548_IG02: ;; offset=0x0007call CORINFO_HELP_RNGCHKFAILmovsxrcx, word ptr [rbx]call[System.Console:WriteLine(int)]int3 ;; size=16 bbWeight=0 PerfScore 0.00; Total bytes of code 23, prolog size 5, PerfScore 0.00, instruction count 7, allocated bytes for code 23 (MethodHash=5b0b9c33) for method ArrayAllocator.Program:Main() (FullOpts); ============================================================

Benchmark on Mandelbrot:

MethodJobMeanErrorStdDevCode SizeAllocated
MandelBrotNoStackAllocationArray199.7 us1.30 us1.22 us1,996 B2.49 KB
MandelBrotStackAllocationArray195.8 us1.16 us1.08 us2,414 B1.14 KB

Diff: https://www.diffchecker.com/bNP4qHdF/

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 15, 2024
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jul 15, 2024
Comment threadsrc/coreclr/jit/objectalloc.h Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated

@AndyAyersMSAndyAyersMS left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For arrays (and also perhaps boxes and ref classes) we ought to have some kind of size limit... possibly similar to the one we use for stackallocs.

We need to be careful we don't allocate a lot of stack for an object that might not be heavily used, as we'll pay per-call prolog zeroing costs.

Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/gentree.cpp Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated
#ifdef FEATURE_READYTORUN
if (comp->opts.IsReadyToRun() && data->IsHelperCall(comp, CORINFO_HELP_READYTORUN_NEWARR_1))
{
len = data->AsCall()->gtArgs.GetArgByIndex(0)->GetNode();

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Dead code?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I removed R2R support more completely

Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
Comment on lines +2849 to +2854
GenTree* const mtStore = gtNewStoreLclFldNode(lclNum, TYP_I_IMPL, 0, mt);
Statement* const mtStmt = gtNewStmt(mtStore);

fgInsertStmtBefore(block, newStmt, mtStmt);
gtSetStmtInfo(mtStmt);
fgSetStmtSeq(mtStmt);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
GenTree* const mtStore = gtNewStoreLclFldNode(lclNum, TYP_I_IMPL, 0, mt);
Statement* const mtStmt = gtNewStmt(mtStore);
fgInsertStmtBefore(block, newStmt, mtStmt);
gtSetStmtInfo(mtStmt);
fgSetStmtSeq(mtStmt);
GenTree* const mtStore = gtNewStoreLclFldNode(lclNum, TYP_I_IMPL, 0, mt);
Statement* const mtStmt = fgNewStmtFromTree(mtStore);
fgInsertStmtBefore(block, newStmt, mtStmt);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ditto below, more things can switch to use fgNewStmtFromTree

Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
@@ -181,6 +254,11 @@ inline bool ObjectAllocator::CanAllocateLclVarOnStack(unsigned int lclNu
return false;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems like we should have a limit on the aggregate stack allocated size. That's somewhat preexisting, but probably more important now.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agree but will do that in a future PR

Comment on lines +6249 to +6260
// If this is a local array, there are no asyncronous modifications, so we can set the
// conservative VN to the liberal VN.
//
VNFuncApp arrFn;
if (vnStore->IsVNNewLocalArr(arrVN, &arrFn))
{
loadTree->gtVNPair.SetConservative(loadValueVN);
}
else
{
loadTree->gtVNPair.SetConservative(vnStore->VNForExpr(compCurBB, loadType));
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does this actually show up as benefits?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In some limited cases, yes... eg we can const propagate through a[2] = 1; y = a[2];

@AndyAyersMS

Copy link
Copy Markdown
Member

@jakobbotsch think I've addressed most of the key points. Overall size limit will come in a future PR.

@jakobbotsch

Copy link
Copy Markdown
Member

@AndyAyersMS Did you push those changes?

@AndyAyersMS

Copy link
Copy Markdown
Member

Ah, I pushed to my fork, but ... this PR is not from my fork.

@AndyAyersMS

Copy link
Copy Markdown
Member

@jakobbotsch changes are there now

GTF_CALL_M_CAST_CAN_BE_EXPANDED = 0x04000000, // this cast (helper call) can be expanded if it's profitable. To be removed.
GTF_CALL_M_CAST_OBJ_NONNULL = 0x08000000, // if we expand this specific cast we don't need to check the input object for null
// NOTE: if needed, this flag can be removed, and we can introduce new _NONNUL cast helpers
GTF_CALL_M_STACK_ARRAY = 0x10000000, // this call is a new array helper for a stack allocated array.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I suppose we weren't able to get rid of this since it's also used in VN.

Comment on lines +14375 to +14380
// modHeap = false;
}
else if (vnf == VNF_JitReadyToRunNewArr)
{
vnf = VNF_JitReadyToRunNewLclArr;
// modHeap = false;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: commented code (feel free to remove as part of a follow up)

}
}

if (isAlloc && ((call->gtCallMoreFlags & GTF_CALL_M_STACK_ARRAY) != 0))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I suppose this could be changed to FindWellKnownArg(StackArrayLocal) != nullptr to get rid of the flag

@jakobbotschjakobbotsch left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@AndyAyersMS

Copy link
Copy Markdown
Member

Thanks. I have more changes in this area so I can handle the last few bits in a subsequent PR.

@AndyAyersMS
AndyAyersMS merged commit 7d75878 into dotnet:mainJan 22, 2025
@AndyAyersMS

Copy link
Copy Markdown
Member

@hez2010 thanks for all the work you did here.

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@hez2010@AndyAyersMS@JulieLeeMSFT@jakobbotsch
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: Extend escape analysis to account for arrays with non-gcref elements - #104906

Merged
AndyAyersMS merged 97 commits into
dotnet:mainfrom
hez2010:value-array-stack-alloc
Jan 22, 2025
Merged

JIT: Extend escape analysis to account for arrays with non-gcref elements#104906
AndyAyersMS merged 97 commits into
dotnet:mainfrom
hez2010:value-array-stack-alloc

Conversation

@hez2010

@hez2010hez2010 commented Jul 15, 2024

Copy link
Copy Markdown
Contributor

Positive case:

varchs=newchar[42];chs[1]='a';Console.WriteLine((int)chs[1]+chs.Length);

Codegen:

; Assembly listing for method ArrayAllocator.Program:Main() (FullOpts); Emitting BLENDED_CODE for X64 with AVX - Windows; FullOpts code; optimized code; rsp based frame; partially interruptible; No PGO data; Final local variable assignments;;* V00 loc0 [V00 ] ( 0, 0 ) long -> zero-ref class-hnd exact <short[]>; V01 OutArgs [V01 ] ( 1, 1 ) struct (32) [rsp+0x00] do-not-enreg[XS] addr-exposed "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) struct (104) zero-ref do-not-enreg[SF] "stack allocated array temp";* V03 tmp2 [V03 ] ( 0, 0 ) long -> zero-ref single-def "V02.[000..008)";* V04 tmp3 [V04 ] ( 0, 0 ) int -> zero-ref single-def "V02.[008..012)";* V05 tmp4 [V05 ] ( 0, 0 ) short -> zero-ref "V02.[018..020)";; Lcl frame size = 40G_M25548_IG01: ;; offset=0x0000subrsp,40 ;; size=4 bbWeight=1 PerfScore 0.25G_M25548_IG02: ;; offset=0x0004movecx,84call[System.Console:WriteLine(int)]nop ;; size=12 bbWeight=1 PerfScore 3.50G_M25548_IG03: ;; offset=0x0010addrsp,40ret ;; size=5 bbWeight=1 PerfScore 1.25; Total bytes of code 21, prolog size 4, PerfScore 5.00, instruction count 6, allocated bytes for code 21 (MethodHash=5b0b9c33) for method ArrayAllocator.Program:Main() (FullOpts)

Negative case:

varchs=newchar[42];chs[1]='a';Console.WriteLine((int)chs[42]+chs.Length);

Codegen:

; Assembly listing for method ArrayAllocator.Program:Main() (FullOpts); Emitting BLENDED_CODE for X64 with AVX - Windows; FullOpts code; optimized code; rsp based frame; partially interruptible; No PGO data; Final local variable assignments;;* V00 loc0 [V00 ] ( 0, 0 ) long -> zero-ref class-hnd exact <short[]>; V01 OutArgs [V01 ] ( 1, 1 ) struct (32) [rsp+0x00] do-not-enreg[XS] addr-exposed "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) struct (104) zero-ref do-not-enreg[SF] "stack allocated array temp"; V03 tmp2 [V03,T00] ( 1, 0 ) byref -> rbx must-init "dummy temp of must thrown exception";* V04 tmp3 [V04 ] ( 0, 0 ) long -> zero-ref single-def "V02.[000..008)";* V05 tmp4 [V05 ] ( 0, 0 ) int -> zero-ref single-def "V02.[008..012)";* V06 tmp5 [V06 ] ( 0, 0 ) short -> zero-ref single-def "V02.[018..020)";; Lcl frame size = 32G_M25548_IG01: ;; offset=0x0000pushrbxsubrsp,32xorebx,ebx ;; size=7 bbWeight=0 PerfScore 0.00G_M25548_IG02: ;; offset=0x0007call CORINFO_HELP_RNGCHKFAILmovsxrcx, word ptr [rbx]call[System.Console:WriteLine(int)]int3 ;; size=16 bbWeight=0 PerfScore 0.00; Total bytes of code 23, prolog size 5, PerfScore 0.00, instruction count 7, allocated bytes for code 23 (MethodHash=5b0b9c33) for method ArrayAllocator.Program:Main() (FullOpts); ============================================================

Benchmark on Mandelbrot:

MethodJobMeanErrorStdDevCode SizeAllocated
MandelBrotNoStackAllocationArray199.7 us1.30 us1.22 us1,996 B2.49 KB
MandelBrotStackAllocationArray195.8 us1.16 us1.08 us2,414 B1.14 KB

Diff: https://www.diffchecker.com/bNP4qHdF/

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 15, 2024
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jul 15, 2024
Comment threadsrc/coreclr/jit/objectalloc.h Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated

@AndyAyersMSAndyAyersMS left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For arrays (and also perhaps boxes and ref classes) we ought to have some kind of size limit... possibly similar to the one we use for stackallocs.

We need to be careful we don't allocate a lot of stack for an object that might not be heavily used, as we'll pay per-call prolog zeroing costs.

Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/gentree.cpp Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated
#ifdef FEATURE_READYTORUN
if (comp->opts.IsReadyToRun() && data->IsHelperCall(comp, CORINFO_HELP_READYTORUN_NEWARR_1))
{
len = data->AsCall()->gtArgs.GetArgByIndex(0)->GetNode();

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Dead code?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I removed R2R support more completely

Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
Comment on lines +2849 to +2854
GenTree* const mtStore = gtNewStoreLclFldNode(lclNum, TYP_I_IMPL, 0, mt);
Statement* const mtStmt = gtNewStmt(mtStore);

fgInsertStmtBefore(block, newStmt, mtStmt);
gtSetStmtInfo(mtStmt);
fgSetStmtSeq(mtStmt);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
GenTree* const mtStore = gtNewStoreLclFldNode(lclNum, TYP_I_IMPL, 0, mt);
Statement* const mtStmt = gtNewStmt(mtStore);
fgInsertStmtBefore(block, newStmt, mtStmt);
gtSetStmtInfo(mtStmt);
fgSetStmtSeq(mtStmt);
GenTree* const mtStore = gtNewStoreLclFldNode(lclNum, TYP_I_IMPL, 0, mt);
Statement* const mtStmt = fgNewStmtFromTree(mtStore);
fgInsertStmtBefore(block, newStmt, mtStmt);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ditto below, more things can switch to use fgNewStmtFromTree

Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
@@ -181,6 +254,11 @@ inline bool ObjectAllocator::CanAllocateLclVarOnStack(unsigned int lclNu
return false;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems like we should have a limit on the aggregate stack allocated size. That's somewhat preexisting, but probably more important now.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agree but will do that in a future PR

Comment on lines +6249 to +6260
// If this is a local array, there are no asyncronous modifications, so we can set the
// conservative VN to the liberal VN.
//
VNFuncApp arrFn;
if (vnStore->IsVNNewLocalArr(arrVN, &arrFn))
{
loadTree->gtVNPair.SetConservative(loadValueVN);
}
else
{
loadTree->gtVNPair.SetConservative(vnStore->VNForExpr(compCurBB, loadType));
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does this actually show up as benefits?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In some limited cases, yes... eg we can const propagate through a[2] = 1; y = a[2];

@AndyAyersMS

Copy link
Copy Markdown
Member

@jakobbotsch think I've addressed most of the key points. Overall size limit will come in a future PR.

@jakobbotsch

Copy link
Copy Markdown
Member

@AndyAyersMS Did you push those changes?

@AndyAyersMS

Copy link
Copy Markdown
Member

Ah, I pushed to my fork, but ... this PR is not from my fork.

@AndyAyersMS

Copy link
Copy Markdown
Member

@jakobbotsch changes are there now

GTF_CALL_M_CAST_CAN_BE_EXPANDED = 0x04000000, // this cast (helper call) can be expanded if it's profitable. To be removed.
GTF_CALL_M_CAST_OBJ_NONNULL = 0x08000000, // if we expand this specific cast we don't need to check the input object for null
// NOTE: if needed, this flag can be removed, and we can introduce new _NONNUL cast helpers
GTF_CALL_M_STACK_ARRAY = 0x10000000, // this call is a new array helper for a stack allocated array.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I suppose we weren't able to get rid of this since it's also used in VN.

Comment on lines +14375 to +14380
// modHeap = false;
}
else if (vnf == VNF_JitReadyToRunNewArr)
{
vnf = VNF_JitReadyToRunNewLclArr;
// modHeap = false;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: commented code (feel free to remove as part of a follow up)

}
}

if (isAlloc && ((call->gtCallMoreFlags & GTF_CALL_M_STACK_ARRAY) != 0))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I suppose this could be changed to FindWellKnownArg(StackArrayLocal) != nullptr to get rid of the flag

@jakobbotschjakobbotsch left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@AndyAyersMS

Copy link
Copy Markdown
Member

Thanks. I have more changes in this area so I can handle the last few bits in a subsequent PR.

@AndyAyersMS
AndyAyersMS merged commit 7d75878 into dotnet:mainJan 22, 2025
@AndyAyersMS

Copy link
Copy Markdown
Member

@hez2010 thanks for all the work you did here.

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@hez2010@AndyAyersMS@JulieLeeMSFT@jakobbotsch
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

JIT: Extend escape analysis to account for arrays with non-gcref elements - #104906

Merged
AndyAyersMS merged 97 commits into
dotnet:mainfrom
hez2010:value-array-stack-alloc
Jan 22, 2025
Merged

JIT: Extend escape analysis to account for arrays with non-gcref elements#104906
AndyAyersMS merged 97 commits into
dotnet:mainfrom
hez2010:value-array-stack-alloc

Conversation

@hez2010

@hez2010hez2010 commented Jul 15, 2024

Copy link
Copy Markdown
Contributor

Positive case:

varchs=newchar[42];chs[1]='a';Console.WriteLine((int)chs[1]+chs.Length);

Codegen:

; Assembly listing for method ArrayAllocator.Program:Main() (FullOpts); Emitting BLENDED_CODE for X64 with AVX - Windows; FullOpts code; optimized code; rsp based frame; partially interruptible; No PGO data; Final local variable assignments;;* V00 loc0 [V00 ] ( 0, 0 ) long -> zero-ref class-hnd exact <short[]>; V01 OutArgs [V01 ] ( 1, 1 ) struct (32) [rsp+0x00] do-not-enreg[XS] addr-exposed "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) struct (104) zero-ref do-not-enreg[SF] "stack allocated array temp";* V03 tmp2 [V03 ] ( 0, 0 ) long -> zero-ref single-def "V02.[000..008)";* V04 tmp3 [V04 ] ( 0, 0 ) int -> zero-ref single-def "V02.[008..012)";* V05 tmp4 [V05 ] ( 0, 0 ) short -> zero-ref "V02.[018..020)";; Lcl frame size = 40G_M25548_IG01: ;; offset=0x0000subrsp,40 ;; size=4 bbWeight=1 PerfScore 0.25G_M25548_IG02: ;; offset=0x0004movecx,84call[System.Console:WriteLine(int)]nop ;; size=12 bbWeight=1 PerfScore 3.50G_M25548_IG03: ;; offset=0x0010addrsp,40ret ;; size=5 bbWeight=1 PerfScore 1.25; Total bytes of code 21, prolog size 4, PerfScore 5.00, instruction count 6, allocated bytes for code 21 (MethodHash=5b0b9c33) for method ArrayAllocator.Program:Main() (FullOpts)

Negative case:

varchs=newchar[42];chs[1]='a';Console.WriteLine((int)chs[42]+chs.Length);

Codegen:

; Assembly listing for method ArrayAllocator.Program:Main() (FullOpts); Emitting BLENDED_CODE for X64 with AVX - Windows; FullOpts code; optimized code; rsp based frame; partially interruptible; No PGO data; Final local variable assignments;;* V00 loc0 [V00 ] ( 0, 0 ) long -> zero-ref class-hnd exact <short[]>; V01 OutArgs [V01 ] ( 1, 1 ) struct (32) [rsp+0x00] do-not-enreg[XS] addr-exposed "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) struct (104) zero-ref do-not-enreg[SF] "stack allocated array temp"; V03 tmp2 [V03,T00] ( 1, 0 ) byref -> rbx must-init "dummy temp of must thrown exception";* V04 tmp3 [V04 ] ( 0, 0 ) long -> zero-ref single-def "V02.[000..008)";* V05 tmp4 [V05 ] ( 0, 0 ) int -> zero-ref single-def "V02.[008..012)";* V06 tmp5 [V06 ] ( 0, 0 ) short -> zero-ref single-def "V02.[018..020)";; Lcl frame size = 32G_M25548_IG01: ;; offset=0x0000pushrbxsubrsp,32xorebx,ebx ;; size=7 bbWeight=0 PerfScore 0.00G_M25548_IG02: ;; offset=0x0007call CORINFO_HELP_RNGCHKFAILmovsxrcx, word ptr [rbx]call[System.Console:WriteLine(int)]int3 ;; size=16 bbWeight=0 PerfScore 0.00; Total bytes of code 23, prolog size 5, PerfScore 0.00, instruction count 7, allocated bytes for code 23 (MethodHash=5b0b9c33) for method ArrayAllocator.Program:Main() (FullOpts); ============================================================

Benchmark on Mandelbrot:

MethodJobMeanErrorStdDevCode SizeAllocated
MandelBrotNoStackAllocationArray199.7 us1.30 us1.22 us1,996 B2.49 KB
MandelBrotStackAllocationArray195.8 us1.16 us1.08 us2,414 B1.14 KB

Diff: https://www.diffchecker.com/bNP4qHdF/

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 15, 2024
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jul 15, 2024
Comment threadsrc/coreclr/jit/objectalloc.h Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated

@AndyAyersMSAndyAyersMS left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For arrays (and also perhaps boxes and ref classes) we ought to have some kind of size limit... possibly similar to the one we use for stackallocs.

We need to be careful we don't allocate a lot of stack for an object that might not be heavily used, as we'll pay per-call prolog zeroing costs.

Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/gentree.cpp Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated
#ifdef FEATURE_READYTORUN
if (comp->opts.IsReadyToRun() && data->IsHelperCall(comp, CORINFO_HELP_READYTORUN_NEWARR_1))
{
len = data->AsCall()->gtArgs.GetArgByIndex(0)->GetNode();

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Dead code?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I removed R2R support more completely

Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
Comment on lines +2849 to +2854
GenTree* const mtStore = gtNewStoreLclFldNode(lclNum, TYP_I_IMPL, 0, mt);
Statement* const mtStmt = gtNewStmt(mtStore);

fgInsertStmtBefore(block, newStmt, mtStmt);
gtSetStmtInfo(mtStmt);
fgSetStmtSeq(mtStmt);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
GenTree* const mtStore = gtNewStoreLclFldNode(lclNum, TYP_I_IMPL, 0, mt);
Statement* const mtStmt = gtNewStmt(mtStore);
fgInsertStmtBefore(block, newStmt, mtStmt);
gtSetStmtInfo(mtStmt);
fgSetStmtSeq(mtStmt);
GenTree* const mtStore = gtNewStoreLclFldNode(lclNum, TYP_I_IMPL, 0, mt);
Statement* const mtStmt = fgNewStmtFromTree(mtStore);
fgInsertStmtBefore(block, newStmt, mtStmt);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ditto below, more things can switch to use fgNewStmtFromTree

Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
@@ -181,6 +254,11 @@ inline bool ObjectAllocator::CanAllocateLclVarOnStack(unsigned int lclNu
return false;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems like we should have a limit on the aggregate stack allocated size. That's somewhat preexisting, but probably more important now.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agree but will do that in a future PR

Comment on lines +6249 to +6260
// If this is a local array, there are no asyncronous modifications, so we can set the
// conservative VN to the liberal VN.
//
VNFuncApp arrFn;
if (vnStore->IsVNNewLocalArr(arrVN, &arrFn))
{
loadTree->gtVNPair.SetConservative(loadValueVN);
}
else
{
loadTree->gtVNPair.SetConservative(vnStore->VNForExpr(compCurBB, loadType));
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does this actually show up as benefits?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In some limited cases, yes... eg we can const propagate through a[2] = 1; y = a[2];

@AndyAyersMS

Copy link
Copy Markdown
Member

@jakobbotsch think I've addressed most of the key points. Overall size limit will come in a future PR.

@jakobbotsch

Copy link
Copy Markdown
Member

@AndyAyersMS Did you push those changes?

@AndyAyersMS

Copy link
Copy Markdown
Member

Ah, I pushed to my fork, but ... this PR is not from my fork.

@AndyAyersMS

Copy link
Copy Markdown
Member

@jakobbotsch changes are there now

GTF_CALL_M_CAST_CAN_BE_EXPANDED = 0x04000000, // this cast (helper call) can be expanded if it's profitable. To be removed.
GTF_CALL_M_CAST_OBJ_NONNULL = 0x08000000, // if we expand this specific cast we don't need to check the input object for null
// NOTE: if needed, this flag can be removed, and we can introduce new _NONNUL cast helpers
GTF_CALL_M_STACK_ARRAY = 0x10000000, // this call is a new array helper for a stack allocated array.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I suppose we weren't able to get rid of this since it's also used in VN.

Comment on lines +14375 to +14380
// modHeap = false;
}
else if (vnf == VNF_JitReadyToRunNewArr)
{
vnf = VNF_JitReadyToRunNewLclArr;
// modHeap = false;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: commented code (feel free to remove as part of a follow up)

}
}

if (isAlloc && ((call->gtCallMoreFlags & GTF_CALL_M_STACK_ARRAY) != 0))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I suppose this could be changed to FindWellKnownArg(StackArrayLocal) != nullptr to get rid of the flag

@jakobbotschjakobbotsch left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@AndyAyersMS

Copy link
Copy Markdown
Member

Thanks. I have more changes in this area so I can handle the last few bits in a subsequent PR.

@AndyAyersMS
AndyAyersMS merged commit 7d75878 into dotnet:mainJan 22, 2025
@AndyAyersMS

Copy link
Copy Markdown
Member

@hez2010 thanks for all the work you did here.

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@hez2010@AndyAyersMS@JulieLeeMSFT@jakobbotsch
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: Extend escape analysis to account for arrays with non-gcref elements - #104906

Merged
AndyAyersMS merged 97 commits into
dotnet:mainfrom
hez2010:value-array-stack-alloc
Jan 22, 2025
Merged

JIT: Extend escape analysis to account for arrays with non-gcref elements#104906
AndyAyersMS merged 97 commits into
dotnet:mainfrom
hez2010:value-array-stack-alloc

Conversation

@hez2010

@hez2010hez2010 commented Jul 15, 2024

Copy link
Copy Markdown
Contributor

Positive case:

varchs=newchar[42];chs[1]='a';Console.WriteLine((int)chs[1]+chs.Length);

Codegen:

; Assembly listing for method ArrayAllocator.Program:Main() (FullOpts); Emitting BLENDED_CODE for X64 with AVX - Windows; FullOpts code; optimized code; rsp based frame; partially interruptible; No PGO data; Final local variable assignments;;* V00 loc0 [V00 ] ( 0, 0 ) long -> zero-ref class-hnd exact <short[]>; V01 OutArgs [V01 ] ( 1, 1 ) struct (32) [rsp+0x00] do-not-enreg[XS] addr-exposed "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) struct (104) zero-ref do-not-enreg[SF] "stack allocated array temp";* V03 tmp2 [V03 ] ( 0, 0 ) long -> zero-ref single-def "V02.[000..008)";* V04 tmp3 [V04 ] ( 0, 0 ) int -> zero-ref single-def "V02.[008..012)";* V05 tmp4 [V05 ] ( 0, 0 ) short -> zero-ref "V02.[018..020)";; Lcl frame size = 40G_M25548_IG01: ;; offset=0x0000subrsp,40 ;; size=4 bbWeight=1 PerfScore 0.25G_M25548_IG02: ;; offset=0x0004movecx,84call[System.Console:WriteLine(int)]nop ;; size=12 bbWeight=1 PerfScore 3.50G_M25548_IG03: ;; offset=0x0010addrsp,40ret ;; size=5 bbWeight=1 PerfScore 1.25; Total bytes of code 21, prolog size 4, PerfScore 5.00, instruction count 6, allocated bytes for code 21 (MethodHash=5b0b9c33) for method ArrayAllocator.Program:Main() (FullOpts)

Negative case:

varchs=newchar[42];chs[1]='a';Console.WriteLine((int)chs[42]+chs.Length);

Codegen:

; Assembly listing for method ArrayAllocator.Program:Main() (FullOpts); Emitting BLENDED_CODE for X64 with AVX - Windows; FullOpts code; optimized code; rsp based frame; partially interruptible; No PGO data; Final local variable assignments;;* V00 loc0 [V00 ] ( 0, 0 ) long -> zero-ref class-hnd exact <short[]>; V01 OutArgs [V01 ] ( 1, 1 ) struct (32) [rsp+0x00] do-not-enreg[XS] addr-exposed "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) struct (104) zero-ref do-not-enreg[SF] "stack allocated array temp"; V03 tmp2 [V03,T00] ( 1, 0 ) byref -> rbx must-init "dummy temp of must thrown exception";* V04 tmp3 [V04 ] ( 0, 0 ) long -> zero-ref single-def "V02.[000..008)";* V05 tmp4 [V05 ] ( 0, 0 ) int -> zero-ref single-def "V02.[008..012)";* V06 tmp5 [V06 ] ( 0, 0 ) short -> zero-ref single-def "V02.[018..020)";; Lcl frame size = 32G_M25548_IG01: ;; offset=0x0000pushrbxsubrsp,32xorebx,ebx ;; size=7 bbWeight=0 PerfScore 0.00G_M25548_IG02: ;; offset=0x0007call CORINFO_HELP_RNGCHKFAILmovsxrcx, word ptr [rbx]call[System.Console:WriteLine(int)]int3 ;; size=16 bbWeight=0 PerfScore 0.00; Total bytes of code 23, prolog size 5, PerfScore 0.00, instruction count 7, allocated bytes for code 23 (MethodHash=5b0b9c33) for method ArrayAllocator.Program:Main() (FullOpts); ============================================================

Benchmark on Mandelbrot:

MethodJobMeanErrorStdDevCode SizeAllocated
MandelBrotNoStackAllocationArray199.7 us1.30 us1.22 us1,996 B2.49 KB
MandelBrotStackAllocationArray195.8 us1.16 us1.08 us2,414 B1.14 KB

Diff: https://www.diffchecker.com/bNP4qHdF/

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 15, 2024
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jul 15, 2024
Comment threadsrc/coreclr/jit/objectalloc.h Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated

@AndyAyersMSAndyAyersMS left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For arrays (and also perhaps boxes and ref classes) we ought to have some kind of size limit... possibly similar to the one we use for stackallocs.

We need to be careful we don't allocate a lot of stack for an object that might not be heavily used, as we'll pay per-call prolog zeroing costs.

Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/gentree.cpp Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated
#ifdef FEATURE_READYTORUN
if (comp->opts.IsReadyToRun() && data->IsHelperCall(comp, CORINFO_HELP_READYTORUN_NEWARR_1))
{
len = data->AsCall()->gtArgs.GetArgByIndex(0)->GetNode();

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Dead code?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I removed R2R support more completely

Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
Comment on lines +2849 to +2854
GenTree* const mtStore = gtNewStoreLclFldNode(lclNum, TYP_I_IMPL, 0, mt);
Statement* const mtStmt = gtNewStmt(mtStore);

fgInsertStmtBefore(block, newStmt, mtStmt);
gtSetStmtInfo(mtStmt);
fgSetStmtSeq(mtStmt);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
GenTree* const mtStore = gtNewStoreLclFldNode(lclNum, TYP_I_IMPL, 0, mt);
Statement* const mtStmt = gtNewStmt(mtStore);
fgInsertStmtBefore(block, newStmt, mtStmt);
gtSetStmtInfo(mtStmt);
fgSetStmtSeq(mtStmt);
GenTree* const mtStore = gtNewStoreLclFldNode(lclNum, TYP_I_IMPL, 0, mt);
Statement* const mtStmt = fgNewStmtFromTree(mtStore);
fgInsertStmtBefore(block, newStmt, mtStmt);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ditto below, more things can switch to use fgNewStmtFromTree

Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
@@ -181,6 +254,11 @@ inline bool ObjectAllocator::CanAllocateLclVarOnStack(unsigned int lclNu
return false;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems like we should have a limit on the aggregate stack allocated size. That's somewhat preexisting, but probably more important now.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agree but will do that in a future PR

Comment on lines +6249 to +6260
// If this is a local array, there are no asyncronous modifications, so we can set the
// conservative VN to the liberal VN.
//
VNFuncApp arrFn;
if (vnStore->IsVNNewLocalArr(arrVN, &arrFn))
{
loadTree->gtVNPair.SetConservative(loadValueVN);
}
else
{
loadTree->gtVNPair.SetConservative(vnStore->VNForExpr(compCurBB, loadType));
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does this actually show up as benefits?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In some limited cases, yes... eg we can const propagate through a[2] = 1; y = a[2];

@AndyAyersMS

Copy link
Copy Markdown
Member

@jakobbotsch think I've addressed most of the key points. Overall size limit will come in a future PR.

@jakobbotsch

Copy link
Copy Markdown
Member

@AndyAyersMS Did you push those changes?

@AndyAyersMS

Copy link
Copy Markdown
Member

Ah, I pushed to my fork, but ... this PR is not from my fork.

@AndyAyersMS

Copy link
Copy Markdown
Member

@jakobbotsch changes are there now

GTF_CALL_M_CAST_CAN_BE_EXPANDED = 0x04000000, // this cast (helper call) can be expanded if it's profitable. To be removed.
GTF_CALL_M_CAST_OBJ_NONNULL = 0x08000000, // if we expand this specific cast we don't need to check the input object for null
// NOTE: if needed, this flag can be removed, and we can introduce new _NONNUL cast helpers
GTF_CALL_M_STACK_ARRAY = 0x10000000, // this call is a new array helper for a stack allocated array.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I suppose we weren't able to get rid of this since it's also used in VN.

Comment on lines +14375 to +14380
// modHeap = false;
}
else if (vnf == VNF_JitReadyToRunNewArr)
{
vnf = VNF_JitReadyToRunNewLclArr;
// modHeap = false;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: commented code (feel free to remove as part of a follow up)

}
}

if (isAlloc && ((call->gtCallMoreFlags & GTF_CALL_M_STACK_ARRAY) != 0))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I suppose this could be changed to FindWellKnownArg(StackArrayLocal) != nullptr to get rid of the flag

@jakobbotschjakobbotsch left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@AndyAyersMS

Copy link
Copy Markdown
Member

Thanks. I have more changes in this area so I can handle the last few bits in a subsequent PR.

@AndyAyersMS
AndyAyersMS merged commit 7d75878 into dotnet:mainJan 22, 2025
@AndyAyersMS

Copy link
Copy Markdown
Member

@hez2010 thanks for all the work you did here.

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@hez2010@AndyAyersMS@JulieLeeMSFT@jakobbotsch
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

JIT: Extend escape analysis to account for arrays with non-gcref elements - #104906

Merged
AndyAyersMS merged 97 commits into
dotnet:mainfrom
hez2010:value-array-stack-alloc
Jan 22, 2025
Merged

JIT: Extend escape analysis to account for arrays with non-gcref elements#104906
AndyAyersMS merged 97 commits into
dotnet:mainfrom
hez2010:value-array-stack-alloc

Conversation

@hez2010

@hez2010hez2010 commented Jul 15, 2024

Copy link
Copy Markdown
Contributor

Positive case:

varchs=newchar[42];chs[1]='a';Console.WriteLine((int)chs[1]+chs.Length);

Codegen:

; Assembly listing for method ArrayAllocator.Program:Main() (FullOpts); Emitting BLENDED_CODE for X64 with AVX - Windows; FullOpts code; optimized code; rsp based frame; partially interruptible; No PGO data; Final local variable assignments;;* V00 loc0 [V00 ] ( 0, 0 ) long -> zero-ref class-hnd exact <short[]>; V01 OutArgs [V01 ] ( 1, 1 ) struct (32) [rsp+0x00] do-not-enreg[XS] addr-exposed "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) struct (104) zero-ref do-not-enreg[SF] "stack allocated array temp";* V03 tmp2 [V03 ] ( 0, 0 ) long -> zero-ref single-def "V02.[000..008)";* V04 tmp3 [V04 ] ( 0, 0 ) int -> zero-ref single-def "V02.[008..012)";* V05 tmp4 [V05 ] ( 0, 0 ) short -> zero-ref "V02.[018..020)";; Lcl frame size = 40G_M25548_IG01: ;; offset=0x0000subrsp,40 ;; size=4 bbWeight=1 PerfScore 0.25G_M25548_IG02: ;; offset=0x0004movecx,84call[System.Console:WriteLine(int)]nop ;; size=12 bbWeight=1 PerfScore 3.50G_M25548_IG03: ;; offset=0x0010addrsp,40ret ;; size=5 bbWeight=1 PerfScore 1.25; Total bytes of code 21, prolog size 4, PerfScore 5.00, instruction count 6, allocated bytes for code 21 (MethodHash=5b0b9c33) for method ArrayAllocator.Program:Main() (FullOpts)

Negative case:

varchs=newchar[42];chs[1]='a';Console.WriteLine((int)chs[42]+chs.Length);

Codegen:

; Assembly listing for method ArrayAllocator.Program:Main() (FullOpts); Emitting BLENDED_CODE for X64 with AVX - Windows; FullOpts code; optimized code; rsp based frame; partially interruptible; No PGO data; Final local variable assignments;;* V00 loc0 [V00 ] ( 0, 0 ) long -> zero-ref class-hnd exact <short[]>; V01 OutArgs [V01 ] ( 1, 1 ) struct (32) [rsp+0x00] do-not-enreg[XS] addr-exposed "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) struct (104) zero-ref do-not-enreg[SF] "stack allocated array temp"; V03 tmp2 [V03,T00] ( 1, 0 ) byref -> rbx must-init "dummy temp of must thrown exception";* V04 tmp3 [V04 ] ( 0, 0 ) long -> zero-ref single-def "V02.[000..008)";* V05 tmp4 [V05 ] ( 0, 0 ) int -> zero-ref single-def "V02.[008..012)";* V06 tmp5 [V06 ] ( 0, 0 ) short -> zero-ref single-def "V02.[018..020)";; Lcl frame size = 32G_M25548_IG01: ;; offset=0x0000pushrbxsubrsp,32xorebx,ebx ;; size=7 bbWeight=0 PerfScore 0.00G_M25548_IG02: ;; offset=0x0007call CORINFO_HELP_RNGCHKFAILmovsxrcx, word ptr [rbx]call[System.Console:WriteLine(int)]int3 ;; size=16 bbWeight=0 PerfScore 0.00; Total bytes of code 23, prolog size 5, PerfScore 0.00, instruction count 7, allocated bytes for code 23 (MethodHash=5b0b9c33) for method ArrayAllocator.Program:Main() (FullOpts); ============================================================

Benchmark on Mandelbrot:

MethodJobMeanErrorStdDevCode SizeAllocated
MandelBrotNoStackAllocationArray199.7 us1.30 us1.22 us1,996 B2.49 KB
MandelBrotStackAllocationArray195.8 us1.16 us1.08 us2,414 B1.14 KB

Diff: https://www.diffchecker.com/bNP4qHdF/

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 15, 2024
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jul 15, 2024
Comment threadsrc/coreclr/jit/objectalloc.h Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated

@AndyAyersMSAndyAyersMS left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For arrays (and also perhaps boxes and ref classes) we ought to have some kind of size limit... possibly similar to the one we use for stackallocs.

We need to be careful we don't allocate a lot of stack for an object that might not be heavily used, as we'll pay per-call prolog zeroing costs.

Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/gentree.cpp Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated
#ifdef FEATURE_READYTORUN
if (comp->opts.IsReadyToRun() && data->IsHelperCall(comp, CORINFO_HELP_READYTORUN_NEWARR_1))
{
len = data->AsCall()->gtArgs.GetArgByIndex(0)->GetNode();

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Dead code?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I removed R2R support more completely

Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
Comment on lines +2849 to +2854
GenTree* const mtStore = gtNewStoreLclFldNode(lclNum, TYP_I_IMPL, 0, mt);
Statement* const mtStmt = gtNewStmt(mtStore);

fgInsertStmtBefore(block, newStmt, mtStmt);
gtSetStmtInfo(mtStmt);
fgSetStmtSeq(mtStmt);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
GenTree* const mtStore = gtNewStoreLclFldNode(lclNum, TYP_I_IMPL, 0, mt);
Statement* const mtStmt = gtNewStmt(mtStore);
fgInsertStmtBefore(block, newStmt, mtStmt);
gtSetStmtInfo(mtStmt);
fgSetStmtSeq(mtStmt);
GenTree* const mtStore = gtNewStoreLclFldNode(lclNum, TYP_I_IMPL, 0, mt);
Statement* const mtStmt = fgNewStmtFromTree(mtStore);
fgInsertStmtBefore(block, newStmt, mtStmt);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ditto below, more things can switch to use fgNewStmtFromTree

Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
@@ -181,6 +254,11 @@ inline bool ObjectAllocator::CanAllocateLclVarOnStack(unsigned int lclNu
return false;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems like we should have a limit on the aggregate stack allocated size. That's somewhat preexisting, but probably more important now.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agree but will do that in a future PR

Comment on lines +6249 to +6260
// If this is a local array, there are no asyncronous modifications, so we can set the
// conservative VN to the liberal VN.
//
VNFuncApp arrFn;
if (vnStore->IsVNNewLocalArr(arrVN, &arrFn))
{
loadTree->gtVNPair.SetConservative(loadValueVN);
}
else
{
loadTree->gtVNPair.SetConservative(vnStore->VNForExpr(compCurBB, loadType));
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does this actually show up as benefits?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In some limited cases, yes... eg we can const propagate through a[2] = 1; y = a[2];

@AndyAyersMS

Copy link
Copy Markdown
Member

@jakobbotsch think I've addressed most of the key points. Overall size limit will come in a future PR.

@jakobbotsch

Copy link
Copy Markdown
Member

@AndyAyersMS Did you push those changes?

@AndyAyersMS

Copy link
Copy Markdown
Member

Ah, I pushed to my fork, but ... this PR is not from my fork.

@AndyAyersMS

Copy link
Copy Markdown
Member

@jakobbotsch changes are there now

GTF_CALL_M_CAST_CAN_BE_EXPANDED = 0x04000000, // this cast (helper call) can be expanded if it's profitable. To be removed.
GTF_CALL_M_CAST_OBJ_NONNULL = 0x08000000, // if we expand this specific cast we don't need to check the input object for null
// NOTE: if needed, this flag can be removed, and we can introduce new _NONNUL cast helpers
GTF_CALL_M_STACK_ARRAY = 0x10000000, // this call is a new array helper for a stack allocated array.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I suppose we weren't able to get rid of this since it's also used in VN.

Comment on lines +14375 to +14380
// modHeap = false;
}
else if (vnf == VNF_JitReadyToRunNewArr)
{
vnf = VNF_JitReadyToRunNewLclArr;
// modHeap = false;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: commented code (feel free to remove as part of a follow up)

}
}

if (isAlloc && ((call->gtCallMoreFlags & GTF_CALL_M_STACK_ARRAY) != 0))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I suppose this could be changed to FindWellKnownArg(StackArrayLocal) != nullptr to get rid of the flag

@jakobbotschjakobbotsch left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@AndyAyersMS

Copy link
Copy Markdown
Member

Thanks. I have more changes in this area so I can handle the last few bits in a subsequent PR.

@AndyAyersMS
AndyAyersMS merged commit 7d75878 into dotnet:mainJan 22, 2025
@AndyAyersMS

Copy link
Copy Markdown
Member

@hez2010 thanks for all the work you did here.

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@hez2010@AndyAyersMS@JulieLeeMSFT@jakobbotsch
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

JIT: Extend escape analysis to account for arrays with non-gcref elements - #104906

Merged
AndyAyersMS merged 97 commits into
dotnet:mainfrom
hez2010:value-array-stack-alloc
Jan 22, 2025
Merged

JIT: Extend escape analysis to account for arrays with non-gcref elements#104906
AndyAyersMS merged 97 commits into
dotnet:mainfrom
hez2010:value-array-stack-alloc

Conversation

@hez2010

@hez2010hez2010 commented Jul 15, 2024

Copy link
Copy Markdown
Contributor

Positive case:

varchs=newchar[42];chs[1]='a';Console.WriteLine((int)chs[1]+chs.Length);

Codegen:

; Assembly listing for method ArrayAllocator.Program:Main() (FullOpts); Emitting BLENDED_CODE for X64 with AVX - Windows; FullOpts code; optimized code; rsp based frame; partially interruptible; No PGO data; Final local variable assignments;;* V00 loc0 [V00 ] ( 0, 0 ) long -> zero-ref class-hnd exact <short[]>; V01 OutArgs [V01 ] ( 1, 1 ) struct (32) [rsp+0x00] do-not-enreg[XS] addr-exposed "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) struct (104) zero-ref do-not-enreg[SF] "stack allocated array temp";* V03 tmp2 [V03 ] ( 0, 0 ) long -> zero-ref single-def "V02.[000..008)";* V04 tmp3 [V04 ] ( 0, 0 ) int -> zero-ref single-def "V02.[008..012)";* V05 tmp4 [V05 ] ( 0, 0 ) short -> zero-ref "V02.[018..020)";; Lcl frame size = 40G_M25548_IG01: ;; offset=0x0000subrsp,40 ;; size=4 bbWeight=1 PerfScore 0.25G_M25548_IG02: ;; offset=0x0004movecx,84call[System.Console:WriteLine(int)]nop ;; size=12 bbWeight=1 PerfScore 3.50G_M25548_IG03: ;; offset=0x0010addrsp,40ret ;; size=5 bbWeight=1 PerfScore 1.25; Total bytes of code 21, prolog size 4, PerfScore 5.00, instruction count 6, allocated bytes for code 21 (MethodHash=5b0b9c33) for method ArrayAllocator.Program:Main() (FullOpts)

Negative case:

varchs=newchar[42];chs[1]='a';Console.WriteLine((int)chs[42]+chs.Length);

Codegen:

; Assembly listing for method ArrayAllocator.Program:Main() (FullOpts); Emitting BLENDED_CODE for X64 with AVX - Windows; FullOpts code; optimized code; rsp based frame; partially interruptible; No PGO data; Final local variable assignments;;* V00 loc0 [V00 ] ( 0, 0 ) long -> zero-ref class-hnd exact <short[]>; V01 OutArgs [V01 ] ( 1, 1 ) struct (32) [rsp+0x00] do-not-enreg[XS] addr-exposed "OutgoingArgSpace";* V02 tmp1 [V02 ] ( 0, 0 ) struct (104) zero-ref do-not-enreg[SF] "stack allocated array temp"; V03 tmp2 [V03,T00] ( 1, 0 ) byref -> rbx must-init "dummy temp of must thrown exception";* V04 tmp3 [V04 ] ( 0, 0 ) long -> zero-ref single-def "V02.[000..008)";* V05 tmp4 [V05 ] ( 0, 0 ) int -> zero-ref single-def "V02.[008..012)";* V06 tmp5 [V06 ] ( 0, 0 ) short -> zero-ref single-def "V02.[018..020)";; Lcl frame size = 32G_M25548_IG01: ;; offset=0x0000pushrbxsubrsp,32xorebx,ebx ;; size=7 bbWeight=0 PerfScore 0.00G_M25548_IG02: ;; offset=0x0007call CORINFO_HELP_RNGCHKFAILmovsxrcx, word ptr [rbx]call[System.Console:WriteLine(int)]int3 ;; size=16 bbWeight=0 PerfScore 0.00; Total bytes of code 23, prolog size 5, PerfScore 0.00, instruction count 7, allocated bytes for code 23 (MethodHash=5b0b9c33) for method ArrayAllocator.Program:Main() (FullOpts); ============================================================

Benchmark on Mandelbrot:

MethodJobMeanErrorStdDevCode SizeAllocated
MandelBrotNoStackAllocationArray199.7 us1.30 us1.22 us1,996 B2.49 KB
MandelBrotStackAllocationArray195.8 us1.16 us1.08 us2,414 B1.14 KB

Diff: https://www.diffchecker.com/bNP4qHdF/

@ghostghost added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Jul 15, 2024
@dotnet-policy-servicedotnet-policy-serviceBot added the community-contribution Indicates that the PR has been added by a community member label Jul 15, 2024
Comment threadsrc/coreclr/jit/objectalloc.h Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated

@AndyAyersMSAndyAyersMS left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For arrays (and also perhaps boxes and ref classes) we ought to have some kind of size limit... possibly similar to the one we use for stackallocs.

We need to be careful we don't allocate a lot of stack for an object that might not be heavily used, as we'll pay per-call prolog zeroing costs.

Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/lclmorph.cpp Outdated
Comment threadsrc/coreclr/jit/gentree.cpp Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated
Comment threadsrc/coreclr/jit/objectalloc.cpp Outdated
#ifdef FEATURE_READYTORUN
if (comp->opts.IsReadyToRun() && data->IsHelperCall(comp, CORINFO_HELP_READYTORUN_NEWARR_1))
{
len = data->AsCall()->gtArgs.GetArgByIndex(0)->GetNode();

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Dead code?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I removed R2R support more completely

Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
Comment on lines +2849 to +2854
GenTree* const mtStore = gtNewStoreLclFldNode(lclNum, TYP_I_IMPL, 0, mt);
Statement* const mtStmt = gtNewStmt(mtStore);

fgInsertStmtBefore(block, newStmt, mtStmt);
gtSetStmtInfo(mtStmt);
fgSetStmtSeq(mtStmt);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
GenTree* const mtStore = gtNewStoreLclFldNode(lclNum, TYP_I_IMPL, 0, mt);
Statement* const mtStmt = gtNewStmt(mtStore);
fgInsertStmtBefore(block, newStmt, mtStmt);
gtSetStmtInfo(mtStmt);
fgSetStmtSeq(mtStmt);
GenTree* const mtStore = gtNewStoreLclFldNode(lclNum, TYP_I_IMPL, 0, mt);
Statement* const mtStmt = fgNewStmtFromTree(mtStore);
fgInsertStmtBefore(block, newStmt, mtStmt);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ditto below, more things can switch to use fgNewStmtFromTree

Comment threadsrc/coreclr/jit/helperexpansion.cpp Outdated
@@ -181,6 +254,11 @@ inline bool ObjectAllocator::CanAllocateLclVarOnStack(unsigned int lclNu
return false;
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems like we should have a limit on the aggregate stack allocated size. That's somewhat preexisting, but probably more important now.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agree but will do that in a future PR

Comment on lines +6249 to +6260
// If this is a local array, there are no asyncronous modifications, so we can set the
// conservative VN to the liberal VN.
//
VNFuncApp arrFn;
if (vnStore->IsVNNewLocalArr(arrVN, &arrFn))
{
loadTree->gtVNPair.SetConservative(loadValueVN);
}
else
{
loadTree->gtVNPair.SetConservative(vnStore->VNForExpr(compCurBB, loadType));
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does this actually show up as benefits?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In some limited cases, yes... eg we can const propagate through a[2] = 1; y = a[2];

@AndyAyersMS

Copy link
Copy Markdown
Member

@jakobbotsch think I've addressed most of the key points. Overall size limit will come in a future PR.

@jakobbotsch

Copy link
Copy Markdown
Member

@AndyAyersMS Did you push those changes?

@AndyAyersMS

Copy link
Copy Markdown
Member

Ah, I pushed to my fork, but ... this PR is not from my fork.

@AndyAyersMS

Copy link
Copy Markdown
Member

@jakobbotsch changes are there now

GTF_CALL_M_CAST_CAN_BE_EXPANDED = 0x04000000, // this cast (helper call) can be expanded if it's profitable. To be removed.
GTF_CALL_M_CAST_OBJ_NONNULL = 0x08000000, // if we expand this specific cast we don't need to check the input object for null
// NOTE: if needed, this flag can be removed, and we can introduce new _NONNUL cast helpers
GTF_CALL_M_STACK_ARRAY = 0x10000000, // this call is a new array helper for a stack allocated array.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I suppose we weren't able to get rid of this since it's also used in VN.

Comment on lines +14375 to +14380
// modHeap = false;
}
else if (vnf == VNF_JitReadyToRunNewArr)
{
vnf = VNF_JitReadyToRunNewLclArr;
// modHeap = false;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: commented code (feel free to remove as part of a follow up)

}
}

if (isAlloc && ((call->gtCallMoreFlags & GTF_CALL_M_STACK_ARRAY) != 0))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I suppose this could be changed to FindWellKnownArg(StackArrayLocal) != nullptr to get rid of the flag

@jakobbotschjakobbotsch left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@AndyAyersMS

Copy link
Copy Markdown
Member

Thanks. I have more changes in this area so I can handle the last few bits in a subsequent PR.

@AndyAyersMS
AndyAyersMS merged commit 7d75878 into dotnet:mainJan 22, 2025
@AndyAyersMS

Copy link
Copy Markdown
Member

@hez2010 thanks for all the work you did here.

Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIcommunity-contributionIndicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@hez2010@AndyAyersMS@JulieLeeMSFT@jakobbotsch