[NativeAOT] Make casting logic closer to CoreCLR - #89548

Closed
VSadov wants to merge 12 commits into
dotnet:mainfrom
VSadov:casts
Closed

[NativeAOT] Make casting logic closer to CoreCLR#89548
VSadov wants to merge 12 commits into
dotnet:mainfrom
VSadov:casts

Conversation

@VSadov

@VSadovVSadov commented Jul 27, 2023

Copy link
Copy Markdown
Member

Fixes: #84464

The actual changes are not as big here as might seem. For the most part this is a refactoring of existing code to have a shape closer to CoreCLR cast helpers, so that similar patterns could be used - in a few places where that has not been done already in earlier changes.
For example cases like CheckCastAny - could start with a cache lookup, since uncached code path can be complex and thus relatively slow. (also addresses some old TODOs in this area)

@ghostghost assigned VSadovJul 27, 2023
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

Issue Details

Fixes: #84464

In progress.

For the most part this is a refactoring of existing code to have a shape closer to CoreCLR. In particular to do cache lookups earlier.
Cases like CheckCastAny - we should basically start with a cache lookup, since uncached path can be relatively slow.

Author:VSadov
Assignees:-
Labels:

area-NativeAOT-coreclr

Milestone:-

@VSadov

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@VSadov

Copy link
Copy Markdown
MemberAuthor

I think this is ready for a review.

Comment threadsrc/coreclr/nativeaot/Runtime.Base/src/System/Runtime/TypeCast.cs Outdated
@VSadov
VSadov requested a review from jkotasAugust 1, 2023 18:33
@jkotas

Copy link
Copy Markdown
Member

Could you please collect numbers for casting microbenchmarks before/after this change?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can delete CORINFO_HELP_CHKCASTARRAY and CORINFO_HELP_ISINSTANCEOFARRAY from JIT/EE interface. They are unnecessary now.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

One possible question here would be - Why not just port/share the CoreCLR casting helpers in their entirety and implement the internal calls in managed code - to have nearly the same implementation?

I think one of the factors for the design of CoreCLR managed casting helpers was to avoid complicated API with the native type system. As a result that API is basically 2 internal calls methods - IsInstanceOfAny_NoCacheLookup and ChkCastAny_NoCacheLookup.

In NativeAOT the type system APIs are easily accessible, so we do not need to minimize the use of those APIs. On the other hand there are some differences, like the way we fetch the base type for arrays, that may stand in the way of code sharing.
Thus I did not consider it is as a goal to share the code or making it maximally similar when possible, but I think it can be done.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

Also in NativeAOT these casting helpers have more users with additional needs, while in CoreClr it is really just type system facade for object casting.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

Could you please collect numbers for casting microbenchmarks before/after this change?

I was thinking of what we could measure here.
The most common cases of casting like regular object/interfaces casts did not change, so unlikely to see any differences. We may see differences in more complex cases - like casting to variant interfaces or arrays.

Perhaps just running the regular casting benchmarks that perf lab uses would be informative enough about what changed perf-wise.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

I have run the same benchmark as in #84430 (comment)

internalclassProgram{constintiters=1000000;staticvoidMain(string[]args){for(;;){Time(TestLStringToIROCstring);}}staticvoidTime(Actiona){varsw=Stopwatch.StartNew();for(inti=0;i<100;i++){a();}sw.Stop();System.Console.WriteLine(sw.ElapsedMilliseconds);}staticobjecto=newList<string>();staticvoidTestLStringToIROCstring(){for(inti=0;i<iters;i++){if(oasIReadOnlyCollection<object>==null)thrownull;if(oasIReadOnlyCollection<string>==null)thrownull;if(oasIEnumerable<object>==null)thrownull;if(oasIEnumerable<string>==null)thrownull;}}}

=== before the change

1451
1460
1456
1458
1470
1463

=== after the change:

799
801
798
799
800
799

The reason for the difference is that original code makes a number of calls. Profiler shows:
TypeCast__IsInstanceOf
MethodTable_get_IsArray
TypeCast__IsInstanceOfVariantType - this one does the cache lookup

making calls and additional checks adds up.

In the new implementation there is only one helper call in the profile:
TypeCast__IsInstanceOfAny - this one does the cache lookup

@VSadov

Copy link
Copy Markdown
MemberAuthor

For comparison the CoreCLR is a bit faster.

The same benchmark as above produces (smaller is better):

558
561
562
557
561

As I see in the debugger the native code that we run after this change is nearly the same between CoreCLR and NativeAOT.
There are minor differences like loading of type pointers/handles.
In JIT code they look like:

movrcx,7FFCA19E79E8h

In NativeAOT loading a type looks like:

learcx,[rip+0x73d65]

I do not see any other significant differences. Maybe it is just these little diffs and some indirect impact on code size or alignment that makes the difference.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

Another thing to notice is that compared to #84430 (comment) and prior to this change it looks like the benchmark in the comment has regressed.

That was possibly caused by #86029 . I suspect it made common/simple cases faster, which is good, but regressed complex cases that rely on caching as more checks like MethodTable_get_IsArray could run before eventually hitting the cache. Just a guess though.

Anyways, it looks like after this PR the complex/cached case is faster than in #84430 (comment)

Comment threadsrc/coreclr/inc/corinfo.h Outdated
Comment threadsrc/coreclr/jit/importer.cpp Outdated
@VSadov

Copy link
Copy Markdown
MemberAuthor

There are some superpmi failures. I can't tell if that simply tells that there are codegen diffs or something is crashing.

@VSadov
VSadovforce-pushed the casts branch 2 times, most recently from c3672cb to ac36d4eCompareAugust 6, 2023 21:51
@jkotas

Copy link
Copy Markdown
Member

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

Comment threadsrc/libraries/System.Reflection/tests/GetTypeTests.cs Outdated
Comment threadsrc/libraries/System.Reflection/tests/GetTypeTests.cs Outdated
@VSadov

Copy link
Copy Markdown
MemberAuthor

I'd be ok with merging this now. I think all concerns have been resolved. Let me know if there is something that may be missing.

Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
_ASSERTE(helper == CORINFO_HELP_ISINSTANCEOFANY);
}
else
if (!clsHnd.IsTypeDesc() && !Nullable::IsNullableType(clsHnd))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The native AOT implementation does not checks for Nullable. Is this check redundant here or is the check for Nullable missing in native AOT?

@VSadovVSadovAug 8, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The difference is that JIT tries to transform simple cases like o is int? into o is int, but can't do that when needs a type lookup. As I understand that is due to limitations of IR.

// ECMA-335 III.4.3: If typeTok is a nullable type, Nullable<T>, it is interpreted as "boxed" T
// We can convert constant-ish tokens of nullable to its underlying type.
// However, when the type is shared generic parameter like Nullable<Struct<__Canon>>, the actual type will require
// runtime lookup. It's too complex to add another level of indirection in op2, fallback to the cast helper instead.
if (isClassExact && !(info.compCompHnd->getClassAttribs(pResolvedToken->hClass) & CORINFO_FLG_SHAREDINST))

NativeAOT does not seem to have a problem with expressing such lookup, so we always cast with nullable stripped

caseReadyToRunHelperId.TypeHandleForCasting:
{
vartype=(TypeDesc)targetOfLookup;
if(type.IsNullable)
targetOfLookup=type.Instantiation[0];
returnNecessaryTypeSymbolIfPossible((TypeDesc)targetOfLookup);

@VSadovVSadovAug 9, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In other words in NativeAOT the type check will always be to an unboxed underlying type, thus CLASS variant is correct, and also optimizable, since T? is isClassExact.

In CoreCLR the type check in complex cases could be to the actual Nullable<SomeStruct<string>>, thus it should use ANY variant. It also means that we may be introducing an optimization bug, since such IsInst cannot be lowered to a type/handle compare.

I will check if that is a case, or if there are some mitigating reasons why it still works correctly.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sadly that is the case. The following works incorrectly

=====Prints:TrueFalseusingSystem.Runtime.CompilerServices;namespaceConsoleApp34{structS1<T>{publicTvalue;}internalclassProgram{staticobjecto;staticvoidMain(string[]args){Test<int>();Test<string>();}[MethodImpl(MethodImplOptions.AggressiveOptimization)]privatestaticvoidTest<T>(){o=newS1<T>();Console.WriteLine(oisS1<T>?);}}}

This PR indirectly enabled casting optimizations for cases like o is int?, but we need to suppress it, since in more general cases it does not work correctly.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Unless there are better ideas, I am thinking of constraining isinst optimization for CORINFO_HELP_ISINSTANCEOFANY only if converting to an array type. So we do not keep finding more broken cases.

That may be too conservative, but we should probably stay closer to the preexisting behavior for now and consider if more cases can work in 9.0

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am thinking that this PR as a whole is too risky for 8.0. Are there parts that we think are critical to get into .NET 8?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think the actual casting helper change was low risk as that is mostly refactoring to get more cases to hit the cache earlier. Touching the JIT appears to be a lot more fragile.

There is nothing really "critical" for 8.0, as in - we are not fixing some complete showstoppers here.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have created a "reduced" version of this. I think that is what we can consider for 8.0 - #90234

I have removed all the JIT changes, but kept the added codegen test.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good.

{
DWORD flags = info.compCompHnd->getClassAttribs(classHnd);
DWORD flagsMask = CORINFO_FLG_FINAL | CORINFO_FLG_VARIANCE | CORINFO_FLG_ARRAY;
DWORD flagsMask = CORINFO_FLG_FINAL | CORINFO_FLG_VARIANCE | CORINFO_FLG_TYPE_EQUIVALENCE | CORINFO_FLG_ARRAY;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both CORINFO_FLG_VARIANCE and CORINFO_FLG_TYPE_EQUIVALENCE are only computed to make the impIsClassExact work. Computing these flags is a waste in all other cases. It is the kind of pattern that calls for introduction of dedicated JIT/EE interface API that replaces the flags.

@VSadov

Copy link
Copy Markdown
MemberAuthor

A reduced version of this affecting only run time behavior has been merged.
For the further improvements for the JIT API in this area a tracking bug has been added - #91016

@VSadovVSadov closed this Aug 23, 2023
@jkotasjkotas mentioned this pull request Aug 23, 2023
@ghostghost locked as resolved and limited conversation to collaborators Sep 23, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[NativeAOT] Possible perf improvements in casting

4 participants

@VSadov@jkotas@jakobbotsch@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

[NativeAOT] Make casting logic closer to CoreCLR - #89548

Closed
VSadov wants to merge 12 commits into
dotnet:mainfrom
VSadov:casts
Closed

[NativeAOT] Make casting logic closer to CoreCLR#89548
VSadov wants to merge 12 commits into
dotnet:mainfrom
VSadov:casts

Conversation

@VSadov

@VSadovVSadov commented Jul 27, 2023

Copy link
Copy Markdown
Member

Fixes: #84464

The actual changes are not as big here as might seem. For the most part this is a refactoring of existing code to have a shape closer to CoreCLR cast helpers, so that similar patterns could be used - in a few places where that has not been done already in earlier changes.
For example cases like CheckCastAny - could start with a cache lookup, since uncached code path can be complex and thus relatively slow. (also addresses some old TODOs in this area)

@ghostghost assigned VSadovJul 27, 2023
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

Issue Details

Fixes: #84464

In progress.

For the most part this is a refactoring of existing code to have a shape closer to CoreCLR. In particular to do cache lookups earlier.
Cases like CheckCastAny - we should basically start with a cache lookup, since uncached path can be relatively slow.

Author:VSadov
Assignees:-
Labels:

area-NativeAOT-coreclr

Milestone:-

@VSadov

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@VSadov

Copy link
Copy Markdown
MemberAuthor

I think this is ready for a review.

Comment threadsrc/coreclr/nativeaot/Runtime.Base/src/System/Runtime/TypeCast.cs Outdated
@VSadov
VSadov requested a review from jkotasAugust 1, 2023 18:33
@jkotas

Copy link
Copy Markdown
Member

Could you please collect numbers for casting microbenchmarks before/after this change?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can delete CORINFO_HELP_CHKCASTARRAY and CORINFO_HELP_ISINSTANCEOFARRAY from JIT/EE interface. They are unnecessary now.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

One possible question here would be - Why not just port/share the CoreCLR casting helpers in their entirety and implement the internal calls in managed code - to have nearly the same implementation?

I think one of the factors for the design of CoreCLR managed casting helpers was to avoid complicated API with the native type system. As a result that API is basically 2 internal calls methods - IsInstanceOfAny_NoCacheLookup and ChkCastAny_NoCacheLookup.

In NativeAOT the type system APIs are easily accessible, so we do not need to minimize the use of those APIs. On the other hand there are some differences, like the way we fetch the base type for arrays, that may stand in the way of code sharing.
Thus I did not consider it is as a goal to share the code or making it maximally similar when possible, but I think it can be done.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

Also in NativeAOT these casting helpers have more users with additional needs, while in CoreClr it is really just type system facade for object casting.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

Could you please collect numbers for casting microbenchmarks before/after this change?

I was thinking of what we could measure here.
The most common cases of casting like regular object/interfaces casts did not change, so unlikely to see any differences. We may see differences in more complex cases - like casting to variant interfaces or arrays.

Perhaps just running the regular casting benchmarks that perf lab uses would be informative enough about what changed perf-wise.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

I have run the same benchmark as in #84430 (comment)

internalclassProgram{constintiters=1000000;staticvoidMain(string[]args){for(;;){Time(TestLStringToIROCstring);}}staticvoidTime(Actiona){varsw=Stopwatch.StartNew();for(inti=0;i<100;i++){a();}sw.Stop();System.Console.WriteLine(sw.ElapsedMilliseconds);}staticobjecto=newList<string>();staticvoidTestLStringToIROCstring(){for(inti=0;i<iters;i++){if(oasIReadOnlyCollection<object>==null)thrownull;if(oasIReadOnlyCollection<string>==null)thrownull;if(oasIEnumerable<object>==null)thrownull;if(oasIEnumerable<string>==null)thrownull;}}}

=== before the change

1451
1460
1456
1458
1470
1463

=== after the change:

799
801
798
799
800
799

The reason for the difference is that original code makes a number of calls. Profiler shows:
TypeCast__IsInstanceOf
MethodTable_get_IsArray
TypeCast__IsInstanceOfVariantType - this one does the cache lookup

making calls and additional checks adds up.

In the new implementation there is only one helper call in the profile:
TypeCast__IsInstanceOfAny - this one does the cache lookup

@VSadov

Copy link
Copy Markdown
MemberAuthor

For comparison the CoreCLR is a bit faster.

The same benchmark as above produces (smaller is better):

558
561
562
557
561

As I see in the debugger the native code that we run after this change is nearly the same between CoreCLR and NativeAOT.
There are minor differences like loading of type pointers/handles.
In JIT code they look like:

movrcx,7FFCA19E79E8h

In NativeAOT loading a type looks like:

learcx,[rip+0x73d65]

I do not see any other significant differences. Maybe it is just these little diffs and some indirect impact on code size or alignment that makes the difference.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

Another thing to notice is that compared to #84430 (comment) and prior to this change it looks like the benchmark in the comment has regressed.

That was possibly caused by #86029 . I suspect it made common/simple cases faster, which is good, but regressed complex cases that rely on caching as more checks like MethodTable_get_IsArray could run before eventually hitting the cache. Just a guess though.

Anyways, it looks like after this PR the complex/cached case is faster than in #84430 (comment)

Comment threadsrc/coreclr/inc/corinfo.h Outdated
Comment threadsrc/coreclr/jit/importer.cpp Outdated
@VSadov

Copy link
Copy Markdown
MemberAuthor

There are some superpmi failures. I can't tell if that simply tells that there are codegen diffs or something is crashing.

@VSadov
VSadovforce-pushed the casts branch 2 times, most recently from c3672cb to ac36d4eCompareAugust 6, 2023 21:51
@jkotas

Copy link
Copy Markdown
Member

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

Comment threadsrc/libraries/System.Reflection/tests/GetTypeTests.cs Outdated
Comment threadsrc/libraries/System.Reflection/tests/GetTypeTests.cs Outdated
@VSadov

Copy link
Copy Markdown
MemberAuthor

I'd be ok with merging this now. I think all concerns have been resolved. Let me know if there is something that may be missing.

Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
_ASSERTE(helper == CORINFO_HELP_ISINSTANCEOFANY);
}
else
if (!clsHnd.IsTypeDesc() && !Nullable::IsNullableType(clsHnd))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The native AOT implementation does not checks for Nullable. Is this check redundant here or is the check for Nullable missing in native AOT?

@VSadovVSadovAug 8, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The difference is that JIT tries to transform simple cases like o is int? into o is int, but can't do that when needs a type lookup. As I understand that is due to limitations of IR.

// ECMA-335 III.4.3: If typeTok is a nullable type, Nullable<T>, it is interpreted as "boxed" T
// We can convert constant-ish tokens of nullable to its underlying type.
// However, when the type is shared generic parameter like Nullable<Struct<__Canon>>, the actual type will require
// runtime lookup. It's too complex to add another level of indirection in op2, fallback to the cast helper instead.
if (isClassExact && !(info.compCompHnd->getClassAttribs(pResolvedToken->hClass) & CORINFO_FLG_SHAREDINST))

NativeAOT does not seem to have a problem with expressing such lookup, so we always cast with nullable stripped

caseReadyToRunHelperId.TypeHandleForCasting:
{
vartype=(TypeDesc)targetOfLookup;
if(type.IsNullable)
targetOfLookup=type.Instantiation[0];
returnNecessaryTypeSymbolIfPossible((TypeDesc)targetOfLookup);

@VSadovVSadovAug 9, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In other words in NativeAOT the type check will always be to an unboxed underlying type, thus CLASS variant is correct, and also optimizable, since T? is isClassExact.

In CoreCLR the type check in complex cases could be to the actual Nullable<SomeStruct<string>>, thus it should use ANY variant. It also means that we may be introducing an optimization bug, since such IsInst cannot be lowered to a type/handle compare.

I will check if that is a case, or if there are some mitigating reasons why it still works correctly.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sadly that is the case. The following works incorrectly

=====Prints:TrueFalseusingSystem.Runtime.CompilerServices;namespaceConsoleApp34{structS1<T>{publicTvalue;}internalclassProgram{staticobjecto;staticvoidMain(string[]args){Test<int>();Test<string>();}[MethodImpl(MethodImplOptions.AggressiveOptimization)]privatestaticvoidTest<T>(){o=newS1<T>();Console.WriteLine(oisS1<T>?);}}}

This PR indirectly enabled casting optimizations for cases like o is int?, but we need to suppress it, since in more general cases it does not work correctly.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Unless there are better ideas, I am thinking of constraining isinst optimization for CORINFO_HELP_ISINSTANCEOFANY only if converting to an array type. So we do not keep finding more broken cases.

That may be too conservative, but we should probably stay closer to the preexisting behavior for now and consider if more cases can work in 9.0

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am thinking that this PR as a whole is too risky for 8.0. Are there parts that we think are critical to get into .NET 8?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think the actual casting helper change was low risk as that is mostly refactoring to get more cases to hit the cache earlier. Touching the JIT appears to be a lot more fragile.

There is nothing really "critical" for 8.0, as in - we are not fixing some complete showstoppers here.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have created a "reduced" version of this. I think that is what we can consider for 8.0 - #90234

I have removed all the JIT changes, but kept the added codegen test.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good.

{
DWORD flags = info.compCompHnd->getClassAttribs(classHnd);
DWORD flagsMask = CORINFO_FLG_FINAL | CORINFO_FLG_VARIANCE | CORINFO_FLG_ARRAY;
DWORD flagsMask = CORINFO_FLG_FINAL | CORINFO_FLG_VARIANCE | CORINFO_FLG_TYPE_EQUIVALENCE | CORINFO_FLG_ARRAY;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both CORINFO_FLG_VARIANCE and CORINFO_FLG_TYPE_EQUIVALENCE are only computed to make the impIsClassExact work. Computing these flags is a waste in all other cases. It is the kind of pattern that calls for introduction of dedicated JIT/EE interface API that replaces the flags.

@VSadov

Copy link
Copy Markdown
MemberAuthor

A reduced version of this affecting only run time behavior has been merged.
For the further improvements for the JIT API in this area a tracking bug has been added - #91016

@VSadovVSadov closed this Aug 23, 2023
@jkotasjkotas mentioned this pull request Aug 23, 2023
@ghostghost locked as resolved and limited conversation to collaborators Sep 23, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[NativeAOT] Possible perf improvements in casting

4 participants

@VSadov@jkotas@jakobbotsch@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[NativeAOT] Make casting logic closer to CoreCLR - #89548

Closed
VSadov wants to merge 12 commits into
dotnet:mainfrom
VSadov:casts
Closed

[NativeAOT] Make casting logic closer to CoreCLR#89548
VSadov wants to merge 12 commits into
dotnet:mainfrom
VSadov:casts

Conversation

@VSadov

@VSadovVSadov commented Jul 27, 2023

Copy link
Copy Markdown
Member

Fixes: #84464

The actual changes are not as big here as might seem. For the most part this is a refactoring of existing code to have a shape closer to CoreCLR cast helpers, so that similar patterns could be used - in a few places where that has not been done already in earlier changes.
For example cases like CheckCastAny - could start with a cache lookup, since uncached code path can be complex and thus relatively slow. (also addresses some old TODOs in this area)

@ghostghost assigned VSadovJul 27, 2023
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

Issue Details

Fixes: #84464

In progress.

For the most part this is a refactoring of existing code to have a shape closer to CoreCLR. In particular to do cache lookups earlier.
Cases like CheckCastAny - we should basically start with a cache lookup, since uncached path can be relatively slow.

Author:VSadov
Assignees:-
Labels:

area-NativeAOT-coreclr

Milestone:-

@VSadov

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@VSadov

Copy link
Copy Markdown
MemberAuthor

I think this is ready for a review.

Comment threadsrc/coreclr/nativeaot/Runtime.Base/src/System/Runtime/TypeCast.cs Outdated
@VSadov
VSadov requested a review from jkotasAugust 1, 2023 18:33
@jkotas

Copy link
Copy Markdown
Member

Could you please collect numbers for casting microbenchmarks before/after this change?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can delete CORINFO_HELP_CHKCASTARRAY and CORINFO_HELP_ISINSTANCEOFARRAY from JIT/EE interface. They are unnecessary now.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

One possible question here would be - Why not just port/share the CoreCLR casting helpers in their entirety and implement the internal calls in managed code - to have nearly the same implementation?

I think one of the factors for the design of CoreCLR managed casting helpers was to avoid complicated API with the native type system. As a result that API is basically 2 internal calls methods - IsInstanceOfAny_NoCacheLookup and ChkCastAny_NoCacheLookup.

In NativeAOT the type system APIs are easily accessible, so we do not need to minimize the use of those APIs. On the other hand there are some differences, like the way we fetch the base type for arrays, that may stand in the way of code sharing.
Thus I did not consider it is as a goal to share the code or making it maximally similar when possible, but I think it can be done.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

Also in NativeAOT these casting helpers have more users with additional needs, while in CoreClr it is really just type system facade for object casting.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

Could you please collect numbers for casting microbenchmarks before/after this change?

I was thinking of what we could measure here.
The most common cases of casting like regular object/interfaces casts did not change, so unlikely to see any differences. We may see differences in more complex cases - like casting to variant interfaces or arrays.

Perhaps just running the regular casting benchmarks that perf lab uses would be informative enough about what changed perf-wise.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

I have run the same benchmark as in #84430 (comment)

internalclassProgram{constintiters=1000000;staticvoidMain(string[]args){for(;;){Time(TestLStringToIROCstring);}}staticvoidTime(Actiona){varsw=Stopwatch.StartNew();for(inti=0;i<100;i++){a();}sw.Stop();System.Console.WriteLine(sw.ElapsedMilliseconds);}staticobjecto=newList<string>();staticvoidTestLStringToIROCstring(){for(inti=0;i<iters;i++){if(oasIReadOnlyCollection<object>==null)thrownull;if(oasIReadOnlyCollection<string>==null)thrownull;if(oasIEnumerable<object>==null)thrownull;if(oasIEnumerable<string>==null)thrownull;}}}

=== before the change

1451
1460
1456
1458
1470
1463

=== after the change:

799
801
798
799
800
799

The reason for the difference is that original code makes a number of calls. Profiler shows:
TypeCast__IsInstanceOf
MethodTable_get_IsArray
TypeCast__IsInstanceOfVariantType - this one does the cache lookup

making calls and additional checks adds up.

In the new implementation there is only one helper call in the profile:
TypeCast__IsInstanceOfAny - this one does the cache lookup

@VSadov

Copy link
Copy Markdown
MemberAuthor

For comparison the CoreCLR is a bit faster.

The same benchmark as above produces (smaller is better):

558
561
562
557
561

As I see in the debugger the native code that we run after this change is nearly the same between CoreCLR and NativeAOT.
There are minor differences like loading of type pointers/handles.
In JIT code they look like:

movrcx,7FFCA19E79E8h

In NativeAOT loading a type looks like:

learcx,[rip+0x73d65]

I do not see any other significant differences. Maybe it is just these little diffs and some indirect impact on code size or alignment that makes the difference.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

Another thing to notice is that compared to #84430 (comment) and prior to this change it looks like the benchmark in the comment has regressed.

That was possibly caused by #86029 . I suspect it made common/simple cases faster, which is good, but regressed complex cases that rely on caching as more checks like MethodTable_get_IsArray could run before eventually hitting the cache. Just a guess though.

Anyways, it looks like after this PR the complex/cached case is faster than in #84430 (comment)

Comment threadsrc/coreclr/inc/corinfo.h Outdated
Comment threadsrc/coreclr/jit/importer.cpp Outdated
@VSadov

Copy link
Copy Markdown
MemberAuthor

There are some superpmi failures. I can't tell if that simply tells that there are codegen diffs or something is crashing.

@VSadov
VSadovforce-pushed the casts branch 2 times, most recently from c3672cb to ac36d4eCompareAugust 6, 2023 21:51
@jkotas

Copy link
Copy Markdown
Member

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

Comment threadsrc/libraries/System.Reflection/tests/GetTypeTests.cs Outdated
Comment threadsrc/libraries/System.Reflection/tests/GetTypeTests.cs Outdated
@VSadov

Copy link
Copy Markdown
MemberAuthor

I'd be ok with merging this now. I think all concerns have been resolved. Let me know if there is something that may be missing.

Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
_ASSERTE(helper == CORINFO_HELP_ISINSTANCEOFANY);
}
else
if (!clsHnd.IsTypeDesc() && !Nullable::IsNullableType(clsHnd))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The native AOT implementation does not checks for Nullable. Is this check redundant here or is the check for Nullable missing in native AOT?

@VSadovVSadovAug 8, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The difference is that JIT tries to transform simple cases like o is int? into o is int, but can't do that when needs a type lookup. As I understand that is due to limitations of IR.

// ECMA-335 III.4.3: If typeTok is a nullable type, Nullable<T>, it is interpreted as "boxed" T
// We can convert constant-ish tokens of nullable to its underlying type.
// However, when the type is shared generic parameter like Nullable<Struct<__Canon>>, the actual type will require
// runtime lookup. It's too complex to add another level of indirection in op2, fallback to the cast helper instead.
if (isClassExact && !(info.compCompHnd->getClassAttribs(pResolvedToken->hClass) & CORINFO_FLG_SHAREDINST))

NativeAOT does not seem to have a problem with expressing such lookup, so we always cast with nullable stripped

caseReadyToRunHelperId.TypeHandleForCasting:
{
vartype=(TypeDesc)targetOfLookup;
if(type.IsNullable)
targetOfLookup=type.Instantiation[0];
returnNecessaryTypeSymbolIfPossible((TypeDesc)targetOfLookup);

@VSadovVSadovAug 9, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In other words in NativeAOT the type check will always be to an unboxed underlying type, thus CLASS variant is correct, and also optimizable, since T? is isClassExact.

In CoreCLR the type check in complex cases could be to the actual Nullable<SomeStruct<string>>, thus it should use ANY variant. It also means that we may be introducing an optimization bug, since such IsInst cannot be lowered to a type/handle compare.

I will check if that is a case, or if there are some mitigating reasons why it still works correctly.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sadly that is the case. The following works incorrectly

=====Prints:TrueFalseusingSystem.Runtime.CompilerServices;namespaceConsoleApp34{structS1<T>{publicTvalue;}internalclassProgram{staticobjecto;staticvoidMain(string[]args){Test<int>();Test<string>();}[MethodImpl(MethodImplOptions.AggressiveOptimization)]privatestaticvoidTest<T>(){o=newS1<T>();Console.WriteLine(oisS1<T>?);}}}

This PR indirectly enabled casting optimizations for cases like o is int?, but we need to suppress it, since in more general cases it does not work correctly.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Unless there are better ideas, I am thinking of constraining isinst optimization for CORINFO_HELP_ISINSTANCEOFANY only if converting to an array type. So we do not keep finding more broken cases.

That may be too conservative, but we should probably stay closer to the preexisting behavior for now and consider if more cases can work in 9.0

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am thinking that this PR as a whole is too risky for 8.0. Are there parts that we think are critical to get into .NET 8?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think the actual casting helper change was low risk as that is mostly refactoring to get more cases to hit the cache earlier. Touching the JIT appears to be a lot more fragile.

There is nothing really "critical" for 8.0, as in - we are not fixing some complete showstoppers here.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have created a "reduced" version of this. I think that is what we can consider for 8.0 - #90234

I have removed all the JIT changes, but kept the added codegen test.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good.

{
DWORD flags = info.compCompHnd->getClassAttribs(classHnd);
DWORD flagsMask = CORINFO_FLG_FINAL | CORINFO_FLG_VARIANCE | CORINFO_FLG_ARRAY;
DWORD flagsMask = CORINFO_FLG_FINAL | CORINFO_FLG_VARIANCE | CORINFO_FLG_TYPE_EQUIVALENCE | CORINFO_FLG_ARRAY;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both CORINFO_FLG_VARIANCE and CORINFO_FLG_TYPE_EQUIVALENCE are only computed to make the impIsClassExact work. Computing these flags is a waste in all other cases. It is the kind of pattern that calls for introduction of dedicated JIT/EE interface API that replaces the flags.

@VSadov

Copy link
Copy Markdown
MemberAuthor

A reduced version of this affecting only run time behavior has been merged.
For the further improvements for the JIT API in this area a tracking bug has been added - #91016

@VSadovVSadov closed this Aug 23, 2023
@jkotasjkotas mentioned this pull request Aug 23, 2023
@ghostghost locked as resolved and limited conversation to collaborators Sep 23, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[NativeAOT] Possible perf improvements in casting

4 participants

@VSadov@jkotas@jakobbotsch@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[NativeAOT] Make casting logic closer to CoreCLR - #89548

Closed
VSadov wants to merge 12 commits into
dotnet:mainfrom
VSadov:casts
Closed

[NativeAOT] Make casting logic closer to CoreCLR#89548
VSadov wants to merge 12 commits into
dotnet:mainfrom
VSadov:casts

Conversation

@VSadov

@VSadovVSadov commented Jul 27, 2023

Copy link
Copy Markdown
Member

Fixes: #84464

The actual changes are not as big here as might seem. For the most part this is a refactoring of existing code to have a shape closer to CoreCLR cast helpers, so that similar patterns could be used - in a few places where that has not been done already in earlier changes.
For example cases like CheckCastAny - could start with a cache lookup, since uncached code path can be complex and thus relatively slow. (also addresses some old TODOs in this area)

@ghostghost assigned VSadovJul 27, 2023
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

Issue Details

Fixes: #84464

In progress.

For the most part this is a refactoring of existing code to have a shape closer to CoreCLR. In particular to do cache lookups earlier.
Cases like CheckCastAny - we should basically start with a cache lookup, since uncached path can be relatively slow.

Author:VSadov
Assignees:-
Labels:

area-NativeAOT-coreclr

Milestone:-

@VSadov

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@VSadov

Copy link
Copy Markdown
MemberAuthor

I think this is ready for a review.

Comment threadsrc/coreclr/nativeaot/Runtime.Base/src/System/Runtime/TypeCast.cs Outdated
@VSadov
VSadov requested a review from jkotasAugust 1, 2023 18:33
@jkotas

Copy link
Copy Markdown
Member

Could you please collect numbers for casting microbenchmarks before/after this change?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can delete CORINFO_HELP_CHKCASTARRAY and CORINFO_HELP_ISINSTANCEOFARRAY from JIT/EE interface. They are unnecessary now.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

One possible question here would be - Why not just port/share the CoreCLR casting helpers in their entirety and implement the internal calls in managed code - to have nearly the same implementation?

I think one of the factors for the design of CoreCLR managed casting helpers was to avoid complicated API with the native type system. As a result that API is basically 2 internal calls methods - IsInstanceOfAny_NoCacheLookup and ChkCastAny_NoCacheLookup.

In NativeAOT the type system APIs are easily accessible, so we do not need to minimize the use of those APIs. On the other hand there are some differences, like the way we fetch the base type for arrays, that may stand in the way of code sharing.
Thus I did not consider it is as a goal to share the code or making it maximally similar when possible, but I think it can be done.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

Also in NativeAOT these casting helpers have more users with additional needs, while in CoreClr it is really just type system facade for object casting.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

Could you please collect numbers for casting microbenchmarks before/after this change?

I was thinking of what we could measure here.
The most common cases of casting like regular object/interfaces casts did not change, so unlikely to see any differences. We may see differences in more complex cases - like casting to variant interfaces or arrays.

Perhaps just running the regular casting benchmarks that perf lab uses would be informative enough about what changed perf-wise.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

I have run the same benchmark as in #84430 (comment)

internalclassProgram{constintiters=1000000;staticvoidMain(string[]args){for(;;){Time(TestLStringToIROCstring);}}staticvoidTime(Actiona){varsw=Stopwatch.StartNew();for(inti=0;i<100;i++){a();}sw.Stop();System.Console.WriteLine(sw.ElapsedMilliseconds);}staticobjecto=newList<string>();staticvoidTestLStringToIROCstring(){for(inti=0;i<iters;i++){if(oasIReadOnlyCollection<object>==null)thrownull;if(oasIReadOnlyCollection<string>==null)thrownull;if(oasIEnumerable<object>==null)thrownull;if(oasIEnumerable<string>==null)thrownull;}}}

=== before the change

1451
1460
1456
1458
1470
1463

=== after the change:

799
801
798
799
800
799

The reason for the difference is that original code makes a number of calls. Profiler shows:
TypeCast__IsInstanceOf
MethodTable_get_IsArray
TypeCast__IsInstanceOfVariantType - this one does the cache lookup

making calls and additional checks adds up.

In the new implementation there is only one helper call in the profile:
TypeCast__IsInstanceOfAny - this one does the cache lookup

@VSadov

Copy link
Copy Markdown
MemberAuthor

For comparison the CoreCLR is a bit faster.

The same benchmark as above produces (smaller is better):

558
561
562
557
561

As I see in the debugger the native code that we run after this change is nearly the same between CoreCLR and NativeAOT.
There are minor differences like loading of type pointers/handles.
In JIT code they look like:

movrcx,7FFCA19E79E8h

In NativeAOT loading a type looks like:

learcx,[rip+0x73d65]

I do not see any other significant differences. Maybe it is just these little diffs and some indirect impact on code size or alignment that makes the difference.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

Another thing to notice is that compared to #84430 (comment) and prior to this change it looks like the benchmark in the comment has regressed.

That was possibly caused by #86029 . I suspect it made common/simple cases faster, which is good, but regressed complex cases that rely on caching as more checks like MethodTable_get_IsArray could run before eventually hitting the cache. Just a guess though.

Anyways, it looks like after this PR the complex/cached case is faster than in #84430 (comment)

Comment threadsrc/coreclr/inc/corinfo.h Outdated
Comment threadsrc/coreclr/jit/importer.cpp Outdated
@VSadov

Copy link
Copy Markdown
MemberAuthor

There are some superpmi failures. I can't tell if that simply tells that there are codegen diffs or something is crashing.

@VSadov
VSadovforce-pushed the casts branch 2 times, most recently from c3672cb to ac36d4eCompareAugust 6, 2023 21:51
@jkotas

Copy link
Copy Markdown
Member

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

Comment threadsrc/libraries/System.Reflection/tests/GetTypeTests.cs Outdated
Comment threadsrc/libraries/System.Reflection/tests/GetTypeTests.cs Outdated
@VSadov

Copy link
Copy Markdown
MemberAuthor

I'd be ok with merging this now. I think all concerns have been resolved. Let me know if there is something that may be missing.

Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
_ASSERTE(helper == CORINFO_HELP_ISINSTANCEOFANY);
}
else
if (!clsHnd.IsTypeDesc() && !Nullable::IsNullableType(clsHnd))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The native AOT implementation does not checks for Nullable. Is this check redundant here or is the check for Nullable missing in native AOT?

@VSadovVSadovAug 8, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The difference is that JIT tries to transform simple cases like o is int? into o is int, but can't do that when needs a type lookup. As I understand that is due to limitations of IR.

// ECMA-335 III.4.3: If typeTok is a nullable type, Nullable<T>, it is interpreted as "boxed" T
// We can convert constant-ish tokens of nullable to its underlying type.
// However, when the type is shared generic parameter like Nullable<Struct<__Canon>>, the actual type will require
// runtime lookup. It's too complex to add another level of indirection in op2, fallback to the cast helper instead.
if (isClassExact && !(info.compCompHnd->getClassAttribs(pResolvedToken->hClass) & CORINFO_FLG_SHAREDINST))

NativeAOT does not seem to have a problem with expressing such lookup, so we always cast with nullable stripped

caseReadyToRunHelperId.TypeHandleForCasting:
{
vartype=(TypeDesc)targetOfLookup;
if(type.IsNullable)
targetOfLookup=type.Instantiation[0];
returnNecessaryTypeSymbolIfPossible((TypeDesc)targetOfLookup);

@VSadovVSadovAug 9, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In other words in NativeAOT the type check will always be to an unboxed underlying type, thus CLASS variant is correct, and also optimizable, since T? is isClassExact.

In CoreCLR the type check in complex cases could be to the actual Nullable<SomeStruct<string>>, thus it should use ANY variant. It also means that we may be introducing an optimization bug, since such IsInst cannot be lowered to a type/handle compare.

I will check if that is a case, or if there are some mitigating reasons why it still works correctly.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sadly that is the case. The following works incorrectly

=====Prints:TrueFalseusingSystem.Runtime.CompilerServices;namespaceConsoleApp34{structS1<T>{publicTvalue;}internalclassProgram{staticobjecto;staticvoidMain(string[]args){Test<int>();Test<string>();}[MethodImpl(MethodImplOptions.AggressiveOptimization)]privatestaticvoidTest<T>(){o=newS1<T>();Console.WriteLine(oisS1<T>?);}}}

This PR indirectly enabled casting optimizations for cases like o is int?, but we need to suppress it, since in more general cases it does not work correctly.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Unless there are better ideas, I am thinking of constraining isinst optimization for CORINFO_HELP_ISINSTANCEOFANY only if converting to an array type. So we do not keep finding more broken cases.

That may be too conservative, but we should probably stay closer to the preexisting behavior for now and consider if more cases can work in 9.0

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am thinking that this PR as a whole is too risky for 8.0. Are there parts that we think are critical to get into .NET 8?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think the actual casting helper change was low risk as that is mostly refactoring to get more cases to hit the cache earlier. Touching the JIT appears to be a lot more fragile.

There is nothing really "critical" for 8.0, as in - we are not fixing some complete showstoppers here.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have created a "reduced" version of this. I think that is what we can consider for 8.0 - #90234

I have removed all the JIT changes, but kept the added codegen test.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good.

{
DWORD flags = info.compCompHnd->getClassAttribs(classHnd);
DWORD flagsMask = CORINFO_FLG_FINAL | CORINFO_FLG_VARIANCE | CORINFO_FLG_ARRAY;
DWORD flagsMask = CORINFO_FLG_FINAL | CORINFO_FLG_VARIANCE | CORINFO_FLG_TYPE_EQUIVALENCE | CORINFO_FLG_ARRAY;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both CORINFO_FLG_VARIANCE and CORINFO_FLG_TYPE_EQUIVALENCE are only computed to make the impIsClassExact work. Computing these flags is a waste in all other cases. It is the kind of pattern that calls for introduction of dedicated JIT/EE interface API that replaces the flags.

@VSadov

Copy link
Copy Markdown
MemberAuthor

A reduced version of this affecting only run time behavior has been merged.
For the further improvements for the JIT API in this area a tracking bug has been added - #91016

@VSadovVSadov closed this Aug 23, 2023
@jkotasjkotas mentioned this pull request Aug 23, 2023
@ghostghost locked as resolved and limited conversation to collaborators Sep 23, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[NativeAOT] Possible perf improvements in casting

4 participants

@VSadov@jkotas@jakobbotsch@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

[NativeAOT] Make casting logic closer to CoreCLR - #89548

Closed
VSadov wants to merge 12 commits into
dotnet:mainfrom
VSadov:casts
Closed

[NativeAOT] Make casting logic closer to CoreCLR#89548
VSadov wants to merge 12 commits into
dotnet:mainfrom
VSadov:casts

Conversation

@VSadov

@VSadovVSadov commented Jul 27, 2023

Copy link
Copy Markdown
Member

Fixes: #84464

The actual changes are not as big here as might seem. For the most part this is a refactoring of existing code to have a shape closer to CoreCLR cast helpers, so that similar patterns could be used - in a few places where that has not been done already in earlier changes.
For example cases like CheckCastAny - could start with a cache lookup, since uncached code path can be complex and thus relatively slow. (also addresses some old TODOs in this area)

@ghostghost assigned VSadovJul 27, 2023
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

Issue Details

Fixes: #84464

In progress.

For the most part this is a refactoring of existing code to have a shape closer to CoreCLR. In particular to do cache lookups earlier.
Cases like CheckCastAny - we should basically start with a cache lookup, since uncached path can be relatively slow.

Author:VSadov
Assignees:-
Labels:

area-NativeAOT-coreclr

Milestone:-

@VSadov

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@VSadov

Copy link
Copy Markdown
MemberAuthor

I think this is ready for a review.

Comment threadsrc/coreclr/nativeaot/Runtime.Base/src/System/Runtime/TypeCast.cs Outdated
@VSadov
VSadov requested a review from jkotasAugust 1, 2023 18:33
@jkotas

Copy link
Copy Markdown
Member

Could you please collect numbers for casting microbenchmarks before/after this change?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can delete CORINFO_HELP_CHKCASTARRAY and CORINFO_HELP_ISINSTANCEOFARRAY from JIT/EE interface. They are unnecessary now.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

One possible question here would be - Why not just port/share the CoreCLR casting helpers in their entirety and implement the internal calls in managed code - to have nearly the same implementation?

I think one of the factors for the design of CoreCLR managed casting helpers was to avoid complicated API with the native type system. As a result that API is basically 2 internal calls methods - IsInstanceOfAny_NoCacheLookup and ChkCastAny_NoCacheLookup.

In NativeAOT the type system APIs are easily accessible, so we do not need to minimize the use of those APIs. On the other hand there are some differences, like the way we fetch the base type for arrays, that may stand in the way of code sharing.
Thus I did not consider it is as a goal to share the code or making it maximally similar when possible, but I think it can be done.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

Also in NativeAOT these casting helpers have more users with additional needs, while in CoreClr it is really just type system facade for object casting.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

Could you please collect numbers for casting microbenchmarks before/after this change?

I was thinking of what we could measure here.
The most common cases of casting like regular object/interfaces casts did not change, so unlikely to see any differences. We may see differences in more complex cases - like casting to variant interfaces or arrays.

Perhaps just running the regular casting benchmarks that perf lab uses would be informative enough about what changed perf-wise.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

I have run the same benchmark as in #84430 (comment)

internalclassProgram{constintiters=1000000;staticvoidMain(string[]args){for(;;){Time(TestLStringToIROCstring);}}staticvoidTime(Actiona){varsw=Stopwatch.StartNew();for(inti=0;i<100;i++){a();}sw.Stop();System.Console.WriteLine(sw.ElapsedMilliseconds);}staticobjecto=newList<string>();staticvoidTestLStringToIROCstring(){for(inti=0;i<iters;i++){if(oasIReadOnlyCollection<object>==null)thrownull;if(oasIReadOnlyCollection<string>==null)thrownull;if(oasIEnumerable<object>==null)thrownull;if(oasIEnumerable<string>==null)thrownull;}}}

=== before the change

1451
1460
1456
1458
1470
1463

=== after the change:

799
801
798
799
800
799

The reason for the difference is that original code makes a number of calls. Profiler shows:
TypeCast__IsInstanceOf
MethodTable_get_IsArray
TypeCast__IsInstanceOfVariantType - this one does the cache lookup

making calls and additional checks adds up.

In the new implementation there is only one helper call in the profile:
TypeCast__IsInstanceOfAny - this one does the cache lookup

@VSadov

Copy link
Copy Markdown
MemberAuthor

For comparison the CoreCLR is a bit faster.

The same benchmark as above produces (smaller is better):

558
561
562
557
561

As I see in the debugger the native code that we run after this change is nearly the same between CoreCLR and NativeAOT.
There are minor differences like loading of type pointers/handles.
In JIT code they look like:

movrcx,7FFCA19E79E8h

In NativeAOT loading a type looks like:

learcx,[rip+0x73d65]

I do not see any other significant differences. Maybe it is just these little diffs and some indirect impact on code size or alignment that makes the difference.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

Another thing to notice is that compared to #84430 (comment) and prior to this change it looks like the benchmark in the comment has regressed.

That was possibly caused by #86029 . I suspect it made common/simple cases faster, which is good, but regressed complex cases that rely on caching as more checks like MethodTable_get_IsArray could run before eventually hitting the cache. Just a guess though.

Anyways, it looks like after this PR the complex/cached case is faster than in #84430 (comment)

Comment threadsrc/coreclr/inc/corinfo.h Outdated
Comment threadsrc/coreclr/jit/importer.cpp Outdated
@VSadov

Copy link
Copy Markdown
MemberAuthor

There are some superpmi failures. I can't tell if that simply tells that there are codegen diffs or something is crashing.

@VSadov
VSadovforce-pushed the casts branch 2 times, most recently from c3672cb to ac36d4eCompareAugust 6, 2023 21:51
@jkotas

Copy link
Copy Markdown
Member

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

Comment threadsrc/libraries/System.Reflection/tests/GetTypeTests.cs Outdated
Comment threadsrc/libraries/System.Reflection/tests/GetTypeTests.cs Outdated
@VSadov

Copy link
Copy Markdown
MemberAuthor

I'd be ok with merging this now. I think all concerns have been resolved. Let me know if there is something that may be missing.

Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
_ASSERTE(helper == CORINFO_HELP_ISINSTANCEOFANY);
}
else
if (!clsHnd.IsTypeDesc() && !Nullable::IsNullableType(clsHnd))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The native AOT implementation does not checks for Nullable. Is this check redundant here or is the check for Nullable missing in native AOT?

@VSadovVSadovAug 8, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The difference is that JIT tries to transform simple cases like o is int? into o is int, but can't do that when needs a type lookup. As I understand that is due to limitations of IR.

// ECMA-335 III.4.3: If typeTok is a nullable type, Nullable<T>, it is interpreted as "boxed" T
// We can convert constant-ish tokens of nullable to its underlying type.
// However, when the type is shared generic parameter like Nullable<Struct<__Canon>>, the actual type will require
// runtime lookup. It's too complex to add another level of indirection in op2, fallback to the cast helper instead.
if (isClassExact && !(info.compCompHnd->getClassAttribs(pResolvedToken->hClass) & CORINFO_FLG_SHAREDINST))

NativeAOT does not seem to have a problem with expressing such lookup, so we always cast with nullable stripped

caseReadyToRunHelperId.TypeHandleForCasting:
{
vartype=(TypeDesc)targetOfLookup;
if(type.IsNullable)
targetOfLookup=type.Instantiation[0];
returnNecessaryTypeSymbolIfPossible((TypeDesc)targetOfLookup);

@VSadovVSadovAug 9, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In other words in NativeAOT the type check will always be to an unboxed underlying type, thus CLASS variant is correct, and also optimizable, since T? is isClassExact.

In CoreCLR the type check in complex cases could be to the actual Nullable<SomeStruct<string>>, thus it should use ANY variant. It also means that we may be introducing an optimization bug, since such IsInst cannot be lowered to a type/handle compare.

I will check if that is a case, or if there are some mitigating reasons why it still works correctly.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sadly that is the case. The following works incorrectly

=====Prints:TrueFalseusingSystem.Runtime.CompilerServices;namespaceConsoleApp34{structS1<T>{publicTvalue;}internalclassProgram{staticobjecto;staticvoidMain(string[]args){Test<int>();Test<string>();}[MethodImpl(MethodImplOptions.AggressiveOptimization)]privatestaticvoidTest<T>(){o=newS1<T>();Console.WriteLine(oisS1<T>?);}}}

This PR indirectly enabled casting optimizations for cases like o is int?, but we need to suppress it, since in more general cases it does not work correctly.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Unless there are better ideas, I am thinking of constraining isinst optimization for CORINFO_HELP_ISINSTANCEOFANY only if converting to an array type. So we do not keep finding more broken cases.

That may be too conservative, but we should probably stay closer to the preexisting behavior for now and consider if more cases can work in 9.0

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am thinking that this PR as a whole is too risky for 8.0. Are there parts that we think are critical to get into .NET 8?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think the actual casting helper change was low risk as that is mostly refactoring to get more cases to hit the cache earlier. Touching the JIT appears to be a lot more fragile.

There is nothing really "critical" for 8.0, as in - we are not fixing some complete showstoppers here.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have created a "reduced" version of this. I think that is what we can consider for 8.0 - #90234

I have removed all the JIT changes, but kept the added codegen test.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good.

{
DWORD flags = info.compCompHnd->getClassAttribs(classHnd);
DWORD flagsMask = CORINFO_FLG_FINAL | CORINFO_FLG_VARIANCE | CORINFO_FLG_ARRAY;
DWORD flagsMask = CORINFO_FLG_FINAL | CORINFO_FLG_VARIANCE | CORINFO_FLG_TYPE_EQUIVALENCE | CORINFO_FLG_ARRAY;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both CORINFO_FLG_VARIANCE and CORINFO_FLG_TYPE_EQUIVALENCE are only computed to make the impIsClassExact work. Computing these flags is a waste in all other cases. It is the kind of pattern that calls for introduction of dedicated JIT/EE interface API that replaces the flags.

@VSadov

Copy link
Copy Markdown
MemberAuthor

A reduced version of this affecting only run time behavior has been merged.
For the further improvements for the JIT API in this area a tracking bug has been added - #91016

@VSadovVSadov closed this Aug 23, 2023
@jkotasjkotas mentioned this pull request Aug 23, 2023
@ghostghost locked as resolved and limited conversation to collaborators Sep 23, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[NativeAOT] Possible perf improvements in casting

4 participants

@VSadov@jkotas@jakobbotsch@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[NativeAOT] Make casting logic closer to CoreCLR - #89548

Closed
VSadov wants to merge 12 commits into
dotnet:mainfrom
VSadov:casts
Closed

[NativeAOT] Make casting logic closer to CoreCLR#89548
VSadov wants to merge 12 commits into
dotnet:mainfrom
VSadov:casts

Conversation

@VSadov

@VSadovVSadov commented Jul 27, 2023

Copy link
Copy Markdown
Member

Fixes: #84464

The actual changes are not as big here as might seem. For the most part this is a refactoring of existing code to have a shape closer to CoreCLR cast helpers, so that similar patterns could be used - in a few places where that has not been done already in earlier changes.
For example cases like CheckCastAny - could start with a cache lookup, since uncached code path can be complex and thus relatively slow. (also addresses some old TODOs in this area)

@ghostghost assigned VSadovJul 27, 2023
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

Issue Details

Fixes: #84464

In progress.

For the most part this is a refactoring of existing code to have a shape closer to CoreCLR. In particular to do cache lookups earlier.
Cases like CheckCastAny - we should basically start with a cache lookup, since uncached path can be relatively slow.

Author:VSadov
Assignees:-
Labels:

area-NativeAOT-coreclr

Milestone:-

@VSadov

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@VSadov

Copy link
Copy Markdown
MemberAuthor

I think this is ready for a review.

Comment threadsrc/coreclr/nativeaot/Runtime.Base/src/System/Runtime/TypeCast.cs Outdated
@VSadov
VSadov requested a review from jkotasAugust 1, 2023 18:33
@jkotas

Copy link
Copy Markdown
Member

Could you please collect numbers for casting microbenchmarks before/after this change?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can delete CORINFO_HELP_CHKCASTARRAY and CORINFO_HELP_ISINSTANCEOFARRAY from JIT/EE interface. They are unnecessary now.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

One possible question here would be - Why not just port/share the CoreCLR casting helpers in their entirety and implement the internal calls in managed code - to have nearly the same implementation?

I think one of the factors for the design of CoreCLR managed casting helpers was to avoid complicated API with the native type system. As a result that API is basically 2 internal calls methods - IsInstanceOfAny_NoCacheLookup and ChkCastAny_NoCacheLookup.

In NativeAOT the type system APIs are easily accessible, so we do not need to minimize the use of those APIs. On the other hand there are some differences, like the way we fetch the base type for arrays, that may stand in the way of code sharing.
Thus I did not consider it is as a goal to share the code or making it maximally similar when possible, but I think it can be done.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

Also in NativeAOT these casting helpers have more users with additional needs, while in CoreClr it is really just type system facade for object casting.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

Could you please collect numbers for casting microbenchmarks before/after this change?

I was thinking of what we could measure here.
The most common cases of casting like regular object/interfaces casts did not change, so unlikely to see any differences. We may see differences in more complex cases - like casting to variant interfaces or arrays.

Perhaps just running the regular casting benchmarks that perf lab uses would be informative enough about what changed perf-wise.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

I have run the same benchmark as in #84430 (comment)

internalclassProgram{constintiters=1000000;staticvoidMain(string[]args){for(;;){Time(TestLStringToIROCstring);}}staticvoidTime(Actiona){varsw=Stopwatch.StartNew();for(inti=0;i<100;i++){a();}sw.Stop();System.Console.WriteLine(sw.ElapsedMilliseconds);}staticobjecto=newList<string>();staticvoidTestLStringToIROCstring(){for(inti=0;i<iters;i++){if(oasIReadOnlyCollection<object>==null)thrownull;if(oasIReadOnlyCollection<string>==null)thrownull;if(oasIEnumerable<object>==null)thrownull;if(oasIEnumerable<string>==null)thrownull;}}}

=== before the change

1451
1460
1456
1458
1470
1463

=== after the change:

799
801
798
799
800
799

The reason for the difference is that original code makes a number of calls. Profiler shows:
TypeCast__IsInstanceOf
MethodTable_get_IsArray
TypeCast__IsInstanceOfVariantType - this one does the cache lookup

making calls and additional checks adds up.

In the new implementation there is only one helper call in the profile:
TypeCast__IsInstanceOfAny - this one does the cache lookup

@VSadov

Copy link
Copy Markdown
MemberAuthor

For comparison the CoreCLR is a bit faster.

The same benchmark as above produces (smaller is better):

558
561
562
557
561

As I see in the debugger the native code that we run after this change is nearly the same between CoreCLR and NativeAOT.
There are minor differences like loading of type pointers/handles.
In JIT code they look like:

movrcx,7FFCA19E79E8h

In NativeAOT loading a type looks like:

learcx,[rip+0x73d65]

I do not see any other significant differences. Maybe it is just these little diffs and some indirect impact on code size or alignment that makes the difference.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

Another thing to notice is that compared to #84430 (comment) and prior to this change it looks like the benchmark in the comment has regressed.

That was possibly caused by #86029 . I suspect it made common/simple cases faster, which is good, but regressed complex cases that rely on caching as more checks like MethodTable_get_IsArray could run before eventually hitting the cache. Just a guess though.

Anyways, it looks like after this PR the complex/cached case is faster than in #84430 (comment)

Comment threadsrc/coreclr/inc/corinfo.h Outdated
Comment threadsrc/coreclr/jit/importer.cpp Outdated
@VSadov

Copy link
Copy Markdown
MemberAuthor

There are some superpmi failures. I can't tell if that simply tells that there are codegen diffs or something is crashing.

@VSadov
VSadovforce-pushed the casts branch 2 times, most recently from c3672cb to ac36d4eCompareAugust 6, 2023 21:51
@jkotas

Copy link
Copy Markdown
Member

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

Comment threadsrc/libraries/System.Reflection/tests/GetTypeTests.cs Outdated
Comment threadsrc/libraries/System.Reflection/tests/GetTypeTests.cs Outdated
@VSadov

Copy link
Copy Markdown
MemberAuthor

I'd be ok with merging this now. I think all concerns have been resolved. Let me know if there is something that may be missing.

Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
_ASSERTE(helper == CORINFO_HELP_ISINSTANCEOFANY);
}
else
if (!clsHnd.IsTypeDesc() && !Nullable::IsNullableType(clsHnd))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The native AOT implementation does not checks for Nullable. Is this check redundant here or is the check for Nullable missing in native AOT?

@VSadovVSadovAug 8, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The difference is that JIT tries to transform simple cases like o is int? into o is int, but can't do that when needs a type lookup. As I understand that is due to limitations of IR.

// ECMA-335 III.4.3: If typeTok is a nullable type, Nullable<T>, it is interpreted as "boxed" T
// We can convert constant-ish tokens of nullable to its underlying type.
// However, when the type is shared generic parameter like Nullable<Struct<__Canon>>, the actual type will require
// runtime lookup. It's too complex to add another level of indirection in op2, fallback to the cast helper instead.
if (isClassExact && !(info.compCompHnd->getClassAttribs(pResolvedToken->hClass) & CORINFO_FLG_SHAREDINST))

NativeAOT does not seem to have a problem with expressing such lookup, so we always cast with nullable stripped

caseReadyToRunHelperId.TypeHandleForCasting:
{
vartype=(TypeDesc)targetOfLookup;
if(type.IsNullable)
targetOfLookup=type.Instantiation[0];
returnNecessaryTypeSymbolIfPossible((TypeDesc)targetOfLookup);

@VSadovVSadovAug 9, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In other words in NativeAOT the type check will always be to an unboxed underlying type, thus CLASS variant is correct, and also optimizable, since T? is isClassExact.

In CoreCLR the type check in complex cases could be to the actual Nullable<SomeStruct<string>>, thus it should use ANY variant. It also means that we may be introducing an optimization bug, since such IsInst cannot be lowered to a type/handle compare.

I will check if that is a case, or if there are some mitigating reasons why it still works correctly.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sadly that is the case. The following works incorrectly

=====Prints:TrueFalseusingSystem.Runtime.CompilerServices;namespaceConsoleApp34{structS1<T>{publicTvalue;}internalclassProgram{staticobjecto;staticvoidMain(string[]args){Test<int>();Test<string>();}[MethodImpl(MethodImplOptions.AggressiveOptimization)]privatestaticvoidTest<T>(){o=newS1<T>();Console.WriteLine(oisS1<T>?);}}}

This PR indirectly enabled casting optimizations for cases like o is int?, but we need to suppress it, since in more general cases it does not work correctly.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Unless there are better ideas, I am thinking of constraining isinst optimization for CORINFO_HELP_ISINSTANCEOFANY only if converting to an array type. So we do not keep finding more broken cases.

That may be too conservative, but we should probably stay closer to the preexisting behavior for now and consider if more cases can work in 9.0

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am thinking that this PR as a whole is too risky for 8.0. Are there parts that we think are critical to get into .NET 8?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think the actual casting helper change was low risk as that is mostly refactoring to get more cases to hit the cache earlier. Touching the JIT appears to be a lot more fragile.

There is nothing really "critical" for 8.0, as in - we are not fixing some complete showstoppers here.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have created a "reduced" version of this. I think that is what we can consider for 8.0 - #90234

I have removed all the JIT changes, but kept the added codegen test.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good.

{
DWORD flags = info.compCompHnd->getClassAttribs(classHnd);
DWORD flagsMask = CORINFO_FLG_FINAL | CORINFO_FLG_VARIANCE | CORINFO_FLG_ARRAY;
DWORD flagsMask = CORINFO_FLG_FINAL | CORINFO_FLG_VARIANCE | CORINFO_FLG_TYPE_EQUIVALENCE | CORINFO_FLG_ARRAY;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both CORINFO_FLG_VARIANCE and CORINFO_FLG_TYPE_EQUIVALENCE are only computed to make the impIsClassExact work. Computing these flags is a waste in all other cases. It is the kind of pattern that calls for introduction of dedicated JIT/EE interface API that replaces the flags.

@VSadov

Copy link
Copy Markdown
MemberAuthor

A reduced version of this affecting only run time behavior has been merged.
For the further improvements for the JIT API in this area a tracking bug has been added - #91016

@VSadovVSadov closed this Aug 23, 2023
@jkotasjkotas mentioned this pull request Aug 23, 2023
@ghostghost locked as resolved and limited conversation to collaborators Sep 23, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[NativeAOT] Possible perf improvements in casting

4 participants

@VSadov@jkotas@jakobbotsch@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[NativeAOT] Make casting logic closer to CoreCLR - #89548

Closed
VSadov wants to merge 12 commits into
dotnet:mainfrom
VSadov:casts
Closed

[NativeAOT] Make casting logic closer to CoreCLR#89548
VSadov wants to merge 12 commits into
dotnet:mainfrom
VSadov:casts

Conversation

@VSadov

@VSadovVSadov commented Jul 27, 2023

Copy link
Copy Markdown
Member

Fixes: #84464

The actual changes are not as big here as might seem. For the most part this is a refactoring of existing code to have a shape closer to CoreCLR cast helpers, so that similar patterns could be used - in a few places where that has not been done already in earlier changes.
For example cases like CheckCastAny - could start with a cache lookup, since uncached code path can be complex and thus relatively slow. (also addresses some old TODOs in this area)

@ghostghost assigned VSadovJul 27, 2023
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

Issue Details

Fixes: #84464

In progress.

For the most part this is a refactoring of existing code to have a shape closer to CoreCLR. In particular to do cache lookups earlier.
Cases like CheckCastAny - we should basically start with a cache lookup, since uncached path can be relatively slow.

Author:VSadov
Assignees:-
Labels:

area-NativeAOT-coreclr

Milestone:-

@VSadov

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@VSadov

Copy link
Copy Markdown
MemberAuthor

I think this is ready for a review.

Comment threadsrc/coreclr/nativeaot/Runtime.Base/src/System/Runtime/TypeCast.cs Outdated
@VSadov
VSadov requested a review from jkotasAugust 1, 2023 18:33
@jkotas

Copy link
Copy Markdown
Member

Could you please collect numbers for casting microbenchmarks before/after this change?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can delete CORINFO_HELP_CHKCASTARRAY and CORINFO_HELP_ISINSTANCEOFARRAY from JIT/EE interface. They are unnecessary now.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

One possible question here would be - Why not just port/share the CoreCLR casting helpers in their entirety and implement the internal calls in managed code - to have nearly the same implementation?

I think one of the factors for the design of CoreCLR managed casting helpers was to avoid complicated API with the native type system. As a result that API is basically 2 internal calls methods - IsInstanceOfAny_NoCacheLookup and ChkCastAny_NoCacheLookup.

In NativeAOT the type system APIs are easily accessible, so we do not need to minimize the use of those APIs. On the other hand there are some differences, like the way we fetch the base type for arrays, that may stand in the way of code sharing.
Thus I did not consider it is as a goal to share the code or making it maximally similar when possible, but I think it can be done.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

Also in NativeAOT these casting helpers have more users with additional needs, while in CoreClr it is really just type system facade for object casting.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

Could you please collect numbers for casting microbenchmarks before/after this change?

I was thinking of what we could measure here.
The most common cases of casting like regular object/interfaces casts did not change, so unlikely to see any differences. We may see differences in more complex cases - like casting to variant interfaces or arrays.

Perhaps just running the regular casting benchmarks that perf lab uses would be informative enough about what changed perf-wise.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

I have run the same benchmark as in #84430 (comment)

internalclassProgram{constintiters=1000000;staticvoidMain(string[]args){for(;;){Time(TestLStringToIROCstring);}}staticvoidTime(Actiona){varsw=Stopwatch.StartNew();for(inti=0;i<100;i++){a();}sw.Stop();System.Console.WriteLine(sw.ElapsedMilliseconds);}staticobjecto=newList<string>();staticvoidTestLStringToIROCstring(){for(inti=0;i<iters;i++){if(oasIReadOnlyCollection<object>==null)thrownull;if(oasIReadOnlyCollection<string>==null)thrownull;if(oasIEnumerable<object>==null)thrownull;if(oasIEnumerable<string>==null)thrownull;}}}

=== before the change

1451
1460
1456
1458
1470
1463

=== after the change:

799
801
798
799
800
799

The reason for the difference is that original code makes a number of calls. Profiler shows:
TypeCast__IsInstanceOf
MethodTable_get_IsArray
TypeCast__IsInstanceOfVariantType - this one does the cache lookup

making calls and additional checks adds up.

In the new implementation there is only one helper call in the profile:
TypeCast__IsInstanceOfAny - this one does the cache lookup

@VSadov

Copy link
Copy Markdown
MemberAuthor

For comparison the CoreCLR is a bit faster.

The same benchmark as above produces (smaller is better):

558
561
562
557
561

As I see in the debugger the native code that we run after this change is nearly the same between CoreCLR and NativeAOT.
There are minor differences like loading of type pointers/handles.
In JIT code they look like:

movrcx,7FFCA19E79E8h

In NativeAOT loading a type looks like:

learcx,[rip+0x73d65]

I do not see any other significant differences. Maybe it is just these little diffs and some indirect impact on code size or alignment that makes the difference.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

Another thing to notice is that compared to #84430 (comment) and prior to this change it looks like the benchmark in the comment has regressed.

That was possibly caused by #86029 . I suspect it made common/simple cases faster, which is good, but regressed complex cases that rely on caching as more checks like MethodTable_get_IsArray could run before eventually hitting the cache. Just a guess though.

Anyways, it looks like after this PR the complex/cached case is faster than in #84430 (comment)

Comment threadsrc/coreclr/inc/corinfo.h Outdated
Comment threadsrc/coreclr/jit/importer.cpp Outdated
@VSadov

Copy link
Copy Markdown
MemberAuthor

There are some superpmi failures. I can't tell if that simply tells that there are codegen diffs or something is crashing.

@VSadov
VSadovforce-pushed the casts branch 2 times, most recently from c3672cb to ac36d4eCompareAugust 6, 2023 21:51
@jkotas

Copy link
Copy Markdown
Member

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

Comment threadsrc/libraries/System.Reflection/tests/GetTypeTests.cs Outdated
Comment threadsrc/libraries/System.Reflection/tests/GetTypeTests.cs Outdated
@VSadov

Copy link
Copy Markdown
MemberAuthor

I'd be ok with merging this now. I think all concerns have been resolved. Let me know if there is something that may be missing.

Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
_ASSERTE(helper == CORINFO_HELP_ISINSTANCEOFANY);
}
else
if (!clsHnd.IsTypeDesc() && !Nullable::IsNullableType(clsHnd))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The native AOT implementation does not checks for Nullable. Is this check redundant here or is the check for Nullable missing in native AOT?

@VSadovVSadovAug 8, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The difference is that JIT tries to transform simple cases like o is int? into o is int, but can't do that when needs a type lookup. As I understand that is due to limitations of IR.

// ECMA-335 III.4.3: If typeTok is a nullable type, Nullable<T>, it is interpreted as "boxed" T
// We can convert constant-ish tokens of nullable to its underlying type.
// However, when the type is shared generic parameter like Nullable<Struct<__Canon>>, the actual type will require
// runtime lookup. It's too complex to add another level of indirection in op2, fallback to the cast helper instead.
if (isClassExact && !(info.compCompHnd->getClassAttribs(pResolvedToken->hClass) & CORINFO_FLG_SHAREDINST))

NativeAOT does not seem to have a problem with expressing such lookup, so we always cast with nullable stripped

caseReadyToRunHelperId.TypeHandleForCasting:
{
vartype=(TypeDesc)targetOfLookup;
if(type.IsNullable)
targetOfLookup=type.Instantiation[0];
returnNecessaryTypeSymbolIfPossible((TypeDesc)targetOfLookup);

@VSadovVSadovAug 9, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In other words in NativeAOT the type check will always be to an unboxed underlying type, thus CLASS variant is correct, and also optimizable, since T? is isClassExact.

In CoreCLR the type check in complex cases could be to the actual Nullable<SomeStruct<string>>, thus it should use ANY variant. It also means that we may be introducing an optimization bug, since such IsInst cannot be lowered to a type/handle compare.

I will check if that is a case, or if there are some mitigating reasons why it still works correctly.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sadly that is the case. The following works incorrectly

=====Prints:TrueFalseusingSystem.Runtime.CompilerServices;namespaceConsoleApp34{structS1<T>{publicTvalue;}internalclassProgram{staticobjecto;staticvoidMain(string[]args){Test<int>();Test<string>();}[MethodImpl(MethodImplOptions.AggressiveOptimization)]privatestaticvoidTest<T>(){o=newS1<T>();Console.WriteLine(oisS1<T>?);}}}

This PR indirectly enabled casting optimizations for cases like o is int?, but we need to suppress it, since in more general cases it does not work correctly.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Unless there are better ideas, I am thinking of constraining isinst optimization for CORINFO_HELP_ISINSTANCEOFANY only if converting to an array type. So we do not keep finding more broken cases.

That may be too conservative, but we should probably stay closer to the preexisting behavior for now and consider if more cases can work in 9.0

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am thinking that this PR as a whole is too risky for 8.0. Are there parts that we think are critical to get into .NET 8?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think the actual casting helper change was low risk as that is mostly refactoring to get more cases to hit the cache earlier. Touching the JIT appears to be a lot more fragile.

There is nothing really "critical" for 8.0, as in - we are not fixing some complete showstoppers here.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have created a "reduced" version of this. I think that is what we can consider for 8.0 - #90234

I have removed all the JIT changes, but kept the added codegen test.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good.

{
DWORD flags = info.compCompHnd->getClassAttribs(classHnd);
DWORD flagsMask = CORINFO_FLG_FINAL | CORINFO_FLG_VARIANCE | CORINFO_FLG_ARRAY;
DWORD flagsMask = CORINFO_FLG_FINAL | CORINFO_FLG_VARIANCE | CORINFO_FLG_TYPE_EQUIVALENCE | CORINFO_FLG_ARRAY;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both CORINFO_FLG_VARIANCE and CORINFO_FLG_TYPE_EQUIVALENCE are only computed to make the impIsClassExact work. Computing these flags is a waste in all other cases. It is the kind of pattern that calls for introduction of dedicated JIT/EE interface API that replaces the flags.

@VSadov

Copy link
Copy Markdown
MemberAuthor

A reduced version of this affecting only run time behavior has been merged.
For the further improvements for the JIT API in this area a tracking bug has been added - #91016

@VSadovVSadov closed this Aug 23, 2023
@jkotasjkotas mentioned this pull request Aug 23, 2023
@ghostghost locked as resolved and limited conversation to collaborators Sep 23, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[NativeAOT] Possible perf improvements in casting

4 participants

@VSadov@jkotas@jakobbotsch@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

[NativeAOT] Make casting logic closer to CoreCLR - #89548

Closed
VSadov wants to merge 12 commits into
dotnet:mainfrom
VSadov:casts
Closed

[NativeAOT] Make casting logic closer to CoreCLR#89548
VSadov wants to merge 12 commits into
dotnet:mainfrom
VSadov:casts

Conversation

@VSadov

@VSadovVSadov commented Jul 27, 2023

Copy link
Copy Markdown
Member

Fixes: #84464

The actual changes are not as big here as might seem. For the most part this is a refactoring of existing code to have a shape closer to CoreCLR cast helpers, so that similar patterns could be used - in a few places where that has not been done already in earlier changes.
For example cases like CheckCastAny - could start with a cache lookup, since uncached code path can be complex and thus relatively slow. (also addresses some old TODOs in this area)

@ghostghost assigned VSadovJul 27, 2023
@ghost

Copy link
Copy Markdown

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

Issue Details

Fixes: #84464

In progress.

For the most part this is a refactoring of existing code to have a shape closer to CoreCLR. In particular to do cache lookups earlier.
Cases like CheckCastAny - we should basically start with a cache lookup, since uncached path can be relatively slow.

Author:VSadov
Assignees:-
Labels:

area-NativeAOT-coreclr

Milestone:-

@VSadov

Copy link
Copy Markdown
MemberAuthor

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@VSadov

Copy link
Copy Markdown
MemberAuthor

I think this is ready for a review.

Comment threadsrc/coreclr/nativeaot/Runtime.Base/src/System/Runtime/TypeCast.cs Outdated
@VSadov
VSadov requested a review from jkotasAugust 1, 2023 18:33
@jkotas

Copy link
Copy Markdown
Member

Could you please collect numbers for casting microbenchmarks before/after this change?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can delete CORINFO_HELP_CHKCASTARRAY and CORINFO_HELP_ISINSTANCEOFARRAY from JIT/EE interface. They are unnecessary now.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

One possible question here would be - Why not just port/share the CoreCLR casting helpers in their entirety and implement the internal calls in managed code - to have nearly the same implementation?

I think one of the factors for the design of CoreCLR managed casting helpers was to avoid complicated API with the native type system. As a result that API is basically 2 internal calls methods - IsInstanceOfAny_NoCacheLookup and ChkCastAny_NoCacheLookup.

In NativeAOT the type system APIs are easily accessible, so we do not need to minimize the use of those APIs. On the other hand there are some differences, like the way we fetch the base type for arrays, that may stand in the way of code sharing.
Thus I did not consider it is as a goal to share the code or making it maximally similar when possible, but I think it can be done.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

Also in NativeAOT these casting helpers have more users with additional needs, while in CoreClr it is really just type system facade for object casting.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

Could you please collect numbers for casting microbenchmarks before/after this change?

I was thinking of what we could measure here.
The most common cases of casting like regular object/interfaces casts did not change, so unlikely to see any differences. We may see differences in more complex cases - like casting to variant interfaces or arrays.

Perhaps just running the regular casting benchmarks that perf lab uses would be informative enough about what changed perf-wise.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

I have run the same benchmark as in #84430 (comment)

internalclassProgram{constintiters=1000000;staticvoidMain(string[]args){for(;;){Time(TestLStringToIROCstring);}}staticvoidTime(Actiona){varsw=Stopwatch.StartNew();for(inti=0;i<100;i++){a();}sw.Stop();System.Console.WriteLine(sw.ElapsedMilliseconds);}staticobjecto=newList<string>();staticvoidTestLStringToIROCstring(){for(inti=0;i<iters;i++){if(oasIReadOnlyCollection<object>==null)thrownull;if(oasIReadOnlyCollection<string>==null)thrownull;if(oasIEnumerable<object>==null)thrownull;if(oasIEnumerable<string>==null)thrownull;}}}

=== before the change

1451
1460
1456
1458
1470
1463

=== after the change:

799
801
798
799
800
799

The reason for the difference is that original code makes a number of calls. Profiler shows:
TypeCast__IsInstanceOf
MethodTable_get_IsArray
TypeCast__IsInstanceOfVariantType - this one does the cache lookup

making calls and additional checks adds up.

In the new implementation there is only one helper call in the profile:
TypeCast__IsInstanceOfAny - this one does the cache lookup

@VSadov

Copy link
Copy Markdown
MemberAuthor

For comparison the CoreCLR is a bit faster.

The same benchmark as above produces (smaller is better):

558
561
562
557
561

As I see in the debugger the native code that we run after this change is nearly the same between CoreCLR and NativeAOT.
There are minor differences like loading of type pointers/handles.
In JIT code they look like:

movrcx,7FFCA19E79E8h

In NativeAOT loading a type looks like:

learcx,[rip+0x73d65]

I do not see any other significant differences. Maybe it is just these little diffs and some indirect impact on code size or alignment that makes the difference.

@VSadov

VSadov commented Aug 1, 2023

Copy link
Copy Markdown
MemberAuthor

Another thing to notice is that compared to #84430 (comment) and prior to this change it looks like the benchmark in the comment has regressed.

That was possibly caused by #86029 . I suspect it made common/simple cases faster, which is good, but regressed complex cases that rely on caching as more checks like MethodTable_get_IsArray could run before eventually hitting the cache. Just a guess though.

Anyways, it looks like after this PR the complex/cached case is faster than in #84430 (comment)

Comment threadsrc/coreclr/inc/corinfo.h Outdated
Comment threadsrc/coreclr/jit/importer.cpp Outdated
@VSadov

Copy link
Copy Markdown
MemberAuthor

There are some superpmi failures. I can't tell if that simply tells that there are codegen diffs or something is crashing.

@VSadov
VSadovforce-pushed the casts branch 2 times, most recently from c3672cb to ac36d4eCompareAugust 6, 2023 21:51
@jkotas

Copy link
Copy Markdown
Member

/azp run runtime-extra-platforms

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

Comment threadsrc/libraries/System.Reflection/tests/GetTypeTests.cs Outdated
Comment threadsrc/libraries/System.Reflection/tests/GetTypeTests.cs Outdated
@VSadov

Copy link
Copy Markdown
MemberAuthor

I'd be ok with merging this now. I think all concerns have been resolved. Let me know if there is something that may be missing.

Comment threadsrc/coreclr/vm/jitinterface.cpp Outdated
_ASSERTE(helper == CORINFO_HELP_ISINSTANCEOFANY);
}
else
if (!clsHnd.IsTypeDesc() && !Nullable::IsNullableType(clsHnd))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The native AOT implementation does not checks for Nullable. Is this check redundant here or is the check for Nullable missing in native AOT?

@VSadovVSadovAug 8, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The difference is that JIT tries to transform simple cases like o is int? into o is int, but can't do that when needs a type lookup. As I understand that is due to limitations of IR.

// ECMA-335 III.4.3: If typeTok is a nullable type, Nullable<T>, it is interpreted as "boxed" T
// We can convert constant-ish tokens of nullable to its underlying type.
// However, when the type is shared generic parameter like Nullable<Struct<__Canon>>, the actual type will require
// runtime lookup. It's too complex to add another level of indirection in op2, fallback to the cast helper instead.
if (isClassExact && !(info.compCompHnd->getClassAttribs(pResolvedToken->hClass) & CORINFO_FLG_SHAREDINST))

NativeAOT does not seem to have a problem with expressing such lookup, so we always cast with nullable stripped

caseReadyToRunHelperId.TypeHandleForCasting:
{
vartype=(TypeDesc)targetOfLookup;
if(type.IsNullable)
targetOfLookup=type.Instantiation[0];
returnNecessaryTypeSymbolIfPossible((TypeDesc)targetOfLookup);

@VSadovVSadovAug 9, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In other words in NativeAOT the type check will always be to an unboxed underlying type, thus CLASS variant is correct, and also optimizable, since T? is isClassExact.

In CoreCLR the type check in complex cases could be to the actual Nullable<SomeStruct<string>>, thus it should use ANY variant. It also means that we may be introducing an optimization bug, since such IsInst cannot be lowered to a type/handle compare.

I will check if that is a case, or if there are some mitigating reasons why it still works correctly.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sadly that is the case. The following works incorrectly

=====Prints:TrueFalseusingSystem.Runtime.CompilerServices;namespaceConsoleApp34{structS1<T>{publicTvalue;}internalclassProgram{staticobjecto;staticvoidMain(string[]args){Test<int>();Test<string>();}[MethodImpl(MethodImplOptions.AggressiveOptimization)]privatestaticvoidTest<T>(){o=newS1<T>();Console.WriteLine(oisS1<T>?);}}}

This PR indirectly enabled casting optimizations for cases like o is int?, but we need to suppress it, since in more general cases it does not work correctly.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Unless there are better ideas, I am thinking of constraining isinst optimization for CORINFO_HELP_ISINSTANCEOFANY only if converting to an array type. So we do not keep finding more broken cases.

That may be too conservative, but we should probably stay closer to the preexisting behavior for now and consider if more cases can work in 9.0

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am thinking that this PR as a whole is too risky for 8.0. Are there parts that we think are critical to get into .NET 8?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think the actual casting helper change was low risk as that is mostly refactoring to get more cases to hit the cache earlier. Touching the JIT appears to be a lot more fragile.

There is nothing really "critical" for 8.0, as in - we are not fixing some complete showstoppers here.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have created a "reduced" version of this. I think that is what we can consider for 8.0 - #90234

I have removed all the JIT changes, but kept the added codegen test.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good.

{
DWORD flags = info.compCompHnd->getClassAttribs(classHnd);
DWORD flagsMask = CORINFO_FLG_FINAL | CORINFO_FLG_VARIANCE | CORINFO_FLG_ARRAY;
DWORD flagsMask = CORINFO_FLG_FINAL | CORINFO_FLG_VARIANCE | CORINFO_FLG_TYPE_EQUIVALENCE | CORINFO_FLG_ARRAY;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both CORINFO_FLG_VARIANCE and CORINFO_FLG_TYPE_EQUIVALENCE are only computed to make the impIsClassExact work. Computing these flags is a waste in all other cases. It is the kind of pattern that calls for introduction of dedicated JIT/EE interface API that replaces the flags.

@VSadov

Copy link
Copy Markdown
MemberAuthor

A reduced version of this affecting only run time behavior has been merged.
For the further improvements for the JIT API in this area a tracking bug has been added - #91016

@VSadovVSadov closed this Aug 23, 2023
@jkotasjkotas mentioned this pull request Aug 23, 2023
@ghostghost locked as resolved and limited conversation to collaborators Sep 23, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[NativeAOT] Possible perf improvements in casting

4 participants

@VSadov@jkotas@jakobbotsch@EgorBo