Uh oh!
There was an error while loading. Please reload this page.
[NativeAOT] Make casting logic closer to CoreCLR - #89548
Conversation
ghost
commented
Jul 27, 2023
Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas Issue DetailsFixes: #84464 In progress. For the most part this is a refactoring of existing code to have a shape closer to CoreCLR. In particular to do cache lookups earlier.
|
VSadov
commented
Jul 27, 2023
/azp run runtime-extra-platforms |
|
Azure Pipelines successfully started running 1 pipeline(s). |
VSadov
commented
Aug 1, 2023
I think this is ready for a review. |
Uh oh!
There was an error while loading. Please reload this page.
jkotas
commented
Aug 1, 2023
Could you please collect numbers for casting microbenchmarks before/after this change? |
There was a problem hiding this comment.
You can delete CORINFO_HELP_CHKCASTARRAY and CORINFO_HELP_ISINSTANCEOFARRAY from JIT/EE interface. They are unnecessary now.
One possible question here would be - Why not just port/share the CoreCLR casting helpers in their entirety and implement the internal calls in managed code - to have nearly the same implementation? I think one of the factors for the design of CoreCLR managed casting helpers was to avoid complicated API with the native type system. As a result that API is basically 2 internal calls methods - In NativeAOT the type system APIs are easily accessible, so we do not need to minimize the use of those APIs. On the other hand there are some differences, like the way we fetch the base type for arrays, that may stand in the way of code sharing. |
Also in NativeAOT these casting helpers have more users with additional needs, while in CoreClr it is really just type system facade for object casting. |
I was thinking of what we could measure here. Perhaps just running the regular casting benchmarks that perf lab uses would be informative enough about what changed perf-wise. |
I have run the same benchmark as in #84430 (comment) internalclassProgram{constintiters=1000000;staticvoidMain(string[]args){for(;;){Time(TestLStringToIROCstring);}}staticvoidTime(Actiona){varsw=Stopwatch.StartNew();for(inti=0;i<100;i++){a();}sw.Stop();System.Console.WriteLine(sw.ElapsedMilliseconds);}staticobjecto=newList<string>();staticvoidTestLStringToIROCstring(){for(inti=0;i<iters;i++){if(oasIReadOnlyCollection<object>==null)thrownull;if(oasIReadOnlyCollection<string>==null)thrownull;if(oasIEnumerable<object>==null)thrownull;if(oasIEnumerable<string>==null)thrownull;}}}=== before the change === after the change: The reason for the difference is that original code makes a number of calls. Profiler shows: making calls and additional checks adds up. In the new implementation there is only one helper call in the profile: |
VSadov
commented
Aug 1, 2023
For comparison the CoreCLR is a bit faster. The same benchmark as above produces (smaller is better): As I see in the debugger the native code that we run after this change is nearly the same between CoreCLR and NativeAOT. movrcx,7FFCA19E79E8hIn NativeAOT loading a type looks like: learcx,[rip+0x73d65]I do not see any other significant differences. Maybe it is just these little diffs and some indirect impact on code size or alignment that makes the difference. |
Another thing to notice is that compared to #84430 (comment) and prior to this change it looks like the benchmark in the comment has regressed. That was possibly caused by #86029 . I suspect it made common/simple cases faster, which is good, but regressed complex cases that rely on caching as more checks like Anyways, it looks like after this PR the complex/cached case is faster than in #84430 (comment) |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
VSadov
commented
Aug 2, 2023
There are some |
c3672cb to
ac36d4eComparejkotas
commented
Aug 6, 2023
/azp run runtime-extra-platforms |
|
Azure Pipelines successfully started running 1 pipeline(s). |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
VSadov
commented
Aug 8, 2023
I'd be ok with merging this now. I think all concerns have been resolved. Let me know if there is something that may be missing. |
Uh oh!
There was an error while loading. Please reload this page.
| _ASSERTE(helper == CORINFO_HELP_ISINSTANCEOFANY); | ||
| } | ||
| else | ||
| if (!clsHnd.IsTypeDesc() && !Nullable::IsNullableType(clsHnd)) |
There was a problem hiding this comment.
The native AOT implementation does not checks for Nullable. Is this check redundant here or is the check for Nullable missing in native AOT?
There was a problem hiding this comment.
The difference is that JIT tries to transform simple cases like o is int? into o is int, but can't do that when needs a type lookup. As I understand that is due to limitations of IR.
runtime/src/coreclr/jit/importer.cpp
Lines 5387 to 5391 in 40b39ff
NativeAOT does not seem to have a problem with expressing such lookup, so we always cast with nullable stripped
runtime/src/coreclr/tools/aot/ILCompiler.Compiler/Compiler/Compilation.cs
Lines 301 to 306 in 17d13fa
There was a problem hiding this comment.
In other words in NativeAOT the type check will always be to an unboxed underlying type, thus CLASS variant is correct, and also optimizable, since T? is isClassExact.
In CoreCLR the type check in complex cases could be to the actual Nullable<SomeStruct<string>>, thus it should use ANY variant. It also means that we may be introducing an optimization bug, since such IsInst cannot be lowered to a type/handle compare.
I will check if that is a case, or if there are some mitigating reasons why it still works correctly.
There was a problem hiding this comment.
Sadly that is the case. The following works incorrectly
=====Prints:TrueFalseusingSystem.Runtime.CompilerServices;namespaceConsoleApp34{structS1<T>{publicTvalue;}internalclassProgram{staticobjecto;staticvoidMain(string[]args){Test<int>();Test<string>();}[MethodImpl(MethodImplOptions.AggressiveOptimization)]privatestaticvoidTest<T>(){o=newS1<T>();Console.WriteLine(oisS1<T>?);}}}This PR indirectly enabled casting optimizations for cases like o is int?, but we need to suppress it, since in more general cases it does not work correctly.
There was a problem hiding this comment.
Unless there are better ideas, I am thinking of constraining isinst optimization for CORINFO_HELP_ISINSTANCEOFANY only if converting to an array type. So we do not keep finding more broken cases.
That may be too conservative, but we should probably stay closer to the preexisting behavior for now and consider if more cases can work in 9.0
There was a problem hiding this comment.
I am thinking that this PR as a whole is too risky for 8.0. Are there parts that we think are critical to get into .NET 8?
There was a problem hiding this comment.
I think the actual casting helper change was low risk as that is mostly refactoring to get more cases to hit the cache earlier. Touching the JIT appears to be a lot more fragile.
There is nothing really "critical" for 8.0, as in - we are not fixing some complete showstoppers here.
There was a problem hiding this comment.
I have created a "reduced" version of this. I think that is what we can consider for 8.0 - #90234
I have removed all the JIT changes, but kept the added codegen test.
| { | ||
| DWORD flags = info.compCompHnd->getClassAttribs(classHnd); | ||
| DWORD flagsMask = CORINFO_FLG_FINAL | CORINFO_FLG_VARIANCE | CORINFO_FLG_ARRAY; | ||
| DWORD flagsMask = CORINFO_FLG_FINAL | CORINFO_FLG_VARIANCE | CORINFO_FLG_TYPE_EQUIVALENCE | CORINFO_FLG_ARRAY; |
There was a problem hiding this comment.
Both CORINFO_FLG_VARIANCE and CORINFO_FLG_TYPE_EQUIVALENCE are only computed to make the impIsClassExact work. Computing these flags is a waste in all other cases. It is the kind of pattern that calls for introduction of dedicated JIT/EE interface API that replaces the flags.
VSadov
commented
Aug 23, 2023
A reduced version of this affecting only run time behavior has been merged. |
Fixes: #84464
The actual changes are not as big here as might seem. For the most part this is a refactoring of existing code to have a shape closer to CoreCLR cast helpers, so that similar patterns could be used - in a few places where that has not been done already in earlier changes.
For example cases like
CheckCastAny- could start with a cache lookup, since uncached code path can be complex and thus relatively slow. (also addresses some old TODOs in this area)