[NativeAOT] Using the same CastCache implementation as in CoreClr - #84430

Merged
VSadov merged 11 commits into
dotnet:mainfrom
VSadov:castCacheN
Apr 7, 2023
Merged

[NativeAOT] Using the same CastCache implementation as in CoreClr#84430
VSadov merged 11 commits into
dotnet:mainfrom
VSadov:castCacheN

Conversation

@VSadov

Copy link
Copy Markdown
Member

Fixes: #75111

Reasons:

  • code sharing
  • CoreClr implementation does not require a Crst , or any kind of locks.

@ghost

ghost commented Apr 6, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

Issue Details

Fixes: #75111

Reasons:

  • code sharing
  • CoreClr implementation does not require a Crst , or any kind of locks.
Author:VSadov
Assignees:-
Labels:

area-NativeAOT-coreclr

Milestone:-


#if TARGET_64BIT
[MethodImpl(MethodImplOptions.AggressiveInlining)]
private static ulong RotateLeft(ulong value, int offset)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This should be deleted and replaced by BitOperations.RotateLeft.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've introduced this because NativeAOT was not building in the test configuration that was missing a lot of things - like Volatile, Interlocked, BitOperations.
Later I moved custom implementations under Test.CoreLib, but missed this one.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We may want to omit this cache for Test.CoreLib. Test.CoreLib has simplistic implementation of number of other subsystems.

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

At this point it would be easy to move these helpers to Test.CoreLib - as custom BitOperations implementations. Then we can have the cache in Test.CorLib and have extra test coverage.

However, if having a minimal implementation is the point, then I can instead add a Test.CoreLib - specific implementation of the cache that does not do anything.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Then we can have the cache in Test.CorLib and have extra test coverage.

There are no relevant tests running against Test.CorLib. All tests that matter for this run against real CoreLib.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reverted additional API implementations for Test.CorLib and added a trivial no-op cache instead.

// - issue a load barrier before reading _version
// benchmarks on available hardware (Jan 2020) show that use of a read barrier is cheaper.

#if CORECLR

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If this is required for correctness, it should not be under ifdef.

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is either a load fence here or the two loads above need to be acquires. As it was measured, when this was implemented, one fence was noticeably cheaper.

We need to add Interlocked.ReadMemoryBarrier() to NativeAOT. And it needs to be an intrinsic - no point in implementing it as an internal call.

I will log an issue on Interlocked.ReadMemoryBarrier(). Once that is added, we can remove ifdefs here and around two acquires above.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Logged an issue for adding Interlocked.ReadMemoryBarrier() intrinsic - #84445

}

// the rest is the support for updating the cache.
// in CoreClr the cache is only updated in the native code

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would it make sense for CoreCLR to just call this managed implementation instead?

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Interesting idea. I think what could prevent that is:

  • if we need to use the cache before we can run managed code.
    JIT calls into casting machinery, so it is hard to tell if we may need the cache too early.

  • if we need to use the cache before managed cctor has run.
    I'd guess the cctor could be forced to run early enough.

  • if we need to use cache from some mode that does not allow managed code.
    Most likely this is not a requirement, but I am not sure.

  • if calling managed code from native has high overhead.
    Adding to the cache is reasonably fast. We do not see the cost of adding stuff even at start up, when we'd expect cache misses.
    I do not have a good sense of how expensive it is to call managed code from the native runtime in this context.

Would we want to call managed TryGet as well?
(that would be a lot more sensitive to the overhead and TryGet is GC_NOTRIGGER, I vaguely remember there were some reasons for that.)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is worth considering, but it should be a separate change.

It does not look like a straightforward tweak and may need some experimenting.
Also the impact of such change will be mostly in CoreClr, while this PR mostly affects NativeAOT, so in terms of watching for failures or stress/perf consequences, it would be better to have separate changes.

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Logged an issue to follow up on switching CoreCLR to use managed TrySet and maybe TryGet - #84448

@VSadov

VSadov commented Apr 6, 2023

Copy link
Copy Markdown
MemberAuthor

I have run some simple scenarios with this change to see if there are no obvious performance regressions.

It looks like cached casting is slightly faster with the new cache, but not by much. The key point is to have a cache in the first place and NativeAOT already had it.

I also noticed that CoreClr is faster than NativeAOT. I think the reason is the code that runs before we hit the cache. That is - the code that lives in nativeaot\Runtime.Base\src\System\Runtime\TypeCast.cs. It looks like the entry points are a bit more complex than in CoreClr cast helpers.
Cached casts are fast and couple extra calls, indirections or branches may have an impact.
Also it could be that some extra complexity is necessary. Dealing with IsCloned appears to be specific to NativeAOT.
There are also couple TODOs about IsIDynamicInterfaceCastable that seem relevant to the overall perf.

It may be worth looking at the cast entry points. Perhaps there are some opportunities there.

The microbenchmark that I used to see impact of the cache in "easy" case:

internalclassProgram{constintiters=1000000;staticvoidMain(string[]args){for(;;){Time(TestLStringToIROCstring);}}staticvoidTime(Actiona){varsw=Stopwatch.StartNew();for(inti=0;i<100;i++){a();}sw.Stop();System.Console.WriteLine(sw.ElapsedMilliseconds);}staticobjecto=newList<string>();staticvoidTestLStringToIROCstring(){for(inti=0;i<iters;i++){if(oasIReadOnlyCollection<object>==null)thrownull;if(oasIReadOnlyCollection<string>==null)thrownull;if(oasIEnumerable<object>==null)thrownull;if(oasIEnumerable<string>==null)thrownull;}}}

On my machine (x64), I see:
(smaller is better)

==== before change:
871
868
874
870
870
870
==== after the change
858
860
863
864
859
856
==== CoreClr (JIT)
586
589
583
587
583
585

@jkotasjkotas left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

More code sharing. Yay!

@jkotas

Copy link
Copy Markdown
Member

Dealing with IsCloned appears to be specific to NativeAOT.

IsCloned support is not used. I would not feel bad about deleting it. I think it is unlikely that we will ever use it for anything.

CoreCLR has the methodtable flags specifically shaped to make casting fast. We may want to copy/unify the shape.

@VSadov

Copy link
Copy Markdown
MemberAuthor

CoreCLR has the methodtable flags specifically shaped to make casting fast. We may want to copy/unify the shape.

Right. Also the CoreClr entry points were carefully crafted to minimize branches, benefit from tail calling, if possible. Perhaps at some costs to readability, but for that code it is acceptable.

I will look at what we can borrow from there.

@MichalStrehovsky

Copy link
Copy Markdown
Member

Right. Also the CoreClr entry points were carefully crafted to minimize branches, benefit from tail calling, if possible. Perhaps at some costs to readability, but for that code it is acceptable.

We also have some unnecessary checks like "is this null" or "is this exactly the same type" for things that RyuJIT already tests for. The code was written for UTC, not RyuJIT. We just need to validate all the non-codegen callees (i.e. when we call into casting from our regular C# code) also do the appropriate checks and delete the ifs.

@VSadov

Copy link
Copy Markdown
MemberAuthor

Thanks!!

@VSadov
VSadov merged commit 38b81ba into dotnet:mainApr 7, 2023
@VSadov
VSadov deleted the castCacheN branch April 7, 2023 02:04
@ghostghost locked as resolved and limited conversation to collaborators May 7, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[NativeAOT] Consider porting/sharing CastCache form CoreClr

3 participants

@VSadov@jkotas@MichalStrehovsky
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

[NativeAOT] Using the same CastCache implementation as in CoreClr - #84430

Merged
VSadov merged 11 commits into
dotnet:mainfrom
VSadov:castCacheN
Apr 7, 2023
Merged

[NativeAOT] Using the same CastCache implementation as in CoreClr#84430
VSadov merged 11 commits into
dotnet:mainfrom
VSadov:castCacheN

Conversation

@VSadov

Copy link
Copy Markdown
Member

Fixes: #75111

Reasons:

  • code sharing
  • CoreClr implementation does not require a Crst , or any kind of locks.

@ghost

ghost commented Apr 6, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

Issue Details

Fixes: #75111

Reasons:

  • code sharing
  • CoreClr implementation does not require a Crst , or any kind of locks.
Author:VSadov
Assignees:-
Labels:

area-NativeAOT-coreclr

Milestone:-


#if TARGET_64BIT
[MethodImpl(MethodImplOptions.AggressiveInlining)]
private static ulong RotateLeft(ulong value, int offset)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This should be deleted and replaced by BitOperations.RotateLeft.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've introduced this because NativeAOT was not building in the test configuration that was missing a lot of things - like Volatile, Interlocked, BitOperations.
Later I moved custom implementations under Test.CoreLib, but missed this one.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We may want to omit this cache for Test.CoreLib. Test.CoreLib has simplistic implementation of number of other subsystems.

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

At this point it would be easy to move these helpers to Test.CoreLib - as custom BitOperations implementations. Then we can have the cache in Test.CorLib and have extra test coverage.

However, if having a minimal implementation is the point, then I can instead add a Test.CoreLib - specific implementation of the cache that does not do anything.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Then we can have the cache in Test.CorLib and have extra test coverage.

There are no relevant tests running against Test.CorLib. All tests that matter for this run against real CoreLib.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reverted additional API implementations for Test.CorLib and added a trivial no-op cache instead.

// - issue a load barrier before reading _version
// benchmarks on available hardware (Jan 2020) show that use of a read barrier is cheaper.

#if CORECLR

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If this is required for correctness, it should not be under ifdef.

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is either a load fence here or the two loads above need to be acquires. As it was measured, when this was implemented, one fence was noticeably cheaper.

We need to add Interlocked.ReadMemoryBarrier() to NativeAOT. And it needs to be an intrinsic - no point in implementing it as an internal call.

I will log an issue on Interlocked.ReadMemoryBarrier(). Once that is added, we can remove ifdefs here and around two acquires above.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Logged an issue for adding Interlocked.ReadMemoryBarrier() intrinsic - #84445

}

// the rest is the support for updating the cache.
// in CoreClr the cache is only updated in the native code

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would it make sense for CoreCLR to just call this managed implementation instead?

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Interesting idea. I think what could prevent that is:

  • if we need to use the cache before we can run managed code.
    JIT calls into casting machinery, so it is hard to tell if we may need the cache too early.

  • if we need to use the cache before managed cctor has run.
    I'd guess the cctor could be forced to run early enough.

  • if we need to use cache from some mode that does not allow managed code.
    Most likely this is not a requirement, but I am not sure.

  • if calling managed code from native has high overhead.
    Adding to the cache is reasonably fast. We do not see the cost of adding stuff even at start up, when we'd expect cache misses.
    I do not have a good sense of how expensive it is to call managed code from the native runtime in this context.

Would we want to call managed TryGet as well?
(that would be a lot more sensitive to the overhead and TryGet is GC_NOTRIGGER, I vaguely remember there were some reasons for that.)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is worth considering, but it should be a separate change.

It does not look like a straightforward tweak and may need some experimenting.
Also the impact of such change will be mostly in CoreClr, while this PR mostly affects NativeAOT, so in terms of watching for failures or stress/perf consequences, it would be better to have separate changes.

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Logged an issue to follow up on switching CoreCLR to use managed TrySet and maybe TryGet - #84448

@VSadov

VSadov commented Apr 6, 2023

Copy link
Copy Markdown
MemberAuthor

I have run some simple scenarios with this change to see if there are no obvious performance regressions.

It looks like cached casting is slightly faster with the new cache, but not by much. The key point is to have a cache in the first place and NativeAOT already had it.

I also noticed that CoreClr is faster than NativeAOT. I think the reason is the code that runs before we hit the cache. That is - the code that lives in nativeaot\Runtime.Base\src\System\Runtime\TypeCast.cs. It looks like the entry points are a bit more complex than in CoreClr cast helpers.
Cached casts are fast and couple extra calls, indirections or branches may have an impact.
Also it could be that some extra complexity is necessary. Dealing with IsCloned appears to be specific to NativeAOT.
There are also couple TODOs about IsIDynamicInterfaceCastable that seem relevant to the overall perf.

It may be worth looking at the cast entry points. Perhaps there are some opportunities there.

The microbenchmark that I used to see impact of the cache in "easy" case:

internalclassProgram{constintiters=1000000;staticvoidMain(string[]args){for(;;){Time(TestLStringToIROCstring);}}staticvoidTime(Actiona){varsw=Stopwatch.StartNew();for(inti=0;i<100;i++){a();}sw.Stop();System.Console.WriteLine(sw.ElapsedMilliseconds);}staticobjecto=newList<string>();staticvoidTestLStringToIROCstring(){for(inti=0;i<iters;i++){if(oasIReadOnlyCollection<object>==null)thrownull;if(oasIReadOnlyCollection<string>==null)thrownull;if(oasIEnumerable<object>==null)thrownull;if(oasIEnumerable<string>==null)thrownull;}}}

On my machine (x64), I see:
(smaller is better)

==== before change:
871
868
874
870
870
870
==== after the change
858
860
863
864
859
856
==== CoreClr (JIT)
586
589
583
587
583
585

@jkotasjkotas left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

More code sharing. Yay!

@jkotas

Copy link
Copy Markdown
Member

Dealing with IsCloned appears to be specific to NativeAOT.

IsCloned support is not used. I would not feel bad about deleting it. I think it is unlikely that we will ever use it for anything.

CoreCLR has the methodtable flags specifically shaped to make casting fast. We may want to copy/unify the shape.

@VSadov

Copy link
Copy Markdown
MemberAuthor

CoreCLR has the methodtable flags specifically shaped to make casting fast. We may want to copy/unify the shape.

Right. Also the CoreClr entry points were carefully crafted to minimize branches, benefit from tail calling, if possible. Perhaps at some costs to readability, but for that code it is acceptable.

I will look at what we can borrow from there.

@MichalStrehovsky

Copy link
Copy Markdown
Member

Right. Also the CoreClr entry points were carefully crafted to minimize branches, benefit from tail calling, if possible. Perhaps at some costs to readability, but for that code it is acceptable.

We also have some unnecessary checks like "is this null" or "is this exactly the same type" for things that RyuJIT already tests for. The code was written for UTC, not RyuJIT. We just need to validate all the non-codegen callees (i.e. when we call into casting from our regular C# code) also do the appropriate checks and delete the ifs.

@VSadov

Copy link
Copy Markdown
MemberAuthor

Thanks!!

@VSadov
VSadov merged commit 38b81ba into dotnet:mainApr 7, 2023
@VSadov
VSadov deleted the castCacheN branch April 7, 2023 02:04
@ghostghost locked as resolved and limited conversation to collaborators May 7, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[NativeAOT] Consider porting/sharing CastCache form CoreClr

3 participants

@VSadov@jkotas@MichalStrehovsky
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[NativeAOT] Using the same CastCache implementation as in CoreClr - #84430

Merged
VSadov merged 11 commits into
dotnet:mainfrom
VSadov:castCacheN
Apr 7, 2023
Merged

[NativeAOT] Using the same CastCache implementation as in CoreClr#84430
VSadov merged 11 commits into
dotnet:mainfrom
VSadov:castCacheN

Conversation

@VSadov

Copy link
Copy Markdown
Member

Fixes: #75111

Reasons:

  • code sharing
  • CoreClr implementation does not require a Crst , or any kind of locks.

@ghost

ghost commented Apr 6, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

Issue Details

Fixes: #75111

Reasons:

  • code sharing
  • CoreClr implementation does not require a Crst , or any kind of locks.
Author:VSadov
Assignees:-
Labels:

area-NativeAOT-coreclr

Milestone:-


#if TARGET_64BIT
[MethodImpl(MethodImplOptions.AggressiveInlining)]
private static ulong RotateLeft(ulong value, int offset)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This should be deleted and replaced by BitOperations.RotateLeft.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've introduced this because NativeAOT was not building in the test configuration that was missing a lot of things - like Volatile, Interlocked, BitOperations.
Later I moved custom implementations under Test.CoreLib, but missed this one.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We may want to omit this cache for Test.CoreLib. Test.CoreLib has simplistic implementation of number of other subsystems.

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

At this point it would be easy to move these helpers to Test.CoreLib - as custom BitOperations implementations. Then we can have the cache in Test.CorLib and have extra test coverage.

However, if having a minimal implementation is the point, then I can instead add a Test.CoreLib - specific implementation of the cache that does not do anything.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Then we can have the cache in Test.CorLib and have extra test coverage.

There are no relevant tests running against Test.CorLib. All tests that matter for this run against real CoreLib.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reverted additional API implementations for Test.CorLib and added a trivial no-op cache instead.

// - issue a load barrier before reading _version
// benchmarks on available hardware (Jan 2020) show that use of a read barrier is cheaper.

#if CORECLR

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If this is required for correctness, it should not be under ifdef.

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is either a load fence here or the two loads above need to be acquires. As it was measured, when this was implemented, one fence was noticeably cheaper.

We need to add Interlocked.ReadMemoryBarrier() to NativeAOT. And it needs to be an intrinsic - no point in implementing it as an internal call.

I will log an issue on Interlocked.ReadMemoryBarrier(). Once that is added, we can remove ifdefs here and around two acquires above.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Logged an issue for adding Interlocked.ReadMemoryBarrier() intrinsic - #84445

}

// the rest is the support for updating the cache.
// in CoreClr the cache is only updated in the native code

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would it make sense for CoreCLR to just call this managed implementation instead?

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Interesting idea. I think what could prevent that is:

  • if we need to use the cache before we can run managed code.
    JIT calls into casting machinery, so it is hard to tell if we may need the cache too early.

  • if we need to use the cache before managed cctor has run.
    I'd guess the cctor could be forced to run early enough.

  • if we need to use cache from some mode that does not allow managed code.
    Most likely this is not a requirement, but I am not sure.

  • if calling managed code from native has high overhead.
    Adding to the cache is reasonably fast. We do not see the cost of adding stuff even at start up, when we'd expect cache misses.
    I do not have a good sense of how expensive it is to call managed code from the native runtime in this context.

Would we want to call managed TryGet as well?
(that would be a lot more sensitive to the overhead and TryGet is GC_NOTRIGGER, I vaguely remember there were some reasons for that.)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is worth considering, but it should be a separate change.

It does not look like a straightforward tweak and may need some experimenting.
Also the impact of such change will be mostly in CoreClr, while this PR mostly affects NativeAOT, so in terms of watching for failures or stress/perf consequences, it would be better to have separate changes.

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Logged an issue to follow up on switching CoreCLR to use managed TrySet and maybe TryGet - #84448

@VSadov

VSadov commented Apr 6, 2023

Copy link
Copy Markdown
MemberAuthor

I have run some simple scenarios with this change to see if there are no obvious performance regressions.

It looks like cached casting is slightly faster with the new cache, but not by much. The key point is to have a cache in the first place and NativeAOT already had it.

I also noticed that CoreClr is faster than NativeAOT. I think the reason is the code that runs before we hit the cache. That is - the code that lives in nativeaot\Runtime.Base\src\System\Runtime\TypeCast.cs. It looks like the entry points are a bit more complex than in CoreClr cast helpers.
Cached casts are fast and couple extra calls, indirections or branches may have an impact.
Also it could be that some extra complexity is necessary. Dealing with IsCloned appears to be specific to NativeAOT.
There are also couple TODOs about IsIDynamicInterfaceCastable that seem relevant to the overall perf.

It may be worth looking at the cast entry points. Perhaps there are some opportunities there.

The microbenchmark that I used to see impact of the cache in "easy" case:

internalclassProgram{constintiters=1000000;staticvoidMain(string[]args){for(;;){Time(TestLStringToIROCstring);}}staticvoidTime(Actiona){varsw=Stopwatch.StartNew();for(inti=0;i<100;i++){a();}sw.Stop();System.Console.WriteLine(sw.ElapsedMilliseconds);}staticobjecto=newList<string>();staticvoidTestLStringToIROCstring(){for(inti=0;i<iters;i++){if(oasIReadOnlyCollection<object>==null)thrownull;if(oasIReadOnlyCollection<string>==null)thrownull;if(oasIEnumerable<object>==null)thrownull;if(oasIEnumerable<string>==null)thrownull;}}}

On my machine (x64), I see:
(smaller is better)

==== before change:
871
868
874
870
870
870
==== after the change
858
860
863
864
859
856
==== CoreClr (JIT)
586
589
583
587
583
585

@jkotasjkotas left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

More code sharing. Yay!

@jkotas

Copy link
Copy Markdown
Member

Dealing with IsCloned appears to be specific to NativeAOT.

IsCloned support is not used. I would not feel bad about deleting it. I think it is unlikely that we will ever use it for anything.

CoreCLR has the methodtable flags specifically shaped to make casting fast. We may want to copy/unify the shape.

@VSadov

Copy link
Copy Markdown
MemberAuthor

CoreCLR has the methodtable flags specifically shaped to make casting fast. We may want to copy/unify the shape.

Right. Also the CoreClr entry points were carefully crafted to minimize branches, benefit from tail calling, if possible. Perhaps at some costs to readability, but for that code it is acceptable.

I will look at what we can borrow from there.

@MichalStrehovsky

Copy link
Copy Markdown
Member

Right. Also the CoreClr entry points were carefully crafted to minimize branches, benefit from tail calling, if possible. Perhaps at some costs to readability, but for that code it is acceptable.

We also have some unnecessary checks like "is this null" or "is this exactly the same type" for things that RyuJIT already tests for. The code was written for UTC, not RyuJIT. We just need to validate all the non-codegen callees (i.e. when we call into casting from our regular C# code) also do the appropriate checks and delete the ifs.

@VSadov

Copy link
Copy Markdown
MemberAuthor

Thanks!!

@VSadov
VSadov merged commit 38b81ba into dotnet:mainApr 7, 2023
@VSadov
VSadov deleted the castCacheN branch April 7, 2023 02:04
@ghostghost locked as resolved and limited conversation to collaborators May 7, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[NativeAOT] Consider porting/sharing CastCache form CoreClr

3 participants

@VSadov@jkotas@MichalStrehovsky
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[NativeAOT] Using the same CastCache implementation as in CoreClr - #84430

Merged
VSadov merged 11 commits into
dotnet:mainfrom
VSadov:castCacheN
Apr 7, 2023
Merged

[NativeAOT] Using the same CastCache implementation as in CoreClr#84430
VSadov merged 11 commits into
dotnet:mainfrom
VSadov:castCacheN

Conversation

@VSadov

Copy link
Copy Markdown
Member

Fixes: #75111

Reasons:

  • code sharing
  • CoreClr implementation does not require a Crst , or any kind of locks.

@ghost

ghost commented Apr 6, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

Issue Details

Fixes: #75111

Reasons:

  • code sharing
  • CoreClr implementation does not require a Crst , or any kind of locks.
Author:VSadov
Assignees:-
Labels:

area-NativeAOT-coreclr

Milestone:-


#if TARGET_64BIT
[MethodImpl(MethodImplOptions.AggressiveInlining)]
private static ulong RotateLeft(ulong value, int offset)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This should be deleted and replaced by BitOperations.RotateLeft.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've introduced this because NativeAOT was not building in the test configuration that was missing a lot of things - like Volatile, Interlocked, BitOperations.
Later I moved custom implementations under Test.CoreLib, but missed this one.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We may want to omit this cache for Test.CoreLib. Test.CoreLib has simplistic implementation of number of other subsystems.

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

At this point it would be easy to move these helpers to Test.CoreLib - as custom BitOperations implementations. Then we can have the cache in Test.CorLib and have extra test coverage.

However, if having a minimal implementation is the point, then I can instead add a Test.CoreLib - specific implementation of the cache that does not do anything.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Then we can have the cache in Test.CorLib and have extra test coverage.

There are no relevant tests running against Test.CorLib. All tests that matter for this run against real CoreLib.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reverted additional API implementations for Test.CorLib and added a trivial no-op cache instead.

// - issue a load barrier before reading _version
// benchmarks on available hardware (Jan 2020) show that use of a read barrier is cheaper.

#if CORECLR

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If this is required for correctness, it should not be under ifdef.

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is either a load fence here or the two loads above need to be acquires. As it was measured, when this was implemented, one fence was noticeably cheaper.

We need to add Interlocked.ReadMemoryBarrier() to NativeAOT. And it needs to be an intrinsic - no point in implementing it as an internal call.

I will log an issue on Interlocked.ReadMemoryBarrier(). Once that is added, we can remove ifdefs here and around two acquires above.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Logged an issue for adding Interlocked.ReadMemoryBarrier() intrinsic - #84445

}

// the rest is the support for updating the cache.
// in CoreClr the cache is only updated in the native code

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would it make sense for CoreCLR to just call this managed implementation instead?

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Interesting idea. I think what could prevent that is:

  • if we need to use the cache before we can run managed code.
    JIT calls into casting machinery, so it is hard to tell if we may need the cache too early.

  • if we need to use the cache before managed cctor has run.
    I'd guess the cctor could be forced to run early enough.

  • if we need to use cache from some mode that does not allow managed code.
    Most likely this is not a requirement, but I am not sure.

  • if calling managed code from native has high overhead.
    Adding to the cache is reasonably fast. We do not see the cost of adding stuff even at start up, when we'd expect cache misses.
    I do not have a good sense of how expensive it is to call managed code from the native runtime in this context.

Would we want to call managed TryGet as well?
(that would be a lot more sensitive to the overhead and TryGet is GC_NOTRIGGER, I vaguely remember there were some reasons for that.)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is worth considering, but it should be a separate change.

It does not look like a straightforward tweak and may need some experimenting.
Also the impact of such change will be mostly in CoreClr, while this PR mostly affects NativeAOT, so in terms of watching for failures or stress/perf consequences, it would be better to have separate changes.

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Logged an issue to follow up on switching CoreCLR to use managed TrySet and maybe TryGet - #84448

@VSadov

VSadov commented Apr 6, 2023

Copy link
Copy Markdown
MemberAuthor

I have run some simple scenarios with this change to see if there are no obvious performance regressions.

It looks like cached casting is slightly faster with the new cache, but not by much. The key point is to have a cache in the first place and NativeAOT already had it.

I also noticed that CoreClr is faster than NativeAOT. I think the reason is the code that runs before we hit the cache. That is - the code that lives in nativeaot\Runtime.Base\src\System\Runtime\TypeCast.cs. It looks like the entry points are a bit more complex than in CoreClr cast helpers.
Cached casts are fast and couple extra calls, indirections or branches may have an impact.
Also it could be that some extra complexity is necessary. Dealing with IsCloned appears to be specific to NativeAOT.
There are also couple TODOs about IsIDynamicInterfaceCastable that seem relevant to the overall perf.

It may be worth looking at the cast entry points. Perhaps there are some opportunities there.

The microbenchmark that I used to see impact of the cache in "easy" case:

internalclassProgram{constintiters=1000000;staticvoidMain(string[]args){for(;;){Time(TestLStringToIROCstring);}}staticvoidTime(Actiona){varsw=Stopwatch.StartNew();for(inti=0;i<100;i++){a();}sw.Stop();System.Console.WriteLine(sw.ElapsedMilliseconds);}staticobjecto=newList<string>();staticvoidTestLStringToIROCstring(){for(inti=0;i<iters;i++){if(oasIReadOnlyCollection<object>==null)thrownull;if(oasIReadOnlyCollection<string>==null)thrownull;if(oasIEnumerable<object>==null)thrownull;if(oasIEnumerable<string>==null)thrownull;}}}

On my machine (x64), I see:
(smaller is better)

==== before change:
871
868
874
870
870
870
==== after the change
858
860
863
864
859
856
==== CoreClr (JIT)
586
589
583
587
583
585

@jkotasjkotas left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

More code sharing. Yay!

@jkotas

Copy link
Copy Markdown
Member

Dealing with IsCloned appears to be specific to NativeAOT.

IsCloned support is not used. I would not feel bad about deleting it. I think it is unlikely that we will ever use it for anything.

CoreCLR has the methodtable flags specifically shaped to make casting fast. We may want to copy/unify the shape.

@VSadov

Copy link
Copy Markdown
MemberAuthor

CoreCLR has the methodtable flags specifically shaped to make casting fast. We may want to copy/unify the shape.

Right. Also the CoreClr entry points were carefully crafted to minimize branches, benefit from tail calling, if possible. Perhaps at some costs to readability, but for that code it is acceptable.

I will look at what we can borrow from there.

@MichalStrehovsky

Copy link
Copy Markdown
Member

Right. Also the CoreClr entry points were carefully crafted to minimize branches, benefit from tail calling, if possible. Perhaps at some costs to readability, but for that code it is acceptable.

We also have some unnecessary checks like "is this null" or "is this exactly the same type" for things that RyuJIT already tests for. The code was written for UTC, not RyuJIT. We just need to validate all the non-codegen callees (i.e. when we call into casting from our regular C# code) also do the appropriate checks and delete the ifs.

@VSadov

Copy link
Copy Markdown
MemberAuthor

Thanks!!

@VSadov
VSadov merged commit 38b81ba into dotnet:mainApr 7, 2023
@VSadov
VSadov deleted the castCacheN branch April 7, 2023 02:04
@ghostghost locked as resolved and limited conversation to collaborators May 7, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[NativeAOT] Consider porting/sharing CastCache form CoreClr

3 participants

@VSadov@jkotas@MichalStrehovsky
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

[NativeAOT] Using the same CastCache implementation as in CoreClr - #84430

Merged
VSadov merged 11 commits into
dotnet:mainfrom
VSadov:castCacheN
Apr 7, 2023
Merged

[NativeAOT] Using the same CastCache implementation as in CoreClr#84430
VSadov merged 11 commits into
dotnet:mainfrom
VSadov:castCacheN

Conversation

@VSadov

Copy link
Copy Markdown
Member

Fixes: #75111

Reasons:

  • code sharing
  • CoreClr implementation does not require a Crst , or any kind of locks.

@ghost

ghost commented Apr 6, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

Issue Details

Fixes: #75111

Reasons:

  • code sharing
  • CoreClr implementation does not require a Crst , or any kind of locks.
Author:VSadov
Assignees:-
Labels:

area-NativeAOT-coreclr

Milestone:-


#if TARGET_64BIT
[MethodImpl(MethodImplOptions.AggressiveInlining)]
private static ulong RotateLeft(ulong value, int offset)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This should be deleted and replaced by BitOperations.RotateLeft.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've introduced this because NativeAOT was not building in the test configuration that was missing a lot of things - like Volatile, Interlocked, BitOperations.
Later I moved custom implementations under Test.CoreLib, but missed this one.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We may want to omit this cache for Test.CoreLib. Test.CoreLib has simplistic implementation of number of other subsystems.

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

At this point it would be easy to move these helpers to Test.CoreLib - as custom BitOperations implementations. Then we can have the cache in Test.CorLib and have extra test coverage.

However, if having a minimal implementation is the point, then I can instead add a Test.CoreLib - specific implementation of the cache that does not do anything.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Then we can have the cache in Test.CorLib and have extra test coverage.

There are no relevant tests running against Test.CorLib. All tests that matter for this run against real CoreLib.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reverted additional API implementations for Test.CorLib and added a trivial no-op cache instead.

// - issue a load barrier before reading _version
// benchmarks on available hardware (Jan 2020) show that use of a read barrier is cheaper.

#if CORECLR

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If this is required for correctness, it should not be under ifdef.

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is either a load fence here or the two loads above need to be acquires. As it was measured, when this was implemented, one fence was noticeably cheaper.

We need to add Interlocked.ReadMemoryBarrier() to NativeAOT. And it needs to be an intrinsic - no point in implementing it as an internal call.

I will log an issue on Interlocked.ReadMemoryBarrier(). Once that is added, we can remove ifdefs here and around two acquires above.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Logged an issue for adding Interlocked.ReadMemoryBarrier() intrinsic - #84445

}

// the rest is the support for updating the cache.
// in CoreClr the cache is only updated in the native code

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would it make sense for CoreCLR to just call this managed implementation instead?

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Interesting idea. I think what could prevent that is:

  • if we need to use the cache before we can run managed code.
    JIT calls into casting machinery, so it is hard to tell if we may need the cache too early.

  • if we need to use the cache before managed cctor has run.
    I'd guess the cctor could be forced to run early enough.

  • if we need to use cache from some mode that does not allow managed code.
    Most likely this is not a requirement, but I am not sure.

  • if calling managed code from native has high overhead.
    Adding to the cache is reasonably fast. We do not see the cost of adding stuff even at start up, when we'd expect cache misses.
    I do not have a good sense of how expensive it is to call managed code from the native runtime in this context.

Would we want to call managed TryGet as well?
(that would be a lot more sensitive to the overhead and TryGet is GC_NOTRIGGER, I vaguely remember there were some reasons for that.)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is worth considering, but it should be a separate change.

It does not look like a straightforward tweak and may need some experimenting.
Also the impact of such change will be mostly in CoreClr, while this PR mostly affects NativeAOT, so in terms of watching for failures or stress/perf consequences, it would be better to have separate changes.

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Logged an issue to follow up on switching CoreCLR to use managed TrySet and maybe TryGet - #84448

@VSadov

VSadov commented Apr 6, 2023

Copy link
Copy Markdown
MemberAuthor

I have run some simple scenarios with this change to see if there are no obvious performance regressions.

It looks like cached casting is slightly faster with the new cache, but not by much. The key point is to have a cache in the first place and NativeAOT already had it.

I also noticed that CoreClr is faster than NativeAOT. I think the reason is the code that runs before we hit the cache. That is - the code that lives in nativeaot\Runtime.Base\src\System\Runtime\TypeCast.cs. It looks like the entry points are a bit more complex than in CoreClr cast helpers.
Cached casts are fast and couple extra calls, indirections or branches may have an impact.
Also it could be that some extra complexity is necessary. Dealing with IsCloned appears to be specific to NativeAOT.
There are also couple TODOs about IsIDynamicInterfaceCastable that seem relevant to the overall perf.

It may be worth looking at the cast entry points. Perhaps there are some opportunities there.

The microbenchmark that I used to see impact of the cache in "easy" case:

internalclassProgram{constintiters=1000000;staticvoidMain(string[]args){for(;;){Time(TestLStringToIROCstring);}}staticvoidTime(Actiona){varsw=Stopwatch.StartNew();for(inti=0;i<100;i++){a();}sw.Stop();System.Console.WriteLine(sw.ElapsedMilliseconds);}staticobjecto=newList<string>();staticvoidTestLStringToIROCstring(){for(inti=0;i<iters;i++){if(oasIReadOnlyCollection<object>==null)thrownull;if(oasIReadOnlyCollection<string>==null)thrownull;if(oasIEnumerable<object>==null)thrownull;if(oasIEnumerable<string>==null)thrownull;}}}

On my machine (x64), I see:
(smaller is better)

==== before change:
871
868
874
870
870
870
==== after the change
858
860
863
864
859
856
==== CoreClr (JIT)
586
589
583
587
583
585

@jkotasjkotas left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

More code sharing. Yay!

@jkotas

Copy link
Copy Markdown
Member

Dealing with IsCloned appears to be specific to NativeAOT.

IsCloned support is not used. I would not feel bad about deleting it. I think it is unlikely that we will ever use it for anything.

CoreCLR has the methodtable flags specifically shaped to make casting fast. We may want to copy/unify the shape.

@VSadov

Copy link
Copy Markdown
MemberAuthor

CoreCLR has the methodtable flags specifically shaped to make casting fast. We may want to copy/unify the shape.

Right. Also the CoreClr entry points were carefully crafted to minimize branches, benefit from tail calling, if possible. Perhaps at some costs to readability, but for that code it is acceptable.

I will look at what we can borrow from there.

@MichalStrehovsky

Copy link
Copy Markdown
Member

Right. Also the CoreClr entry points were carefully crafted to minimize branches, benefit from tail calling, if possible. Perhaps at some costs to readability, but for that code it is acceptable.

We also have some unnecessary checks like "is this null" or "is this exactly the same type" for things that RyuJIT already tests for. The code was written for UTC, not RyuJIT. We just need to validate all the non-codegen callees (i.e. when we call into casting from our regular C# code) also do the appropriate checks and delete the ifs.

@VSadov

Copy link
Copy Markdown
MemberAuthor

Thanks!!

@VSadov
VSadov merged commit 38b81ba into dotnet:mainApr 7, 2023
@VSadov
VSadov deleted the castCacheN branch April 7, 2023 02:04
@ghostghost locked as resolved and limited conversation to collaborators May 7, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[NativeAOT] Consider porting/sharing CastCache form CoreClr

3 participants

@VSadov@jkotas@MichalStrehovsky
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[NativeAOT] Using the same CastCache implementation as in CoreClr - #84430

Merged
VSadov merged 11 commits into
dotnet:mainfrom
VSadov:castCacheN
Apr 7, 2023
Merged

[NativeAOT] Using the same CastCache implementation as in CoreClr#84430
VSadov merged 11 commits into
dotnet:mainfrom
VSadov:castCacheN

Conversation

@VSadov

Copy link
Copy Markdown
Member

Fixes: #75111

Reasons:

  • code sharing
  • CoreClr implementation does not require a Crst , or any kind of locks.

@ghost

ghost commented Apr 6, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

Issue Details

Fixes: #75111

Reasons:

  • code sharing
  • CoreClr implementation does not require a Crst , or any kind of locks.
Author:VSadov
Assignees:-
Labels:

area-NativeAOT-coreclr

Milestone:-


#if TARGET_64BIT
[MethodImpl(MethodImplOptions.AggressiveInlining)]
private static ulong RotateLeft(ulong value, int offset)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This should be deleted and replaced by BitOperations.RotateLeft.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've introduced this because NativeAOT was not building in the test configuration that was missing a lot of things - like Volatile, Interlocked, BitOperations.
Later I moved custom implementations under Test.CoreLib, but missed this one.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We may want to omit this cache for Test.CoreLib. Test.CoreLib has simplistic implementation of number of other subsystems.

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

At this point it would be easy to move these helpers to Test.CoreLib - as custom BitOperations implementations. Then we can have the cache in Test.CorLib and have extra test coverage.

However, if having a minimal implementation is the point, then I can instead add a Test.CoreLib - specific implementation of the cache that does not do anything.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Then we can have the cache in Test.CorLib and have extra test coverage.

There are no relevant tests running against Test.CorLib. All tests that matter for this run against real CoreLib.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reverted additional API implementations for Test.CorLib and added a trivial no-op cache instead.

// - issue a load barrier before reading _version
// benchmarks on available hardware (Jan 2020) show that use of a read barrier is cheaper.

#if CORECLR

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If this is required for correctness, it should not be under ifdef.

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is either a load fence here or the two loads above need to be acquires. As it was measured, when this was implemented, one fence was noticeably cheaper.

We need to add Interlocked.ReadMemoryBarrier() to NativeAOT. And it needs to be an intrinsic - no point in implementing it as an internal call.

I will log an issue on Interlocked.ReadMemoryBarrier(). Once that is added, we can remove ifdefs here and around two acquires above.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Logged an issue for adding Interlocked.ReadMemoryBarrier() intrinsic - #84445

}

// the rest is the support for updating the cache.
// in CoreClr the cache is only updated in the native code

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would it make sense for CoreCLR to just call this managed implementation instead?

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Interesting idea. I think what could prevent that is:

  • if we need to use the cache before we can run managed code.
    JIT calls into casting machinery, so it is hard to tell if we may need the cache too early.

  • if we need to use the cache before managed cctor has run.
    I'd guess the cctor could be forced to run early enough.

  • if we need to use cache from some mode that does not allow managed code.
    Most likely this is not a requirement, but I am not sure.

  • if calling managed code from native has high overhead.
    Adding to the cache is reasonably fast. We do not see the cost of adding stuff even at start up, when we'd expect cache misses.
    I do not have a good sense of how expensive it is to call managed code from the native runtime in this context.

Would we want to call managed TryGet as well?
(that would be a lot more sensitive to the overhead and TryGet is GC_NOTRIGGER, I vaguely remember there were some reasons for that.)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is worth considering, but it should be a separate change.

It does not look like a straightforward tweak and may need some experimenting.
Also the impact of such change will be mostly in CoreClr, while this PR mostly affects NativeAOT, so in terms of watching for failures or stress/perf consequences, it would be better to have separate changes.

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Logged an issue to follow up on switching CoreCLR to use managed TrySet and maybe TryGet - #84448

@VSadov

VSadov commented Apr 6, 2023

Copy link
Copy Markdown
MemberAuthor

I have run some simple scenarios with this change to see if there are no obvious performance regressions.

It looks like cached casting is slightly faster with the new cache, but not by much. The key point is to have a cache in the first place and NativeAOT already had it.

I also noticed that CoreClr is faster than NativeAOT. I think the reason is the code that runs before we hit the cache. That is - the code that lives in nativeaot\Runtime.Base\src\System\Runtime\TypeCast.cs. It looks like the entry points are a bit more complex than in CoreClr cast helpers.
Cached casts are fast and couple extra calls, indirections or branches may have an impact.
Also it could be that some extra complexity is necessary. Dealing with IsCloned appears to be specific to NativeAOT.
There are also couple TODOs about IsIDynamicInterfaceCastable that seem relevant to the overall perf.

It may be worth looking at the cast entry points. Perhaps there are some opportunities there.

The microbenchmark that I used to see impact of the cache in "easy" case:

internalclassProgram{constintiters=1000000;staticvoidMain(string[]args){for(;;){Time(TestLStringToIROCstring);}}staticvoidTime(Actiona){varsw=Stopwatch.StartNew();for(inti=0;i<100;i++){a();}sw.Stop();System.Console.WriteLine(sw.ElapsedMilliseconds);}staticobjecto=newList<string>();staticvoidTestLStringToIROCstring(){for(inti=0;i<iters;i++){if(oasIReadOnlyCollection<object>==null)thrownull;if(oasIReadOnlyCollection<string>==null)thrownull;if(oasIEnumerable<object>==null)thrownull;if(oasIEnumerable<string>==null)thrownull;}}}

On my machine (x64), I see:
(smaller is better)

==== before change:
871
868
874
870
870
870
==== after the change
858
860
863
864
859
856
==== CoreClr (JIT)
586
589
583
587
583
585

@jkotasjkotas left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

More code sharing. Yay!

@jkotas

Copy link
Copy Markdown
Member

Dealing with IsCloned appears to be specific to NativeAOT.

IsCloned support is not used. I would not feel bad about deleting it. I think it is unlikely that we will ever use it for anything.

CoreCLR has the methodtable flags specifically shaped to make casting fast. We may want to copy/unify the shape.

@VSadov

Copy link
Copy Markdown
MemberAuthor

CoreCLR has the methodtable flags specifically shaped to make casting fast. We may want to copy/unify the shape.

Right. Also the CoreClr entry points were carefully crafted to minimize branches, benefit from tail calling, if possible. Perhaps at some costs to readability, but for that code it is acceptable.

I will look at what we can borrow from there.

@MichalStrehovsky

Copy link
Copy Markdown
Member

Right. Also the CoreClr entry points were carefully crafted to minimize branches, benefit from tail calling, if possible. Perhaps at some costs to readability, but for that code it is acceptable.

We also have some unnecessary checks like "is this null" or "is this exactly the same type" for things that RyuJIT already tests for. The code was written for UTC, not RyuJIT. We just need to validate all the non-codegen callees (i.e. when we call into casting from our regular C# code) also do the appropriate checks and delete the ifs.

@VSadov

Copy link
Copy Markdown
MemberAuthor

Thanks!!

@VSadov
VSadov merged commit 38b81ba into dotnet:mainApr 7, 2023
@VSadov
VSadov deleted the castCacheN branch April 7, 2023 02:04
@ghostghost locked as resolved and limited conversation to collaborators May 7, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[NativeAOT] Consider porting/sharing CastCache form CoreClr

3 participants

@VSadov@jkotas@MichalStrehovsky
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[NativeAOT] Using the same CastCache implementation as in CoreClr - #84430

Merged
VSadov merged 11 commits into
dotnet:mainfrom
VSadov:castCacheN
Apr 7, 2023
Merged

[NativeAOT] Using the same CastCache implementation as in CoreClr#84430
VSadov merged 11 commits into
dotnet:mainfrom
VSadov:castCacheN

Conversation

@VSadov

Copy link
Copy Markdown
Member

Fixes: #75111

Reasons:

  • code sharing
  • CoreClr implementation does not require a Crst , or any kind of locks.

@ghost

ghost commented Apr 6, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

Issue Details

Fixes: #75111

Reasons:

  • code sharing
  • CoreClr implementation does not require a Crst , or any kind of locks.
Author:VSadov
Assignees:-
Labels:

area-NativeAOT-coreclr

Milestone:-


#if TARGET_64BIT
[MethodImpl(MethodImplOptions.AggressiveInlining)]
private static ulong RotateLeft(ulong value, int offset)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This should be deleted and replaced by BitOperations.RotateLeft.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've introduced this because NativeAOT was not building in the test configuration that was missing a lot of things - like Volatile, Interlocked, BitOperations.
Later I moved custom implementations under Test.CoreLib, but missed this one.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We may want to omit this cache for Test.CoreLib. Test.CoreLib has simplistic implementation of number of other subsystems.

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

At this point it would be easy to move these helpers to Test.CoreLib - as custom BitOperations implementations. Then we can have the cache in Test.CorLib and have extra test coverage.

However, if having a minimal implementation is the point, then I can instead add a Test.CoreLib - specific implementation of the cache that does not do anything.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Then we can have the cache in Test.CorLib and have extra test coverage.

There are no relevant tests running against Test.CorLib. All tests that matter for this run against real CoreLib.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reverted additional API implementations for Test.CorLib and added a trivial no-op cache instead.

// - issue a load barrier before reading _version
// benchmarks on available hardware (Jan 2020) show that use of a read barrier is cheaper.

#if CORECLR

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If this is required for correctness, it should not be under ifdef.

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is either a load fence here or the two loads above need to be acquires. As it was measured, when this was implemented, one fence was noticeably cheaper.

We need to add Interlocked.ReadMemoryBarrier() to NativeAOT. And it needs to be an intrinsic - no point in implementing it as an internal call.

I will log an issue on Interlocked.ReadMemoryBarrier(). Once that is added, we can remove ifdefs here and around two acquires above.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Logged an issue for adding Interlocked.ReadMemoryBarrier() intrinsic - #84445

}

// the rest is the support for updating the cache.
// in CoreClr the cache is only updated in the native code

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would it make sense for CoreCLR to just call this managed implementation instead?

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Interesting idea. I think what could prevent that is:

  • if we need to use the cache before we can run managed code.
    JIT calls into casting machinery, so it is hard to tell if we may need the cache too early.

  • if we need to use the cache before managed cctor has run.
    I'd guess the cctor could be forced to run early enough.

  • if we need to use cache from some mode that does not allow managed code.
    Most likely this is not a requirement, but I am not sure.

  • if calling managed code from native has high overhead.
    Adding to the cache is reasonably fast. We do not see the cost of adding stuff even at start up, when we'd expect cache misses.
    I do not have a good sense of how expensive it is to call managed code from the native runtime in this context.

Would we want to call managed TryGet as well?
(that would be a lot more sensitive to the overhead and TryGet is GC_NOTRIGGER, I vaguely remember there were some reasons for that.)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is worth considering, but it should be a separate change.

It does not look like a straightforward tweak and may need some experimenting.
Also the impact of such change will be mostly in CoreClr, while this PR mostly affects NativeAOT, so in terms of watching for failures or stress/perf consequences, it would be better to have separate changes.

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Logged an issue to follow up on switching CoreCLR to use managed TrySet and maybe TryGet - #84448

@VSadov

VSadov commented Apr 6, 2023

Copy link
Copy Markdown
MemberAuthor

I have run some simple scenarios with this change to see if there are no obvious performance regressions.

It looks like cached casting is slightly faster with the new cache, but not by much. The key point is to have a cache in the first place and NativeAOT already had it.

I also noticed that CoreClr is faster than NativeAOT. I think the reason is the code that runs before we hit the cache. That is - the code that lives in nativeaot\Runtime.Base\src\System\Runtime\TypeCast.cs. It looks like the entry points are a bit more complex than in CoreClr cast helpers.
Cached casts are fast and couple extra calls, indirections or branches may have an impact.
Also it could be that some extra complexity is necessary. Dealing with IsCloned appears to be specific to NativeAOT.
There are also couple TODOs about IsIDynamicInterfaceCastable that seem relevant to the overall perf.

It may be worth looking at the cast entry points. Perhaps there are some opportunities there.

The microbenchmark that I used to see impact of the cache in "easy" case:

internalclassProgram{constintiters=1000000;staticvoidMain(string[]args){for(;;){Time(TestLStringToIROCstring);}}staticvoidTime(Actiona){varsw=Stopwatch.StartNew();for(inti=0;i<100;i++){a();}sw.Stop();System.Console.WriteLine(sw.ElapsedMilliseconds);}staticobjecto=newList<string>();staticvoidTestLStringToIROCstring(){for(inti=0;i<iters;i++){if(oasIReadOnlyCollection<object>==null)thrownull;if(oasIReadOnlyCollection<string>==null)thrownull;if(oasIEnumerable<object>==null)thrownull;if(oasIEnumerable<string>==null)thrownull;}}}

On my machine (x64), I see:
(smaller is better)

==== before change:
871
868
874
870
870
870
==== after the change
858
860
863
864
859
856
==== CoreClr (JIT)
586
589
583
587
583
585

@jkotasjkotas left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

More code sharing. Yay!

@jkotas

Copy link
Copy Markdown
Member

Dealing with IsCloned appears to be specific to NativeAOT.

IsCloned support is not used. I would not feel bad about deleting it. I think it is unlikely that we will ever use it for anything.

CoreCLR has the methodtable flags specifically shaped to make casting fast. We may want to copy/unify the shape.

@VSadov

Copy link
Copy Markdown
MemberAuthor

CoreCLR has the methodtable flags specifically shaped to make casting fast. We may want to copy/unify the shape.

Right. Also the CoreClr entry points were carefully crafted to minimize branches, benefit from tail calling, if possible. Perhaps at some costs to readability, but for that code it is acceptable.

I will look at what we can borrow from there.

@MichalStrehovsky

Copy link
Copy Markdown
Member

Right. Also the CoreClr entry points were carefully crafted to minimize branches, benefit from tail calling, if possible. Perhaps at some costs to readability, but for that code it is acceptable.

We also have some unnecessary checks like "is this null" or "is this exactly the same type" for things that RyuJIT already tests for. The code was written for UTC, not RyuJIT. We just need to validate all the non-codegen callees (i.e. when we call into casting from our regular C# code) also do the appropriate checks and delete the ifs.

@VSadov

Copy link
Copy Markdown
MemberAuthor

Thanks!!

@VSadov
VSadov merged commit 38b81ba into dotnet:mainApr 7, 2023
@VSadov
VSadov deleted the castCacheN branch April 7, 2023 02:04
@ghostghost locked as resolved and limited conversation to collaborators May 7, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[NativeAOT] Consider porting/sharing CastCache form CoreClr

3 participants

@VSadov@jkotas@MichalStrehovsky
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

[NativeAOT] Using the same CastCache implementation as in CoreClr - #84430

Merged
VSadov merged 11 commits into
dotnet:mainfrom
VSadov:castCacheN
Apr 7, 2023
Merged

[NativeAOT] Using the same CastCache implementation as in CoreClr#84430
VSadov merged 11 commits into
dotnet:mainfrom
VSadov:castCacheN

Conversation

@VSadov

Copy link
Copy Markdown
Member

Fixes: #75111

Reasons:

  • code sharing
  • CoreClr implementation does not require a Crst , or any kind of locks.

@ghost

ghost commented Apr 6, 2023

Copy link
Copy Markdown

Tagging subscribers to this area: @agocke, @MichalStrehovsky, @jkotas
See info in area-owners.md if you want to be subscribed.

Issue Details

Fixes: #75111

Reasons:

  • code sharing
  • CoreClr implementation does not require a Crst , or any kind of locks.
Author:VSadov
Assignees:-
Labels:

area-NativeAOT-coreclr

Milestone:-


#if TARGET_64BIT
[MethodImpl(MethodImplOptions.AggressiveInlining)]
private static ulong RotateLeft(ulong value, int offset)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This should be deleted and replaced by BitOperations.RotateLeft.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've introduced this because NativeAOT was not building in the test configuration that was missing a lot of things - like Volatile, Interlocked, BitOperations.
Later I moved custom implementations under Test.CoreLib, but missed this one.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We may want to omit this cache for Test.CoreLib. Test.CoreLib has simplistic implementation of number of other subsystems.

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

At this point it would be easy to move these helpers to Test.CoreLib - as custom BitOperations implementations. Then we can have the cache in Test.CorLib and have extra test coverage.

However, if having a minimal implementation is the point, then I can instead add a Test.CoreLib - specific implementation of the cache that does not do anything.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Then we can have the cache in Test.CorLib and have extra test coverage.

There are no relevant tests running against Test.CorLib. All tests that matter for this run against real CoreLib.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reverted additional API implementations for Test.CorLib and added a trivial no-op cache instead.

// - issue a load barrier before reading _version
// benchmarks on available hardware (Jan 2020) show that use of a read barrier is cheaper.

#if CORECLR

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If this is required for correctness, it should not be under ifdef.

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is either a load fence here or the two loads above need to be acquires. As it was measured, when this was implemented, one fence was noticeably cheaper.

We need to add Interlocked.ReadMemoryBarrier() to NativeAOT. And it needs to be an intrinsic - no point in implementing it as an internal call.

I will log an issue on Interlocked.ReadMemoryBarrier(). Once that is added, we can remove ifdefs here and around two acquires above.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Logged an issue for adding Interlocked.ReadMemoryBarrier() intrinsic - #84445

}

// the rest is the support for updating the cache.
// in CoreClr the cache is only updated in the native code

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would it make sense for CoreCLR to just call this managed implementation instead?

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Interesting idea. I think what could prevent that is:

  • if we need to use the cache before we can run managed code.
    JIT calls into casting machinery, so it is hard to tell if we may need the cache too early.

  • if we need to use the cache before managed cctor has run.
    I'd guess the cctor could be forced to run early enough.

  • if we need to use cache from some mode that does not allow managed code.
    Most likely this is not a requirement, but I am not sure.

  • if calling managed code from native has high overhead.
    Adding to the cache is reasonably fast. We do not see the cost of adding stuff even at start up, when we'd expect cache misses.
    I do not have a good sense of how expensive it is to call managed code from the native runtime in this context.

Would we want to call managed TryGet as well?
(that would be a lot more sensitive to the overhead and TryGet is GC_NOTRIGGER, I vaguely remember there were some reasons for that.)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is worth considering, but it should be a separate change.

It does not look like a straightforward tweak and may need some experimenting.
Also the impact of such change will be mostly in CoreClr, while this PR mostly affects NativeAOT, so in terms of watching for failures or stress/perf consequences, it would be better to have separate changes.

@VSadovVSadovApr 6, 2023

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Logged an issue to follow up on switching CoreCLR to use managed TrySet and maybe TryGet - #84448

@VSadov

VSadov commented Apr 6, 2023

Copy link
Copy Markdown
MemberAuthor

I have run some simple scenarios with this change to see if there are no obvious performance regressions.

It looks like cached casting is slightly faster with the new cache, but not by much. The key point is to have a cache in the first place and NativeAOT already had it.

I also noticed that CoreClr is faster than NativeAOT. I think the reason is the code that runs before we hit the cache. That is - the code that lives in nativeaot\Runtime.Base\src\System\Runtime\TypeCast.cs. It looks like the entry points are a bit more complex than in CoreClr cast helpers.
Cached casts are fast and couple extra calls, indirections or branches may have an impact.
Also it could be that some extra complexity is necessary. Dealing with IsCloned appears to be specific to NativeAOT.
There are also couple TODOs about IsIDynamicInterfaceCastable that seem relevant to the overall perf.

It may be worth looking at the cast entry points. Perhaps there are some opportunities there.

The microbenchmark that I used to see impact of the cache in "easy" case:

internalclassProgram{constintiters=1000000;staticvoidMain(string[]args){for(;;){Time(TestLStringToIROCstring);}}staticvoidTime(Actiona){varsw=Stopwatch.StartNew();for(inti=0;i<100;i++){a();}sw.Stop();System.Console.WriteLine(sw.ElapsedMilliseconds);}staticobjecto=newList<string>();staticvoidTestLStringToIROCstring(){for(inti=0;i<iters;i++){if(oasIReadOnlyCollection<object>==null)thrownull;if(oasIReadOnlyCollection<string>==null)thrownull;if(oasIEnumerable<object>==null)thrownull;if(oasIEnumerable<string>==null)thrownull;}}}

On my machine (x64), I see:
(smaller is better)

==== before change:
871
868
874
870
870
870
==== after the change
858
860
863
864
859
856
==== CoreClr (JIT)
586
589
583
587
583
585

@jkotasjkotas left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

More code sharing. Yay!

@jkotas

Copy link
Copy Markdown
Member

Dealing with IsCloned appears to be specific to NativeAOT.

IsCloned support is not used. I would not feel bad about deleting it. I think it is unlikely that we will ever use it for anything.

CoreCLR has the methodtable flags specifically shaped to make casting fast. We may want to copy/unify the shape.

@VSadov

Copy link
Copy Markdown
MemberAuthor

CoreCLR has the methodtable flags specifically shaped to make casting fast. We may want to copy/unify the shape.

Right. Also the CoreClr entry points were carefully crafted to minimize branches, benefit from tail calling, if possible. Perhaps at some costs to readability, but for that code it is acceptable.

I will look at what we can borrow from there.

@MichalStrehovsky

Copy link
Copy Markdown
Member

Right. Also the CoreClr entry points were carefully crafted to minimize branches, benefit from tail calling, if possible. Perhaps at some costs to readability, but for that code it is acceptable.

We also have some unnecessary checks like "is this null" or "is this exactly the same type" for things that RyuJIT already tests for. The code was written for UTC, not RyuJIT. We just need to validate all the non-codegen callees (i.e. when we call into casting from our regular C# code) also do the appropriate checks and delete the ifs.

@VSadov

Copy link
Copy Markdown
MemberAuthor

Thanks!!

@VSadov
VSadov merged commit 38b81ba into dotnet:mainApr 7, 2023
@VSadov
VSadov deleted the castCacheN branch April 7, 2023 02:04
@ghostghost locked as resolved and limited conversation to collaborators May 7, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[NativeAOT] Consider porting/sharing CastCache form CoreClr

3 participants

@VSadov@jkotas@MichalStrehovsky