Expose the FusedMultiplyAdd and MultiplyAddEstimate APIs on relevant vector and scalar types - #102181

Merged
tannergooding merged 10 commits into
dotnet:mainfrom
tannergooding:fix-93513
May 16, 2024
Merged

Expose the FusedMultiplyAdd and MultiplyAddEstimate APIs on relevant vector and scalar types#102181
tannergooding merged 10 commits into
dotnet:mainfrom
tannergooding:fix-93513

Conversation

@tannergooding

Copy link
Copy Markdown
Member

This resolves#98053 and makes progress towards #93513

@ghost

Copy link
Copy Markdown

Note regarding the new-api-needs-documentation label:

This serves as a reminder for when your PR is modifying a ref *.cs file and adding/modifying public APIs, please make sure the API implementation in the src *.cs file is documented with triple slash comments, so the PR reviewers can sign off that change.

@tannergooding
tannergooding marked this pull request as draft May 14, 2024 18:31
@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @dotnet/jit-contrib for the JIT side.

This is one of the few remaining cases that should need JIT side work, most of the remaining APIs in #93513 will be implemented purely in managed code. The *Estimate APIs are special as they need to work in the face of R2R due to being intentionally non-deterministic across hardware and FusedMultiplyAdd represents a case that optimizes down to a single instruction and which is a core part of many already deeply complex algorithms, so the benefits of guaranteeing expansion make it good to hook up to the existing intrinsic support.

@EgorBo

Copy link
Copy Markdown
Member

CC. @dotnet/jit-contrib for the JIT side.

This is one of the few remaining cases that should need JIT side work, most of the remaining APIs in #93513 will be implemented purely in managed code. The *Estimate APIs are special as they need to work in the face of R2R due to being intentionally non-deterministic across hardware and FusedMultiplyAdd represents a case that optimizes down to a single instruction and which is a core part of many already deeply complex algorithms, so the benefits of guaranteeing expansion make it good to hook up to the existing intrinsic support.

Should we rather introduce a new attribute to prevent methods from being pre-jitted (and inlined into pre-jitted code)?

// AdvSimd.FusedMultiplyAdd expects (addend, left, right), while the APIs take (left, right, addend)
// We expect op1 and op2 to have already been spilled

std::swap(op1, op3);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am just curious - who's responsible to spill them?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, nvm, I see

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Typically it's the importer code, prior to calling this method. Latter phases that might call such an API (like morph) are responsible for ensuring the swap is safe in the face of potential side effects.

There's, unfortunately, not really a way for us to do such a swap safely in the API itself (at least that I know of) nor to know if the caller has actually done it. So the typical approach has been to do the swap and comment that callers should be doing the validation.

impSpillSideEffect(true, verCurrentState.esStackDepth -
3 DEBUGARG("Spilling op1 side effects for MultiplyAddEstimate"));

impSpillSideEffect(true, verCurrentState.esStackDepth -

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Isn't better/simpler to use impSpillSideEffects(true, CHECK_SPILL_ALL DEBUGARG("spilling side-effects")); ?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's not something we've been doing in other scenarios.

AFAIR it comes down to not spilling values on the stack that aren't impacted. For example, we take 3 args, but the stack could have 4+ on it (spilling these is unnecessary since they are still processed in order with respect to our own op1/op2/op3; otherwise everyone who pops the stack would need to consider the need to spill such additional entries) and we don't need to ever spill the stack top (op3).

So we're doing this to ensure only the minimum number of items that need to be spilled are spilled.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

Should we rather introduce a new attribute to prevent methods from being pre-jitted (and inlined into pre-jitted code)?

We technically have one already, it's BypassReadyToRun (but we don't define or use it anyways as it was a legacy attribute for Crossgen1). The PR where I fixed this for Math.ReciprocalSqrtEstimate it was suggested to do it this other way instead.

I'm also not sure the number of *Estimate APIs is at the point where it's beneficial to define and expose that attribute and there's various other quirks that may need to be considered/solved in that scenario was well, plus the general consideration that some of these APIs are known to be used in hot paths of complex algorithms where inlining can fail, even with the inlining boost that methods using intrinsics get.

-- I do want to get to a point, someday, where we can move a lot of complexity out of the JIT and into the BCL instead. Ideally just the platform specific intrinsics and core xplat APIs which were exposed instead of platform specific APIs (like Create) are handled in the JIT and things like Vector2/3/4 or Vector128.WidenLower and APIs like that are handled entirely in managed code instead. I just don't think we're at that point yet today and need more work in the inliner and other places to ensure that we can get the desired codegen.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @stephentoub for the libraries side.

The one thing to note is that given how TensorPrimitives defines MultiplyAddEstimate and uses it in other APIs, the API was moved down to INumberBase<T> rather than being put on IFloatingPointIeee754<T>. This is still sensible for other types, it just always does the naive thing and isn't exposed as a member that's accessible from non-generic contexts.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

-- I do want to get to a point, someday, where we can move a lot of complexity out of the JIT and into the BCL instead. Ideally just the platform specific intrinsics and core xplat APIs which were exposed instead of platform specific APIs (like Create) are handled in the JIT and things like Vector2/3/4 or Vector128.WidenLower and APIs like that are handled entirely in managed code instead. I just don't think we're at that point yet today and need more work in the inliner and other places to ensure that we can get the desired codegen.

I've logged #102275 to track prototyping this work and seeing if we can identify some of the places that are still preventing it from happening.

This seems like a good thing to try and do around the time .NET 9 RC1 snaps and we have a bit of a small break before .NET 10 work really gets going.

@tannergooding
tannergooding merged commit b7727f5 into dotnet:mainMay 16, 2024
@tannergooding
tannergooding deleted the fix-93513 branch May 16, 2024 21:11
Ruihan-Yin pushed a commit to Ruihan-Yin/runtime that referenced this pull request May 30, 2024
…vector and scalar types (dotnet#102181)
* Expose FusedMultiplyAdd and MultiplyAddEstimate on the scalar and vector types
* Adding tests covering FusedMultiplyAdd and MultiplyAddEstimate for the vector types
* Ensure TensorPrimitives uses the xplat APIs on .NET 9+
* Apply formatting patch
* Fix an accidental change to GenericVectorTests
* Ensure Arm64 passes fma operands in the correct order
* Apply formatting patch
* Ensure all the Arm64 code paths are spilling and swapping operands
* Apply formatting patch
* Don't pop the stack value twice
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jun 16, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[API Proposal]: MultiplyAddEstimate

2 participants

@tannergooding@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Expose the FusedMultiplyAdd and MultiplyAddEstimate APIs on relevant vector and scalar types - #102181

Merged
tannergooding merged 10 commits into
dotnet:mainfrom
tannergooding:fix-93513
May 16, 2024
Merged

Expose the FusedMultiplyAdd and MultiplyAddEstimate APIs on relevant vector and scalar types#102181
tannergooding merged 10 commits into
dotnet:mainfrom
tannergooding:fix-93513

Conversation

@tannergooding

Copy link
Copy Markdown
Member

This resolves#98053 and makes progress towards #93513

@ghost

Copy link
Copy Markdown

Note regarding the new-api-needs-documentation label:

This serves as a reminder for when your PR is modifying a ref *.cs file and adding/modifying public APIs, please make sure the API implementation in the src *.cs file is documented with triple slash comments, so the PR reviewers can sign off that change.

@tannergooding
tannergooding marked this pull request as draft May 14, 2024 18:31
@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @dotnet/jit-contrib for the JIT side.

This is one of the few remaining cases that should need JIT side work, most of the remaining APIs in #93513 will be implemented purely in managed code. The *Estimate APIs are special as they need to work in the face of R2R due to being intentionally non-deterministic across hardware and FusedMultiplyAdd represents a case that optimizes down to a single instruction and which is a core part of many already deeply complex algorithms, so the benefits of guaranteeing expansion make it good to hook up to the existing intrinsic support.

@EgorBo

Copy link
Copy Markdown
Member

CC. @dotnet/jit-contrib for the JIT side.

This is one of the few remaining cases that should need JIT side work, most of the remaining APIs in #93513 will be implemented purely in managed code. The *Estimate APIs are special as they need to work in the face of R2R due to being intentionally non-deterministic across hardware and FusedMultiplyAdd represents a case that optimizes down to a single instruction and which is a core part of many already deeply complex algorithms, so the benefits of guaranteeing expansion make it good to hook up to the existing intrinsic support.

Should we rather introduce a new attribute to prevent methods from being pre-jitted (and inlined into pre-jitted code)?

// AdvSimd.FusedMultiplyAdd expects (addend, left, right), while the APIs take (left, right, addend)
// We expect op1 and op2 to have already been spilled

std::swap(op1, op3);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am just curious - who's responsible to spill them?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, nvm, I see

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Typically it's the importer code, prior to calling this method. Latter phases that might call such an API (like morph) are responsible for ensuring the swap is safe in the face of potential side effects.

There's, unfortunately, not really a way for us to do such a swap safely in the API itself (at least that I know of) nor to know if the caller has actually done it. So the typical approach has been to do the swap and comment that callers should be doing the validation.

impSpillSideEffect(true, verCurrentState.esStackDepth -
3 DEBUGARG("Spilling op1 side effects for MultiplyAddEstimate"));

impSpillSideEffect(true, verCurrentState.esStackDepth -

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Isn't better/simpler to use impSpillSideEffects(true, CHECK_SPILL_ALL DEBUGARG("spilling side-effects")); ?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's not something we've been doing in other scenarios.

AFAIR it comes down to not spilling values on the stack that aren't impacted. For example, we take 3 args, but the stack could have 4+ on it (spilling these is unnecessary since they are still processed in order with respect to our own op1/op2/op3; otherwise everyone who pops the stack would need to consider the need to spill such additional entries) and we don't need to ever spill the stack top (op3).

So we're doing this to ensure only the minimum number of items that need to be spilled are spilled.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

Should we rather introduce a new attribute to prevent methods from being pre-jitted (and inlined into pre-jitted code)?

We technically have one already, it's BypassReadyToRun (but we don't define or use it anyways as it was a legacy attribute for Crossgen1). The PR where I fixed this for Math.ReciprocalSqrtEstimate it was suggested to do it this other way instead.

I'm also not sure the number of *Estimate APIs is at the point where it's beneficial to define and expose that attribute and there's various other quirks that may need to be considered/solved in that scenario was well, plus the general consideration that some of these APIs are known to be used in hot paths of complex algorithms where inlining can fail, even with the inlining boost that methods using intrinsics get.

-- I do want to get to a point, someday, where we can move a lot of complexity out of the JIT and into the BCL instead. Ideally just the platform specific intrinsics and core xplat APIs which were exposed instead of platform specific APIs (like Create) are handled in the JIT and things like Vector2/3/4 or Vector128.WidenLower and APIs like that are handled entirely in managed code instead. I just don't think we're at that point yet today and need more work in the inliner and other places to ensure that we can get the desired codegen.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @stephentoub for the libraries side.

The one thing to note is that given how TensorPrimitives defines MultiplyAddEstimate and uses it in other APIs, the API was moved down to INumberBase<T> rather than being put on IFloatingPointIeee754<T>. This is still sensible for other types, it just always does the naive thing and isn't exposed as a member that's accessible from non-generic contexts.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

-- I do want to get to a point, someday, where we can move a lot of complexity out of the JIT and into the BCL instead. Ideally just the platform specific intrinsics and core xplat APIs which were exposed instead of platform specific APIs (like Create) are handled in the JIT and things like Vector2/3/4 or Vector128.WidenLower and APIs like that are handled entirely in managed code instead. I just don't think we're at that point yet today and need more work in the inliner and other places to ensure that we can get the desired codegen.

I've logged #102275 to track prototyping this work and seeing if we can identify some of the places that are still preventing it from happening.

This seems like a good thing to try and do around the time .NET 9 RC1 snaps and we have a bit of a small break before .NET 10 work really gets going.

@tannergooding
tannergooding merged commit b7727f5 into dotnet:mainMay 16, 2024
@tannergooding
tannergooding deleted the fix-93513 branch May 16, 2024 21:11
Ruihan-Yin pushed a commit to Ruihan-Yin/runtime that referenced this pull request May 30, 2024
…vector and scalar types (dotnet#102181)
* Expose FusedMultiplyAdd and MultiplyAddEstimate on the scalar and vector types
* Adding tests covering FusedMultiplyAdd and MultiplyAddEstimate for the vector types
* Ensure TensorPrimitives uses the xplat APIs on .NET 9+
* Apply formatting patch
* Fix an accidental change to GenericVectorTests
* Ensure Arm64 passes fma operands in the correct order
* Apply formatting patch
* Ensure all the Arm64 code paths are spilling and swapping operands
* Apply formatting patch
* Don't pop the stack value twice
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jun 16, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[API Proposal]: MultiplyAddEstimate

2 participants

@tannergooding@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Expose the FusedMultiplyAdd and MultiplyAddEstimate APIs on relevant vector and scalar types - #102181

Merged
tannergooding merged 10 commits into
dotnet:mainfrom
tannergooding:fix-93513
May 16, 2024
Merged

Expose the FusedMultiplyAdd and MultiplyAddEstimate APIs on relevant vector and scalar types#102181
tannergooding merged 10 commits into
dotnet:mainfrom
tannergooding:fix-93513

Conversation

@tannergooding

Copy link
Copy Markdown
Member

This resolves#98053 and makes progress towards #93513

@ghost

Copy link
Copy Markdown

Note regarding the new-api-needs-documentation label:

This serves as a reminder for when your PR is modifying a ref *.cs file and adding/modifying public APIs, please make sure the API implementation in the src *.cs file is documented with triple slash comments, so the PR reviewers can sign off that change.

@tannergooding
tannergooding marked this pull request as draft May 14, 2024 18:31
@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @dotnet/jit-contrib for the JIT side.

This is one of the few remaining cases that should need JIT side work, most of the remaining APIs in #93513 will be implemented purely in managed code. The *Estimate APIs are special as they need to work in the face of R2R due to being intentionally non-deterministic across hardware and FusedMultiplyAdd represents a case that optimizes down to a single instruction and which is a core part of many already deeply complex algorithms, so the benefits of guaranteeing expansion make it good to hook up to the existing intrinsic support.

@EgorBo

Copy link
Copy Markdown
Member

CC. @dotnet/jit-contrib for the JIT side.

This is one of the few remaining cases that should need JIT side work, most of the remaining APIs in #93513 will be implemented purely in managed code. The *Estimate APIs are special as they need to work in the face of R2R due to being intentionally non-deterministic across hardware and FusedMultiplyAdd represents a case that optimizes down to a single instruction and which is a core part of many already deeply complex algorithms, so the benefits of guaranteeing expansion make it good to hook up to the existing intrinsic support.

Should we rather introduce a new attribute to prevent methods from being pre-jitted (and inlined into pre-jitted code)?

// AdvSimd.FusedMultiplyAdd expects (addend, left, right), while the APIs take (left, right, addend)
// We expect op1 and op2 to have already been spilled

std::swap(op1, op3);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am just curious - who's responsible to spill them?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, nvm, I see

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Typically it's the importer code, prior to calling this method. Latter phases that might call such an API (like morph) are responsible for ensuring the swap is safe in the face of potential side effects.

There's, unfortunately, not really a way for us to do such a swap safely in the API itself (at least that I know of) nor to know if the caller has actually done it. So the typical approach has been to do the swap and comment that callers should be doing the validation.

impSpillSideEffect(true, verCurrentState.esStackDepth -
3 DEBUGARG("Spilling op1 side effects for MultiplyAddEstimate"));

impSpillSideEffect(true, verCurrentState.esStackDepth -

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Isn't better/simpler to use impSpillSideEffects(true, CHECK_SPILL_ALL DEBUGARG("spilling side-effects")); ?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's not something we've been doing in other scenarios.

AFAIR it comes down to not spilling values on the stack that aren't impacted. For example, we take 3 args, but the stack could have 4+ on it (spilling these is unnecessary since they are still processed in order with respect to our own op1/op2/op3; otherwise everyone who pops the stack would need to consider the need to spill such additional entries) and we don't need to ever spill the stack top (op3).

So we're doing this to ensure only the minimum number of items that need to be spilled are spilled.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

Should we rather introduce a new attribute to prevent methods from being pre-jitted (and inlined into pre-jitted code)?

We technically have one already, it's BypassReadyToRun (but we don't define or use it anyways as it was a legacy attribute for Crossgen1). The PR where I fixed this for Math.ReciprocalSqrtEstimate it was suggested to do it this other way instead.

I'm also not sure the number of *Estimate APIs is at the point where it's beneficial to define and expose that attribute and there's various other quirks that may need to be considered/solved in that scenario was well, plus the general consideration that some of these APIs are known to be used in hot paths of complex algorithms where inlining can fail, even with the inlining boost that methods using intrinsics get.

-- I do want to get to a point, someday, where we can move a lot of complexity out of the JIT and into the BCL instead. Ideally just the platform specific intrinsics and core xplat APIs which were exposed instead of platform specific APIs (like Create) are handled in the JIT and things like Vector2/3/4 or Vector128.WidenLower and APIs like that are handled entirely in managed code instead. I just don't think we're at that point yet today and need more work in the inliner and other places to ensure that we can get the desired codegen.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @stephentoub for the libraries side.

The one thing to note is that given how TensorPrimitives defines MultiplyAddEstimate and uses it in other APIs, the API was moved down to INumberBase<T> rather than being put on IFloatingPointIeee754<T>. This is still sensible for other types, it just always does the naive thing and isn't exposed as a member that's accessible from non-generic contexts.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

-- I do want to get to a point, someday, where we can move a lot of complexity out of the JIT and into the BCL instead. Ideally just the platform specific intrinsics and core xplat APIs which were exposed instead of platform specific APIs (like Create) are handled in the JIT and things like Vector2/3/4 or Vector128.WidenLower and APIs like that are handled entirely in managed code instead. I just don't think we're at that point yet today and need more work in the inliner and other places to ensure that we can get the desired codegen.

I've logged #102275 to track prototyping this work and seeing if we can identify some of the places that are still preventing it from happening.

This seems like a good thing to try and do around the time .NET 9 RC1 snaps and we have a bit of a small break before .NET 10 work really gets going.

@tannergooding
tannergooding merged commit b7727f5 into dotnet:mainMay 16, 2024
@tannergooding
tannergooding deleted the fix-93513 branch May 16, 2024 21:11
Ruihan-Yin pushed a commit to Ruihan-Yin/runtime that referenced this pull request May 30, 2024
…vector and scalar types (dotnet#102181)
* Expose FusedMultiplyAdd and MultiplyAddEstimate on the scalar and vector types
* Adding tests covering FusedMultiplyAdd and MultiplyAddEstimate for the vector types
* Ensure TensorPrimitives uses the xplat APIs on .NET 9+
* Apply formatting patch
* Fix an accidental change to GenericVectorTests
* Ensure Arm64 passes fma operands in the correct order
* Apply formatting patch
* Ensure all the Arm64 code paths are spilling and swapping operands
* Apply formatting patch
* Don't pop the stack value twice
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jun 16, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[API Proposal]: MultiplyAddEstimate

2 participants

@tannergooding@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Expose the FusedMultiplyAdd and MultiplyAddEstimate APIs on relevant vector and scalar types - #102181

Merged
tannergooding merged 10 commits into
dotnet:mainfrom
tannergooding:fix-93513
May 16, 2024
Merged

Expose the FusedMultiplyAdd and MultiplyAddEstimate APIs on relevant vector and scalar types#102181
tannergooding merged 10 commits into
dotnet:mainfrom
tannergooding:fix-93513

Conversation

@tannergooding

Copy link
Copy Markdown
Member

This resolves#98053 and makes progress towards #93513

@ghost

Copy link
Copy Markdown

Note regarding the new-api-needs-documentation label:

This serves as a reminder for when your PR is modifying a ref *.cs file and adding/modifying public APIs, please make sure the API implementation in the src *.cs file is documented with triple slash comments, so the PR reviewers can sign off that change.

@tannergooding
tannergooding marked this pull request as draft May 14, 2024 18:31
@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @dotnet/jit-contrib for the JIT side.

This is one of the few remaining cases that should need JIT side work, most of the remaining APIs in #93513 will be implemented purely in managed code. The *Estimate APIs are special as they need to work in the face of R2R due to being intentionally non-deterministic across hardware and FusedMultiplyAdd represents a case that optimizes down to a single instruction and which is a core part of many already deeply complex algorithms, so the benefits of guaranteeing expansion make it good to hook up to the existing intrinsic support.

@EgorBo

Copy link
Copy Markdown
Member

CC. @dotnet/jit-contrib for the JIT side.

This is one of the few remaining cases that should need JIT side work, most of the remaining APIs in #93513 will be implemented purely in managed code. The *Estimate APIs are special as they need to work in the face of R2R due to being intentionally non-deterministic across hardware and FusedMultiplyAdd represents a case that optimizes down to a single instruction and which is a core part of many already deeply complex algorithms, so the benefits of guaranteeing expansion make it good to hook up to the existing intrinsic support.

Should we rather introduce a new attribute to prevent methods from being pre-jitted (and inlined into pre-jitted code)?

// AdvSimd.FusedMultiplyAdd expects (addend, left, right), while the APIs take (left, right, addend)
// We expect op1 and op2 to have already been spilled

std::swap(op1, op3);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am just curious - who's responsible to spill them?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, nvm, I see

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Typically it's the importer code, prior to calling this method. Latter phases that might call such an API (like morph) are responsible for ensuring the swap is safe in the face of potential side effects.

There's, unfortunately, not really a way for us to do such a swap safely in the API itself (at least that I know of) nor to know if the caller has actually done it. So the typical approach has been to do the swap and comment that callers should be doing the validation.

impSpillSideEffect(true, verCurrentState.esStackDepth -
3 DEBUGARG("Spilling op1 side effects for MultiplyAddEstimate"));

impSpillSideEffect(true, verCurrentState.esStackDepth -

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Isn't better/simpler to use impSpillSideEffects(true, CHECK_SPILL_ALL DEBUGARG("spilling side-effects")); ?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's not something we've been doing in other scenarios.

AFAIR it comes down to not spilling values on the stack that aren't impacted. For example, we take 3 args, but the stack could have 4+ on it (spilling these is unnecessary since they are still processed in order with respect to our own op1/op2/op3; otherwise everyone who pops the stack would need to consider the need to spill such additional entries) and we don't need to ever spill the stack top (op3).

So we're doing this to ensure only the minimum number of items that need to be spilled are spilled.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

Should we rather introduce a new attribute to prevent methods from being pre-jitted (and inlined into pre-jitted code)?

We technically have one already, it's BypassReadyToRun (but we don't define or use it anyways as it was a legacy attribute for Crossgen1). The PR where I fixed this for Math.ReciprocalSqrtEstimate it was suggested to do it this other way instead.

I'm also not sure the number of *Estimate APIs is at the point where it's beneficial to define and expose that attribute and there's various other quirks that may need to be considered/solved in that scenario was well, plus the general consideration that some of these APIs are known to be used in hot paths of complex algorithms where inlining can fail, even with the inlining boost that methods using intrinsics get.

-- I do want to get to a point, someday, where we can move a lot of complexity out of the JIT and into the BCL instead. Ideally just the platform specific intrinsics and core xplat APIs which were exposed instead of platform specific APIs (like Create) are handled in the JIT and things like Vector2/3/4 or Vector128.WidenLower and APIs like that are handled entirely in managed code instead. I just don't think we're at that point yet today and need more work in the inliner and other places to ensure that we can get the desired codegen.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @stephentoub for the libraries side.

The one thing to note is that given how TensorPrimitives defines MultiplyAddEstimate and uses it in other APIs, the API was moved down to INumberBase<T> rather than being put on IFloatingPointIeee754<T>. This is still sensible for other types, it just always does the naive thing and isn't exposed as a member that's accessible from non-generic contexts.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

-- I do want to get to a point, someday, where we can move a lot of complexity out of the JIT and into the BCL instead. Ideally just the platform specific intrinsics and core xplat APIs which were exposed instead of platform specific APIs (like Create) are handled in the JIT and things like Vector2/3/4 or Vector128.WidenLower and APIs like that are handled entirely in managed code instead. I just don't think we're at that point yet today and need more work in the inliner and other places to ensure that we can get the desired codegen.

I've logged #102275 to track prototyping this work and seeing if we can identify some of the places that are still preventing it from happening.

This seems like a good thing to try and do around the time .NET 9 RC1 snaps and we have a bit of a small break before .NET 10 work really gets going.

@tannergooding
tannergooding merged commit b7727f5 into dotnet:mainMay 16, 2024
@tannergooding
tannergooding deleted the fix-93513 branch May 16, 2024 21:11
Ruihan-Yin pushed a commit to Ruihan-Yin/runtime that referenced this pull request May 30, 2024
…vector and scalar types (dotnet#102181)
* Expose FusedMultiplyAdd and MultiplyAddEstimate on the scalar and vector types
* Adding tests covering FusedMultiplyAdd and MultiplyAddEstimate for the vector types
* Ensure TensorPrimitives uses the xplat APIs on .NET 9+
* Apply formatting patch
* Fix an accidental change to GenericVectorTests
* Ensure Arm64 passes fma operands in the correct order
* Apply formatting patch
* Ensure all the Arm64 code paths are spilling and swapping operands
* Apply formatting patch
* Don't pop the stack value twice
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jun 16, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[API Proposal]: MultiplyAddEstimate

2 participants

@tannergooding@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Expose the FusedMultiplyAdd and MultiplyAddEstimate APIs on relevant vector and scalar types - #102181

Merged
tannergooding merged 10 commits into
dotnet:mainfrom
tannergooding:fix-93513
May 16, 2024
Merged

Expose the FusedMultiplyAdd and MultiplyAddEstimate APIs on relevant vector and scalar types#102181
tannergooding merged 10 commits into
dotnet:mainfrom
tannergooding:fix-93513

Conversation

@tannergooding

Copy link
Copy Markdown
Member

This resolves#98053 and makes progress towards #93513

@ghost

Copy link
Copy Markdown

Note regarding the new-api-needs-documentation label:

This serves as a reminder for when your PR is modifying a ref *.cs file and adding/modifying public APIs, please make sure the API implementation in the src *.cs file is documented with triple slash comments, so the PR reviewers can sign off that change.

@tannergooding
tannergooding marked this pull request as draft May 14, 2024 18:31
@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @dotnet/jit-contrib for the JIT side.

This is one of the few remaining cases that should need JIT side work, most of the remaining APIs in #93513 will be implemented purely in managed code. The *Estimate APIs are special as they need to work in the face of R2R due to being intentionally non-deterministic across hardware and FusedMultiplyAdd represents a case that optimizes down to a single instruction and which is a core part of many already deeply complex algorithms, so the benefits of guaranteeing expansion make it good to hook up to the existing intrinsic support.

@EgorBo

Copy link
Copy Markdown
Member

CC. @dotnet/jit-contrib for the JIT side.

This is one of the few remaining cases that should need JIT side work, most of the remaining APIs in #93513 will be implemented purely in managed code. The *Estimate APIs are special as they need to work in the face of R2R due to being intentionally non-deterministic across hardware and FusedMultiplyAdd represents a case that optimizes down to a single instruction and which is a core part of many already deeply complex algorithms, so the benefits of guaranteeing expansion make it good to hook up to the existing intrinsic support.

Should we rather introduce a new attribute to prevent methods from being pre-jitted (and inlined into pre-jitted code)?

// AdvSimd.FusedMultiplyAdd expects (addend, left, right), while the APIs take (left, right, addend)
// We expect op1 and op2 to have already been spilled

std::swap(op1, op3);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am just curious - who's responsible to spill them?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, nvm, I see

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Typically it's the importer code, prior to calling this method. Latter phases that might call such an API (like morph) are responsible for ensuring the swap is safe in the face of potential side effects.

There's, unfortunately, not really a way for us to do such a swap safely in the API itself (at least that I know of) nor to know if the caller has actually done it. So the typical approach has been to do the swap and comment that callers should be doing the validation.

impSpillSideEffect(true, verCurrentState.esStackDepth -
3 DEBUGARG("Spilling op1 side effects for MultiplyAddEstimate"));

impSpillSideEffect(true, verCurrentState.esStackDepth -

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Isn't better/simpler to use impSpillSideEffects(true, CHECK_SPILL_ALL DEBUGARG("spilling side-effects")); ?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's not something we've been doing in other scenarios.

AFAIR it comes down to not spilling values on the stack that aren't impacted. For example, we take 3 args, but the stack could have 4+ on it (spilling these is unnecessary since they are still processed in order with respect to our own op1/op2/op3; otherwise everyone who pops the stack would need to consider the need to spill such additional entries) and we don't need to ever spill the stack top (op3).

So we're doing this to ensure only the minimum number of items that need to be spilled are spilled.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

Should we rather introduce a new attribute to prevent methods from being pre-jitted (and inlined into pre-jitted code)?

We technically have one already, it's BypassReadyToRun (but we don't define or use it anyways as it was a legacy attribute for Crossgen1). The PR where I fixed this for Math.ReciprocalSqrtEstimate it was suggested to do it this other way instead.

I'm also not sure the number of *Estimate APIs is at the point where it's beneficial to define and expose that attribute and there's various other quirks that may need to be considered/solved in that scenario was well, plus the general consideration that some of these APIs are known to be used in hot paths of complex algorithms where inlining can fail, even with the inlining boost that methods using intrinsics get.

-- I do want to get to a point, someday, where we can move a lot of complexity out of the JIT and into the BCL instead. Ideally just the platform specific intrinsics and core xplat APIs which were exposed instead of platform specific APIs (like Create) are handled in the JIT and things like Vector2/3/4 or Vector128.WidenLower and APIs like that are handled entirely in managed code instead. I just don't think we're at that point yet today and need more work in the inliner and other places to ensure that we can get the desired codegen.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @stephentoub for the libraries side.

The one thing to note is that given how TensorPrimitives defines MultiplyAddEstimate and uses it in other APIs, the API was moved down to INumberBase<T> rather than being put on IFloatingPointIeee754<T>. This is still sensible for other types, it just always does the naive thing and isn't exposed as a member that's accessible from non-generic contexts.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

-- I do want to get to a point, someday, where we can move a lot of complexity out of the JIT and into the BCL instead. Ideally just the platform specific intrinsics and core xplat APIs which were exposed instead of platform specific APIs (like Create) are handled in the JIT and things like Vector2/3/4 or Vector128.WidenLower and APIs like that are handled entirely in managed code instead. I just don't think we're at that point yet today and need more work in the inliner and other places to ensure that we can get the desired codegen.

I've logged #102275 to track prototyping this work and seeing if we can identify some of the places that are still preventing it from happening.

This seems like a good thing to try and do around the time .NET 9 RC1 snaps and we have a bit of a small break before .NET 10 work really gets going.

@tannergooding
tannergooding merged commit b7727f5 into dotnet:mainMay 16, 2024
@tannergooding
tannergooding deleted the fix-93513 branch May 16, 2024 21:11
Ruihan-Yin pushed a commit to Ruihan-Yin/runtime that referenced this pull request May 30, 2024
…vector and scalar types (dotnet#102181)
* Expose FusedMultiplyAdd and MultiplyAddEstimate on the scalar and vector types
* Adding tests covering FusedMultiplyAdd and MultiplyAddEstimate for the vector types
* Ensure TensorPrimitives uses the xplat APIs on .NET 9+
* Apply formatting patch
* Fix an accidental change to GenericVectorTests
* Ensure Arm64 passes fma operands in the correct order
* Apply formatting patch
* Ensure all the Arm64 code paths are spilling and swapping operands
* Apply formatting patch
* Don't pop the stack value twice
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jun 16, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[API Proposal]: MultiplyAddEstimate

2 participants

@tannergooding@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Expose the FusedMultiplyAdd and MultiplyAddEstimate APIs on relevant vector and scalar types - #102181

Merged
tannergooding merged 10 commits into
dotnet:mainfrom
tannergooding:fix-93513
May 16, 2024
Merged

Expose the FusedMultiplyAdd and MultiplyAddEstimate APIs on relevant vector and scalar types#102181
tannergooding merged 10 commits into
dotnet:mainfrom
tannergooding:fix-93513

Conversation

@tannergooding

Copy link
Copy Markdown
Member

This resolves#98053 and makes progress towards #93513

@ghost

Copy link
Copy Markdown

Note regarding the new-api-needs-documentation label:

This serves as a reminder for when your PR is modifying a ref *.cs file and adding/modifying public APIs, please make sure the API implementation in the src *.cs file is documented with triple slash comments, so the PR reviewers can sign off that change.

@tannergooding
tannergooding marked this pull request as draft May 14, 2024 18:31
@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @dotnet/jit-contrib for the JIT side.

This is one of the few remaining cases that should need JIT side work, most of the remaining APIs in #93513 will be implemented purely in managed code. The *Estimate APIs are special as they need to work in the face of R2R due to being intentionally non-deterministic across hardware and FusedMultiplyAdd represents a case that optimizes down to a single instruction and which is a core part of many already deeply complex algorithms, so the benefits of guaranteeing expansion make it good to hook up to the existing intrinsic support.

@EgorBo

Copy link
Copy Markdown
Member

CC. @dotnet/jit-contrib for the JIT side.

This is one of the few remaining cases that should need JIT side work, most of the remaining APIs in #93513 will be implemented purely in managed code. The *Estimate APIs are special as they need to work in the face of R2R due to being intentionally non-deterministic across hardware and FusedMultiplyAdd represents a case that optimizes down to a single instruction and which is a core part of many already deeply complex algorithms, so the benefits of guaranteeing expansion make it good to hook up to the existing intrinsic support.

Should we rather introduce a new attribute to prevent methods from being pre-jitted (and inlined into pre-jitted code)?

// AdvSimd.FusedMultiplyAdd expects (addend, left, right), while the APIs take (left, right, addend)
// We expect op1 and op2 to have already been spilled

std::swap(op1, op3);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am just curious - who's responsible to spill them?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, nvm, I see

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Typically it's the importer code, prior to calling this method. Latter phases that might call such an API (like morph) are responsible for ensuring the swap is safe in the face of potential side effects.

There's, unfortunately, not really a way for us to do such a swap safely in the API itself (at least that I know of) nor to know if the caller has actually done it. So the typical approach has been to do the swap and comment that callers should be doing the validation.

impSpillSideEffect(true, verCurrentState.esStackDepth -
3 DEBUGARG("Spilling op1 side effects for MultiplyAddEstimate"));

impSpillSideEffect(true, verCurrentState.esStackDepth -

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Isn't better/simpler to use impSpillSideEffects(true, CHECK_SPILL_ALL DEBUGARG("spilling side-effects")); ?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's not something we've been doing in other scenarios.

AFAIR it comes down to not spilling values on the stack that aren't impacted. For example, we take 3 args, but the stack could have 4+ on it (spilling these is unnecessary since they are still processed in order with respect to our own op1/op2/op3; otherwise everyone who pops the stack would need to consider the need to spill such additional entries) and we don't need to ever spill the stack top (op3).

So we're doing this to ensure only the minimum number of items that need to be spilled are spilled.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

Should we rather introduce a new attribute to prevent methods from being pre-jitted (and inlined into pre-jitted code)?

We technically have one already, it's BypassReadyToRun (but we don't define or use it anyways as it was a legacy attribute for Crossgen1). The PR where I fixed this for Math.ReciprocalSqrtEstimate it was suggested to do it this other way instead.

I'm also not sure the number of *Estimate APIs is at the point where it's beneficial to define and expose that attribute and there's various other quirks that may need to be considered/solved in that scenario was well, plus the general consideration that some of these APIs are known to be used in hot paths of complex algorithms where inlining can fail, even with the inlining boost that methods using intrinsics get.

-- I do want to get to a point, someday, where we can move a lot of complexity out of the JIT and into the BCL instead. Ideally just the platform specific intrinsics and core xplat APIs which were exposed instead of platform specific APIs (like Create) are handled in the JIT and things like Vector2/3/4 or Vector128.WidenLower and APIs like that are handled entirely in managed code instead. I just don't think we're at that point yet today and need more work in the inliner and other places to ensure that we can get the desired codegen.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @stephentoub for the libraries side.

The one thing to note is that given how TensorPrimitives defines MultiplyAddEstimate and uses it in other APIs, the API was moved down to INumberBase<T> rather than being put on IFloatingPointIeee754<T>. This is still sensible for other types, it just always does the naive thing and isn't exposed as a member that's accessible from non-generic contexts.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

-- I do want to get to a point, someday, where we can move a lot of complexity out of the JIT and into the BCL instead. Ideally just the platform specific intrinsics and core xplat APIs which were exposed instead of platform specific APIs (like Create) are handled in the JIT and things like Vector2/3/4 or Vector128.WidenLower and APIs like that are handled entirely in managed code instead. I just don't think we're at that point yet today and need more work in the inliner and other places to ensure that we can get the desired codegen.

I've logged #102275 to track prototyping this work and seeing if we can identify some of the places that are still preventing it from happening.

This seems like a good thing to try and do around the time .NET 9 RC1 snaps and we have a bit of a small break before .NET 10 work really gets going.

@tannergooding
tannergooding merged commit b7727f5 into dotnet:mainMay 16, 2024
@tannergooding
tannergooding deleted the fix-93513 branch May 16, 2024 21:11
Ruihan-Yin pushed a commit to Ruihan-Yin/runtime that referenced this pull request May 30, 2024
…vector and scalar types (dotnet#102181)
* Expose FusedMultiplyAdd and MultiplyAddEstimate on the scalar and vector types
* Adding tests covering FusedMultiplyAdd and MultiplyAddEstimate for the vector types
* Ensure TensorPrimitives uses the xplat APIs on .NET 9+
* Apply formatting patch
* Fix an accidental change to GenericVectorTests
* Ensure Arm64 passes fma operands in the correct order
* Apply formatting patch
* Ensure all the Arm64 code paths are spilling and swapping operands
* Apply formatting patch
* Don't pop the stack value twice
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jun 16, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[API Proposal]: MultiplyAddEstimate

2 participants

@tannergooding@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Expose the FusedMultiplyAdd and MultiplyAddEstimate APIs on relevant vector and scalar types - #102181

Merged
tannergooding merged 10 commits into
dotnet:mainfrom
tannergooding:fix-93513
May 16, 2024
Merged

Expose the FusedMultiplyAdd and MultiplyAddEstimate APIs on relevant vector and scalar types#102181
tannergooding merged 10 commits into
dotnet:mainfrom
tannergooding:fix-93513

Conversation

@tannergooding

Copy link
Copy Markdown
Member

This resolves#98053 and makes progress towards #93513

@ghost

Copy link
Copy Markdown

Note regarding the new-api-needs-documentation label:

This serves as a reminder for when your PR is modifying a ref *.cs file and adding/modifying public APIs, please make sure the API implementation in the src *.cs file is documented with triple slash comments, so the PR reviewers can sign off that change.

@tannergooding
tannergooding marked this pull request as draft May 14, 2024 18:31
@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @dotnet/jit-contrib for the JIT side.

This is one of the few remaining cases that should need JIT side work, most of the remaining APIs in #93513 will be implemented purely in managed code. The *Estimate APIs are special as they need to work in the face of R2R due to being intentionally non-deterministic across hardware and FusedMultiplyAdd represents a case that optimizes down to a single instruction and which is a core part of many already deeply complex algorithms, so the benefits of guaranteeing expansion make it good to hook up to the existing intrinsic support.

@EgorBo

Copy link
Copy Markdown
Member

CC. @dotnet/jit-contrib for the JIT side.

This is one of the few remaining cases that should need JIT side work, most of the remaining APIs in #93513 will be implemented purely in managed code. The *Estimate APIs are special as they need to work in the face of R2R due to being intentionally non-deterministic across hardware and FusedMultiplyAdd represents a case that optimizes down to a single instruction and which is a core part of many already deeply complex algorithms, so the benefits of guaranteeing expansion make it good to hook up to the existing intrinsic support.

Should we rather introduce a new attribute to prevent methods from being pre-jitted (and inlined into pre-jitted code)?

// AdvSimd.FusedMultiplyAdd expects (addend, left, right), while the APIs take (left, right, addend)
// We expect op1 and op2 to have already been spilled

std::swap(op1, op3);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am just curious - who's responsible to spill them?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, nvm, I see

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Typically it's the importer code, prior to calling this method. Latter phases that might call such an API (like morph) are responsible for ensuring the swap is safe in the face of potential side effects.

There's, unfortunately, not really a way for us to do such a swap safely in the API itself (at least that I know of) nor to know if the caller has actually done it. So the typical approach has been to do the swap and comment that callers should be doing the validation.

impSpillSideEffect(true, verCurrentState.esStackDepth -
3 DEBUGARG("Spilling op1 side effects for MultiplyAddEstimate"));

impSpillSideEffect(true, verCurrentState.esStackDepth -

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Isn't better/simpler to use impSpillSideEffects(true, CHECK_SPILL_ALL DEBUGARG("spilling side-effects")); ?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's not something we've been doing in other scenarios.

AFAIR it comes down to not spilling values on the stack that aren't impacted. For example, we take 3 args, but the stack could have 4+ on it (spilling these is unnecessary since they are still processed in order with respect to our own op1/op2/op3; otherwise everyone who pops the stack would need to consider the need to spill such additional entries) and we don't need to ever spill the stack top (op3).

So we're doing this to ensure only the minimum number of items that need to be spilled are spilled.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

Should we rather introduce a new attribute to prevent methods from being pre-jitted (and inlined into pre-jitted code)?

We technically have one already, it's BypassReadyToRun (but we don't define or use it anyways as it was a legacy attribute for Crossgen1). The PR where I fixed this for Math.ReciprocalSqrtEstimate it was suggested to do it this other way instead.

I'm also not sure the number of *Estimate APIs is at the point where it's beneficial to define and expose that attribute and there's various other quirks that may need to be considered/solved in that scenario was well, plus the general consideration that some of these APIs are known to be used in hot paths of complex algorithms where inlining can fail, even with the inlining boost that methods using intrinsics get.

-- I do want to get to a point, someday, where we can move a lot of complexity out of the JIT and into the BCL instead. Ideally just the platform specific intrinsics and core xplat APIs which were exposed instead of platform specific APIs (like Create) are handled in the JIT and things like Vector2/3/4 or Vector128.WidenLower and APIs like that are handled entirely in managed code instead. I just don't think we're at that point yet today and need more work in the inliner and other places to ensure that we can get the desired codegen.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @stephentoub for the libraries side.

The one thing to note is that given how TensorPrimitives defines MultiplyAddEstimate and uses it in other APIs, the API was moved down to INumberBase<T> rather than being put on IFloatingPointIeee754<T>. This is still sensible for other types, it just always does the naive thing and isn't exposed as a member that's accessible from non-generic contexts.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

-- I do want to get to a point, someday, where we can move a lot of complexity out of the JIT and into the BCL instead. Ideally just the platform specific intrinsics and core xplat APIs which were exposed instead of platform specific APIs (like Create) are handled in the JIT and things like Vector2/3/4 or Vector128.WidenLower and APIs like that are handled entirely in managed code instead. I just don't think we're at that point yet today and need more work in the inliner and other places to ensure that we can get the desired codegen.

I've logged #102275 to track prototyping this work and seeing if we can identify some of the places that are still preventing it from happening.

This seems like a good thing to try and do around the time .NET 9 RC1 snaps and we have a bit of a small break before .NET 10 work really gets going.

@tannergooding
tannergooding merged commit b7727f5 into dotnet:mainMay 16, 2024
@tannergooding
tannergooding deleted the fix-93513 branch May 16, 2024 21:11
Ruihan-Yin pushed a commit to Ruihan-Yin/runtime that referenced this pull request May 30, 2024
…vector and scalar types (dotnet#102181)
* Expose FusedMultiplyAdd and MultiplyAddEstimate on the scalar and vector types
* Adding tests covering FusedMultiplyAdd and MultiplyAddEstimate for the vector types
* Ensure TensorPrimitives uses the xplat APIs on .NET 9+
* Apply formatting patch
* Fix an accidental change to GenericVectorTests
* Ensure Arm64 passes fma operands in the correct order
* Apply formatting patch
* Ensure all the Arm64 code paths are spilling and swapping operands
* Apply formatting patch
* Don't pop the stack value twice
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jun 16, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[API Proposal]: MultiplyAddEstimate

2 participants

@tannergooding@EgorBo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Expose the FusedMultiplyAdd and MultiplyAddEstimate APIs on relevant vector and scalar types - #102181

Merged
tannergooding merged 10 commits into
dotnet:mainfrom
tannergooding:fix-93513
May 16, 2024
Merged

Expose the FusedMultiplyAdd and MultiplyAddEstimate APIs on relevant vector and scalar types#102181
tannergooding merged 10 commits into
dotnet:mainfrom
tannergooding:fix-93513

Conversation

@tannergooding

Copy link
Copy Markdown
Member

This resolves#98053 and makes progress towards #93513

@ghost

Copy link
Copy Markdown

Note regarding the new-api-needs-documentation label:

This serves as a reminder for when your PR is modifying a ref *.cs file and adding/modifying public APIs, please make sure the API implementation in the src *.cs file is documented with triple slash comments, so the PR reviewers can sign off that change.

@tannergooding
tannergooding marked this pull request as draft May 14, 2024 18:31
@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @dotnet/jit-contrib for the JIT side.

This is one of the few remaining cases that should need JIT side work, most of the remaining APIs in #93513 will be implemented purely in managed code. The *Estimate APIs are special as they need to work in the face of R2R due to being intentionally non-deterministic across hardware and FusedMultiplyAdd represents a case that optimizes down to a single instruction and which is a core part of many already deeply complex algorithms, so the benefits of guaranteeing expansion make it good to hook up to the existing intrinsic support.

@EgorBo

Copy link
Copy Markdown
Member

CC. @dotnet/jit-contrib for the JIT side.

This is one of the few remaining cases that should need JIT side work, most of the remaining APIs in #93513 will be implemented purely in managed code. The *Estimate APIs are special as they need to work in the face of R2R due to being intentionally non-deterministic across hardware and FusedMultiplyAdd represents a case that optimizes down to a single instruction and which is a core part of many already deeply complex algorithms, so the benefits of guaranteeing expansion make it good to hook up to the existing intrinsic support.

Should we rather introduce a new attribute to prevent methods from being pre-jitted (and inlined into pre-jitted code)?

// AdvSimd.FusedMultiplyAdd expects (addend, left, right), while the APIs take (left, right, addend)
// We expect op1 and op2 to have already been spilled

std::swap(op1, op3);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am just curious - who's responsible to spill them?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, nvm, I see

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Typically it's the importer code, prior to calling this method. Latter phases that might call such an API (like morph) are responsible for ensuring the swap is safe in the face of potential side effects.

There's, unfortunately, not really a way for us to do such a swap safely in the API itself (at least that I know of) nor to know if the caller has actually done it. So the typical approach has been to do the swap and comment that callers should be doing the validation.

impSpillSideEffect(true, verCurrentState.esStackDepth -
3 DEBUGARG("Spilling op1 side effects for MultiplyAddEstimate"));

impSpillSideEffect(true, verCurrentState.esStackDepth -

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Isn't better/simpler to use impSpillSideEffects(true, CHECK_SPILL_ALL DEBUGARG("spilling side-effects")); ?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's not something we've been doing in other scenarios.

AFAIR it comes down to not spilling values on the stack that aren't impacted. For example, we take 3 args, but the stack could have 4+ on it (spilling these is unnecessary since they are still processed in order with respect to our own op1/op2/op3; otherwise everyone who pops the stack would need to consider the need to spill such additional entries) and we don't need to ever spill the stack top (op3).

So we're doing this to ensure only the minimum number of items that need to be spilled are spilled.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

Should we rather introduce a new attribute to prevent methods from being pre-jitted (and inlined into pre-jitted code)?

We technically have one already, it's BypassReadyToRun (but we don't define or use it anyways as it was a legacy attribute for Crossgen1). The PR where I fixed this for Math.ReciprocalSqrtEstimate it was suggested to do it this other way instead.

I'm also not sure the number of *Estimate APIs is at the point where it's beneficial to define and expose that attribute and there's various other quirks that may need to be considered/solved in that scenario was well, plus the general consideration that some of these APIs are known to be used in hot paths of complex algorithms where inlining can fail, even with the inlining boost that methods using intrinsics get.

-- I do want to get to a point, someday, where we can move a lot of complexity out of the JIT and into the BCL instead. Ideally just the platform specific intrinsics and core xplat APIs which were exposed instead of platform specific APIs (like Create) are handled in the JIT and things like Vector2/3/4 or Vector128.WidenLower and APIs like that are handled entirely in managed code instead. I just don't think we're at that point yet today and need more work in the inliner and other places to ensure that we can get the desired codegen.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

CC. @stephentoub for the libraries side.

The one thing to note is that given how TensorPrimitives defines MultiplyAddEstimate and uses it in other APIs, the API was moved down to INumberBase<T> rather than being put on IFloatingPointIeee754<T>. This is still sensible for other types, it just always does the naive thing and isn't exposed as a member that's accessible from non-generic contexts.

@tannergooding

Copy link
Copy Markdown
MemberAuthor

-- I do want to get to a point, someday, where we can move a lot of complexity out of the JIT and into the BCL instead. Ideally just the platform specific intrinsics and core xplat APIs which were exposed instead of platform specific APIs (like Create) are handled in the JIT and things like Vector2/3/4 or Vector128.WidenLower and APIs like that are handled entirely in managed code instead. I just don't think we're at that point yet today and need more work in the inliner and other places to ensure that we can get the desired codegen.

I've logged #102275 to track prototyping this work and seeing if we can identify some of the places that are still preventing it from happening.

This seems like a good thing to try and do around the time .NET 9 RC1 snaps and we have a bit of a small break before .NET 10 work really gets going.

@tannergooding
tannergooding merged commit b7727f5 into dotnet:mainMay 16, 2024
@tannergooding
tannergooding deleted the fix-93513 branch May 16, 2024 21:11
Ruihan-Yin pushed a commit to Ruihan-Yin/runtime that referenced this pull request May 30, 2024
…vector and scalar types (dotnet#102181)
* Expose FusedMultiplyAdd and MultiplyAddEstimate on the scalar and vector types
* Adding tests covering FusedMultiplyAdd and MultiplyAddEstimate for the vector types
* Ensure TensorPrimitives uses the xplat APIs on .NET 9+
* Apply formatting patch
* Fix an accidental change to GenericVectorTests
* Ensure Arm64 passes fma operands in the correct order
* Apply formatting patch
* Ensure all the Arm64 code paths are spilling and swapping operands
* Apply formatting patch
* Don't pop the stack value twice
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Jun 16, 2024
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[API Proposal]: MultiplyAddEstimate

2 participants

@tannergooding@EgorBo