Changes some of the CPU Math implemenation from our current version to use the new TensorPrimitives package. - #6875

Merged
michaelgsharp merged 27 commits into
dotnet:mainfrom
michaelgsharp:tensor-math2
Nov 15, 2023
Merged

Changes some of the CPU Math implemenation from our current version to use the new TensorPrimitives package.#6875
michaelgsharp merged 27 commits into
dotnet:mainfrom
michaelgsharp:tensor-math2

Conversation

@michaelgsharp

@michaelgsharpmichaelgsharp commented Nov 2, 2023

Copy link
Copy Markdown
Contributor

This changes some of the CPU Math implementation from our current version to use the new TensorPrimitives package.

Currently we are pointing to the rc2 version, but the following benchmarks have been done with a local copy of the GA version.

This also changes CPUMath to target .NET 8 instead of .NET 6. Did we want that for this version? Or should I change it back to 6 for this release? @ericstj@jeffhandley

The following is a summary of the methods in CPUMath, the old vs new benchmarks, and whether I updated it to use the new TensorPrimitives package. @tannergooding@stephentoub@jeffhandley@ericstj@luisquintanilla This is where we need to discuss. Is any performance hit worth taking? Or should anything that is slower be kept on the existing code?

NET 8

MethodarrayLengthMean - OriginalMean - New% FasterComments
AddScalarU51225.30 ns20.32 ns25%
Scale51219.91 ns19.29 ns3%
ScaleSrcU51227.58 ns20.74 ns33%
ScaleAddU51228.46 ns29.05 nsMethod Unchanged, composite function so slower with new code
AddScaleU51229.74 ns28.59 ns4%
AddScaleSU512345.92 ns327.68 ns6%Method Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
AddScaleCopyU51234.01 ns27.03 ns26%
AddU51229.80 ns26.71 ns12%
AddSU512325.32 ns349.46 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
MulElementWiseU51233.92 ns27.29 ns24%
Sum51236.57 ns34.34 ns6%
SumSqU51237.50 ns39.34 ns-5%
SumSqDiffU51241.23 ns43.38 nsMethod Unchanged, composite function so slower with new code
SumAbsU51243.74 ns39.27 ns11%
SumAbsDiffU51247.23 ns37.48 ns26%
MaxAbsU51242.30 ns43.26 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release
MaxAbsDiffU51246.94 ns47.73 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release. Is composite function.
DotU51250.34 ns43.20 ns17%
DotSU512212.19 ns213.18 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
Dist251255.48 ns47.43 ns17%

Framework

MethodarrayLengthMean - OriginalMean - New% FasterComments
AddScalarU25648.48 ns29.88 ns62%
Scale25643.45 ns28.55 ns52%
ScaleSrcU25649.87 ns38.13 ns31%
ScaleAddU25647.87 ns45.76 nsMethod Unchanged, composite function so slower with new code
AddScaleU25652.63 ns62.58 ns-16%Slightly slower in new code. Do we want to keep it?
AddScaleSU256151.00 ns152.77 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
AddScaleCopyU25648.35 ns63.94 ns-24%Slightly slower in new code. Do we want to keep it?
AddU25649.68 ns59.32 ns-16%Slightly slower in new code. Do we want to keep it?
AddSU256150.34 ns153.89 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
MulElementWiseU25648.26 ns69.89 ns-31%
Sum25668.05 ns59.74 ns14%
SumSqU25668.21 ns62.08 ns10%
SumSqDiffU25657.52 ns57.64 nsMethod Unchanged, composite function so slower with new code
SumAbsU25672.88 ns65.01 ns12%
SumAbsDiffU25659.51 ns68.23 ns-13%Slightly slower in new code. Do we want to keep it?
MaxAbsU25672.26 ns71.48 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release
MaxAbsDiffU25659.30 ns58.87 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release. Is composite function.
DotU25658.93 ns68.42 ns-14%Slightly slower in new code. Do we want to keep it?
DotSU256109.76 ns113.78 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
Dist225659.49 ns86.97 ns-32%Slightly slower in new code. Do we want to keep it?

I think that even if we don't want to keep the TensorPrimitives code in the cases where its slower, at least for .NET Framework we should add a check and if the native code doesn't exist to run these accelerated, we should fallback to the TensorPrimitives approach. That would have to be added in though.

All this was done with AVX256.

Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs Outdated
Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs Outdated
Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs
Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netstandard.cs Outdated
@codecov

codecovBot commented Nov 14, 2023

Copy link
Copy Markdown

Codecov Report

Merging #6875 (54e876a) into main (796cb35) will decrease coverage by 0.60%.
Report is 1 commits behind head on main.
The diff coverage is 100.00%.

Additional details and impacted files
@@ Coverage Diff @@## main #6875 +/- ##
==========================================
- Coverage 69.40% 68.80% -0.60% 
==========================================
Files 1238 1240 +2 Lines 249462 249392 -70 Branches 25522 25493 -29 ==========================================
- Hits 173139 171599 -1540 - Misses 69578 71196 +1618 + Partials 6745 6597 -148 
FlagCoverage Δ
Debug68.80% <100.00%> (-0.60%)⬇️
production63.26% <100.00%> (-0.67%)⬇️
test88.49% <ø> (-0.41%)⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

FilesCoverage Δ
src/Microsoft.ML.CpuMath/AvxIntrinsics.cs58.18% <ø> (-38.51%)⬇️
src/Microsoft.ML.CpuMath/CpuMathUtils.cs100.00% <100.00%> (ø)
...rc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs97.80% <100.00%> (-0.84%)⬇️
src/Microsoft.ML.CpuMath/SseIntrinsics.cs54.80% <ø> (-41.55%)⬇️

... and 48 files with indirect coverage changes

@michaelgsharp

Copy link
Copy Markdown
ContributorAuthor

/azp run

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

Comment threadsrc/Microsoft.ML.CpuMath/AvxIntrinsics.cs
@michaelgsharp
michaelgsharp merged commit d2cf997 into dotnet:mainNov 15, 2023
@michaelgsharp
michaelgsharp deleted the tensor-math2 branch November 15, 2023 05:46
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Dec 15, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@michaelgsharp@stephentoub@tannergooding@JakeRadMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Changes some of the CPU Math implemenation from our current version to use the new TensorPrimitives package. - #6875

Merged
michaelgsharp merged 27 commits into
dotnet:mainfrom
michaelgsharp:tensor-math2
Nov 15, 2023
Merged

Changes some of the CPU Math implemenation from our current version to use the new TensorPrimitives package.#6875
michaelgsharp merged 27 commits into
dotnet:mainfrom
michaelgsharp:tensor-math2

Conversation

@michaelgsharp

@michaelgsharpmichaelgsharp commented Nov 2, 2023

Copy link
Copy Markdown
Contributor

This changes some of the CPU Math implementation from our current version to use the new TensorPrimitives package.

Currently we are pointing to the rc2 version, but the following benchmarks have been done with a local copy of the GA version.

This also changes CPUMath to target .NET 8 instead of .NET 6. Did we want that for this version? Or should I change it back to 6 for this release? @ericstj@jeffhandley

The following is a summary of the methods in CPUMath, the old vs new benchmarks, and whether I updated it to use the new TensorPrimitives package. @tannergooding@stephentoub@jeffhandley@ericstj@luisquintanilla This is where we need to discuss. Is any performance hit worth taking? Or should anything that is slower be kept on the existing code?

NET 8

MethodarrayLengthMean - OriginalMean - New% FasterComments
AddScalarU51225.30 ns20.32 ns25%
Scale51219.91 ns19.29 ns3%
ScaleSrcU51227.58 ns20.74 ns33%
ScaleAddU51228.46 ns29.05 nsMethod Unchanged, composite function so slower with new code
AddScaleU51229.74 ns28.59 ns4%
AddScaleSU512345.92 ns327.68 ns6%Method Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
AddScaleCopyU51234.01 ns27.03 ns26%
AddU51229.80 ns26.71 ns12%
AddSU512325.32 ns349.46 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
MulElementWiseU51233.92 ns27.29 ns24%
Sum51236.57 ns34.34 ns6%
SumSqU51237.50 ns39.34 ns-5%
SumSqDiffU51241.23 ns43.38 nsMethod Unchanged, composite function so slower with new code
SumAbsU51243.74 ns39.27 ns11%
SumAbsDiffU51247.23 ns37.48 ns26%
MaxAbsU51242.30 ns43.26 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release
MaxAbsDiffU51246.94 ns47.73 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release. Is composite function.
DotU51250.34 ns43.20 ns17%
DotSU512212.19 ns213.18 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
Dist251255.48 ns47.43 ns17%

Framework

MethodarrayLengthMean - OriginalMean - New% FasterComments
AddScalarU25648.48 ns29.88 ns62%
Scale25643.45 ns28.55 ns52%
ScaleSrcU25649.87 ns38.13 ns31%
ScaleAddU25647.87 ns45.76 nsMethod Unchanged, composite function so slower with new code
AddScaleU25652.63 ns62.58 ns-16%Slightly slower in new code. Do we want to keep it?
AddScaleSU256151.00 ns152.77 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
AddScaleCopyU25648.35 ns63.94 ns-24%Slightly slower in new code. Do we want to keep it?
AddU25649.68 ns59.32 ns-16%Slightly slower in new code. Do we want to keep it?
AddSU256150.34 ns153.89 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
MulElementWiseU25648.26 ns69.89 ns-31%
Sum25668.05 ns59.74 ns14%
SumSqU25668.21 ns62.08 ns10%
SumSqDiffU25657.52 ns57.64 nsMethod Unchanged, composite function so slower with new code
SumAbsU25672.88 ns65.01 ns12%
SumAbsDiffU25659.51 ns68.23 ns-13%Slightly slower in new code. Do we want to keep it?
MaxAbsU25672.26 ns71.48 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release
MaxAbsDiffU25659.30 ns58.87 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release. Is composite function.
DotU25658.93 ns68.42 ns-14%Slightly slower in new code. Do we want to keep it?
DotSU256109.76 ns113.78 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
Dist225659.49 ns86.97 ns-32%Slightly slower in new code. Do we want to keep it?

I think that even if we don't want to keep the TensorPrimitives code in the cases where its slower, at least for .NET Framework we should add a check and if the native code doesn't exist to run these accelerated, we should fallback to the TensorPrimitives approach. That would have to be added in though.

All this was done with AVX256.

Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs Outdated
Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs Outdated
Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs
Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netstandard.cs Outdated
@codecov

codecovBot commented Nov 14, 2023

Copy link
Copy Markdown

Codecov Report

Merging #6875 (54e876a) into main (796cb35) will decrease coverage by 0.60%.
Report is 1 commits behind head on main.
The diff coverage is 100.00%.

Additional details and impacted files
@@ Coverage Diff @@## main #6875 +/- ##
==========================================
- Coverage 69.40% 68.80% -0.60% 
==========================================
Files 1238 1240 +2 Lines 249462 249392 -70 Branches 25522 25493 -29 ==========================================
- Hits 173139 171599 -1540 - Misses 69578 71196 +1618 + Partials 6745 6597 -148 
FlagCoverage Δ
Debug68.80% <100.00%> (-0.60%)⬇️
production63.26% <100.00%> (-0.67%)⬇️
test88.49% <ø> (-0.41%)⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

FilesCoverage Δ
src/Microsoft.ML.CpuMath/AvxIntrinsics.cs58.18% <ø> (-38.51%)⬇️
src/Microsoft.ML.CpuMath/CpuMathUtils.cs100.00% <100.00%> (ø)
...rc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs97.80% <100.00%> (-0.84%)⬇️
src/Microsoft.ML.CpuMath/SseIntrinsics.cs54.80% <ø> (-41.55%)⬇️

... and 48 files with indirect coverage changes

@michaelgsharp

Copy link
Copy Markdown
ContributorAuthor

/azp run

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

Comment threadsrc/Microsoft.ML.CpuMath/AvxIntrinsics.cs
@michaelgsharp
michaelgsharp merged commit d2cf997 into dotnet:mainNov 15, 2023
@michaelgsharp
michaelgsharp deleted the tensor-math2 branch November 15, 2023 05:46
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Dec 15, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@michaelgsharp@stephentoub@tannergooding@JakeRadMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Changes some of the CPU Math implemenation from our current version to use the new TensorPrimitives package. - #6875

Merged
michaelgsharp merged 27 commits into
dotnet:mainfrom
michaelgsharp:tensor-math2
Nov 15, 2023
Merged

Changes some of the CPU Math implemenation from our current version to use the new TensorPrimitives package.#6875
michaelgsharp merged 27 commits into
dotnet:mainfrom
michaelgsharp:tensor-math2

Conversation

@michaelgsharp

@michaelgsharpmichaelgsharp commented Nov 2, 2023

Copy link
Copy Markdown
Contributor

This changes some of the CPU Math implementation from our current version to use the new TensorPrimitives package.

Currently we are pointing to the rc2 version, but the following benchmarks have been done with a local copy of the GA version.

This also changes CPUMath to target .NET 8 instead of .NET 6. Did we want that for this version? Or should I change it back to 6 for this release? @ericstj@jeffhandley

The following is a summary of the methods in CPUMath, the old vs new benchmarks, and whether I updated it to use the new TensorPrimitives package. @tannergooding@stephentoub@jeffhandley@ericstj@luisquintanilla This is where we need to discuss. Is any performance hit worth taking? Or should anything that is slower be kept on the existing code?

NET 8

MethodarrayLengthMean - OriginalMean - New% FasterComments
AddScalarU51225.30 ns20.32 ns25%
Scale51219.91 ns19.29 ns3%
ScaleSrcU51227.58 ns20.74 ns33%
ScaleAddU51228.46 ns29.05 nsMethod Unchanged, composite function so slower with new code
AddScaleU51229.74 ns28.59 ns4%
AddScaleSU512345.92 ns327.68 ns6%Method Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
AddScaleCopyU51234.01 ns27.03 ns26%
AddU51229.80 ns26.71 ns12%
AddSU512325.32 ns349.46 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
MulElementWiseU51233.92 ns27.29 ns24%
Sum51236.57 ns34.34 ns6%
SumSqU51237.50 ns39.34 ns-5%
SumSqDiffU51241.23 ns43.38 nsMethod Unchanged, composite function so slower with new code
SumAbsU51243.74 ns39.27 ns11%
SumAbsDiffU51247.23 ns37.48 ns26%
MaxAbsU51242.30 ns43.26 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release
MaxAbsDiffU51246.94 ns47.73 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release. Is composite function.
DotU51250.34 ns43.20 ns17%
DotSU512212.19 ns213.18 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
Dist251255.48 ns47.43 ns17%

Framework

MethodarrayLengthMean - OriginalMean - New% FasterComments
AddScalarU25648.48 ns29.88 ns62%
Scale25643.45 ns28.55 ns52%
ScaleSrcU25649.87 ns38.13 ns31%
ScaleAddU25647.87 ns45.76 nsMethod Unchanged, composite function so slower with new code
AddScaleU25652.63 ns62.58 ns-16%Slightly slower in new code. Do we want to keep it?
AddScaleSU256151.00 ns152.77 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
AddScaleCopyU25648.35 ns63.94 ns-24%Slightly slower in new code. Do we want to keep it?
AddU25649.68 ns59.32 ns-16%Slightly slower in new code. Do we want to keep it?
AddSU256150.34 ns153.89 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
MulElementWiseU25648.26 ns69.89 ns-31%
Sum25668.05 ns59.74 ns14%
SumSqU25668.21 ns62.08 ns10%
SumSqDiffU25657.52 ns57.64 nsMethod Unchanged, composite function so slower with new code
SumAbsU25672.88 ns65.01 ns12%
SumAbsDiffU25659.51 ns68.23 ns-13%Slightly slower in new code. Do we want to keep it?
MaxAbsU25672.26 ns71.48 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release
MaxAbsDiffU25659.30 ns58.87 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release. Is composite function.
DotU25658.93 ns68.42 ns-14%Slightly slower in new code. Do we want to keep it?
DotSU256109.76 ns113.78 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
Dist225659.49 ns86.97 ns-32%Slightly slower in new code. Do we want to keep it?

I think that even if we don't want to keep the TensorPrimitives code in the cases where its slower, at least for .NET Framework we should add a check and if the native code doesn't exist to run these accelerated, we should fallback to the TensorPrimitives approach. That would have to be added in though.

All this was done with AVX256.

Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs Outdated
Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs Outdated
Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs
Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netstandard.cs Outdated
@codecov

codecovBot commented Nov 14, 2023

Copy link
Copy Markdown

Codecov Report

Merging #6875 (54e876a) into main (796cb35) will decrease coverage by 0.60%.
Report is 1 commits behind head on main.
The diff coverage is 100.00%.

Additional details and impacted files
@@ Coverage Diff @@## main #6875 +/- ##
==========================================
- Coverage 69.40% 68.80% -0.60% 
==========================================
Files 1238 1240 +2 Lines 249462 249392 -70 Branches 25522 25493 -29 ==========================================
- Hits 173139 171599 -1540 - Misses 69578 71196 +1618 + Partials 6745 6597 -148 
FlagCoverage Δ
Debug68.80% <100.00%> (-0.60%)⬇️
production63.26% <100.00%> (-0.67%)⬇️
test88.49% <ø> (-0.41%)⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

FilesCoverage Δ
src/Microsoft.ML.CpuMath/AvxIntrinsics.cs58.18% <ø> (-38.51%)⬇️
src/Microsoft.ML.CpuMath/CpuMathUtils.cs100.00% <100.00%> (ø)
...rc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs97.80% <100.00%> (-0.84%)⬇️
src/Microsoft.ML.CpuMath/SseIntrinsics.cs54.80% <ø> (-41.55%)⬇️

... and 48 files with indirect coverage changes

@michaelgsharp

Copy link
Copy Markdown
ContributorAuthor

/azp run

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

Comment threadsrc/Microsoft.ML.CpuMath/AvxIntrinsics.cs
@michaelgsharp
michaelgsharp merged commit d2cf997 into dotnet:mainNov 15, 2023
@michaelgsharp
michaelgsharp deleted the tensor-math2 branch November 15, 2023 05:46
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Dec 15, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@michaelgsharp@stephentoub@tannergooding@JakeRadMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Changes some of the CPU Math implemenation from our current version to use the new TensorPrimitives package. - #6875

Merged
michaelgsharp merged 27 commits into
dotnet:mainfrom
michaelgsharp:tensor-math2
Nov 15, 2023
Merged

Changes some of the CPU Math implemenation from our current version to use the new TensorPrimitives package.#6875
michaelgsharp merged 27 commits into
dotnet:mainfrom
michaelgsharp:tensor-math2

Conversation

@michaelgsharp

@michaelgsharpmichaelgsharp commented Nov 2, 2023

Copy link
Copy Markdown
Contributor

This changes some of the CPU Math implementation from our current version to use the new TensorPrimitives package.

Currently we are pointing to the rc2 version, but the following benchmarks have been done with a local copy of the GA version.

This also changes CPUMath to target .NET 8 instead of .NET 6. Did we want that for this version? Or should I change it back to 6 for this release? @ericstj@jeffhandley

The following is a summary of the methods in CPUMath, the old vs new benchmarks, and whether I updated it to use the new TensorPrimitives package. @tannergooding@stephentoub@jeffhandley@ericstj@luisquintanilla This is where we need to discuss. Is any performance hit worth taking? Or should anything that is slower be kept on the existing code?

NET 8

MethodarrayLengthMean - OriginalMean - New% FasterComments
AddScalarU51225.30 ns20.32 ns25%
Scale51219.91 ns19.29 ns3%
ScaleSrcU51227.58 ns20.74 ns33%
ScaleAddU51228.46 ns29.05 nsMethod Unchanged, composite function so slower with new code
AddScaleU51229.74 ns28.59 ns4%
AddScaleSU512345.92 ns327.68 ns6%Method Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
AddScaleCopyU51234.01 ns27.03 ns26%
AddU51229.80 ns26.71 ns12%
AddSU512325.32 ns349.46 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
MulElementWiseU51233.92 ns27.29 ns24%
Sum51236.57 ns34.34 ns6%
SumSqU51237.50 ns39.34 ns-5%
SumSqDiffU51241.23 ns43.38 nsMethod Unchanged, composite function so slower with new code
SumAbsU51243.74 ns39.27 ns11%
SumAbsDiffU51247.23 ns37.48 ns26%
MaxAbsU51242.30 ns43.26 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release
MaxAbsDiffU51246.94 ns47.73 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release. Is composite function.
DotU51250.34 ns43.20 ns17%
DotSU512212.19 ns213.18 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
Dist251255.48 ns47.43 ns17%

Framework

MethodarrayLengthMean - OriginalMean - New% FasterComments
AddScalarU25648.48 ns29.88 ns62%
Scale25643.45 ns28.55 ns52%
ScaleSrcU25649.87 ns38.13 ns31%
ScaleAddU25647.87 ns45.76 nsMethod Unchanged, composite function so slower with new code
AddScaleU25652.63 ns62.58 ns-16%Slightly slower in new code. Do we want to keep it?
AddScaleSU256151.00 ns152.77 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
AddScaleCopyU25648.35 ns63.94 ns-24%Slightly slower in new code. Do we want to keep it?
AddU25649.68 ns59.32 ns-16%Slightly slower in new code. Do we want to keep it?
AddSU256150.34 ns153.89 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
MulElementWiseU25648.26 ns69.89 ns-31%
Sum25668.05 ns59.74 ns14%
SumSqU25668.21 ns62.08 ns10%
SumSqDiffU25657.52 ns57.64 nsMethod Unchanged, composite function so slower with new code
SumAbsU25672.88 ns65.01 ns12%
SumAbsDiffU25659.51 ns68.23 ns-13%Slightly slower in new code. Do we want to keep it?
MaxAbsU25672.26 ns71.48 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release
MaxAbsDiffU25659.30 ns58.87 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release. Is composite function.
DotU25658.93 ns68.42 ns-14%Slightly slower in new code. Do we want to keep it?
DotSU256109.76 ns113.78 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
Dist225659.49 ns86.97 ns-32%Slightly slower in new code. Do we want to keep it?

I think that even if we don't want to keep the TensorPrimitives code in the cases where its slower, at least for .NET Framework we should add a check and if the native code doesn't exist to run these accelerated, we should fallback to the TensorPrimitives approach. That would have to be added in though.

All this was done with AVX256.

Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs Outdated
Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs Outdated
Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs
Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netstandard.cs Outdated
@codecov

codecovBot commented Nov 14, 2023

Copy link
Copy Markdown

Codecov Report

Merging #6875 (54e876a) into main (796cb35) will decrease coverage by 0.60%.
Report is 1 commits behind head on main.
The diff coverage is 100.00%.

Additional details and impacted files
@@ Coverage Diff @@## main #6875 +/- ##
==========================================
- Coverage 69.40% 68.80% -0.60% 
==========================================
Files 1238 1240 +2 Lines 249462 249392 -70 Branches 25522 25493 -29 ==========================================
- Hits 173139 171599 -1540 - Misses 69578 71196 +1618 + Partials 6745 6597 -148 
FlagCoverage Δ
Debug68.80% <100.00%> (-0.60%)⬇️
production63.26% <100.00%> (-0.67%)⬇️
test88.49% <ø> (-0.41%)⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

FilesCoverage Δ
src/Microsoft.ML.CpuMath/AvxIntrinsics.cs58.18% <ø> (-38.51%)⬇️
src/Microsoft.ML.CpuMath/CpuMathUtils.cs100.00% <100.00%> (ø)
...rc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs97.80% <100.00%> (-0.84%)⬇️
src/Microsoft.ML.CpuMath/SseIntrinsics.cs54.80% <ø> (-41.55%)⬇️

... and 48 files with indirect coverage changes

@michaelgsharp

Copy link
Copy Markdown
ContributorAuthor

/azp run

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

Comment threadsrc/Microsoft.ML.CpuMath/AvxIntrinsics.cs
@michaelgsharp
michaelgsharp merged commit d2cf997 into dotnet:mainNov 15, 2023
@michaelgsharp
michaelgsharp deleted the tensor-math2 branch November 15, 2023 05:46
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Dec 15, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@michaelgsharp@stephentoub@tannergooding@JakeRadMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Changes some of the CPU Math implemenation from our current version to use the new TensorPrimitives package. - #6875

Merged
michaelgsharp merged 27 commits into
dotnet:mainfrom
michaelgsharp:tensor-math2
Nov 15, 2023
Merged

Changes some of the CPU Math implemenation from our current version to use the new TensorPrimitives package.#6875
michaelgsharp merged 27 commits into
dotnet:mainfrom
michaelgsharp:tensor-math2

Conversation

@michaelgsharp

@michaelgsharpmichaelgsharp commented Nov 2, 2023

Copy link
Copy Markdown
Contributor

This changes some of the CPU Math implementation from our current version to use the new TensorPrimitives package.

Currently we are pointing to the rc2 version, but the following benchmarks have been done with a local copy of the GA version.

This also changes CPUMath to target .NET 8 instead of .NET 6. Did we want that for this version? Or should I change it back to 6 for this release? @ericstj@jeffhandley

The following is a summary of the methods in CPUMath, the old vs new benchmarks, and whether I updated it to use the new TensorPrimitives package. @tannergooding@stephentoub@jeffhandley@ericstj@luisquintanilla This is where we need to discuss. Is any performance hit worth taking? Or should anything that is slower be kept on the existing code?

NET 8

MethodarrayLengthMean - OriginalMean - New% FasterComments
AddScalarU51225.30 ns20.32 ns25%
Scale51219.91 ns19.29 ns3%
ScaleSrcU51227.58 ns20.74 ns33%
ScaleAddU51228.46 ns29.05 nsMethod Unchanged, composite function so slower with new code
AddScaleU51229.74 ns28.59 ns4%
AddScaleSU512345.92 ns327.68 ns6%Method Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
AddScaleCopyU51234.01 ns27.03 ns26%
AddU51229.80 ns26.71 ns12%
AddSU512325.32 ns349.46 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
MulElementWiseU51233.92 ns27.29 ns24%
Sum51236.57 ns34.34 ns6%
SumSqU51237.50 ns39.34 ns-5%
SumSqDiffU51241.23 ns43.38 nsMethod Unchanged, composite function so slower with new code
SumAbsU51243.74 ns39.27 ns11%
SumAbsDiffU51247.23 ns37.48 ns26%
MaxAbsU51242.30 ns43.26 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release
MaxAbsDiffU51246.94 ns47.73 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release. Is composite function.
DotU51250.34 ns43.20 ns17%
DotSU512212.19 ns213.18 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
Dist251255.48 ns47.43 ns17%

Framework

MethodarrayLengthMean - OriginalMean - New% FasterComments
AddScalarU25648.48 ns29.88 ns62%
Scale25643.45 ns28.55 ns52%
ScaleSrcU25649.87 ns38.13 ns31%
ScaleAddU25647.87 ns45.76 nsMethod Unchanged, composite function so slower with new code
AddScaleU25652.63 ns62.58 ns-16%Slightly slower in new code. Do we want to keep it?
AddScaleSU256151.00 ns152.77 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
AddScaleCopyU25648.35 ns63.94 ns-24%Slightly slower in new code. Do we want to keep it?
AddU25649.68 ns59.32 ns-16%Slightly slower in new code. Do we want to keep it?
AddSU256150.34 ns153.89 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
MulElementWiseU25648.26 ns69.89 ns-31%
Sum25668.05 ns59.74 ns14%
SumSqU25668.21 ns62.08 ns10%
SumSqDiffU25657.52 ns57.64 nsMethod Unchanged, composite function so slower with new code
SumAbsU25672.88 ns65.01 ns12%
SumAbsDiffU25659.51 ns68.23 ns-13%Slightly slower in new code. Do we want to keep it?
MaxAbsU25672.26 ns71.48 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release
MaxAbsDiffU25659.30 ns58.87 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release. Is composite function.
DotU25658.93 ns68.42 ns-14%Slightly slower in new code. Do we want to keep it?
DotSU256109.76 ns113.78 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
Dist225659.49 ns86.97 ns-32%Slightly slower in new code. Do we want to keep it?

I think that even if we don't want to keep the TensorPrimitives code in the cases where its slower, at least for .NET Framework we should add a check and if the native code doesn't exist to run these accelerated, we should fallback to the TensorPrimitives approach. That would have to be added in though.

All this was done with AVX256.

Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs Outdated
Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs Outdated
Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs
Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netstandard.cs Outdated
@codecov

codecovBot commented Nov 14, 2023

Copy link
Copy Markdown

Codecov Report

Merging #6875 (54e876a) into main (796cb35) will decrease coverage by 0.60%.
Report is 1 commits behind head on main.
The diff coverage is 100.00%.

Additional details and impacted files
@@ Coverage Diff @@## main #6875 +/- ##
==========================================
- Coverage 69.40% 68.80% -0.60% 
==========================================
Files 1238 1240 +2 Lines 249462 249392 -70 Branches 25522 25493 -29 ==========================================
- Hits 173139 171599 -1540 - Misses 69578 71196 +1618 + Partials 6745 6597 -148 
FlagCoverage Δ
Debug68.80% <100.00%> (-0.60%)⬇️
production63.26% <100.00%> (-0.67%)⬇️
test88.49% <ø> (-0.41%)⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

FilesCoverage Δ
src/Microsoft.ML.CpuMath/AvxIntrinsics.cs58.18% <ø> (-38.51%)⬇️
src/Microsoft.ML.CpuMath/CpuMathUtils.cs100.00% <100.00%> (ø)
...rc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs97.80% <100.00%> (-0.84%)⬇️
src/Microsoft.ML.CpuMath/SseIntrinsics.cs54.80% <ø> (-41.55%)⬇️

... and 48 files with indirect coverage changes

@michaelgsharp

Copy link
Copy Markdown
ContributorAuthor

/azp run

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

Comment threadsrc/Microsoft.ML.CpuMath/AvxIntrinsics.cs
@michaelgsharp
michaelgsharp merged commit d2cf997 into dotnet:mainNov 15, 2023
@michaelgsharp
michaelgsharp deleted the tensor-math2 branch November 15, 2023 05:46
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Dec 15, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@michaelgsharp@stephentoub@tannergooding@JakeRadMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Changes some of the CPU Math implemenation from our current version to use the new TensorPrimitives package. - #6875

Merged
michaelgsharp merged 27 commits into
dotnet:mainfrom
michaelgsharp:tensor-math2
Nov 15, 2023
Merged

Changes some of the CPU Math implemenation from our current version to use the new TensorPrimitives package.#6875
michaelgsharp merged 27 commits into
dotnet:mainfrom
michaelgsharp:tensor-math2

Conversation

@michaelgsharp

@michaelgsharpmichaelgsharp commented Nov 2, 2023

Copy link
Copy Markdown
Contributor

This changes some of the CPU Math implementation from our current version to use the new TensorPrimitives package.

Currently we are pointing to the rc2 version, but the following benchmarks have been done with a local copy of the GA version.

This also changes CPUMath to target .NET 8 instead of .NET 6. Did we want that for this version? Or should I change it back to 6 for this release? @ericstj@jeffhandley

The following is a summary of the methods in CPUMath, the old vs new benchmarks, and whether I updated it to use the new TensorPrimitives package. @tannergooding@stephentoub@jeffhandley@ericstj@luisquintanilla This is where we need to discuss. Is any performance hit worth taking? Or should anything that is slower be kept on the existing code?

NET 8

MethodarrayLengthMean - OriginalMean - New% FasterComments
AddScalarU51225.30 ns20.32 ns25%
Scale51219.91 ns19.29 ns3%
ScaleSrcU51227.58 ns20.74 ns33%
ScaleAddU51228.46 ns29.05 nsMethod Unchanged, composite function so slower with new code
AddScaleU51229.74 ns28.59 ns4%
AddScaleSU512345.92 ns327.68 ns6%Method Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
AddScaleCopyU51234.01 ns27.03 ns26%
AddU51229.80 ns26.71 ns12%
AddSU512325.32 ns349.46 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
MulElementWiseU51233.92 ns27.29 ns24%
Sum51236.57 ns34.34 ns6%
SumSqU51237.50 ns39.34 ns-5%
SumSqDiffU51241.23 ns43.38 nsMethod Unchanged, composite function so slower with new code
SumAbsU51243.74 ns39.27 ns11%
SumAbsDiffU51247.23 ns37.48 ns26%
MaxAbsU51242.30 ns43.26 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release
MaxAbsDiffU51246.94 ns47.73 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release. Is composite function.
DotU51250.34 ns43.20 ns17%
DotSU512212.19 ns213.18 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
Dist251255.48 ns47.43 ns17%

Framework

MethodarrayLengthMean - OriginalMean - New% FasterComments
AddScalarU25648.48 ns29.88 ns62%
Scale25643.45 ns28.55 ns52%
ScaleSrcU25649.87 ns38.13 ns31%
ScaleAddU25647.87 ns45.76 nsMethod Unchanged, composite function so slower with new code
AddScaleU25652.63 ns62.58 ns-16%Slightly slower in new code. Do we want to keep it?
AddScaleSU256151.00 ns152.77 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
AddScaleCopyU25648.35 ns63.94 ns-24%Slightly slower in new code. Do we want to keep it?
AddU25649.68 ns59.32 ns-16%Slightly slower in new code. Do we want to keep it?
AddSU256150.34 ns153.89 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
MulElementWiseU25648.26 ns69.89 ns-31%
Sum25668.05 ns59.74 ns14%
SumSqU25668.21 ns62.08 ns10%
SumSqDiffU25657.52 ns57.64 nsMethod Unchanged, composite function so slower with new code
SumAbsU25672.88 ns65.01 ns12%
SumAbsDiffU25659.51 ns68.23 ns-13%Slightly slower in new code. Do we want to keep it?
MaxAbsU25672.26 ns71.48 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release
MaxAbsDiffU25659.30 ns58.87 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release. Is composite function.
DotU25658.93 ns68.42 ns-14%Slightly slower in new code. Do we want to keep it?
DotSU256109.76 ns113.78 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
Dist225659.49 ns86.97 ns-32%Slightly slower in new code. Do we want to keep it?

I think that even if we don't want to keep the TensorPrimitives code in the cases where its slower, at least for .NET Framework we should add a check and if the native code doesn't exist to run these accelerated, we should fallback to the TensorPrimitives approach. That would have to be added in though.

All this was done with AVX256.

Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs Outdated
Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs Outdated
Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs
Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netstandard.cs Outdated
@codecov

codecovBot commented Nov 14, 2023

Copy link
Copy Markdown

Codecov Report

Merging #6875 (54e876a) into main (796cb35) will decrease coverage by 0.60%.
Report is 1 commits behind head on main.
The diff coverage is 100.00%.

Additional details and impacted files
@@ Coverage Diff @@## main #6875 +/- ##
==========================================
- Coverage 69.40% 68.80% -0.60% 
==========================================
Files 1238 1240 +2 Lines 249462 249392 -70 Branches 25522 25493 -29 ==========================================
- Hits 173139 171599 -1540 - Misses 69578 71196 +1618 + Partials 6745 6597 -148 
FlagCoverage Δ
Debug68.80% <100.00%> (-0.60%)⬇️
production63.26% <100.00%> (-0.67%)⬇️
test88.49% <ø> (-0.41%)⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

FilesCoverage Δ
src/Microsoft.ML.CpuMath/AvxIntrinsics.cs58.18% <ø> (-38.51%)⬇️
src/Microsoft.ML.CpuMath/CpuMathUtils.cs100.00% <100.00%> (ø)
...rc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs97.80% <100.00%> (-0.84%)⬇️
src/Microsoft.ML.CpuMath/SseIntrinsics.cs54.80% <ø> (-41.55%)⬇️

... and 48 files with indirect coverage changes

@michaelgsharp

Copy link
Copy Markdown
ContributorAuthor

/azp run

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

Comment threadsrc/Microsoft.ML.CpuMath/AvxIntrinsics.cs
@michaelgsharp
michaelgsharp merged commit d2cf997 into dotnet:mainNov 15, 2023
@michaelgsharp
michaelgsharp deleted the tensor-math2 branch November 15, 2023 05:46
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Dec 15, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@michaelgsharp@stephentoub@tannergooding@JakeRadMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Changes some of the CPU Math implemenation from our current version to use the new TensorPrimitives package. - #6875

Merged
michaelgsharp merged 27 commits into
dotnet:mainfrom
michaelgsharp:tensor-math2
Nov 15, 2023
Merged

Changes some of the CPU Math implemenation from our current version to use the new TensorPrimitives package.#6875
michaelgsharp merged 27 commits into
dotnet:mainfrom
michaelgsharp:tensor-math2

Conversation

@michaelgsharp

@michaelgsharpmichaelgsharp commented Nov 2, 2023

Copy link
Copy Markdown
Contributor

This changes some of the CPU Math implementation from our current version to use the new TensorPrimitives package.

Currently we are pointing to the rc2 version, but the following benchmarks have been done with a local copy of the GA version.

This also changes CPUMath to target .NET 8 instead of .NET 6. Did we want that for this version? Or should I change it back to 6 for this release? @ericstj@jeffhandley

The following is a summary of the methods in CPUMath, the old vs new benchmarks, and whether I updated it to use the new TensorPrimitives package. @tannergooding@stephentoub@jeffhandley@ericstj@luisquintanilla This is where we need to discuss. Is any performance hit worth taking? Or should anything that is slower be kept on the existing code?

NET 8

MethodarrayLengthMean - OriginalMean - New% FasterComments
AddScalarU51225.30 ns20.32 ns25%
Scale51219.91 ns19.29 ns3%
ScaleSrcU51227.58 ns20.74 ns33%
ScaleAddU51228.46 ns29.05 nsMethod Unchanged, composite function so slower with new code
AddScaleU51229.74 ns28.59 ns4%
AddScaleSU512345.92 ns327.68 ns6%Method Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
AddScaleCopyU51234.01 ns27.03 ns26%
AddU51229.80 ns26.71 ns12%
AddSU512325.32 ns349.46 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
MulElementWiseU51233.92 ns27.29 ns24%
Sum51236.57 ns34.34 ns6%
SumSqU51237.50 ns39.34 ns-5%
SumSqDiffU51241.23 ns43.38 nsMethod Unchanged, composite function so slower with new code
SumAbsU51243.74 ns39.27 ns11%
SumAbsDiffU51247.23 ns37.48 ns26%
MaxAbsU51242.30 ns43.26 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release
MaxAbsDiffU51246.94 ns47.73 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release. Is composite function.
DotU51250.34 ns43.20 ns17%
DotSU512212.19 ns213.18 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
Dist251255.48 ns47.43 ns17%

Framework

MethodarrayLengthMean - OriginalMean - New% FasterComments
AddScalarU25648.48 ns29.88 ns62%
Scale25643.45 ns28.55 ns52%
ScaleSrcU25649.87 ns38.13 ns31%
ScaleAddU25647.87 ns45.76 nsMethod Unchanged, composite function so slower with new code
AddScaleU25652.63 ns62.58 ns-16%Slightly slower in new code. Do we want to keep it?
AddScaleSU256151.00 ns152.77 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
AddScaleCopyU25648.35 ns63.94 ns-24%Slightly slower in new code. Do we want to keep it?
AddU25649.68 ns59.32 ns-16%Slightly slower in new code. Do we want to keep it?
AddSU256150.34 ns153.89 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
MulElementWiseU25648.26 ns69.89 ns-31%
Sum25668.05 ns59.74 ns14%
SumSqU25668.21 ns62.08 ns10%
SumSqDiffU25657.52 ns57.64 nsMethod Unchanged, composite function so slower with new code
SumAbsU25672.88 ns65.01 ns12%
SumAbsDiffU25659.51 ns68.23 ns-13%Slightly slower in new code. Do we want to keep it?
MaxAbsU25672.26 ns71.48 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release
MaxAbsDiffU25659.30 ns58.87 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release. Is composite function.
DotU25658.93 ns68.42 ns-14%Slightly slower in new code. Do we want to keep it?
DotSU256109.76 ns113.78 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
Dist225659.49 ns86.97 ns-32%Slightly slower in new code. Do we want to keep it?

I think that even if we don't want to keep the TensorPrimitives code in the cases where its slower, at least for .NET Framework we should add a check and if the native code doesn't exist to run these accelerated, we should fallback to the TensorPrimitives approach. That would have to be added in though.

All this was done with AVX256.

Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs Outdated
Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs Outdated
Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs
Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netstandard.cs Outdated
@codecov

codecovBot commented Nov 14, 2023

Copy link
Copy Markdown

Codecov Report

Merging #6875 (54e876a) into main (796cb35) will decrease coverage by 0.60%.
Report is 1 commits behind head on main.
The diff coverage is 100.00%.

Additional details and impacted files
@@ Coverage Diff @@## main #6875 +/- ##
==========================================
- Coverage 69.40% 68.80% -0.60% 
==========================================
Files 1238 1240 +2 Lines 249462 249392 -70 Branches 25522 25493 -29 ==========================================
- Hits 173139 171599 -1540 - Misses 69578 71196 +1618 + Partials 6745 6597 -148 
FlagCoverage Δ
Debug68.80% <100.00%> (-0.60%)⬇️
production63.26% <100.00%> (-0.67%)⬇️
test88.49% <ø> (-0.41%)⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

FilesCoverage Δ
src/Microsoft.ML.CpuMath/AvxIntrinsics.cs58.18% <ø> (-38.51%)⬇️
src/Microsoft.ML.CpuMath/CpuMathUtils.cs100.00% <100.00%> (ø)
...rc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs97.80% <100.00%> (-0.84%)⬇️
src/Microsoft.ML.CpuMath/SseIntrinsics.cs54.80% <ø> (-41.55%)⬇️

... and 48 files with indirect coverage changes

@michaelgsharp

Copy link
Copy Markdown
ContributorAuthor

/azp run

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

Comment threadsrc/Microsoft.ML.CpuMath/AvxIntrinsics.cs
@michaelgsharp
michaelgsharp merged commit d2cf997 into dotnet:mainNov 15, 2023
@michaelgsharp
michaelgsharp deleted the tensor-math2 branch November 15, 2023 05:46
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Dec 15, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@michaelgsharp@stephentoub@tannergooding@JakeRadMSFT
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Changes some of the CPU Math implemenation from our current version to use the new TensorPrimitives package. - #6875

Merged
michaelgsharp merged 27 commits into
dotnet:mainfrom
michaelgsharp:tensor-math2
Nov 15, 2023
Merged

Changes some of the CPU Math implemenation from our current version to use the new TensorPrimitives package.#6875
michaelgsharp merged 27 commits into
dotnet:mainfrom
michaelgsharp:tensor-math2

Conversation

@michaelgsharp

@michaelgsharpmichaelgsharp commented Nov 2, 2023

Copy link
Copy Markdown
Contributor

This changes some of the CPU Math implementation from our current version to use the new TensorPrimitives package.

Currently we are pointing to the rc2 version, but the following benchmarks have been done with a local copy of the GA version.

This also changes CPUMath to target .NET 8 instead of .NET 6. Did we want that for this version? Or should I change it back to 6 for this release? @ericstj@jeffhandley

The following is a summary of the methods in CPUMath, the old vs new benchmarks, and whether I updated it to use the new TensorPrimitives package. @tannergooding@stephentoub@jeffhandley@ericstj@luisquintanilla This is where we need to discuss. Is any performance hit worth taking? Or should anything that is slower be kept on the existing code?

NET 8

MethodarrayLengthMean - OriginalMean - New% FasterComments
AddScalarU51225.30 ns20.32 ns25%
Scale51219.91 ns19.29 ns3%
ScaleSrcU51227.58 ns20.74 ns33%
ScaleAddU51228.46 ns29.05 nsMethod Unchanged, composite function so slower with new code
AddScaleU51229.74 ns28.59 ns4%
AddScaleSU512345.92 ns327.68 ns6%Method Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
AddScaleCopyU51234.01 ns27.03 ns26%
AddU51229.80 ns26.71 ns12%
AddSU512325.32 ns349.46 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
MulElementWiseU51233.92 ns27.29 ns24%
Sum51236.57 ns34.34 ns6%
SumSqU51237.50 ns39.34 ns-5%
SumSqDiffU51241.23 ns43.38 nsMethod Unchanged, composite function so slower with new code
SumAbsU51243.74 ns39.27 ns11%
SumAbsDiffU51247.23 ns37.48 ns26%
MaxAbsU51242.30 ns43.26 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release
MaxAbsDiffU51246.94 ns47.73 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release. Is composite function.
DotU51250.34 ns43.20 ns17%
DotSU512212.19 ns213.18 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
Dist251255.48 ns47.43 ns17%

Framework

MethodarrayLengthMean - OriginalMean - New% FasterComments
AddScalarU25648.48 ns29.88 ns62%
Scale25643.45 ns28.55 ns52%
ScaleSrcU25649.87 ns38.13 ns31%
ScaleAddU25647.87 ns45.76 nsMethod Unchanged, composite function so slower with new code
AddScaleU25652.63 ns62.58 ns-16%Slightly slower in new code. Do we want to keep it?
AddScaleSU256151.00 ns152.77 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
AddScaleCopyU25648.35 ns63.94 ns-24%Slightly slower in new code. Do we want to keep it?
AddU25649.68 ns59.32 ns-16%Slightly slower in new code. Do we want to keep it?
AddSU256150.34 ns153.89 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
MulElementWiseU25648.26 ns69.89 ns-31%
Sum25668.05 ns59.74 ns14%
SumSqU25668.21 ns62.08 ns10%
SumSqDiffU25657.52 ns57.64 nsMethod Unchanged, composite function so slower with new code
SumAbsU25672.88 ns65.01 ns12%
SumAbsDiffU25659.51 ns68.23 ns-13%Slightly slower in new code. Do we want to keep it?
MaxAbsU25672.26 ns71.48 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release
MaxAbsDiffU25659.30 ns58.87 nsMethod Unchanged, in GA MaxMagnitude is slow, has been fixed for next release. Is composite function.
DotU25658.93 ns68.42 ns-14%Slightly slower in new code. Do we want to keep it?
DotSU256109.76 ns113.78 nsMethod Unchanged, dont have Sparse in Tensor Primitives. Can simulate but is slower.
Dist225659.49 ns86.97 ns-32%Slightly slower in new code. Do we want to keep it?

I think that even if we don't want to keep the TensorPrimitives code in the cases where its slower, at least for .NET Framework we should add a check and if the native code doesn't exist to run these accelerated, we should fallback to the TensorPrimitives approach. That would have to be added in though.

All this was done with AVX256.

Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs Outdated
Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs Outdated
Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs
Comment threadsrc/Microsoft.ML.CpuMath/CpuMathUtils.netstandard.cs Outdated
@codecov

codecovBot commented Nov 14, 2023

Copy link
Copy Markdown

Codecov Report

Merging #6875 (54e876a) into main (796cb35) will decrease coverage by 0.60%.
Report is 1 commits behind head on main.
The diff coverage is 100.00%.

Additional details and impacted files
@@ Coverage Diff @@## main #6875 +/- ##
==========================================
- Coverage 69.40% 68.80% -0.60% 
==========================================
Files 1238 1240 +2 Lines 249462 249392 -70 Branches 25522 25493 -29 ==========================================
- Hits 173139 171599 -1540 - Misses 69578 71196 +1618 + Partials 6745 6597 -148 
FlagCoverage Δ
Debug68.80% <100.00%> (-0.60%)⬇️
production63.26% <100.00%> (-0.67%)⬇️
test88.49% <ø> (-0.41%)⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

FilesCoverage Δ
src/Microsoft.ML.CpuMath/AvxIntrinsics.cs58.18% <ø> (-38.51%)⬇️
src/Microsoft.ML.CpuMath/CpuMathUtils.cs100.00% <100.00%> (ø)
...rc/Microsoft.ML.CpuMath/CpuMathUtils.netcoreapp.cs97.80% <100.00%> (-0.84%)⬇️
src/Microsoft.ML.CpuMath/SseIntrinsics.cs54.80% <ø> (-41.55%)⬇️

... and 48 files with indirect coverage changes

@michaelgsharp

Copy link
Copy Markdown
ContributorAuthor

/azp run

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 2 pipeline(s).

Comment threadsrc/Microsoft.ML.CpuMath/AvxIntrinsics.cs
@michaelgsharp
michaelgsharp merged commit d2cf997 into dotnet:mainNov 15, 2023
@michaelgsharp
michaelgsharp deleted the tensor-math2 branch November 15, 2023 05:46
@github-actionsgithub-actionsBot locked and limited conversation to collaborators Dec 15, 2023
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@michaelgsharp@stephentoub@tannergooding@JakeRadMSFT