Change the _maxCalibrationExamples default on CalibratorUtils - #5415

Merged
antoniovs1029 merged 2 commits into
dotnet:masterfrom
antoniovs1029:platt-TLC
Sep 30, 2020
Merged

Change the _maxCalibrationExamples default on CalibratorUtils#5415
antoniovs1029 merged 2 commits into
dotnet:masterfrom
antoniovs1029:platt-TLC

Conversation

@antoniovs1029

@antoniovs1029antoniovs1029 commented Sep 30, 2020

Copy link
Copy Markdown
Contributor

As reported offline, ML.NET yielded different results than TLC when training a PlattCalibrator with the same dataset.

Upon further investigation, it turns out that it only happened on datasets over 1 million rows, and the reason was that when porting the CalibratorUtils class from TLC, a "_maxCalibrationExamples = 1000000" default parameter was added.

Upon reading through the code (in particular CalibratorTrainingBase's ProcessingTrainingExample) it turns out that on TLC TrainCalibrator was called with maxRows = 0, and this made that when training the PlattCalibrator, all the dataset was seen, but only 1M rows where selected randomly to be added to the DataStore. In contrast, on ML.NET that same method was called with maxRows = 1M, and this made that only the first 1M rows were added to the DataStore (instead of randomly selecting them from the complete dataset). This caused bias and undesired results.

@antoniovs1029
antoniovs1029 requested a review from a team as a code ownerSeptember 30, 2020 08:18
@codecov

codecovBot commented Sep 30, 2020

Copy link
Copy Markdown

Codecov Report

Merging #5415 into master will decrease coverage by 0.06%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #5415 +/- ##
==========================================
- Coverage 74.08% 74.02% -0.07% 
==========================================
Files 1019 1019 Lines 190355 190363 +8 Branches 20469 20469 ==========================================
- Hits 141033 140914 -119 - Misses 43791 43905 +114 - Partials 5531 5544 +13 
FlagCoverage Δ
#Debug74.02% <ø> (-0.07%)⬇️
#production69.77% <ø> (-0.09%)⬇️
#test87.71% <ø> (+<0.01%)⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Impacted FilesCoverage Δ
src/Microsoft.ML.Data/Prediction/Calibrator.cs81.29% <ø> (+0.08%)⬆️
...osoft.ML.KMeansClustering/KMeansPlusPlusTrainer.cs83.60% <0.00%> (-7.27%)⬇️
src/Microsoft.ML.Data/Training/TrainerUtils.cs66.86% <0.00%> (-3.82%)⬇️
...crosoft.ML.StandardTrainers/Standard/SdcaBinary.cs85.23% <0.00%> (-3.33%)⬇️
...crosoft.ML.StandardTrainers/Optimizer/Optimizer.cs71.96% <0.00%> (-1.16%)⬇️
...oft.ML.StandardTrainers/Standard/SdcaMulticlass.cs91.46% <0.00%> (-1.03%)⬇️
src/Microsoft.ML.Data/Utils/LossFunctions.cs66.83% <0.00%> (-0.52%)⬇️
...StandardTrainers/Standard/LinearModelParameters.cs66.32% <0.00%> (-0.26%)⬇️
test/Microsoft.ML.Functional.Tests/ONNX.cs100.00% <0.00%> (ø)
test/Microsoft.ML.Tests/OnnxConversionTest.cs96.17% <0.00%> (+<0.01%)⬆️
... and 5 more

// maximum number of rows passed to the calibrator.
private const int _maxCalibrationExamples = 1000000;
// if 0, we'll actually look through the whole dataset to
// when training the calibrator

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Earlier, you had explained to me that if this value is zero, we would look through the whole dataset, but still only use a million rows (randomly selected) for calibration. Is that correct?

Can you please clarify the exact behavior in comments?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is correct for PlattCalibrator, but depending of the calibrator the behavior is different. If this is 0, the only thing that does happen for all the calibrators is that we'll look through all the dataset. I explained this further on the other comment I left, so I don't think it's necessary to clarify it more in here.

if (maxRows > 0 && ++num >= maxRows)
// If maxRows was 0, we'll process all of the rows in the dataset
// Notice that depending of the calibrator, "processing" means
// only using N random rows of the ones that where processed

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit, typo: "depending on the calibrator"

@harishskharishsk left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@antoniovs1029
antoniovs1029 merged commit 57be476 into dotnet:masterSep 30, 2020
frank-dong-ms-zz added a commit that referenced this pull request Oct 8, 2020
* Update to Onnxruntime 1.5.1 (#5406)
* Added variables to tests to control Gpu settings
* Added dependency to prerelease
* Updated to 1.5.1
* Remove prerelease feed
* Nit on GPU variables
* Change the _maxCalibrationExamples default on CalibratorUtils (#5415)
* Change the _maxCalibrationExamples default
* Improving comments
* Fix perf regression in ShuffleRows (#5417)
RowShufflingTransformer is using ChannelReader incorrectly. It needs to block waiting for items to read and was Thread.Sleeping in order to wait, but not spin the current core. This caused a major perf regression.
The fix is to block synchronously correctly - by calling AsTask() on the ValueTask that is returned from the ChannelReader and block on the Task.
Fix#5416
Co-authored-by: Antonio Velázquez <38739674+antoniovs1029@users.noreply.github.com>
Co-authored-by: Eric Erhardt <eric.erhardt@microsoft.com>
mstfbl pushed a commit to mstfbl/machinelearning that referenced this pull request Nov 12, 2020
* Update to Onnxruntime 1.5.1 (dotnet#5406)
* Added variables to tests to control Gpu settings
* Added dependency to prerelease
* Updated to 1.5.1
* Remove prerelease feed
* Nit on GPU variables
* Change the _maxCalibrationExamples default on CalibratorUtils (dotnet#5415)
* Change the _maxCalibrationExamples default
* Improving comments
* Fix perf regression in ShuffleRows (dotnet#5417)
RowShufflingTransformer is using ChannelReader incorrectly. It needs to block waiting for items to read and was Thread.Sleeping in order to wait, but not spin the current core. This caused a major perf regression.
The fix is to block synchronously correctly - by calling AsTask() on the ValueTask that is returned from the ChannelReader and block on the Task.
Fixdotnet#5416
Co-authored-by: Antonio Velázquez <38739674+antoniovs1029@users.noreply.github.com>
Co-authored-by: Eric Erhardt <eric.erhardt@microsoft.com>
mstfbl pushed a commit that referenced this pull request Nov 12, 2020
* Update to Onnxruntime 1.5.1 (#5406)
* Added variables to tests to control Gpu settings
* Added dependency to prerelease
* Updated to 1.5.1
* Remove prerelease feed
* Nit on GPU variables
* Change the _maxCalibrationExamples default on CalibratorUtils (#5415)
* Change the _maxCalibrationExamples default
* Improving comments
* Fix perf regression in ShuffleRows (#5417)
RowShufflingTransformer is using ChannelReader incorrectly. It needs to block waiting for items to read and was Thread.Sleeping in order to wait, but not spin the current core. This caused a major perf regression.
The fix is to block synchronously correctly - by calling AsTask() on the ValueTask that is returned from the ChannelReader and block on the Task.
Fix#5416
Co-authored-by: Antonio Velázquez <38739674+antoniovs1029@users.noreply.github.com>
Co-authored-by: Eric Erhardt <eric.erhardt@microsoft.com>
@ghostghost locked as resolved and limited conversation to collaborators Mar 17, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@antoniovs1029@harishsk
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Change the _maxCalibrationExamples default on CalibratorUtils - #5415

Merged
antoniovs1029 merged 2 commits into
dotnet:masterfrom
antoniovs1029:platt-TLC
Sep 30, 2020
Merged

Change the _maxCalibrationExamples default on CalibratorUtils#5415
antoniovs1029 merged 2 commits into
dotnet:masterfrom
antoniovs1029:platt-TLC

Conversation

@antoniovs1029

@antoniovs1029antoniovs1029 commented Sep 30, 2020

Copy link
Copy Markdown
Contributor

As reported offline, ML.NET yielded different results than TLC when training a PlattCalibrator with the same dataset.

Upon further investigation, it turns out that it only happened on datasets over 1 million rows, and the reason was that when porting the CalibratorUtils class from TLC, a "_maxCalibrationExamples = 1000000" default parameter was added.

Upon reading through the code (in particular CalibratorTrainingBase's ProcessingTrainingExample) it turns out that on TLC TrainCalibrator was called with maxRows = 0, and this made that when training the PlattCalibrator, all the dataset was seen, but only 1M rows where selected randomly to be added to the DataStore. In contrast, on ML.NET that same method was called with maxRows = 1M, and this made that only the first 1M rows were added to the DataStore (instead of randomly selecting them from the complete dataset). This caused bias and undesired results.

@antoniovs1029
antoniovs1029 requested a review from a team as a code ownerSeptember 30, 2020 08:18
@codecov

codecovBot commented Sep 30, 2020

Copy link
Copy Markdown

Codecov Report

Merging #5415 into master will decrease coverage by 0.06%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #5415 +/- ##
==========================================
- Coverage 74.08% 74.02% -0.07% 
==========================================
Files 1019 1019 Lines 190355 190363 +8 Branches 20469 20469 ==========================================
- Hits 141033 140914 -119 - Misses 43791 43905 +114 - Partials 5531 5544 +13 
FlagCoverage Δ
#Debug74.02% <ø> (-0.07%)⬇️
#production69.77% <ø> (-0.09%)⬇️
#test87.71% <ø> (+<0.01%)⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Impacted FilesCoverage Δ
src/Microsoft.ML.Data/Prediction/Calibrator.cs81.29% <ø> (+0.08%)⬆️
...osoft.ML.KMeansClustering/KMeansPlusPlusTrainer.cs83.60% <0.00%> (-7.27%)⬇️
src/Microsoft.ML.Data/Training/TrainerUtils.cs66.86% <0.00%> (-3.82%)⬇️
...crosoft.ML.StandardTrainers/Standard/SdcaBinary.cs85.23% <0.00%> (-3.33%)⬇️
...crosoft.ML.StandardTrainers/Optimizer/Optimizer.cs71.96% <0.00%> (-1.16%)⬇️
...oft.ML.StandardTrainers/Standard/SdcaMulticlass.cs91.46% <0.00%> (-1.03%)⬇️
src/Microsoft.ML.Data/Utils/LossFunctions.cs66.83% <0.00%> (-0.52%)⬇️
...StandardTrainers/Standard/LinearModelParameters.cs66.32% <0.00%> (-0.26%)⬇️
test/Microsoft.ML.Functional.Tests/ONNX.cs100.00% <0.00%> (ø)
test/Microsoft.ML.Tests/OnnxConversionTest.cs96.17% <0.00%> (+<0.01%)⬆️
... and 5 more

// maximum number of rows passed to the calibrator.
private const int _maxCalibrationExamples = 1000000;
// if 0, we'll actually look through the whole dataset to
// when training the calibrator

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Earlier, you had explained to me that if this value is zero, we would look through the whole dataset, but still only use a million rows (randomly selected) for calibration. Is that correct?

Can you please clarify the exact behavior in comments?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is correct for PlattCalibrator, but depending of the calibrator the behavior is different. If this is 0, the only thing that does happen for all the calibrators is that we'll look through all the dataset. I explained this further on the other comment I left, so I don't think it's necessary to clarify it more in here.

if (maxRows > 0 && ++num >= maxRows)
// If maxRows was 0, we'll process all of the rows in the dataset
// Notice that depending of the calibrator, "processing" means
// only using N random rows of the ones that where processed

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit, typo: "depending on the calibrator"

@harishskharishsk left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@antoniovs1029
antoniovs1029 merged commit 57be476 into dotnet:masterSep 30, 2020
frank-dong-ms-zz added a commit that referenced this pull request Oct 8, 2020
* Update to Onnxruntime 1.5.1 (#5406)
* Added variables to tests to control Gpu settings
* Added dependency to prerelease
* Updated to 1.5.1
* Remove prerelease feed
* Nit on GPU variables
* Change the _maxCalibrationExamples default on CalibratorUtils (#5415)
* Change the _maxCalibrationExamples default
* Improving comments
* Fix perf regression in ShuffleRows (#5417)
RowShufflingTransformer is using ChannelReader incorrectly. It needs to block waiting for items to read and was Thread.Sleeping in order to wait, but not spin the current core. This caused a major perf regression.
The fix is to block synchronously correctly - by calling AsTask() on the ValueTask that is returned from the ChannelReader and block on the Task.
Fix#5416
Co-authored-by: Antonio Velázquez <38739674+antoniovs1029@users.noreply.github.com>
Co-authored-by: Eric Erhardt <eric.erhardt@microsoft.com>
mstfbl pushed a commit to mstfbl/machinelearning that referenced this pull request Nov 12, 2020
* Update to Onnxruntime 1.5.1 (dotnet#5406)
* Added variables to tests to control Gpu settings
* Added dependency to prerelease
* Updated to 1.5.1
* Remove prerelease feed
* Nit on GPU variables
* Change the _maxCalibrationExamples default on CalibratorUtils (dotnet#5415)
* Change the _maxCalibrationExamples default
* Improving comments
* Fix perf regression in ShuffleRows (dotnet#5417)
RowShufflingTransformer is using ChannelReader incorrectly. It needs to block waiting for items to read and was Thread.Sleeping in order to wait, but not spin the current core. This caused a major perf regression.
The fix is to block synchronously correctly - by calling AsTask() on the ValueTask that is returned from the ChannelReader and block on the Task.
Fixdotnet#5416
Co-authored-by: Antonio Velázquez <38739674+antoniovs1029@users.noreply.github.com>
Co-authored-by: Eric Erhardt <eric.erhardt@microsoft.com>
mstfbl pushed a commit that referenced this pull request Nov 12, 2020
* Update to Onnxruntime 1.5.1 (#5406)
* Added variables to tests to control Gpu settings
* Added dependency to prerelease
* Updated to 1.5.1
* Remove prerelease feed
* Nit on GPU variables
* Change the _maxCalibrationExamples default on CalibratorUtils (#5415)
* Change the _maxCalibrationExamples default
* Improving comments
* Fix perf regression in ShuffleRows (#5417)
RowShufflingTransformer is using ChannelReader incorrectly. It needs to block waiting for items to read and was Thread.Sleeping in order to wait, but not spin the current core. This caused a major perf regression.
The fix is to block synchronously correctly - by calling AsTask() on the ValueTask that is returned from the ChannelReader and block on the Task.
Fix#5416
Co-authored-by: Antonio Velázquez <38739674+antoniovs1029@users.noreply.github.com>
Co-authored-by: Eric Erhardt <eric.erhardt@microsoft.com>
@ghostghost locked as resolved and limited conversation to collaborators Mar 17, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@antoniovs1029@harishsk
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Change the _maxCalibrationExamples default on CalibratorUtils - #5415

Merged
antoniovs1029 merged 2 commits into
dotnet:masterfrom
antoniovs1029:platt-TLC
Sep 30, 2020
Merged

Change the _maxCalibrationExamples default on CalibratorUtils#5415
antoniovs1029 merged 2 commits into
dotnet:masterfrom
antoniovs1029:platt-TLC

Conversation

@antoniovs1029

@antoniovs1029antoniovs1029 commented Sep 30, 2020

Copy link
Copy Markdown
Contributor

As reported offline, ML.NET yielded different results than TLC when training a PlattCalibrator with the same dataset.

Upon further investigation, it turns out that it only happened on datasets over 1 million rows, and the reason was that when porting the CalibratorUtils class from TLC, a "_maxCalibrationExamples = 1000000" default parameter was added.

Upon reading through the code (in particular CalibratorTrainingBase's ProcessingTrainingExample) it turns out that on TLC TrainCalibrator was called with maxRows = 0, and this made that when training the PlattCalibrator, all the dataset was seen, but only 1M rows where selected randomly to be added to the DataStore. In contrast, on ML.NET that same method was called with maxRows = 1M, and this made that only the first 1M rows were added to the DataStore (instead of randomly selecting them from the complete dataset). This caused bias and undesired results.

@antoniovs1029
antoniovs1029 requested a review from a team as a code ownerSeptember 30, 2020 08:18
@codecov

codecovBot commented Sep 30, 2020

Copy link
Copy Markdown

Codecov Report

Merging #5415 into master will decrease coverage by 0.06%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #5415 +/- ##
==========================================
- Coverage 74.08% 74.02% -0.07% 
==========================================
Files 1019 1019 Lines 190355 190363 +8 Branches 20469 20469 ==========================================
- Hits 141033 140914 -119 - Misses 43791 43905 +114 - Partials 5531 5544 +13 
FlagCoverage Δ
#Debug74.02% <ø> (-0.07%)⬇️
#production69.77% <ø> (-0.09%)⬇️
#test87.71% <ø> (+<0.01%)⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Impacted FilesCoverage Δ
src/Microsoft.ML.Data/Prediction/Calibrator.cs81.29% <ø> (+0.08%)⬆️
...osoft.ML.KMeansClustering/KMeansPlusPlusTrainer.cs83.60% <0.00%> (-7.27%)⬇️
src/Microsoft.ML.Data/Training/TrainerUtils.cs66.86% <0.00%> (-3.82%)⬇️
...crosoft.ML.StandardTrainers/Standard/SdcaBinary.cs85.23% <0.00%> (-3.33%)⬇️
...crosoft.ML.StandardTrainers/Optimizer/Optimizer.cs71.96% <0.00%> (-1.16%)⬇️
...oft.ML.StandardTrainers/Standard/SdcaMulticlass.cs91.46% <0.00%> (-1.03%)⬇️
src/Microsoft.ML.Data/Utils/LossFunctions.cs66.83% <0.00%> (-0.52%)⬇️
...StandardTrainers/Standard/LinearModelParameters.cs66.32% <0.00%> (-0.26%)⬇️
test/Microsoft.ML.Functional.Tests/ONNX.cs100.00% <0.00%> (ø)
test/Microsoft.ML.Tests/OnnxConversionTest.cs96.17% <0.00%> (+<0.01%)⬆️
... and 5 more

// maximum number of rows passed to the calibrator.
private const int _maxCalibrationExamples = 1000000;
// if 0, we'll actually look through the whole dataset to
// when training the calibrator

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Earlier, you had explained to me that if this value is zero, we would look through the whole dataset, but still only use a million rows (randomly selected) for calibration. Is that correct?

Can you please clarify the exact behavior in comments?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is correct for PlattCalibrator, but depending of the calibrator the behavior is different. If this is 0, the only thing that does happen for all the calibrators is that we'll look through all the dataset. I explained this further on the other comment I left, so I don't think it's necessary to clarify it more in here.

if (maxRows > 0 && ++num >= maxRows)
// If maxRows was 0, we'll process all of the rows in the dataset
// Notice that depending of the calibrator, "processing" means
// only using N random rows of the ones that where processed

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit, typo: "depending on the calibrator"

@harishskharishsk left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@antoniovs1029
antoniovs1029 merged commit 57be476 into dotnet:masterSep 30, 2020
frank-dong-ms-zz added a commit that referenced this pull request Oct 8, 2020
* Update to Onnxruntime 1.5.1 (#5406)
* Added variables to tests to control Gpu settings
* Added dependency to prerelease
* Updated to 1.5.1
* Remove prerelease feed
* Nit on GPU variables
* Change the _maxCalibrationExamples default on CalibratorUtils (#5415)
* Change the _maxCalibrationExamples default
* Improving comments
* Fix perf regression in ShuffleRows (#5417)
RowShufflingTransformer is using ChannelReader incorrectly. It needs to block waiting for items to read and was Thread.Sleeping in order to wait, but not spin the current core. This caused a major perf regression.
The fix is to block synchronously correctly - by calling AsTask() on the ValueTask that is returned from the ChannelReader and block on the Task.
Fix#5416
Co-authored-by: Antonio Velázquez <38739674+antoniovs1029@users.noreply.github.com>
Co-authored-by: Eric Erhardt <eric.erhardt@microsoft.com>
mstfbl pushed a commit to mstfbl/machinelearning that referenced this pull request Nov 12, 2020
* Update to Onnxruntime 1.5.1 (dotnet#5406)
* Added variables to tests to control Gpu settings
* Added dependency to prerelease
* Updated to 1.5.1
* Remove prerelease feed
* Nit on GPU variables
* Change the _maxCalibrationExamples default on CalibratorUtils (dotnet#5415)
* Change the _maxCalibrationExamples default
* Improving comments
* Fix perf regression in ShuffleRows (dotnet#5417)
RowShufflingTransformer is using ChannelReader incorrectly. It needs to block waiting for items to read and was Thread.Sleeping in order to wait, but not spin the current core. This caused a major perf regression.
The fix is to block synchronously correctly - by calling AsTask() on the ValueTask that is returned from the ChannelReader and block on the Task.
Fixdotnet#5416
Co-authored-by: Antonio Velázquez <38739674+antoniovs1029@users.noreply.github.com>
Co-authored-by: Eric Erhardt <eric.erhardt@microsoft.com>
mstfbl pushed a commit that referenced this pull request Nov 12, 2020
* Update to Onnxruntime 1.5.1 (#5406)
* Added variables to tests to control Gpu settings
* Added dependency to prerelease
* Updated to 1.5.1
* Remove prerelease feed
* Nit on GPU variables
* Change the _maxCalibrationExamples default on CalibratorUtils (#5415)
* Change the _maxCalibrationExamples default
* Improving comments
* Fix perf regression in ShuffleRows (#5417)
RowShufflingTransformer is using ChannelReader incorrectly. It needs to block waiting for items to read and was Thread.Sleeping in order to wait, but not spin the current core. This caused a major perf regression.
The fix is to block synchronously correctly - by calling AsTask() on the ValueTask that is returned from the ChannelReader and block on the Task.
Fix#5416
Co-authored-by: Antonio Velázquez <38739674+antoniovs1029@users.noreply.github.com>
Co-authored-by: Eric Erhardt <eric.erhardt@microsoft.com>
@ghostghost locked as resolved and limited conversation to collaborators Mar 17, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@antoniovs1029@harishsk
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Change the _maxCalibrationExamples default on CalibratorUtils - #5415

Merged
antoniovs1029 merged 2 commits into
dotnet:masterfrom
antoniovs1029:platt-TLC
Sep 30, 2020
Merged

Change the _maxCalibrationExamples default on CalibratorUtils#5415
antoniovs1029 merged 2 commits into
dotnet:masterfrom
antoniovs1029:platt-TLC

Conversation

@antoniovs1029

@antoniovs1029antoniovs1029 commented Sep 30, 2020

Copy link
Copy Markdown
Contributor

As reported offline, ML.NET yielded different results than TLC when training a PlattCalibrator with the same dataset.

Upon further investigation, it turns out that it only happened on datasets over 1 million rows, and the reason was that when porting the CalibratorUtils class from TLC, a "_maxCalibrationExamples = 1000000" default parameter was added.

Upon reading through the code (in particular CalibratorTrainingBase's ProcessingTrainingExample) it turns out that on TLC TrainCalibrator was called with maxRows = 0, and this made that when training the PlattCalibrator, all the dataset was seen, but only 1M rows where selected randomly to be added to the DataStore. In contrast, on ML.NET that same method was called with maxRows = 1M, and this made that only the first 1M rows were added to the DataStore (instead of randomly selecting them from the complete dataset). This caused bias and undesired results.

@antoniovs1029
antoniovs1029 requested a review from a team as a code ownerSeptember 30, 2020 08:18
@codecov

codecovBot commented Sep 30, 2020

Copy link
Copy Markdown

Codecov Report

Merging #5415 into master will decrease coverage by 0.06%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #5415 +/- ##
==========================================
- Coverage 74.08% 74.02% -0.07% 
==========================================
Files 1019 1019 Lines 190355 190363 +8 Branches 20469 20469 ==========================================
- Hits 141033 140914 -119 - Misses 43791 43905 +114 - Partials 5531 5544 +13 
FlagCoverage Δ
#Debug74.02% <ø> (-0.07%)⬇️
#production69.77% <ø> (-0.09%)⬇️
#test87.71% <ø> (+<0.01%)⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Impacted FilesCoverage Δ
src/Microsoft.ML.Data/Prediction/Calibrator.cs81.29% <ø> (+0.08%)⬆️
...osoft.ML.KMeansClustering/KMeansPlusPlusTrainer.cs83.60% <0.00%> (-7.27%)⬇️
src/Microsoft.ML.Data/Training/TrainerUtils.cs66.86% <0.00%> (-3.82%)⬇️
...crosoft.ML.StandardTrainers/Standard/SdcaBinary.cs85.23% <0.00%> (-3.33%)⬇️
...crosoft.ML.StandardTrainers/Optimizer/Optimizer.cs71.96% <0.00%> (-1.16%)⬇️
...oft.ML.StandardTrainers/Standard/SdcaMulticlass.cs91.46% <0.00%> (-1.03%)⬇️
src/Microsoft.ML.Data/Utils/LossFunctions.cs66.83% <0.00%> (-0.52%)⬇️
...StandardTrainers/Standard/LinearModelParameters.cs66.32% <0.00%> (-0.26%)⬇️
test/Microsoft.ML.Functional.Tests/ONNX.cs100.00% <0.00%> (ø)
test/Microsoft.ML.Tests/OnnxConversionTest.cs96.17% <0.00%> (+<0.01%)⬆️
... and 5 more

// maximum number of rows passed to the calibrator.
private const int _maxCalibrationExamples = 1000000;
// if 0, we'll actually look through the whole dataset to
// when training the calibrator

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Earlier, you had explained to me that if this value is zero, we would look through the whole dataset, but still only use a million rows (randomly selected) for calibration. Is that correct?

Can you please clarify the exact behavior in comments?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is correct for PlattCalibrator, but depending of the calibrator the behavior is different. If this is 0, the only thing that does happen for all the calibrators is that we'll look through all the dataset. I explained this further on the other comment I left, so I don't think it's necessary to clarify it more in here.

if (maxRows > 0 && ++num >= maxRows)
// If maxRows was 0, we'll process all of the rows in the dataset
// Notice that depending of the calibrator, "processing" means
// only using N random rows of the ones that where processed

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit, typo: "depending on the calibrator"

@harishskharishsk left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@antoniovs1029
antoniovs1029 merged commit 57be476 into dotnet:masterSep 30, 2020
frank-dong-ms-zz added a commit that referenced this pull request Oct 8, 2020
* Update to Onnxruntime 1.5.1 (#5406)
* Added variables to tests to control Gpu settings
* Added dependency to prerelease
* Updated to 1.5.1
* Remove prerelease feed
* Nit on GPU variables
* Change the _maxCalibrationExamples default on CalibratorUtils (#5415)
* Change the _maxCalibrationExamples default
* Improving comments
* Fix perf regression in ShuffleRows (#5417)
RowShufflingTransformer is using ChannelReader incorrectly. It needs to block waiting for items to read and was Thread.Sleeping in order to wait, but not spin the current core. This caused a major perf regression.
The fix is to block synchronously correctly - by calling AsTask() on the ValueTask that is returned from the ChannelReader and block on the Task.
Fix#5416
Co-authored-by: Antonio Velázquez <38739674+antoniovs1029@users.noreply.github.com>
Co-authored-by: Eric Erhardt <eric.erhardt@microsoft.com>
mstfbl pushed a commit to mstfbl/machinelearning that referenced this pull request Nov 12, 2020
* Update to Onnxruntime 1.5.1 (dotnet#5406)
* Added variables to tests to control Gpu settings
* Added dependency to prerelease
* Updated to 1.5.1
* Remove prerelease feed
* Nit on GPU variables
* Change the _maxCalibrationExamples default on CalibratorUtils (dotnet#5415)
* Change the _maxCalibrationExamples default
* Improving comments
* Fix perf regression in ShuffleRows (dotnet#5417)
RowShufflingTransformer is using ChannelReader incorrectly. It needs to block waiting for items to read and was Thread.Sleeping in order to wait, but not spin the current core. This caused a major perf regression.
The fix is to block synchronously correctly - by calling AsTask() on the ValueTask that is returned from the ChannelReader and block on the Task.
Fixdotnet#5416
Co-authored-by: Antonio Velázquez <38739674+antoniovs1029@users.noreply.github.com>
Co-authored-by: Eric Erhardt <eric.erhardt@microsoft.com>
mstfbl pushed a commit that referenced this pull request Nov 12, 2020
* Update to Onnxruntime 1.5.1 (#5406)
* Added variables to tests to control Gpu settings
* Added dependency to prerelease
* Updated to 1.5.1
* Remove prerelease feed
* Nit on GPU variables
* Change the _maxCalibrationExamples default on CalibratorUtils (#5415)
* Change the _maxCalibrationExamples default
* Improving comments
* Fix perf regression in ShuffleRows (#5417)
RowShufflingTransformer is using ChannelReader incorrectly. It needs to block waiting for items to read and was Thread.Sleeping in order to wait, but not spin the current core. This caused a major perf regression.
The fix is to block synchronously correctly - by calling AsTask() on the ValueTask that is returned from the ChannelReader and block on the Task.
Fix#5416
Co-authored-by: Antonio Velázquez <38739674+antoniovs1029@users.noreply.github.com>
Co-authored-by: Eric Erhardt <eric.erhardt@microsoft.com>
@ghostghost locked as resolved and limited conversation to collaborators Mar 17, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@antoniovs1029@harishsk
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Change the _maxCalibrationExamples default on CalibratorUtils - #5415

Merged
antoniovs1029 merged 2 commits into
dotnet:masterfrom
antoniovs1029:platt-TLC
Sep 30, 2020
Merged

Change the _maxCalibrationExamples default on CalibratorUtils#5415
antoniovs1029 merged 2 commits into
dotnet:masterfrom
antoniovs1029:platt-TLC

Conversation

@antoniovs1029

@antoniovs1029antoniovs1029 commented Sep 30, 2020

Copy link
Copy Markdown
Contributor

As reported offline, ML.NET yielded different results than TLC when training a PlattCalibrator with the same dataset.

Upon further investigation, it turns out that it only happened on datasets over 1 million rows, and the reason was that when porting the CalibratorUtils class from TLC, a "_maxCalibrationExamples = 1000000" default parameter was added.

Upon reading through the code (in particular CalibratorTrainingBase's ProcessingTrainingExample) it turns out that on TLC TrainCalibrator was called with maxRows = 0, and this made that when training the PlattCalibrator, all the dataset was seen, but only 1M rows where selected randomly to be added to the DataStore. In contrast, on ML.NET that same method was called with maxRows = 1M, and this made that only the first 1M rows were added to the DataStore (instead of randomly selecting them from the complete dataset). This caused bias and undesired results.

@antoniovs1029
antoniovs1029 requested a review from a team as a code ownerSeptember 30, 2020 08:18
@codecov

codecovBot commented Sep 30, 2020

Copy link
Copy Markdown

Codecov Report

Merging #5415 into master will decrease coverage by 0.06%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #5415 +/- ##
==========================================
- Coverage 74.08% 74.02% -0.07% 
==========================================
Files 1019 1019 Lines 190355 190363 +8 Branches 20469 20469 ==========================================
- Hits 141033 140914 -119 - Misses 43791 43905 +114 - Partials 5531 5544 +13 
FlagCoverage Δ
#Debug74.02% <ø> (-0.07%)⬇️
#production69.77% <ø> (-0.09%)⬇️
#test87.71% <ø> (+<0.01%)⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Impacted FilesCoverage Δ
src/Microsoft.ML.Data/Prediction/Calibrator.cs81.29% <ø> (+0.08%)⬆️
...osoft.ML.KMeansClustering/KMeansPlusPlusTrainer.cs83.60% <0.00%> (-7.27%)⬇️
src/Microsoft.ML.Data/Training/TrainerUtils.cs66.86% <0.00%> (-3.82%)⬇️
...crosoft.ML.StandardTrainers/Standard/SdcaBinary.cs85.23% <0.00%> (-3.33%)⬇️
...crosoft.ML.StandardTrainers/Optimizer/Optimizer.cs71.96% <0.00%> (-1.16%)⬇️
...oft.ML.StandardTrainers/Standard/SdcaMulticlass.cs91.46% <0.00%> (-1.03%)⬇️
src/Microsoft.ML.Data/Utils/LossFunctions.cs66.83% <0.00%> (-0.52%)⬇️
...StandardTrainers/Standard/LinearModelParameters.cs66.32% <0.00%> (-0.26%)⬇️
test/Microsoft.ML.Functional.Tests/ONNX.cs100.00% <0.00%> (ø)
test/Microsoft.ML.Tests/OnnxConversionTest.cs96.17% <0.00%> (+<0.01%)⬆️
... and 5 more

// maximum number of rows passed to the calibrator.
private const int _maxCalibrationExamples = 1000000;
// if 0, we'll actually look through the whole dataset to
// when training the calibrator

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Earlier, you had explained to me that if this value is zero, we would look through the whole dataset, but still only use a million rows (randomly selected) for calibration. Is that correct?

Can you please clarify the exact behavior in comments?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is correct for PlattCalibrator, but depending of the calibrator the behavior is different. If this is 0, the only thing that does happen for all the calibrators is that we'll look through all the dataset. I explained this further on the other comment I left, so I don't think it's necessary to clarify it more in here.

if (maxRows > 0 && ++num >= maxRows)
// If maxRows was 0, we'll process all of the rows in the dataset
// Notice that depending of the calibrator, "processing" means
// only using N random rows of the ones that where processed

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit, typo: "depending on the calibrator"

@harishskharishsk left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@antoniovs1029
antoniovs1029 merged commit 57be476 into dotnet:masterSep 30, 2020
frank-dong-ms-zz added a commit that referenced this pull request Oct 8, 2020
* Update to Onnxruntime 1.5.1 (#5406)
* Added variables to tests to control Gpu settings
* Added dependency to prerelease
* Updated to 1.5.1
* Remove prerelease feed
* Nit on GPU variables
* Change the _maxCalibrationExamples default on CalibratorUtils (#5415)
* Change the _maxCalibrationExamples default
* Improving comments
* Fix perf regression in ShuffleRows (#5417)
RowShufflingTransformer is using ChannelReader incorrectly. It needs to block waiting for items to read and was Thread.Sleeping in order to wait, but not spin the current core. This caused a major perf regression.
The fix is to block synchronously correctly - by calling AsTask() on the ValueTask that is returned from the ChannelReader and block on the Task.
Fix#5416
Co-authored-by: Antonio Velázquez <38739674+antoniovs1029@users.noreply.github.com>
Co-authored-by: Eric Erhardt <eric.erhardt@microsoft.com>
mstfbl pushed a commit to mstfbl/machinelearning that referenced this pull request Nov 12, 2020
* Update to Onnxruntime 1.5.1 (dotnet#5406)
* Added variables to tests to control Gpu settings
* Added dependency to prerelease
* Updated to 1.5.1
* Remove prerelease feed
* Nit on GPU variables
* Change the _maxCalibrationExamples default on CalibratorUtils (dotnet#5415)
* Change the _maxCalibrationExamples default
* Improving comments
* Fix perf regression in ShuffleRows (dotnet#5417)
RowShufflingTransformer is using ChannelReader incorrectly. It needs to block waiting for items to read and was Thread.Sleeping in order to wait, but not spin the current core. This caused a major perf regression.
The fix is to block synchronously correctly - by calling AsTask() on the ValueTask that is returned from the ChannelReader and block on the Task.
Fixdotnet#5416
Co-authored-by: Antonio Velázquez <38739674+antoniovs1029@users.noreply.github.com>
Co-authored-by: Eric Erhardt <eric.erhardt@microsoft.com>
mstfbl pushed a commit that referenced this pull request Nov 12, 2020
* Update to Onnxruntime 1.5.1 (#5406)
* Added variables to tests to control Gpu settings
* Added dependency to prerelease
* Updated to 1.5.1
* Remove prerelease feed
* Nit on GPU variables
* Change the _maxCalibrationExamples default on CalibratorUtils (#5415)
* Change the _maxCalibrationExamples default
* Improving comments
* Fix perf regression in ShuffleRows (#5417)
RowShufflingTransformer is using ChannelReader incorrectly. It needs to block waiting for items to read and was Thread.Sleeping in order to wait, but not spin the current core. This caused a major perf regression.
The fix is to block synchronously correctly - by calling AsTask() on the ValueTask that is returned from the ChannelReader and block on the Task.
Fix#5416
Co-authored-by: Antonio Velázquez <38739674+antoniovs1029@users.noreply.github.com>
Co-authored-by: Eric Erhardt <eric.erhardt@microsoft.com>
@ghostghost locked as resolved and limited conversation to collaborators Mar 17, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@antoniovs1029@harishsk
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Change the _maxCalibrationExamples default on CalibratorUtils - #5415

Merged
antoniovs1029 merged 2 commits into
dotnet:masterfrom
antoniovs1029:platt-TLC
Sep 30, 2020
Merged

Change the _maxCalibrationExamples default on CalibratorUtils#5415
antoniovs1029 merged 2 commits into
dotnet:masterfrom
antoniovs1029:platt-TLC

Conversation

@antoniovs1029

@antoniovs1029antoniovs1029 commented Sep 30, 2020

Copy link
Copy Markdown
Contributor

As reported offline, ML.NET yielded different results than TLC when training a PlattCalibrator with the same dataset.

Upon further investigation, it turns out that it only happened on datasets over 1 million rows, and the reason was that when porting the CalibratorUtils class from TLC, a "_maxCalibrationExamples = 1000000" default parameter was added.

Upon reading through the code (in particular CalibratorTrainingBase's ProcessingTrainingExample) it turns out that on TLC TrainCalibrator was called with maxRows = 0, and this made that when training the PlattCalibrator, all the dataset was seen, but only 1M rows where selected randomly to be added to the DataStore. In contrast, on ML.NET that same method was called with maxRows = 1M, and this made that only the first 1M rows were added to the DataStore (instead of randomly selecting them from the complete dataset). This caused bias and undesired results.

@antoniovs1029
antoniovs1029 requested a review from a team as a code ownerSeptember 30, 2020 08:18
@codecov

codecovBot commented Sep 30, 2020

Copy link
Copy Markdown

Codecov Report

Merging #5415 into master will decrease coverage by 0.06%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #5415 +/- ##
==========================================
- Coverage 74.08% 74.02% -0.07% 
==========================================
Files 1019 1019 Lines 190355 190363 +8 Branches 20469 20469 ==========================================
- Hits 141033 140914 -119 - Misses 43791 43905 +114 - Partials 5531 5544 +13 
FlagCoverage Δ
#Debug74.02% <ø> (-0.07%)⬇️
#production69.77% <ø> (-0.09%)⬇️
#test87.71% <ø> (+<0.01%)⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Impacted FilesCoverage Δ
src/Microsoft.ML.Data/Prediction/Calibrator.cs81.29% <ø> (+0.08%)⬆️
...osoft.ML.KMeansClustering/KMeansPlusPlusTrainer.cs83.60% <0.00%> (-7.27%)⬇️
src/Microsoft.ML.Data/Training/TrainerUtils.cs66.86% <0.00%> (-3.82%)⬇️
...crosoft.ML.StandardTrainers/Standard/SdcaBinary.cs85.23% <0.00%> (-3.33%)⬇️
...crosoft.ML.StandardTrainers/Optimizer/Optimizer.cs71.96% <0.00%> (-1.16%)⬇️
...oft.ML.StandardTrainers/Standard/SdcaMulticlass.cs91.46% <0.00%> (-1.03%)⬇️
src/Microsoft.ML.Data/Utils/LossFunctions.cs66.83% <0.00%> (-0.52%)⬇️
...StandardTrainers/Standard/LinearModelParameters.cs66.32% <0.00%> (-0.26%)⬇️
test/Microsoft.ML.Functional.Tests/ONNX.cs100.00% <0.00%> (ø)
test/Microsoft.ML.Tests/OnnxConversionTest.cs96.17% <0.00%> (+<0.01%)⬆️
... and 5 more

// maximum number of rows passed to the calibrator.
private const int _maxCalibrationExamples = 1000000;
// if 0, we'll actually look through the whole dataset to
// when training the calibrator

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Earlier, you had explained to me that if this value is zero, we would look through the whole dataset, but still only use a million rows (randomly selected) for calibration. Is that correct?

Can you please clarify the exact behavior in comments?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is correct for PlattCalibrator, but depending of the calibrator the behavior is different. If this is 0, the only thing that does happen for all the calibrators is that we'll look through all the dataset. I explained this further on the other comment I left, so I don't think it's necessary to clarify it more in here.

if (maxRows > 0 && ++num >= maxRows)
// If maxRows was 0, we'll process all of the rows in the dataset
// Notice that depending of the calibrator, "processing" means
// only using N random rows of the ones that where processed

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit, typo: "depending on the calibrator"

@harishskharishsk left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@antoniovs1029
antoniovs1029 merged commit 57be476 into dotnet:masterSep 30, 2020
frank-dong-ms-zz added a commit that referenced this pull request Oct 8, 2020
* Update to Onnxruntime 1.5.1 (#5406)
* Added variables to tests to control Gpu settings
* Added dependency to prerelease
* Updated to 1.5.1
* Remove prerelease feed
* Nit on GPU variables
* Change the _maxCalibrationExamples default on CalibratorUtils (#5415)
* Change the _maxCalibrationExamples default
* Improving comments
* Fix perf regression in ShuffleRows (#5417)
RowShufflingTransformer is using ChannelReader incorrectly. It needs to block waiting for items to read and was Thread.Sleeping in order to wait, but not spin the current core. This caused a major perf regression.
The fix is to block synchronously correctly - by calling AsTask() on the ValueTask that is returned from the ChannelReader and block on the Task.
Fix#5416
Co-authored-by: Antonio Velázquez <38739674+antoniovs1029@users.noreply.github.com>
Co-authored-by: Eric Erhardt <eric.erhardt@microsoft.com>
mstfbl pushed a commit to mstfbl/machinelearning that referenced this pull request Nov 12, 2020
* Update to Onnxruntime 1.5.1 (dotnet#5406)
* Added variables to tests to control Gpu settings
* Added dependency to prerelease
* Updated to 1.5.1
* Remove prerelease feed
* Nit on GPU variables
* Change the _maxCalibrationExamples default on CalibratorUtils (dotnet#5415)
* Change the _maxCalibrationExamples default
* Improving comments
* Fix perf regression in ShuffleRows (dotnet#5417)
RowShufflingTransformer is using ChannelReader incorrectly. It needs to block waiting for items to read and was Thread.Sleeping in order to wait, but not spin the current core. This caused a major perf regression.
The fix is to block synchronously correctly - by calling AsTask() on the ValueTask that is returned from the ChannelReader and block on the Task.
Fixdotnet#5416
Co-authored-by: Antonio Velázquez <38739674+antoniovs1029@users.noreply.github.com>
Co-authored-by: Eric Erhardt <eric.erhardt@microsoft.com>
mstfbl pushed a commit that referenced this pull request Nov 12, 2020
* Update to Onnxruntime 1.5.1 (#5406)
* Added variables to tests to control Gpu settings
* Added dependency to prerelease
* Updated to 1.5.1
* Remove prerelease feed
* Nit on GPU variables
* Change the _maxCalibrationExamples default on CalibratorUtils (#5415)
* Change the _maxCalibrationExamples default
* Improving comments
* Fix perf regression in ShuffleRows (#5417)
RowShufflingTransformer is using ChannelReader incorrectly. It needs to block waiting for items to read and was Thread.Sleeping in order to wait, but not spin the current core. This caused a major perf regression.
The fix is to block synchronously correctly - by calling AsTask() on the ValueTask that is returned from the ChannelReader and block on the Task.
Fix#5416
Co-authored-by: Antonio Velázquez <38739674+antoniovs1029@users.noreply.github.com>
Co-authored-by: Eric Erhardt <eric.erhardt@microsoft.com>
@ghostghost locked as resolved and limited conversation to collaborators Mar 17, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@antoniovs1029@harishsk
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Change the _maxCalibrationExamples default on CalibratorUtils - #5415

Merged
antoniovs1029 merged 2 commits into
dotnet:masterfrom
antoniovs1029:platt-TLC
Sep 30, 2020
Merged

Change the _maxCalibrationExamples default on CalibratorUtils#5415
antoniovs1029 merged 2 commits into
dotnet:masterfrom
antoniovs1029:platt-TLC

Conversation

@antoniovs1029

@antoniovs1029antoniovs1029 commented Sep 30, 2020

Copy link
Copy Markdown
Contributor

As reported offline, ML.NET yielded different results than TLC when training a PlattCalibrator with the same dataset.

Upon further investigation, it turns out that it only happened on datasets over 1 million rows, and the reason was that when porting the CalibratorUtils class from TLC, a "_maxCalibrationExamples = 1000000" default parameter was added.

Upon reading through the code (in particular CalibratorTrainingBase's ProcessingTrainingExample) it turns out that on TLC TrainCalibrator was called with maxRows = 0, and this made that when training the PlattCalibrator, all the dataset was seen, but only 1M rows where selected randomly to be added to the DataStore. In contrast, on ML.NET that same method was called with maxRows = 1M, and this made that only the first 1M rows were added to the DataStore (instead of randomly selecting them from the complete dataset). This caused bias and undesired results.

@antoniovs1029
antoniovs1029 requested a review from a team as a code ownerSeptember 30, 2020 08:18
@codecov

codecovBot commented Sep 30, 2020

Copy link
Copy Markdown

Codecov Report

Merging #5415 into master will decrease coverage by 0.06%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #5415 +/- ##
==========================================
- Coverage 74.08% 74.02% -0.07% 
==========================================
Files 1019 1019 Lines 190355 190363 +8 Branches 20469 20469 ==========================================
- Hits 141033 140914 -119 - Misses 43791 43905 +114 - Partials 5531 5544 +13 
FlagCoverage Δ
#Debug74.02% <ø> (-0.07%)⬇️
#production69.77% <ø> (-0.09%)⬇️
#test87.71% <ø> (+<0.01%)⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Impacted FilesCoverage Δ
src/Microsoft.ML.Data/Prediction/Calibrator.cs81.29% <ø> (+0.08%)⬆️
...osoft.ML.KMeansClustering/KMeansPlusPlusTrainer.cs83.60% <0.00%> (-7.27%)⬇️
src/Microsoft.ML.Data/Training/TrainerUtils.cs66.86% <0.00%> (-3.82%)⬇️
...crosoft.ML.StandardTrainers/Standard/SdcaBinary.cs85.23% <0.00%> (-3.33%)⬇️
...crosoft.ML.StandardTrainers/Optimizer/Optimizer.cs71.96% <0.00%> (-1.16%)⬇️
...oft.ML.StandardTrainers/Standard/SdcaMulticlass.cs91.46% <0.00%> (-1.03%)⬇️
src/Microsoft.ML.Data/Utils/LossFunctions.cs66.83% <0.00%> (-0.52%)⬇️
...StandardTrainers/Standard/LinearModelParameters.cs66.32% <0.00%> (-0.26%)⬇️
test/Microsoft.ML.Functional.Tests/ONNX.cs100.00% <0.00%> (ø)
test/Microsoft.ML.Tests/OnnxConversionTest.cs96.17% <0.00%> (+<0.01%)⬆️
... and 5 more

// maximum number of rows passed to the calibrator.
private const int _maxCalibrationExamples = 1000000;
// if 0, we'll actually look through the whole dataset to
// when training the calibrator

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Earlier, you had explained to me that if this value is zero, we would look through the whole dataset, but still only use a million rows (randomly selected) for calibration. Is that correct?

Can you please clarify the exact behavior in comments?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is correct for PlattCalibrator, but depending of the calibrator the behavior is different. If this is 0, the only thing that does happen for all the calibrators is that we'll look through all the dataset. I explained this further on the other comment I left, so I don't think it's necessary to clarify it more in here.

if (maxRows > 0 && ++num >= maxRows)
// If maxRows was 0, we'll process all of the rows in the dataset
// Notice that depending of the calibrator, "processing" means
// only using N random rows of the ones that where processed

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit, typo: "depending on the calibrator"

@harishskharishsk left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@antoniovs1029
antoniovs1029 merged commit 57be476 into dotnet:masterSep 30, 2020
frank-dong-ms-zz added a commit that referenced this pull request Oct 8, 2020
* Update to Onnxruntime 1.5.1 (#5406)
* Added variables to tests to control Gpu settings
* Added dependency to prerelease
* Updated to 1.5.1
* Remove prerelease feed
* Nit on GPU variables
* Change the _maxCalibrationExamples default on CalibratorUtils (#5415)
* Change the _maxCalibrationExamples default
* Improving comments
* Fix perf regression in ShuffleRows (#5417)
RowShufflingTransformer is using ChannelReader incorrectly. It needs to block waiting for items to read and was Thread.Sleeping in order to wait, but not spin the current core. This caused a major perf regression.
The fix is to block synchronously correctly - by calling AsTask() on the ValueTask that is returned from the ChannelReader and block on the Task.
Fix#5416
Co-authored-by: Antonio Velázquez <38739674+antoniovs1029@users.noreply.github.com>
Co-authored-by: Eric Erhardt <eric.erhardt@microsoft.com>
mstfbl pushed a commit to mstfbl/machinelearning that referenced this pull request Nov 12, 2020
* Update to Onnxruntime 1.5.1 (dotnet#5406)
* Added variables to tests to control Gpu settings
* Added dependency to prerelease
* Updated to 1.5.1
* Remove prerelease feed
* Nit on GPU variables
* Change the _maxCalibrationExamples default on CalibratorUtils (dotnet#5415)
* Change the _maxCalibrationExamples default
* Improving comments
* Fix perf regression in ShuffleRows (dotnet#5417)
RowShufflingTransformer is using ChannelReader incorrectly. It needs to block waiting for items to read and was Thread.Sleeping in order to wait, but not spin the current core. This caused a major perf regression.
The fix is to block synchronously correctly - by calling AsTask() on the ValueTask that is returned from the ChannelReader and block on the Task.
Fixdotnet#5416
Co-authored-by: Antonio Velázquez <38739674+antoniovs1029@users.noreply.github.com>
Co-authored-by: Eric Erhardt <eric.erhardt@microsoft.com>
mstfbl pushed a commit that referenced this pull request Nov 12, 2020
* Update to Onnxruntime 1.5.1 (#5406)
* Added variables to tests to control Gpu settings
* Added dependency to prerelease
* Updated to 1.5.1
* Remove prerelease feed
* Nit on GPU variables
* Change the _maxCalibrationExamples default on CalibratorUtils (#5415)
* Change the _maxCalibrationExamples default
* Improving comments
* Fix perf regression in ShuffleRows (#5417)
RowShufflingTransformer is using ChannelReader incorrectly. It needs to block waiting for items to read and was Thread.Sleeping in order to wait, but not spin the current core. This caused a major perf regression.
The fix is to block synchronously correctly - by calling AsTask() on the ValueTask that is returned from the ChannelReader and block on the Task.
Fix#5416
Co-authored-by: Antonio Velázquez <38739674+antoniovs1029@users.noreply.github.com>
Co-authored-by: Eric Erhardt <eric.erhardt@microsoft.com>
@ghostghost locked as resolved and limited conversation to collaborators Mar 17, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@antoniovs1029@harishsk
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Change the _maxCalibrationExamples default on CalibratorUtils - #5415

Merged
antoniovs1029 merged 2 commits into
dotnet:masterfrom
antoniovs1029:platt-TLC
Sep 30, 2020
Merged

Change the _maxCalibrationExamples default on CalibratorUtils#5415
antoniovs1029 merged 2 commits into
dotnet:masterfrom
antoniovs1029:platt-TLC

Conversation

@antoniovs1029

@antoniovs1029antoniovs1029 commented Sep 30, 2020

Copy link
Copy Markdown
Contributor

As reported offline, ML.NET yielded different results than TLC when training a PlattCalibrator with the same dataset.

Upon further investigation, it turns out that it only happened on datasets over 1 million rows, and the reason was that when porting the CalibratorUtils class from TLC, a "_maxCalibrationExamples = 1000000" default parameter was added.

Upon reading through the code (in particular CalibratorTrainingBase's ProcessingTrainingExample) it turns out that on TLC TrainCalibrator was called with maxRows = 0, and this made that when training the PlattCalibrator, all the dataset was seen, but only 1M rows where selected randomly to be added to the DataStore. In contrast, on ML.NET that same method was called with maxRows = 1M, and this made that only the first 1M rows were added to the DataStore (instead of randomly selecting them from the complete dataset). This caused bias and undesired results.

@antoniovs1029
antoniovs1029 requested a review from a team as a code ownerSeptember 30, 2020 08:18
@codecov

codecovBot commented Sep 30, 2020

Copy link
Copy Markdown

Codecov Report

Merging #5415 into master will decrease coverage by 0.06%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #5415 +/- ##
==========================================
- Coverage 74.08% 74.02% -0.07% 
==========================================
Files 1019 1019 Lines 190355 190363 +8 Branches 20469 20469 ==========================================
- Hits 141033 140914 -119 - Misses 43791 43905 +114 - Partials 5531 5544 +13 
FlagCoverage Δ
#Debug74.02% <ø> (-0.07%)⬇️
#production69.77% <ø> (-0.09%)⬇️
#test87.71% <ø> (+<0.01%)⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Impacted FilesCoverage Δ
src/Microsoft.ML.Data/Prediction/Calibrator.cs81.29% <ø> (+0.08%)⬆️
...osoft.ML.KMeansClustering/KMeansPlusPlusTrainer.cs83.60% <0.00%> (-7.27%)⬇️
src/Microsoft.ML.Data/Training/TrainerUtils.cs66.86% <0.00%> (-3.82%)⬇️
...crosoft.ML.StandardTrainers/Standard/SdcaBinary.cs85.23% <0.00%> (-3.33%)⬇️
...crosoft.ML.StandardTrainers/Optimizer/Optimizer.cs71.96% <0.00%> (-1.16%)⬇️
...oft.ML.StandardTrainers/Standard/SdcaMulticlass.cs91.46% <0.00%> (-1.03%)⬇️
src/Microsoft.ML.Data/Utils/LossFunctions.cs66.83% <0.00%> (-0.52%)⬇️
...StandardTrainers/Standard/LinearModelParameters.cs66.32% <0.00%> (-0.26%)⬇️
test/Microsoft.ML.Functional.Tests/ONNX.cs100.00% <0.00%> (ø)
test/Microsoft.ML.Tests/OnnxConversionTest.cs96.17% <0.00%> (+<0.01%)⬆️
... and 5 more

// maximum number of rows passed to the calibrator.
private const int _maxCalibrationExamples = 1000000;
// if 0, we'll actually look through the whole dataset to
// when training the calibrator

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Earlier, you had explained to me that if this value is zero, we would look through the whole dataset, but still only use a million rows (randomly selected) for calibration. Is that correct?

Can you please clarify the exact behavior in comments?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is correct for PlattCalibrator, but depending of the calibrator the behavior is different. If this is 0, the only thing that does happen for all the calibrators is that we'll look through all the dataset. I explained this further on the other comment I left, so I don't think it's necessary to clarify it more in here.

if (maxRows > 0 && ++num >= maxRows)
// If maxRows was 0, we'll process all of the rows in the dataset
// Notice that depending of the calibrator, "processing" means
// only using N random rows of the ones that where processed

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit, typo: "depending on the calibrator"

@harishskharishsk left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@antoniovs1029
antoniovs1029 merged commit 57be476 into dotnet:masterSep 30, 2020
frank-dong-ms-zz added a commit that referenced this pull request Oct 8, 2020
* Update to Onnxruntime 1.5.1 (#5406)
* Added variables to tests to control Gpu settings
* Added dependency to prerelease
* Updated to 1.5.1
* Remove prerelease feed
* Nit on GPU variables
* Change the _maxCalibrationExamples default on CalibratorUtils (#5415)
* Change the _maxCalibrationExamples default
* Improving comments
* Fix perf regression in ShuffleRows (#5417)
RowShufflingTransformer is using ChannelReader incorrectly. It needs to block waiting for items to read and was Thread.Sleeping in order to wait, but not spin the current core. This caused a major perf regression.
The fix is to block synchronously correctly - by calling AsTask() on the ValueTask that is returned from the ChannelReader and block on the Task.
Fix#5416
Co-authored-by: Antonio Velázquez <38739674+antoniovs1029@users.noreply.github.com>
Co-authored-by: Eric Erhardt <eric.erhardt@microsoft.com>
mstfbl pushed a commit to mstfbl/machinelearning that referenced this pull request Nov 12, 2020
* Update to Onnxruntime 1.5.1 (dotnet#5406)
* Added variables to tests to control Gpu settings
* Added dependency to prerelease
* Updated to 1.5.1
* Remove prerelease feed
* Nit on GPU variables
* Change the _maxCalibrationExamples default on CalibratorUtils (dotnet#5415)
* Change the _maxCalibrationExamples default
* Improving comments
* Fix perf regression in ShuffleRows (dotnet#5417)
RowShufflingTransformer is using ChannelReader incorrectly. It needs to block waiting for items to read and was Thread.Sleeping in order to wait, but not spin the current core. This caused a major perf regression.
The fix is to block synchronously correctly - by calling AsTask() on the ValueTask that is returned from the ChannelReader and block on the Task.
Fixdotnet#5416
Co-authored-by: Antonio Velázquez <38739674+antoniovs1029@users.noreply.github.com>
Co-authored-by: Eric Erhardt <eric.erhardt@microsoft.com>
mstfbl pushed a commit that referenced this pull request Nov 12, 2020
* Update to Onnxruntime 1.5.1 (#5406)
* Added variables to tests to control Gpu settings
* Added dependency to prerelease
* Updated to 1.5.1
* Remove prerelease feed
* Nit on GPU variables
* Change the _maxCalibrationExamples default on CalibratorUtils (#5415)
* Change the _maxCalibrationExamples default
* Improving comments
* Fix perf regression in ShuffleRows (#5417)
RowShufflingTransformer is using ChannelReader incorrectly. It needs to block waiting for items to read and was Thread.Sleeping in order to wait, but not spin the current core. This caused a major perf regression.
The fix is to block synchronously correctly - by calling AsTask() on the ValueTask that is returned from the ChannelReader and block on the Task.
Fix#5416
Co-authored-by: Antonio Velázquez <38739674+antoniovs1029@users.noreply.github.com>
Co-authored-by: Eric Erhardt <eric.erhardt@microsoft.com>
@ghostghost locked as resolved and limited conversation to collaborators Mar 17, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@antoniovs1029@harishsk