Debugging hanging AutoFitImageClassificationTrainTest - #4893

Closed
mstfbl wants to merge 28 commits into
dotnet:masterfrom
mstfbl:AutoFitTests-Debugging
Closed

Debugging hanging AutoFitImageClassificationTrainTest#4893
mstfbl wants to merge 28 commits into
dotnet:masterfrom
mstfbl:AutoFitTests-Debugging

Conversation

@mstfbl

@mstfblmstfbl commented Feb 26, 2020

Copy link
Copy Markdown
Contributor

Will be using this draft PR for general debugging purposes on CI

Notes:
Windows builds have 7,168 MBs of RAM

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

Testing of AutoFitImageClassificationTrainTest is taking too long per test, so tesitng it with 1000 iterations isn't feasible.

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

The tests AutoFitRecommendationTest and AutoFitRegressionTest are passing. AutoFitImageClassificationTrainTest is displaying errors every now and then.

@mstfbl

mstfbl commented Feb 26, 2020

Copy link
Copy Markdown
ContributorAuthor

The reason why AutoFitImageClassificationTrainTest is crashing is after running var result = context.Auto().CreateMulticlassClassificationExperiment(0).Execute(trainDataset, testDataset, columnInference.ColumnInformation), result.Best run might not always be updated. I saw this as I caught result.BestRun having a null value right when it is being called. Exact location where error is thrown.

@mstfbl

mstfbl commented Feb 27, 2020

Copy link
Copy Markdown
ContributorAuthor

The original bug with AutoFitImageClassificationTrainTest is occuring due to the null returned value here:


For some reason, sometimes the validationMetrics of an IEnumerable<(RunDetail) is null.

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

There's an issue with this Evaluation function:

publicMulticlassClassificationMetricsEvaluate(IDataViewdata,stringlabel,stringscore,stringpredictedLabel)
{
Host.CheckValue(data,nameof(data));
Host.CheckNonEmpty(label,nameof(label));
Host.CheckNonEmpty(score,nameof(score));
Host.CheckNonEmpty(predictedLabel,nameof(predictedLabel));
varroles=newRoleMappedData(data,opt:false,
RoleMappedSchema.ColumnRole.Label.Bind(label),
RoleMappedSchema.CreatePair(AnnotationUtils.Const.ScoreValueKind.Score,score),
RoleMappedSchema.CreatePair(AnnotationUtils.Const.ScoreValueKind.PredictedLabel,predictedLabel));
varresultDict=((IEvaluator)this).Evaluate(roles);
Host.Assert(resultDict.ContainsKey(MetricKinds.OverallMetrics));
varoverall=resultDict[MetricKinds.OverallMetrics];
varconfusionMatrix=resultDict[MetricKinds.ConfusionMatrix];
MulticlassClassificationMetricsresult;
using(varcursor=overall.GetRowCursorForAllColumns())
{
varmoved=cursor.MoveNext();
Host.Assert(moved);
result=newMulticlassClassificationMetrics(Host,cursor,_outputTopKAcc??0,confusionMatrix);
moved=cursor.MoveNext();
Host.Assert(!moved);
}
returnresult;
}
}

The returned result value can sometimes be null, which is what is causing AutoFitImageClassificationTrainTest to sometimes fail.

@mstfbl

mstfbl commented Feb 28, 2020

Copy link
Copy Markdown
ContributorAuthor

I figured out the cause of the occasional crash of AutoFitImageClassificationTrainTest. When any exception occurs in RunnerUtil.TrainAndScorePipeline, instead of throwing the error, it is instead caught and ignored while a null metrics value (in line 49) is sent up through the call stack instead.

try
{
varestimator=pipeline.ToEstimator(trainData,validData);
varmodel=estimator.Fit(trainData);
varscoredData=model.Transform(validData);
varmetrics=metricsAgent.EvaluateMetrics(scoredData,labelColumn);
varscore=metricsAgent.GetScore(metrics);
if(preprocessorTransform!=null)
{
model=preprocessorTransform.Append(model);
}
// Build container for model
varmodelContainer=modelFileInfo==null?
newModelContainer(context,model):
newModelContainer(context,modelFileInfo,model,modelInputSchema);
return(modelContainer,metrics,null,score);
}
catch(Exceptionex)
{
logger.Error($"Pipeline crashed: {pipeline.ToString()} . Exception: {ex}");
return(null,null,ex,double.NaN);

This is the exception being caught:

System.ArgumentException : PIPELINE CRASHES - Line 55 - RunnerUtil.cs - Pipeline crash string: xf=ValueToKeyMapping{ col=Label:Label} xf=RawByteImageLoading{ col=ImagePath_featurized:ImagePath imageFolder=} xf=ColumnCopying{ col=Features:ImagePath_featurized} tr=ImageClassification{} xf=KeyToValueMapping{ col=PredictedLabel:PredictedLabel} cache=- - Exception String: System.FormatException: Tensorflow exception triggered while loading model. ---> System.Runtime.InteropServices.SEHException: External component has thrown an exception.

When reproduced locally, the exception string is:

"Could not find a part of the path 'C:\Users\mubal\AppData\Local\Temp\Microsoft.ML.AutoML\experiment_y1gbdyum.xdu\Model1.zip'."

@mstfbl

mstfbl commented Mar 2, 2020

Copy link
Copy Markdown
ContributorAuthor

Update: AutoFitImageClassificationTrainTest with 100 iterations fail on Windows x64 builds with:

System.Runtime.InteropServices.SEHException (0x80004005): External component has thrown an exception

but also I get:
System.FormatException: Tensorflow exception triggered while loading model. ---> System.OutOfMemoryException
and:
System.AccessViolationException: Attempted to read or write protected memory. This is often an indication that other memory is corrupt.

in Tensorflow.c_api.TF_SessionRun, which is the C++ implementation of TensorFlow's training code. This seems related to Issue SciSharp/TensorFlow.NET#485

As mentioned in this PR #4755, we still cannot see details about the crash in Tensorflow.c_api.TF_SessionRun.

@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from 4520530 to 65a72efCompareMarch 19, 2020 05:17
@mstfblmstfbl closed this Mar 20, 2020
@mstfblmstfbl changed the title Auto fit tests debuggingDebugging PRMar 22, 2020
@mstfblmstfbl reopened this Mar 22, 2020
@mstfblmstfbl closed this Mar 22, 2020
@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from c1f8231 to c1e422dCompareMarch 22, 2020 08:39
@mstfblmstfbl reopened this Mar 22, 2020
@mstfblmstfbl changed the title Debugging PRDebugging hanging AutoFitImageClassificationTrainTestMar 26, 2020
@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from df5f642 to 5a7ad17CompareMarch 26, 2020 03:32
@mstfbl

Copy link
Copy Markdown
ContributorAuthor

Will be using this PR to debug AutoFitImageClassificationTrainTest hanging occasionally on Windows builds.

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

AutoFitImageClassificationTrainTest is still occasionally hanging, mostly due to indisposed Tensorflow objects after the test is complete. I found this comment in Microsoft.ML.Vision/ImageClassificationTrainer.TrainModelCore to be of interest:

// Leave the ownership of _session so that it is not disposed/closed when this object goes out of scope
// since it will be used by ImageClassificationModelParameters class (new owner that will take care of
// disposing).
varsession=_session;
_session=null;
returnnewImageClassificationModelParameters(Host,session,_classCount,_jpegDataTensorName,
_resizedImageTensorName,_inputTensorName,_softmaxTensorName);

@mstfbl

mstfbl commented Mar 27, 2020

Copy link
Copy Markdown
ContributorAuthor

Adding the fix (model as IDisposable)?.Dispose(); for freeing Tensor objects worked! This fix is necessary, as these Tensor objects made in the C TensorFlow libraries are not automatically cleaned up by C#'s Garbage Collector.

Edit: While this fix works, it is not safe to assume that this model can be disposed in RunnerUtil.cs. The user might be accessing this model during disposal, which would result in use-after-free and/or null reference errors.

@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from 78a65c2 to 087c0d5CompareApril 16, 2020 05:13
@dotnetdotnet deleted a comment from azure-pipelinesBotApr 16, 2020
@mstfbl

Copy link
Copy Markdown
ContributorAuthor

Freeing Tensor objects in model in a finally statement in TrainAndScorePipeline works in fixing memory bug, and is safe to do when model is never saved in memory and written to disk always.

@mstfblmstfbl closed this Apr 26, 2020
@ghostghost locked as resolved and limited conversation to collaborators Mar 19, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@mstfbl
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Debugging hanging AutoFitImageClassificationTrainTest - #4893

Closed
mstfbl wants to merge 28 commits into
dotnet:masterfrom
mstfbl:AutoFitTests-Debugging
Closed

Debugging hanging AutoFitImageClassificationTrainTest#4893
mstfbl wants to merge 28 commits into
dotnet:masterfrom
mstfbl:AutoFitTests-Debugging

Conversation

@mstfbl

@mstfblmstfbl commented Feb 26, 2020

Copy link
Copy Markdown
Contributor

Will be using this draft PR for general debugging purposes on CI

Notes:
Windows builds have 7,168 MBs of RAM

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

Testing of AutoFitImageClassificationTrainTest is taking too long per test, so tesitng it with 1000 iterations isn't feasible.

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

The tests AutoFitRecommendationTest and AutoFitRegressionTest are passing. AutoFitImageClassificationTrainTest is displaying errors every now and then.

@mstfbl

mstfbl commented Feb 26, 2020

Copy link
Copy Markdown
ContributorAuthor

The reason why AutoFitImageClassificationTrainTest is crashing is after running var result = context.Auto().CreateMulticlassClassificationExperiment(0).Execute(trainDataset, testDataset, columnInference.ColumnInformation), result.Best run might not always be updated. I saw this as I caught result.BestRun having a null value right when it is being called. Exact location where error is thrown.

@mstfbl

mstfbl commented Feb 27, 2020

Copy link
Copy Markdown
ContributorAuthor

The original bug with AutoFitImageClassificationTrainTest is occuring due to the null returned value here:


For some reason, sometimes the validationMetrics of an IEnumerable<(RunDetail) is null.

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

There's an issue with this Evaluation function:

publicMulticlassClassificationMetricsEvaluate(IDataViewdata,stringlabel,stringscore,stringpredictedLabel)
{
Host.CheckValue(data,nameof(data));
Host.CheckNonEmpty(label,nameof(label));
Host.CheckNonEmpty(score,nameof(score));
Host.CheckNonEmpty(predictedLabel,nameof(predictedLabel));
varroles=newRoleMappedData(data,opt:false,
RoleMappedSchema.ColumnRole.Label.Bind(label),
RoleMappedSchema.CreatePair(AnnotationUtils.Const.ScoreValueKind.Score,score),
RoleMappedSchema.CreatePair(AnnotationUtils.Const.ScoreValueKind.PredictedLabel,predictedLabel));
varresultDict=((IEvaluator)this).Evaluate(roles);
Host.Assert(resultDict.ContainsKey(MetricKinds.OverallMetrics));
varoverall=resultDict[MetricKinds.OverallMetrics];
varconfusionMatrix=resultDict[MetricKinds.ConfusionMatrix];
MulticlassClassificationMetricsresult;
using(varcursor=overall.GetRowCursorForAllColumns())
{
varmoved=cursor.MoveNext();
Host.Assert(moved);
result=newMulticlassClassificationMetrics(Host,cursor,_outputTopKAcc??0,confusionMatrix);
moved=cursor.MoveNext();
Host.Assert(!moved);
}
returnresult;
}
}

The returned result value can sometimes be null, which is what is causing AutoFitImageClassificationTrainTest to sometimes fail.

@mstfbl

mstfbl commented Feb 28, 2020

Copy link
Copy Markdown
ContributorAuthor

I figured out the cause of the occasional crash of AutoFitImageClassificationTrainTest. When any exception occurs in RunnerUtil.TrainAndScorePipeline, instead of throwing the error, it is instead caught and ignored while a null metrics value (in line 49) is sent up through the call stack instead.

try
{
varestimator=pipeline.ToEstimator(trainData,validData);
varmodel=estimator.Fit(trainData);
varscoredData=model.Transform(validData);
varmetrics=metricsAgent.EvaluateMetrics(scoredData,labelColumn);
varscore=metricsAgent.GetScore(metrics);
if(preprocessorTransform!=null)
{
model=preprocessorTransform.Append(model);
}
// Build container for model
varmodelContainer=modelFileInfo==null?
newModelContainer(context,model):
newModelContainer(context,modelFileInfo,model,modelInputSchema);
return(modelContainer,metrics,null,score);
}
catch(Exceptionex)
{
logger.Error($"Pipeline crashed: {pipeline.ToString()} . Exception: {ex}");
return(null,null,ex,double.NaN);

This is the exception being caught:

System.ArgumentException : PIPELINE CRASHES - Line 55 - RunnerUtil.cs - Pipeline crash string: xf=ValueToKeyMapping{ col=Label:Label} xf=RawByteImageLoading{ col=ImagePath_featurized:ImagePath imageFolder=} xf=ColumnCopying{ col=Features:ImagePath_featurized} tr=ImageClassification{} xf=KeyToValueMapping{ col=PredictedLabel:PredictedLabel} cache=- - Exception String: System.FormatException: Tensorflow exception triggered while loading model. ---> System.Runtime.InteropServices.SEHException: External component has thrown an exception.

When reproduced locally, the exception string is:

"Could not find a part of the path 'C:\Users\mubal\AppData\Local\Temp\Microsoft.ML.AutoML\experiment_y1gbdyum.xdu\Model1.zip'."

@mstfbl

mstfbl commented Mar 2, 2020

Copy link
Copy Markdown
ContributorAuthor

Update: AutoFitImageClassificationTrainTest with 100 iterations fail on Windows x64 builds with:

System.Runtime.InteropServices.SEHException (0x80004005): External component has thrown an exception

but also I get:
System.FormatException: Tensorflow exception triggered while loading model. ---> System.OutOfMemoryException
and:
System.AccessViolationException: Attempted to read or write protected memory. This is often an indication that other memory is corrupt.

in Tensorflow.c_api.TF_SessionRun, which is the C++ implementation of TensorFlow's training code. This seems related to Issue SciSharp/TensorFlow.NET#485

As mentioned in this PR #4755, we still cannot see details about the crash in Tensorflow.c_api.TF_SessionRun.

@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from 4520530 to 65a72efCompareMarch 19, 2020 05:17
@mstfblmstfbl closed this Mar 20, 2020
@mstfblmstfbl changed the title Auto fit tests debuggingDebugging PRMar 22, 2020
@mstfblmstfbl reopened this Mar 22, 2020
@mstfblmstfbl closed this Mar 22, 2020
@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from c1f8231 to c1e422dCompareMarch 22, 2020 08:39
@mstfblmstfbl reopened this Mar 22, 2020
@mstfblmstfbl changed the title Debugging PRDebugging hanging AutoFitImageClassificationTrainTestMar 26, 2020
@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from df5f642 to 5a7ad17CompareMarch 26, 2020 03:32
@mstfbl

Copy link
Copy Markdown
ContributorAuthor

Will be using this PR to debug AutoFitImageClassificationTrainTest hanging occasionally on Windows builds.

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

AutoFitImageClassificationTrainTest is still occasionally hanging, mostly due to indisposed Tensorflow objects after the test is complete. I found this comment in Microsoft.ML.Vision/ImageClassificationTrainer.TrainModelCore to be of interest:

// Leave the ownership of _session so that it is not disposed/closed when this object goes out of scope
// since it will be used by ImageClassificationModelParameters class (new owner that will take care of
// disposing).
varsession=_session;
_session=null;
returnnewImageClassificationModelParameters(Host,session,_classCount,_jpegDataTensorName,
_resizedImageTensorName,_inputTensorName,_softmaxTensorName);

@mstfbl

mstfbl commented Mar 27, 2020

Copy link
Copy Markdown
ContributorAuthor

Adding the fix (model as IDisposable)?.Dispose(); for freeing Tensor objects worked! This fix is necessary, as these Tensor objects made in the C TensorFlow libraries are not automatically cleaned up by C#'s Garbage Collector.

Edit: While this fix works, it is not safe to assume that this model can be disposed in RunnerUtil.cs. The user might be accessing this model during disposal, which would result in use-after-free and/or null reference errors.

@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from 78a65c2 to 087c0d5CompareApril 16, 2020 05:13
@dotnetdotnet deleted a comment from azure-pipelinesBotApr 16, 2020
@mstfbl

Copy link
Copy Markdown
ContributorAuthor

Freeing Tensor objects in model in a finally statement in TrainAndScorePipeline works in fixing memory bug, and is safe to do when model is never saved in memory and written to disk always.

@mstfblmstfbl closed this Apr 26, 2020
@ghostghost locked as resolved and limited conversation to collaborators Mar 19, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@mstfbl
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Debugging hanging AutoFitImageClassificationTrainTest - #4893

Closed
mstfbl wants to merge 28 commits into
dotnet:masterfrom
mstfbl:AutoFitTests-Debugging
Closed

Debugging hanging AutoFitImageClassificationTrainTest#4893
mstfbl wants to merge 28 commits into
dotnet:masterfrom
mstfbl:AutoFitTests-Debugging

Conversation

@mstfbl

@mstfblmstfbl commented Feb 26, 2020

Copy link
Copy Markdown
Contributor

Will be using this draft PR for general debugging purposes on CI

Notes:
Windows builds have 7,168 MBs of RAM

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

Testing of AutoFitImageClassificationTrainTest is taking too long per test, so tesitng it with 1000 iterations isn't feasible.

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

The tests AutoFitRecommendationTest and AutoFitRegressionTest are passing. AutoFitImageClassificationTrainTest is displaying errors every now and then.

@mstfbl

mstfbl commented Feb 26, 2020

Copy link
Copy Markdown
ContributorAuthor

The reason why AutoFitImageClassificationTrainTest is crashing is after running var result = context.Auto().CreateMulticlassClassificationExperiment(0).Execute(trainDataset, testDataset, columnInference.ColumnInformation), result.Best run might not always be updated. I saw this as I caught result.BestRun having a null value right when it is being called. Exact location where error is thrown.

@mstfbl

mstfbl commented Feb 27, 2020

Copy link
Copy Markdown
ContributorAuthor

The original bug with AutoFitImageClassificationTrainTest is occuring due to the null returned value here:


For some reason, sometimes the validationMetrics of an IEnumerable<(RunDetail) is null.

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

There's an issue with this Evaluation function:

publicMulticlassClassificationMetricsEvaluate(IDataViewdata,stringlabel,stringscore,stringpredictedLabel)
{
Host.CheckValue(data,nameof(data));
Host.CheckNonEmpty(label,nameof(label));
Host.CheckNonEmpty(score,nameof(score));
Host.CheckNonEmpty(predictedLabel,nameof(predictedLabel));
varroles=newRoleMappedData(data,opt:false,
RoleMappedSchema.ColumnRole.Label.Bind(label),
RoleMappedSchema.CreatePair(AnnotationUtils.Const.ScoreValueKind.Score,score),
RoleMappedSchema.CreatePair(AnnotationUtils.Const.ScoreValueKind.PredictedLabel,predictedLabel));
varresultDict=((IEvaluator)this).Evaluate(roles);
Host.Assert(resultDict.ContainsKey(MetricKinds.OverallMetrics));
varoverall=resultDict[MetricKinds.OverallMetrics];
varconfusionMatrix=resultDict[MetricKinds.ConfusionMatrix];
MulticlassClassificationMetricsresult;
using(varcursor=overall.GetRowCursorForAllColumns())
{
varmoved=cursor.MoveNext();
Host.Assert(moved);
result=newMulticlassClassificationMetrics(Host,cursor,_outputTopKAcc??0,confusionMatrix);
moved=cursor.MoveNext();
Host.Assert(!moved);
}
returnresult;
}
}

The returned result value can sometimes be null, which is what is causing AutoFitImageClassificationTrainTest to sometimes fail.

@mstfbl

mstfbl commented Feb 28, 2020

Copy link
Copy Markdown
ContributorAuthor

I figured out the cause of the occasional crash of AutoFitImageClassificationTrainTest. When any exception occurs in RunnerUtil.TrainAndScorePipeline, instead of throwing the error, it is instead caught and ignored while a null metrics value (in line 49) is sent up through the call stack instead.

try
{
varestimator=pipeline.ToEstimator(trainData,validData);
varmodel=estimator.Fit(trainData);
varscoredData=model.Transform(validData);
varmetrics=metricsAgent.EvaluateMetrics(scoredData,labelColumn);
varscore=metricsAgent.GetScore(metrics);
if(preprocessorTransform!=null)
{
model=preprocessorTransform.Append(model);
}
// Build container for model
varmodelContainer=modelFileInfo==null?
newModelContainer(context,model):
newModelContainer(context,modelFileInfo,model,modelInputSchema);
return(modelContainer,metrics,null,score);
}
catch(Exceptionex)
{
logger.Error($"Pipeline crashed: {pipeline.ToString()} . Exception: {ex}");
return(null,null,ex,double.NaN);

This is the exception being caught:

System.ArgumentException : PIPELINE CRASHES - Line 55 - RunnerUtil.cs - Pipeline crash string: xf=ValueToKeyMapping{ col=Label:Label} xf=RawByteImageLoading{ col=ImagePath_featurized:ImagePath imageFolder=} xf=ColumnCopying{ col=Features:ImagePath_featurized} tr=ImageClassification{} xf=KeyToValueMapping{ col=PredictedLabel:PredictedLabel} cache=- - Exception String: System.FormatException: Tensorflow exception triggered while loading model. ---> System.Runtime.InteropServices.SEHException: External component has thrown an exception.

When reproduced locally, the exception string is:

"Could not find a part of the path 'C:\Users\mubal\AppData\Local\Temp\Microsoft.ML.AutoML\experiment_y1gbdyum.xdu\Model1.zip'."

@mstfbl

mstfbl commented Mar 2, 2020

Copy link
Copy Markdown
ContributorAuthor

Update: AutoFitImageClassificationTrainTest with 100 iterations fail on Windows x64 builds with:

System.Runtime.InteropServices.SEHException (0x80004005): External component has thrown an exception

but also I get:
System.FormatException: Tensorflow exception triggered while loading model. ---> System.OutOfMemoryException
and:
System.AccessViolationException: Attempted to read or write protected memory. This is often an indication that other memory is corrupt.

in Tensorflow.c_api.TF_SessionRun, which is the C++ implementation of TensorFlow's training code. This seems related to Issue SciSharp/TensorFlow.NET#485

As mentioned in this PR #4755, we still cannot see details about the crash in Tensorflow.c_api.TF_SessionRun.

@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from 4520530 to 65a72efCompareMarch 19, 2020 05:17
@mstfblmstfbl closed this Mar 20, 2020
@mstfblmstfbl changed the title Auto fit tests debuggingDebugging PRMar 22, 2020
@mstfblmstfbl reopened this Mar 22, 2020
@mstfblmstfbl closed this Mar 22, 2020
@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from c1f8231 to c1e422dCompareMarch 22, 2020 08:39
@mstfblmstfbl reopened this Mar 22, 2020
@mstfblmstfbl changed the title Debugging PRDebugging hanging AutoFitImageClassificationTrainTestMar 26, 2020
@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from df5f642 to 5a7ad17CompareMarch 26, 2020 03:32
@mstfbl

Copy link
Copy Markdown
ContributorAuthor

Will be using this PR to debug AutoFitImageClassificationTrainTest hanging occasionally on Windows builds.

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

AutoFitImageClassificationTrainTest is still occasionally hanging, mostly due to indisposed Tensorflow objects after the test is complete. I found this comment in Microsoft.ML.Vision/ImageClassificationTrainer.TrainModelCore to be of interest:

// Leave the ownership of _session so that it is not disposed/closed when this object goes out of scope
// since it will be used by ImageClassificationModelParameters class (new owner that will take care of
// disposing).
varsession=_session;
_session=null;
returnnewImageClassificationModelParameters(Host,session,_classCount,_jpegDataTensorName,
_resizedImageTensorName,_inputTensorName,_softmaxTensorName);

@mstfbl

mstfbl commented Mar 27, 2020

Copy link
Copy Markdown
ContributorAuthor

Adding the fix (model as IDisposable)?.Dispose(); for freeing Tensor objects worked! This fix is necessary, as these Tensor objects made in the C TensorFlow libraries are not automatically cleaned up by C#'s Garbage Collector.

Edit: While this fix works, it is not safe to assume that this model can be disposed in RunnerUtil.cs. The user might be accessing this model during disposal, which would result in use-after-free and/or null reference errors.

@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from 78a65c2 to 087c0d5CompareApril 16, 2020 05:13
@dotnetdotnet deleted a comment from azure-pipelinesBotApr 16, 2020
@mstfbl

Copy link
Copy Markdown
ContributorAuthor

Freeing Tensor objects in model in a finally statement in TrainAndScorePipeline works in fixing memory bug, and is safe to do when model is never saved in memory and written to disk always.

@mstfblmstfbl closed this Apr 26, 2020
@ghostghost locked as resolved and limited conversation to collaborators Mar 19, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@mstfbl
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Debugging hanging AutoFitImageClassificationTrainTest - #4893

Closed
mstfbl wants to merge 28 commits into
dotnet:masterfrom
mstfbl:AutoFitTests-Debugging
Closed

Debugging hanging AutoFitImageClassificationTrainTest#4893
mstfbl wants to merge 28 commits into
dotnet:masterfrom
mstfbl:AutoFitTests-Debugging

Conversation

@mstfbl

@mstfblmstfbl commented Feb 26, 2020

Copy link
Copy Markdown
Contributor

Will be using this draft PR for general debugging purposes on CI

Notes:
Windows builds have 7,168 MBs of RAM

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

Testing of AutoFitImageClassificationTrainTest is taking too long per test, so tesitng it with 1000 iterations isn't feasible.

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

The tests AutoFitRecommendationTest and AutoFitRegressionTest are passing. AutoFitImageClassificationTrainTest is displaying errors every now and then.

@mstfbl

mstfbl commented Feb 26, 2020

Copy link
Copy Markdown
ContributorAuthor

The reason why AutoFitImageClassificationTrainTest is crashing is after running var result = context.Auto().CreateMulticlassClassificationExperiment(0).Execute(trainDataset, testDataset, columnInference.ColumnInformation), result.Best run might not always be updated. I saw this as I caught result.BestRun having a null value right when it is being called. Exact location where error is thrown.

@mstfbl

mstfbl commented Feb 27, 2020

Copy link
Copy Markdown
ContributorAuthor

The original bug with AutoFitImageClassificationTrainTest is occuring due to the null returned value here:


For some reason, sometimes the validationMetrics of an IEnumerable<(RunDetail) is null.

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

There's an issue with this Evaluation function:

publicMulticlassClassificationMetricsEvaluate(IDataViewdata,stringlabel,stringscore,stringpredictedLabel)
{
Host.CheckValue(data,nameof(data));
Host.CheckNonEmpty(label,nameof(label));
Host.CheckNonEmpty(score,nameof(score));
Host.CheckNonEmpty(predictedLabel,nameof(predictedLabel));
varroles=newRoleMappedData(data,opt:false,
RoleMappedSchema.ColumnRole.Label.Bind(label),
RoleMappedSchema.CreatePair(AnnotationUtils.Const.ScoreValueKind.Score,score),
RoleMappedSchema.CreatePair(AnnotationUtils.Const.ScoreValueKind.PredictedLabel,predictedLabel));
varresultDict=((IEvaluator)this).Evaluate(roles);
Host.Assert(resultDict.ContainsKey(MetricKinds.OverallMetrics));
varoverall=resultDict[MetricKinds.OverallMetrics];
varconfusionMatrix=resultDict[MetricKinds.ConfusionMatrix];
MulticlassClassificationMetricsresult;
using(varcursor=overall.GetRowCursorForAllColumns())
{
varmoved=cursor.MoveNext();
Host.Assert(moved);
result=newMulticlassClassificationMetrics(Host,cursor,_outputTopKAcc??0,confusionMatrix);
moved=cursor.MoveNext();
Host.Assert(!moved);
}
returnresult;
}
}

The returned result value can sometimes be null, which is what is causing AutoFitImageClassificationTrainTest to sometimes fail.

@mstfbl

mstfbl commented Feb 28, 2020

Copy link
Copy Markdown
ContributorAuthor

I figured out the cause of the occasional crash of AutoFitImageClassificationTrainTest. When any exception occurs in RunnerUtil.TrainAndScorePipeline, instead of throwing the error, it is instead caught and ignored while a null metrics value (in line 49) is sent up through the call stack instead.

try
{
varestimator=pipeline.ToEstimator(trainData,validData);
varmodel=estimator.Fit(trainData);
varscoredData=model.Transform(validData);
varmetrics=metricsAgent.EvaluateMetrics(scoredData,labelColumn);
varscore=metricsAgent.GetScore(metrics);
if(preprocessorTransform!=null)
{
model=preprocessorTransform.Append(model);
}
// Build container for model
varmodelContainer=modelFileInfo==null?
newModelContainer(context,model):
newModelContainer(context,modelFileInfo,model,modelInputSchema);
return(modelContainer,metrics,null,score);
}
catch(Exceptionex)
{
logger.Error($"Pipeline crashed: {pipeline.ToString()} . Exception: {ex}");
return(null,null,ex,double.NaN);

This is the exception being caught:

System.ArgumentException : PIPELINE CRASHES - Line 55 - RunnerUtil.cs - Pipeline crash string: xf=ValueToKeyMapping{ col=Label:Label} xf=RawByteImageLoading{ col=ImagePath_featurized:ImagePath imageFolder=} xf=ColumnCopying{ col=Features:ImagePath_featurized} tr=ImageClassification{} xf=KeyToValueMapping{ col=PredictedLabel:PredictedLabel} cache=- - Exception String: System.FormatException: Tensorflow exception triggered while loading model. ---> System.Runtime.InteropServices.SEHException: External component has thrown an exception.

When reproduced locally, the exception string is:

"Could not find a part of the path 'C:\Users\mubal\AppData\Local\Temp\Microsoft.ML.AutoML\experiment_y1gbdyum.xdu\Model1.zip'."

@mstfbl

mstfbl commented Mar 2, 2020

Copy link
Copy Markdown
ContributorAuthor

Update: AutoFitImageClassificationTrainTest with 100 iterations fail on Windows x64 builds with:

System.Runtime.InteropServices.SEHException (0x80004005): External component has thrown an exception

but also I get:
System.FormatException: Tensorflow exception triggered while loading model. ---> System.OutOfMemoryException
and:
System.AccessViolationException: Attempted to read or write protected memory. This is often an indication that other memory is corrupt.

in Tensorflow.c_api.TF_SessionRun, which is the C++ implementation of TensorFlow's training code. This seems related to Issue SciSharp/TensorFlow.NET#485

As mentioned in this PR #4755, we still cannot see details about the crash in Tensorflow.c_api.TF_SessionRun.

@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from 4520530 to 65a72efCompareMarch 19, 2020 05:17
@mstfblmstfbl closed this Mar 20, 2020
@mstfblmstfbl changed the title Auto fit tests debuggingDebugging PRMar 22, 2020
@mstfblmstfbl reopened this Mar 22, 2020
@mstfblmstfbl closed this Mar 22, 2020
@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from c1f8231 to c1e422dCompareMarch 22, 2020 08:39
@mstfblmstfbl reopened this Mar 22, 2020
@mstfblmstfbl changed the title Debugging PRDebugging hanging AutoFitImageClassificationTrainTestMar 26, 2020
@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from df5f642 to 5a7ad17CompareMarch 26, 2020 03:32
@mstfbl

Copy link
Copy Markdown
ContributorAuthor

Will be using this PR to debug AutoFitImageClassificationTrainTest hanging occasionally on Windows builds.

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

AutoFitImageClassificationTrainTest is still occasionally hanging, mostly due to indisposed Tensorflow objects after the test is complete. I found this comment in Microsoft.ML.Vision/ImageClassificationTrainer.TrainModelCore to be of interest:

// Leave the ownership of _session so that it is not disposed/closed when this object goes out of scope
// since it will be used by ImageClassificationModelParameters class (new owner that will take care of
// disposing).
varsession=_session;
_session=null;
returnnewImageClassificationModelParameters(Host,session,_classCount,_jpegDataTensorName,
_resizedImageTensorName,_inputTensorName,_softmaxTensorName);

@mstfbl

mstfbl commented Mar 27, 2020

Copy link
Copy Markdown
ContributorAuthor

Adding the fix (model as IDisposable)?.Dispose(); for freeing Tensor objects worked! This fix is necessary, as these Tensor objects made in the C TensorFlow libraries are not automatically cleaned up by C#'s Garbage Collector.

Edit: While this fix works, it is not safe to assume that this model can be disposed in RunnerUtil.cs. The user might be accessing this model during disposal, which would result in use-after-free and/or null reference errors.

@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from 78a65c2 to 087c0d5CompareApril 16, 2020 05:13
@dotnetdotnet deleted a comment from azure-pipelinesBotApr 16, 2020
@mstfbl

Copy link
Copy Markdown
ContributorAuthor

Freeing Tensor objects in model in a finally statement in TrainAndScorePipeline works in fixing memory bug, and is safe to do when model is never saved in memory and written to disk always.

@mstfblmstfbl closed this Apr 26, 2020
@ghostghost locked as resolved and limited conversation to collaborators Mar 19, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@mstfbl
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Debugging hanging AutoFitImageClassificationTrainTest - #4893

Closed
mstfbl wants to merge 28 commits into
dotnet:masterfrom
mstfbl:AutoFitTests-Debugging
Closed

Debugging hanging AutoFitImageClassificationTrainTest#4893
mstfbl wants to merge 28 commits into
dotnet:masterfrom
mstfbl:AutoFitTests-Debugging

Conversation

@mstfbl

@mstfblmstfbl commented Feb 26, 2020

Copy link
Copy Markdown
Contributor

Will be using this draft PR for general debugging purposes on CI

Notes:
Windows builds have 7,168 MBs of RAM

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

Testing of AutoFitImageClassificationTrainTest is taking too long per test, so tesitng it with 1000 iterations isn't feasible.

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

The tests AutoFitRecommendationTest and AutoFitRegressionTest are passing. AutoFitImageClassificationTrainTest is displaying errors every now and then.

@mstfbl

mstfbl commented Feb 26, 2020

Copy link
Copy Markdown
ContributorAuthor

The reason why AutoFitImageClassificationTrainTest is crashing is after running var result = context.Auto().CreateMulticlassClassificationExperiment(0).Execute(trainDataset, testDataset, columnInference.ColumnInformation), result.Best run might not always be updated. I saw this as I caught result.BestRun having a null value right when it is being called. Exact location where error is thrown.

@mstfbl

mstfbl commented Feb 27, 2020

Copy link
Copy Markdown
ContributorAuthor

The original bug with AutoFitImageClassificationTrainTest is occuring due to the null returned value here:


For some reason, sometimes the validationMetrics of an IEnumerable<(RunDetail) is null.

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

There's an issue with this Evaluation function:

publicMulticlassClassificationMetricsEvaluate(IDataViewdata,stringlabel,stringscore,stringpredictedLabel)
{
Host.CheckValue(data,nameof(data));
Host.CheckNonEmpty(label,nameof(label));
Host.CheckNonEmpty(score,nameof(score));
Host.CheckNonEmpty(predictedLabel,nameof(predictedLabel));
varroles=newRoleMappedData(data,opt:false,
RoleMappedSchema.ColumnRole.Label.Bind(label),
RoleMappedSchema.CreatePair(AnnotationUtils.Const.ScoreValueKind.Score,score),
RoleMappedSchema.CreatePair(AnnotationUtils.Const.ScoreValueKind.PredictedLabel,predictedLabel));
varresultDict=((IEvaluator)this).Evaluate(roles);
Host.Assert(resultDict.ContainsKey(MetricKinds.OverallMetrics));
varoverall=resultDict[MetricKinds.OverallMetrics];
varconfusionMatrix=resultDict[MetricKinds.ConfusionMatrix];
MulticlassClassificationMetricsresult;
using(varcursor=overall.GetRowCursorForAllColumns())
{
varmoved=cursor.MoveNext();
Host.Assert(moved);
result=newMulticlassClassificationMetrics(Host,cursor,_outputTopKAcc??0,confusionMatrix);
moved=cursor.MoveNext();
Host.Assert(!moved);
}
returnresult;
}
}

The returned result value can sometimes be null, which is what is causing AutoFitImageClassificationTrainTest to sometimes fail.

@mstfbl

mstfbl commented Feb 28, 2020

Copy link
Copy Markdown
ContributorAuthor

I figured out the cause of the occasional crash of AutoFitImageClassificationTrainTest. When any exception occurs in RunnerUtil.TrainAndScorePipeline, instead of throwing the error, it is instead caught and ignored while a null metrics value (in line 49) is sent up through the call stack instead.

try
{
varestimator=pipeline.ToEstimator(trainData,validData);
varmodel=estimator.Fit(trainData);
varscoredData=model.Transform(validData);
varmetrics=metricsAgent.EvaluateMetrics(scoredData,labelColumn);
varscore=metricsAgent.GetScore(metrics);
if(preprocessorTransform!=null)
{
model=preprocessorTransform.Append(model);
}
// Build container for model
varmodelContainer=modelFileInfo==null?
newModelContainer(context,model):
newModelContainer(context,modelFileInfo,model,modelInputSchema);
return(modelContainer,metrics,null,score);
}
catch(Exceptionex)
{
logger.Error($"Pipeline crashed: {pipeline.ToString()} . Exception: {ex}");
return(null,null,ex,double.NaN);

This is the exception being caught:

System.ArgumentException : PIPELINE CRASHES - Line 55 - RunnerUtil.cs - Pipeline crash string: xf=ValueToKeyMapping{ col=Label:Label} xf=RawByteImageLoading{ col=ImagePath_featurized:ImagePath imageFolder=} xf=ColumnCopying{ col=Features:ImagePath_featurized} tr=ImageClassification{} xf=KeyToValueMapping{ col=PredictedLabel:PredictedLabel} cache=- - Exception String: System.FormatException: Tensorflow exception triggered while loading model. ---> System.Runtime.InteropServices.SEHException: External component has thrown an exception.

When reproduced locally, the exception string is:

"Could not find a part of the path 'C:\Users\mubal\AppData\Local\Temp\Microsoft.ML.AutoML\experiment_y1gbdyum.xdu\Model1.zip'."

@mstfbl

mstfbl commented Mar 2, 2020

Copy link
Copy Markdown
ContributorAuthor

Update: AutoFitImageClassificationTrainTest with 100 iterations fail on Windows x64 builds with:

System.Runtime.InteropServices.SEHException (0x80004005): External component has thrown an exception

but also I get:
System.FormatException: Tensorflow exception triggered while loading model. ---> System.OutOfMemoryException
and:
System.AccessViolationException: Attempted to read or write protected memory. This is often an indication that other memory is corrupt.

in Tensorflow.c_api.TF_SessionRun, which is the C++ implementation of TensorFlow's training code. This seems related to Issue SciSharp/TensorFlow.NET#485

As mentioned in this PR #4755, we still cannot see details about the crash in Tensorflow.c_api.TF_SessionRun.

@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from 4520530 to 65a72efCompareMarch 19, 2020 05:17
@mstfblmstfbl closed this Mar 20, 2020
@mstfblmstfbl changed the title Auto fit tests debuggingDebugging PRMar 22, 2020
@mstfblmstfbl reopened this Mar 22, 2020
@mstfblmstfbl closed this Mar 22, 2020
@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from c1f8231 to c1e422dCompareMarch 22, 2020 08:39
@mstfblmstfbl reopened this Mar 22, 2020
@mstfblmstfbl changed the title Debugging PRDebugging hanging AutoFitImageClassificationTrainTestMar 26, 2020
@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from df5f642 to 5a7ad17CompareMarch 26, 2020 03:32
@mstfbl

Copy link
Copy Markdown
ContributorAuthor

Will be using this PR to debug AutoFitImageClassificationTrainTest hanging occasionally on Windows builds.

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

AutoFitImageClassificationTrainTest is still occasionally hanging, mostly due to indisposed Tensorflow objects after the test is complete. I found this comment in Microsoft.ML.Vision/ImageClassificationTrainer.TrainModelCore to be of interest:

// Leave the ownership of _session so that it is not disposed/closed when this object goes out of scope
// since it will be used by ImageClassificationModelParameters class (new owner that will take care of
// disposing).
varsession=_session;
_session=null;
returnnewImageClassificationModelParameters(Host,session,_classCount,_jpegDataTensorName,
_resizedImageTensorName,_inputTensorName,_softmaxTensorName);

@mstfbl

mstfbl commented Mar 27, 2020

Copy link
Copy Markdown
ContributorAuthor

Adding the fix (model as IDisposable)?.Dispose(); for freeing Tensor objects worked! This fix is necessary, as these Tensor objects made in the C TensorFlow libraries are not automatically cleaned up by C#'s Garbage Collector.

Edit: While this fix works, it is not safe to assume that this model can be disposed in RunnerUtil.cs. The user might be accessing this model during disposal, which would result in use-after-free and/or null reference errors.

@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from 78a65c2 to 087c0d5CompareApril 16, 2020 05:13
@dotnetdotnet deleted a comment from azure-pipelinesBotApr 16, 2020
@mstfbl

Copy link
Copy Markdown
ContributorAuthor

Freeing Tensor objects in model in a finally statement in TrainAndScorePipeline works in fixing memory bug, and is safe to do when model is never saved in memory and written to disk always.

@mstfblmstfbl closed this Apr 26, 2020
@ghostghost locked as resolved and limited conversation to collaborators Mar 19, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@mstfbl
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Debugging hanging AutoFitImageClassificationTrainTest - #4893

Closed
mstfbl wants to merge 28 commits into
dotnet:masterfrom
mstfbl:AutoFitTests-Debugging
Closed

Debugging hanging AutoFitImageClassificationTrainTest#4893
mstfbl wants to merge 28 commits into
dotnet:masterfrom
mstfbl:AutoFitTests-Debugging

Conversation

@mstfbl

@mstfblmstfbl commented Feb 26, 2020

Copy link
Copy Markdown
Contributor

Will be using this draft PR for general debugging purposes on CI

Notes:
Windows builds have 7,168 MBs of RAM

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

Testing of AutoFitImageClassificationTrainTest is taking too long per test, so tesitng it with 1000 iterations isn't feasible.

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

The tests AutoFitRecommendationTest and AutoFitRegressionTest are passing. AutoFitImageClassificationTrainTest is displaying errors every now and then.

@mstfbl

mstfbl commented Feb 26, 2020

Copy link
Copy Markdown
ContributorAuthor

The reason why AutoFitImageClassificationTrainTest is crashing is after running var result = context.Auto().CreateMulticlassClassificationExperiment(0).Execute(trainDataset, testDataset, columnInference.ColumnInformation), result.Best run might not always be updated. I saw this as I caught result.BestRun having a null value right when it is being called. Exact location where error is thrown.

@mstfbl

mstfbl commented Feb 27, 2020

Copy link
Copy Markdown
ContributorAuthor

The original bug with AutoFitImageClassificationTrainTest is occuring due to the null returned value here:


For some reason, sometimes the validationMetrics of an IEnumerable<(RunDetail) is null.

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

There's an issue with this Evaluation function:

publicMulticlassClassificationMetricsEvaluate(IDataViewdata,stringlabel,stringscore,stringpredictedLabel)
{
Host.CheckValue(data,nameof(data));
Host.CheckNonEmpty(label,nameof(label));
Host.CheckNonEmpty(score,nameof(score));
Host.CheckNonEmpty(predictedLabel,nameof(predictedLabel));
varroles=newRoleMappedData(data,opt:false,
RoleMappedSchema.ColumnRole.Label.Bind(label),
RoleMappedSchema.CreatePair(AnnotationUtils.Const.ScoreValueKind.Score,score),
RoleMappedSchema.CreatePair(AnnotationUtils.Const.ScoreValueKind.PredictedLabel,predictedLabel));
varresultDict=((IEvaluator)this).Evaluate(roles);
Host.Assert(resultDict.ContainsKey(MetricKinds.OverallMetrics));
varoverall=resultDict[MetricKinds.OverallMetrics];
varconfusionMatrix=resultDict[MetricKinds.ConfusionMatrix];
MulticlassClassificationMetricsresult;
using(varcursor=overall.GetRowCursorForAllColumns())
{
varmoved=cursor.MoveNext();
Host.Assert(moved);
result=newMulticlassClassificationMetrics(Host,cursor,_outputTopKAcc??0,confusionMatrix);
moved=cursor.MoveNext();
Host.Assert(!moved);
}
returnresult;
}
}

The returned result value can sometimes be null, which is what is causing AutoFitImageClassificationTrainTest to sometimes fail.

@mstfbl

mstfbl commented Feb 28, 2020

Copy link
Copy Markdown
ContributorAuthor

I figured out the cause of the occasional crash of AutoFitImageClassificationTrainTest. When any exception occurs in RunnerUtil.TrainAndScorePipeline, instead of throwing the error, it is instead caught and ignored while a null metrics value (in line 49) is sent up through the call stack instead.

try
{
varestimator=pipeline.ToEstimator(trainData,validData);
varmodel=estimator.Fit(trainData);
varscoredData=model.Transform(validData);
varmetrics=metricsAgent.EvaluateMetrics(scoredData,labelColumn);
varscore=metricsAgent.GetScore(metrics);
if(preprocessorTransform!=null)
{
model=preprocessorTransform.Append(model);
}
// Build container for model
varmodelContainer=modelFileInfo==null?
newModelContainer(context,model):
newModelContainer(context,modelFileInfo,model,modelInputSchema);
return(modelContainer,metrics,null,score);
}
catch(Exceptionex)
{
logger.Error($"Pipeline crashed: {pipeline.ToString()} . Exception: {ex}");
return(null,null,ex,double.NaN);

This is the exception being caught:

System.ArgumentException : PIPELINE CRASHES - Line 55 - RunnerUtil.cs - Pipeline crash string: xf=ValueToKeyMapping{ col=Label:Label} xf=RawByteImageLoading{ col=ImagePath_featurized:ImagePath imageFolder=} xf=ColumnCopying{ col=Features:ImagePath_featurized} tr=ImageClassification{} xf=KeyToValueMapping{ col=PredictedLabel:PredictedLabel} cache=- - Exception String: System.FormatException: Tensorflow exception triggered while loading model. ---> System.Runtime.InteropServices.SEHException: External component has thrown an exception.

When reproduced locally, the exception string is:

"Could not find a part of the path 'C:\Users\mubal\AppData\Local\Temp\Microsoft.ML.AutoML\experiment_y1gbdyum.xdu\Model1.zip'."

@mstfbl

mstfbl commented Mar 2, 2020

Copy link
Copy Markdown
ContributorAuthor

Update: AutoFitImageClassificationTrainTest with 100 iterations fail on Windows x64 builds with:

System.Runtime.InteropServices.SEHException (0x80004005): External component has thrown an exception

but also I get:
System.FormatException: Tensorflow exception triggered while loading model. ---> System.OutOfMemoryException
and:
System.AccessViolationException: Attempted to read or write protected memory. This is often an indication that other memory is corrupt.

in Tensorflow.c_api.TF_SessionRun, which is the C++ implementation of TensorFlow's training code. This seems related to Issue SciSharp/TensorFlow.NET#485

As mentioned in this PR #4755, we still cannot see details about the crash in Tensorflow.c_api.TF_SessionRun.

@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from 4520530 to 65a72efCompareMarch 19, 2020 05:17
@mstfblmstfbl closed this Mar 20, 2020
@mstfblmstfbl changed the title Auto fit tests debuggingDebugging PRMar 22, 2020
@mstfblmstfbl reopened this Mar 22, 2020
@mstfblmstfbl closed this Mar 22, 2020
@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from c1f8231 to c1e422dCompareMarch 22, 2020 08:39
@mstfblmstfbl reopened this Mar 22, 2020
@mstfblmstfbl changed the title Debugging PRDebugging hanging AutoFitImageClassificationTrainTestMar 26, 2020
@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from df5f642 to 5a7ad17CompareMarch 26, 2020 03:32
@mstfbl

Copy link
Copy Markdown
ContributorAuthor

Will be using this PR to debug AutoFitImageClassificationTrainTest hanging occasionally on Windows builds.

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

AutoFitImageClassificationTrainTest is still occasionally hanging, mostly due to indisposed Tensorflow objects after the test is complete. I found this comment in Microsoft.ML.Vision/ImageClassificationTrainer.TrainModelCore to be of interest:

// Leave the ownership of _session so that it is not disposed/closed when this object goes out of scope
// since it will be used by ImageClassificationModelParameters class (new owner that will take care of
// disposing).
varsession=_session;
_session=null;
returnnewImageClassificationModelParameters(Host,session,_classCount,_jpegDataTensorName,
_resizedImageTensorName,_inputTensorName,_softmaxTensorName);

@mstfbl

mstfbl commented Mar 27, 2020

Copy link
Copy Markdown
ContributorAuthor

Adding the fix (model as IDisposable)?.Dispose(); for freeing Tensor objects worked! This fix is necessary, as these Tensor objects made in the C TensorFlow libraries are not automatically cleaned up by C#'s Garbage Collector.

Edit: While this fix works, it is not safe to assume that this model can be disposed in RunnerUtil.cs. The user might be accessing this model during disposal, which would result in use-after-free and/or null reference errors.

@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from 78a65c2 to 087c0d5CompareApril 16, 2020 05:13
@dotnetdotnet deleted a comment from azure-pipelinesBotApr 16, 2020
@mstfbl

Copy link
Copy Markdown
ContributorAuthor

Freeing Tensor objects in model in a finally statement in TrainAndScorePipeline works in fixing memory bug, and is safe to do when model is never saved in memory and written to disk always.

@mstfblmstfbl closed this Apr 26, 2020
@ghostghost locked as resolved and limited conversation to collaborators Mar 19, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@mstfbl
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Debugging hanging AutoFitImageClassificationTrainTest - #4893

Closed
mstfbl wants to merge 28 commits into
dotnet:masterfrom
mstfbl:AutoFitTests-Debugging
Closed

Debugging hanging AutoFitImageClassificationTrainTest#4893
mstfbl wants to merge 28 commits into
dotnet:masterfrom
mstfbl:AutoFitTests-Debugging

Conversation

@mstfbl

@mstfblmstfbl commented Feb 26, 2020

Copy link
Copy Markdown
Contributor

Will be using this draft PR for general debugging purposes on CI

Notes:
Windows builds have 7,168 MBs of RAM

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

Testing of AutoFitImageClassificationTrainTest is taking too long per test, so tesitng it with 1000 iterations isn't feasible.

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

The tests AutoFitRecommendationTest and AutoFitRegressionTest are passing. AutoFitImageClassificationTrainTest is displaying errors every now and then.

@mstfbl

mstfbl commented Feb 26, 2020

Copy link
Copy Markdown
ContributorAuthor

The reason why AutoFitImageClassificationTrainTest is crashing is after running var result = context.Auto().CreateMulticlassClassificationExperiment(0).Execute(trainDataset, testDataset, columnInference.ColumnInformation), result.Best run might not always be updated. I saw this as I caught result.BestRun having a null value right when it is being called. Exact location where error is thrown.

@mstfbl

mstfbl commented Feb 27, 2020

Copy link
Copy Markdown
ContributorAuthor

The original bug with AutoFitImageClassificationTrainTest is occuring due to the null returned value here:


For some reason, sometimes the validationMetrics of an IEnumerable<(RunDetail) is null.

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

There's an issue with this Evaluation function:

publicMulticlassClassificationMetricsEvaluate(IDataViewdata,stringlabel,stringscore,stringpredictedLabel)
{
Host.CheckValue(data,nameof(data));
Host.CheckNonEmpty(label,nameof(label));
Host.CheckNonEmpty(score,nameof(score));
Host.CheckNonEmpty(predictedLabel,nameof(predictedLabel));
varroles=newRoleMappedData(data,opt:false,
RoleMappedSchema.ColumnRole.Label.Bind(label),
RoleMappedSchema.CreatePair(AnnotationUtils.Const.ScoreValueKind.Score,score),
RoleMappedSchema.CreatePair(AnnotationUtils.Const.ScoreValueKind.PredictedLabel,predictedLabel));
varresultDict=((IEvaluator)this).Evaluate(roles);
Host.Assert(resultDict.ContainsKey(MetricKinds.OverallMetrics));
varoverall=resultDict[MetricKinds.OverallMetrics];
varconfusionMatrix=resultDict[MetricKinds.ConfusionMatrix];
MulticlassClassificationMetricsresult;
using(varcursor=overall.GetRowCursorForAllColumns())
{
varmoved=cursor.MoveNext();
Host.Assert(moved);
result=newMulticlassClassificationMetrics(Host,cursor,_outputTopKAcc??0,confusionMatrix);
moved=cursor.MoveNext();
Host.Assert(!moved);
}
returnresult;
}
}

The returned result value can sometimes be null, which is what is causing AutoFitImageClassificationTrainTest to sometimes fail.

@mstfbl

mstfbl commented Feb 28, 2020

Copy link
Copy Markdown
ContributorAuthor

I figured out the cause of the occasional crash of AutoFitImageClassificationTrainTest. When any exception occurs in RunnerUtil.TrainAndScorePipeline, instead of throwing the error, it is instead caught and ignored while a null metrics value (in line 49) is sent up through the call stack instead.

try
{
varestimator=pipeline.ToEstimator(trainData,validData);
varmodel=estimator.Fit(trainData);
varscoredData=model.Transform(validData);
varmetrics=metricsAgent.EvaluateMetrics(scoredData,labelColumn);
varscore=metricsAgent.GetScore(metrics);
if(preprocessorTransform!=null)
{
model=preprocessorTransform.Append(model);
}
// Build container for model
varmodelContainer=modelFileInfo==null?
newModelContainer(context,model):
newModelContainer(context,modelFileInfo,model,modelInputSchema);
return(modelContainer,metrics,null,score);
}
catch(Exceptionex)
{
logger.Error($"Pipeline crashed: {pipeline.ToString()} . Exception: {ex}");
return(null,null,ex,double.NaN);

This is the exception being caught:

System.ArgumentException : PIPELINE CRASHES - Line 55 - RunnerUtil.cs - Pipeline crash string: xf=ValueToKeyMapping{ col=Label:Label} xf=RawByteImageLoading{ col=ImagePath_featurized:ImagePath imageFolder=} xf=ColumnCopying{ col=Features:ImagePath_featurized} tr=ImageClassification{} xf=KeyToValueMapping{ col=PredictedLabel:PredictedLabel} cache=- - Exception String: System.FormatException: Tensorflow exception triggered while loading model. ---> System.Runtime.InteropServices.SEHException: External component has thrown an exception.

When reproduced locally, the exception string is:

"Could not find a part of the path 'C:\Users\mubal\AppData\Local\Temp\Microsoft.ML.AutoML\experiment_y1gbdyum.xdu\Model1.zip'."

@mstfbl

mstfbl commented Mar 2, 2020

Copy link
Copy Markdown
ContributorAuthor

Update: AutoFitImageClassificationTrainTest with 100 iterations fail on Windows x64 builds with:

System.Runtime.InteropServices.SEHException (0x80004005): External component has thrown an exception

but also I get:
System.FormatException: Tensorflow exception triggered while loading model. ---> System.OutOfMemoryException
and:
System.AccessViolationException: Attempted to read or write protected memory. This is often an indication that other memory is corrupt.

in Tensorflow.c_api.TF_SessionRun, which is the C++ implementation of TensorFlow's training code. This seems related to Issue SciSharp/TensorFlow.NET#485

As mentioned in this PR #4755, we still cannot see details about the crash in Tensorflow.c_api.TF_SessionRun.

@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from 4520530 to 65a72efCompareMarch 19, 2020 05:17
@mstfblmstfbl closed this Mar 20, 2020
@mstfblmstfbl changed the title Auto fit tests debuggingDebugging PRMar 22, 2020
@mstfblmstfbl reopened this Mar 22, 2020
@mstfblmstfbl closed this Mar 22, 2020
@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from c1f8231 to c1e422dCompareMarch 22, 2020 08:39
@mstfblmstfbl reopened this Mar 22, 2020
@mstfblmstfbl changed the title Debugging PRDebugging hanging AutoFitImageClassificationTrainTestMar 26, 2020
@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from df5f642 to 5a7ad17CompareMarch 26, 2020 03:32
@mstfbl

Copy link
Copy Markdown
ContributorAuthor

Will be using this PR to debug AutoFitImageClassificationTrainTest hanging occasionally on Windows builds.

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

AutoFitImageClassificationTrainTest is still occasionally hanging, mostly due to indisposed Tensorflow objects after the test is complete. I found this comment in Microsoft.ML.Vision/ImageClassificationTrainer.TrainModelCore to be of interest:

// Leave the ownership of _session so that it is not disposed/closed when this object goes out of scope
// since it will be used by ImageClassificationModelParameters class (new owner that will take care of
// disposing).
varsession=_session;
_session=null;
returnnewImageClassificationModelParameters(Host,session,_classCount,_jpegDataTensorName,
_resizedImageTensorName,_inputTensorName,_softmaxTensorName);

@mstfbl

mstfbl commented Mar 27, 2020

Copy link
Copy Markdown
ContributorAuthor

Adding the fix (model as IDisposable)?.Dispose(); for freeing Tensor objects worked! This fix is necessary, as these Tensor objects made in the C TensorFlow libraries are not automatically cleaned up by C#'s Garbage Collector.

Edit: While this fix works, it is not safe to assume that this model can be disposed in RunnerUtil.cs. The user might be accessing this model during disposal, which would result in use-after-free and/or null reference errors.

@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from 78a65c2 to 087c0d5CompareApril 16, 2020 05:13
@dotnetdotnet deleted a comment from azure-pipelinesBotApr 16, 2020
@mstfbl

Copy link
Copy Markdown
ContributorAuthor

Freeing Tensor objects in model in a finally statement in TrainAndScorePipeline works in fixing memory bug, and is safe to do when model is never saved in memory and written to disk always.

@mstfblmstfbl closed this Apr 26, 2020
@ghostghost locked as resolved and limited conversation to collaborators Mar 19, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@mstfbl
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Debugging hanging AutoFitImageClassificationTrainTest - #4893

Closed
mstfbl wants to merge 28 commits into
dotnet:masterfrom
mstfbl:AutoFitTests-Debugging
Closed

Debugging hanging AutoFitImageClassificationTrainTest#4893
mstfbl wants to merge 28 commits into
dotnet:masterfrom
mstfbl:AutoFitTests-Debugging

Conversation

@mstfbl

@mstfblmstfbl commented Feb 26, 2020

Copy link
Copy Markdown
Contributor

Will be using this draft PR for general debugging purposes on CI

Notes:
Windows builds have 7,168 MBs of RAM

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

Testing of AutoFitImageClassificationTrainTest is taking too long per test, so tesitng it with 1000 iterations isn't feasible.

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

The tests AutoFitRecommendationTest and AutoFitRegressionTest are passing. AutoFitImageClassificationTrainTest is displaying errors every now and then.

@mstfbl

mstfbl commented Feb 26, 2020

Copy link
Copy Markdown
ContributorAuthor

The reason why AutoFitImageClassificationTrainTest is crashing is after running var result = context.Auto().CreateMulticlassClassificationExperiment(0).Execute(trainDataset, testDataset, columnInference.ColumnInformation), result.Best run might not always be updated. I saw this as I caught result.BestRun having a null value right when it is being called. Exact location where error is thrown.

@mstfbl

mstfbl commented Feb 27, 2020

Copy link
Copy Markdown
ContributorAuthor

The original bug with AutoFitImageClassificationTrainTest is occuring due to the null returned value here:


For some reason, sometimes the validationMetrics of an IEnumerable<(RunDetail) is null.

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

There's an issue with this Evaluation function:

publicMulticlassClassificationMetricsEvaluate(IDataViewdata,stringlabel,stringscore,stringpredictedLabel)
{
Host.CheckValue(data,nameof(data));
Host.CheckNonEmpty(label,nameof(label));
Host.CheckNonEmpty(score,nameof(score));
Host.CheckNonEmpty(predictedLabel,nameof(predictedLabel));
varroles=newRoleMappedData(data,opt:false,
RoleMappedSchema.ColumnRole.Label.Bind(label),
RoleMappedSchema.CreatePair(AnnotationUtils.Const.ScoreValueKind.Score,score),
RoleMappedSchema.CreatePair(AnnotationUtils.Const.ScoreValueKind.PredictedLabel,predictedLabel));
varresultDict=((IEvaluator)this).Evaluate(roles);
Host.Assert(resultDict.ContainsKey(MetricKinds.OverallMetrics));
varoverall=resultDict[MetricKinds.OverallMetrics];
varconfusionMatrix=resultDict[MetricKinds.ConfusionMatrix];
MulticlassClassificationMetricsresult;
using(varcursor=overall.GetRowCursorForAllColumns())
{
varmoved=cursor.MoveNext();
Host.Assert(moved);
result=newMulticlassClassificationMetrics(Host,cursor,_outputTopKAcc??0,confusionMatrix);
moved=cursor.MoveNext();
Host.Assert(!moved);
}
returnresult;
}
}

The returned result value can sometimes be null, which is what is causing AutoFitImageClassificationTrainTest to sometimes fail.

@mstfbl

mstfbl commented Feb 28, 2020

Copy link
Copy Markdown
ContributorAuthor

I figured out the cause of the occasional crash of AutoFitImageClassificationTrainTest. When any exception occurs in RunnerUtil.TrainAndScorePipeline, instead of throwing the error, it is instead caught and ignored while a null metrics value (in line 49) is sent up through the call stack instead.

try
{
varestimator=pipeline.ToEstimator(trainData,validData);
varmodel=estimator.Fit(trainData);
varscoredData=model.Transform(validData);
varmetrics=metricsAgent.EvaluateMetrics(scoredData,labelColumn);
varscore=metricsAgent.GetScore(metrics);
if(preprocessorTransform!=null)
{
model=preprocessorTransform.Append(model);
}
// Build container for model
varmodelContainer=modelFileInfo==null?
newModelContainer(context,model):
newModelContainer(context,modelFileInfo,model,modelInputSchema);
return(modelContainer,metrics,null,score);
}
catch(Exceptionex)
{
logger.Error($"Pipeline crashed: {pipeline.ToString()} . Exception: {ex}");
return(null,null,ex,double.NaN);

This is the exception being caught:

System.ArgumentException : PIPELINE CRASHES - Line 55 - RunnerUtil.cs - Pipeline crash string: xf=ValueToKeyMapping{ col=Label:Label} xf=RawByteImageLoading{ col=ImagePath_featurized:ImagePath imageFolder=} xf=ColumnCopying{ col=Features:ImagePath_featurized} tr=ImageClassification{} xf=KeyToValueMapping{ col=PredictedLabel:PredictedLabel} cache=- - Exception String: System.FormatException: Tensorflow exception triggered while loading model. ---> System.Runtime.InteropServices.SEHException: External component has thrown an exception.

When reproduced locally, the exception string is:

"Could not find a part of the path 'C:\Users\mubal\AppData\Local\Temp\Microsoft.ML.AutoML\experiment_y1gbdyum.xdu\Model1.zip'."

@mstfbl

mstfbl commented Mar 2, 2020

Copy link
Copy Markdown
ContributorAuthor

Update: AutoFitImageClassificationTrainTest with 100 iterations fail on Windows x64 builds with:

System.Runtime.InteropServices.SEHException (0x80004005): External component has thrown an exception

but also I get:
System.FormatException: Tensorflow exception triggered while loading model. ---> System.OutOfMemoryException
and:
System.AccessViolationException: Attempted to read or write protected memory. This is often an indication that other memory is corrupt.

in Tensorflow.c_api.TF_SessionRun, which is the C++ implementation of TensorFlow's training code. This seems related to Issue SciSharp/TensorFlow.NET#485

As mentioned in this PR #4755, we still cannot see details about the crash in Tensorflow.c_api.TF_SessionRun.

@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from 4520530 to 65a72efCompareMarch 19, 2020 05:17
@mstfblmstfbl closed this Mar 20, 2020
@mstfblmstfbl changed the title Auto fit tests debuggingDebugging PRMar 22, 2020
@mstfblmstfbl reopened this Mar 22, 2020
@mstfblmstfbl closed this Mar 22, 2020
@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from c1f8231 to c1e422dCompareMarch 22, 2020 08:39
@mstfblmstfbl reopened this Mar 22, 2020
@mstfblmstfbl changed the title Debugging PRDebugging hanging AutoFitImageClassificationTrainTestMar 26, 2020
@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from df5f642 to 5a7ad17CompareMarch 26, 2020 03:32
@mstfbl

Copy link
Copy Markdown
ContributorAuthor

Will be using this PR to debug AutoFitImageClassificationTrainTest hanging occasionally on Windows builds.

@mstfbl

Copy link
Copy Markdown
ContributorAuthor

AutoFitImageClassificationTrainTest is still occasionally hanging, mostly due to indisposed Tensorflow objects after the test is complete. I found this comment in Microsoft.ML.Vision/ImageClassificationTrainer.TrainModelCore to be of interest:

// Leave the ownership of _session so that it is not disposed/closed when this object goes out of scope
// since it will be used by ImageClassificationModelParameters class (new owner that will take care of
// disposing).
varsession=_session;
_session=null;
returnnewImageClassificationModelParameters(Host,session,_classCount,_jpegDataTensorName,
_resizedImageTensorName,_inputTensorName,_softmaxTensorName);

@mstfbl

mstfbl commented Mar 27, 2020

Copy link
Copy Markdown
ContributorAuthor

Adding the fix (model as IDisposable)?.Dispose(); for freeing Tensor objects worked! This fix is necessary, as these Tensor objects made in the C TensorFlow libraries are not automatically cleaned up by C#'s Garbage Collector.

Edit: While this fix works, it is not safe to assume that this model can be disposed in RunnerUtil.cs. The user might be accessing this model during disposal, which would result in use-after-free and/or null reference errors.

@mstfbl
mstfblforce-pushed the AutoFitTests-Debugging branch from 78a65c2 to 087c0d5CompareApril 16, 2020 05:13
@dotnetdotnet deleted a comment from azure-pipelinesBotApr 16, 2020
@mstfbl

Copy link
Copy Markdown
ContributorAuthor

Freeing Tensor objects in model in a finally statement in TrainAndScorePipeline works in fixing memory bug, and is safe to do when model is never saved in memory and written to disk always.

@mstfblmstfbl closed this Apr 26, 2020
@ghostghost locked as resolved and limited conversation to collaborators Mar 19, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@mstfbl