Tree estimators - #855

Merged
sfilipi merged 19 commits into
dotnet:masterfrom
sfilipi:fastTreeEstimators
Sep 19, 2018
Merged

Tree estimators#855
sfilipi merged 19 commits into
dotnet:masterfrom
sfilipi:fastTreeEstimators

Conversation

@sfilipi

Copy link
Copy Markdown
Member

Ongoing work on converting the trainers to estimators. This PR converts the Tree -type Predictors.

@sfilipi

sfilipi commented Sep 7, 2018

Copy link
Copy Markdown
MemberAuthor

I will add tests next. We don't seem to have many ranking tests enabled :( #Resolved

@sfilipisfilipi self-assigned this Sep 7, 2018
@sfilipisfilipi added the API Issues pertaining the friendly API label Sep 7, 2018
@sfilipisfilipi added this to the 0918 milestone Sep 7, 2018
@sfilipisfilipi changed the title WIP: Fast tree estimatorsWIP: Tree estimatorsSep 7, 2018
@Zruty0Zruty0 mentioned this pull request Sep 7, 2018
}

protected override RankingPredictionTransformer<FastTreeRankingPredictor> MakeTransformer(FastTreeRankingPredictor model, ISchema trainSchema)
=> new RankingPredictionTransformer<FastTreeRankingPredictor>(Host, model, trainSchema, FeatureColumn.Name);

@sfilipisfilipiSep 7, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FeatureColumn.Name); [](start = 96, length = 20)

should add the GroupID to the base constructor #Resolved

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

GroupID? why?


In reply to: 216093781 [](ancestors = 216093781)

Changing the behavior for the creation of the weight column, based on whether it is explicit, or implicit.

private static SchemaShape.Column MakeWeightColumn(Optional<string> weightColumn)
{
if (weightColumn == null || !weightColumn.IsExplicit)

@sfilipisfilipiSep 8, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

|| !weightColumn.IsExplicit [](start = 37, length = 27)

this is not entirely correct either. It won't create the column when the user doesn't specify the weight colum, because it already had the name weight in the data.
we can't peak at the data at this time.

@tfinley@gmail.com@Zruty0 can we move from the Optional to just string for the weight, name, group ID and enforce the user typing in the names? is there another way around it, now that we need to know the information before seeing the data? #Resolved

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we cannot do this really, can we?


In reply to: 216135820 [](ancestors = 216135820)

/// (e.g., the prediction does not happen over a file as it did during training).
/// </summary>
[Fact]
public void New_SimpleTrainAndPredictWithFT()

@Zruty0Zruty0Sep 13, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

New_SimpleTrainAndPredictWithFT [](start = 20, length = 31)

move this test somewhere else #Resolved

using Microsoft.ML.Runtime.Internal.Utilities;
using Microsoft.ML.Runtime.Model;
using Microsoft.ML.Runtime.Internal.Internallearn;
using Microsoft.ML.Core.Data;

@Zruty0Zruty0Sep 13, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

using [](start = 0, length = 5)

sort #Resolved

{
new SchemaShape.Column(DefaultColumnNames.Score, SchemaShape.Column.VectorKind.Scalar, NumberType.R4, false),
new SchemaShape.Column(DefaultColumnNames.Probability, SchemaShape.Column.VectorKind.Scalar, NumberType.R4, false),
new SchemaShape.Column(DefaultColumnNames.PredictedLabel, SchemaShape.Column.VectorKind.Scalar, BoolType.Instance, false)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

double-check this is correct

Making use of dataset definitions
adding Iris.data and the adult.tiny files to TestDatasets
adding regression and ranking tests
/// FastTreeBinaryClassification TrainerEstimator test
/// </summary>
[Fact]
public void FastTreeRankerEstimator()

@sfilipisfilipiSep 14, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

public void FastTreeRankerEstimator() [](start = 7, length = 38)

this is currently failing. #Resolved

}
}

public sealed class RankingPredictionTransformer<TModel> : PredictionTransformerBase<TModel>

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

RankingPredictionTransformer [](start = 24, length = 28)

Is the reason why we have two types that are identical in practically everything but name, so we can identify ranking estimators vs. regression estimators in a statically typed way?

@Zruty0Zruty0Sep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this transformer should also expose the group ID column name, at least that would be my belief


In reply to: 218214277 [](ancestors = 218214277)

@TomFinleyTomFinleySep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actually thought about this, like labels group ids are only needed for training, right? So for prediction I don't think they should be.


In reply to: 218216192 [](ancestors = 218216192,218214277)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So keep it, or make the Regression one Generic and use it for both?


In reply to: 218216839 [](ancestors = 218216839,218216192,218214277)

{
PredictorType = ComponentFactoryUtils.CreateFromFunction(
e => new AveragedPerceptronTrainer(e, new AveragedPerceptronTrainer.Arguments()))
e => new FastTreeBinaryClassificationTrainer(e, DefaultColumnNames.Label, DefaultColumnNames.Features))

@TomFinleyTomFinleySep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FastTreeBinaryClassificationTrainer [](start = 37, length = 35)

I'd really rather we didn't. This seems to fit into the same bucket as the discussion on #682. That ensembling should have a dependency on FastTree merely because we have a default does not make sense to me. If someone wants to use stacking, that's great, but they need to specify the learners. #Pending

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

But maybe we can hold off for right now.


In reply to: 218215145 [](ancestors = 218215145)

@sfilipisfilipiSep 17, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, let's do that separately, when we shape the ensembles to take in the arguments in the constructor.


In reply to: 218215323 [](ancestors = 218215323,218215145)

using Microsoft.ML.Runtime.TreePredictor;
using Newtonsoft.Json.Linq;
using Microsoft.ML.Core.Data;
using Microsoft.ML.Runtime.EntryPoints;

@TomFinleyTomFinleySep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm probably just missing something obvious, but why does this now depend on entry-points namespace?

Also sorting. #Resolved

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you! Oversight


In reply to: 218216150 [](ancestors = 218216150)

@TomFinleyTomFinley left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@TomFinley

Copy link
Copy Markdown
Contributor

Is omission of Pigsty extensions deliberate?

@sfilipi

Copy link
Copy Markdown
MemberAuthor

Did i misunderstand that for trainers we should hold on to doing the Pigsty extensions until we get the ml task, so we could extend on that, rather than the label? @tfinley@gmail.com@Zruty0, let me know if i should actually work on them in the same PR.


In reply to: 422161730 [](ancestors = 422161730)

@Zruty0Zruty0 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@Zruty0

Copy link
Copy Markdown
Contributor

I have the same (mis)understanding. In any case, let's do it after this oner


In reply to: 422174998 [](ancestors = 422174998,422161730)

@sfilipisfilipi changed the title WIP: Tree estimatorsTree estimatorsSep 18, 2018
@sfilipi
sfilipi merged commit d13b415 into dotnet:masterSep 19, 2018
@sfilipisfilipi mentioned this pull request Sep 21, 2018
@sfilipi
sfilipi deleted the fastTreeEstimators branch October 22, 2018 16:57
@ghostghost locked as resolved and limited conversation to collaborators Mar 29, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

APIIssues pertaining the friendly API

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@sfilipi@TomFinley@Zruty0
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Tree estimators - #855

Merged
sfilipi merged 19 commits into
dotnet:masterfrom
sfilipi:fastTreeEstimators
Sep 19, 2018
Merged

Tree estimators#855
sfilipi merged 19 commits into
dotnet:masterfrom
sfilipi:fastTreeEstimators

Conversation

@sfilipi

Copy link
Copy Markdown
Member

Ongoing work on converting the trainers to estimators. This PR converts the Tree -type Predictors.

@sfilipi

sfilipi commented Sep 7, 2018

Copy link
Copy Markdown
MemberAuthor

I will add tests next. We don't seem to have many ranking tests enabled :( #Resolved

@sfilipisfilipi self-assigned this Sep 7, 2018
@sfilipisfilipi added the API Issues pertaining the friendly API label Sep 7, 2018
@sfilipisfilipi added this to the 0918 milestone Sep 7, 2018
@sfilipisfilipi changed the title WIP: Fast tree estimatorsWIP: Tree estimatorsSep 7, 2018
@Zruty0Zruty0 mentioned this pull request Sep 7, 2018
}

protected override RankingPredictionTransformer<FastTreeRankingPredictor> MakeTransformer(FastTreeRankingPredictor model, ISchema trainSchema)
=> new RankingPredictionTransformer<FastTreeRankingPredictor>(Host, model, trainSchema, FeatureColumn.Name);

@sfilipisfilipiSep 7, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FeatureColumn.Name); [](start = 96, length = 20)

should add the GroupID to the base constructor #Resolved

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

GroupID? why?


In reply to: 216093781 [](ancestors = 216093781)

Changing the behavior for the creation of the weight column, based on whether it is explicit, or implicit.

private static SchemaShape.Column MakeWeightColumn(Optional<string> weightColumn)
{
if (weightColumn == null || !weightColumn.IsExplicit)

@sfilipisfilipiSep 8, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

|| !weightColumn.IsExplicit [](start = 37, length = 27)

this is not entirely correct either. It won't create the column when the user doesn't specify the weight colum, because it already had the name weight in the data.
we can't peak at the data at this time.

@tfinley@gmail.com@Zruty0 can we move from the Optional to just string for the weight, name, group ID and enforce the user typing in the names? is there another way around it, now that we need to know the information before seeing the data? #Resolved

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we cannot do this really, can we?


In reply to: 216135820 [](ancestors = 216135820)

/// (e.g., the prediction does not happen over a file as it did during training).
/// </summary>
[Fact]
public void New_SimpleTrainAndPredictWithFT()

@Zruty0Zruty0Sep 13, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

New_SimpleTrainAndPredictWithFT [](start = 20, length = 31)

move this test somewhere else #Resolved

using Microsoft.ML.Runtime.Internal.Utilities;
using Microsoft.ML.Runtime.Model;
using Microsoft.ML.Runtime.Internal.Internallearn;
using Microsoft.ML.Core.Data;

@Zruty0Zruty0Sep 13, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

using [](start = 0, length = 5)

sort #Resolved

{
new SchemaShape.Column(DefaultColumnNames.Score, SchemaShape.Column.VectorKind.Scalar, NumberType.R4, false),
new SchemaShape.Column(DefaultColumnNames.Probability, SchemaShape.Column.VectorKind.Scalar, NumberType.R4, false),
new SchemaShape.Column(DefaultColumnNames.PredictedLabel, SchemaShape.Column.VectorKind.Scalar, BoolType.Instance, false)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

double-check this is correct

Making use of dataset definitions
adding Iris.data and the adult.tiny files to TestDatasets
adding regression and ranking tests
/// FastTreeBinaryClassification TrainerEstimator test
/// </summary>
[Fact]
public void FastTreeRankerEstimator()

@sfilipisfilipiSep 14, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

public void FastTreeRankerEstimator() [](start = 7, length = 38)

this is currently failing. #Resolved

}
}

public sealed class RankingPredictionTransformer<TModel> : PredictionTransformerBase<TModel>

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

RankingPredictionTransformer [](start = 24, length = 28)

Is the reason why we have two types that are identical in practically everything but name, so we can identify ranking estimators vs. regression estimators in a statically typed way?

@Zruty0Zruty0Sep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this transformer should also expose the group ID column name, at least that would be my belief


In reply to: 218214277 [](ancestors = 218214277)

@TomFinleyTomFinleySep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actually thought about this, like labels group ids are only needed for training, right? So for prediction I don't think they should be.


In reply to: 218216192 [](ancestors = 218216192,218214277)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So keep it, or make the Regression one Generic and use it for both?


In reply to: 218216839 [](ancestors = 218216839,218216192,218214277)

{
PredictorType = ComponentFactoryUtils.CreateFromFunction(
e => new AveragedPerceptronTrainer(e, new AveragedPerceptronTrainer.Arguments()))
e => new FastTreeBinaryClassificationTrainer(e, DefaultColumnNames.Label, DefaultColumnNames.Features))

@TomFinleyTomFinleySep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FastTreeBinaryClassificationTrainer [](start = 37, length = 35)

I'd really rather we didn't. This seems to fit into the same bucket as the discussion on #682. That ensembling should have a dependency on FastTree merely because we have a default does not make sense to me. If someone wants to use stacking, that's great, but they need to specify the learners. #Pending

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

But maybe we can hold off for right now.


In reply to: 218215145 [](ancestors = 218215145)

@sfilipisfilipiSep 17, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, let's do that separately, when we shape the ensembles to take in the arguments in the constructor.


In reply to: 218215323 [](ancestors = 218215323,218215145)

using Microsoft.ML.Runtime.TreePredictor;
using Newtonsoft.Json.Linq;
using Microsoft.ML.Core.Data;
using Microsoft.ML.Runtime.EntryPoints;

@TomFinleyTomFinleySep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm probably just missing something obvious, but why does this now depend on entry-points namespace?

Also sorting. #Resolved

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you! Oversight


In reply to: 218216150 [](ancestors = 218216150)

@TomFinleyTomFinley left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@TomFinley

Copy link
Copy Markdown
Contributor

Is omission of Pigsty extensions deliberate?

@sfilipi

Copy link
Copy Markdown
MemberAuthor

Did i misunderstand that for trainers we should hold on to doing the Pigsty extensions until we get the ml task, so we could extend on that, rather than the label? @tfinley@gmail.com@Zruty0, let me know if i should actually work on them in the same PR.


In reply to: 422161730 [](ancestors = 422161730)

@Zruty0Zruty0 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@Zruty0

Copy link
Copy Markdown
Contributor

I have the same (mis)understanding. In any case, let's do it after this oner


In reply to: 422174998 [](ancestors = 422174998,422161730)

@sfilipisfilipi changed the title WIP: Tree estimatorsTree estimatorsSep 18, 2018
@sfilipi
sfilipi merged commit d13b415 into dotnet:masterSep 19, 2018
@sfilipisfilipi mentioned this pull request Sep 21, 2018
@sfilipi
sfilipi deleted the fastTreeEstimators branch October 22, 2018 16:57
@ghostghost locked as resolved and limited conversation to collaborators Mar 29, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

APIIssues pertaining the friendly API

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@sfilipi@TomFinley@Zruty0
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Tree estimators - #855

Merged
sfilipi merged 19 commits into
dotnet:masterfrom
sfilipi:fastTreeEstimators
Sep 19, 2018
Merged

Tree estimators#855
sfilipi merged 19 commits into
dotnet:masterfrom
sfilipi:fastTreeEstimators

Conversation

@sfilipi

Copy link
Copy Markdown
Member

Ongoing work on converting the trainers to estimators. This PR converts the Tree -type Predictors.

@sfilipi

sfilipi commented Sep 7, 2018

Copy link
Copy Markdown
MemberAuthor

I will add tests next. We don't seem to have many ranking tests enabled :( #Resolved

@sfilipisfilipi self-assigned this Sep 7, 2018
@sfilipisfilipi added the API Issues pertaining the friendly API label Sep 7, 2018
@sfilipisfilipi added this to the 0918 milestone Sep 7, 2018
@sfilipisfilipi changed the title WIP: Fast tree estimatorsWIP: Tree estimatorsSep 7, 2018
@Zruty0Zruty0 mentioned this pull request Sep 7, 2018
}

protected override RankingPredictionTransformer<FastTreeRankingPredictor> MakeTransformer(FastTreeRankingPredictor model, ISchema trainSchema)
=> new RankingPredictionTransformer<FastTreeRankingPredictor>(Host, model, trainSchema, FeatureColumn.Name);

@sfilipisfilipiSep 7, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FeatureColumn.Name); [](start = 96, length = 20)

should add the GroupID to the base constructor #Resolved

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

GroupID? why?


In reply to: 216093781 [](ancestors = 216093781)

Changing the behavior for the creation of the weight column, based on whether it is explicit, or implicit.

private static SchemaShape.Column MakeWeightColumn(Optional<string> weightColumn)
{
if (weightColumn == null || !weightColumn.IsExplicit)

@sfilipisfilipiSep 8, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

|| !weightColumn.IsExplicit [](start = 37, length = 27)

this is not entirely correct either. It won't create the column when the user doesn't specify the weight colum, because it already had the name weight in the data.
we can't peak at the data at this time.

@tfinley@gmail.com@Zruty0 can we move from the Optional to just string for the weight, name, group ID and enforce the user typing in the names? is there another way around it, now that we need to know the information before seeing the data? #Resolved

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we cannot do this really, can we?


In reply to: 216135820 [](ancestors = 216135820)

/// (e.g., the prediction does not happen over a file as it did during training).
/// </summary>
[Fact]
public void New_SimpleTrainAndPredictWithFT()

@Zruty0Zruty0Sep 13, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

New_SimpleTrainAndPredictWithFT [](start = 20, length = 31)

move this test somewhere else #Resolved

using Microsoft.ML.Runtime.Internal.Utilities;
using Microsoft.ML.Runtime.Model;
using Microsoft.ML.Runtime.Internal.Internallearn;
using Microsoft.ML.Core.Data;

@Zruty0Zruty0Sep 13, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

using [](start = 0, length = 5)

sort #Resolved

{
new SchemaShape.Column(DefaultColumnNames.Score, SchemaShape.Column.VectorKind.Scalar, NumberType.R4, false),
new SchemaShape.Column(DefaultColumnNames.Probability, SchemaShape.Column.VectorKind.Scalar, NumberType.R4, false),
new SchemaShape.Column(DefaultColumnNames.PredictedLabel, SchemaShape.Column.VectorKind.Scalar, BoolType.Instance, false)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

double-check this is correct

Making use of dataset definitions
adding Iris.data and the adult.tiny files to TestDatasets
adding regression and ranking tests
/// FastTreeBinaryClassification TrainerEstimator test
/// </summary>
[Fact]
public void FastTreeRankerEstimator()

@sfilipisfilipiSep 14, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

public void FastTreeRankerEstimator() [](start = 7, length = 38)

this is currently failing. #Resolved

}
}

public sealed class RankingPredictionTransformer<TModel> : PredictionTransformerBase<TModel>

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

RankingPredictionTransformer [](start = 24, length = 28)

Is the reason why we have two types that are identical in practically everything but name, so we can identify ranking estimators vs. regression estimators in a statically typed way?

@Zruty0Zruty0Sep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this transformer should also expose the group ID column name, at least that would be my belief


In reply to: 218214277 [](ancestors = 218214277)

@TomFinleyTomFinleySep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actually thought about this, like labels group ids are only needed for training, right? So for prediction I don't think they should be.


In reply to: 218216192 [](ancestors = 218216192,218214277)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So keep it, or make the Regression one Generic and use it for both?


In reply to: 218216839 [](ancestors = 218216839,218216192,218214277)

{
PredictorType = ComponentFactoryUtils.CreateFromFunction(
e => new AveragedPerceptronTrainer(e, new AveragedPerceptronTrainer.Arguments()))
e => new FastTreeBinaryClassificationTrainer(e, DefaultColumnNames.Label, DefaultColumnNames.Features))

@TomFinleyTomFinleySep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FastTreeBinaryClassificationTrainer [](start = 37, length = 35)

I'd really rather we didn't. This seems to fit into the same bucket as the discussion on #682. That ensembling should have a dependency on FastTree merely because we have a default does not make sense to me. If someone wants to use stacking, that's great, but they need to specify the learners. #Pending

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

But maybe we can hold off for right now.


In reply to: 218215145 [](ancestors = 218215145)

@sfilipisfilipiSep 17, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, let's do that separately, when we shape the ensembles to take in the arguments in the constructor.


In reply to: 218215323 [](ancestors = 218215323,218215145)

using Microsoft.ML.Runtime.TreePredictor;
using Newtonsoft.Json.Linq;
using Microsoft.ML.Core.Data;
using Microsoft.ML.Runtime.EntryPoints;

@TomFinleyTomFinleySep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm probably just missing something obvious, but why does this now depend on entry-points namespace?

Also sorting. #Resolved

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you! Oversight


In reply to: 218216150 [](ancestors = 218216150)

@TomFinleyTomFinley left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@TomFinley

Copy link
Copy Markdown
Contributor

Is omission of Pigsty extensions deliberate?

@sfilipi

Copy link
Copy Markdown
MemberAuthor

Did i misunderstand that for trainers we should hold on to doing the Pigsty extensions until we get the ml task, so we could extend on that, rather than the label? @tfinley@gmail.com@Zruty0, let me know if i should actually work on them in the same PR.


In reply to: 422161730 [](ancestors = 422161730)

@Zruty0Zruty0 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@Zruty0

Copy link
Copy Markdown
Contributor

I have the same (mis)understanding. In any case, let's do it after this oner


In reply to: 422174998 [](ancestors = 422174998,422161730)

@sfilipisfilipi changed the title WIP: Tree estimatorsTree estimatorsSep 18, 2018
@sfilipi
sfilipi merged commit d13b415 into dotnet:masterSep 19, 2018
@sfilipisfilipi mentioned this pull request Sep 21, 2018
@sfilipi
sfilipi deleted the fastTreeEstimators branch October 22, 2018 16:57
@ghostghost locked as resolved and limited conversation to collaborators Mar 29, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

APIIssues pertaining the friendly API

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@sfilipi@TomFinley@Zruty0
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Tree estimators - #855

Merged
sfilipi merged 19 commits into
dotnet:masterfrom
sfilipi:fastTreeEstimators
Sep 19, 2018
Merged

Tree estimators#855
sfilipi merged 19 commits into
dotnet:masterfrom
sfilipi:fastTreeEstimators

Conversation

@sfilipi

Copy link
Copy Markdown
Member

Ongoing work on converting the trainers to estimators. This PR converts the Tree -type Predictors.

@sfilipi

sfilipi commented Sep 7, 2018

Copy link
Copy Markdown
MemberAuthor

I will add tests next. We don't seem to have many ranking tests enabled :( #Resolved

@sfilipisfilipi self-assigned this Sep 7, 2018
@sfilipisfilipi added the API Issues pertaining the friendly API label Sep 7, 2018
@sfilipisfilipi added this to the 0918 milestone Sep 7, 2018
@sfilipisfilipi changed the title WIP: Fast tree estimatorsWIP: Tree estimatorsSep 7, 2018
@Zruty0Zruty0 mentioned this pull request Sep 7, 2018
}

protected override RankingPredictionTransformer<FastTreeRankingPredictor> MakeTransformer(FastTreeRankingPredictor model, ISchema trainSchema)
=> new RankingPredictionTransformer<FastTreeRankingPredictor>(Host, model, trainSchema, FeatureColumn.Name);

@sfilipisfilipiSep 7, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FeatureColumn.Name); [](start = 96, length = 20)

should add the GroupID to the base constructor #Resolved

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

GroupID? why?


In reply to: 216093781 [](ancestors = 216093781)

Changing the behavior for the creation of the weight column, based on whether it is explicit, or implicit.

private static SchemaShape.Column MakeWeightColumn(Optional<string> weightColumn)
{
if (weightColumn == null || !weightColumn.IsExplicit)

@sfilipisfilipiSep 8, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

|| !weightColumn.IsExplicit [](start = 37, length = 27)

this is not entirely correct either. It won't create the column when the user doesn't specify the weight colum, because it already had the name weight in the data.
we can't peak at the data at this time.

@tfinley@gmail.com@Zruty0 can we move from the Optional to just string for the weight, name, group ID and enforce the user typing in the names? is there another way around it, now that we need to know the information before seeing the data? #Resolved

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we cannot do this really, can we?


In reply to: 216135820 [](ancestors = 216135820)

/// (e.g., the prediction does not happen over a file as it did during training).
/// </summary>
[Fact]
public void New_SimpleTrainAndPredictWithFT()

@Zruty0Zruty0Sep 13, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

New_SimpleTrainAndPredictWithFT [](start = 20, length = 31)

move this test somewhere else #Resolved

using Microsoft.ML.Runtime.Internal.Utilities;
using Microsoft.ML.Runtime.Model;
using Microsoft.ML.Runtime.Internal.Internallearn;
using Microsoft.ML.Core.Data;

@Zruty0Zruty0Sep 13, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

using [](start = 0, length = 5)

sort #Resolved

{
new SchemaShape.Column(DefaultColumnNames.Score, SchemaShape.Column.VectorKind.Scalar, NumberType.R4, false),
new SchemaShape.Column(DefaultColumnNames.Probability, SchemaShape.Column.VectorKind.Scalar, NumberType.R4, false),
new SchemaShape.Column(DefaultColumnNames.PredictedLabel, SchemaShape.Column.VectorKind.Scalar, BoolType.Instance, false)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

double-check this is correct

Making use of dataset definitions
adding Iris.data and the adult.tiny files to TestDatasets
adding regression and ranking tests
/// FastTreeBinaryClassification TrainerEstimator test
/// </summary>
[Fact]
public void FastTreeRankerEstimator()

@sfilipisfilipiSep 14, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

public void FastTreeRankerEstimator() [](start = 7, length = 38)

this is currently failing. #Resolved

}
}

public sealed class RankingPredictionTransformer<TModel> : PredictionTransformerBase<TModel>

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

RankingPredictionTransformer [](start = 24, length = 28)

Is the reason why we have two types that are identical in practically everything but name, so we can identify ranking estimators vs. regression estimators in a statically typed way?

@Zruty0Zruty0Sep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this transformer should also expose the group ID column name, at least that would be my belief


In reply to: 218214277 [](ancestors = 218214277)

@TomFinleyTomFinleySep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actually thought about this, like labels group ids are only needed for training, right? So for prediction I don't think they should be.


In reply to: 218216192 [](ancestors = 218216192,218214277)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So keep it, or make the Regression one Generic and use it for both?


In reply to: 218216839 [](ancestors = 218216839,218216192,218214277)

{
PredictorType = ComponentFactoryUtils.CreateFromFunction(
e => new AveragedPerceptronTrainer(e, new AveragedPerceptronTrainer.Arguments()))
e => new FastTreeBinaryClassificationTrainer(e, DefaultColumnNames.Label, DefaultColumnNames.Features))

@TomFinleyTomFinleySep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FastTreeBinaryClassificationTrainer [](start = 37, length = 35)

I'd really rather we didn't. This seems to fit into the same bucket as the discussion on #682. That ensembling should have a dependency on FastTree merely because we have a default does not make sense to me. If someone wants to use stacking, that's great, but they need to specify the learners. #Pending

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

But maybe we can hold off for right now.


In reply to: 218215145 [](ancestors = 218215145)

@sfilipisfilipiSep 17, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, let's do that separately, when we shape the ensembles to take in the arguments in the constructor.


In reply to: 218215323 [](ancestors = 218215323,218215145)

using Microsoft.ML.Runtime.TreePredictor;
using Newtonsoft.Json.Linq;
using Microsoft.ML.Core.Data;
using Microsoft.ML.Runtime.EntryPoints;

@TomFinleyTomFinleySep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm probably just missing something obvious, but why does this now depend on entry-points namespace?

Also sorting. #Resolved

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you! Oversight


In reply to: 218216150 [](ancestors = 218216150)

@TomFinleyTomFinley left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@TomFinley

Copy link
Copy Markdown
Contributor

Is omission of Pigsty extensions deliberate?

@sfilipi

Copy link
Copy Markdown
MemberAuthor

Did i misunderstand that for trainers we should hold on to doing the Pigsty extensions until we get the ml task, so we could extend on that, rather than the label? @tfinley@gmail.com@Zruty0, let me know if i should actually work on them in the same PR.


In reply to: 422161730 [](ancestors = 422161730)

@Zruty0Zruty0 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@Zruty0

Copy link
Copy Markdown
Contributor

I have the same (mis)understanding. In any case, let's do it after this oner


In reply to: 422174998 [](ancestors = 422174998,422161730)

@sfilipisfilipi changed the title WIP: Tree estimatorsTree estimatorsSep 18, 2018
@sfilipi
sfilipi merged commit d13b415 into dotnet:masterSep 19, 2018
@sfilipisfilipi mentioned this pull request Sep 21, 2018
@sfilipi
sfilipi deleted the fastTreeEstimators branch October 22, 2018 16:57
@ghostghost locked as resolved and limited conversation to collaborators Mar 29, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

APIIssues pertaining the friendly API

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@sfilipi@TomFinley@Zruty0
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Tree estimators - #855

Merged
sfilipi merged 19 commits into
dotnet:masterfrom
sfilipi:fastTreeEstimators
Sep 19, 2018
Merged

Tree estimators#855
sfilipi merged 19 commits into
dotnet:masterfrom
sfilipi:fastTreeEstimators

Conversation

@sfilipi

Copy link
Copy Markdown
Member

Ongoing work on converting the trainers to estimators. This PR converts the Tree -type Predictors.

@sfilipi

sfilipi commented Sep 7, 2018

Copy link
Copy Markdown
MemberAuthor

I will add tests next. We don't seem to have many ranking tests enabled :( #Resolved

@sfilipisfilipi self-assigned this Sep 7, 2018
@sfilipisfilipi added the API Issues pertaining the friendly API label Sep 7, 2018
@sfilipisfilipi added this to the 0918 milestone Sep 7, 2018
@sfilipisfilipi changed the title WIP: Fast tree estimatorsWIP: Tree estimatorsSep 7, 2018
@Zruty0Zruty0 mentioned this pull request Sep 7, 2018
}

protected override RankingPredictionTransformer<FastTreeRankingPredictor> MakeTransformer(FastTreeRankingPredictor model, ISchema trainSchema)
=> new RankingPredictionTransformer<FastTreeRankingPredictor>(Host, model, trainSchema, FeatureColumn.Name);

@sfilipisfilipiSep 7, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FeatureColumn.Name); [](start = 96, length = 20)

should add the GroupID to the base constructor #Resolved

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

GroupID? why?


In reply to: 216093781 [](ancestors = 216093781)

Changing the behavior for the creation of the weight column, based on whether it is explicit, or implicit.

private static SchemaShape.Column MakeWeightColumn(Optional<string> weightColumn)
{
if (weightColumn == null || !weightColumn.IsExplicit)

@sfilipisfilipiSep 8, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

|| !weightColumn.IsExplicit [](start = 37, length = 27)

this is not entirely correct either. It won't create the column when the user doesn't specify the weight colum, because it already had the name weight in the data.
we can't peak at the data at this time.

@tfinley@gmail.com@Zruty0 can we move from the Optional to just string for the weight, name, group ID and enforce the user typing in the names? is there another way around it, now that we need to know the information before seeing the data? #Resolved

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we cannot do this really, can we?


In reply to: 216135820 [](ancestors = 216135820)

/// (e.g., the prediction does not happen over a file as it did during training).
/// </summary>
[Fact]
public void New_SimpleTrainAndPredictWithFT()

@Zruty0Zruty0Sep 13, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

New_SimpleTrainAndPredictWithFT [](start = 20, length = 31)

move this test somewhere else #Resolved

using Microsoft.ML.Runtime.Internal.Utilities;
using Microsoft.ML.Runtime.Model;
using Microsoft.ML.Runtime.Internal.Internallearn;
using Microsoft.ML.Core.Data;

@Zruty0Zruty0Sep 13, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

using [](start = 0, length = 5)

sort #Resolved

{
new SchemaShape.Column(DefaultColumnNames.Score, SchemaShape.Column.VectorKind.Scalar, NumberType.R4, false),
new SchemaShape.Column(DefaultColumnNames.Probability, SchemaShape.Column.VectorKind.Scalar, NumberType.R4, false),
new SchemaShape.Column(DefaultColumnNames.PredictedLabel, SchemaShape.Column.VectorKind.Scalar, BoolType.Instance, false)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

double-check this is correct

Making use of dataset definitions
adding Iris.data and the adult.tiny files to TestDatasets
adding regression and ranking tests
/// FastTreeBinaryClassification TrainerEstimator test
/// </summary>
[Fact]
public void FastTreeRankerEstimator()

@sfilipisfilipiSep 14, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

public void FastTreeRankerEstimator() [](start = 7, length = 38)

this is currently failing. #Resolved

}
}

public sealed class RankingPredictionTransformer<TModel> : PredictionTransformerBase<TModel>

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

RankingPredictionTransformer [](start = 24, length = 28)

Is the reason why we have two types that are identical in practically everything but name, so we can identify ranking estimators vs. regression estimators in a statically typed way?

@Zruty0Zruty0Sep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this transformer should also expose the group ID column name, at least that would be my belief


In reply to: 218214277 [](ancestors = 218214277)

@TomFinleyTomFinleySep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actually thought about this, like labels group ids are only needed for training, right? So for prediction I don't think they should be.


In reply to: 218216192 [](ancestors = 218216192,218214277)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So keep it, or make the Regression one Generic and use it for both?


In reply to: 218216839 [](ancestors = 218216839,218216192,218214277)

{
PredictorType = ComponentFactoryUtils.CreateFromFunction(
e => new AveragedPerceptronTrainer(e, new AveragedPerceptronTrainer.Arguments()))
e => new FastTreeBinaryClassificationTrainer(e, DefaultColumnNames.Label, DefaultColumnNames.Features))

@TomFinleyTomFinleySep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FastTreeBinaryClassificationTrainer [](start = 37, length = 35)

I'd really rather we didn't. This seems to fit into the same bucket as the discussion on #682. That ensembling should have a dependency on FastTree merely because we have a default does not make sense to me. If someone wants to use stacking, that's great, but they need to specify the learners. #Pending

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

But maybe we can hold off for right now.


In reply to: 218215145 [](ancestors = 218215145)

@sfilipisfilipiSep 17, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, let's do that separately, when we shape the ensembles to take in the arguments in the constructor.


In reply to: 218215323 [](ancestors = 218215323,218215145)

using Microsoft.ML.Runtime.TreePredictor;
using Newtonsoft.Json.Linq;
using Microsoft.ML.Core.Data;
using Microsoft.ML.Runtime.EntryPoints;

@TomFinleyTomFinleySep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm probably just missing something obvious, but why does this now depend on entry-points namespace?

Also sorting. #Resolved

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you! Oversight


In reply to: 218216150 [](ancestors = 218216150)

@TomFinleyTomFinley left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@TomFinley

Copy link
Copy Markdown
Contributor

Is omission of Pigsty extensions deliberate?

@sfilipi

Copy link
Copy Markdown
MemberAuthor

Did i misunderstand that for trainers we should hold on to doing the Pigsty extensions until we get the ml task, so we could extend on that, rather than the label? @tfinley@gmail.com@Zruty0, let me know if i should actually work on them in the same PR.


In reply to: 422161730 [](ancestors = 422161730)

@Zruty0Zruty0 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@Zruty0

Copy link
Copy Markdown
Contributor

I have the same (mis)understanding. In any case, let's do it after this oner


In reply to: 422174998 [](ancestors = 422174998,422161730)

@sfilipisfilipi changed the title WIP: Tree estimatorsTree estimatorsSep 18, 2018
@sfilipi
sfilipi merged commit d13b415 into dotnet:masterSep 19, 2018
@sfilipisfilipi mentioned this pull request Sep 21, 2018
@sfilipi
sfilipi deleted the fastTreeEstimators branch October 22, 2018 16:57
@ghostghost locked as resolved and limited conversation to collaborators Mar 29, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

APIIssues pertaining the friendly API

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@sfilipi@TomFinley@Zruty0
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Tree estimators - #855

Merged
sfilipi merged 19 commits into
dotnet:masterfrom
sfilipi:fastTreeEstimators
Sep 19, 2018
Merged

Tree estimators#855
sfilipi merged 19 commits into
dotnet:masterfrom
sfilipi:fastTreeEstimators

Conversation

@sfilipi

Copy link
Copy Markdown
Member

Ongoing work on converting the trainers to estimators. This PR converts the Tree -type Predictors.

@sfilipi

sfilipi commented Sep 7, 2018

Copy link
Copy Markdown
MemberAuthor

I will add tests next. We don't seem to have many ranking tests enabled :( #Resolved

@sfilipisfilipi self-assigned this Sep 7, 2018
@sfilipisfilipi added the API Issues pertaining the friendly API label Sep 7, 2018
@sfilipisfilipi added this to the 0918 milestone Sep 7, 2018
@sfilipisfilipi changed the title WIP: Fast tree estimatorsWIP: Tree estimatorsSep 7, 2018
@Zruty0Zruty0 mentioned this pull request Sep 7, 2018
}

protected override RankingPredictionTransformer<FastTreeRankingPredictor> MakeTransformer(FastTreeRankingPredictor model, ISchema trainSchema)
=> new RankingPredictionTransformer<FastTreeRankingPredictor>(Host, model, trainSchema, FeatureColumn.Name);

@sfilipisfilipiSep 7, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FeatureColumn.Name); [](start = 96, length = 20)

should add the GroupID to the base constructor #Resolved

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

GroupID? why?


In reply to: 216093781 [](ancestors = 216093781)

Changing the behavior for the creation of the weight column, based on whether it is explicit, or implicit.

private static SchemaShape.Column MakeWeightColumn(Optional<string> weightColumn)
{
if (weightColumn == null || !weightColumn.IsExplicit)

@sfilipisfilipiSep 8, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

|| !weightColumn.IsExplicit [](start = 37, length = 27)

this is not entirely correct either. It won't create the column when the user doesn't specify the weight colum, because it already had the name weight in the data.
we can't peak at the data at this time.

@tfinley@gmail.com@Zruty0 can we move from the Optional to just string for the weight, name, group ID and enforce the user typing in the names? is there another way around it, now that we need to know the information before seeing the data? #Resolved

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we cannot do this really, can we?


In reply to: 216135820 [](ancestors = 216135820)

/// (e.g., the prediction does not happen over a file as it did during training).
/// </summary>
[Fact]
public void New_SimpleTrainAndPredictWithFT()

@Zruty0Zruty0Sep 13, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

New_SimpleTrainAndPredictWithFT [](start = 20, length = 31)

move this test somewhere else #Resolved

using Microsoft.ML.Runtime.Internal.Utilities;
using Microsoft.ML.Runtime.Model;
using Microsoft.ML.Runtime.Internal.Internallearn;
using Microsoft.ML.Core.Data;

@Zruty0Zruty0Sep 13, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

using [](start = 0, length = 5)

sort #Resolved

{
new SchemaShape.Column(DefaultColumnNames.Score, SchemaShape.Column.VectorKind.Scalar, NumberType.R4, false),
new SchemaShape.Column(DefaultColumnNames.Probability, SchemaShape.Column.VectorKind.Scalar, NumberType.R4, false),
new SchemaShape.Column(DefaultColumnNames.PredictedLabel, SchemaShape.Column.VectorKind.Scalar, BoolType.Instance, false)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

double-check this is correct

Making use of dataset definitions
adding Iris.data and the adult.tiny files to TestDatasets
adding regression and ranking tests
/// FastTreeBinaryClassification TrainerEstimator test
/// </summary>
[Fact]
public void FastTreeRankerEstimator()

@sfilipisfilipiSep 14, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

public void FastTreeRankerEstimator() [](start = 7, length = 38)

this is currently failing. #Resolved

}
}

public sealed class RankingPredictionTransformer<TModel> : PredictionTransformerBase<TModel>

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

RankingPredictionTransformer [](start = 24, length = 28)

Is the reason why we have two types that are identical in practically everything but name, so we can identify ranking estimators vs. regression estimators in a statically typed way?

@Zruty0Zruty0Sep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this transformer should also expose the group ID column name, at least that would be my belief


In reply to: 218214277 [](ancestors = 218214277)

@TomFinleyTomFinleySep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actually thought about this, like labels group ids are only needed for training, right? So for prediction I don't think they should be.


In reply to: 218216192 [](ancestors = 218216192,218214277)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So keep it, or make the Regression one Generic and use it for both?


In reply to: 218216839 [](ancestors = 218216839,218216192,218214277)

{
PredictorType = ComponentFactoryUtils.CreateFromFunction(
e => new AveragedPerceptronTrainer(e, new AveragedPerceptronTrainer.Arguments()))
e => new FastTreeBinaryClassificationTrainer(e, DefaultColumnNames.Label, DefaultColumnNames.Features))

@TomFinleyTomFinleySep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FastTreeBinaryClassificationTrainer [](start = 37, length = 35)

I'd really rather we didn't. This seems to fit into the same bucket as the discussion on #682. That ensembling should have a dependency on FastTree merely because we have a default does not make sense to me. If someone wants to use stacking, that's great, but they need to specify the learners. #Pending

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

But maybe we can hold off for right now.


In reply to: 218215145 [](ancestors = 218215145)

@sfilipisfilipiSep 17, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, let's do that separately, when we shape the ensembles to take in the arguments in the constructor.


In reply to: 218215323 [](ancestors = 218215323,218215145)

using Microsoft.ML.Runtime.TreePredictor;
using Newtonsoft.Json.Linq;
using Microsoft.ML.Core.Data;
using Microsoft.ML.Runtime.EntryPoints;

@TomFinleyTomFinleySep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm probably just missing something obvious, but why does this now depend on entry-points namespace?

Also sorting. #Resolved

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you! Oversight


In reply to: 218216150 [](ancestors = 218216150)

@TomFinleyTomFinley left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@TomFinley

Copy link
Copy Markdown
Contributor

Is omission of Pigsty extensions deliberate?

@sfilipi

Copy link
Copy Markdown
MemberAuthor

Did i misunderstand that for trainers we should hold on to doing the Pigsty extensions until we get the ml task, so we could extend on that, rather than the label? @tfinley@gmail.com@Zruty0, let me know if i should actually work on them in the same PR.


In reply to: 422161730 [](ancestors = 422161730)

@Zruty0Zruty0 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@Zruty0

Copy link
Copy Markdown
Contributor

I have the same (mis)understanding. In any case, let's do it after this oner


In reply to: 422174998 [](ancestors = 422174998,422161730)

@sfilipisfilipi changed the title WIP: Tree estimatorsTree estimatorsSep 18, 2018
@sfilipi
sfilipi merged commit d13b415 into dotnet:masterSep 19, 2018
@sfilipisfilipi mentioned this pull request Sep 21, 2018
@sfilipi
sfilipi deleted the fastTreeEstimators branch October 22, 2018 16:57
@ghostghost locked as resolved and limited conversation to collaborators Mar 29, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

APIIssues pertaining the friendly API

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@sfilipi@TomFinley@Zruty0
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Tree estimators - #855

Merged
sfilipi merged 19 commits into
dotnet:masterfrom
sfilipi:fastTreeEstimators
Sep 19, 2018
Merged

Tree estimators#855
sfilipi merged 19 commits into
dotnet:masterfrom
sfilipi:fastTreeEstimators

Conversation

@sfilipi

Copy link
Copy Markdown
Member

Ongoing work on converting the trainers to estimators. This PR converts the Tree -type Predictors.

@sfilipi

sfilipi commented Sep 7, 2018

Copy link
Copy Markdown
MemberAuthor

I will add tests next. We don't seem to have many ranking tests enabled :( #Resolved

@sfilipisfilipi self-assigned this Sep 7, 2018
@sfilipisfilipi added the API Issues pertaining the friendly API label Sep 7, 2018
@sfilipisfilipi added this to the 0918 milestone Sep 7, 2018
@sfilipisfilipi changed the title WIP: Fast tree estimatorsWIP: Tree estimatorsSep 7, 2018
@Zruty0Zruty0 mentioned this pull request Sep 7, 2018
}

protected override RankingPredictionTransformer<FastTreeRankingPredictor> MakeTransformer(FastTreeRankingPredictor model, ISchema trainSchema)
=> new RankingPredictionTransformer<FastTreeRankingPredictor>(Host, model, trainSchema, FeatureColumn.Name);

@sfilipisfilipiSep 7, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FeatureColumn.Name); [](start = 96, length = 20)

should add the GroupID to the base constructor #Resolved

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

GroupID? why?


In reply to: 216093781 [](ancestors = 216093781)

Changing the behavior for the creation of the weight column, based on whether it is explicit, or implicit.

private static SchemaShape.Column MakeWeightColumn(Optional<string> weightColumn)
{
if (weightColumn == null || !weightColumn.IsExplicit)

@sfilipisfilipiSep 8, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

|| !weightColumn.IsExplicit [](start = 37, length = 27)

this is not entirely correct either. It won't create the column when the user doesn't specify the weight colum, because it already had the name weight in the data.
we can't peak at the data at this time.

@tfinley@gmail.com@Zruty0 can we move from the Optional to just string for the weight, name, group ID and enforce the user typing in the names? is there another way around it, now that we need to know the information before seeing the data? #Resolved

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we cannot do this really, can we?


In reply to: 216135820 [](ancestors = 216135820)

/// (e.g., the prediction does not happen over a file as it did during training).
/// </summary>
[Fact]
public void New_SimpleTrainAndPredictWithFT()

@Zruty0Zruty0Sep 13, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

New_SimpleTrainAndPredictWithFT [](start = 20, length = 31)

move this test somewhere else #Resolved

using Microsoft.ML.Runtime.Internal.Utilities;
using Microsoft.ML.Runtime.Model;
using Microsoft.ML.Runtime.Internal.Internallearn;
using Microsoft.ML.Core.Data;

@Zruty0Zruty0Sep 13, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

using [](start = 0, length = 5)

sort #Resolved

{
new SchemaShape.Column(DefaultColumnNames.Score, SchemaShape.Column.VectorKind.Scalar, NumberType.R4, false),
new SchemaShape.Column(DefaultColumnNames.Probability, SchemaShape.Column.VectorKind.Scalar, NumberType.R4, false),
new SchemaShape.Column(DefaultColumnNames.PredictedLabel, SchemaShape.Column.VectorKind.Scalar, BoolType.Instance, false)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

double-check this is correct

Making use of dataset definitions
adding Iris.data and the adult.tiny files to TestDatasets
adding regression and ranking tests
/// FastTreeBinaryClassification TrainerEstimator test
/// </summary>
[Fact]
public void FastTreeRankerEstimator()

@sfilipisfilipiSep 14, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

public void FastTreeRankerEstimator() [](start = 7, length = 38)

this is currently failing. #Resolved

}
}

public sealed class RankingPredictionTransformer<TModel> : PredictionTransformerBase<TModel>

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

RankingPredictionTransformer [](start = 24, length = 28)

Is the reason why we have two types that are identical in practically everything but name, so we can identify ranking estimators vs. regression estimators in a statically typed way?

@Zruty0Zruty0Sep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this transformer should also expose the group ID column name, at least that would be my belief


In reply to: 218214277 [](ancestors = 218214277)

@TomFinleyTomFinleySep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actually thought about this, like labels group ids are only needed for training, right? So for prediction I don't think they should be.


In reply to: 218216192 [](ancestors = 218216192,218214277)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So keep it, or make the Regression one Generic and use it for both?


In reply to: 218216839 [](ancestors = 218216839,218216192,218214277)

{
PredictorType = ComponentFactoryUtils.CreateFromFunction(
e => new AveragedPerceptronTrainer(e, new AveragedPerceptronTrainer.Arguments()))
e => new FastTreeBinaryClassificationTrainer(e, DefaultColumnNames.Label, DefaultColumnNames.Features))

@TomFinleyTomFinleySep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FastTreeBinaryClassificationTrainer [](start = 37, length = 35)

I'd really rather we didn't. This seems to fit into the same bucket as the discussion on #682. That ensembling should have a dependency on FastTree merely because we have a default does not make sense to me. If someone wants to use stacking, that's great, but they need to specify the learners. #Pending

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

But maybe we can hold off for right now.


In reply to: 218215145 [](ancestors = 218215145)

@sfilipisfilipiSep 17, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, let's do that separately, when we shape the ensembles to take in the arguments in the constructor.


In reply to: 218215323 [](ancestors = 218215323,218215145)

using Microsoft.ML.Runtime.TreePredictor;
using Newtonsoft.Json.Linq;
using Microsoft.ML.Core.Data;
using Microsoft.ML.Runtime.EntryPoints;

@TomFinleyTomFinleySep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm probably just missing something obvious, but why does this now depend on entry-points namespace?

Also sorting. #Resolved

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you! Oversight


In reply to: 218216150 [](ancestors = 218216150)

@TomFinleyTomFinley left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@TomFinley

Copy link
Copy Markdown
Contributor

Is omission of Pigsty extensions deliberate?

@sfilipi

Copy link
Copy Markdown
MemberAuthor

Did i misunderstand that for trainers we should hold on to doing the Pigsty extensions until we get the ml task, so we could extend on that, rather than the label? @tfinley@gmail.com@Zruty0, let me know if i should actually work on them in the same PR.


In reply to: 422161730 [](ancestors = 422161730)

@Zruty0Zruty0 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@Zruty0

Copy link
Copy Markdown
Contributor

I have the same (mis)understanding. In any case, let's do it after this oner


In reply to: 422174998 [](ancestors = 422174998,422161730)

@sfilipisfilipi changed the title WIP: Tree estimatorsTree estimatorsSep 18, 2018
@sfilipi
sfilipi merged commit d13b415 into dotnet:masterSep 19, 2018
@sfilipisfilipi mentioned this pull request Sep 21, 2018
@sfilipi
sfilipi deleted the fastTreeEstimators branch October 22, 2018 16:57
@ghostghost locked as resolved and limited conversation to collaborators Mar 29, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

APIIssues pertaining the friendly API

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@sfilipi@TomFinley@Zruty0
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Tree estimators - #855

Merged
sfilipi merged 19 commits into
dotnet:masterfrom
sfilipi:fastTreeEstimators
Sep 19, 2018
Merged

Tree estimators#855
sfilipi merged 19 commits into
dotnet:masterfrom
sfilipi:fastTreeEstimators

Conversation

@sfilipi

Copy link
Copy Markdown
Member

Ongoing work on converting the trainers to estimators. This PR converts the Tree -type Predictors.

@sfilipi

sfilipi commented Sep 7, 2018

Copy link
Copy Markdown
MemberAuthor

I will add tests next. We don't seem to have many ranking tests enabled :( #Resolved

@sfilipisfilipi self-assigned this Sep 7, 2018
@sfilipisfilipi added the API Issues pertaining the friendly API label Sep 7, 2018
@sfilipisfilipi added this to the 0918 milestone Sep 7, 2018
@sfilipisfilipi changed the title WIP: Fast tree estimatorsWIP: Tree estimatorsSep 7, 2018
@Zruty0Zruty0 mentioned this pull request Sep 7, 2018
}

protected override RankingPredictionTransformer<FastTreeRankingPredictor> MakeTransformer(FastTreeRankingPredictor model, ISchema trainSchema)
=> new RankingPredictionTransformer<FastTreeRankingPredictor>(Host, model, trainSchema, FeatureColumn.Name);

@sfilipisfilipiSep 7, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FeatureColumn.Name); [](start = 96, length = 20)

should add the GroupID to the base constructor #Resolved

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

GroupID? why?


In reply to: 216093781 [](ancestors = 216093781)

Changing the behavior for the creation of the weight column, based on whether it is explicit, or implicit.

private static SchemaShape.Column MakeWeightColumn(Optional<string> weightColumn)
{
if (weightColumn == null || !weightColumn.IsExplicit)

@sfilipisfilipiSep 8, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

|| !weightColumn.IsExplicit [](start = 37, length = 27)

this is not entirely correct either. It won't create the column when the user doesn't specify the weight colum, because it already had the name weight in the data.
we can't peak at the data at this time.

@tfinley@gmail.com@Zruty0 can we move from the Optional to just string for the weight, name, group ID and enforce the user typing in the names? is there another way around it, now that we need to know the information before seeing the data? #Resolved

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we cannot do this really, can we?


In reply to: 216135820 [](ancestors = 216135820)

/// (e.g., the prediction does not happen over a file as it did during training).
/// </summary>
[Fact]
public void New_SimpleTrainAndPredictWithFT()

@Zruty0Zruty0Sep 13, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

New_SimpleTrainAndPredictWithFT [](start = 20, length = 31)

move this test somewhere else #Resolved

using Microsoft.ML.Runtime.Internal.Utilities;
using Microsoft.ML.Runtime.Model;
using Microsoft.ML.Runtime.Internal.Internallearn;
using Microsoft.ML.Core.Data;

@Zruty0Zruty0Sep 13, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

using [](start = 0, length = 5)

sort #Resolved

{
new SchemaShape.Column(DefaultColumnNames.Score, SchemaShape.Column.VectorKind.Scalar, NumberType.R4, false),
new SchemaShape.Column(DefaultColumnNames.Probability, SchemaShape.Column.VectorKind.Scalar, NumberType.R4, false),
new SchemaShape.Column(DefaultColumnNames.PredictedLabel, SchemaShape.Column.VectorKind.Scalar, BoolType.Instance, false)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

double-check this is correct

Making use of dataset definitions
adding Iris.data and the adult.tiny files to TestDatasets
adding regression and ranking tests
/// FastTreeBinaryClassification TrainerEstimator test
/// </summary>
[Fact]
public void FastTreeRankerEstimator()

@sfilipisfilipiSep 14, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

public void FastTreeRankerEstimator() [](start = 7, length = 38)

this is currently failing. #Resolved

}
}

public sealed class RankingPredictionTransformer<TModel> : PredictionTransformerBase<TModel>

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

RankingPredictionTransformer [](start = 24, length = 28)

Is the reason why we have two types that are identical in practically everything but name, so we can identify ranking estimators vs. regression estimators in a statically typed way?

@Zruty0Zruty0Sep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this transformer should also expose the group ID column name, at least that would be my belief


In reply to: 218214277 [](ancestors = 218214277)

@TomFinleyTomFinleySep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actually thought about this, like labels group ids are only needed for training, right? So for prediction I don't think they should be.


In reply to: 218216192 [](ancestors = 218216192,218214277)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So keep it, or make the Regression one Generic and use it for both?


In reply to: 218216839 [](ancestors = 218216839,218216192,218214277)

{
PredictorType = ComponentFactoryUtils.CreateFromFunction(
e => new AveragedPerceptronTrainer(e, new AveragedPerceptronTrainer.Arguments()))
e => new FastTreeBinaryClassificationTrainer(e, DefaultColumnNames.Label, DefaultColumnNames.Features))

@TomFinleyTomFinleySep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FastTreeBinaryClassificationTrainer [](start = 37, length = 35)

I'd really rather we didn't. This seems to fit into the same bucket as the discussion on #682. That ensembling should have a dependency on FastTree merely because we have a default does not make sense to me. If someone wants to use stacking, that's great, but they need to specify the learners. #Pending

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

But maybe we can hold off for right now.


In reply to: 218215145 [](ancestors = 218215145)

@sfilipisfilipiSep 17, 2018

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, let's do that separately, when we shape the ensembles to take in the arguments in the constructor.


In reply to: 218215323 [](ancestors = 218215323,218215145)

using Microsoft.ML.Runtime.TreePredictor;
using Newtonsoft.Json.Linq;
using Microsoft.ML.Core.Data;
using Microsoft.ML.Runtime.EntryPoints;

@TomFinleyTomFinleySep 17, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm probably just missing something obvious, but why does this now depend on entry-points namespace?

Also sorting. #Resolved

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you! Oversight


In reply to: 218216150 [](ancestors = 218216150)

@TomFinleyTomFinley left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@TomFinley

Copy link
Copy Markdown
Contributor

Is omission of Pigsty extensions deliberate?

@sfilipi

Copy link
Copy Markdown
MemberAuthor

Did i misunderstand that for trainers we should hold on to doing the Pigsty extensions until we get the ml task, so we could extend on that, rather than the label? @tfinley@gmail.com@Zruty0, let me know if i should actually work on them in the same PR.


In reply to: 422161730 [](ancestors = 422161730)

@Zruty0Zruty0 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@Zruty0

Copy link
Copy Markdown
Contributor

I have the same (mis)understanding. In any case, let's do it after this oner


In reply to: 422174998 [](ancestors = 422174998,422161730)

@sfilipisfilipi changed the title WIP: Tree estimatorsTree estimatorsSep 18, 2018
@sfilipi
sfilipi merged commit d13b415 into dotnet:masterSep 19, 2018
@sfilipisfilipi mentioned this pull request Sep 21, 2018
@sfilipi
sfilipi deleted the fastTreeEstimators branch October 22, 2018 16:57
@ghostghost locked as resolved and limited conversation to collaborators Mar 29, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

APIIssues pertaining the friendly API

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@sfilipi@TomFinley@Zruty0