Schema comprehension doc - #572

Merged
Zruty0 merged 2 commits into
dotnet:masterfrom
Zruty0:feature/554-schema-doc
Jul 24, 2018
Merged

Schema comprehension doc#572
Zruty0 merged 2 commits into
dotnet:masterfrom
Zruty0:feature/554-schema-doc

Conversation

@Zruty0

Copy link
Copy Markdown
Contributor

Added a document that describe typed schema comprehension.

Fixes#554

@dnfclas

dnfclas commented Jul 23, 2018

Copy link
Copy Markdown

CLA assistant check
All CLA requirements met. #Closed

@Zruty0

Zruty0 commented Jul 23, 2018

Copy link
Copy Markdown
ContributorAuthor

@dotnet-bot test Linux Release
#Closed

@eerhardt

eerhardt commented Jul 23, 2018

Copy link
Copy Markdown
Member

@Zruty0 - we are in the middle of moving our CI system from Jenkins to VSTS. You can ignore those 2 failed runs. (Plus you are just modifying .md files anyway.) #Resolved

@eerhardteerhardt left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for this helpful document, @Zruty0. It looks really good.

Comment threaddocs/code/SchemaComprehension.md Outdated

### Streaming data views

What if the original data doesn't support seeking, kile if it's some form of `IEnumerable<IrisData>` instead of `IList<IrisData>`? Well, we can simply use another helper function:

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(type-o) kile #Resolved

Comment threaddocs/code/SchemaComprehension.md Outdated

Let's see how we can create a new `IDataView` out of an in-memory array, run some operations on it, and then read it back into the array.

```(csharp)

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure what exactly about your string isn't working, but I don't get syntax highlighting when viewing the document.

Typically, I use the format ```C# instead. #Resolved

Comment threaddocs/code/SchemaComprehension.md Outdated
Below are the most notable examples of the differences:

* `IDataView` vector columns may have a fixed (and known) size, C# arrays can not. You can use `[VectorType(N)]` attribute to an array field to specify that the column is a vector of fixed size N. This is often necessary: most ML components don't work with variable-size vectors, they require fixed-size ones.
* `IDataView`'s **key types** don't have an underlying C# type either. To declare a key-type column, you need to make your field an `uint`, and decorate it with `[KeyType(Min=A, Count=B)]` to denote that the field is a key with the specified range of values.

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This makes it sound like Min and Count are required values, but they do not appear to be. (And I'm assuming there are plenty of scenarios where a user doesn't know up front what are all the possible values. #Resolved

Comment threaddocs/code/SchemaComprehension.md Outdated
var predictionEngine = env.CreatePredictionEngine<IrisData, IrisVectorData>(dv, outputSchemaDefinition: schemaDef);
```

In addition to the above, you can use `SchemaDefinition` to add per-column metadata, or even a 'value generator' (so that the column value is not read from the field, but computed using a delegate).

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It might be interesting to have a code snippet example for this scenario? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In fact, I think the 'generator' bit didn't make it into ML.NET


In reply to: 204536819 [](ancestors = 204536819)

Comment threaddocs/code/SchemaComprehension.md Outdated
* Reading a different subset of columns on every row: the cursor always populates the entire row object.
* Reading column metadata from the data view.
* Accessing the 'hidden' data view columns by index.
* Creating 'cursor sets'.

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A link or definition of cursor sets may be helpful here. #Resolved

@Zruty0

Copy link
Copy Markdown
ContributorAuthor

I hope they will still go away, because they block the merging.


In reply to: 407177294 [](ancestors = 407177294)

@@ -0,0 +1,210 @@
# Schema comprehension in ML.NET

This document describes in detail the under-the-hood mechanism that ML.NET uses to automate the creation of `IDataView` schema, with the goal to make it as convenient to the end user as possible, while not incurring extra computational costs.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

IDataView` schem [](start = 109, length = 16)

Might be useful to link to the IDV doc. #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated

## Introduction

Every dataset in ML.NET is an `IDataView`, which is, for the purposes of this document, a collection of rows that share the same columns. The set of columns, their names, types and other metadata is known as the *schema* of the `IDataView`, and it's represented as an `ISchema` object.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is an [](start = 24, length = 5)

would it be more clear if it says "gets loaded into an IDV" #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm, I think it's more correct to say 'is represented as', because you don't necessarily LOAD a dataset.


In reply to: 204819663 [](ancestors = 204819663)

Comment threaddocs/code/SchemaComprehension.md Outdated

## Introduction

Every dataset in ML.NET is an `IDataView`, which is, for the purposes of this document, a collection of rows that share the same columns. The set of columns, their names, types and other metadata is known as the *schema* of the `IDataView`, and it's represented as an `ISchema` object.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

schema [](start = 212, length = 9)

link to the schema section of the IDV Design Principles #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added one above


In reply to: 204820706 [](ancestors = 204820706)

These items above are very similar to the definition of fields in a C# class: names and types of columns correspond to names and types of fields, and metadata can correspond to field attributes.
Because of this similarity, ML.NET offers a common convenient mechanism for creating a schema: it is done via defining a C# class.

For example, the below class definition can be used to define a data view with 5 float columns:

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

data view [](start = 64, length = 9)

wondering if it would help to state at the beginning that IDataView and 'data view' are interchangeable, because you give the definition of one, and use the other term for it. #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated
.ToArray();
}
```
After this code runs, `arr` will contain two `IrisVectorData` objects, each having `Features` filled with the actual values of the features.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

features [](start = 131, length = 8)

I'd add (the 4 concatenated columns) after features, to make it more explicit. #Closed

```(csharp)
var streamingDv = env.CreateStreamingDataView<IrisData>(dataEnumerable);
```
The only subtle difference is, the resulting `streamingDv` will not support shuffling (a property that's useful to some ML application).

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

shuffling [](start = 76, length = 9)

Maybe link to what data shuffling is.
#Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure what to link to...


In reply to: 204823733 [](ancestors = 204823733)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah, the Wikipedia one isn't great..


In reply to: 204855447 [](ancestors = 204855447,204823733)

Comment threaddocs/code/SchemaComprehension.md Outdated
`IDataView` [type system](IDataViewTypeSystem.md) differs slightly from the C# type system, so a 1-1 mapping between column types and C# types is not always feasible.
Below are the most notable examples of the differences:

* `IDataView` vector columns may have a fixed (and known) size, C# arrays can not. You can use `[VectorType(N)]` attribute to an array field to specify that the column is a vector of fixed size N. This is often necessary: most ML components don't work with variable-size vectors, they require fixed-size ones.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

C# arrays can not [](start = 64, length = 17)

this might get confusing if you think about initialized arrays. #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tried to clarify


In reply to: 204831222 [](ancestors = 204831222)

Comment threaddocs/code/SchemaComprehension.md Outdated
Below are the most notable examples of the differences:

* `IDataView` vector columns may have a fixed (and known) size, C# arrays can not. You can use `[VectorType(N)]` attribute to an array field to specify that the column is a vector of fixed size N. This is often necessary: most ML components don't work with variable-size vectors, they require fixed-size ones.
* `IDataView`'s **key types** don't have an underlying C# type either. To declare a key-type column, you need to make your field an `uint`, and decorate it with `[KeyType(Min=A, Count=B)]` to denote that the field is a key with the specified range of values.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

*key types [](start = 16, length = 12)

link maybe #Closed

| `BL` | `DvBool` | `bool`, `bool?` |
| `TS` | `DvTimeSpan` | |
| `DT` | `DvDateTime` | |
| `DZ` | `DvDateTimeZone` | |

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

They don't map to TimeSpan, DateTime and DataTimeZone? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No


In reply to: 204836126 [](ancestors = 204836126)

It was our design decision to not allow these scenarios, thus simplifying the other, more common scenarios.

Here is the list of things that are only possible via the low-level interface:
* Creating or reading a data view, where even column *types* are not known at compile time (so you cannot create a C# class to define the schema)

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Creating or reading a data view, where even column types are not known at compile time (so you cannot create a C# class to define the schema) [](start = 2, length = 143)

example of scenario when this might occur #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated
* Creating or reading a data view, where even column *types* are not known at compile time (so you cannot create a C# class to define the schema)
* Reading a different subset of columns on every row: the cursor always populates the entire row object.
* Reading column metadata from the data view.
* Accessing the 'hidden' data view columns by index.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

'hidden' data view column [](start = 16, length = 25)

link or define "hidden" #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated

Here is the list of things that are only possible via the low-level interface:
* Creating or reading a data view, where even column *types* are not known at compile time (so you cannot create a C# class to define the schema)
* Reading a different subset of columns on every row: the cursor always populates the entire row object.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

different subset of columns on every row [](start = 12, length = 40)

what does 'different' mean here? #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tried to rephrase


In reply to: 204839399 [](ancestors = 204839399)

@sfilipisfilipi left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@Zruty0
Zruty0 merged commit 8cfa2ed into dotnet:masterJul 24, 2018
@Zruty0
Zruty0 deleted the feature/554-schema-doc branch July 24, 2018 19:32
eerhardt pushed a commit to eerhardt/machinelearning that referenced this pull request Jul 27, 2018
* Added a doc for schema comprehension
codemzs pushed a commit to codemzs/machinelearning that referenced this pull request Aug 1, 2018
* Added a doc for schema comprehension
@ghostghost locked as resolved and limited conversation to collaborators Mar 29, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Zruty0@dnfclas@eerhardt@sfilipi
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Schema comprehension doc - #572

Merged
Zruty0 merged 2 commits into
dotnet:masterfrom
Zruty0:feature/554-schema-doc
Jul 24, 2018
Merged

Schema comprehension doc#572
Zruty0 merged 2 commits into
dotnet:masterfrom
Zruty0:feature/554-schema-doc

Conversation

@Zruty0

Copy link
Copy Markdown
Contributor

Added a document that describe typed schema comprehension.

Fixes#554

@dnfclas

dnfclas commented Jul 23, 2018

Copy link
Copy Markdown

CLA assistant check
All CLA requirements met. #Closed

@Zruty0

Zruty0 commented Jul 23, 2018

Copy link
Copy Markdown
ContributorAuthor

@dotnet-bot test Linux Release
#Closed

@eerhardt

eerhardt commented Jul 23, 2018

Copy link
Copy Markdown
Member

@Zruty0 - we are in the middle of moving our CI system from Jenkins to VSTS. You can ignore those 2 failed runs. (Plus you are just modifying .md files anyway.) #Resolved

@eerhardteerhardt left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for this helpful document, @Zruty0. It looks really good.

Comment threaddocs/code/SchemaComprehension.md Outdated

### Streaming data views

What if the original data doesn't support seeking, kile if it's some form of `IEnumerable<IrisData>` instead of `IList<IrisData>`? Well, we can simply use another helper function:

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(type-o) kile #Resolved

Comment threaddocs/code/SchemaComprehension.md Outdated

Let's see how we can create a new `IDataView` out of an in-memory array, run some operations on it, and then read it back into the array.

```(csharp)

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure what exactly about your string isn't working, but I don't get syntax highlighting when viewing the document.

Typically, I use the format ```C# instead. #Resolved

Comment threaddocs/code/SchemaComprehension.md Outdated
Below are the most notable examples of the differences:

* `IDataView` vector columns may have a fixed (and known) size, C# arrays can not. You can use `[VectorType(N)]` attribute to an array field to specify that the column is a vector of fixed size N. This is often necessary: most ML components don't work with variable-size vectors, they require fixed-size ones.
* `IDataView`'s **key types** don't have an underlying C# type either. To declare a key-type column, you need to make your field an `uint`, and decorate it with `[KeyType(Min=A, Count=B)]` to denote that the field is a key with the specified range of values.

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This makes it sound like Min and Count are required values, but they do not appear to be. (And I'm assuming there are plenty of scenarios where a user doesn't know up front what are all the possible values. #Resolved

Comment threaddocs/code/SchemaComprehension.md Outdated
var predictionEngine = env.CreatePredictionEngine<IrisData, IrisVectorData>(dv, outputSchemaDefinition: schemaDef);
```

In addition to the above, you can use `SchemaDefinition` to add per-column metadata, or even a 'value generator' (so that the column value is not read from the field, but computed using a delegate).

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It might be interesting to have a code snippet example for this scenario? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In fact, I think the 'generator' bit didn't make it into ML.NET


In reply to: 204536819 [](ancestors = 204536819)

Comment threaddocs/code/SchemaComprehension.md Outdated
* Reading a different subset of columns on every row: the cursor always populates the entire row object.
* Reading column metadata from the data view.
* Accessing the 'hidden' data view columns by index.
* Creating 'cursor sets'.

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A link or definition of cursor sets may be helpful here. #Resolved

@Zruty0

Copy link
Copy Markdown
ContributorAuthor

I hope they will still go away, because they block the merging.


In reply to: 407177294 [](ancestors = 407177294)

@@ -0,0 +1,210 @@
# Schema comprehension in ML.NET

This document describes in detail the under-the-hood mechanism that ML.NET uses to automate the creation of `IDataView` schema, with the goal to make it as convenient to the end user as possible, while not incurring extra computational costs.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

IDataView` schem [](start = 109, length = 16)

Might be useful to link to the IDV doc. #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated

## Introduction

Every dataset in ML.NET is an `IDataView`, which is, for the purposes of this document, a collection of rows that share the same columns. The set of columns, their names, types and other metadata is known as the *schema* of the `IDataView`, and it's represented as an `ISchema` object.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is an [](start = 24, length = 5)

would it be more clear if it says "gets loaded into an IDV" #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm, I think it's more correct to say 'is represented as', because you don't necessarily LOAD a dataset.


In reply to: 204819663 [](ancestors = 204819663)

Comment threaddocs/code/SchemaComprehension.md Outdated

## Introduction

Every dataset in ML.NET is an `IDataView`, which is, for the purposes of this document, a collection of rows that share the same columns. The set of columns, their names, types and other metadata is known as the *schema* of the `IDataView`, and it's represented as an `ISchema` object.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

schema [](start = 212, length = 9)

link to the schema section of the IDV Design Principles #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added one above


In reply to: 204820706 [](ancestors = 204820706)

These items above are very similar to the definition of fields in a C# class: names and types of columns correspond to names and types of fields, and metadata can correspond to field attributes.
Because of this similarity, ML.NET offers a common convenient mechanism for creating a schema: it is done via defining a C# class.

For example, the below class definition can be used to define a data view with 5 float columns:

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

data view [](start = 64, length = 9)

wondering if it would help to state at the beginning that IDataView and 'data view' are interchangeable, because you give the definition of one, and use the other term for it. #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated
.ToArray();
}
```
After this code runs, `arr` will contain two `IrisVectorData` objects, each having `Features` filled with the actual values of the features.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

features [](start = 131, length = 8)

I'd add (the 4 concatenated columns) after features, to make it more explicit. #Closed

```(csharp)
var streamingDv = env.CreateStreamingDataView<IrisData>(dataEnumerable);
```
The only subtle difference is, the resulting `streamingDv` will not support shuffling (a property that's useful to some ML application).

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

shuffling [](start = 76, length = 9)

Maybe link to what data shuffling is.
#Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure what to link to...


In reply to: 204823733 [](ancestors = 204823733)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah, the Wikipedia one isn't great..


In reply to: 204855447 [](ancestors = 204855447,204823733)

Comment threaddocs/code/SchemaComprehension.md Outdated
`IDataView` [type system](IDataViewTypeSystem.md) differs slightly from the C# type system, so a 1-1 mapping between column types and C# types is not always feasible.
Below are the most notable examples of the differences:

* `IDataView` vector columns may have a fixed (and known) size, C# arrays can not. You can use `[VectorType(N)]` attribute to an array field to specify that the column is a vector of fixed size N. This is often necessary: most ML components don't work with variable-size vectors, they require fixed-size ones.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

C# arrays can not [](start = 64, length = 17)

this might get confusing if you think about initialized arrays. #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tried to clarify


In reply to: 204831222 [](ancestors = 204831222)

Comment threaddocs/code/SchemaComprehension.md Outdated
Below are the most notable examples of the differences:

* `IDataView` vector columns may have a fixed (and known) size, C# arrays can not. You can use `[VectorType(N)]` attribute to an array field to specify that the column is a vector of fixed size N. This is often necessary: most ML components don't work with variable-size vectors, they require fixed-size ones.
* `IDataView`'s **key types** don't have an underlying C# type either. To declare a key-type column, you need to make your field an `uint`, and decorate it with `[KeyType(Min=A, Count=B)]` to denote that the field is a key with the specified range of values.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

*key types [](start = 16, length = 12)

link maybe #Closed

| `BL` | `DvBool` | `bool`, `bool?` |
| `TS` | `DvTimeSpan` | |
| `DT` | `DvDateTime` | |
| `DZ` | `DvDateTimeZone` | |

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

They don't map to TimeSpan, DateTime and DataTimeZone? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No


In reply to: 204836126 [](ancestors = 204836126)

It was our design decision to not allow these scenarios, thus simplifying the other, more common scenarios.

Here is the list of things that are only possible via the low-level interface:
* Creating or reading a data view, where even column *types* are not known at compile time (so you cannot create a C# class to define the schema)

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Creating or reading a data view, where even column types are not known at compile time (so you cannot create a C# class to define the schema) [](start = 2, length = 143)

example of scenario when this might occur #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated
* Creating or reading a data view, where even column *types* are not known at compile time (so you cannot create a C# class to define the schema)
* Reading a different subset of columns on every row: the cursor always populates the entire row object.
* Reading column metadata from the data view.
* Accessing the 'hidden' data view columns by index.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

'hidden' data view column [](start = 16, length = 25)

link or define "hidden" #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated

Here is the list of things that are only possible via the low-level interface:
* Creating or reading a data view, where even column *types* are not known at compile time (so you cannot create a C# class to define the schema)
* Reading a different subset of columns on every row: the cursor always populates the entire row object.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

different subset of columns on every row [](start = 12, length = 40)

what does 'different' mean here? #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tried to rephrase


In reply to: 204839399 [](ancestors = 204839399)

@sfilipisfilipi left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@Zruty0
Zruty0 merged commit 8cfa2ed into dotnet:masterJul 24, 2018
@Zruty0
Zruty0 deleted the feature/554-schema-doc branch July 24, 2018 19:32
eerhardt pushed a commit to eerhardt/machinelearning that referenced this pull request Jul 27, 2018
* Added a doc for schema comprehension
codemzs pushed a commit to codemzs/machinelearning that referenced this pull request Aug 1, 2018
* Added a doc for schema comprehension
@ghostghost locked as resolved and limited conversation to collaborators Mar 29, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Zruty0@dnfclas@eerhardt@sfilipi
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Schema comprehension doc - #572

Merged
Zruty0 merged 2 commits into
dotnet:masterfrom
Zruty0:feature/554-schema-doc
Jul 24, 2018
Merged

Schema comprehension doc#572
Zruty0 merged 2 commits into
dotnet:masterfrom
Zruty0:feature/554-schema-doc

Conversation

@Zruty0

Copy link
Copy Markdown
Contributor

Added a document that describe typed schema comprehension.

Fixes#554

@dnfclas

dnfclas commented Jul 23, 2018

Copy link
Copy Markdown

CLA assistant check
All CLA requirements met. #Closed

@Zruty0

Zruty0 commented Jul 23, 2018

Copy link
Copy Markdown
ContributorAuthor

@dotnet-bot test Linux Release
#Closed

@eerhardt

eerhardt commented Jul 23, 2018

Copy link
Copy Markdown
Member

@Zruty0 - we are in the middle of moving our CI system from Jenkins to VSTS. You can ignore those 2 failed runs. (Plus you are just modifying .md files anyway.) #Resolved

@eerhardteerhardt left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for this helpful document, @Zruty0. It looks really good.

Comment threaddocs/code/SchemaComprehension.md Outdated

### Streaming data views

What if the original data doesn't support seeking, kile if it's some form of `IEnumerable<IrisData>` instead of `IList<IrisData>`? Well, we can simply use another helper function:

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(type-o) kile #Resolved

Comment threaddocs/code/SchemaComprehension.md Outdated

Let's see how we can create a new `IDataView` out of an in-memory array, run some operations on it, and then read it back into the array.

```(csharp)

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure what exactly about your string isn't working, but I don't get syntax highlighting when viewing the document.

Typically, I use the format ```C# instead. #Resolved

Comment threaddocs/code/SchemaComprehension.md Outdated
Below are the most notable examples of the differences:

* `IDataView` vector columns may have a fixed (and known) size, C# arrays can not. You can use `[VectorType(N)]` attribute to an array field to specify that the column is a vector of fixed size N. This is often necessary: most ML components don't work with variable-size vectors, they require fixed-size ones.
* `IDataView`'s **key types** don't have an underlying C# type either. To declare a key-type column, you need to make your field an `uint`, and decorate it with `[KeyType(Min=A, Count=B)]` to denote that the field is a key with the specified range of values.

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This makes it sound like Min and Count are required values, but they do not appear to be. (And I'm assuming there are plenty of scenarios where a user doesn't know up front what are all the possible values. #Resolved

Comment threaddocs/code/SchemaComprehension.md Outdated
var predictionEngine = env.CreatePredictionEngine<IrisData, IrisVectorData>(dv, outputSchemaDefinition: schemaDef);
```

In addition to the above, you can use `SchemaDefinition` to add per-column metadata, or even a 'value generator' (so that the column value is not read from the field, but computed using a delegate).

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It might be interesting to have a code snippet example for this scenario? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In fact, I think the 'generator' bit didn't make it into ML.NET


In reply to: 204536819 [](ancestors = 204536819)

Comment threaddocs/code/SchemaComprehension.md Outdated
* Reading a different subset of columns on every row: the cursor always populates the entire row object.
* Reading column metadata from the data view.
* Accessing the 'hidden' data view columns by index.
* Creating 'cursor sets'.

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A link or definition of cursor sets may be helpful here. #Resolved

@Zruty0

Copy link
Copy Markdown
ContributorAuthor

I hope they will still go away, because they block the merging.


In reply to: 407177294 [](ancestors = 407177294)

@@ -0,0 +1,210 @@
# Schema comprehension in ML.NET

This document describes in detail the under-the-hood mechanism that ML.NET uses to automate the creation of `IDataView` schema, with the goal to make it as convenient to the end user as possible, while not incurring extra computational costs.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

IDataView` schem [](start = 109, length = 16)

Might be useful to link to the IDV doc. #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated

## Introduction

Every dataset in ML.NET is an `IDataView`, which is, for the purposes of this document, a collection of rows that share the same columns. The set of columns, their names, types and other metadata is known as the *schema* of the `IDataView`, and it's represented as an `ISchema` object.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is an [](start = 24, length = 5)

would it be more clear if it says "gets loaded into an IDV" #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm, I think it's more correct to say 'is represented as', because you don't necessarily LOAD a dataset.


In reply to: 204819663 [](ancestors = 204819663)

Comment threaddocs/code/SchemaComprehension.md Outdated

## Introduction

Every dataset in ML.NET is an `IDataView`, which is, for the purposes of this document, a collection of rows that share the same columns. The set of columns, their names, types and other metadata is known as the *schema* of the `IDataView`, and it's represented as an `ISchema` object.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

schema [](start = 212, length = 9)

link to the schema section of the IDV Design Principles #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added one above


In reply to: 204820706 [](ancestors = 204820706)

These items above are very similar to the definition of fields in a C# class: names and types of columns correspond to names and types of fields, and metadata can correspond to field attributes.
Because of this similarity, ML.NET offers a common convenient mechanism for creating a schema: it is done via defining a C# class.

For example, the below class definition can be used to define a data view with 5 float columns:

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

data view [](start = 64, length = 9)

wondering if it would help to state at the beginning that IDataView and 'data view' are interchangeable, because you give the definition of one, and use the other term for it. #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated
.ToArray();
}
```
After this code runs, `arr` will contain two `IrisVectorData` objects, each having `Features` filled with the actual values of the features.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

features [](start = 131, length = 8)

I'd add (the 4 concatenated columns) after features, to make it more explicit. #Closed

```(csharp)
var streamingDv = env.CreateStreamingDataView<IrisData>(dataEnumerable);
```
The only subtle difference is, the resulting `streamingDv` will not support shuffling (a property that's useful to some ML application).

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

shuffling [](start = 76, length = 9)

Maybe link to what data shuffling is.
#Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure what to link to...


In reply to: 204823733 [](ancestors = 204823733)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah, the Wikipedia one isn't great..


In reply to: 204855447 [](ancestors = 204855447,204823733)

Comment threaddocs/code/SchemaComprehension.md Outdated
`IDataView` [type system](IDataViewTypeSystem.md) differs slightly from the C# type system, so a 1-1 mapping between column types and C# types is not always feasible.
Below are the most notable examples of the differences:

* `IDataView` vector columns may have a fixed (and known) size, C# arrays can not. You can use `[VectorType(N)]` attribute to an array field to specify that the column is a vector of fixed size N. This is often necessary: most ML components don't work with variable-size vectors, they require fixed-size ones.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

C# arrays can not [](start = 64, length = 17)

this might get confusing if you think about initialized arrays. #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tried to clarify


In reply to: 204831222 [](ancestors = 204831222)

Comment threaddocs/code/SchemaComprehension.md Outdated
Below are the most notable examples of the differences:

* `IDataView` vector columns may have a fixed (and known) size, C# arrays can not. You can use `[VectorType(N)]` attribute to an array field to specify that the column is a vector of fixed size N. This is often necessary: most ML components don't work with variable-size vectors, they require fixed-size ones.
* `IDataView`'s **key types** don't have an underlying C# type either. To declare a key-type column, you need to make your field an `uint`, and decorate it with `[KeyType(Min=A, Count=B)]` to denote that the field is a key with the specified range of values.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

*key types [](start = 16, length = 12)

link maybe #Closed

| `BL` | `DvBool` | `bool`, `bool?` |
| `TS` | `DvTimeSpan` | |
| `DT` | `DvDateTime` | |
| `DZ` | `DvDateTimeZone` | |

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

They don't map to TimeSpan, DateTime and DataTimeZone? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No


In reply to: 204836126 [](ancestors = 204836126)

It was our design decision to not allow these scenarios, thus simplifying the other, more common scenarios.

Here is the list of things that are only possible via the low-level interface:
* Creating or reading a data view, where even column *types* are not known at compile time (so you cannot create a C# class to define the schema)

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Creating or reading a data view, where even column types are not known at compile time (so you cannot create a C# class to define the schema) [](start = 2, length = 143)

example of scenario when this might occur #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated
* Creating or reading a data view, where even column *types* are not known at compile time (so you cannot create a C# class to define the schema)
* Reading a different subset of columns on every row: the cursor always populates the entire row object.
* Reading column metadata from the data view.
* Accessing the 'hidden' data view columns by index.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

'hidden' data view column [](start = 16, length = 25)

link or define "hidden" #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated

Here is the list of things that are only possible via the low-level interface:
* Creating or reading a data view, where even column *types* are not known at compile time (so you cannot create a C# class to define the schema)
* Reading a different subset of columns on every row: the cursor always populates the entire row object.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

different subset of columns on every row [](start = 12, length = 40)

what does 'different' mean here? #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tried to rephrase


In reply to: 204839399 [](ancestors = 204839399)

@sfilipisfilipi left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@Zruty0
Zruty0 merged commit 8cfa2ed into dotnet:masterJul 24, 2018
@Zruty0
Zruty0 deleted the feature/554-schema-doc branch July 24, 2018 19:32
eerhardt pushed a commit to eerhardt/machinelearning that referenced this pull request Jul 27, 2018
* Added a doc for schema comprehension
codemzs pushed a commit to codemzs/machinelearning that referenced this pull request Aug 1, 2018
* Added a doc for schema comprehension
@ghostghost locked as resolved and limited conversation to collaborators Mar 29, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Zruty0@dnfclas@eerhardt@sfilipi
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Schema comprehension doc - #572

Merged
Zruty0 merged 2 commits into
dotnet:masterfrom
Zruty0:feature/554-schema-doc
Jul 24, 2018
Merged

Schema comprehension doc#572
Zruty0 merged 2 commits into
dotnet:masterfrom
Zruty0:feature/554-schema-doc

Conversation

@Zruty0

Copy link
Copy Markdown
Contributor

Added a document that describe typed schema comprehension.

Fixes#554

@dnfclas

dnfclas commented Jul 23, 2018

Copy link
Copy Markdown

CLA assistant check
All CLA requirements met. #Closed

@Zruty0

Zruty0 commented Jul 23, 2018

Copy link
Copy Markdown
ContributorAuthor

@dotnet-bot test Linux Release
#Closed

@eerhardt

eerhardt commented Jul 23, 2018

Copy link
Copy Markdown
Member

@Zruty0 - we are in the middle of moving our CI system from Jenkins to VSTS. You can ignore those 2 failed runs. (Plus you are just modifying .md files anyway.) #Resolved

@eerhardteerhardt left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for this helpful document, @Zruty0. It looks really good.

Comment threaddocs/code/SchemaComprehension.md Outdated

### Streaming data views

What if the original data doesn't support seeking, kile if it's some form of `IEnumerable<IrisData>` instead of `IList<IrisData>`? Well, we can simply use another helper function:

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(type-o) kile #Resolved

Comment threaddocs/code/SchemaComprehension.md Outdated

Let's see how we can create a new `IDataView` out of an in-memory array, run some operations on it, and then read it back into the array.

```(csharp)

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure what exactly about your string isn't working, but I don't get syntax highlighting when viewing the document.

Typically, I use the format ```C# instead. #Resolved

Comment threaddocs/code/SchemaComprehension.md Outdated
Below are the most notable examples of the differences:

* `IDataView` vector columns may have a fixed (and known) size, C# arrays can not. You can use `[VectorType(N)]` attribute to an array field to specify that the column is a vector of fixed size N. This is often necessary: most ML components don't work with variable-size vectors, they require fixed-size ones.
* `IDataView`'s **key types** don't have an underlying C# type either. To declare a key-type column, you need to make your field an `uint`, and decorate it with `[KeyType(Min=A, Count=B)]` to denote that the field is a key with the specified range of values.

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This makes it sound like Min and Count are required values, but they do not appear to be. (And I'm assuming there are plenty of scenarios where a user doesn't know up front what are all the possible values. #Resolved

Comment threaddocs/code/SchemaComprehension.md Outdated
var predictionEngine = env.CreatePredictionEngine<IrisData, IrisVectorData>(dv, outputSchemaDefinition: schemaDef);
```

In addition to the above, you can use `SchemaDefinition` to add per-column metadata, or even a 'value generator' (so that the column value is not read from the field, but computed using a delegate).

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It might be interesting to have a code snippet example for this scenario? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In fact, I think the 'generator' bit didn't make it into ML.NET


In reply to: 204536819 [](ancestors = 204536819)

Comment threaddocs/code/SchemaComprehension.md Outdated
* Reading a different subset of columns on every row: the cursor always populates the entire row object.
* Reading column metadata from the data view.
* Accessing the 'hidden' data view columns by index.
* Creating 'cursor sets'.

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A link or definition of cursor sets may be helpful here. #Resolved

@Zruty0

Copy link
Copy Markdown
ContributorAuthor

I hope they will still go away, because they block the merging.


In reply to: 407177294 [](ancestors = 407177294)

@@ -0,0 +1,210 @@
# Schema comprehension in ML.NET

This document describes in detail the under-the-hood mechanism that ML.NET uses to automate the creation of `IDataView` schema, with the goal to make it as convenient to the end user as possible, while not incurring extra computational costs.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

IDataView` schem [](start = 109, length = 16)

Might be useful to link to the IDV doc. #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated

## Introduction

Every dataset in ML.NET is an `IDataView`, which is, for the purposes of this document, a collection of rows that share the same columns. The set of columns, their names, types and other metadata is known as the *schema* of the `IDataView`, and it's represented as an `ISchema` object.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is an [](start = 24, length = 5)

would it be more clear if it says "gets loaded into an IDV" #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm, I think it's more correct to say 'is represented as', because you don't necessarily LOAD a dataset.


In reply to: 204819663 [](ancestors = 204819663)

Comment threaddocs/code/SchemaComprehension.md Outdated

## Introduction

Every dataset in ML.NET is an `IDataView`, which is, for the purposes of this document, a collection of rows that share the same columns. The set of columns, their names, types and other metadata is known as the *schema* of the `IDataView`, and it's represented as an `ISchema` object.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

schema [](start = 212, length = 9)

link to the schema section of the IDV Design Principles #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added one above


In reply to: 204820706 [](ancestors = 204820706)

These items above are very similar to the definition of fields in a C# class: names and types of columns correspond to names and types of fields, and metadata can correspond to field attributes.
Because of this similarity, ML.NET offers a common convenient mechanism for creating a schema: it is done via defining a C# class.

For example, the below class definition can be used to define a data view with 5 float columns:

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

data view [](start = 64, length = 9)

wondering if it would help to state at the beginning that IDataView and 'data view' are interchangeable, because you give the definition of one, and use the other term for it. #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated
.ToArray();
}
```
After this code runs, `arr` will contain two `IrisVectorData` objects, each having `Features` filled with the actual values of the features.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

features [](start = 131, length = 8)

I'd add (the 4 concatenated columns) after features, to make it more explicit. #Closed

```(csharp)
var streamingDv = env.CreateStreamingDataView<IrisData>(dataEnumerable);
```
The only subtle difference is, the resulting `streamingDv` will not support shuffling (a property that's useful to some ML application).

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

shuffling [](start = 76, length = 9)

Maybe link to what data shuffling is.
#Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure what to link to...


In reply to: 204823733 [](ancestors = 204823733)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah, the Wikipedia one isn't great..


In reply to: 204855447 [](ancestors = 204855447,204823733)

Comment threaddocs/code/SchemaComprehension.md Outdated
`IDataView` [type system](IDataViewTypeSystem.md) differs slightly from the C# type system, so a 1-1 mapping between column types and C# types is not always feasible.
Below are the most notable examples of the differences:

* `IDataView` vector columns may have a fixed (and known) size, C# arrays can not. You can use `[VectorType(N)]` attribute to an array field to specify that the column is a vector of fixed size N. This is often necessary: most ML components don't work with variable-size vectors, they require fixed-size ones.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

C# arrays can not [](start = 64, length = 17)

this might get confusing if you think about initialized arrays. #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tried to clarify


In reply to: 204831222 [](ancestors = 204831222)

Comment threaddocs/code/SchemaComprehension.md Outdated
Below are the most notable examples of the differences:

* `IDataView` vector columns may have a fixed (and known) size, C# arrays can not. You can use `[VectorType(N)]` attribute to an array field to specify that the column is a vector of fixed size N. This is often necessary: most ML components don't work with variable-size vectors, they require fixed-size ones.
* `IDataView`'s **key types** don't have an underlying C# type either. To declare a key-type column, you need to make your field an `uint`, and decorate it with `[KeyType(Min=A, Count=B)]` to denote that the field is a key with the specified range of values.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

*key types [](start = 16, length = 12)

link maybe #Closed

| `BL` | `DvBool` | `bool`, `bool?` |
| `TS` | `DvTimeSpan` | |
| `DT` | `DvDateTime` | |
| `DZ` | `DvDateTimeZone` | |

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

They don't map to TimeSpan, DateTime and DataTimeZone? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No


In reply to: 204836126 [](ancestors = 204836126)

It was our design decision to not allow these scenarios, thus simplifying the other, more common scenarios.

Here is the list of things that are only possible via the low-level interface:
* Creating or reading a data view, where even column *types* are not known at compile time (so you cannot create a C# class to define the schema)

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Creating or reading a data view, where even column types are not known at compile time (so you cannot create a C# class to define the schema) [](start = 2, length = 143)

example of scenario when this might occur #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated
* Creating or reading a data view, where even column *types* are not known at compile time (so you cannot create a C# class to define the schema)
* Reading a different subset of columns on every row: the cursor always populates the entire row object.
* Reading column metadata from the data view.
* Accessing the 'hidden' data view columns by index.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

'hidden' data view column [](start = 16, length = 25)

link or define "hidden" #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated

Here is the list of things that are only possible via the low-level interface:
* Creating or reading a data view, where even column *types* are not known at compile time (so you cannot create a C# class to define the schema)
* Reading a different subset of columns on every row: the cursor always populates the entire row object.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

different subset of columns on every row [](start = 12, length = 40)

what does 'different' mean here? #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tried to rephrase


In reply to: 204839399 [](ancestors = 204839399)

@sfilipisfilipi left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@Zruty0
Zruty0 merged commit 8cfa2ed into dotnet:masterJul 24, 2018
@Zruty0
Zruty0 deleted the feature/554-schema-doc branch July 24, 2018 19:32
eerhardt pushed a commit to eerhardt/machinelearning that referenced this pull request Jul 27, 2018
* Added a doc for schema comprehension
codemzs pushed a commit to codemzs/machinelearning that referenced this pull request Aug 1, 2018
* Added a doc for schema comprehension
@ghostghost locked as resolved and limited conversation to collaborators Mar 29, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Zruty0@dnfclas@eerhardt@sfilipi
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Schema comprehension doc - #572

Merged
Zruty0 merged 2 commits into
dotnet:masterfrom
Zruty0:feature/554-schema-doc
Jul 24, 2018
Merged

Schema comprehension doc#572
Zruty0 merged 2 commits into
dotnet:masterfrom
Zruty0:feature/554-schema-doc

Conversation

@Zruty0

Copy link
Copy Markdown
Contributor

Added a document that describe typed schema comprehension.

Fixes#554

@dnfclas

dnfclas commented Jul 23, 2018

Copy link
Copy Markdown

CLA assistant check
All CLA requirements met. #Closed

@Zruty0

Zruty0 commented Jul 23, 2018

Copy link
Copy Markdown
ContributorAuthor

@dotnet-bot test Linux Release
#Closed

@eerhardt

eerhardt commented Jul 23, 2018

Copy link
Copy Markdown
Member

@Zruty0 - we are in the middle of moving our CI system from Jenkins to VSTS. You can ignore those 2 failed runs. (Plus you are just modifying .md files anyway.) #Resolved

@eerhardteerhardt left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for this helpful document, @Zruty0. It looks really good.

Comment threaddocs/code/SchemaComprehension.md Outdated

### Streaming data views

What if the original data doesn't support seeking, kile if it's some form of `IEnumerable<IrisData>` instead of `IList<IrisData>`? Well, we can simply use another helper function:

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(type-o) kile #Resolved

Comment threaddocs/code/SchemaComprehension.md Outdated

Let's see how we can create a new `IDataView` out of an in-memory array, run some operations on it, and then read it back into the array.

```(csharp)

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure what exactly about your string isn't working, but I don't get syntax highlighting when viewing the document.

Typically, I use the format ```C# instead. #Resolved

Comment threaddocs/code/SchemaComprehension.md Outdated
Below are the most notable examples of the differences:

* `IDataView` vector columns may have a fixed (and known) size, C# arrays can not. You can use `[VectorType(N)]` attribute to an array field to specify that the column is a vector of fixed size N. This is often necessary: most ML components don't work with variable-size vectors, they require fixed-size ones.
* `IDataView`'s **key types** don't have an underlying C# type either. To declare a key-type column, you need to make your field an `uint`, and decorate it with `[KeyType(Min=A, Count=B)]` to denote that the field is a key with the specified range of values.

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This makes it sound like Min and Count are required values, but they do not appear to be. (And I'm assuming there are plenty of scenarios where a user doesn't know up front what are all the possible values. #Resolved

Comment threaddocs/code/SchemaComprehension.md Outdated
var predictionEngine = env.CreatePredictionEngine<IrisData, IrisVectorData>(dv, outputSchemaDefinition: schemaDef);
```

In addition to the above, you can use `SchemaDefinition` to add per-column metadata, or even a 'value generator' (so that the column value is not read from the field, but computed using a delegate).

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It might be interesting to have a code snippet example for this scenario? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In fact, I think the 'generator' bit didn't make it into ML.NET


In reply to: 204536819 [](ancestors = 204536819)

Comment threaddocs/code/SchemaComprehension.md Outdated
* Reading a different subset of columns on every row: the cursor always populates the entire row object.
* Reading column metadata from the data view.
* Accessing the 'hidden' data view columns by index.
* Creating 'cursor sets'.

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A link or definition of cursor sets may be helpful here. #Resolved

@Zruty0

Copy link
Copy Markdown
ContributorAuthor

I hope they will still go away, because they block the merging.


In reply to: 407177294 [](ancestors = 407177294)

@@ -0,0 +1,210 @@
# Schema comprehension in ML.NET

This document describes in detail the under-the-hood mechanism that ML.NET uses to automate the creation of `IDataView` schema, with the goal to make it as convenient to the end user as possible, while not incurring extra computational costs.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

IDataView` schem [](start = 109, length = 16)

Might be useful to link to the IDV doc. #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated

## Introduction

Every dataset in ML.NET is an `IDataView`, which is, for the purposes of this document, a collection of rows that share the same columns. The set of columns, their names, types and other metadata is known as the *schema* of the `IDataView`, and it's represented as an `ISchema` object.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is an [](start = 24, length = 5)

would it be more clear if it says "gets loaded into an IDV" #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm, I think it's more correct to say 'is represented as', because you don't necessarily LOAD a dataset.


In reply to: 204819663 [](ancestors = 204819663)

Comment threaddocs/code/SchemaComprehension.md Outdated

## Introduction

Every dataset in ML.NET is an `IDataView`, which is, for the purposes of this document, a collection of rows that share the same columns. The set of columns, their names, types and other metadata is known as the *schema* of the `IDataView`, and it's represented as an `ISchema` object.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

schema [](start = 212, length = 9)

link to the schema section of the IDV Design Principles #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added one above


In reply to: 204820706 [](ancestors = 204820706)

These items above are very similar to the definition of fields in a C# class: names and types of columns correspond to names and types of fields, and metadata can correspond to field attributes.
Because of this similarity, ML.NET offers a common convenient mechanism for creating a schema: it is done via defining a C# class.

For example, the below class definition can be used to define a data view with 5 float columns:

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

data view [](start = 64, length = 9)

wondering if it would help to state at the beginning that IDataView and 'data view' are interchangeable, because you give the definition of one, and use the other term for it. #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated
.ToArray();
}
```
After this code runs, `arr` will contain two `IrisVectorData` objects, each having `Features` filled with the actual values of the features.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

features [](start = 131, length = 8)

I'd add (the 4 concatenated columns) after features, to make it more explicit. #Closed

```(csharp)
var streamingDv = env.CreateStreamingDataView<IrisData>(dataEnumerable);
```
The only subtle difference is, the resulting `streamingDv` will not support shuffling (a property that's useful to some ML application).

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

shuffling [](start = 76, length = 9)

Maybe link to what data shuffling is.
#Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure what to link to...


In reply to: 204823733 [](ancestors = 204823733)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah, the Wikipedia one isn't great..


In reply to: 204855447 [](ancestors = 204855447,204823733)

Comment threaddocs/code/SchemaComprehension.md Outdated
`IDataView` [type system](IDataViewTypeSystem.md) differs slightly from the C# type system, so a 1-1 mapping between column types and C# types is not always feasible.
Below are the most notable examples of the differences:

* `IDataView` vector columns may have a fixed (and known) size, C# arrays can not. You can use `[VectorType(N)]` attribute to an array field to specify that the column is a vector of fixed size N. This is often necessary: most ML components don't work with variable-size vectors, they require fixed-size ones.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

C# arrays can not [](start = 64, length = 17)

this might get confusing if you think about initialized arrays. #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tried to clarify


In reply to: 204831222 [](ancestors = 204831222)

Comment threaddocs/code/SchemaComprehension.md Outdated
Below are the most notable examples of the differences:

* `IDataView` vector columns may have a fixed (and known) size, C# arrays can not. You can use `[VectorType(N)]` attribute to an array field to specify that the column is a vector of fixed size N. This is often necessary: most ML components don't work with variable-size vectors, they require fixed-size ones.
* `IDataView`'s **key types** don't have an underlying C# type either. To declare a key-type column, you need to make your field an `uint`, and decorate it with `[KeyType(Min=A, Count=B)]` to denote that the field is a key with the specified range of values.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

*key types [](start = 16, length = 12)

link maybe #Closed

| `BL` | `DvBool` | `bool`, `bool?` |
| `TS` | `DvTimeSpan` | |
| `DT` | `DvDateTime` | |
| `DZ` | `DvDateTimeZone` | |

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

They don't map to TimeSpan, DateTime and DataTimeZone? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No


In reply to: 204836126 [](ancestors = 204836126)

It was our design decision to not allow these scenarios, thus simplifying the other, more common scenarios.

Here is the list of things that are only possible via the low-level interface:
* Creating or reading a data view, where even column *types* are not known at compile time (so you cannot create a C# class to define the schema)

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Creating or reading a data view, where even column types are not known at compile time (so you cannot create a C# class to define the schema) [](start = 2, length = 143)

example of scenario when this might occur #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated
* Creating or reading a data view, where even column *types* are not known at compile time (so you cannot create a C# class to define the schema)
* Reading a different subset of columns on every row: the cursor always populates the entire row object.
* Reading column metadata from the data view.
* Accessing the 'hidden' data view columns by index.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

'hidden' data view column [](start = 16, length = 25)

link or define "hidden" #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated

Here is the list of things that are only possible via the low-level interface:
* Creating or reading a data view, where even column *types* are not known at compile time (so you cannot create a C# class to define the schema)
* Reading a different subset of columns on every row: the cursor always populates the entire row object.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

different subset of columns on every row [](start = 12, length = 40)

what does 'different' mean here? #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tried to rephrase


In reply to: 204839399 [](ancestors = 204839399)

@sfilipisfilipi left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@Zruty0
Zruty0 merged commit 8cfa2ed into dotnet:masterJul 24, 2018
@Zruty0
Zruty0 deleted the feature/554-schema-doc branch July 24, 2018 19:32
eerhardt pushed a commit to eerhardt/machinelearning that referenced this pull request Jul 27, 2018
* Added a doc for schema comprehension
codemzs pushed a commit to codemzs/machinelearning that referenced this pull request Aug 1, 2018
* Added a doc for schema comprehension
@ghostghost locked as resolved and limited conversation to collaborators Mar 29, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Zruty0@dnfclas@eerhardt@sfilipi
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Schema comprehension doc - #572

Merged
Zruty0 merged 2 commits into
dotnet:masterfrom
Zruty0:feature/554-schema-doc
Jul 24, 2018
Merged

Schema comprehension doc#572
Zruty0 merged 2 commits into
dotnet:masterfrom
Zruty0:feature/554-schema-doc

Conversation

@Zruty0

Copy link
Copy Markdown
Contributor

Added a document that describe typed schema comprehension.

Fixes#554

@dnfclas

dnfclas commented Jul 23, 2018

Copy link
Copy Markdown

CLA assistant check
All CLA requirements met. #Closed

@Zruty0

Zruty0 commented Jul 23, 2018

Copy link
Copy Markdown
ContributorAuthor

@dotnet-bot test Linux Release
#Closed

@eerhardt

eerhardt commented Jul 23, 2018

Copy link
Copy Markdown
Member

@Zruty0 - we are in the middle of moving our CI system from Jenkins to VSTS. You can ignore those 2 failed runs. (Plus you are just modifying .md files anyway.) #Resolved

@eerhardteerhardt left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for this helpful document, @Zruty0. It looks really good.

Comment threaddocs/code/SchemaComprehension.md Outdated

### Streaming data views

What if the original data doesn't support seeking, kile if it's some form of `IEnumerable<IrisData>` instead of `IList<IrisData>`? Well, we can simply use another helper function:

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(type-o) kile #Resolved

Comment threaddocs/code/SchemaComprehension.md Outdated

Let's see how we can create a new `IDataView` out of an in-memory array, run some operations on it, and then read it back into the array.

```(csharp)

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure what exactly about your string isn't working, but I don't get syntax highlighting when viewing the document.

Typically, I use the format ```C# instead. #Resolved

Comment threaddocs/code/SchemaComprehension.md Outdated
Below are the most notable examples of the differences:

* `IDataView` vector columns may have a fixed (and known) size, C# arrays can not. You can use `[VectorType(N)]` attribute to an array field to specify that the column is a vector of fixed size N. This is often necessary: most ML components don't work with variable-size vectors, they require fixed-size ones.
* `IDataView`'s **key types** don't have an underlying C# type either. To declare a key-type column, you need to make your field an `uint`, and decorate it with `[KeyType(Min=A, Count=B)]` to denote that the field is a key with the specified range of values.

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This makes it sound like Min and Count are required values, but they do not appear to be. (And I'm assuming there are plenty of scenarios where a user doesn't know up front what are all the possible values. #Resolved

Comment threaddocs/code/SchemaComprehension.md Outdated
var predictionEngine = env.CreatePredictionEngine<IrisData, IrisVectorData>(dv, outputSchemaDefinition: schemaDef);
```

In addition to the above, you can use `SchemaDefinition` to add per-column metadata, or even a 'value generator' (so that the column value is not read from the field, but computed using a delegate).

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It might be interesting to have a code snippet example for this scenario? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In fact, I think the 'generator' bit didn't make it into ML.NET


In reply to: 204536819 [](ancestors = 204536819)

Comment threaddocs/code/SchemaComprehension.md Outdated
* Reading a different subset of columns on every row: the cursor always populates the entire row object.
* Reading column metadata from the data view.
* Accessing the 'hidden' data view columns by index.
* Creating 'cursor sets'.

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A link or definition of cursor sets may be helpful here. #Resolved

@Zruty0

Copy link
Copy Markdown
ContributorAuthor

I hope they will still go away, because they block the merging.


In reply to: 407177294 [](ancestors = 407177294)

@@ -0,0 +1,210 @@
# Schema comprehension in ML.NET

This document describes in detail the under-the-hood mechanism that ML.NET uses to automate the creation of `IDataView` schema, with the goal to make it as convenient to the end user as possible, while not incurring extra computational costs.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

IDataView` schem [](start = 109, length = 16)

Might be useful to link to the IDV doc. #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated

## Introduction

Every dataset in ML.NET is an `IDataView`, which is, for the purposes of this document, a collection of rows that share the same columns. The set of columns, their names, types and other metadata is known as the *schema* of the `IDataView`, and it's represented as an `ISchema` object.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is an [](start = 24, length = 5)

would it be more clear if it says "gets loaded into an IDV" #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm, I think it's more correct to say 'is represented as', because you don't necessarily LOAD a dataset.


In reply to: 204819663 [](ancestors = 204819663)

Comment threaddocs/code/SchemaComprehension.md Outdated

## Introduction

Every dataset in ML.NET is an `IDataView`, which is, for the purposes of this document, a collection of rows that share the same columns. The set of columns, their names, types and other metadata is known as the *schema* of the `IDataView`, and it's represented as an `ISchema` object.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

schema [](start = 212, length = 9)

link to the schema section of the IDV Design Principles #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added one above


In reply to: 204820706 [](ancestors = 204820706)

These items above are very similar to the definition of fields in a C# class: names and types of columns correspond to names and types of fields, and metadata can correspond to field attributes.
Because of this similarity, ML.NET offers a common convenient mechanism for creating a schema: it is done via defining a C# class.

For example, the below class definition can be used to define a data view with 5 float columns:

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

data view [](start = 64, length = 9)

wondering if it would help to state at the beginning that IDataView and 'data view' are interchangeable, because you give the definition of one, and use the other term for it. #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated
.ToArray();
}
```
After this code runs, `arr` will contain two `IrisVectorData` objects, each having `Features` filled with the actual values of the features.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

features [](start = 131, length = 8)

I'd add (the 4 concatenated columns) after features, to make it more explicit. #Closed

```(csharp)
var streamingDv = env.CreateStreamingDataView<IrisData>(dataEnumerable);
```
The only subtle difference is, the resulting `streamingDv` will not support shuffling (a property that's useful to some ML application).

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

shuffling [](start = 76, length = 9)

Maybe link to what data shuffling is.
#Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure what to link to...


In reply to: 204823733 [](ancestors = 204823733)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah, the Wikipedia one isn't great..


In reply to: 204855447 [](ancestors = 204855447,204823733)

Comment threaddocs/code/SchemaComprehension.md Outdated
`IDataView` [type system](IDataViewTypeSystem.md) differs slightly from the C# type system, so a 1-1 mapping between column types and C# types is not always feasible.
Below are the most notable examples of the differences:

* `IDataView` vector columns may have a fixed (and known) size, C# arrays can not. You can use `[VectorType(N)]` attribute to an array field to specify that the column is a vector of fixed size N. This is often necessary: most ML components don't work with variable-size vectors, they require fixed-size ones.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

C# arrays can not [](start = 64, length = 17)

this might get confusing if you think about initialized arrays. #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tried to clarify


In reply to: 204831222 [](ancestors = 204831222)

Comment threaddocs/code/SchemaComprehension.md Outdated
Below are the most notable examples of the differences:

* `IDataView` vector columns may have a fixed (and known) size, C# arrays can not. You can use `[VectorType(N)]` attribute to an array field to specify that the column is a vector of fixed size N. This is often necessary: most ML components don't work with variable-size vectors, they require fixed-size ones.
* `IDataView`'s **key types** don't have an underlying C# type either. To declare a key-type column, you need to make your field an `uint`, and decorate it with `[KeyType(Min=A, Count=B)]` to denote that the field is a key with the specified range of values.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

*key types [](start = 16, length = 12)

link maybe #Closed

| `BL` | `DvBool` | `bool`, `bool?` |
| `TS` | `DvTimeSpan` | |
| `DT` | `DvDateTime` | |
| `DZ` | `DvDateTimeZone` | |

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

They don't map to TimeSpan, DateTime and DataTimeZone? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No


In reply to: 204836126 [](ancestors = 204836126)

It was our design decision to not allow these scenarios, thus simplifying the other, more common scenarios.

Here is the list of things that are only possible via the low-level interface:
* Creating or reading a data view, where even column *types* are not known at compile time (so you cannot create a C# class to define the schema)

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Creating or reading a data view, where even column types are not known at compile time (so you cannot create a C# class to define the schema) [](start = 2, length = 143)

example of scenario when this might occur #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated
* Creating or reading a data view, where even column *types* are not known at compile time (so you cannot create a C# class to define the schema)
* Reading a different subset of columns on every row: the cursor always populates the entire row object.
* Reading column metadata from the data view.
* Accessing the 'hidden' data view columns by index.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

'hidden' data view column [](start = 16, length = 25)

link or define "hidden" #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated

Here is the list of things that are only possible via the low-level interface:
* Creating or reading a data view, where even column *types* are not known at compile time (so you cannot create a C# class to define the schema)
* Reading a different subset of columns on every row: the cursor always populates the entire row object.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

different subset of columns on every row [](start = 12, length = 40)

what does 'different' mean here? #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tried to rephrase


In reply to: 204839399 [](ancestors = 204839399)

@sfilipisfilipi left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@Zruty0
Zruty0 merged commit 8cfa2ed into dotnet:masterJul 24, 2018
@Zruty0
Zruty0 deleted the feature/554-schema-doc branch July 24, 2018 19:32
eerhardt pushed a commit to eerhardt/machinelearning that referenced this pull request Jul 27, 2018
* Added a doc for schema comprehension
codemzs pushed a commit to codemzs/machinelearning that referenced this pull request Aug 1, 2018
* Added a doc for schema comprehension
@ghostghost locked as resolved and limited conversation to collaborators Mar 29, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Zruty0@dnfclas@eerhardt@sfilipi
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Schema comprehension doc - #572

Merged
Zruty0 merged 2 commits into
dotnet:masterfrom
Zruty0:feature/554-schema-doc
Jul 24, 2018
Merged

Schema comprehension doc#572
Zruty0 merged 2 commits into
dotnet:masterfrom
Zruty0:feature/554-schema-doc

Conversation

@Zruty0

Copy link
Copy Markdown
Contributor

Added a document that describe typed schema comprehension.

Fixes#554

@dnfclas

dnfclas commented Jul 23, 2018

Copy link
Copy Markdown

CLA assistant check
All CLA requirements met. #Closed

@Zruty0

Zruty0 commented Jul 23, 2018

Copy link
Copy Markdown
ContributorAuthor

@dotnet-bot test Linux Release
#Closed

@eerhardt

eerhardt commented Jul 23, 2018

Copy link
Copy Markdown
Member

@Zruty0 - we are in the middle of moving our CI system from Jenkins to VSTS. You can ignore those 2 failed runs. (Plus you are just modifying .md files anyway.) #Resolved

@eerhardteerhardt left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for this helpful document, @Zruty0. It looks really good.

Comment threaddocs/code/SchemaComprehension.md Outdated

### Streaming data views

What if the original data doesn't support seeking, kile if it's some form of `IEnumerable<IrisData>` instead of `IList<IrisData>`? Well, we can simply use another helper function:

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(type-o) kile #Resolved

Comment threaddocs/code/SchemaComprehension.md Outdated

Let's see how we can create a new `IDataView` out of an in-memory array, run some operations on it, and then read it back into the array.

```(csharp)

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure what exactly about your string isn't working, but I don't get syntax highlighting when viewing the document.

Typically, I use the format ```C# instead. #Resolved

Comment threaddocs/code/SchemaComprehension.md Outdated
Below are the most notable examples of the differences:

* `IDataView` vector columns may have a fixed (and known) size, C# arrays can not. You can use `[VectorType(N)]` attribute to an array field to specify that the column is a vector of fixed size N. This is often necessary: most ML components don't work with variable-size vectors, they require fixed-size ones.
* `IDataView`'s **key types** don't have an underlying C# type either. To declare a key-type column, you need to make your field an `uint`, and decorate it with `[KeyType(Min=A, Count=B)]` to denote that the field is a key with the specified range of values.

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This makes it sound like Min and Count are required values, but they do not appear to be. (And I'm assuming there are plenty of scenarios where a user doesn't know up front what are all the possible values. #Resolved

Comment threaddocs/code/SchemaComprehension.md Outdated
var predictionEngine = env.CreatePredictionEngine<IrisData, IrisVectorData>(dv, outputSchemaDefinition: schemaDef);
```

In addition to the above, you can use `SchemaDefinition` to add per-column metadata, or even a 'value generator' (so that the column value is not read from the field, but computed using a delegate).

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It might be interesting to have a code snippet example for this scenario? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In fact, I think the 'generator' bit didn't make it into ML.NET


In reply to: 204536819 [](ancestors = 204536819)

Comment threaddocs/code/SchemaComprehension.md Outdated
* Reading a different subset of columns on every row: the cursor always populates the entire row object.
* Reading column metadata from the data view.
* Accessing the 'hidden' data view columns by index.
* Creating 'cursor sets'.

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A link or definition of cursor sets may be helpful here. #Resolved

@Zruty0

Copy link
Copy Markdown
ContributorAuthor

I hope they will still go away, because they block the merging.


In reply to: 407177294 [](ancestors = 407177294)

@@ -0,0 +1,210 @@
# Schema comprehension in ML.NET

This document describes in detail the under-the-hood mechanism that ML.NET uses to automate the creation of `IDataView` schema, with the goal to make it as convenient to the end user as possible, while not incurring extra computational costs.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

IDataView` schem [](start = 109, length = 16)

Might be useful to link to the IDV doc. #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated

## Introduction

Every dataset in ML.NET is an `IDataView`, which is, for the purposes of this document, a collection of rows that share the same columns. The set of columns, their names, types and other metadata is known as the *schema* of the `IDataView`, and it's represented as an `ISchema` object.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is an [](start = 24, length = 5)

would it be more clear if it says "gets loaded into an IDV" #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm, I think it's more correct to say 'is represented as', because you don't necessarily LOAD a dataset.


In reply to: 204819663 [](ancestors = 204819663)

Comment threaddocs/code/SchemaComprehension.md Outdated

## Introduction

Every dataset in ML.NET is an `IDataView`, which is, for the purposes of this document, a collection of rows that share the same columns. The set of columns, their names, types and other metadata is known as the *schema* of the `IDataView`, and it's represented as an `ISchema` object.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

schema [](start = 212, length = 9)

link to the schema section of the IDV Design Principles #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added one above


In reply to: 204820706 [](ancestors = 204820706)

These items above are very similar to the definition of fields in a C# class: names and types of columns correspond to names and types of fields, and metadata can correspond to field attributes.
Because of this similarity, ML.NET offers a common convenient mechanism for creating a schema: it is done via defining a C# class.

For example, the below class definition can be used to define a data view with 5 float columns:

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

data view [](start = 64, length = 9)

wondering if it would help to state at the beginning that IDataView and 'data view' are interchangeable, because you give the definition of one, and use the other term for it. #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated
.ToArray();
}
```
After this code runs, `arr` will contain two `IrisVectorData` objects, each having `Features` filled with the actual values of the features.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

features [](start = 131, length = 8)

I'd add (the 4 concatenated columns) after features, to make it more explicit. #Closed

```(csharp)
var streamingDv = env.CreateStreamingDataView<IrisData>(dataEnumerable);
```
The only subtle difference is, the resulting `streamingDv` will not support shuffling (a property that's useful to some ML application).

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

shuffling [](start = 76, length = 9)

Maybe link to what data shuffling is.
#Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure what to link to...


In reply to: 204823733 [](ancestors = 204823733)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah, the Wikipedia one isn't great..


In reply to: 204855447 [](ancestors = 204855447,204823733)

Comment threaddocs/code/SchemaComprehension.md Outdated
`IDataView` [type system](IDataViewTypeSystem.md) differs slightly from the C# type system, so a 1-1 mapping between column types and C# types is not always feasible.
Below are the most notable examples of the differences:

* `IDataView` vector columns may have a fixed (and known) size, C# arrays can not. You can use `[VectorType(N)]` attribute to an array field to specify that the column is a vector of fixed size N. This is often necessary: most ML components don't work with variable-size vectors, they require fixed-size ones.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

C# arrays can not [](start = 64, length = 17)

this might get confusing if you think about initialized arrays. #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tried to clarify


In reply to: 204831222 [](ancestors = 204831222)

Comment threaddocs/code/SchemaComprehension.md Outdated
Below are the most notable examples of the differences:

* `IDataView` vector columns may have a fixed (and known) size, C# arrays can not. You can use `[VectorType(N)]` attribute to an array field to specify that the column is a vector of fixed size N. This is often necessary: most ML components don't work with variable-size vectors, they require fixed-size ones.
* `IDataView`'s **key types** don't have an underlying C# type either. To declare a key-type column, you need to make your field an `uint`, and decorate it with `[KeyType(Min=A, Count=B)]` to denote that the field is a key with the specified range of values.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

*key types [](start = 16, length = 12)

link maybe #Closed

| `BL` | `DvBool` | `bool`, `bool?` |
| `TS` | `DvTimeSpan` | |
| `DT` | `DvDateTime` | |
| `DZ` | `DvDateTimeZone` | |

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

They don't map to TimeSpan, DateTime and DataTimeZone? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No


In reply to: 204836126 [](ancestors = 204836126)

It was our design decision to not allow these scenarios, thus simplifying the other, more common scenarios.

Here is the list of things that are only possible via the low-level interface:
* Creating or reading a data view, where even column *types* are not known at compile time (so you cannot create a C# class to define the schema)

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Creating or reading a data view, where even column types are not known at compile time (so you cannot create a C# class to define the schema) [](start = 2, length = 143)

example of scenario when this might occur #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated
* Creating or reading a data view, where even column *types* are not known at compile time (so you cannot create a C# class to define the schema)
* Reading a different subset of columns on every row: the cursor always populates the entire row object.
* Reading column metadata from the data view.
* Accessing the 'hidden' data view columns by index.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

'hidden' data view column [](start = 16, length = 25)

link or define "hidden" #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated

Here is the list of things that are only possible via the low-level interface:
* Creating or reading a data view, where even column *types* are not known at compile time (so you cannot create a C# class to define the schema)
* Reading a different subset of columns on every row: the cursor always populates the entire row object.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

different subset of columns on every row [](start = 12, length = 40)

what does 'different' mean here? #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tried to rephrase


In reply to: 204839399 [](ancestors = 204839399)

@sfilipisfilipi left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@Zruty0
Zruty0 merged commit 8cfa2ed into dotnet:masterJul 24, 2018
@Zruty0
Zruty0 deleted the feature/554-schema-doc branch July 24, 2018 19:32
eerhardt pushed a commit to eerhardt/machinelearning that referenced this pull request Jul 27, 2018
* Added a doc for schema comprehension
codemzs pushed a commit to codemzs/machinelearning that referenced this pull request Aug 1, 2018
* Added a doc for schema comprehension
@ghostghost locked as resolved and limited conversation to collaborators Mar 29, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Zruty0@dnfclas@eerhardt@sfilipi
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Schema comprehension doc - #572

Merged
Zruty0 merged 2 commits into
dotnet:masterfrom
Zruty0:feature/554-schema-doc
Jul 24, 2018
Merged

Schema comprehension doc#572
Zruty0 merged 2 commits into
dotnet:masterfrom
Zruty0:feature/554-schema-doc

Conversation

@Zruty0

Copy link
Copy Markdown
Contributor

Added a document that describe typed schema comprehension.

Fixes#554

@dnfclas

dnfclas commented Jul 23, 2018

Copy link
Copy Markdown

CLA assistant check
All CLA requirements met. #Closed

@Zruty0

Zruty0 commented Jul 23, 2018

Copy link
Copy Markdown
ContributorAuthor

@dotnet-bot test Linux Release
#Closed

@eerhardt

eerhardt commented Jul 23, 2018

Copy link
Copy Markdown
Member

@Zruty0 - we are in the middle of moving our CI system from Jenkins to VSTS. You can ignore those 2 failed runs. (Plus you are just modifying .md files anyway.) #Resolved

@eerhardteerhardt left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for this helpful document, @Zruty0. It looks really good.

Comment threaddocs/code/SchemaComprehension.md Outdated

### Streaming data views

What if the original data doesn't support seeking, kile if it's some form of `IEnumerable<IrisData>` instead of `IList<IrisData>`? Well, we can simply use another helper function:

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(type-o) kile #Resolved

Comment threaddocs/code/SchemaComprehension.md Outdated

Let's see how we can create a new `IDataView` out of an in-memory array, run some operations on it, and then read it back into the array.

```(csharp)

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure what exactly about your string isn't working, but I don't get syntax highlighting when viewing the document.

Typically, I use the format ```C# instead. #Resolved

Comment threaddocs/code/SchemaComprehension.md Outdated
Below are the most notable examples of the differences:

* `IDataView` vector columns may have a fixed (and known) size, C# arrays can not. You can use `[VectorType(N)]` attribute to an array field to specify that the column is a vector of fixed size N. This is often necessary: most ML components don't work with variable-size vectors, they require fixed-size ones.
* `IDataView`'s **key types** don't have an underlying C# type either. To declare a key-type column, you need to make your field an `uint`, and decorate it with `[KeyType(Min=A, Count=B)]` to denote that the field is a key with the specified range of values.

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This makes it sound like Min and Count are required values, but they do not appear to be. (And I'm assuming there are plenty of scenarios where a user doesn't know up front what are all the possible values. #Resolved

Comment threaddocs/code/SchemaComprehension.md Outdated
var predictionEngine = env.CreatePredictionEngine<IrisData, IrisVectorData>(dv, outputSchemaDefinition: schemaDef);
```

In addition to the above, you can use `SchemaDefinition` to add per-column metadata, or even a 'value generator' (so that the column value is not read from the field, but computed using a delegate).

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It might be interesting to have a code snippet example for this scenario? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In fact, I think the 'generator' bit didn't make it into ML.NET


In reply to: 204536819 [](ancestors = 204536819)

Comment threaddocs/code/SchemaComprehension.md Outdated
* Reading a different subset of columns on every row: the cursor always populates the entire row object.
* Reading column metadata from the data view.
* Accessing the 'hidden' data view columns by index.
* Creating 'cursor sets'.

@eerhardteerhardtJul 23, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A link or definition of cursor sets may be helpful here. #Resolved

@Zruty0

Copy link
Copy Markdown
ContributorAuthor

I hope they will still go away, because they block the merging.


In reply to: 407177294 [](ancestors = 407177294)

@@ -0,0 +1,210 @@
# Schema comprehension in ML.NET

This document describes in detail the under-the-hood mechanism that ML.NET uses to automate the creation of `IDataView` schema, with the goal to make it as convenient to the end user as possible, while not incurring extra computational costs.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

IDataView` schem [](start = 109, length = 16)

Might be useful to link to the IDV doc. #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated

## Introduction

Every dataset in ML.NET is an `IDataView`, which is, for the purposes of this document, a collection of rows that share the same columns. The set of columns, their names, types and other metadata is known as the *schema* of the `IDataView`, and it's represented as an `ISchema` object.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is an [](start = 24, length = 5)

would it be more clear if it says "gets loaded into an IDV" #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm, I think it's more correct to say 'is represented as', because you don't necessarily LOAD a dataset.


In reply to: 204819663 [](ancestors = 204819663)

Comment threaddocs/code/SchemaComprehension.md Outdated

## Introduction

Every dataset in ML.NET is an `IDataView`, which is, for the purposes of this document, a collection of rows that share the same columns. The set of columns, their names, types and other metadata is known as the *schema* of the `IDataView`, and it's represented as an `ISchema` object.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

schema [](start = 212, length = 9)

link to the schema section of the IDV Design Principles #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added one above


In reply to: 204820706 [](ancestors = 204820706)

These items above are very similar to the definition of fields in a C# class: names and types of columns correspond to names and types of fields, and metadata can correspond to field attributes.
Because of this similarity, ML.NET offers a common convenient mechanism for creating a schema: it is done via defining a C# class.

For example, the below class definition can be used to define a data view with 5 float columns:

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

data view [](start = 64, length = 9)

wondering if it would help to state at the beginning that IDataView and 'data view' are interchangeable, because you give the definition of one, and use the other term for it. #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated
.ToArray();
}
```
After this code runs, `arr` will contain two `IrisVectorData` objects, each having `Features` filled with the actual values of the features.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

features [](start = 131, length = 8)

I'd add (the 4 concatenated columns) after features, to make it more explicit. #Closed

```(csharp)
var streamingDv = env.CreateStreamingDataView<IrisData>(dataEnumerable);
```
The only subtle difference is, the resulting `streamingDv` will not support shuffling (a property that's useful to some ML application).

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

shuffling [](start = 76, length = 9)

Maybe link to what data shuffling is.
#Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure what to link to...


In reply to: 204823733 [](ancestors = 204823733)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah, the Wikipedia one isn't great..


In reply to: 204855447 [](ancestors = 204855447,204823733)

Comment threaddocs/code/SchemaComprehension.md Outdated
`IDataView` [type system](IDataViewTypeSystem.md) differs slightly from the C# type system, so a 1-1 mapping between column types and C# types is not always feasible.
Below are the most notable examples of the differences:

* `IDataView` vector columns may have a fixed (and known) size, C# arrays can not. You can use `[VectorType(N)]` attribute to an array field to specify that the column is a vector of fixed size N. This is often necessary: most ML components don't work with variable-size vectors, they require fixed-size ones.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

C# arrays can not [](start = 64, length = 17)

this might get confusing if you think about initialized arrays. #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tried to clarify


In reply to: 204831222 [](ancestors = 204831222)

Comment threaddocs/code/SchemaComprehension.md Outdated
Below are the most notable examples of the differences:

* `IDataView` vector columns may have a fixed (and known) size, C# arrays can not. You can use `[VectorType(N)]` attribute to an array field to specify that the column is a vector of fixed size N. This is often necessary: most ML components don't work with variable-size vectors, they require fixed-size ones.
* `IDataView`'s **key types** don't have an underlying C# type either. To declare a key-type column, you need to make your field an `uint`, and decorate it with `[KeyType(Min=A, Count=B)]` to denote that the field is a key with the specified range of values.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

*key types [](start = 16, length = 12)

link maybe #Closed

| `BL` | `DvBool` | `bool`, `bool?` |
| `TS` | `DvTimeSpan` | |
| `DT` | `DvDateTime` | |
| `DZ` | `DvDateTimeZone` | |

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

They don't map to TimeSpan, DateTime and DataTimeZone? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No


In reply to: 204836126 [](ancestors = 204836126)

It was our design decision to not allow these scenarios, thus simplifying the other, more common scenarios.

Here is the list of things that are only possible via the low-level interface:
* Creating or reading a data view, where even column *types* are not known at compile time (so you cannot create a C# class to define the schema)

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Creating or reading a data view, where even column types are not known at compile time (so you cannot create a C# class to define the schema) [](start = 2, length = 143)

example of scenario when this might occur #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated
* Creating or reading a data view, where even column *types* are not known at compile time (so you cannot create a C# class to define the schema)
* Reading a different subset of columns on every row: the cursor always populates the entire row object.
* Reading column metadata from the data view.
* Accessing the 'hidden' data view columns by index.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

'hidden' data view column [](start = 16, length = 25)

link or define "hidden" #Closed

Comment threaddocs/code/SchemaComprehension.md Outdated

Here is the list of things that are only possible via the low-level interface:
* Creating or reading a data view, where even column *types* are not known at compile time (so you cannot create a C# class to define the schema)
* Reading a different subset of columns on every row: the cursor always populates the entire row object.

@sfilipisfilipiJul 24, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

different subset of columns on every row [](start = 12, length = 40)

what does 'different' mean here? #Closed

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tried to rephrase


In reply to: 204839399 [](ancestors = 204839399)

@sfilipisfilipi left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@Zruty0
Zruty0 merged commit 8cfa2ed into dotnet:masterJul 24, 2018
@Zruty0
Zruty0 deleted the feature/554-schema-doc branch July 24, 2018 19:32
eerhardt pushed a commit to eerhardt/machinelearning that referenced this pull request Jul 27, 2018
* Added a doc for schema comprehension
codemzs pushed a commit to codemzs/machinelearning that referenced this pull request Aug 1, 2018
* Added a doc for schema comprehension
@ghostghost locked as resolved and limited conversation to collaborators Mar 29, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@Zruty0@dnfclas@eerhardt@sfilipi