Created samples for 'ProduceNgrams' and 'ProduceHashedNgrams' APIs. - #3177

Merged
zeahmed merged 8 commits into
dotnet:masterfrom
zeahmed:ngram_samples
Apr 5, 2019
Merged

Created samples for 'ProduceNgrams' and 'ProduceHashedNgrams' APIs.#3177
zeahmed merged 8 commits into
dotnet:masterfrom
zeahmed:ngram_samples

Conversation

@zeahmed

Copy link
Copy Markdown
Contributor

Related to #1209.

foreach (var item in featureRow.Items())
Console.Write($"{slots[item.Key]} ");
Console.WriteLine();
}

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is the concern for me right now. There is no way to get this meta data through the Transformer or through the prediction engine. The only way is through IDataView obtained from .Transform call.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you have to transform the data anyway, why not just print off the transformed IDV instead of off a prediction? You can limit it to one row with a TakeRows filter.


In reply to: 271476465 [](ancestors = 271476465)

new TextData(){ Text = "The value at each position corresponds to," },
new TextData(){ Text = "the number of times Ngram occured in the data (Tf), or" },
new TextData(){ Text = "the inverse of the number of documents that contain the Ngram (Idf), or." },
new TextData(){ Text = "or compute both and multipy together (Tf-Idf)." },

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

love it! #Closed

var textPipeline = mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "Text")
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))
.Append(mlContext.Transforms.Text.ProduceNgrams("NgramFeatures", "Tokens",
ngramLength: 3, useAllLengths: false, weighting: NgramExtractingEstimator.WeightingCriteria.Tf));

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ngramLength: 3, useAllLengths: false, weighting: NgramExtractingEstimator.WeightingCriteria.Tf [](start = 16, length = 94)

one parameter in one line might be more presentable. #Resolved

// This is acheived by calling 'TokenizeIntoWords' first followed by 'ProduceNgrams'.
// Please note that the length of the output feature vector depends on the Ngram settings.
var textPipeline = mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "Text")
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens")) [](start = 14, length = 66)

add one line of comment on why this is here. #Resolved

}

// Print the first 10 feature values.
Console.Write("Features: ");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Console.Write("Features: "); [](start = 11, length = 29)

i'd remove unnecessary printings

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actually, I'd like to keep it because its more clear to read each line with this prefix on console.


In reply to: 271816037 [](ancestors = 271816037)

transformedDataView.Schema["NgramFeatures"].GetSlotNames(ref slotNames);
var NgramFeaturesColumn = transformedDataView.GetColumn<VBuffer<float>>(transformedDataView.Schema["NgramFeatures"]);
var slots = slotNames.GetValues();
Console.Write("Ngrams: ");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Console.Write("Ngrams: "); [](start = 10, length = 28)

i'd remove this too.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actually, I'd like to keep it because its more clear to read each line with this prefix on console.


In reply to: 271816588 [](ancestors = 271816588)

// Print the length of the feature vector.
Console.WriteLine($"Number of Features: {prediction.NgramFeatures.Length}");

// Preview of the produced .

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

// Preview of the produced . [](start = 11, length = 29)

comment about slot names #Resolved

var prediction = predictionEngine.Predict(samples[0]);

// Print the length of the feature vector.
Console.WriteLine($"Number of Features: {prediction.NgramFeatures.Length}");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is this necessary? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes.


In reply to: 271817449 [](ancestors = 271817449)

var prediction = predictionEngine.Predict(samples[0]);

// Print the length of the feature vector.
Console.WriteLine($"Number of Features: {prediction.NgramFeatures.Length}");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

similar comment to the other file, is this needed? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes.


In reply to: 271819189 [](ancestors = 271819189)

@codecov

codecovBot commented Apr 3, 2019

Copy link
Copy Markdown

Codecov Report

Merging #3177 into master will increase coverage by 0.04%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #3177 +/- ##
==========================================
+ Coverage 72.54% 72.58% +0.04% 
==========================================
Files 807 807 Lines 144774 144956 +182 Branches 16208 16212 +4 ==========================================
+ Hits 105022 105215 +193 + Misses 35338 35325 -13 - Partials 4414 4416 +2
FlagCoverage Δ
#Debug72.58% <ø> (+0.04%)⬆️
#production68.14% <ø> (+0.01%)⬆️
#test88.88% <ø> (+0.04%)⬆️
Impacted FilesCoverage Δ
src/Microsoft.ML.Transforms/Text/TextCatalog.cs41.66% <ø> (ø)⬆️
src/Microsoft.ML.DataView/KeyDataViewType.cs74.57% <0%> (-3.76%)⬇️
src/Microsoft.ML.Transforms/Text/LdaTransform.cs89.26% <0%> (-0.63%)⬇️
...soft.ML.TestFramework/DataPipe/TestDataPipeBase.cs73.7% <0%> (-0.34%)⬇️
test/Microsoft.ML.Tests/ImagesTests.cs98.69% <0%> (-0.13%)⬇️
...Microsoft.ML.Tests/Transformers/NormalizerTests.cs100% <0%> (ø)⬆️
...ML.Data/Transforms/ConversionsExtensionsCatalog.cs44.87% <0%> (ø)⬆️
src/Microsoft.ML.Maml/MAML.cs26.21% <0%> (+1.45%)⬆️
...rosoft.ML.ImageAnalytics/VectorToImageTransform.cs76.77% <0%> (+4.53%)⬆️
... and 3 more

new TextData(){ Text = "Each position in the vector corresponds to a particular Ngram." },
new TextData(){ Text = "The value at each position corresponds to," },
new TextData(){ Text = "the number of times Ngram occured in the data (Tf), or" },
new TextData(){ Text = "the inverse of the number of documents that contain the Ngram (Idf), or." },

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

or. [](start = 109, length = 3)

omit this one. #Resolved

new TextData(){ Text = "This is an example to compute Ngrams using hashing." },
new TextData(){ Text = "Ngram is a sequence of 'N' consecutive words/tokens." },
new TextData(){ Text = "ML.NET's ProduceHashedNgrams API produces count of Ngrams and hashes it as an index into a vector of given bit length." },
new TextData(){ Text = "The hashing schem reduces the size of the output feature vector" },

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

schem [](start = 52, length = 5)

process? #Resolved

Console.Write($"{prediction.NgramFeatures[i]:F4} ");

// Expected output:
// Number of Features: 256

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

// Number of Features: 256 [](start = 12, length = 28)

if you use maximumNumberOfInverts you can also show up slot names as in just ngrams. #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I tried that but its bit complex to represent it on console as many ngram can fall into one slot.


In reply to: 271903797 [](ancestors = 271903797)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Well you can control with that parameter how many you actually want to remember, right? So if you put 1, you will have only one of them


In reply to: 271929298 [](ancestors = 271929298,271903797)

@Ivanidzo4kaIvanidzo4ka left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

// Please note that the length of the output feature vector depends on the n-gram settings.
var textPipeline = mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "Text")
// 'ProduceNgrams' takes key type as input. Converting the tokens into key type using 'MapValueToKey'.
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens")) [](start = 16, length = 64)

This seems like a holdover from the internal codebase. I wonder if we should consider doing a breaking change to move this keytype conversion into the ProduceNGrams operation. The question is what other use cases do we expect to see?

@rogancarrrogancarr left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved with comments.

@zeahmed
zeahmed merged commit 854f154 into dotnet:masterApr 5, 2019
zeahmed added a commit to zeahmed/machinelearning that referenced this pull request Apr 8, 2019
@zeahmedzeahmed mentioned this pull request Apr 8, 2019
@ghostghost locked as resolved and limited conversation to collaborators Mar 23, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@zeahmed@Ivanidzo4ka@sfilipi@rogancarr
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Created samples for 'ProduceNgrams' and 'ProduceHashedNgrams' APIs. - #3177

Merged
zeahmed merged 8 commits into
dotnet:masterfrom
zeahmed:ngram_samples
Apr 5, 2019
Merged

Created samples for 'ProduceNgrams' and 'ProduceHashedNgrams' APIs.#3177
zeahmed merged 8 commits into
dotnet:masterfrom
zeahmed:ngram_samples

Conversation

@zeahmed

Copy link
Copy Markdown
Contributor

Related to #1209.

foreach (var item in featureRow.Items())
Console.Write($"{slots[item.Key]} ");
Console.WriteLine();
}

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is the concern for me right now. There is no way to get this meta data through the Transformer or through the prediction engine. The only way is through IDataView obtained from .Transform call.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you have to transform the data anyway, why not just print off the transformed IDV instead of off a prediction? You can limit it to one row with a TakeRows filter.


In reply to: 271476465 [](ancestors = 271476465)

new TextData(){ Text = "The value at each position corresponds to," },
new TextData(){ Text = "the number of times Ngram occured in the data (Tf), or" },
new TextData(){ Text = "the inverse of the number of documents that contain the Ngram (Idf), or." },
new TextData(){ Text = "or compute both and multipy together (Tf-Idf)." },

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

love it! #Closed

var textPipeline = mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "Text")
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))
.Append(mlContext.Transforms.Text.ProduceNgrams("NgramFeatures", "Tokens",
ngramLength: 3, useAllLengths: false, weighting: NgramExtractingEstimator.WeightingCriteria.Tf));

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ngramLength: 3, useAllLengths: false, weighting: NgramExtractingEstimator.WeightingCriteria.Tf [](start = 16, length = 94)

one parameter in one line might be more presentable. #Resolved

// This is acheived by calling 'TokenizeIntoWords' first followed by 'ProduceNgrams'.
// Please note that the length of the output feature vector depends on the Ngram settings.
var textPipeline = mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "Text")
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens")) [](start = 14, length = 66)

add one line of comment on why this is here. #Resolved

}

// Print the first 10 feature values.
Console.Write("Features: ");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Console.Write("Features: "); [](start = 11, length = 29)

i'd remove unnecessary printings

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actually, I'd like to keep it because its more clear to read each line with this prefix on console.


In reply to: 271816037 [](ancestors = 271816037)

transformedDataView.Schema["NgramFeatures"].GetSlotNames(ref slotNames);
var NgramFeaturesColumn = transformedDataView.GetColumn<VBuffer<float>>(transformedDataView.Schema["NgramFeatures"]);
var slots = slotNames.GetValues();
Console.Write("Ngrams: ");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Console.Write("Ngrams: "); [](start = 10, length = 28)

i'd remove this too.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actually, I'd like to keep it because its more clear to read each line with this prefix on console.


In reply to: 271816588 [](ancestors = 271816588)

// Print the length of the feature vector.
Console.WriteLine($"Number of Features: {prediction.NgramFeatures.Length}");

// Preview of the produced .

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

// Preview of the produced . [](start = 11, length = 29)

comment about slot names #Resolved

var prediction = predictionEngine.Predict(samples[0]);

// Print the length of the feature vector.
Console.WriteLine($"Number of Features: {prediction.NgramFeatures.Length}");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is this necessary? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes.


In reply to: 271817449 [](ancestors = 271817449)

var prediction = predictionEngine.Predict(samples[0]);

// Print the length of the feature vector.
Console.WriteLine($"Number of Features: {prediction.NgramFeatures.Length}");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

similar comment to the other file, is this needed? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes.


In reply to: 271819189 [](ancestors = 271819189)

@codecov

codecovBot commented Apr 3, 2019

Copy link
Copy Markdown

Codecov Report

Merging #3177 into master will increase coverage by 0.04%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #3177 +/- ##
==========================================
+ Coverage 72.54% 72.58% +0.04% 
==========================================
Files 807 807 Lines 144774 144956 +182 Branches 16208 16212 +4 ==========================================
+ Hits 105022 105215 +193 + Misses 35338 35325 -13 - Partials 4414 4416 +2
FlagCoverage Δ
#Debug72.58% <ø> (+0.04%)⬆️
#production68.14% <ø> (+0.01%)⬆️
#test88.88% <ø> (+0.04%)⬆️
Impacted FilesCoverage Δ
src/Microsoft.ML.Transforms/Text/TextCatalog.cs41.66% <ø> (ø)⬆️
src/Microsoft.ML.DataView/KeyDataViewType.cs74.57% <0%> (-3.76%)⬇️
src/Microsoft.ML.Transforms/Text/LdaTransform.cs89.26% <0%> (-0.63%)⬇️
...soft.ML.TestFramework/DataPipe/TestDataPipeBase.cs73.7% <0%> (-0.34%)⬇️
test/Microsoft.ML.Tests/ImagesTests.cs98.69% <0%> (-0.13%)⬇️
...Microsoft.ML.Tests/Transformers/NormalizerTests.cs100% <0%> (ø)⬆️
...ML.Data/Transforms/ConversionsExtensionsCatalog.cs44.87% <0%> (ø)⬆️
src/Microsoft.ML.Maml/MAML.cs26.21% <0%> (+1.45%)⬆️
...rosoft.ML.ImageAnalytics/VectorToImageTransform.cs76.77% <0%> (+4.53%)⬆️
... and 3 more

new TextData(){ Text = "Each position in the vector corresponds to a particular Ngram." },
new TextData(){ Text = "The value at each position corresponds to," },
new TextData(){ Text = "the number of times Ngram occured in the data (Tf), or" },
new TextData(){ Text = "the inverse of the number of documents that contain the Ngram (Idf), or." },

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

or. [](start = 109, length = 3)

omit this one. #Resolved

new TextData(){ Text = "This is an example to compute Ngrams using hashing." },
new TextData(){ Text = "Ngram is a sequence of 'N' consecutive words/tokens." },
new TextData(){ Text = "ML.NET's ProduceHashedNgrams API produces count of Ngrams and hashes it as an index into a vector of given bit length." },
new TextData(){ Text = "The hashing schem reduces the size of the output feature vector" },

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

schem [](start = 52, length = 5)

process? #Resolved

Console.Write($"{prediction.NgramFeatures[i]:F4} ");

// Expected output:
// Number of Features: 256

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

// Number of Features: 256 [](start = 12, length = 28)

if you use maximumNumberOfInverts you can also show up slot names as in just ngrams. #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I tried that but its bit complex to represent it on console as many ngram can fall into one slot.


In reply to: 271903797 [](ancestors = 271903797)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Well you can control with that parameter how many you actually want to remember, right? So if you put 1, you will have only one of them


In reply to: 271929298 [](ancestors = 271929298,271903797)

@Ivanidzo4kaIvanidzo4ka left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

// Please note that the length of the output feature vector depends on the n-gram settings.
var textPipeline = mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "Text")
// 'ProduceNgrams' takes key type as input. Converting the tokens into key type using 'MapValueToKey'.
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens")) [](start = 16, length = 64)

This seems like a holdover from the internal codebase. I wonder if we should consider doing a breaking change to move this keytype conversion into the ProduceNGrams operation. The question is what other use cases do we expect to see?

@rogancarrrogancarr left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved with comments.

@zeahmed
zeahmed merged commit 854f154 into dotnet:masterApr 5, 2019
zeahmed added a commit to zeahmed/machinelearning that referenced this pull request Apr 8, 2019
@zeahmedzeahmed mentioned this pull request Apr 8, 2019
@ghostghost locked as resolved and limited conversation to collaborators Mar 23, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@zeahmed@Ivanidzo4ka@sfilipi@rogancarr
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Created samples for 'ProduceNgrams' and 'ProduceHashedNgrams' APIs. - #3177

Merged
zeahmed merged 8 commits into
dotnet:masterfrom
zeahmed:ngram_samples
Apr 5, 2019
Merged

Created samples for 'ProduceNgrams' and 'ProduceHashedNgrams' APIs.#3177
zeahmed merged 8 commits into
dotnet:masterfrom
zeahmed:ngram_samples

Conversation

@zeahmed

Copy link
Copy Markdown
Contributor

Related to #1209.

foreach (var item in featureRow.Items())
Console.Write($"{slots[item.Key]} ");
Console.WriteLine();
}

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is the concern for me right now. There is no way to get this meta data through the Transformer or through the prediction engine. The only way is through IDataView obtained from .Transform call.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you have to transform the data anyway, why not just print off the transformed IDV instead of off a prediction? You can limit it to one row with a TakeRows filter.


In reply to: 271476465 [](ancestors = 271476465)

new TextData(){ Text = "The value at each position corresponds to," },
new TextData(){ Text = "the number of times Ngram occured in the data (Tf), or" },
new TextData(){ Text = "the inverse of the number of documents that contain the Ngram (Idf), or." },
new TextData(){ Text = "or compute both and multipy together (Tf-Idf)." },

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

love it! #Closed

var textPipeline = mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "Text")
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))
.Append(mlContext.Transforms.Text.ProduceNgrams("NgramFeatures", "Tokens",
ngramLength: 3, useAllLengths: false, weighting: NgramExtractingEstimator.WeightingCriteria.Tf));

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ngramLength: 3, useAllLengths: false, weighting: NgramExtractingEstimator.WeightingCriteria.Tf [](start = 16, length = 94)

one parameter in one line might be more presentable. #Resolved

// This is acheived by calling 'TokenizeIntoWords' first followed by 'ProduceNgrams'.
// Please note that the length of the output feature vector depends on the Ngram settings.
var textPipeline = mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "Text")
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens")) [](start = 14, length = 66)

add one line of comment on why this is here. #Resolved

}

// Print the first 10 feature values.
Console.Write("Features: ");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Console.Write("Features: "); [](start = 11, length = 29)

i'd remove unnecessary printings

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actually, I'd like to keep it because its more clear to read each line with this prefix on console.


In reply to: 271816037 [](ancestors = 271816037)

transformedDataView.Schema["NgramFeatures"].GetSlotNames(ref slotNames);
var NgramFeaturesColumn = transformedDataView.GetColumn<VBuffer<float>>(transformedDataView.Schema["NgramFeatures"]);
var slots = slotNames.GetValues();
Console.Write("Ngrams: ");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Console.Write("Ngrams: "); [](start = 10, length = 28)

i'd remove this too.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actually, I'd like to keep it because its more clear to read each line with this prefix on console.


In reply to: 271816588 [](ancestors = 271816588)

// Print the length of the feature vector.
Console.WriteLine($"Number of Features: {prediction.NgramFeatures.Length}");

// Preview of the produced .

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

// Preview of the produced . [](start = 11, length = 29)

comment about slot names #Resolved

var prediction = predictionEngine.Predict(samples[0]);

// Print the length of the feature vector.
Console.WriteLine($"Number of Features: {prediction.NgramFeatures.Length}");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is this necessary? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes.


In reply to: 271817449 [](ancestors = 271817449)

var prediction = predictionEngine.Predict(samples[0]);

// Print the length of the feature vector.
Console.WriteLine($"Number of Features: {prediction.NgramFeatures.Length}");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

similar comment to the other file, is this needed? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes.


In reply to: 271819189 [](ancestors = 271819189)

@codecov

codecovBot commented Apr 3, 2019

Copy link
Copy Markdown

Codecov Report

Merging #3177 into master will increase coverage by 0.04%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #3177 +/- ##
==========================================
+ Coverage 72.54% 72.58% +0.04% 
==========================================
Files 807 807 Lines 144774 144956 +182 Branches 16208 16212 +4 ==========================================
+ Hits 105022 105215 +193 + Misses 35338 35325 -13 - Partials 4414 4416 +2
FlagCoverage Δ
#Debug72.58% <ø> (+0.04%)⬆️
#production68.14% <ø> (+0.01%)⬆️
#test88.88% <ø> (+0.04%)⬆️
Impacted FilesCoverage Δ
src/Microsoft.ML.Transforms/Text/TextCatalog.cs41.66% <ø> (ø)⬆️
src/Microsoft.ML.DataView/KeyDataViewType.cs74.57% <0%> (-3.76%)⬇️
src/Microsoft.ML.Transforms/Text/LdaTransform.cs89.26% <0%> (-0.63%)⬇️
...soft.ML.TestFramework/DataPipe/TestDataPipeBase.cs73.7% <0%> (-0.34%)⬇️
test/Microsoft.ML.Tests/ImagesTests.cs98.69% <0%> (-0.13%)⬇️
...Microsoft.ML.Tests/Transformers/NormalizerTests.cs100% <0%> (ø)⬆️
...ML.Data/Transforms/ConversionsExtensionsCatalog.cs44.87% <0%> (ø)⬆️
src/Microsoft.ML.Maml/MAML.cs26.21% <0%> (+1.45%)⬆️
...rosoft.ML.ImageAnalytics/VectorToImageTransform.cs76.77% <0%> (+4.53%)⬆️
... and 3 more

new TextData(){ Text = "Each position in the vector corresponds to a particular Ngram." },
new TextData(){ Text = "The value at each position corresponds to," },
new TextData(){ Text = "the number of times Ngram occured in the data (Tf), or" },
new TextData(){ Text = "the inverse of the number of documents that contain the Ngram (Idf), or." },

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

or. [](start = 109, length = 3)

omit this one. #Resolved

new TextData(){ Text = "This is an example to compute Ngrams using hashing." },
new TextData(){ Text = "Ngram is a sequence of 'N' consecutive words/tokens." },
new TextData(){ Text = "ML.NET's ProduceHashedNgrams API produces count of Ngrams and hashes it as an index into a vector of given bit length." },
new TextData(){ Text = "The hashing schem reduces the size of the output feature vector" },

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

schem [](start = 52, length = 5)

process? #Resolved

Console.Write($"{prediction.NgramFeatures[i]:F4} ");

// Expected output:
// Number of Features: 256

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

// Number of Features: 256 [](start = 12, length = 28)

if you use maximumNumberOfInverts you can also show up slot names as in just ngrams. #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I tried that but its bit complex to represent it on console as many ngram can fall into one slot.


In reply to: 271903797 [](ancestors = 271903797)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Well you can control with that parameter how many you actually want to remember, right? So if you put 1, you will have only one of them


In reply to: 271929298 [](ancestors = 271929298,271903797)

@Ivanidzo4kaIvanidzo4ka left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

// Please note that the length of the output feature vector depends on the n-gram settings.
var textPipeline = mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "Text")
// 'ProduceNgrams' takes key type as input. Converting the tokens into key type using 'MapValueToKey'.
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens")) [](start = 16, length = 64)

This seems like a holdover from the internal codebase. I wonder if we should consider doing a breaking change to move this keytype conversion into the ProduceNGrams operation. The question is what other use cases do we expect to see?

@rogancarrrogancarr left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved with comments.

@zeahmed
zeahmed merged commit 854f154 into dotnet:masterApr 5, 2019
zeahmed added a commit to zeahmed/machinelearning that referenced this pull request Apr 8, 2019
@zeahmedzeahmed mentioned this pull request Apr 8, 2019
@ghostghost locked as resolved and limited conversation to collaborators Mar 23, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@zeahmed@Ivanidzo4ka@sfilipi@rogancarr
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Created samples for 'ProduceNgrams' and 'ProduceHashedNgrams' APIs. - #3177

Merged
zeahmed merged 8 commits into
dotnet:masterfrom
zeahmed:ngram_samples
Apr 5, 2019
Merged

Created samples for 'ProduceNgrams' and 'ProduceHashedNgrams' APIs.#3177
zeahmed merged 8 commits into
dotnet:masterfrom
zeahmed:ngram_samples

Conversation

@zeahmed

Copy link
Copy Markdown
Contributor

Related to #1209.

foreach (var item in featureRow.Items())
Console.Write($"{slots[item.Key]} ");
Console.WriteLine();
}

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is the concern for me right now. There is no way to get this meta data through the Transformer or through the prediction engine. The only way is through IDataView obtained from .Transform call.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you have to transform the data anyway, why not just print off the transformed IDV instead of off a prediction? You can limit it to one row with a TakeRows filter.


In reply to: 271476465 [](ancestors = 271476465)

new TextData(){ Text = "The value at each position corresponds to," },
new TextData(){ Text = "the number of times Ngram occured in the data (Tf), or" },
new TextData(){ Text = "the inverse of the number of documents that contain the Ngram (Idf), or." },
new TextData(){ Text = "or compute both and multipy together (Tf-Idf)." },

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

love it! #Closed

var textPipeline = mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "Text")
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))
.Append(mlContext.Transforms.Text.ProduceNgrams("NgramFeatures", "Tokens",
ngramLength: 3, useAllLengths: false, weighting: NgramExtractingEstimator.WeightingCriteria.Tf));

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ngramLength: 3, useAllLengths: false, weighting: NgramExtractingEstimator.WeightingCriteria.Tf [](start = 16, length = 94)

one parameter in one line might be more presentable. #Resolved

// This is acheived by calling 'TokenizeIntoWords' first followed by 'ProduceNgrams'.
// Please note that the length of the output feature vector depends on the Ngram settings.
var textPipeline = mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "Text")
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens")) [](start = 14, length = 66)

add one line of comment on why this is here. #Resolved

}

// Print the first 10 feature values.
Console.Write("Features: ");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Console.Write("Features: "); [](start = 11, length = 29)

i'd remove unnecessary printings

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actually, I'd like to keep it because its more clear to read each line with this prefix on console.


In reply to: 271816037 [](ancestors = 271816037)

transformedDataView.Schema["NgramFeatures"].GetSlotNames(ref slotNames);
var NgramFeaturesColumn = transformedDataView.GetColumn<VBuffer<float>>(transformedDataView.Schema["NgramFeatures"]);
var slots = slotNames.GetValues();
Console.Write("Ngrams: ");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Console.Write("Ngrams: "); [](start = 10, length = 28)

i'd remove this too.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actually, I'd like to keep it because its more clear to read each line with this prefix on console.


In reply to: 271816588 [](ancestors = 271816588)

// Print the length of the feature vector.
Console.WriteLine($"Number of Features: {prediction.NgramFeatures.Length}");

// Preview of the produced .

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

// Preview of the produced . [](start = 11, length = 29)

comment about slot names #Resolved

var prediction = predictionEngine.Predict(samples[0]);

// Print the length of the feature vector.
Console.WriteLine($"Number of Features: {prediction.NgramFeatures.Length}");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is this necessary? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes.


In reply to: 271817449 [](ancestors = 271817449)

var prediction = predictionEngine.Predict(samples[0]);

// Print the length of the feature vector.
Console.WriteLine($"Number of Features: {prediction.NgramFeatures.Length}");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

similar comment to the other file, is this needed? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes.


In reply to: 271819189 [](ancestors = 271819189)

@codecov

codecovBot commented Apr 3, 2019

Copy link
Copy Markdown

Codecov Report

Merging #3177 into master will increase coverage by 0.04%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #3177 +/- ##
==========================================
+ Coverage 72.54% 72.58% +0.04% 
==========================================
Files 807 807 Lines 144774 144956 +182 Branches 16208 16212 +4 ==========================================
+ Hits 105022 105215 +193 + Misses 35338 35325 -13 - Partials 4414 4416 +2
FlagCoverage Δ
#Debug72.58% <ø> (+0.04%)⬆️
#production68.14% <ø> (+0.01%)⬆️
#test88.88% <ø> (+0.04%)⬆️
Impacted FilesCoverage Δ
src/Microsoft.ML.Transforms/Text/TextCatalog.cs41.66% <ø> (ø)⬆️
src/Microsoft.ML.DataView/KeyDataViewType.cs74.57% <0%> (-3.76%)⬇️
src/Microsoft.ML.Transforms/Text/LdaTransform.cs89.26% <0%> (-0.63%)⬇️
...soft.ML.TestFramework/DataPipe/TestDataPipeBase.cs73.7% <0%> (-0.34%)⬇️
test/Microsoft.ML.Tests/ImagesTests.cs98.69% <0%> (-0.13%)⬇️
...Microsoft.ML.Tests/Transformers/NormalizerTests.cs100% <0%> (ø)⬆️
...ML.Data/Transforms/ConversionsExtensionsCatalog.cs44.87% <0%> (ø)⬆️
src/Microsoft.ML.Maml/MAML.cs26.21% <0%> (+1.45%)⬆️
...rosoft.ML.ImageAnalytics/VectorToImageTransform.cs76.77% <0%> (+4.53%)⬆️
... and 3 more

new TextData(){ Text = "Each position in the vector corresponds to a particular Ngram." },
new TextData(){ Text = "The value at each position corresponds to," },
new TextData(){ Text = "the number of times Ngram occured in the data (Tf), or" },
new TextData(){ Text = "the inverse of the number of documents that contain the Ngram (Idf), or." },

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

or. [](start = 109, length = 3)

omit this one. #Resolved

new TextData(){ Text = "This is an example to compute Ngrams using hashing." },
new TextData(){ Text = "Ngram is a sequence of 'N' consecutive words/tokens." },
new TextData(){ Text = "ML.NET's ProduceHashedNgrams API produces count of Ngrams and hashes it as an index into a vector of given bit length." },
new TextData(){ Text = "The hashing schem reduces the size of the output feature vector" },

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

schem [](start = 52, length = 5)

process? #Resolved

Console.Write($"{prediction.NgramFeatures[i]:F4} ");

// Expected output:
// Number of Features: 256

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

// Number of Features: 256 [](start = 12, length = 28)

if you use maximumNumberOfInverts you can also show up slot names as in just ngrams. #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I tried that but its bit complex to represent it on console as many ngram can fall into one slot.


In reply to: 271903797 [](ancestors = 271903797)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Well you can control with that parameter how many you actually want to remember, right? So if you put 1, you will have only one of them


In reply to: 271929298 [](ancestors = 271929298,271903797)

@Ivanidzo4kaIvanidzo4ka left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

// Please note that the length of the output feature vector depends on the n-gram settings.
var textPipeline = mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "Text")
// 'ProduceNgrams' takes key type as input. Converting the tokens into key type using 'MapValueToKey'.
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens")) [](start = 16, length = 64)

This seems like a holdover from the internal codebase. I wonder if we should consider doing a breaking change to move this keytype conversion into the ProduceNGrams operation. The question is what other use cases do we expect to see?

@rogancarrrogancarr left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved with comments.

@zeahmed
zeahmed merged commit 854f154 into dotnet:masterApr 5, 2019
zeahmed added a commit to zeahmed/machinelearning that referenced this pull request Apr 8, 2019
@zeahmedzeahmed mentioned this pull request Apr 8, 2019
@ghostghost locked as resolved and limited conversation to collaborators Mar 23, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@zeahmed@Ivanidzo4ka@sfilipi@rogancarr
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Created samples for 'ProduceNgrams' and 'ProduceHashedNgrams' APIs. - #3177

Merged
zeahmed merged 8 commits into
dotnet:masterfrom
zeahmed:ngram_samples
Apr 5, 2019
Merged

Created samples for 'ProduceNgrams' and 'ProduceHashedNgrams' APIs.#3177
zeahmed merged 8 commits into
dotnet:masterfrom
zeahmed:ngram_samples

Conversation

@zeahmed

Copy link
Copy Markdown
Contributor

Related to #1209.

foreach (var item in featureRow.Items())
Console.Write($"{slots[item.Key]} ");
Console.WriteLine();
}

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is the concern for me right now. There is no way to get this meta data through the Transformer or through the prediction engine. The only way is through IDataView obtained from .Transform call.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you have to transform the data anyway, why not just print off the transformed IDV instead of off a prediction? You can limit it to one row with a TakeRows filter.


In reply to: 271476465 [](ancestors = 271476465)

new TextData(){ Text = "The value at each position corresponds to," },
new TextData(){ Text = "the number of times Ngram occured in the data (Tf), or" },
new TextData(){ Text = "the inverse of the number of documents that contain the Ngram (Idf), or." },
new TextData(){ Text = "or compute both and multipy together (Tf-Idf)." },

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

love it! #Closed

var textPipeline = mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "Text")
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))
.Append(mlContext.Transforms.Text.ProduceNgrams("NgramFeatures", "Tokens",
ngramLength: 3, useAllLengths: false, weighting: NgramExtractingEstimator.WeightingCriteria.Tf));

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ngramLength: 3, useAllLengths: false, weighting: NgramExtractingEstimator.WeightingCriteria.Tf [](start = 16, length = 94)

one parameter in one line might be more presentable. #Resolved

// This is acheived by calling 'TokenizeIntoWords' first followed by 'ProduceNgrams'.
// Please note that the length of the output feature vector depends on the Ngram settings.
var textPipeline = mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "Text")
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens")) [](start = 14, length = 66)

add one line of comment on why this is here. #Resolved

}

// Print the first 10 feature values.
Console.Write("Features: ");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Console.Write("Features: "); [](start = 11, length = 29)

i'd remove unnecessary printings

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actually, I'd like to keep it because its more clear to read each line with this prefix on console.


In reply to: 271816037 [](ancestors = 271816037)

transformedDataView.Schema["NgramFeatures"].GetSlotNames(ref slotNames);
var NgramFeaturesColumn = transformedDataView.GetColumn<VBuffer<float>>(transformedDataView.Schema["NgramFeatures"]);
var slots = slotNames.GetValues();
Console.Write("Ngrams: ");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Console.Write("Ngrams: "); [](start = 10, length = 28)

i'd remove this too.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actually, I'd like to keep it because its more clear to read each line with this prefix on console.


In reply to: 271816588 [](ancestors = 271816588)

// Print the length of the feature vector.
Console.WriteLine($"Number of Features: {prediction.NgramFeatures.Length}");

// Preview of the produced .

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

// Preview of the produced . [](start = 11, length = 29)

comment about slot names #Resolved

var prediction = predictionEngine.Predict(samples[0]);

// Print the length of the feature vector.
Console.WriteLine($"Number of Features: {prediction.NgramFeatures.Length}");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is this necessary? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes.


In reply to: 271817449 [](ancestors = 271817449)

var prediction = predictionEngine.Predict(samples[0]);

// Print the length of the feature vector.
Console.WriteLine($"Number of Features: {prediction.NgramFeatures.Length}");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

similar comment to the other file, is this needed? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes.


In reply to: 271819189 [](ancestors = 271819189)

@codecov

codecovBot commented Apr 3, 2019

Copy link
Copy Markdown

Codecov Report

Merging #3177 into master will increase coverage by 0.04%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #3177 +/- ##
==========================================
+ Coverage 72.54% 72.58% +0.04% 
==========================================
Files 807 807 Lines 144774 144956 +182 Branches 16208 16212 +4 ==========================================
+ Hits 105022 105215 +193 + Misses 35338 35325 -13 - Partials 4414 4416 +2
FlagCoverage Δ
#Debug72.58% <ø> (+0.04%)⬆️
#production68.14% <ø> (+0.01%)⬆️
#test88.88% <ø> (+0.04%)⬆️
Impacted FilesCoverage Δ
src/Microsoft.ML.Transforms/Text/TextCatalog.cs41.66% <ø> (ø)⬆️
src/Microsoft.ML.DataView/KeyDataViewType.cs74.57% <0%> (-3.76%)⬇️
src/Microsoft.ML.Transforms/Text/LdaTransform.cs89.26% <0%> (-0.63%)⬇️
...soft.ML.TestFramework/DataPipe/TestDataPipeBase.cs73.7% <0%> (-0.34%)⬇️
test/Microsoft.ML.Tests/ImagesTests.cs98.69% <0%> (-0.13%)⬇️
...Microsoft.ML.Tests/Transformers/NormalizerTests.cs100% <0%> (ø)⬆️
...ML.Data/Transforms/ConversionsExtensionsCatalog.cs44.87% <0%> (ø)⬆️
src/Microsoft.ML.Maml/MAML.cs26.21% <0%> (+1.45%)⬆️
...rosoft.ML.ImageAnalytics/VectorToImageTransform.cs76.77% <0%> (+4.53%)⬆️
... and 3 more

new TextData(){ Text = "Each position in the vector corresponds to a particular Ngram." },
new TextData(){ Text = "The value at each position corresponds to," },
new TextData(){ Text = "the number of times Ngram occured in the data (Tf), or" },
new TextData(){ Text = "the inverse of the number of documents that contain the Ngram (Idf), or." },

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

or. [](start = 109, length = 3)

omit this one. #Resolved

new TextData(){ Text = "This is an example to compute Ngrams using hashing." },
new TextData(){ Text = "Ngram is a sequence of 'N' consecutive words/tokens." },
new TextData(){ Text = "ML.NET's ProduceHashedNgrams API produces count of Ngrams and hashes it as an index into a vector of given bit length." },
new TextData(){ Text = "The hashing schem reduces the size of the output feature vector" },

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

schem [](start = 52, length = 5)

process? #Resolved

Console.Write($"{prediction.NgramFeatures[i]:F4} ");

// Expected output:
// Number of Features: 256

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

// Number of Features: 256 [](start = 12, length = 28)

if you use maximumNumberOfInverts you can also show up slot names as in just ngrams. #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I tried that but its bit complex to represent it on console as many ngram can fall into one slot.


In reply to: 271903797 [](ancestors = 271903797)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Well you can control with that parameter how many you actually want to remember, right? So if you put 1, you will have only one of them


In reply to: 271929298 [](ancestors = 271929298,271903797)

@Ivanidzo4kaIvanidzo4ka left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

// Please note that the length of the output feature vector depends on the n-gram settings.
var textPipeline = mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "Text")
// 'ProduceNgrams' takes key type as input. Converting the tokens into key type using 'MapValueToKey'.
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens")) [](start = 16, length = 64)

This seems like a holdover from the internal codebase. I wonder if we should consider doing a breaking change to move this keytype conversion into the ProduceNGrams operation. The question is what other use cases do we expect to see?

@rogancarrrogancarr left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved with comments.

@zeahmed
zeahmed merged commit 854f154 into dotnet:masterApr 5, 2019
zeahmed added a commit to zeahmed/machinelearning that referenced this pull request Apr 8, 2019
@zeahmedzeahmed mentioned this pull request Apr 8, 2019
@ghostghost locked as resolved and limited conversation to collaborators Mar 23, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@zeahmed@Ivanidzo4ka@sfilipi@rogancarr
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Created samples for 'ProduceNgrams' and 'ProduceHashedNgrams' APIs. - #3177

Merged
zeahmed merged 8 commits into
dotnet:masterfrom
zeahmed:ngram_samples
Apr 5, 2019
Merged

Created samples for 'ProduceNgrams' and 'ProduceHashedNgrams' APIs.#3177
zeahmed merged 8 commits into
dotnet:masterfrom
zeahmed:ngram_samples

Conversation

@zeahmed

Copy link
Copy Markdown
Contributor

Related to #1209.

foreach (var item in featureRow.Items())
Console.Write($"{slots[item.Key]} ");
Console.WriteLine();
}

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is the concern for me right now. There is no way to get this meta data through the Transformer or through the prediction engine. The only way is through IDataView obtained from .Transform call.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you have to transform the data anyway, why not just print off the transformed IDV instead of off a prediction? You can limit it to one row with a TakeRows filter.


In reply to: 271476465 [](ancestors = 271476465)

new TextData(){ Text = "The value at each position corresponds to," },
new TextData(){ Text = "the number of times Ngram occured in the data (Tf), or" },
new TextData(){ Text = "the inverse of the number of documents that contain the Ngram (Idf), or." },
new TextData(){ Text = "or compute both and multipy together (Tf-Idf)." },

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

love it! #Closed

var textPipeline = mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "Text")
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))
.Append(mlContext.Transforms.Text.ProduceNgrams("NgramFeatures", "Tokens",
ngramLength: 3, useAllLengths: false, weighting: NgramExtractingEstimator.WeightingCriteria.Tf));

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ngramLength: 3, useAllLengths: false, weighting: NgramExtractingEstimator.WeightingCriteria.Tf [](start = 16, length = 94)

one parameter in one line might be more presentable. #Resolved

// This is acheived by calling 'TokenizeIntoWords' first followed by 'ProduceNgrams'.
// Please note that the length of the output feature vector depends on the Ngram settings.
var textPipeline = mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "Text")
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens")) [](start = 14, length = 66)

add one line of comment on why this is here. #Resolved

}

// Print the first 10 feature values.
Console.Write("Features: ");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Console.Write("Features: "); [](start = 11, length = 29)

i'd remove unnecessary printings

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actually, I'd like to keep it because its more clear to read each line with this prefix on console.


In reply to: 271816037 [](ancestors = 271816037)

transformedDataView.Schema["NgramFeatures"].GetSlotNames(ref slotNames);
var NgramFeaturesColumn = transformedDataView.GetColumn<VBuffer<float>>(transformedDataView.Schema["NgramFeatures"]);
var slots = slotNames.GetValues();
Console.Write("Ngrams: ");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Console.Write("Ngrams: "); [](start = 10, length = 28)

i'd remove this too.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actually, I'd like to keep it because its more clear to read each line with this prefix on console.


In reply to: 271816588 [](ancestors = 271816588)

// Print the length of the feature vector.
Console.WriteLine($"Number of Features: {prediction.NgramFeatures.Length}");

// Preview of the produced .

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

// Preview of the produced . [](start = 11, length = 29)

comment about slot names #Resolved

var prediction = predictionEngine.Predict(samples[0]);

// Print the length of the feature vector.
Console.WriteLine($"Number of Features: {prediction.NgramFeatures.Length}");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is this necessary? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes.


In reply to: 271817449 [](ancestors = 271817449)

var prediction = predictionEngine.Predict(samples[0]);

// Print the length of the feature vector.
Console.WriteLine($"Number of Features: {prediction.NgramFeatures.Length}");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

similar comment to the other file, is this needed? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes.


In reply to: 271819189 [](ancestors = 271819189)

@codecov

codecovBot commented Apr 3, 2019

Copy link
Copy Markdown

Codecov Report

Merging #3177 into master will increase coverage by 0.04%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #3177 +/- ##
==========================================
+ Coverage 72.54% 72.58% +0.04% 
==========================================
Files 807 807 Lines 144774 144956 +182 Branches 16208 16212 +4 ==========================================
+ Hits 105022 105215 +193 + Misses 35338 35325 -13 - Partials 4414 4416 +2
FlagCoverage Δ
#Debug72.58% <ø> (+0.04%)⬆️
#production68.14% <ø> (+0.01%)⬆️
#test88.88% <ø> (+0.04%)⬆️
Impacted FilesCoverage Δ
src/Microsoft.ML.Transforms/Text/TextCatalog.cs41.66% <ø> (ø)⬆️
src/Microsoft.ML.DataView/KeyDataViewType.cs74.57% <0%> (-3.76%)⬇️
src/Microsoft.ML.Transforms/Text/LdaTransform.cs89.26% <0%> (-0.63%)⬇️
...soft.ML.TestFramework/DataPipe/TestDataPipeBase.cs73.7% <0%> (-0.34%)⬇️
test/Microsoft.ML.Tests/ImagesTests.cs98.69% <0%> (-0.13%)⬇️
...Microsoft.ML.Tests/Transformers/NormalizerTests.cs100% <0%> (ø)⬆️
...ML.Data/Transforms/ConversionsExtensionsCatalog.cs44.87% <0%> (ø)⬆️
src/Microsoft.ML.Maml/MAML.cs26.21% <0%> (+1.45%)⬆️
...rosoft.ML.ImageAnalytics/VectorToImageTransform.cs76.77% <0%> (+4.53%)⬆️
... and 3 more

new TextData(){ Text = "Each position in the vector corresponds to a particular Ngram." },
new TextData(){ Text = "The value at each position corresponds to," },
new TextData(){ Text = "the number of times Ngram occured in the data (Tf), or" },
new TextData(){ Text = "the inverse of the number of documents that contain the Ngram (Idf), or." },

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

or. [](start = 109, length = 3)

omit this one. #Resolved

new TextData(){ Text = "This is an example to compute Ngrams using hashing." },
new TextData(){ Text = "Ngram is a sequence of 'N' consecutive words/tokens." },
new TextData(){ Text = "ML.NET's ProduceHashedNgrams API produces count of Ngrams and hashes it as an index into a vector of given bit length." },
new TextData(){ Text = "The hashing schem reduces the size of the output feature vector" },

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

schem [](start = 52, length = 5)

process? #Resolved

Console.Write($"{prediction.NgramFeatures[i]:F4} ");

// Expected output:
// Number of Features: 256

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

// Number of Features: 256 [](start = 12, length = 28)

if you use maximumNumberOfInverts you can also show up slot names as in just ngrams. #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I tried that but its bit complex to represent it on console as many ngram can fall into one slot.


In reply to: 271903797 [](ancestors = 271903797)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Well you can control with that parameter how many you actually want to remember, right? So if you put 1, you will have only one of them


In reply to: 271929298 [](ancestors = 271929298,271903797)

@Ivanidzo4kaIvanidzo4ka left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

// Please note that the length of the output feature vector depends on the n-gram settings.
var textPipeline = mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "Text")
// 'ProduceNgrams' takes key type as input. Converting the tokens into key type using 'MapValueToKey'.
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens")) [](start = 16, length = 64)

This seems like a holdover from the internal codebase. I wonder if we should consider doing a breaking change to move this keytype conversion into the ProduceNGrams operation. The question is what other use cases do we expect to see?

@rogancarrrogancarr left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved with comments.

@zeahmed
zeahmed merged commit 854f154 into dotnet:masterApr 5, 2019
zeahmed added a commit to zeahmed/machinelearning that referenced this pull request Apr 8, 2019
@zeahmedzeahmed mentioned this pull request Apr 8, 2019
@ghostghost locked as resolved and limited conversation to collaborators Mar 23, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@zeahmed@Ivanidzo4ka@sfilipi@rogancarr
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Created samples for 'ProduceNgrams' and 'ProduceHashedNgrams' APIs. - #3177

Merged
zeahmed merged 8 commits into
dotnet:masterfrom
zeahmed:ngram_samples
Apr 5, 2019
Merged

Created samples for 'ProduceNgrams' and 'ProduceHashedNgrams' APIs.#3177
zeahmed merged 8 commits into
dotnet:masterfrom
zeahmed:ngram_samples

Conversation

@zeahmed

Copy link
Copy Markdown
Contributor

Related to #1209.

foreach (var item in featureRow.Items())
Console.Write($"{slots[item.Key]} ");
Console.WriteLine();
}

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is the concern for me right now. There is no way to get this meta data through the Transformer or through the prediction engine. The only way is through IDataView obtained from .Transform call.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you have to transform the data anyway, why not just print off the transformed IDV instead of off a prediction? You can limit it to one row with a TakeRows filter.


In reply to: 271476465 [](ancestors = 271476465)

new TextData(){ Text = "The value at each position corresponds to," },
new TextData(){ Text = "the number of times Ngram occured in the data (Tf), or" },
new TextData(){ Text = "the inverse of the number of documents that contain the Ngram (Idf), or." },
new TextData(){ Text = "or compute both and multipy together (Tf-Idf)." },

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

love it! #Closed

var textPipeline = mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "Text")
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))
.Append(mlContext.Transforms.Text.ProduceNgrams("NgramFeatures", "Tokens",
ngramLength: 3, useAllLengths: false, weighting: NgramExtractingEstimator.WeightingCriteria.Tf));

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ngramLength: 3, useAllLengths: false, weighting: NgramExtractingEstimator.WeightingCriteria.Tf [](start = 16, length = 94)

one parameter in one line might be more presentable. #Resolved

// This is acheived by calling 'TokenizeIntoWords' first followed by 'ProduceNgrams'.
// Please note that the length of the output feature vector depends on the Ngram settings.
var textPipeline = mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "Text")
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens")) [](start = 14, length = 66)

add one line of comment on why this is here. #Resolved

}

// Print the first 10 feature values.
Console.Write("Features: ");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Console.Write("Features: "); [](start = 11, length = 29)

i'd remove unnecessary printings

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actually, I'd like to keep it because its more clear to read each line with this prefix on console.


In reply to: 271816037 [](ancestors = 271816037)

transformedDataView.Schema["NgramFeatures"].GetSlotNames(ref slotNames);
var NgramFeaturesColumn = transformedDataView.GetColumn<VBuffer<float>>(transformedDataView.Schema["NgramFeatures"]);
var slots = slotNames.GetValues();
Console.Write("Ngrams: ");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Console.Write("Ngrams: "); [](start = 10, length = 28)

i'd remove this too.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actually, I'd like to keep it because its more clear to read each line with this prefix on console.


In reply to: 271816588 [](ancestors = 271816588)

// Print the length of the feature vector.
Console.WriteLine($"Number of Features: {prediction.NgramFeatures.Length}");

// Preview of the produced .

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

// Preview of the produced . [](start = 11, length = 29)

comment about slot names #Resolved

var prediction = predictionEngine.Predict(samples[0]);

// Print the length of the feature vector.
Console.WriteLine($"Number of Features: {prediction.NgramFeatures.Length}");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is this necessary? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes.


In reply to: 271817449 [](ancestors = 271817449)

var prediction = predictionEngine.Predict(samples[0]);

// Print the length of the feature vector.
Console.WriteLine($"Number of Features: {prediction.NgramFeatures.Length}");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

similar comment to the other file, is this needed? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes.


In reply to: 271819189 [](ancestors = 271819189)

@codecov

codecovBot commented Apr 3, 2019

Copy link
Copy Markdown

Codecov Report

Merging #3177 into master will increase coverage by 0.04%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #3177 +/- ##
==========================================
+ Coverage 72.54% 72.58% +0.04% 
==========================================
Files 807 807 Lines 144774 144956 +182 Branches 16208 16212 +4 ==========================================
+ Hits 105022 105215 +193 + Misses 35338 35325 -13 - Partials 4414 4416 +2
FlagCoverage Δ
#Debug72.58% <ø> (+0.04%)⬆️
#production68.14% <ø> (+0.01%)⬆️
#test88.88% <ø> (+0.04%)⬆️
Impacted FilesCoverage Δ
src/Microsoft.ML.Transforms/Text/TextCatalog.cs41.66% <ø> (ø)⬆️
src/Microsoft.ML.DataView/KeyDataViewType.cs74.57% <0%> (-3.76%)⬇️
src/Microsoft.ML.Transforms/Text/LdaTransform.cs89.26% <0%> (-0.63%)⬇️
...soft.ML.TestFramework/DataPipe/TestDataPipeBase.cs73.7% <0%> (-0.34%)⬇️
test/Microsoft.ML.Tests/ImagesTests.cs98.69% <0%> (-0.13%)⬇️
...Microsoft.ML.Tests/Transformers/NormalizerTests.cs100% <0%> (ø)⬆️
...ML.Data/Transforms/ConversionsExtensionsCatalog.cs44.87% <0%> (ø)⬆️
src/Microsoft.ML.Maml/MAML.cs26.21% <0%> (+1.45%)⬆️
...rosoft.ML.ImageAnalytics/VectorToImageTransform.cs76.77% <0%> (+4.53%)⬆️
... and 3 more

new TextData(){ Text = "Each position in the vector corresponds to a particular Ngram." },
new TextData(){ Text = "The value at each position corresponds to," },
new TextData(){ Text = "the number of times Ngram occured in the data (Tf), or" },
new TextData(){ Text = "the inverse of the number of documents that contain the Ngram (Idf), or." },

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

or. [](start = 109, length = 3)

omit this one. #Resolved

new TextData(){ Text = "This is an example to compute Ngrams using hashing." },
new TextData(){ Text = "Ngram is a sequence of 'N' consecutive words/tokens." },
new TextData(){ Text = "ML.NET's ProduceHashedNgrams API produces count of Ngrams and hashes it as an index into a vector of given bit length." },
new TextData(){ Text = "The hashing schem reduces the size of the output feature vector" },

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

schem [](start = 52, length = 5)

process? #Resolved

Console.Write($"{prediction.NgramFeatures[i]:F4} ");

// Expected output:
// Number of Features: 256

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

// Number of Features: 256 [](start = 12, length = 28)

if you use maximumNumberOfInverts you can also show up slot names as in just ngrams. #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I tried that but its bit complex to represent it on console as many ngram can fall into one slot.


In reply to: 271903797 [](ancestors = 271903797)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Well you can control with that parameter how many you actually want to remember, right? So if you put 1, you will have only one of them


In reply to: 271929298 [](ancestors = 271929298,271903797)

@Ivanidzo4kaIvanidzo4ka left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

// Please note that the length of the output feature vector depends on the n-gram settings.
var textPipeline = mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "Text")
// 'ProduceNgrams' takes key type as input. Converting the tokens into key type using 'MapValueToKey'.
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens")) [](start = 16, length = 64)

This seems like a holdover from the internal codebase. I wonder if we should consider doing a breaking change to move this keytype conversion into the ProduceNGrams operation. The question is what other use cases do we expect to see?

@rogancarrrogancarr left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved with comments.

@zeahmed
zeahmed merged commit 854f154 into dotnet:masterApr 5, 2019
zeahmed added a commit to zeahmed/machinelearning that referenced this pull request Apr 8, 2019
@zeahmedzeahmed mentioned this pull request Apr 8, 2019
@ghostghost locked as resolved and limited conversation to collaborators Mar 23, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@zeahmed@Ivanidzo4ka@sfilipi@rogancarr
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Created samples for 'ProduceNgrams' and 'ProduceHashedNgrams' APIs. - #3177

Merged
zeahmed merged 8 commits into
dotnet:masterfrom
zeahmed:ngram_samples
Apr 5, 2019
Merged

Created samples for 'ProduceNgrams' and 'ProduceHashedNgrams' APIs.#3177
zeahmed merged 8 commits into
dotnet:masterfrom
zeahmed:ngram_samples

Conversation

@zeahmed

Copy link
Copy Markdown
Contributor

Related to #1209.

foreach (var item in featureRow.Items())
Console.Write($"{slots[item.Key]} ");
Console.WriteLine();
}

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is the concern for me right now. There is no way to get this meta data through the Transformer or through the prediction engine. The only way is through IDataView obtained from .Transform call.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you have to transform the data anyway, why not just print off the transformed IDV instead of off a prediction? You can limit it to one row with a TakeRows filter.


In reply to: 271476465 [](ancestors = 271476465)

new TextData(){ Text = "The value at each position corresponds to," },
new TextData(){ Text = "the number of times Ngram occured in the data (Tf), or" },
new TextData(){ Text = "the inverse of the number of documents that contain the Ngram (Idf), or." },
new TextData(){ Text = "or compute both and multipy together (Tf-Idf)." },

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

love it! #Closed

var textPipeline = mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "Text")
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))
.Append(mlContext.Transforms.Text.ProduceNgrams("NgramFeatures", "Tokens",
ngramLength: 3, useAllLengths: false, weighting: NgramExtractingEstimator.WeightingCriteria.Tf));

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ngramLength: 3, useAllLengths: false, weighting: NgramExtractingEstimator.WeightingCriteria.Tf [](start = 16, length = 94)

one parameter in one line might be more presentable. #Resolved

// This is acheived by calling 'TokenizeIntoWords' first followed by 'ProduceNgrams'.
// Please note that the length of the output feature vector depends on the Ngram settings.
var textPipeline = mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "Text")
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens")) [](start = 14, length = 66)

add one line of comment on why this is here. #Resolved

}

// Print the first 10 feature values.
Console.Write("Features: ");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Console.Write("Features: "); [](start = 11, length = 29)

i'd remove unnecessary printings

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actually, I'd like to keep it because its more clear to read each line with this prefix on console.


In reply to: 271816037 [](ancestors = 271816037)

transformedDataView.Schema["NgramFeatures"].GetSlotNames(ref slotNames);
var NgramFeaturesColumn = transformedDataView.GetColumn<VBuffer<float>>(transformedDataView.Schema["NgramFeatures"]);
var slots = slotNames.GetValues();
Console.Write("Ngrams: ");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Console.Write("Ngrams: "); [](start = 10, length = 28)

i'd remove this too.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actually, I'd like to keep it because its more clear to read each line with this prefix on console.


In reply to: 271816588 [](ancestors = 271816588)

// Print the length of the feature vector.
Console.WriteLine($"Number of Features: {prediction.NgramFeatures.Length}");

// Preview of the produced .

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

// Preview of the produced . [](start = 11, length = 29)

comment about slot names #Resolved

var prediction = predictionEngine.Predict(samples[0]);

// Print the length of the feature vector.
Console.WriteLine($"Number of Features: {prediction.NgramFeatures.Length}");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is this necessary? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes.


In reply to: 271817449 [](ancestors = 271817449)

var prediction = predictionEngine.Predict(samples[0]);

// Print the length of the feature vector.
Console.WriteLine($"Number of Features: {prediction.NgramFeatures.Length}");

@sfilipisfilipiApr 3, 2019

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

similar comment to the other file, is this needed? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes.


In reply to: 271819189 [](ancestors = 271819189)

@codecov

codecovBot commented Apr 3, 2019

Copy link
Copy Markdown

Codecov Report

Merging #3177 into master will increase coverage by 0.04%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #3177 +/- ##
==========================================
+ Coverage 72.54% 72.58% +0.04% 
==========================================
Files 807 807 Lines 144774 144956 +182 Branches 16208 16212 +4 ==========================================
+ Hits 105022 105215 +193 + Misses 35338 35325 -13 - Partials 4414 4416 +2
FlagCoverage Δ
#Debug72.58% <ø> (+0.04%)⬆️
#production68.14% <ø> (+0.01%)⬆️
#test88.88% <ø> (+0.04%)⬆️
Impacted FilesCoverage Δ
src/Microsoft.ML.Transforms/Text/TextCatalog.cs41.66% <ø> (ø)⬆️
src/Microsoft.ML.DataView/KeyDataViewType.cs74.57% <0%> (-3.76%)⬇️
src/Microsoft.ML.Transforms/Text/LdaTransform.cs89.26% <0%> (-0.63%)⬇️
...soft.ML.TestFramework/DataPipe/TestDataPipeBase.cs73.7% <0%> (-0.34%)⬇️
test/Microsoft.ML.Tests/ImagesTests.cs98.69% <0%> (-0.13%)⬇️
...Microsoft.ML.Tests/Transformers/NormalizerTests.cs100% <0%> (ø)⬆️
...ML.Data/Transforms/ConversionsExtensionsCatalog.cs44.87% <0%> (ø)⬆️
src/Microsoft.ML.Maml/MAML.cs26.21% <0%> (+1.45%)⬆️
...rosoft.ML.ImageAnalytics/VectorToImageTransform.cs76.77% <0%> (+4.53%)⬆️
... and 3 more

new TextData(){ Text = "Each position in the vector corresponds to a particular Ngram." },
new TextData(){ Text = "The value at each position corresponds to," },
new TextData(){ Text = "the number of times Ngram occured in the data (Tf), or" },
new TextData(){ Text = "the inverse of the number of documents that contain the Ngram (Idf), or." },

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

or. [](start = 109, length = 3)

omit this one. #Resolved

new TextData(){ Text = "This is an example to compute Ngrams using hashing." },
new TextData(){ Text = "Ngram is a sequence of 'N' consecutive words/tokens." },
new TextData(){ Text = "ML.NET's ProduceHashedNgrams API produces count of Ngrams and hashes it as an index into a vector of given bit length." },
new TextData(){ Text = "The hashing schem reduces the size of the output feature vector" },

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

schem [](start = 52, length = 5)

process? #Resolved

Console.Write($"{prediction.NgramFeatures[i]:F4} ");

// Expected output:
// Number of Features: 256

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

// Number of Features: 256 [](start = 12, length = 28)

if you use maximumNumberOfInverts you can also show up slot names as in just ngrams. #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I tried that but its bit complex to represent it on console as many ngram can fall into one slot.


In reply to: 271903797 [](ancestors = 271903797)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Well you can control with that parameter how many you actually want to remember, right? So if you put 1, you will have only one of them


In reply to: 271929298 [](ancestors = 271929298,271903797)

@Ivanidzo4kaIvanidzo4ka left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

// Please note that the length of the output feature vector depends on the n-gram settings.
var textPipeline = mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "Text")
// 'ProduceNgrams' takes key type as input. Converting the tokens into key type using 'MapValueToKey'.
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens")) [](start = 16, length = 64)

This seems like a holdover from the internal codebase. I wonder if we should consider doing a breaking change to move this keytype conversion into the ProduceNGrams operation. The question is what other use cases do we expect to see?

@rogancarrrogancarr left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved with comments.

@zeahmed
zeahmed merged commit 854f154 into dotnet:masterApr 5, 2019
zeahmed added a commit to zeahmed/machinelearning that referenced this pull request Apr 8, 2019
@zeahmedzeahmed mentioned this pull request Apr 8, 2019
@ghostghost locked as resolved and limited conversation to collaborators Mar 23, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@zeahmed@Ivanidzo4ka@sfilipi@rogancarr