Created sample for 'LatentDirichletAllocation' API. - #3191

Merged
zeahmed merged 5 commits into
dotnet:masterfrom
zeahmed:lda_sample
Apr 5, 2019
Merged

Created sample for 'LatentDirichletAllocation' API.#3191
zeahmed merged 5 commits into
dotnet:masterfrom
zeahmed:lda_sample

Conversation

@zeahmed

Copy link
Copy Markdown
Contributor

Related to #1209.

// before passing tokens to LatentDirichletAllocation.
var pipeline = mlContext.Transforms.Text.NormalizeText("normText", "Text")
.Append(mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "normText"))
.Append(mlContext.Transforms.Text.RemoveStopWords("Tokens"))

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

RemoveStopWords [](start = 50, length = 15)

Funny, this is custom stop words remover with no stop words.
So it does nothing.

I guess we need to remove param from params string[] stopwords) in RemoveStopWords #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

opss...That's why I was not getting the output what I was expecting...:)


In reply to: 271928104 [](ancestors = 271928104)

.Append(mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "normText"))
.Append(mlContext.Transforms.Text.RemoveStopWords("Tokens"))
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))
.Append(mlContext.Transforms.Text.ProduceNgrams("Tokens"))

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ProduceNgrams [](start = 50, length = 13)

Do we actually want to run LDA on top of 2 ngrams since 2 is default value for ProduceNgrams or we should recommend to use ngrams:1 ? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

2 is fine as it backs off to unigram ( useAllLengths=true) . I think higher is better in case there is a lot of data available.


In reply to: 271930804 [](ancestors = 271930804)

@codecov

codecovBot commented Apr 3, 2019

Copy link
Copy Markdown

Codecov Report

Merging #3191 into master will decrease coverage by 0.01%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #3191 +/- ##
==========================================
- Coverage 72.6% 72.58% -0.02% 
==========================================
Files 807 807 Lines 144957 144957 Branches 16211 16211 ==========================================
- Hits 105240 105215 -25 - Misses 35298 35326 +28 + Partials 4419 4416 -3
FlagCoverage Δ
#Debug72.58% <ø> (-0.02%)⬇️
#production68.14% <ø> (-0.03%)⬇️
#test88.88% <ø> (ø)⬆️
Impacted FilesCoverage Δ
src/Microsoft.ML.Transforms/Text/TextCatalog.cs41.66% <ø> (ø)⬆️
src/Microsoft.ML.Core/Data/ProgressReporter.cs70.95% <0%> (-6.99%)⬇️
src/Microsoft.ML.Maml/MAML.cs24.75% <0%> (-1.46%)⬇️
src/Microsoft.ML.Transforms/Text/LdaTransform.cs89.26% <0%> (-0.63%)⬇️
...ML.Transforms/Text/StopWordsRemovingTransformer.cs86.26% <0%> (+0.15%)⬆️

@codecov

codecovBot commented Apr 3, 2019

Copy link
Copy Markdown

Codecov Report

Merging #3191 into master will decrease coverage by 0.01%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #3191 +/- ##
==========================================
- Coverage 72.6% 72.58% -0.02% 
==========================================
Files 807 807 Lines 144957 144957 Branches 16211 16211 ==========================================
- Hits 105240 105221 -19 - Misses 35298 35321 +23 + Partials 4419 4415 -4
FlagCoverage Δ
#Debug72.58% <ø> (-0.02%)⬇️
#production68.15% <ø> (-0.02%)⬇️
#test88.88% <ø> (ø)⬆️
Impacted FilesCoverage Δ
src/Microsoft.ML.Transforms/Text/TextCatalog.cs41.66% <ø> (ø)⬆️
src/Microsoft.ML.Core/Data/ProgressReporter.cs70.95% <0%> (-6.99%)⬇️

// Create a small dataset as an IEnumerable.
var samples = new List<TextData>()
{
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API computes topic model." },

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

topic model [](start = 88, length = 11)

models with an s #Resolved

{
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API computes topic model." },
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API is the best for topic model." },
new TextData(){ Text = "I like to eat broccoli and banana." },

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

banana [](start = 67, length = 6)

bananas #Resolved

new TextData(){ Text = "ML.NET's LatentDirichletAllocation API computes topic model." },
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API is the best for topic model." },
new TextData(){ Text = "I like to eat broccoli and banana." },
new TextData(){ Text = "I eat a banana in the breakfast." },

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

in the [](start = 55, length = 6)

for #Resolved


// A pipeline for featurizing the text/string using LatentDirichletAllocation API.
// To be more accurate in computing the LDA features, the pipeline first normalizes text and removes stop words
// before passing tokens to LatentDirichletAllocation.

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tokens [](start = 30, length = 6)

"tokens (the individual words, lower cased, with common words removed)"

Many people won't be familiar with the specific language of NLP. #Resolved

// A pipeline for featurizing the text/string using LatentDirichletAllocation API.
// To be more accurate in computing the LDA features, the pipeline first normalizes text and removes stop words
// before passing tokens to LatentDirichletAllocation.
var pipeline = mlContext.Transforms.Text.NormalizeText("normText", "Text")

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

normText [](start = 68, length = 8)

I would spell it out for the example. #Resolved

var transformer = pipeline.Fit(dataview);

// Create the prediction engine to get the LDA features extracted from the text.
var predictionEngine = mlContext.Model.CreatePredictionEngine<TextData, TransformedTextData>(transformer);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

predictionEngine [](start = 16, length = 16)

Similar to the other PR, I wonder if we should stay entirely within IDataView and not create a prediction engine. That is, use a TakeRows filter followed by a CreateEnumerable.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is done because some of the transforms related to text processing such as (NormalizeText. TokenizeIntoWords etc.) don't need training data. In such cases, prediction engine seems more appropriate. But we can definitely have consensus on this. I will follow up.


In reply to: 271981536 [](ancestors = 271981536)

@rogancarrrogancarr left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved with comments.

🚴

// 0.5455 0.1818 0.2727
}

private static void PrintPredictions(TransformedTextData prediction)

@wschinwschinApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
privatestaticvoidPrintPredictions(TransformedTextDataprediction)
privatestaticvoidPrintLdaFeatures(TransformedTextDataprediction)
``` #Resolved

Console.WriteLine();
}

public class TextData

@wschinwschinApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
publicclass TextData
privateclass TextData
``` #Resolved

public string Text { get; set; }
}

public class TransformedTextData : TextData

@wschinwschinApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
publicclassTransformedTextData:TextData
privateclassTransformedTextData:TextData
``` #Resolved

@Ivanidzo4kaIvanidzo4ka left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@zeahmed
zeahmed merged commit a8915f4 into dotnet:masterApr 5, 2019
zeahmed added a commit to zeahmed/machinelearning that referenced this pull request Apr 8, 2019
@zeahmedzeahmed mentioned this pull request Apr 8, 2019
@ghostghost locked as resolved and limited conversation to collaborators Mar 23, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@zeahmed@Ivanidzo4ka@wschin@rogancarr
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Created sample for 'LatentDirichletAllocation' API. - #3191

Merged
zeahmed merged 5 commits into
dotnet:masterfrom
zeahmed:lda_sample
Apr 5, 2019
Merged

Created sample for 'LatentDirichletAllocation' API.#3191
zeahmed merged 5 commits into
dotnet:masterfrom
zeahmed:lda_sample

Conversation

@zeahmed

Copy link
Copy Markdown
Contributor

Related to #1209.

// before passing tokens to LatentDirichletAllocation.
var pipeline = mlContext.Transforms.Text.NormalizeText("normText", "Text")
.Append(mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "normText"))
.Append(mlContext.Transforms.Text.RemoveStopWords("Tokens"))

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

RemoveStopWords [](start = 50, length = 15)

Funny, this is custom stop words remover with no stop words.
So it does nothing.

I guess we need to remove param from params string[] stopwords) in RemoveStopWords #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

opss...That's why I was not getting the output what I was expecting...:)


In reply to: 271928104 [](ancestors = 271928104)

.Append(mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "normText"))
.Append(mlContext.Transforms.Text.RemoveStopWords("Tokens"))
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))
.Append(mlContext.Transforms.Text.ProduceNgrams("Tokens"))

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ProduceNgrams [](start = 50, length = 13)

Do we actually want to run LDA on top of 2 ngrams since 2 is default value for ProduceNgrams or we should recommend to use ngrams:1 ? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

2 is fine as it backs off to unigram ( useAllLengths=true) . I think higher is better in case there is a lot of data available.


In reply to: 271930804 [](ancestors = 271930804)

@codecov

codecovBot commented Apr 3, 2019

Copy link
Copy Markdown

Codecov Report

Merging #3191 into master will decrease coverage by 0.01%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #3191 +/- ##
==========================================
- Coverage 72.6% 72.58% -0.02% 
==========================================
Files 807 807 Lines 144957 144957 Branches 16211 16211 ==========================================
- Hits 105240 105215 -25 - Misses 35298 35326 +28 + Partials 4419 4416 -3
FlagCoverage Δ
#Debug72.58% <ø> (-0.02%)⬇️
#production68.14% <ø> (-0.03%)⬇️
#test88.88% <ø> (ø)⬆️
Impacted FilesCoverage Δ
src/Microsoft.ML.Transforms/Text/TextCatalog.cs41.66% <ø> (ø)⬆️
src/Microsoft.ML.Core/Data/ProgressReporter.cs70.95% <0%> (-6.99%)⬇️
src/Microsoft.ML.Maml/MAML.cs24.75% <0%> (-1.46%)⬇️
src/Microsoft.ML.Transforms/Text/LdaTransform.cs89.26% <0%> (-0.63%)⬇️
...ML.Transforms/Text/StopWordsRemovingTransformer.cs86.26% <0%> (+0.15%)⬆️

@codecov

codecovBot commented Apr 3, 2019

Copy link
Copy Markdown

Codecov Report

Merging #3191 into master will decrease coverage by 0.01%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #3191 +/- ##
==========================================
- Coverage 72.6% 72.58% -0.02% 
==========================================
Files 807 807 Lines 144957 144957 Branches 16211 16211 ==========================================
- Hits 105240 105221 -19 - Misses 35298 35321 +23 + Partials 4419 4415 -4
FlagCoverage Δ
#Debug72.58% <ø> (-0.02%)⬇️
#production68.15% <ø> (-0.02%)⬇️
#test88.88% <ø> (ø)⬆️
Impacted FilesCoverage Δ
src/Microsoft.ML.Transforms/Text/TextCatalog.cs41.66% <ø> (ø)⬆️
src/Microsoft.ML.Core/Data/ProgressReporter.cs70.95% <0%> (-6.99%)⬇️

// Create a small dataset as an IEnumerable.
var samples = new List<TextData>()
{
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API computes topic model." },

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

topic model [](start = 88, length = 11)

models with an s #Resolved

{
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API computes topic model." },
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API is the best for topic model." },
new TextData(){ Text = "I like to eat broccoli and banana." },

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

banana [](start = 67, length = 6)

bananas #Resolved

new TextData(){ Text = "ML.NET's LatentDirichletAllocation API computes topic model." },
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API is the best for topic model." },
new TextData(){ Text = "I like to eat broccoli and banana." },
new TextData(){ Text = "I eat a banana in the breakfast." },

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

in the [](start = 55, length = 6)

for #Resolved


// A pipeline for featurizing the text/string using LatentDirichletAllocation API.
// To be more accurate in computing the LDA features, the pipeline first normalizes text and removes stop words
// before passing tokens to LatentDirichletAllocation.

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tokens [](start = 30, length = 6)

"tokens (the individual words, lower cased, with common words removed)"

Many people won't be familiar with the specific language of NLP. #Resolved

// A pipeline for featurizing the text/string using LatentDirichletAllocation API.
// To be more accurate in computing the LDA features, the pipeline first normalizes text and removes stop words
// before passing tokens to LatentDirichletAllocation.
var pipeline = mlContext.Transforms.Text.NormalizeText("normText", "Text")

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

normText [](start = 68, length = 8)

I would spell it out for the example. #Resolved

var transformer = pipeline.Fit(dataview);

// Create the prediction engine to get the LDA features extracted from the text.
var predictionEngine = mlContext.Model.CreatePredictionEngine<TextData, TransformedTextData>(transformer);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

predictionEngine [](start = 16, length = 16)

Similar to the other PR, I wonder if we should stay entirely within IDataView and not create a prediction engine. That is, use a TakeRows filter followed by a CreateEnumerable.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is done because some of the transforms related to text processing such as (NormalizeText. TokenizeIntoWords etc.) don't need training data. In such cases, prediction engine seems more appropriate. But we can definitely have consensus on this. I will follow up.


In reply to: 271981536 [](ancestors = 271981536)

@rogancarrrogancarr left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved with comments.

🚴

// 0.5455 0.1818 0.2727
}

private static void PrintPredictions(TransformedTextData prediction)

@wschinwschinApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
privatestaticvoidPrintPredictions(TransformedTextDataprediction)
privatestaticvoidPrintLdaFeatures(TransformedTextDataprediction)
``` #Resolved

Console.WriteLine();
}

public class TextData

@wschinwschinApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
publicclass TextData
privateclass TextData
``` #Resolved

public string Text { get; set; }
}

public class TransformedTextData : TextData

@wschinwschinApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
publicclassTransformedTextData:TextData
privateclassTransformedTextData:TextData
``` #Resolved

@Ivanidzo4kaIvanidzo4ka left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@zeahmed
zeahmed merged commit a8915f4 into dotnet:masterApr 5, 2019
zeahmed added a commit to zeahmed/machinelearning that referenced this pull request Apr 8, 2019
@zeahmedzeahmed mentioned this pull request Apr 8, 2019
@ghostghost locked as resolved and limited conversation to collaborators Mar 23, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@zeahmed@Ivanidzo4ka@wschin@rogancarr
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Created sample for 'LatentDirichletAllocation' API. - #3191

Merged
zeahmed merged 5 commits into
dotnet:masterfrom
zeahmed:lda_sample
Apr 5, 2019
Merged

Created sample for 'LatentDirichletAllocation' API.#3191
zeahmed merged 5 commits into
dotnet:masterfrom
zeahmed:lda_sample

Conversation

@zeahmed

Copy link
Copy Markdown
Contributor

Related to #1209.

// before passing tokens to LatentDirichletAllocation.
var pipeline = mlContext.Transforms.Text.NormalizeText("normText", "Text")
.Append(mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "normText"))
.Append(mlContext.Transforms.Text.RemoveStopWords("Tokens"))

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

RemoveStopWords [](start = 50, length = 15)

Funny, this is custom stop words remover with no stop words.
So it does nothing.

I guess we need to remove param from params string[] stopwords) in RemoveStopWords #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

opss...That's why I was not getting the output what I was expecting...:)


In reply to: 271928104 [](ancestors = 271928104)

.Append(mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "normText"))
.Append(mlContext.Transforms.Text.RemoveStopWords("Tokens"))
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))
.Append(mlContext.Transforms.Text.ProduceNgrams("Tokens"))

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ProduceNgrams [](start = 50, length = 13)

Do we actually want to run LDA on top of 2 ngrams since 2 is default value for ProduceNgrams or we should recommend to use ngrams:1 ? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

2 is fine as it backs off to unigram ( useAllLengths=true) . I think higher is better in case there is a lot of data available.


In reply to: 271930804 [](ancestors = 271930804)

@codecov

codecovBot commented Apr 3, 2019

Copy link
Copy Markdown

Codecov Report

Merging #3191 into master will decrease coverage by 0.01%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #3191 +/- ##
==========================================
- Coverage 72.6% 72.58% -0.02% 
==========================================
Files 807 807 Lines 144957 144957 Branches 16211 16211 ==========================================
- Hits 105240 105215 -25 - Misses 35298 35326 +28 + Partials 4419 4416 -3
FlagCoverage Δ
#Debug72.58% <ø> (-0.02%)⬇️
#production68.14% <ø> (-0.03%)⬇️
#test88.88% <ø> (ø)⬆️
Impacted FilesCoverage Δ
src/Microsoft.ML.Transforms/Text/TextCatalog.cs41.66% <ø> (ø)⬆️
src/Microsoft.ML.Core/Data/ProgressReporter.cs70.95% <0%> (-6.99%)⬇️
src/Microsoft.ML.Maml/MAML.cs24.75% <0%> (-1.46%)⬇️
src/Microsoft.ML.Transforms/Text/LdaTransform.cs89.26% <0%> (-0.63%)⬇️
...ML.Transforms/Text/StopWordsRemovingTransformer.cs86.26% <0%> (+0.15%)⬆️

@codecov

codecovBot commented Apr 3, 2019

Copy link
Copy Markdown

Codecov Report

Merging #3191 into master will decrease coverage by 0.01%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #3191 +/- ##
==========================================
- Coverage 72.6% 72.58% -0.02% 
==========================================
Files 807 807 Lines 144957 144957 Branches 16211 16211 ==========================================
- Hits 105240 105221 -19 - Misses 35298 35321 +23 + Partials 4419 4415 -4
FlagCoverage Δ
#Debug72.58% <ø> (-0.02%)⬇️
#production68.15% <ø> (-0.02%)⬇️
#test88.88% <ø> (ø)⬆️
Impacted FilesCoverage Δ
src/Microsoft.ML.Transforms/Text/TextCatalog.cs41.66% <ø> (ø)⬆️
src/Microsoft.ML.Core/Data/ProgressReporter.cs70.95% <0%> (-6.99%)⬇️

// Create a small dataset as an IEnumerable.
var samples = new List<TextData>()
{
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API computes topic model." },

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

topic model [](start = 88, length = 11)

models with an s #Resolved

{
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API computes topic model." },
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API is the best for topic model." },
new TextData(){ Text = "I like to eat broccoli and banana." },

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

banana [](start = 67, length = 6)

bananas #Resolved

new TextData(){ Text = "ML.NET's LatentDirichletAllocation API computes topic model." },
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API is the best for topic model." },
new TextData(){ Text = "I like to eat broccoli and banana." },
new TextData(){ Text = "I eat a banana in the breakfast." },

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

in the [](start = 55, length = 6)

for #Resolved


// A pipeline for featurizing the text/string using LatentDirichletAllocation API.
// To be more accurate in computing the LDA features, the pipeline first normalizes text and removes stop words
// before passing tokens to LatentDirichletAllocation.

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tokens [](start = 30, length = 6)

"tokens (the individual words, lower cased, with common words removed)"

Many people won't be familiar with the specific language of NLP. #Resolved

// A pipeline for featurizing the text/string using LatentDirichletAllocation API.
// To be more accurate in computing the LDA features, the pipeline first normalizes text and removes stop words
// before passing tokens to LatentDirichletAllocation.
var pipeline = mlContext.Transforms.Text.NormalizeText("normText", "Text")

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

normText [](start = 68, length = 8)

I would spell it out for the example. #Resolved

var transformer = pipeline.Fit(dataview);

// Create the prediction engine to get the LDA features extracted from the text.
var predictionEngine = mlContext.Model.CreatePredictionEngine<TextData, TransformedTextData>(transformer);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

predictionEngine [](start = 16, length = 16)

Similar to the other PR, I wonder if we should stay entirely within IDataView and not create a prediction engine. That is, use a TakeRows filter followed by a CreateEnumerable.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is done because some of the transforms related to text processing such as (NormalizeText. TokenizeIntoWords etc.) don't need training data. In such cases, prediction engine seems more appropriate. But we can definitely have consensus on this. I will follow up.


In reply to: 271981536 [](ancestors = 271981536)

@rogancarrrogancarr left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved with comments.

🚴

// 0.5455 0.1818 0.2727
}

private static void PrintPredictions(TransformedTextData prediction)

@wschinwschinApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
privatestaticvoidPrintPredictions(TransformedTextDataprediction)
privatestaticvoidPrintLdaFeatures(TransformedTextDataprediction)
``` #Resolved

Console.WriteLine();
}

public class TextData

@wschinwschinApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
publicclass TextData
privateclass TextData
``` #Resolved

public string Text { get; set; }
}

public class TransformedTextData : TextData

@wschinwschinApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
publicclassTransformedTextData:TextData
privateclassTransformedTextData:TextData
``` #Resolved

@Ivanidzo4kaIvanidzo4ka left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@zeahmed
zeahmed merged commit a8915f4 into dotnet:masterApr 5, 2019
zeahmed added a commit to zeahmed/machinelearning that referenced this pull request Apr 8, 2019
@zeahmedzeahmed mentioned this pull request Apr 8, 2019
@ghostghost locked as resolved and limited conversation to collaborators Mar 23, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@zeahmed@Ivanidzo4ka@wschin@rogancarr
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Created sample for 'LatentDirichletAllocation' API. - #3191

Merged
zeahmed merged 5 commits into
dotnet:masterfrom
zeahmed:lda_sample
Apr 5, 2019
Merged

Created sample for 'LatentDirichletAllocation' API.#3191
zeahmed merged 5 commits into
dotnet:masterfrom
zeahmed:lda_sample

Conversation

@zeahmed

Copy link
Copy Markdown
Contributor

Related to #1209.

// before passing tokens to LatentDirichletAllocation.
var pipeline = mlContext.Transforms.Text.NormalizeText("normText", "Text")
.Append(mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "normText"))
.Append(mlContext.Transforms.Text.RemoveStopWords("Tokens"))

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

RemoveStopWords [](start = 50, length = 15)

Funny, this is custom stop words remover with no stop words.
So it does nothing.

I guess we need to remove param from params string[] stopwords) in RemoveStopWords #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

opss...That's why I was not getting the output what I was expecting...:)


In reply to: 271928104 [](ancestors = 271928104)

.Append(mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "normText"))
.Append(mlContext.Transforms.Text.RemoveStopWords("Tokens"))
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))
.Append(mlContext.Transforms.Text.ProduceNgrams("Tokens"))

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ProduceNgrams [](start = 50, length = 13)

Do we actually want to run LDA on top of 2 ngrams since 2 is default value for ProduceNgrams or we should recommend to use ngrams:1 ? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

2 is fine as it backs off to unigram ( useAllLengths=true) . I think higher is better in case there is a lot of data available.


In reply to: 271930804 [](ancestors = 271930804)

@codecov

codecovBot commented Apr 3, 2019

Copy link
Copy Markdown

Codecov Report

Merging #3191 into master will decrease coverage by 0.01%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #3191 +/- ##
==========================================
- Coverage 72.6% 72.58% -0.02% 
==========================================
Files 807 807 Lines 144957 144957 Branches 16211 16211 ==========================================
- Hits 105240 105215 -25 - Misses 35298 35326 +28 + Partials 4419 4416 -3
FlagCoverage Δ
#Debug72.58% <ø> (-0.02%)⬇️
#production68.14% <ø> (-0.03%)⬇️
#test88.88% <ø> (ø)⬆️
Impacted FilesCoverage Δ
src/Microsoft.ML.Transforms/Text/TextCatalog.cs41.66% <ø> (ø)⬆️
src/Microsoft.ML.Core/Data/ProgressReporter.cs70.95% <0%> (-6.99%)⬇️
src/Microsoft.ML.Maml/MAML.cs24.75% <0%> (-1.46%)⬇️
src/Microsoft.ML.Transforms/Text/LdaTransform.cs89.26% <0%> (-0.63%)⬇️
...ML.Transforms/Text/StopWordsRemovingTransformer.cs86.26% <0%> (+0.15%)⬆️

@codecov

codecovBot commented Apr 3, 2019

Copy link
Copy Markdown

Codecov Report

Merging #3191 into master will decrease coverage by 0.01%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #3191 +/- ##
==========================================
- Coverage 72.6% 72.58% -0.02% 
==========================================
Files 807 807 Lines 144957 144957 Branches 16211 16211 ==========================================
- Hits 105240 105221 -19 - Misses 35298 35321 +23 + Partials 4419 4415 -4
FlagCoverage Δ
#Debug72.58% <ø> (-0.02%)⬇️
#production68.15% <ø> (-0.02%)⬇️
#test88.88% <ø> (ø)⬆️
Impacted FilesCoverage Δ
src/Microsoft.ML.Transforms/Text/TextCatalog.cs41.66% <ø> (ø)⬆️
src/Microsoft.ML.Core/Data/ProgressReporter.cs70.95% <0%> (-6.99%)⬇️

// Create a small dataset as an IEnumerable.
var samples = new List<TextData>()
{
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API computes topic model." },

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

topic model [](start = 88, length = 11)

models with an s #Resolved

{
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API computes topic model." },
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API is the best for topic model." },
new TextData(){ Text = "I like to eat broccoli and banana." },

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

banana [](start = 67, length = 6)

bananas #Resolved

new TextData(){ Text = "ML.NET's LatentDirichletAllocation API computes topic model." },
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API is the best for topic model." },
new TextData(){ Text = "I like to eat broccoli and banana." },
new TextData(){ Text = "I eat a banana in the breakfast." },

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

in the [](start = 55, length = 6)

for #Resolved


// A pipeline for featurizing the text/string using LatentDirichletAllocation API.
// To be more accurate in computing the LDA features, the pipeline first normalizes text and removes stop words
// before passing tokens to LatentDirichletAllocation.

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tokens [](start = 30, length = 6)

"tokens (the individual words, lower cased, with common words removed)"

Many people won't be familiar with the specific language of NLP. #Resolved

// A pipeline for featurizing the text/string using LatentDirichletAllocation API.
// To be more accurate in computing the LDA features, the pipeline first normalizes text and removes stop words
// before passing tokens to LatentDirichletAllocation.
var pipeline = mlContext.Transforms.Text.NormalizeText("normText", "Text")

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

normText [](start = 68, length = 8)

I would spell it out for the example. #Resolved

var transformer = pipeline.Fit(dataview);

// Create the prediction engine to get the LDA features extracted from the text.
var predictionEngine = mlContext.Model.CreatePredictionEngine<TextData, TransformedTextData>(transformer);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

predictionEngine [](start = 16, length = 16)

Similar to the other PR, I wonder if we should stay entirely within IDataView and not create a prediction engine. That is, use a TakeRows filter followed by a CreateEnumerable.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is done because some of the transforms related to text processing such as (NormalizeText. TokenizeIntoWords etc.) don't need training data. In such cases, prediction engine seems more appropriate. But we can definitely have consensus on this. I will follow up.


In reply to: 271981536 [](ancestors = 271981536)

@rogancarrrogancarr left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved with comments.

🚴

// 0.5455 0.1818 0.2727
}

private static void PrintPredictions(TransformedTextData prediction)

@wschinwschinApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
privatestaticvoidPrintPredictions(TransformedTextDataprediction)
privatestaticvoidPrintLdaFeatures(TransformedTextDataprediction)
``` #Resolved

Console.WriteLine();
}

public class TextData

@wschinwschinApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
publicclass TextData
privateclass TextData
``` #Resolved

public string Text { get; set; }
}

public class TransformedTextData : TextData

@wschinwschinApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
publicclassTransformedTextData:TextData
privateclassTransformedTextData:TextData
``` #Resolved

@Ivanidzo4kaIvanidzo4ka left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@zeahmed
zeahmed merged commit a8915f4 into dotnet:masterApr 5, 2019
zeahmed added a commit to zeahmed/machinelearning that referenced this pull request Apr 8, 2019
@zeahmedzeahmed mentioned this pull request Apr 8, 2019
@ghostghost locked as resolved and limited conversation to collaborators Mar 23, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@zeahmed@Ivanidzo4ka@wschin@rogancarr
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Created sample for 'LatentDirichletAllocation' API. - #3191

Merged
zeahmed merged 5 commits into
dotnet:masterfrom
zeahmed:lda_sample
Apr 5, 2019
Merged

Created sample for 'LatentDirichletAllocation' API.#3191
zeahmed merged 5 commits into
dotnet:masterfrom
zeahmed:lda_sample

Conversation

@zeahmed

Copy link
Copy Markdown
Contributor

Related to #1209.

// before passing tokens to LatentDirichletAllocation.
var pipeline = mlContext.Transforms.Text.NormalizeText("normText", "Text")
.Append(mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "normText"))
.Append(mlContext.Transforms.Text.RemoveStopWords("Tokens"))

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

RemoveStopWords [](start = 50, length = 15)

Funny, this is custom stop words remover with no stop words.
So it does nothing.

I guess we need to remove param from params string[] stopwords) in RemoveStopWords #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

opss...That's why I was not getting the output what I was expecting...:)


In reply to: 271928104 [](ancestors = 271928104)

.Append(mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "normText"))
.Append(mlContext.Transforms.Text.RemoveStopWords("Tokens"))
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))
.Append(mlContext.Transforms.Text.ProduceNgrams("Tokens"))

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ProduceNgrams [](start = 50, length = 13)

Do we actually want to run LDA on top of 2 ngrams since 2 is default value for ProduceNgrams or we should recommend to use ngrams:1 ? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

2 is fine as it backs off to unigram ( useAllLengths=true) . I think higher is better in case there is a lot of data available.


In reply to: 271930804 [](ancestors = 271930804)

@codecov

codecovBot commented Apr 3, 2019

Copy link
Copy Markdown

Codecov Report

Merging #3191 into master will decrease coverage by 0.01%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #3191 +/- ##
==========================================
- Coverage 72.6% 72.58% -0.02% 
==========================================
Files 807 807 Lines 144957 144957 Branches 16211 16211 ==========================================
- Hits 105240 105215 -25 - Misses 35298 35326 +28 + Partials 4419 4416 -3
FlagCoverage Δ
#Debug72.58% <ø> (-0.02%)⬇️
#production68.14% <ø> (-0.03%)⬇️
#test88.88% <ø> (ø)⬆️
Impacted FilesCoverage Δ
src/Microsoft.ML.Transforms/Text/TextCatalog.cs41.66% <ø> (ø)⬆️
src/Microsoft.ML.Core/Data/ProgressReporter.cs70.95% <0%> (-6.99%)⬇️
src/Microsoft.ML.Maml/MAML.cs24.75% <0%> (-1.46%)⬇️
src/Microsoft.ML.Transforms/Text/LdaTransform.cs89.26% <0%> (-0.63%)⬇️
...ML.Transforms/Text/StopWordsRemovingTransformer.cs86.26% <0%> (+0.15%)⬆️

@codecov

codecovBot commented Apr 3, 2019

Copy link
Copy Markdown

Codecov Report

Merging #3191 into master will decrease coverage by 0.01%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #3191 +/- ##
==========================================
- Coverage 72.6% 72.58% -0.02% 
==========================================
Files 807 807 Lines 144957 144957 Branches 16211 16211 ==========================================
- Hits 105240 105221 -19 - Misses 35298 35321 +23 + Partials 4419 4415 -4
FlagCoverage Δ
#Debug72.58% <ø> (-0.02%)⬇️
#production68.15% <ø> (-0.02%)⬇️
#test88.88% <ø> (ø)⬆️
Impacted FilesCoverage Δ
src/Microsoft.ML.Transforms/Text/TextCatalog.cs41.66% <ø> (ø)⬆️
src/Microsoft.ML.Core/Data/ProgressReporter.cs70.95% <0%> (-6.99%)⬇️

// Create a small dataset as an IEnumerable.
var samples = new List<TextData>()
{
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API computes topic model." },

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

topic model [](start = 88, length = 11)

models with an s #Resolved

{
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API computes topic model." },
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API is the best for topic model." },
new TextData(){ Text = "I like to eat broccoli and banana." },

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

banana [](start = 67, length = 6)

bananas #Resolved

new TextData(){ Text = "ML.NET's LatentDirichletAllocation API computes topic model." },
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API is the best for topic model." },
new TextData(){ Text = "I like to eat broccoli and banana." },
new TextData(){ Text = "I eat a banana in the breakfast." },

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

in the [](start = 55, length = 6)

for #Resolved


// A pipeline for featurizing the text/string using LatentDirichletAllocation API.
// To be more accurate in computing the LDA features, the pipeline first normalizes text and removes stop words
// before passing tokens to LatentDirichletAllocation.

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tokens [](start = 30, length = 6)

"tokens (the individual words, lower cased, with common words removed)"

Many people won't be familiar with the specific language of NLP. #Resolved

// A pipeline for featurizing the text/string using LatentDirichletAllocation API.
// To be more accurate in computing the LDA features, the pipeline first normalizes text and removes stop words
// before passing tokens to LatentDirichletAllocation.
var pipeline = mlContext.Transforms.Text.NormalizeText("normText", "Text")

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

normText [](start = 68, length = 8)

I would spell it out for the example. #Resolved

var transformer = pipeline.Fit(dataview);

// Create the prediction engine to get the LDA features extracted from the text.
var predictionEngine = mlContext.Model.CreatePredictionEngine<TextData, TransformedTextData>(transformer);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

predictionEngine [](start = 16, length = 16)

Similar to the other PR, I wonder if we should stay entirely within IDataView and not create a prediction engine. That is, use a TakeRows filter followed by a CreateEnumerable.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is done because some of the transforms related to text processing such as (NormalizeText. TokenizeIntoWords etc.) don't need training data. In such cases, prediction engine seems more appropriate. But we can definitely have consensus on this. I will follow up.


In reply to: 271981536 [](ancestors = 271981536)

@rogancarrrogancarr left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved with comments.

🚴

// 0.5455 0.1818 0.2727
}

private static void PrintPredictions(TransformedTextData prediction)

@wschinwschinApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
privatestaticvoidPrintPredictions(TransformedTextDataprediction)
privatestaticvoidPrintLdaFeatures(TransformedTextDataprediction)
``` #Resolved

Console.WriteLine();
}

public class TextData

@wschinwschinApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
publicclass TextData
privateclass TextData
``` #Resolved

public string Text { get; set; }
}

public class TransformedTextData : TextData

@wschinwschinApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
publicclassTransformedTextData:TextData
privateclassTransformedTextData:TextData
``` #Resolved

@Ivanidzo4kaIvanidzo4ka left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@zeahmed
zeahmed merged commit a8915f4 into dotnet:masterApr 5, 2019
zeahmed added a commit to zeahmed/machinelearning that referenced this pull request Apr 8, 2019
@zeahmedzeahmed mentioned this pull request Apr 8, 2019
@ghostghost locked as resolved and limited conversation to collaborators Mar 23, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@zeahmed@Ivanidzo4ka@wschin@rogancarr
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Created sample for 'LatentDirichletAllocation' API. - #3191

Merged
zeahmed merged 5 commits into
dotnet:masterfrom
zeahmed:lda_sample
Apr 5, 2019
Merged

Created sample for 'LatentDirichletAllocation' API.#3191
zeahmed merged 5 commits into
dotnet:masterfrom
zeahmed:lda_sample

Conversation

@zeahmed

Copy link
Copy Markdown
Contributor

Related to #1209.

// before passing tokens to LatentDirichletAllocation.
var pipeline = mlContext.Transforms.Text.NormalizeText("normText", "Text")
.Append(mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "normText"))
.Append(mlContext.Transforms.Text.RemoveStopWords("Tokens"))

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

RemoveStopWords [](start = 50, length = 15)

Funny, this is custom stop words remover with no stop words.
So it does nothing.

I guess we need to remove param from params string[] stopwords) in RemoveStopWords #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

opss...That's why I was not getting the output what I was expecting...:)


In reply to: 271928104 [](ancestors = 271928104)

.Append(mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "normText"))
.Append(mlContext.Transforms.Text.RemoveStopWords("Tokens"))
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))
.Append(mlContext.Transforms.Text.ProduceNgrams("Tokens"))

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ProduceNgrams [](start = 50, length = 13)

Do we actually want to run LDA on top of 2 ngrams since 2 is default value for ProduceNgrams or we should recommend to use ngrams:1 ? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

2 is fine as it backs off to unigram ( useAllLengths=true) . I think higher is better in case there is a lot of data available.


In reply to: 271930804 [](ancestors = 271930804)

@codecov

codecovBot commented Apr 3, 2019

Copy link
Copy Markdown

Codecov Report

Merging #3191 into master will decrease coverage by 0.01%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #3191 +/- ##
==========================================
- Coverage 72.6% 72.58% -0.02% 
==========================================
Files 807 807 Lines 144957 144957 Branches 16211 16211 ==========================================
- Hits 105240 105215 -25 - Misses 35298 35326 +28 + Partials 4419 4416 -3
FlagCoverage Δ
#Debug72.58% <ø> (-0.02%)⬇️
#production68.14% <ø> (-0.03%)⬇️
#test88.88% <ø> (ø)⬆️
Impacted FilesCoverage Δ
src/Microsoft.ML.Transforms/Text/TextCatalog.cs41.66% <ø> (ø)⬆️
src/Microsoft.ML.Core/Data/ProgressReporter.cs70.95% <0%> (-6.99%)⬇️
src/Microsoft.ML.Maml/MAML.cs24.75% <0%> (-1.46%)⬇️
src/Microsoft.ML.Transforms/Text/LdaTransform.cs89.26% <0%> (-0.63%)⬇️
...ML.Transforms/Text/StopWordsRemovingTransformer.cs86.26% <0%> (+0.15%)⬆️

@codecov

codecovBot commented Apr 3, 2019

Copy link
Copy Markdown

Codecov Report

Merging #3191 into master will decrease coverage by 0.01%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #3191 +/- ##
==========================================
- Coverage 72.6% 72.58% -0.02% 
==========================================
Files 807 807 Lines 144957 144957 Branches 16211 16211 ==========================================
- Hits 105240 105221 -19 - Misses 35298 35321 +23 + Partials 4419 4415 -4
FlagCoverage Δ
#Debug72.58% <ø> (-0.02%)⬇️
#production68.15% <ø> (-0.02%)⬇️
#test88.88% <ø> (ø)⬆️
Impacted FilesCoverage Δ
src/Microsoft.ML.Transforms/Text/TextCatalog.cs41.66% <ø> (ø)⬆️
src/Microsoft.ML.Core/Data/ProgressReporter.cs70.95% <0%> (-6.99%)⬇️

// Create a small dataset as an IEnumerable.
var samples = new List<TextData>()
{
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API computes topic model." },

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

topic model [](start = 88, length = 11)

models with an s #Resolved

{
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API computes topic model." },
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API is the best for topic model." },
new TextData(){ Text = "I like to eat broccoli and banana." },

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

banana [](start = 67, length = 6)

bananas #Resolved

new TextData(){ Text = "ML.NET's LatentDirichletAllocation API computes topic model." },
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API is the best for topic model." },
new TextData(){ Text = "I like to eat broccoli and banana." },
new TextData(){ Text = "I eat a banana in the breakfast." },

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

in the [](start = 55, length = 6)

for #Resolved


// A pipeline for featurizing the text/string using LatentDirichletAllocation API.
// To be more accurate in computing the LDA features, the pipeline first normalizes text and removes stop words
// before passing tokens to LatentDirichletAllocation.

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tokens [](start = 30, length = 6)

"tokens (the individual words, lower cased, with common words removed)"

Many people won't be familiar with the specific language of NLP. #Resolved

// A pipeline for featurizing the text/string using LatentDirichletAllocation API.
// To be more accurate in computing the LDA features, the pipeline first normalizes text and removes stop words
// before passing tokens to LatentDirichletAllocation.
var pipeline = mlContext.Transforms.Text.NormalizeText("normText", "Text")

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

normText [](start = 68, length = 8)

I would spell it out for the example. #Resolved

var transformer = pipeline.Fit(dataview);

// Create the prediction engine to get the LDA features extracted from the text.
var predictionEngine = mlContext.Model.CreatePredictionEngine<TextData, TransformedTextData>(transformer);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

predictionEngine [](start = 16, length = 16)

Similar to the other PR, I wonder if we should stay entirely within IDataView and not create a prediction engine. That is, use a TakeRows filter followed by a CreateEnumerable.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is done because some of the transforms related to text processing such as (NormalizeText. TokenizeIntoWords etc.) don't need training data. In such cases, prediction engine seems more appropriate. But we can definitely have consensus on this. I will follow up.


In reply to: 271981536 [](ancestors = 271981536)

@rogancarrrogancarr left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved with comments.

🚴

// 0.5455 0.1818 0.2727
}

private static void PrintPredictions(TransformedTextData prediction)

@wschinwschinApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
privatestaticvoidPrintPredictions(TransformedTextDataprediction)
privatestaticvoidPrintLdaFeatures(TransformedTextDataprediction)
``` #Resolved

Console.WriteLine();
}

public class TextData

@wschinwschinApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
publicclass TextData
privateclass TextData
``` #Resolved

public string Text { get; set; }
}

public class TransformedTextData : TextData

@wschinwschinApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
publicclassTransformedTextData:TextData
privateclassTransformedTextData:TextData
``` #Resolved

@Ivanidzo4kaIvanidzo4ka left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@zeahmed
zeahmed merged commit a8915f4 into dotnet:masterApr 5, 2019
zeahmed added a commit to zeahmed/machinelearning that referenced this pull request Apr 8, 2019
@zeahmedzeahmed mentioned this pull request Apr 8, 2019
@ghostghost locked as resolved and limited conversation to collaborators Mar 23, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@zeahmed@Ivanidzo4ka@wschin@rogancarr
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Created sample for 'LatentDirichletAllocation' API. - #3191

Merged
zeahmed merged 5 commits into
dotnet:masterfrom
zeahmed:lda_sample
Apr 5, 2019
Merged

Created sample for 'LatentDirichletAllocation' API.#3191
zeahmed merged 5 commits into
dotnet:masterfrom
zeahmed:lda_sample

Conversation

@zeahmed

Copy link
Copy Markdown
Contributor

Related to #1209.

// before passing tokens to LatentDirichletAllocation.
var pipeline = mlContext.Transforms.Text.NormalizeText("normText", "Text")
.Append(mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "normText"))
.Append(mlContext.Transforms.Text.RemoveStopWords("Tokens"))

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

RemoveStopWords [](start = 50, length = 15)

Funny, this is custom stop words remover with no stop words.
So it does nothing.

I guess we need to remove param from params string[] stopwords) in RemoveStopWords #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

opss...That's why I was not getting the output what I was expecting...:)


In reply to: 271928104 [](ancestors = 271928104)

.Append(mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "normText"))
.Append(mlContext.Transforms.Text.RemoveStopWords("Tokens"))
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))
.Append(mlContext.Transforms.Text.ProduceNgrams("Tokens"))

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ProduceNgrams [](start = 50, length = 13)

Do we actually want to run LDA on top of 2 ngrams since 2 is default value for ProduceNgrams or we should recommend to use ngrams:1 ? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

2 is fine as it backs off to unigram ( useAllLengths=true) . I think higher is better in case there is a lot of data available.


In reply to: 271930804 [](ancestors = 271930804)

@codecov

codecovBot commented Apr 3, 2019

Copy link
Copy Markdown

Codecov Report

Merging #3191 into master will decrease coverage by 0.01%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #3191 +/- ##
==========================================
- Coverage 72.6% 72.58% -0.02% 
==========================================
Files 807 807 Lines 144957 144957 Branches 16211 16211 ==========================================
- Hits 105240 105215 -25 - Misses 35298 35326 +28 + Partials 4419 4416 -3
FlagCoverage Δ
#Debug72.58% <ø> (-0.02%)⬇️
#production68.14% <ø> (-0.03%)⬇️
#test88.88% <ø> (ø)⬆️
Impacted FilesCoverage Δ
src/Microsoft.ML.Transforms/Text/TextCatalog.cs41.66% <ø> (ø)⬆️
src/Microsoft.ML.Core/Data/ProgressReporter.cs70.95% <0%> (-6.99%)⬇️
src/Microsoft.ML.Maml/MAML.cs24.75% <0%> (-1.46%)⬇️
src/Microsoft.ML.Transforms/Text/LdaTransform.cs89.26% <0%> (-0.63%)⬇️
...ML.Transforms/Text/StopWordsRemovingTransformer.cs86.26% <0%> (+0.15%)⬆️

@codecov

codecovBot commented Apr 3, 2019

Copy link
Copy Markdown

Codecov Report

Merging #3191 into master will decrease coverage by 0.01%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #3191 +/- ##
==========================================
- Coverage 72.6% 72.58% -0.02% 
==========================================
Files 807 807 Lines 144957 144957 Branches 16211 16211 ==========================================
- Hits 105240 105221 -19 - Misses 35298 35321 +23 + Partials 4419 4415 -4
FlagCoverage Δ
#Debug72.58% <ø> (-0.02%)⬇️
#production68.15% <ø> (-0.02%)⬇️
#test88.88% <ø> (ø)⬆️
Impacted FilesCoverage Δ
src/Microsoft.ML.Transforms/Text/TextCatalog.cs41.66% <ø> (ø)⬆️
src/Microsoft.ML.Core/Data/ProgressReporter.cs70.95% <0%> (-6.99%)⬇️

// Create a small dataset as an IEnumerable.
var samples = new List<TextData>()
{
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API computes topic model." },

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

topic model [](start = 88, length = 11)

models with an s #Resolved

{
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API computes topic model." },
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API is the best for topic model." },
new TextData(){ Text = "I like to eat broccoli and banana." },

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

banana [](start = 67, length = 6)

bananas #Resolved

new TextData(){ Text = "ML.NET's LatentDirichletAllocation API computes topic model." },
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API is the best for topic model." },
new TextData(){ Text = "I like to eat broccoli and banana." },
new TextData(){ Text = "I eat a banana in the breakfast." },

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

in the [](start = 55, length = 6)

for #Resolved


// A pipeline for featurizing the text/string using LatentDirichletAllocation API.
// To be more accurate in computing the LDA features, the pipeline first normalizes text and removes stop words
// before passing tokens to LatentDirichletAllocation.

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tokens [](start = 30, length = 6)

"tokens (the individual words, lower cased, with common words removed)"

Many people won't be familiar with the specific language of NLP. #Resolved

// A pipeline for featurizing the text/string using LatentDirichletAllocation API.
// To be more accurate in computing the LDA features, the pipeline first normalizes text and removes stop words
// before passing tokens to LatentDirichletAllocation.
var pipeline = mlContext.Transforms.Text.NormalizeText("normText", "Text")

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

normText [](start = 68, length = 8)

I would spell it out for the example. #Resolved

var transformer = pipeline.Fit(dataview);

// Create the prediction engine to get the LDA features extracted from the text.
var predictionEngine = mlContext.Model.CreatePredictionEngine<TextData, TransformedTextData>(transformer);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

predictionEngine [](start = 16, length = 16)

Similar to the other PR, I wonder if we should stay entirely within IDataView and not create a prediction engine. That is, use a TakeRows filter followed by a CreateEnumerable.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is done because some of the transforms related to text processing such as (NormalizeText. TokenizeIntoWords etc.) don't need training data. In such cases, prediction engine seems more appropriate. But we can definitely have consensus on this. I will follow up.


In reply to: 271981536 [](ancestors = 271981536)

@rogancarrrogancarr left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved with comments.

🚴

// 0.5455 0.1818 0.2727
}

private static void PrintPredictions(TransformedTextData prediction)

@wschinwschinApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
privatestaticvoidPrintPredictions(TransformedTextDataprediction)
privatestaticvoidPrintLdaFeatures(TransformedTextDataprediction)
``` #Resolved

Console.WriteLine();
}

public class TextData

@wschinwschinApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
publicclass TextData
privateclass TextData
``` #Resolved

public string Text { get; set; }
}

public class TransformedTextData : TextData

@wschinwschinApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
publicclassTransformedTextData:TextData
privateclassTransformedTextData:TextData
``` #Resolved

@Ivanidzo4kaIvanidzo4ka left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@zeahmed
zeahmed merged commit a8915f4 into dotnet:masterApr 5, 2019
zeahmed added a commit to zeahmed/machinelearning that referenced this pull request Apr 8, 2019
@zeahmedzeahmed mentioned this pull request Apr 8, 2019
@ghostghost locked as resolved and limited conversation to collaborators Mar 23, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@zeahmed@Ivanidzo4ka@wschin@rogancarr
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Created sample for 'LatentDirichletAllocation' API. - #3191

Merged
zeahmed merged 5 commits into
dotnet:masterfrom
zeahmed:lda_sample
Apr 5, 2019
Merged

Created sample for 'LatentDirichletAllocation' API.#3191
zeahmed merged 5 commits into
dotnet:masterfrom
zeahmed:lda_sample

Conversation

@zeahmed

Copy link
Copy Markdown
Contributor

Related to #1209.

// before passing tokens to LatentDirichletAllocation.
var pipeline = mlContext.Transforms.Text.NormalizeText("normText", "Text")
.Append(mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "normText"))
.Append(mlContext.Transforms.Text.RemoveStopWords("Tokens"))

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

RemoveStopWords [](start = 50, length = 15)

Funny, this is custom stop words remover with no stop words.
So it does nothing.

I guess we need to remove param from params string[] stopwords) in RemoveStopWords #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

opss...That's why I was not getting the output what I was expecting...:)


In reply to: 271928104 [](ancestors = 271928104)

.Append(mlContext.Transforms.Text.TokenizeIntoWords("Tokens", "normText"))
.Append(mlContext.Transforms.Text.RemoveStopWords("Tokens"))
.Append(mlContext.Transforms.Conversion.MapValueToKey("Tokens"))
.Append(mlContext.Transforms.Text.ProduceNgrams("Tokens"))

@Ivanidzo4kaIvanidzo4kaApr 3, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ProduceNgrams [](start = 50, length = 13)

Do we actually want to run LDA on top of 2 ngrams since 2 is default value for ProduceNgrams or we should recommend to use ngrams:1 ? #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

2 is fine as it backs off to unigram ( useAllLengths=true) . I think higher is better in case there is a lot of data available.


In reply to: 271930804 [](ancestors = 271930804)

@codecov

codecovBot commented Apr 3, 2019

Copy link
Copy Markdown

Codecov Report

Merging #3191 into master will decrease coverage by 0.01%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #3191 +/- ##
==========================================
- Coverage 72.6% 72.58% -0.02% 
==========================================
Files 807 807 Lines 144957 144957 Branches 16211 16211 ==========================================
- Hits 105240 105215 -25 - Misses 35298 35326 +28 + Partials 4419 4416 -3
FlagCoverage Δ
#Debug72.58% <ø> (-0.02%)⬇️
#production68.14% <ø> (-0.03%)⬇️
#test88.88% <ø> (ø)⬆️
Impacted FilesCoverage Δ
src/Microsoft.ML.Transforms/Text/TextCatalog.cs41.66% <ø> (ø)⬆️
src/Microsoft.ML.Core/Data/ProgressReporter.cs70.95% <0%> (-6.99%)⬇️
src/Microsoft.ML.Maml/MAML.cs24.75% <0%> (-1.46%)⬇️
src/Microsoft.ML.Transforms/Text/LdaTransform.cs89.26% <0%> (-0.63%)⬇️
...ML.Transforms/Text/StopWordsRemovingTransformer.cs86.26% <0%> (+0.15%)⬆️

@codecov

codecovBot commented Apr 3, 2019

Copy link
Copy Markdown

Codecov Report

Merging #3191 into master will decrease coverage by 0.01%.
The diff coverage is n/a.

@@ Coverage Diff @@## master #3191 +/- ##
==========================================
- Coverage 72.6% 72.58% -0.02% 
==========================================
Files 807 807 Lines 144957 144957 Branches 16211 16211 ==========================================
- Hits 105240 105221 -19 - Misses 35298 35321 +23 + Partials 4419 4415 -4
FlagCoverage Δ
#Debug72.58% <ø> (-0.02%)⬇️
#production68.15% <ø> (-0.02%)⬇️
#test88.88% <ø> (ø)⬆️
Impacted FilesCoverage Δ
src/Microsoft.ML.Transforms/Text/TextCatalog.cs41.66% <ø> (ø)⬆️
src/Microsoft.ML.Core/Data/ProgressReporter.cs70.95% <0%> (-6.99%)⬇️

// Create a small dataset as an IEnumerable.
var samples = new List<TextData>()
{
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API computes topic model." },

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

topic model [](start = 88, length = 11)

models with an s #Resolved

{
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API computes topic model." },
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API is the best for topic model." },
new TextData(){ Text = "I like to eat broccoli and banana." },

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

banana [](start = 67, length = 6)

bananas #Resolved

new TextData(){ Text = "ML.NET's LatentDirichletAllocation API computes topic model." },
new TextData(){ Text = "ML.NET's LatentDirichletAllocation API is the best for topic model." },
new TextData(){ Text = "I like to eat broccoli and banana." },
new TextData(){ Text = "I eat a banana in the breakfast." },

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

in the [](start = 55, length = 6)

for #Resolved


// A pipeline for featurizing the text/string using LatentDirichletAllocation API.
// To be more accurate in computing the LDA features, the pipeline first normalizes text and removes stop words
// before passing tokens to LatentDirichletAllocation.

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tokens [](start = 30, length = 6)

"tokens (the individual words, lower cased, with common words removed)"

Many people won't be familiar with the specific language of NLP. #Resolved

// A pipeline for featurizing the text/string using LatentDirichletAllocation API.
// To be more accurate in computing the LDA features, the pipeline first normalizes text and removes stop words
// before passing tokens to LatentDirichletAllocation.
var pipeline = mlContext.Transforms.Text.NormalizeText("normText", "Text")

@rogancarrrogancarrApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

normText [](start = 68, length = 8)

I would spell it out for the example. #Resolved

var transformer = pipeline.Fit(dataview);

// Create the prediction engine to get the LDA features extracted from the text.
var predictionEngine = mlContext.Model.CreatePredictionEngine<TextData, TransformedTextData>(transformer);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

predictionEngine [](start = 16, length = 16)

Similar to the other PR, I wonder if we should stay entirely within IDataView and not create a prediction engine. That is, use a TakeRows filter followed by a CreateEnumerable.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is done because some of the transforms related to text processing such as (NormalizeText. TokenizeIntoWords etc.) don't need training data. In such cases, prediction engine seems more appropriate. But we can definitely have consensus on this. I will follow up.


In reply to: 271981536 [](ancestors = 271981536)

@rogancarrrogancarr left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved with comments.

🚴

// 0.5455 0.1818 0.2727
}

private static void PrintPredictions(TransformedTextData prediction)

@wschinwschinApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
privatestaticvoidPrintPredictions(TransformedTextDataprediction)
privatestaticvoidPrintLdaFeatures(TransformedTextDataprediction)
``` #Resolved

Console.WriteLine();
}

public class TextData

@wschinwschinApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
publicclass TextData
privateclass TextData
``` #Resolved

public string Text { get; set; }
}

public class TransformedTextData : TextData

@wschinwschinApr 4, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
publicclassTransformedTextData:TextData
privateclassTransformedTextData:TextData
``` #Resolved

@Ivanidzo4kaIvanidzo4ka left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@zeahmed
zeahmed merged commit a8915f4 into dotnet:masterApr 5, 2019
zeahmed added a commit to zeahmed/machinelearning that referenced this pull request Apr 8, 2019
@zeahmedzeahmed mentioned this pull request Apr 8, 2019
@ghostghost locked as resolved and limited conversation to collaborators Mar 23, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@zeahmed@Ivanidzo4ka@wschin@rogancarr