word embedding transform - #545

Merged
Ivanidzo4ka merged 29 commits into
dotnet:masterfrom
Ivanidzo4ka:ivanidze/wordembedding
Jul 31, 2018
Merged

word embedding transform#545
Ivanidzo4ka merged 29 commits into
dotnet:masterfrom
Ivanidzo4ka:ivanidze/wordembedding

Conversation

@Ivanidzo4ka

@Ivanidzo4kaIvanidzo4ka commented Jul 17, 2018

Copy link
Copy Markdown
Contributor

I heard word embedding can be nice thing for Text classification

  • Create issue
  • Put legal attributes for fastText files
  • Put legal attributes for GloVe files

(edited by @justinormont to fix model type names)
closes#615

@shauheen

Copy link
Copy Markdown
Contributor

Thanks @Ivanidzo4ka , can you please create an issue, and explain what is missing from ML.NET. 👍

if (string.IsNullOrWhiteSpace(_modelFileNameWithPath))
{
throw Host.Except("Model file for Word Embedding transform could not be found! " +
@"Please copy the model file '{0}' from '\\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVectors\' " +

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVectors' [](start = 61, length = 67)

what should this look like now?

if (_modelFileNameWithPath == null)
{
throw Host.Except("Model file for Word Embedding transform could not be found! " +
@"Please copy the model file '{0}' from '\\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVectors\' " +

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVect [](start = 61, length = 62)

and here


private static Dictionary<PretrainedModelKind, string> _modelsMetaData = new Dictionary<PretrainedModelKind, string>()
{
{ PretrainedModelKind.GloVe50D, "glove.6B.50d.txt" },

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

GloVe50D [](start = 35, length = 8)

do we have public locations for all of this?
Do we want to be the ones storing them?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we have aka.ms/tlc-resources/


In reply to: 204168101 [](ancestors = 204168101)

WordEmbeddings wrap different embedding models, such as GloVe. Users can specify which embedding to use.
The available options are various versions of <a href="https://nlp.stanford.edu/projects/glove/">GloVe Models</a>, <a href="https://en.wikipedia.org/wiki/FastText">FastText</a>, and <a href="http://anthology.aclweb.org/P/P14/P14-1146.pdf">Sswe</a>.
<para>
Note: As WordEmbedding requires a column with text vector, e.g. %3C'This', 'is', 'good'%3E, users need to create an input column by:

@sfilipisfilipiJul 20, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

' [](start = 85, length = 1)

apostrophes need be encoded too: ' #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hope %27 will work


In reply to: 204168630 [](ancestors = 204168630)

</item>
</list>
In the following example, after the NGramFeaturizer, features named ngram.__ are generated. A new column named ngram_TransformedText is
also created with the text vector, similar as running .split(' '). However, due to the variable length of this column it cannot be properly

@sfilipisfilipiJul 20, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

' ' [](start = 71, length = 3)

encode #Resolved

@sfilipi

sfilipi commented Jul 20, 2018

Copy link
Copy Markdown
Member
 pipeline.Add(new LightLda(("InTextCol" , "OutTextCol")));

bummer! #Resolved


Refers to: src/Microsoft.ML.Transforms/Text/doc.xml:182 in 68696f4. [](commit_id = 68696f4, deletion_comment = False)

<example name="WordEmbeddings">
<example>
<code language="csharp">
pipeline.Add(new WordEmbeddings(("InTextCol" , "OutTextCol")));

@sfilipisfilipiJul 20, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

" [](start = 43, length = 1)

encode those all. #Resolved

Ivan Matantsev added 2 commits July 20, 2018 14:29
remove links to internal storage
WordEmbedding. The output from WordEmbedding is named ngram_TransformedText.__
</para>
<para>
License attributes for pretrained models:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

License attributes for pretrained models: [](start = 9, length = 42)

@GalOshri is this wording looks ok for you?

</summary>
<remarks>
WordEmbeddings wrap different embedding models, such as GloVe. Users can specify which embedding to use.
The available options are various versions of <a href="https://nlp.stanford.edu/projects/glove/">GloVe Models</a>, <a href="https://en.wikipedia.org/wiki/FastText">FastText</a>, and <a href="http://anthology.aclweb.org/P/P14/P14-1146.pdf">Sswe</a>.

@justinormontjustinormontJul 20, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

fastText should be lower camel cased [1] #Resolved

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and SSWE is upper cased #Resolved

<para>
Note: As WordEmbedding requires a column with text vector, e.g. %3C%27This%27, %27is%27, %27good%27%3E, users need to create an input column by:
<list type="bullet">
<item><description>concatenating columns with TX type,</description></item>

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Concatenating columns of single unigrams is very unlikely.

We should be recommending only to use output_tokens=True in NGramFeaturizer(). All of our current pre-trained models require tokens which are lowercased unigrams w/ diacritics removed (which are the defaults for NGramFeaturizer). #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for clarification, didn't knew about this.
Will update documentation in later PR.


In reply to: 204193596 [](ancestors = 204193596)

</list>
In the following example, after the NGramFeaturizer, features named ngram.__ are generated. A new column named ngram_TransformedText is
also created with the text vector, similar as running .split(%27 %27). However, due to the variable length of this column it cannot be properly
converted to pandas dataframe, thus any pipelines/transforms output this text vector column will throw errors. However, we use

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"pandas dataframe" won't apply to the ML.NET #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The curse of copy paste!


In reply to: 204193641 [](ancestors = 204193641)

</item>
<item>
<description>
Glove models by Stanford University, or (Jeffrey Pennington, Richard Socher, Christopher D. Manning) is licensed under <a href="https://opendatacommons.org/licenses/pddl/1.0/">PDDL</a>.

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Glove should be capitalized as GloVe #Resolved

<item>
<description>
Glove models by Stanford University, or (Jeffrey Pennington, Richard Socher, Christopher D. Manning) is licensed under <a href="https://opendatacommons.org/licenses/pddl/1.0/">PDDL</a>.
More information can be found <a href="https://nlp.stanford.edu/projects/glove/">here</a>.

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Linking to their repo would also be nice: https://github.com/stanfordnlp/GloVe #Resolved

</item>
<item>
<description>
Glove models by Stanford University, or (Jeffrey Pennington, Richard Socher, Christopher D. Manning) is licensed under <a href="https://opendatacommons.org/licenses/pddl/1.0/">PDDL</a>.

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Recommend adding, Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014. [GloVe: Global Vectors for Word Representation](https://nlp.stanford.edu/pubs/glove.pdf)., which is their asked for citation format, as per: https://nlp.stanford.edu/projects/glove/
#Resolved

@Ivanidzo4ka

Copy link
Copy Markdown
ContributorAuthor

@dotnet-bot test OSX10.13 Release

@Ivanidzo4ka

Copy link
Copy Markdown
ContributorAuthor

@dotnet-bot test OSX10.13 Debug

int deno = 0;
srcGetter(ref src);
var values = dst.Values;
Utils.EnsureSize(ref values, 3 * dimension, keepOld: false);

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With keepOld: false, the Utils.EnsureSize() function always allocates a new array. Would it be faster to perform the size check and just zero the existing?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm actually not sure why we allocate new array all the time, (considering what we clean it immediately after) so I rather remove this "keepOld"


In reply to: 204564374 [](ancestors = 204564374)

@TomFinleyTomFinleyJul 27, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reactivating. Read implementation or documentation of EnsureSize a bit more carefully. THe keepOld: false certainly does not always allocate a new array. The difference is, in the situation where it is necessary to resize the array, whether that is created through Array.Resize or new... the latter is faster if we can get away with it, which we certainly can here.


In reply to: 204578950 [](ancestors = 204578950,204564374)

for (int i = 0; i < dimension; i++)
{
float currentTerm = wordVector[i];
if (values[i] > currentTerm)

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you point out the code fix for #545 (review)? I'm not seeing it. #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I removed Array.Clear and replace it with setting MaxValue,0,MinValue.


In reply to: 204566178 [](ancestors = 204566178)

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks. I see it now. #Resolved

int offset = 2 * dimension;
for (int i = 0; i < dimension; i++)
{
values[i] = int.MaxValue;

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Array is of type float, so float.MaxValue will be better #Resolved

@justinormontjustinormont left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

<para>
Note: As WordEmbedding requires a column with text vector, e.g. %3C%27this%27, %27is%27, %27good%27%3E, users need to create an input column by
using the output_tokens=True for TextTransform to convert a column with sentences like "This is good" into %3C%27this%27, %27is%27, %27good%27 %3E.
The column for the output token column is renamed original column with a prefix of %27_TranformedText%27.

@justinormontjustinormontJul 24, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can word smith this a bit..
The suffix of %27_TransformedText%27 is added to the original column name to create the output token column. For instance if the input column is %27body%27, the output tokens column is named %27body_TransformedText%27. #Resolved

if (model == null)
model = new Model(dimension);
if (model.Dimension != dimension)
ch.Warning($"Dimension mismatch while reading model file: '{_modelFileNameWithPath}', line number 1, expected dimension = {model.Dimension}, received dimension = {dimension}");

@justinormontjustinormontJul 24, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can remove this warning. The purpose of this block of code is to allow the 1st line to be of different length (and ignored if so). Hence the warning is superfluous. Currently this warning is displayed for all fastText models, even though the model is read correctly (by ignoring the 1st line).

Background: In fastText models, the 1st line is: . Other word embedding models don't have a header line. #Resolved

@Ivanidzo4ka
Ivanidzo4ka requested a review from Zruty0July 25, 2018 19:55
@justinormont

Copy link
Copy Markdown
Contributor

:shipit:

var name = Path.GetFileName(errorResult.FileName);
throw ch.Except($"{errorMessage}\nModel file for Word Embedding transform could not be found! " +
$@"Please copy the model file '{name}' from '{url}' to '{directory}'.");
}

@TomFinleyTomFinleyJul 27, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The user story here doesn't seem great. As in, I'm not sure this transform is actually usable at all. This ResourceManagerUtils points people to the problematically named "https://aka.ms/tlc-resources/". From an outside user's perspective, does that make it useless?

This transform is clearly written with the expectation of a console application not a library. If it were a library, we might expect these sorts of "optional things" would be brought in via nuget dependencies... that is, if you want to use this or that, you subscribe to the appropriate nuget, then in the Arguments of this thing assign the relevant resource as an actual object published by that nuget. (It would be some variety of IComponentFactory.)

I feel like this whole approach needs some deeper thought.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

problematically named resource has it's own issue #546

Model files which we use (except SSWE which is a 70mb) in range from 160MB(glove50d) to 6 GB(fastText). Which I think impossible to fit into any nuget due to size limitation.


In reply to: 205848439 [](ancestors = 205848439)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh boy. That's pretty big. Hmmm... hmmm... this is pretty complicated then. All right let's punt on that for now.

@TomFinleyTomFinley left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you @Ivanidzo4ka ! But seriously, could you create an issue?

@Ivanidzo4ka
Ivanidzo4ka merged commit b727d10 into dotnet:masterJul 31, 2018
codemzs pushed a commit to codemzs/machinelearning that referenced this pull request Aug 1, 2018
Introduce word embedding transform
@ghostghost locked as resolved and limited conversation to collaborators Mar 29, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Word embeddings

5 participants

@Ivanidzo4ka@shauheen@sfilipi@justinormont@TomFinley
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

word embedding transform - #545

Merged
Ivanidzo4ka merged 29 commits into
dotnet:masterfrom
Ivanidzo4ka:ivanidze/wordembedding
Jul 31, 2018
Merged

word embedding transform#545
Ivanidzo4ka merged 29 commits into
dotnet:masterfrom
Ivanidzo4ka:ivanidze/wordembedding

Conversation

@Ivanidzo4ka

@Ivanidzo4kaIvanidzo4ka commented Jul 17, 2018

Copy link
Copy Markdown
Contributor

I heard word embedding can be nice thing for Text classification

  • Create issue
  • Put legal attributes for fastText files
  • Put legal attributes for GloVe files

(edited by @justinormont to fix model type names)
closes#615

@shauheen

Copy link
Copy Markdown
Contributor

Thanks @Ivanidzo4ka , can you please create an issue, and explain what is missing from ML.NET. 👍

if (string.IsNullOrWhiteSpace(_modelFileNameWithPath))
{
throw Host.Except("Model file for Word Embedding transform could not be found! " +
@"Please copy the model file '{0}' from '\\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVectors\' " +

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVectors' [](start = 61, length = 67)

what should this look like now?

if (_modelFileNameWithPath == null)
{
throw Host.Except("Model file for Word Embedding transform could not be found! " +
@"Please copy the model file '{0}' from '\\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVectors\' " +

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVect [](start = 61, length = 62)

and here


private static Dictionary<PretrainedModelKind, string> _modelsMetaData = new Dictionary<PretrainedModelKind, string>()
{
{ PretrainedModelKind.GloVe50D, "glove.6B.50d.txt" },

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

GloVe50D [](start = 35, length = 8)

do we have public locations for all of this?
Do we want to be the ones storing them?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we have aka.ms/tlc-resources/


In reply to: 204168101 [](ancestors = 204168101)

WordEmbeddings wrap different embedding models, such as GloVe. Users can specify which embedding to use.
The available options are various versions of <a href="https://nlp.stanford.edu/projects/glove/">GloVe Models</a>, <a href="https://en.wikipedia.org/wiki/FastText">FastText</a>, and <a href="http://anthology.aclweb.org/P/P14/P14-1146.pdf">Sswe</a>.
<para>
Note: As WordEmbedding requires a column with text vector, e.g. %3C'This', 'is', 'good'%3E, users need to create an input column by:

@sfilipisfilipiJul 20, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

' [](start = 85, length = 1)

apostrophes need be encoded too: ' #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hope %27 will work


In reply to: 204168630 [](ancestors = 204168630)

</item>
</list>
In the following example, after the NGramFeaturizer, features named ngram.__ are generated. A new column named ngram_TransformedText is
also created with the text vector, similar as running .split(' '). However, due to the variable length of this column it cannot be properly

@sfilipisfilipiJul 20, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

' ' [](start = 71, length = 3)

encode #Resolved

@sfilipi

sfilipi commented Jul 20, 2018

Copy link
Copy Markdown
Member
 pipeline.Add(new LightLda(("InTextCol" , "OutTextCol")));

bummer! #Resolved


Refers to: src/Microsoft.ML.Transforms/Text/doc.xml:182 in 68696f4. [](commit_id = 68696f4, deletion_comment = False)

<example name="WordEmbeddings">
<example>
<code language="csharp">
pipeline.Add(new WordEmbeddings(("InTextCol" , "OutTextCol")));

@sfilipisfilipiJul 20, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

" [](start = 43, length = 1)

encode those all. #Resolved

Ivan Matantsev added 2 commits July 20, 2018 14:29
remove links to internal storage
WordEmbedding. The output from WordEmbedding is named ngram_TransformedText.__
</para>
<para>
License attributes for pretrained models:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

License attributes for pretrained models: [](start = 9, length = 42)

@GalOshri is this wording looks ok for you?

</summary>
<remarks>
WordEmbeddings wrap different embedding models, such as GloVe. Users can specify which embedding to use.
The available options are various versions of <a href="https://nlp.stanford.edu/projects/glove/">GloVe Models</a>, <a href="https://en.wikipedia.org/wiki/FastText">FastText</a>, and <a href="http://anthology.aclweb.org/P/P14/P14-1146.pdf">Sswe</a>.

@justinormontjustinormontJul 20, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

fastText should be lower camel cased [1] #Resolved

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and SSWE is upper cased #Resolved

<para>
Note: As WordEmbedding requires a column with text vector, e.g. %3C%27This%27, %27is%27, %27good%27%3E, users need to create an input column by:
<list type="bullet">
<item><description>concatenating columns with TX type,</description></item>

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Concatenating columns of single unigrams is very unlikely.

We should be recommending only to use output_tokens=True in NGramFeaturizer(). All of our current pre-trained models require tokens which are lowercased unigrams w/ diacritics removed (which are the defaults for NGramFeaturizer). #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for clarification, didn't knew about this.
Will update documentation in later PR.


In reply to: 204193596 [](ancestors = 204193596)

</list>
In the following example, after the NGramFeaturizer, features named ngram.__ are generated. A new column named ngram_TransformedText is
also created with the text vector, similar as running .split(%27 %27). However, due to the variable length of this column it cannot be properly
converted to pandas dataframe, thus any pipelines/transforms output this text vector column will throw errors. However, we use

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"pandas dataframe" won't apply to the ML.NET #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The curse of copy paste!


In reply to: 204193641 [](ancestors = 204193641)

</item>
<item>
<description>
Glove models by Stanford University, or (Jeffrey Pennington, Richard Socher, Christopher D. Manning) is licensed under <a href="https://opendatacommons.org/licenses/pddl/1.0/">PDDL</a>.

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Glove should be capitalized as GloVe #Resolved

<item>
<description>
Glove models by Stanford University, or (Jeffrey Pennington, Richard Socher, Christopher D. Manning) is licensed under <a href="https://opendatacommons.org/licenses/pddl/1.0/">PDDL</a>.
More information can be found <a href="https://nlp.stanford.edu/projects/glove/">here</a>.

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Linking to their repo would also be nice: https://github.com/stanfordnlp/GloVe #Resolved

</item>
<item>
<description>
Glove models by Stanford University, or (Jeffrey Pennington, Richard Socher, Christopher D. Manning) is licensed under <a href="https://opendatacommons.org/licenses/pddl/1.0/">PDDL</a>.

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Recommend adding, Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014. [GloVe: Global Vectors for Word Representation](https://nlp.stanford.edu/pubs/glove.pdf)., which is their asked for citation format, as per: https://nlp.stanford.edu/projects/glove/
#Resolved

@Ivanidzo4ka

Copy link
Copy Markdown
ContributorAuthor

@dotnet-bot test OSX10.13 Release

@Ivanidzo4ka

Copy link
Copy Markdown
ContributorAuthor

@dotnet-bot test OSX10.13 Debug

int deno = 0;
srcGetter(ref src);
var values = dst.Values;
Utils.EnsureSize(ref values, 3 * dimension, keepOld: false);

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With keepOld: false, the Utils.EnsureSize() function always allocates a new array. Would it be faster to perform the size check and just zero the existing?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm actually not sure why we allocate new array all the time, (considering what we clean it immediately after) so I rather remove this "keepOld"


In reply to: 204564374 [](ancestors = 204564374)

@TomFinleyTomFinleyJul 27, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reactivating. Read implementation or documentation of EnsureSize a bit more carefully. THe keepOld: false certainly does not always allocate a new array. The difference is, in the situation where it is necessary to resize the array, whether that is created through Array.Resize or new... the latter is faster if we can get away with it, which we certainly can here.


In reply to: 204578950 [](ancestors = 204578950,204564374)

for (int i = 0; i < dimension; i++)
{
float currentTerm = wordVector[i];
if (values[i] > currentTerm)

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you point out the code fix for #545 (review)? I'm not seeing it. #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I removed Array.Clear and replace it with setting MaxValue,0,MinValue.


In reply to: 204566178 [](ancestors = 204566178)

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks. I see it now. #Resolved

int offset = 2 * dimension;
for (int i = 0; i < dimension; i++)
{
values[i] = int.MaxValue;

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Array is of type float, so float.MaxValue will be better #Resolved

@justinormontjustinormont left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

<para>
Note: As WordEmbedding requires a column with text vector, e.g. %3C%27this%27, %27is%27, %27good%27%3E, users need to create an input column by
using the output_tokens=True for TextTransform to convert a column with sentences like "This is good" into %3C%27this%27, %27is%27, %27good%27 %3E.
The column for the output token column is renamed original column with a prefix of %27_TranformedText%27.

@justinormontjustinormontJul 24, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can word smith this a bit..
The suffix of %27_TransformedText%27 is added to the original column name to create the output token column. For instance if the input column is %27body%27, the output tokens column is named %27body_TransformedText%27. #Resolved

if (model == null)
model = new Model(dimension);
if (model.Dimension != dimension)
ch.Warning($"Dimension mismatch while reading model file: '{_modelFileNameWithPath}', line number 1, expected dimension = {model.Dimension}, received dimension = {dimension}");

@justinormontjustinormontJul 24, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can remove this warning. The purpose of this block of code is to allow the 1st line to be of different length (and ignored if so). Hence the warning is superfluous. Currently this warning is displayed for all fastText models, even though the model is read correctly (by ignoring the 1st line).

Background: In fastText models, the 1st line is: . Other word embedding models don't have a header line. #Resolved

@Ivanidzo4ka
Ivanidzo4ka requested a review from Zruty0July 25, 2018 19:55
@justinormont

Copy link
Copy Markdown
Contributor

:shipit:

var name = Path.GetFileName(errorResult.FileName);
throw ch.Except($"{errorMessage}\nModel file for Word Embedding transform could not be found! " +
$@"Please copy the model file '{name}' from '{url}' to '{directory}'.");
}

@TomFinleyTomFinleyJul 27, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The user story here doesn't seem great. As in, I'm not sure this transform is actually usable at all. This ResourceManagerUtils points people to the problematically named "https://aka.ms/tlc-resources/". From an outside user's perspective, does that make it useless?

This transform is clearly written with the expectation of a console application not a library. If it were a library, we might expect these sorts of "optional things" would be brought in via nuget dependencies... that is, if you want to use this or that, you subscribe to the appropriate nuget, then in the Arguments of this thing assign the relevant resource as an actual object published by that nuget. (It would be some variety of IComponentFactory.)

I feel like this whole approach needs some deeper thought.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

problematically named resource has it's own issue #546

Model files which we use (except SSWE which is a 70mb) in range from 160MB(glove50d) to 6 GB(fastText). Which I think impossible to fit into any nuget due to size limitation.


In reply to: 205848439 [](ancestors = 205848439)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh boy. That's pretty big. Hmmm... hmmm... this is pretty complicated then. All right let's punt on that for now.

@TomFinleyTomFinley left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you @Ivanidzo4ka ! But seriously, could you create an issue?

@Ivanidzo4ka
Ivanidzo4ka merged commit b727d10 into dotnet:masterJul 31, 2018
codemzs pushed a commit to codemzs/machinelearning that referenced this pull request Aug 1, 2018
Introduce word embedding transform
@ghostghost locked as resolved and limited conversation to collaborators Mar 29, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Word embeddings

5 participants

@Ivanidzo4ka@shauheen@sfilipi@justinormont@TomFinley
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

word embedding transform - #545

Merged
Ivanidzo4ka merged 29 commits into
dotnet:masterfrom
Ivanidzo4ka:ivanidze/wordembedding
Jul 31, 2018
Merged

word embedding transform#545
Ivanidzo4ka merged 29 commits into
dotnet:masterfrom
Ivanidzo4ka:ivanidze/wordembedding

Conversation

@Ivanidzo4ka

@Ivanidzo4kaIvanidzo4ka commented Jul 17, 2018

Copy link
Copy Markdown
Contributor

I heard word embedding can be nice thing for Text classification

  • Create issue
  • Put legal attributes for fastText files
  • Put legal attributes for GloVe files

(edited by @justinormont to fix model type names)
closes#615

@shauheen

Copy link
Copy Markdown
Contributor

Thanks @Ivanidzo4ka , can you please create an issue, and explain what is missing from ML.NET. 👍

if (string.IsNullOrWhiteSpace(_modelFileNameWithPath))
{
throw Host.Except("Model file for Word Embedding transform could not be found! " +
@"Please copy the model file '{0}' from '\\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVectors\' " +

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVectors' [](start = 61, length = 67)

what should this look like now?

if (_modelFileNameWithPath == null)
{
throw Host.Except("Model file for Word Embedding transform could not be found! " +
@"Please copy the model file '{0}' from '\\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVectors\' " +

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVect [](start = 61, length = 62)

and here


private static Dictionary<PretrainedModelKind, string> _modelsMetaData = new Dictionary<PretrainedModelKind, string>()
{
{ PretrainedModelKind.GloVe50D, "glove.6B.50d.txt" },

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

GloVe50D [](start = 35, length = 8)

do we have public locations for all of this?
Do we want to be the ones storing them?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we have aka.ms/tlc-resources/


In reply to: 204168101 [](ancestors = 204168101)

WordEmbeddings wrap different embedding models, such as GloVe. Users can specify which embedding to use.
The available options are various versions of <a href="https://nlp.stanford.edu/projects/glove/">GloVe Models</a>, <a href="https://en.wikipedia.org/wiki/FastText">FastText</a>, and <a href="http://anthology.aclweb.org/P/P14/P14-1146.pdf">Sswe</a>.
<para>
Note: As WordEmbedding requires a column with text vector, e.g. %3C'This', 'is', 'good'%3E, users need to create an input column by:

@sfilipisfilipiJul 20, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

' [](start = 85, length = 1)

apostrophes need be encoded too: ' #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hope %27 will work


In reply to: 204168630 [](ancestors = 204168630)

</item>
</list>
In the following example, after the NGramFeaturizer, features named ngram.__ are generated. A new column named ngram_TransformedText is
also created with the text vector, similar as running .split(' '). However, due to the variable length of this column it cannot be properly

@sfilipisfilipiJul 20, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

' ' [](start = 71, length = 3)

encode #Resolved

@sfilipi

sfilipi commented Jul 20, 2018

Copy link
Copy Markdown
Member
 pipeline.Add(new LightLda(("InTextCol" , "OutTextCol")));

bummer! #Resolved


Refers to: src/Microsoft.ML.Transforms/Text/doc.xml:182 in 68696f4. [](commit_id = 68696f4, deletion_comment = False)

<example name="WordEmbeddings">
<example>
<code language="csharp">
pipeline.Add(new WordEmbeddings(("InTextCol" , "OutTextCol")));

@sfilipisfilipiJul 20, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

" [](start = 43, length = 1)

encode those all. #Resolved

Ivan Matantsev added 2 commits July 20, 2018 14:29
remove links to internal storage
WordEmbedding. The output from WordEmbedding is named ngram_TransformedText.__
</para>
<para>
License attributes for pretrained models:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

License attributes for pretrained models: [](start = 9, length = 42)

@GalOshri is this wording looks ok for you?

</summary>
<remarks>
WordEmbeddings wrap different embedding models, such as GloVe. Users can specify which embedding to use.
The available options are various versions of <a href="https://nlp.stanford.edu/projects/glove/">GloVe Models</a>, <a href="https://en.wikipedia.org/wiki/FastText">FastText</a>, and <a href="http://anthology.aclweb.org/P/P14/P14-1146.pdf">Sswe</a>.

@justinormontjustinormontJul 20, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

fastText should be lower camel cased [1] #Resolved

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and SSWE is upper cased #Resolved

<para>
Note: As WordEmbedding requires a column with text vector, e.g. %3C%27This%27, %27is%27, %27good%27%3E, users need to create an input column by:
<list type="bullet">
<item><description>concatenating columns with TX type,</description></item>

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Concatenating columns of single unigrams is very unlikely.

We should be recommending only to use output_tokens=True in NGramFeaturizer(). All of our current pre-trained models require tokens which are lowercased unigrams w/ diacritics removed (which are the defaults for NGramFeaturizer). #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for clarification, didn't knew about this.
Will update documentation in later PR.


In reply to: 204193596 [](ancestors = 204193596)

</list>
In the following example, after the NGramFeaturizer, features named ngram.__ are generated. A new column named ngram_TransformedText is
also created with the text vector, similar as running .split(%27 %27). However, due to the variable length of this column it cannot be properly
converted to pandas dataframe, thus any pipelines/transforms output this text vector column will throw errors. However, we use

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"pandas dataframe" won't apply to the ML.NET #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The curse of copy paste!


In reply to: 204193641 [](ancestors = 204193641)

</item>
<item>
<description>
Glove models by Stanford University, or (Jeffrey Pennington, Richard Socher, Christopher D. Manning) is licensed under <a href="https://opendatacommons.org/licenses/pddl/1.0/">PDDL</a>.

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Glove should be capitalized as GloVe #Resolved

<item>
<description>
Glove models by Stanford University, or (Jeffrey Pennington, Richard Socher, Christopher D. Manning) is licensed under <a href="https://opendatacommons.org/licenses/pddl/1.0/">PDDL</a>.
More information can be found <a href="https://nlp.stanford.edu/projects/glove/">here</a>.

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Linking to their repo would also be nice: https://github.com/stanfordnlp/GloVe #Resolved

</item>
<item>
<description>
Glove models by Stanford University, or (Jeffrey Pennington, Richard Socher, Christopher D. Manning) is licensed under <a href="https://opendatacommons.org/licenses/pddl/1.0/">PDDL</a>.

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Recommend adding, Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014. [GloVe: Global Vectors for Word Representation](https://nlp.stanford.edu/pubs/glove.pdf)., which is their asked for citation format, as per: https://nlp.stanford.edu/projects/glove/
#Resolved

@Ivanidzo4ka

Copy link
Copy Markdown
ContributorAuthor

@dotnet-bot test OSX10.13 Release

@Ivanidzo4ka

Copy link
Copy Markdown
ContributorAuthor

@dotnet-bot test OSX10.13 Debug

int deno = 0;
srcGetter(ref src);
var values = dst.Values;
Utils.EnsureSize(ref values, 3 * dimension, keepOld: false);

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With keepOld: false, the Utils.EnsureSize() function always allocates a new array. Would it be faster to perform the size check and just zero the existing?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm actually not sure why we allocate new array all the time, (considering what we clean it immediately after) so I rather remove this "keepOld"


In reply to: 204564374 [](ancestors = 204564374)

@TomFinleyTomFinleyJul 27, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reactivating. Read implementation or documentation of EnsureSize a bit more carefully. THe keepOld: false certainly does not always allocate a new array. The difference is, in the situation where it is necessary to resize the array, whether that is created through Array.Resize or new... the latter is faster if we can get away with it, which we certainly can here.


In reply to: 204578950 [](ancestors = 204578950,204564374)

for (int i = 0; i < dimension; i++)
{
float currentTerm = wordVector[i];
if (values[i] > currentTerm)

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you point out the code fix for #545 (review)? I'm not seeing it. #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I removed Array.Clear and replace it with setting MaxValue,0,MinValue.


In reply to: 204566178 [](ancestors = 204566178)

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks. I see it now. #Resolved

int offset = 2 * dimension;
for (int i = 0; i < dimension; i++)
{
values[i] = int.MaxValue;

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Array is of type float, so float.MaxValue will be better #Resolved

@justinormontjustinormont left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

<para>
Note: As WordEmbedding requires a column with text vector, e.g. %3C%27this%27, %27is%27, %27good%27%3E, users need to create an input column by
using the output_tokens=True for TextTransform to convert a column with sentences like "This is good" into %3C%27this%27, %27is%27, %27good%27 %3E.
The column for the output token column is renamed original column with a prefix of %27_TranformedText%27.

@justinormontjustinormontJul 24, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can word smith this a bit..
The suffix of %27_TransformedText%27 is added to the original column name to create the output token column. For instance if the input column is %27body%27, the output tokens column is named %27body_TransformedText%27. #Resolved

if (model == null)
model = new Model(dimension);
if (model.Dimension != dimension)
ch.Warning($"Dimension mismatch while reading model file: '{_modelFileNameWithPath}', line number 1, expected dimension = {model.Dimension}, received dimension = {dimension}");

@justinormontjustinormontJul 24, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can remove this warning. The purpose of this block of code is to allow the 1st line to be of different length (and ignored if so). Hence the warning is superfluous. Currently this warning is displayed for all fastText models, even though the model is read correctly (by ignoring the 1st line).

Background: In fastText models, the 1st line is: . Other word embedding models don't have a header line. #Resolved

@Ivanidzo4ka
Ivanidzo4ka requested a review from Zruty0July 25, 2018 19:55
@justinormont

Copy link
Copy Markdown
Contributor

:shipit:

var name = Path.GetFileName(errorResult.FileName);
throw ch.Except($"{errorMessage}\nModel file for Word Embedding transform could not be found! " +
$@"Please copy the model file '{name}' from '{url}' to '{directory}'.");
}

@TomFinleyTomFinleyJul 27, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The user story here doesn't seem great. As in, I'm not sure this transform is actually usable at all. This ResourceManagerUtils points people to the problematically named "https://aka.ms/tlc-resources/". From an outside user's perspective, does that make it useless?

This transform is clearly written with the expectation of a console application not a library. If it were a library, we might expect these sorts of "optional things" would be brought in via nuget dependencies... that is, if you want to use this or that, you subscribe to the appropriate nuget, then in the Arguments of this thing assign the relevant resource as an actual object published by that nuget. (It would be some variety of IComponentFactory.)

I feel like this whole approach needs some deeper thought.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

problematically named resource has it's own issue #546

Model files which we use (except SSWE which is a 70mb) in range from 160MB(glove50d) to 6 GB(fastText). Which I think impossible to fit into any nuget due to size limitation.


In reply to: 205848439 [](ancestors = 205848439)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh boy. That's pretty big. Hmmm... hmmm... this is pretty complicated then. All right let's punt on that for now.

@TomFinleyTomFinley left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you @Ivanidzo4ka ! But seriously, could you create an issue?

@Ivanidzo4ka
Ivanidzo4ka merged commit b727d10 into dotnet:masterJul 31, 2018
codemzs pushed a commit to codemzs/machinelearning that referenced this pull request Aug 1, 2018
Introduce word embedding transform
@ghostghost locked as resolved and limited conversation to collaborators Mar 29, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Word embeddings

5 participants

@Ivanidzo4ka@shauheen@sfilipi@justinormont@TomFinley
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

word embedding transform - #545

Merged
Ivanidzo4ka merged 29 commits into
dotnet:masterfrom
Ivanidzo4ka:ivanidze/wordembedding
Jul 31, 2018
Merged

word embedding transform#545
Ivanidzo4ka merged 29 commits into
dotnet:masterfrom
Ivanidzo4ka:ivanidze/wordembedding

Conversation

@Ivanidzo4ka

@Ivanidzo4kaIvanidzo4ka commented Jul 17, 2018

Copy link
Copy Markdown
Contributor

I heard word embedding can be nice thing for Text classification

  • Create issue
  • Put legal attributes for fastText files
  • Put legal attributes for GloVe files

(edited by @justinormont to fix model type names)
closes#615

@shauheen

Copy link
Copy Markdown
Contributor

Thanks @Ivanidzo4ka , can you please create an issue, and explain what is missing from ML.NET. 👍

if (string.IsNullOrWhiteSpace(_modelFileNameWithPath))
{
throw Host.Except("Model file for Word Embedding transform could not be found! " +
@"Please copy the model file '{0}' from '\\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVectors\' " +

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVectors' [](start = 61, length = 67)

what should this look like now?

if (_modelFileNameWithPath == null)
{
throw Host.Except("Model file for Word Embedding transform could not be found! " +
@"Please copy the model file '{0}' from '\\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVectors\' " +

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVect [](start = 61, length = 62)

and here


private static Dictionary<PretrainedModelKind, string> _modelsMetaData = new Dictionary<PretrainedModelKind, string>()
{
{ PretrainedModelKind.GloVe50D, "glove.6B.50d.txt" },

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

GloVe50D [](start = 35, length = 8)

do we have public locations for all of this?
Do we want to be the ones storing them?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we have aka.ms/tlc-resources/


In reply to: 204168101 [](ancestors = 204168101)

WordEmbeddings wrap different embedding models, such as GloVe. Users can specify which embedding to use.
The available options are various versions of <a href="https://nlp.stanford.edu/projects/glove/">GloVe Models</a>, <a href="https://en.wikipedia.org/wiki/FastText">FastText</a>, and <a href="http://anthology.aclweb.org/P/P14/P14-1146.pdf">Sswe</a>.
<para>
Note: As WordEmbedding requires a column with text vector, e.g. %3C'This', 'is', 'good'%3E, users need to create an input column by:

@sfilipisfilipiJul 20, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

' [](start = 85, length = 1)

apostrophes need be encoded too: ' #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hope %27 will work


In reply to: 204168630 [](ancestors = 204168630)

</item>
</list>
In the following example, after the NGramFeaturizer, features named ngram.__ are generated. A new column named ngram_TransformedText is
also created with the text vector, similar as running .split(' '). However, due to the variable length of this column it cannot be properly

@sfilipisfilipiJul 20, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

' ' [](start = 71, length = 3)

encode #Resolved

@sfilipi

sfilipi commented Jul 20, 2018

Copy link
Copy Markdown
Member
 pipeline.Add(new LightLda(("InTextCol" , "OutTextCol")));

bummer! #Resolved


Refers to: src/Microsoft.ML.Transforms/Text/doc.xml:182 in 68696f4. [](commit_id = 68696f4, deletion_comment = False)

<example name="WordEmbeddings">
<example>
<code language="csharp">
pipeline.Add(new WordEmbeddings(("InTextCol" , "OutTextCol")));

@sfilipisfilipiJul 20, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

" [](start = 43, length = 1)

encode those all. #Resolved

Ivan Matantsev added 2 commits July 20, 2018 14:29
remove links to internal storage
WordEmbedding. The output from WordEmbedding is named ngram_TransformedText.__
</para>
<para>
License attributes for pretrained models:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

License attributes for pretrained models: [](start = 9, length = 42)

@GalOshri is this wording looks ok for you?

</summary>
<remarks>
WordEmbeddings wrap different embedding models, such as GloVe. Users can specify which embedding to use.
The available options are various versions of <a href="https://nlp.stanford.edu/projects/glove/">GloVe Models</a>, <a href="https://en.wikipedia.org/wiki/FastText">FastText</a>, and <a href="http://anthology.aclweb.org/P/P14/P14-1146.pdf">Sswe</a>.

@justinormontjustinormontJul 20, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

fastText should be lower camel cased [1] #Resolved

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and SSWE is upper cased #Resolved

<para>
Note: As WordEmbedding requires a column with text vector, e.g. %3C%27This%27, %27is%27, %27good%27%3E, users need to create an input column by:
<list type="bullet">
<item><description>concatenating columns with TX type,</description></item>

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Concatenating columns of single unigrams is very unlikely.

We should be recommending only to use output_tokens=True in NGramFeaturizer(). All of our current pre-trained models require tokens which are lowercased unigrams w/ diacritics removed (which are the defaults for NGramFeaturizer). #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for clarification, didn't knew about this.
Will update documentation in later PR.


In reply to: 204193596 [](ancestors = 204193596)

</list>
In the following example, after the NGramFeaturizer, features named ngram.__ are generated. A new column named ngram_TransformedText is
also created with the text vector, similar as running .split(%27 %27). However, due to the variable length of this column it cannot be properly
converted to pandas dataframe, thus any pipelines/transforms output this text vector column will throw errors. However, we use

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"pandas dataframe" won't apply to the ML.NET #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The curse of copy paste!


In reply to: 204193641 [](ancestors = 204193641)

</item>
<item>
<description>
Glove models by Stanford University, or (Jeffrey Pennington, Richard Socher, Christopher D. Manning) is licensed under <a href="https://opendatacommons.org/licenses/pddl/1.0/">PDDL</a>.

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Glove should be capitalized as GloVe #Resolved

<item>
<description>
Glove models by Stanford University, or (Jeffrey Pennington, Richard Socher, Christopher D. Manning) is licensed under <a href="https://opendatacommons.org/licenses/pddl/1.0/">PDDL</a>.
More information can be found <a href="https://nlp.stanford.edu/projects/glove/">here</a>.

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Linking to their repo would also be nice: https://github.com/stanfordnlp/GloVe #Resolved

</item>
<item>
<description>
Glove models by Stanford University, or (Jeffrey Pennington, Richard Socher, Christopher D. Manning) is licensed under <a href="https://opendatacommons.org/licenses/pddl/1.0/">PDDL</a>.

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Recommend adding, Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014. [GloVe: Global Vectors for Word Representation](https://nlp.stanford.edu/pubs/glove.pdf)., which is their asked for citation format, as per: https://nlp.stanford.edu/projects/glove/
#Resolved

@Ivanidzo4ka

Copy link
Copy Markdown
ContributorAuthor

@dotnet-bot test OSX10.13 Release

@Ivanidzo4ka

Copy link
Copy Markdown
ContributorAuthor

@dotnet-bot test OSX10.13 Debug

int deno = 0;
srcGetter(ref src);
var values = dst.Values;
Utils.EnsureSize(ref values, 3 * dimension, keepOld: false);

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With keepOld: false, the Utils.EnsureSize() function always allocates a new array. Would it be faster to perform the size check and just zero the existing?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm actually not sure why we allocate new array all the time, (considering what we clean it immediately after) so I rather remove this "keepOld"


In reply to: 204564374 [](ancestors = 204564374)

@TomFinleyTomFinleyJul 27, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reactivating. Read implementation or documentation of EnsureSize a bit more carefully. THe keepOld: false certainly does not always allocate a new array. The difference is, in the situation where it is necessary to resize the array, whether that is created through Array.Resize or new... the latter is faster if we can get away with it, which we certainly can here.


In reply to: 204578950 [](ancestors = 204578950,204564374)

for (int i = 0; i < dimension; i++)
{
float currentTerm = wordVector[i];
if (values[i] > currentTerm)

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you point out the code fix for #545 (review)? I'm not seeing it. #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I removed Array.Clear and replace it with setting MaxValue,0,MinValue.


In reply to: 204566178 [](ancestors = 204566178)

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks. I see it now. #Resolved

int offset = 2 * dimension;
for (int i = 0; i < dimension; i++)
{
values[i] = int.MaxValue;

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Array is of type float, so float.MaxValue will be better #Resolved

@justinormontjustinormont left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

<para>
Note: As WordEmbedding requires a column with text vector, e.g. %3C%27this%27, %27is%27, %27good%27%3E, users need to create an input column by
using the output_tokens=True for TextTransform to convert a column with sentences like "This is good" into %3C%27this%27, %27is%27, %27good%27 %3E.
The column for the output token column is renamed original column with a prefix of %27_TranformedText%27.

@justinormontjustinormontJul 24, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can word smith this a bit..
The suffix of %27_TransformedText%27 is added to the original column name to create the output token column. For instance if the input column is %27body%27, the output tokens column is named %27body_TransformedText%27. #Resolved

if (model == null)
model = new Model(dimension);
if (model.Dimension != dimension)
ch.Warning($"Dimension mismatch while reading model file: '{_modelFileNameWithPath}', line number 1, expected dimension = {model.Dimension}, received dimension = {dimension}");

@justinormontjustinormontJul 24, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can remove this warning. The purpose of this block of code is to allow the 1st line to be of different length (and ignored if so). Hence the warning is superfluous. Currently this warning is displayed for all fastText models, even though the model is read correctly (by ignoring the 1st line).

Background: In fastText models, the 1st line is: . Other word embedding models don't have a header line. #Resolved

@Ivanidzo4ka
Ivanidzo4ka requested a review from Zruty0July 25, 2018 19:55
@justinormont

Copy link
Copy Markdown
Contributor

:shipit:

var name = Path.GetFileName(errorResult.FileName);
throw ch.Except($"{errorMessage}\nModel file for Word Embedding transform could not be found! " +
$@"Please copy the model file '{name}' from '{url}' to '{directory}'.");
}

@TomFinleyTomFinleyJul 27, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The user story here doesn't seem great. As in, I'm not sure this transform is actually usable at all. This ResourceManagerUtils points people to the problematically named "https://aka.ms/tlc-resources/". From an outside user's perspective, does that make it useless?

This transform is clearly written with the expectation of a console application not a library. If it were a library, we might expect these sorts of "optional things" would be brought in via nuget dependencies... that is, if you want to use this or that, you subscribe to the appropriate nuget, then in the Arguments of this thing assign the relevant resource as an actual object published by that nuget. (It would be some variety of IComponentFactory.)

I feel like this whole approach needs some deeper thought.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

problematically named resource has it's own issue #546

Model files which we use (except SSWE which is a 70mb) in range from 160MB(glove50d) to 6 GB(fastText). Which I think impossible to fit into any nuget due to size limitation.


In reply to: 205848439 [](ancestors = 205848439)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh boy. That's pretty big. Hmmm... hmmm... this is pretty complicated then. All right let's punt on that for now.

@TomFinleyTomFinley left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you @Ivanidzo4ka ! But seriously, could you create an issue?

@Ivanidzo4ka
Ivanidzo4ka merged commit b727d10 into dotnet:masterJul 31, 2018
codemzs pushed a commit to codemzs/machinelearning that referenced this pull request Aug 1, 2018
Introduce word embedding transform
@ghostghost locked as resolved and limited conversation to collaborators Mar 29, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Word embeddings

5 participants

@Ivanidzo4ka@shauheen@sfilipi@justinormont@TomFinley
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

word embedding transform - #545

Merged
Ivanidzo4ka merged 29 commits into
dotnet:masterfrom
Ivanidzo4ka:ivanidze/wordembedding
Jul 31, 2018
Merged

word embedding transform#545
Ivanidzo4ka merged 29 commits into
dotnet:masterfrom
Ivanidzo4ka:ivanidze/wordembedding

Conversation

@Ivanidzo4ka

@Ivanidzo4kaIvanidzo4ka commented Jul 17, 2018

Copy link
Copy Markdown
Contributor

I heard word embedding can be nice thing for Text classification

  • Create issue
  • Put legal attributes for fastText files
  • Put legal attributes for GloVe files

(edited by @justinormont to fix model type names)
closes#615

@shauheen

Copy link
Copy Markdown
Contributor

Thanks @Ivanidzo4ka , can you please create an issue, and explain what is missing from ML.NET. 👍

if (string.IsNullOrWhiteSpace(_modelFileNameWithPath))
{
throw Host.Except("Model file for Word Embedding transform could not be found! " +
@"Please copy the model file '{0}' from '\\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVectors\' " +

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVectors' [](start = 61, length = 67)

what should this look like now?

if (_modelFileNameWithPath == null)
{
throw Host.Except("Model file for Word Embedding transform could not be found! " +
@"Please copy the model file '{0}' from '\\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVectors\' " +

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVect [](start = 61, length = 62)

and here


private static Dictionary<PretrainedModelKind, string> _modelsMetaData = new Dictionary<PretrainedModelKind, string>()
{
{ PretrainedModelKind.GloVe50D, "glove.6B.50d.txt" },

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

GloVe50D [](start = 35, length = 8)

do we have public locations for all of this?
Do we want to be the ones storing them?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we have aka.ms/tlc-resources/


In reply to: 204168101 [](ancestors = 204168101)

WordEmbeddings wrap different embedding models, such as GloVe. Users can specify which embedding to use.
The available options are various versions of <a href="https://nlp.stanford.edu/projects/glove/">GloVe Models</a>, <a href="https://en.wikipedia.org/wiki/FastText">FastText</a>, and <a href="http://anthology.aclweb.org/P/P14/P14-1146.pdf">Sswe</a>.
<para>
Note: As WordEmbedding requires a column with text vector, e.g. %3C'This', 'is', 'good'%3E, users need to create an input column by:

@sfilipisfilipiJul 20, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

' [](start = 85, length = 1)

apostrophes need be encoded too: ' #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hope %27 will work


In reply to: 204168630 [](ancestors = 204168630)

</item>
</list>
In the following example, after the NGramFeaturizer, features named ngram.__ are generated. A new column named ngram_TransformedText is
also created with the text vector, similar as running .split(' '). However, due to the variable length of this column it cannot be properly

@sfilipisfilipiJul 20, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

' ' [](start = 71, length = 3)

encode #Resolved

@sfilipi

sfilipi commented Jul 20, 2018

Copy link
Copy Markdown
Member
 pipeline.Add(new LightLda(("InTextCol" , "OutTextCol")));

bummer! #Resolved


Refers to: src/Microsoft.ML.Transforms/Text/doc.xml:182 in 68696f4. [](commit_id = 68696f4, deletion_comment = False)

<example name="WordEmbeddings">
<example>
<code language="csharp">
pipeline.Add(new WordEmbeddings(("InTextCol" , "OutTextCol")));

@sfilipisfilipiJul 20, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

" [](start = 43, length = 1)

encode those all. #Resolved

Ivan Matantsev added 2 commits July 20, 2018 14:29
remove links to internal storage
WordEmbedding. The output from WordEmbedding is named ngram_TransformedText.__
</para>
<para>
License attributes for pretrained models:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

License attributes for pretrained models: [](start = 9, length = 42)

@GalOshri is this wording looks ok for you?

</summary>
<remarks>
WordEmbeddings wrap different embedding models, such as GloVe. Users can specify which embedding to use.
The available options are various versions of <a href="https://nlp.stanford.edu/projects/glove/">GloVe Models</a>, <a href="https://en.wikipedia.org/wiki/FastText">FastText</a>, and <a href="http://anthology.aclweb.org/P/P14/P14-1146.pdf">Sswe</a>.

@justinormontjustinormontJul 20, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

fastText should be lower camel cased [1] #Resolved

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and SSWE is upper cased #Resolved

<para>
Note: As WordEmbedding requires a column with text vector, e.g. %3C%27This%27, %27is%27, %27good%27%3E, users need to create an input column by:
<list type="bullet">
<item><description>concatenating columns with TX type,</description></item>

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Concatenating columns of single unigrams is very unlikely.

We should be recommending only to use output_tokens=True in NGramFeaturizer(). All of our current pre-trained models require tokens which are lowercased unigrams w/ diacritics removed (which are the defaults for NGramFeaturizer). #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for clarification, didn't knew about this.
Will update documentation in later PR.


In reply to: 204193596 [](ancestors = 204193596)

</list>
In the following example, after the NGramFeaturizer, features named ngram.__ are generated. A new column named ngram_TransformedText is
also created with the text vector, similar as running .split(%27 %27). However, due to the variable length of this column it cannot be properly
converted to pandas dataframe, thus any pipelines/transforms output this text vector column will throw errors. However, we use

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"pandas dataframe" won't apply to the ML.NET #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The curse of copy paste!


In reply to: 204193641 [](ancestors = 204193641)

</item>
<item>
<description>
Glove models by Stanford University, or (Jeffrey Pennington, Richard Socher, Christopher D. Manning) is licensed under <a href="https://opendatacommons.org/licenses/pddl/1.0/">PDDL</a>.

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Glove should be capitalized as GloVe #Resolved

<item>
<description>
Glove models by Stanford University, or (Jeffrey Pennington, Richard Socher, Christopher D. Manning) is licensed under <a href="https://opendatacommons.org/licenses/pddl/1.0/">PDDL</a>.
More information can be found <a href="https://nlp.stanford.edu/projects/glove/">here</a>.

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Linking to their repo would also be nice: https://github.com/stanfordnlp/GloVe #Resolved

</item>
<item>
<description>
Glove models by Stanford University, or (Jeffrey Pennington, Richard Socher, Christopher D. Manning) is licensed under <a href="https://opendatacommons.org/licenses/pddl/1.0/">PDDL</a>.

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Recommend adding, Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014. [GloVe: Global Vectors for Word Representation](https://nlp.stanford.edu/pubs/glove.pdf)., which is their asked for citation format, as per: https://nlp.stanford.edu/projects/glove/
#Resolved

@Ivanidzo4ka

Copy link
Copy Markdown
ContributorAuthor

@dotnet-bot test OSX10.13 Release

@Ivanidzo4ka

Copy link
Copy Markdown
ContributorAuthor

@dotnet-bot test OSX10.13 Debug

int deno = 0;
srcGetter(ref src);
var values = dst.Values;
Utils.EnsureSize(ref values, 3 * dimension, keepOld: false);

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With keepOld: false, the Utils.EnsureSize() function always allocates a new array. Would it be faster to perform the size check and just zero the existing?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm actually not sure why we allocate new array all the time, (considering what we clean it immediately after) so I rather remove this "keepOld"


In reply to: 204564374 [](ancestors = 204564374)

@TomFinleyTomFinleyJul 27, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reactivating. Read implementation or documentation of EnsureSize a bit more carefully. THe keepOld: false certainly does not always allocate a new array. The difference is, in the situation where it is necessary to resize the array, whether that is created through Array.Resize or new... the latter is faster if we can get away with it, which we certainly can here.


In reply to: 204578950 [](ancestors = 204578950,204564374)

for (int i = 0; i < dimension; i++)
{
float currentTerm = wordVector[i];
if (values[i] > currentTerm)

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you point out the code fix for #545 (review)? I'm not seeing it. #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I removed Array.Clear and replace it with setting MaxValue,0,MinValue.


In reply to: 204566178 [](ancestors = 204566178)

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks. I see it now. #Resolved

int offset = 2 * dimension;
for (int i = 0; i < dimension; i++)
{
values[i] = int.MaxValue;

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Array is of type float, so float.MaxValue will be better #Resolved

@justinormontjustinormont left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

<para>
Note: As WordEmbedding requires a column with text vector, e.g. %3C%27this%27, %27is%27, %27good%27%3E, users need to create an input column by
using the output_tokens=True for TextTransform to convert a column with sentences like "This is good" into %3C%27this%27, %27is%27, %27good%27 %3E.
The column for the output token column is renamed original column with a prefix of %27_TranformedText%27.

@justinormontjustinormontJul 24, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can word smith this a bit..
The suffix of %27_TransformedText%27 is added to the original column name to create the output token column. For instance if the input column is %27body%27, the output tokens column is named %27body_TransformedText%27. #Resolved

if (model == null)
model = new Model(dimension);
if (model.Dimension != dimension)
ch.Warning($"Dimension mismatch while reading model file: '{_modelFileNameWithPath}', line number 1, expected dimension = {model.Dimension}, received dimension = {dimension}");

@justinormontjustinormontJul 24, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can remove this warning. The purpose of this block of code is to allow the 1st line to be of different length (and ignored if so). Hence the warning is superfluous. Currently this warning is displayed for all fastText models, even though the model is read correctly (by ignoring the 1st line).

Background: In fastText models, the 1st line is: . Other word embedding models don't have a header line. #Resolved

@Ivanidzo4ka
Ivanidzo4ka requested a review from Zruty0July 25, 2018 19:55
@justinormont

Copy link
Copy Markdown
Contributor

:shipit:

var name = Path.GetFileName(errorResult.FileName);
throw ch.Except($"{errorMessage}\nModel file for Word Embedding transform could not be found! " +
$@"Please copy the model file '{name}' from '{url}' to '{directory}'.");
}

@TomFinleyTomFinleyJul 27, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The user story here doesn't seem great. As in, I'm not sure this transform is actually usable at all. This ResourceManagerUtils points people to the problematically named "https://aka.ms/tlc-resources/". From an outside user's perspective, does that make it useless?

This transform is clearly written with the expectation of a console application not a library. If it were a library, we might expect these sorts of "optional things" would be brought in via nuget dependencies... that is, if you want to use this or that, you subscribe to the appropriate nuget, then in the Arguments of this thing assign the relevant resource as an actual object published by that nuget. (It would be some variety of IComponentFactory.)

I feel like this whole approach needs some deeper thought.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

problematically named resource has it's own issue #546

Model files which we use (except SSWE which is a 70mb) in range from 160MB(glove50d) to 6 GB(fastText). Which I think impossible to fit into any nuget due to size limitation.


In reply to: 205848439 [](ancestors = 205848439)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh boy. That's pretty big. Hmmm... hmmm... this is pretty complicated then. All right let's punt on that for now.

@TomFinleyTomFinley left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you @Ivanidzo4ka ! But seriously, could you create an issue?

@Ivanidzo4ka
Ivanidzo4ka merged commit b727d10 into dotnet:masterJul 31, 2018
codemzs pushed a commit to codemzs/machinelearning that referenced this pull request Aug 1, 2018
Introduce word embedding transform
@ghostghost locked as resolved and limited conversation to collaborators Mar 29, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Word embeddings

5 participants

@Ivanidzo4ka@shauheen@sfilipi@justinormont@TomFinley
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

word embedding transform - #545

Merged
Ivanidzo4ka merged 29 commits into
dotnet:masterfrom
Ivanidzo4ka:ivanidze/wordembedding
Jul 31, 2018
Merged

word embedding transform#545
Ivanidzo4ka merged 29 commits into
dotnet:masterfrom
Ivanidzo4ka:ivanidze/wordembedding

Conversation

@Ivanidzo4ka

@Ivanidzo4kaIvanidzo4ka commented Jul 17, 2018

Copy link
Copy Markdown
Contributor

I heard word embedding can be nice thing for Text classification

  • Create issue
  • Put legal attributes for fastText files
  • Put legal attributes for GloVe files

(edited by @justinormont to fix model type names)
closes#615

@shauheen

Copy link
Copy Markdown
Contributor

Thanks @Ivanidzo4ka , can you please create an issue, and explain what is missing from ML.NET. 👍

if (string.IsNullOrWhiteSpace(_modelFileNameWithPath))
{
throw Host.Except("Model file for Word Embedding transform could not be found! " +
@"Please copy the model file '{0}' from '\\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVectors\' " +

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVectors' [](start = 61, length = 67)

what should this look like now?

if (_modelFileNameWithPath == null)
{
throw Host.Except("Model file for Word Embedding transform could not be found! " +
@"Please copy the model file '{0}' from '\\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVectors\' " +

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVect [](start = 61, length = 62)

and here


private static Dictionary<PretrainedModelKind, string> _modelsMetaData = new Dictionary<PretrainedModelKind, string>()
{
{ PretrainedModelKind.GloVe50D, "glove.6B.50d.txt" },

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

GloVe50D [](start = 35, length = 8)

do we have public locations for all of this?
Do we want to be the ones storing them?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we have aka.ms/tlc-resources/


In reply to: 204168101 [](ancestors = 204168101)

WordEmbeddings wrap different embedding models, such as GloVe. Users can specify which embedding to use.
The available options are various versions of <a href="https://nlp.stanford.edu/projects/glove/">GloVe Models</a>, <a href="https://en.wikipedia.org/wiki/FastText">FastText</a>, and <a href="http://anthology.aclweb.org/P/P14/P14-1146.pdf">Sswe</a>.
<para>
Note: As WordEmbedding requires a column with text vector, e.g. %3C'This', 'is', 'good'%3E, users need to create an input column by:

@sfilipisfilipiJul 20, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

' [](start = 85, length = 1)

apostrophes need be encoded too: ' #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hope %27 will work


In reply to: 204168630 [](ancestors = 204168630)

</item>
</list>
In the following example, after the NGramFeaturizer, features named ngram.__ are generated. A new column named ngram_TransformedText is
also created with the text vector, similar as running .split(' '). However, due to the variable length of this column it cannot be properly

@sfilipisfilipiJul 20, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

' ' [](start = 71, length = 3)

encode #Resolved

@sfilipi

sfilipi commented Jul 20, 2018

Copy link
Copy Markdown
Member
 pipeline.Add(new LightLda(("InTextCol" , "OutTextCol")));

bummer! #Resolved


Refers to: src/Microsoft.ML.Transforms/Text/doc.xml:182 in 68696f4. [](commit_id = 68696f4, deletion_comment = False)

<example name="WordEmbeddings">
<example>
<code language="csharp">
pipeline.Add(new WordEmbeddings(("InTextCol" , "OutTextCol")));

@sfilipisfilipiJul 20, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

" [](start = 43, length = 1)

encode those all. #Resolved

Ivan Matantsev added 2 commits July 20, 2018 14:29
remove links to internal storage
WordEmbedding. The output from WordEmbedding is named ngram_TransformedText.__
</para>
<para>
License attributes for pretrained models:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

License attributes for pretrained models: [](start = 9, length = 42)

@GalOshri is this wording looks ok for you?

</summary>
<remarks>
WordEmbeddings wrap different embedding models, such as GloVe. Users can specify which embedding to use.
The available options are various versions of <a href="https://nlp.stanford.edu/projects/glove/">GloVe Models</a>, <a href="https://en.wikipedia.org/wiki/FastText">FastText</a>, and <a href="http://anthology.aclweb.org/P/P14/P14-1146.pdf">Sswe</a>.

@justinormontjustinormontJul 20, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

fastText should be lower camel cased [1] #Resolved

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and SSWE is upper cased #Resolved

<para>
Note: As WordEmbedding requires a column with text vector, e.g. %3C%27This%27, %27is%27, %27good%27%3E, users need to create an input column by:
<list type="bullet">
<item><description>concatenating columns with TX type,</description></item>

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Concatenating columns of single unigrams is very unlikely.

We should be recommending only to use output_tokens=True in NGramFeaturizer(). All of our current pre-trained models require tokens which are lowercased unigrams w/ diacritics removed (which are the defaults for NGramFeaturizer). #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for clarification, didn't knew about this.
Will update documentation in later PR.


In reply to: 204193596 [](ancestors = 204193596)

</list>
In the following example, after the NGramFeaturizer, features named ngram.__ are generated. A new column named ngram_TransformedText is
also created with the text vector, similar as running .split(%27 %27). However, due to the variable length of this column it cannot be properly
converted to pandas dataframe, thus any pipelines/transforms output this text vector column will throw errors. However, we use

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"pandas dataframe" won't apply to the ML.NET #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The curse of copy paste!


In reply to: 204193641 [](ancestors = 204193641)

</item>
<item>
<description>
Glove models by Stanford University, or (Jeffrey Pennington, Richard Socher, Christopher D. Manning) is licensed under <a href="https://opendatacommons.org/licenses/pddl/1.0/">PDDL</a>.

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Glove should be capitalized as GloVe #Resolved

<item>
<description>
Glove models by Stanford University, or (Jeffrey Pennington, Richard Socher, Christopher D. Manning) is licensed under <a href="https://opendatacommons.org/licenses/pddl/1.0/">PDDL</a>.
More information can be found <a href="https://nlp.stanford.edu/projects/glove/">here</a>.

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Linking to their repo would also be nice: https://github.com/stanfordnlp/GloVe #Resolved

</item>
<item>
<description>
Glove models by Stanford University, or (Jeffrey Pennington, Richard Socher, Christopher D. Manning) is licensed under <a href="https://opendatacommons.org/licenses/pddl/1.0/">PDDL</a>.

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Recommend adding, Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014. [GloVe: Global Vectors for Word Representation](https://nlp.stanford.edu/pubs/glove.pdf)., which is their asked for citation format, as per: https://nlp.stanford.edu/projects/glove/
#Resolved

@Ivanidzo4ka

Copy link
Copy Markdown
ContributorAuthor

@dotnet-bot test OSX10.13 Release

@Ivanidzo4ka

Copy link
Copy Markdown
ContributorAuthor

@dotnet-bot test OSX10.13 Debug

int deno = 0;
srcGetter(ref src);
var values = dst.Values;
Utils.EnsureSize(ref values, 3 * dimension, keepOld: false);

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With keepOld: false, the Utils.EnsureSize() function always allocates a new array. Would it be faster to perform the size check and just zero the existing?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm actually not sure why we allocate new array all the time, (considering what we clean it immediately after) so I rather remove this "keepOld"


In reply to: 204564374 [](ancestors = 204564374)

@TomFinleyTomFinleyJul 27, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reactivating. Read implementation or documentation of EnsureSize a bit more carefully. THe keepOld: false certainly does not always allocate a new array. The difference is, in the situation where it is necessary to resize the array, whether that is created through Array.Resize or new... the latter is faster if we can get away with it, which we certainly can here.


In reply to: 204578950 [](ancestors = 204578950,204564374)

for (int i = 0; i < dimension; i++)
{
float currentTerm = wordVector[i];
if (values[i] > currentTerm)

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you point out the code fix for #545 (review)? I'm not seeing it. #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I removed Array.Clear and replace it with setting MaxValue,0,MinValue.


In reply to: 204566178 [](ancestors = 204566178)

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks. I see it now. #Resolved

int offset = 2 * dimension;
for (int i = 0; i < dimension; i++)
{
values[i] = int.MaxValue;

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Array is of type float, so float.MaxValue will be better #Resolved

@justinormontjustinormont left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

<para>
Note: As WordEmbedding requires a column with text vector, e.g. %3C%27this%27, %27is%27, %27good%27%3E, users need to create an input column by
using the output_tokens=True for TextTransform to convert a column with sentences like "This is good" into %3C%27this%27, %27is%27, %27good%27 %3E.
The column for the output token column is renamed original column with a prefix of %27_TranformedText%27.

@justinormontjustinormontJul 24, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can word smith this a bit..
The suffix of %27_TransformedText%27 is added to the original column name to create the output token column. For instance if the input column is %27body%27, the output tokens column is named %27body_TransformedText%27. #Resolved

if (model == null)
model = new Model(dimension);
if (model.Dimension != dimension)
ch.Warning($"Dimension mismatch while reading model file: '{_modelFileNameWithPath}', line number 1, expected dimension = {model.Dimension}, received dimension = {dimension}");

@justinormontjustinormontJul 24, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can remove this warning. The purpose of this block of code is to allow the 1st line to be of different length (and ignored if so). Hence the warning is superfluous. Currently this warning is displayed for all fastText models, even though the model is read correctly (by ignoring the 1st line).

Background: In fastText models, the 1st line is: . Other word embedding models don't have a header line. #Resolved

@Ivanidzo4ka
Ivanidzo4ka requested a review from Zruty0July 25, 2018 19:55
@justinormont

Copy link
Copy Markdown
Contributor

:shipit:

var name = Path.GetFileName(errorResult.FileName);
throw ch.Except($"{errorMessage}\nModel file for Word Embedding transform could not be found! " +
$@"Please copy the model file '{name}' from '{url}' to '{directory}'.");
}

@TomFinleyTomFinleyJul 27, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The user story here doesn't seem great. As in, I'm not sure this transform is actually usable at all. This ResourceManagerUtils points people to the problematically named "https://aka.ms/tlc-resources/". From an outside user's perspective, does that make it useless?

This transform is clearly written with the expectation of a console application not a library. If it were a library, we might expect these sorts of "optional things" would be brought in via nuget dependencies... that is, if you want to use this or that, you subscribe to the appropriate nuget, then in the Arguments of this thing assign the relevant resource as an actual object published by that nuget. (It would be some variety of IComponentFactory.)

I feel like this whole approach needs some deeper thought.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

problematically named resource has it's own issue #546

Model files which we use (except SSWE which is a 70mb) in range from 160MB(glove50d) to 6 GB(fastText). Which I think impossible to fit into any nuget due to size limitation.


In reply to: 205848439 [](ancestors = 205848439)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh boy. That's pretty big. Hmmm... hmmm... this is pretty complicated then. All right let's punt on that for now.

@TomFinleyTomFinley left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you @Ivanidzo4ka ! But seriously, could you create an issue?

@Ivanidzo4ka
Ivanidzo4ka merged commit b727d10 into dotnet:masterJul 31, 2018
codemzs pushed a commit to codemzs/machinelearning that referenced this pull request Aug 1, 2018
Introduce word embedding transform
@ghostghost locked as resolved and limited conversation to collaborators Mar 29, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Word embeddings

5 participants

@Ivanidzo4ka@shauheen@sfilipi@justinormont@TomFinley
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

word embedding transform - #545

Merged
Ivanidzo4ka merged 29 commits into
dotnet:masterfrom
Ivanidzo4ka:ivanidze/wordembedding
Jul 31, 2018
Merged

word embedding transform#545
Ivanidzo4ka merged 29 commits into
dotnet:masterfrom
Ivanidzo4ka:ivanidze/wordembedding

Conversation

@Ivanidzo4ka

@Ivanidzo4kaIvanidzo4ka commented Jul 17, 2018

Copy link
Copy Markdown
Contributor

I heard word embedding can be nice thing for Text classification

  • Create issue
  • Put legal attributes for fastText files
  • Put legal attributes for GloVe files

(edited by @justinormont to fix model type names)
closes#615

@shauheen

Copy link
Copy Markdown
Contributor

Thanks @Ivanidzo4ka , can you please create an issue, and explain what is missing from ML.NET. 👍

if (string.IsNullOrWhiteSpace(_modelFileNameWithPath))
{
throw Host.Except("Model file for Word Embedding transform could not be found! " +
@"Please copy the model file '{0}' from '\\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVectors\' " +

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVectors' [](start = 61, length = 67)

what should this look like now?

if (_modelFileNameWithPath == null)
{
throw Host.Except("Model file for Word Embedding transform could not be found! " +
@"Please copy the model file '{0}' from '\\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVectors\' " +

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVect [](start = 61, length = 62)

and here


private static Dictionary<PretrainedModelKind, string> _modelsMetaData = new Dictionary<PretrainedModelKind, string>()
{
{ PretrainedModelKind.GloVe50D, "glove.6B.50d.txt" },

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

GloVe50D [](start = 35, length = 8)

do we have public locations for all of this?
Do we want to be the ones storing them?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we have aka.ms/tlc-resources/


In reply to: 204168101 [](ancestors = 204168101)

WordEmbeddings wrap different embedding models, such as GloVe. Users can specify which embedding to use.
The available options are various versions of <a href="https://nlp.stanford.edu/projects/glove/">GloVe Models</a>, <a href="https://en.wikipedia.org/wiki/FastText">FastText</a>, and <a href="http://anthology.aclweb.org/P/P14/P14-1146.pdf">Sswe</a>.
<para>
Note: As WordEmbedding requires a column with text vector, e.g. %3C'This', 'is', 'good'%3E, users need to create an input column by:

@sfilipisfilipiJul 20, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

' [](start = 85, length = 1)

apostrophes need be encoded too: ' #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hope %27 will work


In reply to: 204168630 [](ancestors = 204168630)

</item>
</list>
In the following example, after the NGramFeaturizer, features named ngram.__ are generated. A new column named ngram_TransformedText is
also created with the text vector, similar as running .split(' '). However, due to the variable length of this column it cannot be properly

@sfilipisfilipiJul 20, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

' ' [](start = 71, length = 3)

encode #Resolved

@sfilipi

sfilipi commented Jul 20, 2018

Copy link
Copy Markdown
Member
 pipeline.Add(new LightLda(("InTextCol" , "OutTextCol")));

bummer! #Resolved


Refers to: src/Microsoft.ML.Transforms/Text/doc.xml:182 in 68696f4. [](commit_id = 68696f4, deletion_comment = False)

<example name="WordEmbeddings">
<example>
<code language="csharp">
pipeline.Add(new WordEmbeddings(("InTextCol" , "OutTextCol")));

@sfilipisfilipiJul 20, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

" [](start = 43, length = 1)

encode those all. #Resolved

Ivan Matantsev added 2 commits July 20, 2018 14:29
remove links to internal storage
WordEmbedding. The output from WordEmbedding is named ngram_TransformedText.__
</para>
<para>
License attributes for pretrained models:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

License attributes for pretrained models: [](start = 9, length = 42)

@GalOshri is this wording looks ok for you?

</summary>
<remarks>
WordEmbeddings wrap different embedding models, such as GloVe. Users can specify which embedding to use.
The available options are various versions of <a href="https://nlp.stanford.edu/projects/glove/">GloVe Models</a>, <a href="https://en.wikipedia.org/wiki/FastText">FastText</a>, and <a href="http://anthology.aclweb.org/P/P14/P14-1146.pdf">Sswe</a>.

@justinormontjustinormontJul 20, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

fastText should be lower camel cased [1] #Resolved

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and SSWE is upper cased #Resolved

<para>
Note: As WordEmbedding requires a column with text vector, e.g. %3C%27This%27, %27is%27, %27good%27%3E, users need to create an input column by:
<list type="bullet">
<item><description>concatenating columns with TX type,</description></item>

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Concatenating columns of single unigrams is very unlikely.

We should be recommending only to use output_tokens=True in NGramFeaturizer(). All of our current pre-trained models require tokens which are lowercased unigrams w/ diacritics removed (which are the defaults for NGramFeaturizer). #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for clarification, didn't knew about this.
Will update documentation in later PR.


In reply to: 204193596 [](ancestors = 204193596)

</list>
In the following example, after the NGramFeaturizer, features named ngram.__ are generated. A new column named ngram_TransformedText is
also created with the text vector, similar as running .split(%27 %27). However, due to the variable length of this column it cannot be properly
converted to pandas dataframe, thus any pipelines/transforms output this text vector column will throw errors. However, we use

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"pandas dataframe" won't apply to the ML.NET #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The curse of copy paste!


In reply to: 204193641 [](ancestors = 204193641)

</item>
<item>
<description>
Glove models by Stanford University, or (Jeffrey Pennington, Richard Socher, Christopher D. Manning) is licensed under <a href="https://opendatacommons.org/licenses/pddl/1.0/">PDDL</a>.

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Glove should be capitalized as GloVe #Resolved

<item>
<description>
Glove models by Stanford University, or (Jeffrey Pennington, Richard Socher, Christopher D. Manning) is licensed under <a href="https://opendatacommons.org/licenses/pddl/1.0/">PDDL</a>.
More information can be found <a href="https://nlp.stanford.edu/projects/glove/">here</a>.

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Linking to their repo would also be nice: https://github.com/stanfordnlp/GloVe #Resolved

</item>
<item>
<description>
Glove models by Stanford University, or (Jeffrey Pennington, Richard Socher, Christopher D. Manning) is licensed under <a href="https://opendatacommons.org/licenses/pddl/1.0/">PDDL</a>.

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Recommend adding, Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014. [GloVe: Global Vectors for Word Representation](https://nlp.stanford.edu/pubs/glove.pdf)., which is their asked for citation format, as per: https://nlp.stanford.edu/projects/glove/
#Resolved

@Ivanidzo4ka

Copy link
Copy Markdown
ContributorAuthor

@dotnet-bot test OSX10.13 Release

@Ivanidzo4ka

Copy link
Copy Markdown
ContributorAuthor

@dotnet-bot test OSX10.13 Debug

int deno = 0;
srcGetter(ref src);
var values = dst.Values;
Utils.EnsureSize(ref values, 3 * dimension, keepOld: false);

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With keepOld: false, the Utils.EnsureSize() function always allocates a new array. Would it be faster to perform the size check and just zero the existing?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm actually not sure why we allocate new array all the time, (considering what we clean it immediately after) so I rather remove this "keepOld"


In reply to: 204564374 [](ancestors = 204564374)

@TomFinleyTomFinleyJul 27, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reactivating. Read implementation or documentation of EnsureSize a bit more carefully. THe keepOld: false certainly does not always allocate a new array. The difference is, in the situation where it is necessary to resize the array, whether that is created through Array.Resize or new... the latter is faster if we can get away with it, which we certainly can here.


In reply to: 204578950 [](ancestors = 204578950,204564374)

for (int i = 0; i < dimension; i++)
{
float currentTerm = wordVector[i];
if (values[i] > currentTerm)

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you point out the code fix for #545 (review)? I'm not seeing it. #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I removed Array.Clear and replace it with setting MaxValue,0,MinValue.


In reply to: 204566178 [](ancestors = 204566178)

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks. I see it now. #Resolved

int offset = 2 * dimension;
for (int i = 0; i < dimension; i++)
{
values[i] = int.MaxValue;

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Array is of type float, so float.MaxValue will be better #Resolved

@justinormontjustinormont left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

<para>
Note: As WordEmbedding requires a column with text vector, e.g. %3C%27this%27, %27is%27, %27good%27%3E, users need to create an input column by
using the output_tokens=True for TextTransform to convert a column with sentences like "This is good" into %3C%27this%27, %27is%27, %27good%27 %3E.
The column for the output token column is renamed original column with a prefix of %27_TranformedText%27.

@justinormontjustinormontJul 24, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can word smith this a bit..
The suffix of %27_TransformedText%27 is added to the original column name to create the output token column. For instance if the input column is %27body%27, the output tokens column is named %27body_TransformedText%27. #Resolved

if (model == null)
model = new Model(dimension);
if (model.Dimension != dimension)
ch.Warning($"Dimension mismatch while reading model file: '{_modelFileNameWithPath}', line number 1, expected dimension = {model.Dimension}, received dimension = {dimension}");

@justinormontjustinormontJul 24, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can remove this warning. The purpose of this block of code is to allow the 1st line to be of different length (and ignored if so). Hence the warning is superfluous. Currently this warning is displayed for all fastText models, even though the model is read correctly (by ignoring the 1st line).

Background: In fastText models, the 1st line is: . Other word embedding models don't have a header line. #Resolved

@Ivanidzo4ka
Ivanidzo4ka requested a review from Zruty0July 25, 2018 19:55
@justinormont

Copy link
Copy Markdown
Contributor

:shipit:

var name = Path.GetFileName(errorResult.FileName);
throw ch.Except($"{errorMessage}\nModel file for Word Embedding transform could not be found! " +
$@"Please copy the model file '{name}' from '{url}' to '{directory}'.");
}

@TomFinleyTomFinleyJul 27, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The user story here doesn't seem great. As in, I'm not sure this transform is actually usable at all. This ResourceManagerUtils points people to the problematically named "https://aka.ms/tlc-resources/". From an outside user's perspective, does that make it useless?

This transform is clearly written with the expectation of a console application not a library. If it were a library, we might expect these sorts of "optional things" would be brought in via nuget dependencies... that is, if you want to use this or that, you subscribe to the appropriate nuget, then in the Arguments of this thing assign the relevant resource as an actual object published by that nuget. (It would be some variety of IComponentFactory.)

I feel like this whole approach needs some deeper thought.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

problematically named resource has it's own issue #546

Model files which we use (except SSWE which is a 70mb) in range from 160MB(glove50d) to 6 GB(fastText). Which I think impossible to fit into any nuget due to size limitation.


In reply to: 205848439 [](ancestors = 205848439)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh boy. That's pretty big. Hmmm... hmmm... this is pretty complicated then. All right let's punt on that for now.

@TomFinleyTomFinley left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you @Ivanidzo4ka ! But seriously, could you create an issue?

@Ivanidzo4ka
Ivanidzo4ka merged commit b727d10 into dotnet:masterJul 31, 2018
codemzs pushed a commit to codemzs/machinelearning that referenced this pull request Aug 1, 2018
Introduce word embedding transform
@ghostghost locked as resolved and limited conversation to collaborators Mar 29, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Word embeddings

5 participants

@Ivanidzo4ka@shauheen@sfilipi@justinormont@TomFinley
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

word embedding transform - #545

Merged
Ivanidzo4ka merged 29 commits into
dotnet:masterfrom
Ivanidzo4ka:ivanidze/wordembedding
Jul 31, 2018
Merged

word embedding transform#545
Ivanidzo4ka merged 29 commits into
dotnet:masterfrom
Ivanidzo4ka:ivanidze/wordembedding

Conversation

@Ivanidzo4ka

@Ivanidzo4kaIvanidzo4ka commented Jul 17, 2018

Copy link
Copy Markdown
Contributor

I heard word embedding can be nice thing for Text classification

  • Create issue
  • Put legal attributes for fastText files
  • Put legal attributes for GloVe files

(edited by @justinormont to fix model type names)
closes#615

@shauheen

Copy link
Copy Markdown
Contributor

Thanks @Ivanidzo4ka , can you please create an issue, and explain what is missing from ML.NET. 👍

if (string.IsNullOrWhiteSpace(_modelFileNameWithPath))
{
throw Host.Except("Model file for Word Embedding transform could not be found! " +
@"Please copy the model file '{0}' from '\\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVectors\' " +

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVectors' [](start = 61, length = 67)

what should this look like now?

if (_modelFileNameWithPath == null)
{
throw Host.Except("Model file for Word Embedding transform could not be found! " +
@"Please copy the model file '{0}' from '\\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVectors\' " +

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

\cloudmltlc\TLC\releases\Resources\3.8\TextAnalytics\WordVect [](start = 61, length = 62)

and here


private static Dictionary<PretrainedModelKind, string> _modelsMetaData = new Dictionary<PretrainedModelKind, string>()
{
{ PretrainedModelKind.GloVe50D, "glove.6B.50d.txt" },

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

GloVe50D [](start = 35, length = 8)

do we have public locations for all of this?
Do we want to be the ones storing them?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we have aka.ms/tlc-resources/


In reply to: 204168101 [](ancestors = 204168101)

WordEmbeddings wrap different embedding models, such as GloVe. Users can specify which embedding to use.
The available options are various versions of <a href="https://nlp.stanford.edu/projects/glove/">GloVe Models</a>, <a href="https://en.wikipedia.org/wiki/FastText">FastText</a>, and <a href="http://anthology.aclweb.org/P/P14/P14-1146.pdf">Sswe</a>.
<para>
Note: As WordEmbedding requires a column with text vector, e.g. %3C'This', 'is', 'good'%3E, users need to create an input column by:

@sfilipisfilipiJul 20, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

' [](start = 85, length = 1)

apostrophes need be encoded too: ' #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hope %27 will work


In reply to: 204168630 [](ancestors = 204168630)

</item>
</list>
In the following example, after the NGramFeaturizer, features named ngram.__ are generated. A new column named ngram_TransformedText is
also created with the text vector, similar as running .split(' '). However, due to the variable length of this column it cannot be properly

@sfilipisfilipiJul 20, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

' ' [](start = 71, length = 3)

encode #Resolved

@sfilipi

sfilipi commented Jul 20, 2018

Copy link
Copy Markdown
Member
 pipeline.Add(new LightLda(("InTextCol" , "OutTextCol")));

bummer! #Resolved


Refers to: src/Microsoft.ML.Transforms/Text/doc.xml:182 in 68696f4. [](commit_id = 68696f4, deletion_comment = False)

<example name="WordEmbeddings">
<example>
<code language="csharp">
pipeline.Add(new WordEmbeddings(("InTextCol" , "OutTextCol")));

@sfilipisfilipiJul 20, 2018

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

" [](start = 43, length = 1)

encode those all. #Resolved

Ivan Matantsev added 2 commits July 20, 2018 14:29
remove links to internal storage
WordEmbedding. The output from WordEmbedding is named ngram_TransformedText.__
</para>
<para>
License attributes for pretrained models:

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

License attributes for pretrained models: [](start = 9, length = 42)

@GalOshri is this wording looks ok for you?

</summary>
<remarks>
WordEmbeddings wrap different embedding models, such as GloVe. Users can specify which embedding to use.
The available options are various versions of <a href="https://nlp.stanford.edu/projects/glove/">GloVe Models</a>, <a href="https://en.wikipedia.org/wiki/FastText">FastText</a>, and <a href="http://anthology.aclweb.org/P/P14/P14-1146.pdf">Sswe</a>.

@justinormontjustinormontJul 20, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

fastText should be lower camel cased [1] #Resolved

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and SSWE is upper cased #Resolved

<para>
Note: As WordEmbedding requires a column with text vector, e.g. %3C%27This%27, %27is%27, %27good%27%3E, users need to create an input column by:
<list type="bullet">
<item><description>concatenating columns with TX type,</description></item>

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Concatenating columns of single unigrams is very unlikely.

We should be recommending only to use output_tokens=True in NGramFeaturizer(). All of our current pre-trained models require tokens which are lowercased unigrams w/ diacritics removed (which are the defaults for NGramFeaturizer). #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for clarification, didn't knew about this.
Will update documentation in later PR.


In reply to: 204193596 [](ancestors = 204193596)

</list>
In the following example, after the NGramFeaturizer, features named ngram.__ are generated. A new column named ngram_TransformedText is
also created with the text vector, similar as running .split(%27 %27). However, due to the variable length of this column it cannot be properly
converted to pandas dataframe, thus any pipelines/transforms output this text vector column will throw errors. However, we use

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"pandas dataframe" won't apply to the ML.NET #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The curse of copy paste!


In reply to: 204193641 [](ancestors = 204193641)

</item>
<item>
<description>
Glove models by Stanford University, or (Jeffrey Pennington, Richard Socher, Christopher D. Manning) is licensed under <a href="https://opendatacommons.org/licenses/pddl/1.0/">PDDL</a>.

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Glove should be capitalized as GloVe #Resolved

<item>
<description>
Glove models by Stanford University, or (Jeffrey Pennington, Richard Socher, Christopher D. Manning) is licensed under <a href="https://opendatacommons.org/licenses/pddl/1.0/">PDDL</a>.
More information can be found <a href="https://nlp.stanford.edu/projects/glove/">here</a>.

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Linking to their repo would also be nice: https://github.com/stanfordnlp/GloVe #Resolved

</item>
<item>
<description>
Glove models by Stanford University, or (Jeffrey Pennington, Richard Socher, Christopher D. Manning) is licensed under <a href="https://opendatacommons.org/licenses/pddl/1.0/">PDDL</a>.

@justinormontjustinormontJul 21, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Recommend adding, Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014. [GloVe: Global Vectors for Word Representation](https://nlp.stanford.edu/pubs/glove.pdf)., which is their asked for citation format, as per: https://nlp.stanford.edu/projects/glove/
#Resolved

@Ivanidzo4ka

Copy link
Copy Markdown
ContributorAuthor

@dotnet-bot test OSX10.13 Release

@Ivanidzo4ka

Copy link
Copy Markdown
ContributorAuthor

@dotnet-bot test OSX10.13 Debug

int deno = 0;
srcGetter(ref src);
var values = dst.Values;
Utils.EnsureSize(ref values, 3 * dimension, keepOld: false);

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With keepOld: false, the Utils.EnsureSize() function always allocates a new array. Would it be faster to perform the size check and just zero the existing?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm actually not sure why we allocate new array all the time, (considering what we clean it immediately after) so I rather remove this "keepOld"


In reply to: 204564374 [](ancestors = 204564374)

@TomFinleyTomFinleyJul 27, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reactivating. Read implementation or documentation of EnsureSize a bit more carefully. THe keepOld: false certainly does not always allocate a new array. The difference is, in the situation where it is necessary to resize the array, whether that is created through Array.Resize or new... the latter is faster if we can get away with it, which we certainly can here.


In reply to: 204578950 [](ancestors = 204578950,204564374)

for (int i = 0; i < dimension; i++)
{
float currentTerm = wordVector[i];
if (values[i] > currentTerm)

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you point out the code fix for #545 (review)? I'm not seeing it. #Resolved

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I removed Array.Clear and replace it with setting MaxValue,0,MinValue.


In reply to: 204566178 [](ancestors = 204566178)

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks. I see it now. #Resolved

int offset = 2 * dimension;
for (int i = 0; i < dimension; i++)
{
values[i] = int.MaxValue;

@justinormontjustinormontJul 23, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Array is of type float, so float.MaxValue will be better #Resolved

@justinormontjustinormont left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

<para>
Note: As WordEmbedding requires a column with text vector, e.g. %3C%27this%27, %27is%27, %27good%27%3E, users need to create an input column by
using the output_tokens=True for TextTransform to convert a column with sentences like "This is good" into %3C%27this%27, %27is%27, %27good%27 %3E.
The column for the output token column is renamed original column with a prefix of %27_TranformedText%27.

@justinormontjustinormontJul 24, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can word smith this a bit..
The suffix of %27_TransformedText%27 is added to the original column name to create the output token column. For instance if the input column is %27body%27, the output tokens column is named %27body_TransformedText%27. #Resolved

if (model == null)
model = new Model(dimension);
if (model.Dimension != dimension)
ch.Warning($"Dimension mismatch while reading model file: '{_modelFileNameWithPath}', line number 1, expected dimension = {model.Dimension}, received dimension = {dimension}");

@justinormontjustinormontJul 24, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can remove this warning. The purpose of this block of code is to allow the 1st line to be of different length (and ignored if so). Hence the warning is superfluous. Currently this warning is displayed for all fastText models, even though the model is read correctly (by ignoring the 1st line).

Background: In fastText models, the 1st line is: . Other word embedding models don't have a header line. #Resolved

@Ivanidzo4ka
Ivanidzo4ka requested a review from Zruty0July 25, 2018 19:55
@justinormont

Copy link
Copy Markdown
Contributor

:shipit:

var name = Path.GetFileName(errorResult.FileName);
throw ch.Except($"{errorMessage}\nModel file for Word Embedding transform could not be found! " +
$@"Please copy the model file '{name}' from '{url}' to '{directory}'.");
}

@TomFinleyTomFinleyJul 27, 2018

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The user story here doesn't seem great. As in, I'm not sure this transform is actually usable at all. This ResourceManagerUtils points people to the problematically named "https://aka.ms/tlc-resources/". From an outside user's perspective, does that make it useless?

This transform is clearly written with the expectation of a console application not a library. If it were a library, we might expect these sorts of "optional things" would be brought in via nuget dependencies... that is, if you want to use this or that, you subscribe to the appropriate nuget, then in the Arguments of this thing assign the relevant resource as an actual object published by that nuget. (It would be some variety of IComponentFactory.)

I feel like this whole approach needs some deeper thought.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

problematically named resource has it's own issue #546

Model files which we use (except SSWE which is a 70mb) in range from 160MB(glove50d) to 6 GB(fastText). Which I think impossible to fit into any nuget due to size limitation.


In reply to: 205848439 [](ancestors = 205848439)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh boy. That's pretty big. Hmmm... hmmm... this is pretty complicated then. All right let's punt on that for now.

@TomFinleyTomFinley left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you @Ivanidzo4ka ! But seriously, could you create an issue?

@Ivanidzo4ka
Ivanidzo4ka merged commit b727d10 into dotnet:masterJul 31, 2018
codemzs pushed a commit to codemzs/machinelearning that referenced this pull request Aug 1, 2018
Introduce word embedding transform
@ghostghost locked as resolved and limited conversation to collaborators Mar 29, 2022
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Word embeddings

5 participants

@Ivanidzo4ka@shauheen@sfilipi@justinormont@TomFinley