Repository files navigation

For testing just run
python LanguageDetect.py
and then, enter an example
The model is stored in the .pickle file.
For training, change the read_from_file_and_test flag in the main function to False and run.
Network Design
I chose to make my neural network with objects as I felt it lead to a cleaner albeit slower
implementation.
The class Neuron has activation and a list of weights. There as many weights in a neuron as there
are neurons in the previous layer. The feed_forward method abstracts the weighted sum and
sigmoid at the neurons. Every neuron can thus compute its own activation.
The class NeuralNetwork has the learning rate alpha, and two lists for hidden and output neurons.
My network has 5 input neurons, 5 hidden neurons and 3 output neurons.
Feature Selection
I decided to use the following features for the three languages when I evaluate the training data.
The number of words that end in vowels is a very strong indicator for Italian.
Presence of the bigram “ij” for Dutch. “ij” has a heavier weight in my features as it is an alphabet in
Dutch and almost never occurs in the others.
Number of words with more than 8 characters indicating Dutch as dutch has frequent long words.
The number of occurrences of the most frequent bigram “th” as part of the most frequent word
“the” is a good feature to pin down English. As is the bigram “ed” at the end of words for English in
the past tense as most Wikipedia articles are.
The raw feature counts are scaled to the length of the example to represent the fraction of the text
the feature occurs in giving me a feature list in the range of [0, 1].
Data gathering and data set construction.
I used the wikis for the three languages as instructed. But only about 25% of the data is Wikipedia.
The rest if from Project Gutenberg. I got three novels in text format for the three languages. A small
python script and some unix commands gave me my data set. The python script makes examples of
random word count between 10 and a 200 words.
I have a training file for each language. There is an example on each line. The test and validation
set were made in the same way. Each language thus has training, validation and testing sets.
Each training set has 600 examples for a total of 1800 examples for the three languages.
The validation set has 300 examples each for a total of 900 examples in the set for all three
languages.
The testing set is the same size as the validation set, but has different examples.
Training process:
All three training files are read and evaluated. The evaluation takes in a block of text and returns a
list of features that represent the text. This is used as input to the neural network. The output is a list
where the correct class is a 1 and the rest are 0s.
For instance, the text:
Led Zeppelin's next album, Houses of the Holy, was released in March 1973. It featured further
experimentation by the band, who expanded their use of synthesisers and mellotron orchestration.
The predominately orange album cover, designed by the London-based design group Hipgnosis,
depicts images of nude children climbing the Giant's Causeway in Northern Ireland. Although the
children are not shown from the front, the cover was controversial at the time of the album's
release. As with the band's fourth album, neither their name nor the album title was printed on the
sleeve
is evaluated as [0.2527, 0.0329, 0.0329, 0.4615] and the corresponding output classs for training is
[1, 0, 0] for English.
All the examples containing (input, output) are shuffled so the network is not trained on the classes sequentially.
The feed forward mechanism is implemented in a separate function so I can test novel examples
easily.
Backpropogation:
The number of epochs is given as input to the back_propogation method. In each epoch, the
network is trained on all examples and the error is recorded. After each epoch, the neural network is
validated on the validation set to keep track of how well is the network leaning. Ideally we'd like to
see a decrease in the errors after each epoch.
I use sum of squared error as the metric for error.
To avoid overfitting, I use a threshold value of validation error and stop training once it drops below
the threshold.
I tried a lot of other things. I let it run for different values of alpha, epochs and number of hidden
layers before settling for one set of values.
I did three random restarts and chose the best out of the three models based on error in the validtion
set and test set accuracy.
I used matplotlib to help me with these decisions.
Testing:
Once training is done, the user is prompted for the novel input through standard input. The user
must enter a piece of text and hit Enter. The neural network returns what language it thinks the text
belongs to.
The user can also enter “test” at the input followed by three test files in the order English, Italian
and Dutch. The network classifies all these and prints a confusion matrix for the given test set. The
user can also enter “default” to test with files named en_test, it_test and nl_test included with the
code.
I let my model classify the text of 15 randomly selected novels, 5 from each language to give a total
of more than 11,000 examples and got an accuracy of 98.87%.
That is a sufficiently high accuracy for a test set that is 6 times larger than the training set.
I have included these big test files. They have “_big” in their names.
The accuracy of the model over the test set is also printed.

About

Neural Network that can classify English, Italian and Dutch

Resources

Stars

5 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

For testing just run
python LanguageDetect.py
and then, enter an example
The model is stored in the .pickle file.
For training, change the read_from_file_and_test flag in the main function to False and run.
Network Design
I chose to make my neural network with objects as I felt it lead to a cleaner albeit slower
implementation.
The class Neuron has activation and a list of weights. There as many weights in a neuron as there
are neurons in the previous layer. The feed_forward method abstracts the weighted sum and
sigmoid at the neurons. Every neuron can thus compute its own activation.
The class NeuralNetwork has the learning rate alpha, and two lists for hidden and output neurons.
My network has 5 input neurons, 5 hidden neurons and 3 output neurons.
Feature Selection
I decided to use the following features for the three languages when I evaluate the training data.
The number of words that end in vowels is a very strong indicator for Italian.
Presence of the bigram “ij” for Dutch. “ij” has a heavier weight in my features as it is an alphabet in
Dutch and almost never occurs in the others.
Number of words with more than 8 characters indicating Dutch as dutch has frequent long words.
The number of occurrences of the most frequent bigram “th” as part of the most frequent word
“the” is a good feature to pin down English. As is the bigram “ed” at the end of words for English in
the past tense as most Wikipedia articles are.
The raw feature counts are scaled to the length of the example to represent the fraction of the text
the feature occurs in giving me a feature list in the range of [0, 1].
Data gathering and data set construction.
I used the wikis for the three languages as instructed. But only about 25% of the data is Wikipedia.
The rest if from Project Gutenberg. I got three novels in text format for the three languages. A small
python script and some unix commands gave me my data set. The python script makes examples of
random word count between 10 and a 200 words.
I have a training file for each language. There is an example on each line. The test and validation
set were made in the same way. Each language thus has training, validation and testing sets.
Each training set has 600 examples for a total of 1800 examples for the three languages.
The validation set has 300 examples each for a total of 900 examples in the set for all three
languages.
The testing set is the same size as the validation set, but has different examples.
Training process:
All three training files are read and evaluated. The evaluation takes in a block of text and returns a
list of features that represent the text. This is used as input to the neural network. The output is a list
where the correct class is a 1 and the rest are 0s.
For instance, the text:
Led Zeppelin's next album, Houses of the Holy, was released in March 1973. It featured further
experimentation by the band, who expanded their use of synthesisers and mellotron orchestration.
The predominately orange album cover, designed by the London-based design group Hipgnosis,
depicts images of nude children climbing the Giant's Causeway in Northern Ireland. Although the
children are not shown from the front, the cover was controversial at the time of the album's
release. As with the band's fourth album, neither their name nor the album title was printed on the
sleeve
is evaluated as [0.2527, 0.0329, 0.0329, 0.4615] and the corresponding output classs for training is
[1, 0, 0] for English.
All the examples containing (input, output) are shuffled so the network is not trained on the classes sequentially.
The feed forward mechanism is implemented in a separate function so I can test novel examples
easily.
Backpropogation:
The number of epochs is given as input to the back_propogation method. In each epoch, the
network is trained on all examples and the error is recorded. After each epoch, the neural network is
validated on the validation set to keep track of how well is the network leaning. Ideally we'd like to
see a decrease in the errors after each epoch.
I use sum of squared error as the metric for error.
To avoid overfitting, I use a threshold value of validation error and stop training once it drops below
the threshold.
I tried a lot of other things. I let it run for different values of alpha, epochs and number of hidden
layers before settling for one set of values.
I did three random restarts and chose the best out of the three models based on error in the validtion
set and test set accuracy.
I used matplotlib to help me with these decisions.
Testing:
Once training is done, the user is prompted for the novel input through standard input. The user
must enter a piece of text and hit Enter. The neural network returns what language it thinks the text
belongs to.
The user can also enter “test” at the input followed by three test files in the order English, Italian
and Dutch. The network classifies all these and prints a confusion matrix for the given test set. The
user can also enter “default” to test with files named en_test, it_test and nl_test included with the
code.
I let my model classify the text of 15 randomly selected novels, 5 from each language to give a total
of more than 11,000 examples and got an accuracy of 98.87%.
That is a sufficiently high accuracy for a test set that is 6 times larger than the training set.
I have included these big test files. They have “_big” in their names.
The accuracy of the model over the test set is also printed.

About

Neural Network that can classify English, Italian and Dutch

Resources

Stars

5 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

For testing just run
python LanguageDetect.py
and then, enter an example
The model is stored in the .pickle file.
For training, change the read_from_file_and_test flag in the main function to False and run.
Network Design
I chose to make my neural network with objects as I felt it lead to a cleaner albeit slower
implementation.
The class Neuron has activation and a list of weights. There as many weights in a neuron as there
are neurons in the previous layer. The feed_forward method abstracts the weighted sum and
sigmoid at the neurons. Every neuron can thus compute its own activation.
The class NeuralNetwork has the learning rate alpha, and two lists for hidden and output neurons.
My network has 5 input neurons, 5 hidden neurons and 3 output neurons.
Feature Selection
I decided to use the following features for the three languages when I evaluate the training data.
The number of words that end in vowels is a very strong indicator for Italian.
Presence of the bigram “ij” for Dutch. “ij” has a heavier weight in my features as it is an alphabet in
Dutch and almost never occurs in the others.
Number of words with more than 8 characters indicating Dutch as dutch has frequent long words.
The number of occurrences of the most frequent bigram “th” as part of the most frequent word
“the” is a good feature to pin down English. As is the bigram “ed” at the end of words for English in
the past tense as most Wikipedia articles are.
The raw feature counts are scaled to the length of the example to represent the fraction of the text
the feature occurs in giving me a feature list in the range of [0, 1].
Data gathering and data set construction.
I used the wikis for the three languages as instructed. But only about 25% of the data is Wikipedia.
The rest if from Project Gutenberg. I got three novels in text format for the three languages. A small
python script and some unix commands gave me my data set. The python script makes examples of
random word count between 10 and a 200 words.
I have a training file for each language. There is an example on each line. The test and validation
set were made in the same way. Each language thus has training, validation and testing sets.
Each training set has 600 examples for a total of 1800 examples for the three languages.
The validation set has 300 examples each for a total of 900 examples in the set for all three
languages.
The testing set is the same size as the validation set, but has different examples.
Training process:
All three training files are read and evaluated. The evaluation takes in a block of text and returns a
list of features that represent the text. This is used as input to the neural network. The output is a list
where the correct class is a 1 and the rest are 0s.
For instance, the text:
Led Zeppelin's next album, Houses of the Holy, was released in March 1973. It featured further
experimentation by the band, who expanded their use of synthesisers and mellotron orchestration.
The predominately orange album cover, designed by the London-based design group Hipgnosis,
depicts images of nude children climbing the Giant's Causeway in Northern Ireland. Although the
children are not shown from the front, the cover was controversial at the time of the album's
release. As with the band's fourth album, neither their name nor the album title was printed on the
sleeve
is evaluated as [0.2527, 0.0329, 0.0329, 0.4615] and the corresponding output classs for training is
[1, 0, 0] for English.
All the examples containing (input, output) are shuffled so the network is not trained on the classes sequentially.
The feed forward mechanism is implemented in a separate function so I can test novel examples
easily.
Backpropogation:
The number of epochs is given as input to the back_propogation method. In each epoch, the
network is trained on all examples and the error is recorded. After each epoch, the neural network is
validated on the validation set to keep track of how well is the network leaning. Ideally we'd like to
see a decrease in the errors after each epoch.
I use sum of squared error as the metric for error.
To avoid overfitting, I use a threshold value of validation error and stop training once it drops below
the threshold.
I tried a lot of other things. I let it run for different values of alpha, epochs and number of hidden
layers before settling for one set of values.
I did three random restarts and chose the best out of the three models based on error in the validtion
set and test set accuracy.
I used matplotlib to help me with these decisions.
Testing:
Once training is done, the user is prompted for the novel input through standard input. The user
must enter a piece of text and hit Enter. The neural network returns what language it thinks the text
belongs to.
The user can also enter “test” at the input followed by three test files in the order English, Italian
and Dutch. The network classifies all these and prints a confusion matrix for the given test set. The
user can also enter “default” to test with files named en_test, it_test and nl_test included with the
code.
I let my model classify the text of 15 randomly selected novels, 5 from each language to give a total
of more than 11,000 examples and got an accuracy of 98.87%.
That is a sufficiently high accuracy for a test set that is 6 times larger than the training set.
I have included these big test files. They have “_big” in their names.
The accuracy of the model over the test set is also printed.

About

Neural Network that can classify English, Italian and Dutch

Resources

Stars

5 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

For testing just run
python LanguageDetect.py
and then, enter an example
The model is stored in the .pickle file.
For training, change the read_from_file_and_test flag in the main function to False and run.
Network Design
I chose to make my neural network with objects as I felt it lead to a cleaner albeit slower
implementation.
The class Neuron has activation and a list of weights. There as many weights in a neuron as there
are neurons in the previous layer. The feed_forward method abstracts the weighted sum and
sigmoid at the neurons. Every neuron can thus compute its own activation.
The class NeuralNetwork has the learning rate alpha, and two lists for hidden and output neurons.
My network has 5 input neurons, 5 hidden neurons and 3 output neurons.
Feature Selection
I decided to use the following features for the three languages when I evaluate the training data.
The number of words that end in vowels is a very strong indicator for Italian.
Presence of the bigram “ij” for Dutch. “ij” has a heavier weight in my features as it is an alphabet in
Dutch and almost never occurs in the others.
Number of words with more than 8 characters indicating Dutch as dutch has frequent long words.
The number of occurrences of the most frequent bigram “th” as part of the most frequent word
“the” is a good feature to pin down English. As is the bigram “ed” at the end of words for English in
the past tense as most Wikipedia articles are.
The raw feature counts are scaled to the length of the example to represent the fraction of the text
the feature occurs in giving me a feature list in the range of [0, 1].
Data gathering and data set construction.
I used the wikis for the three languages as instructed. But only about 25% of the data is Wikipedia.
The rest if from Project Gutenberg. I got three novels in text format for the three languages. A small
python script and some unix commands gave me my data set. The python script makes examples of
random word count between 10 and a 200 words.
I have a training file for each language. There is an example on each line. The test and validation
set were made in the same way. Each language thus has training, validation and testing sets.
Each training set has 600 examples for a total of 1800 examples for the three languages.
The validation set has 300 examples each for a total of 900 examples in the set for all three
languages.
The testing set is the same size as the validation set, but has different examples.
Training process:
All three training files are read and evaluated. The evaluation takes in a block of text and returns a
list of features that represent the text. This is used as input to the neural network. The output is a list
where the correct class is a 1 and the rest are 0s.
For instance, the text:
Led Zeppelin's next album, Houses of the Holy, was released in March 1973. It featured further
experimentation by the band, who expanded their use of synthesisers and mellotron orchestration.
The predominately orange album cover, designed by the London-based design group Hipgnosis,
depicts images of nude children climbing the Giant's Causeway in Northern Ireland. Although the
children are not shown from the front, the cover was controversial at the time of the album's
release. As with the band's fourth album, neither their name nor the album title was printed on the
sleeve
is evaluated as [0.2527, 0.0329, 0.0329, 0.4615] and the corresponding output classs for training is
[1, 0, 0] for English.
All the examples containing (input, output) are shuffled so the network is not trained on the classes sequentially.
The feed forward mechanism is implemented in a separate function so I can test novel examples
easily.
Backpropogation:
The number of epochs is given as input to the back_propogation method. In each epoch, the
network is trained on all examples and the error is recorded. After each epoch, the neural network is
validated on the validation set to keep track of how well is the network leaning. Ideally we'd like to
see a decrease in the errors after each epoch.
I use sum of squared error as the metric for error.
To avoid overfitting, I use a threshold value of validation error and stop training once it drops below
the threshold.
I tried a lot of other things. I let it run for different values of alpha, epochs and number of hidden
layers before settling for one set of values.
I did three random restarts and chose the best out of the three models based on error in the validtion
set and test set accuracy.
I used matplotlib to help me with these decisions.
Testing:
Once training is done, the user is prompted for the novel input through standard input. The user
must enter a piece of text and hit Enter. The neural network returns what language it thinks the text
belongs to.
The user can also enter “test” at the input followed by three test files in the order English, Italian
and Dutch. The network classifies all these and prints a confusion matrix for the given test set. The
user can also enter “default” to test with files named en_test, it_test and nl_test included with the
code.
I let my model classify the text of 15 randomly selected novels, 5 from each language to give a total
of more than 11,000 examples and got an accuracy of 98.87%.
That is a sufficiently high accuracy for a test set that is 6 times larger than the training set.
I have included these big test files. They have “_big” in their names.
The accuracy of the model over the test set is also printed.

About

Neural Network that can classify English, Italian and Dutch

Resources

Stars

5 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

For testing just run
python LanguageDetect.py
and then, enter an example
The model is stored in the .pickle file.
For training, change the read_from_file_and_test flag in the main function to False and run.
Network Design
I chose to make my neural network with objects as I felt it lead to a cleaner albeit slower
implementation.
The class Neuron has activation and a list of weights. There as many weights in a neuron as there
are neurons in the previous layer. The feed_forward method abstracts the weighted sum and
sigmoid at the neurons. Every neuron can thus compute its own activation.
The class NeuralNetwork has the learning rate alpha, and two lists for hidden and output neurons.
My network has 5 input neurons, 5 hidden neurons and 3 output neurons.
Feature Selection
I decided to use the following features for the three languages when I evaluate the training data.
The number of words that end in vowels is a very strong indicator for Italian.
Presence of the bigram “ij” for Dutch. “ij” has a heavier weight in my features as it is an alphabet in
Dutch and almost never occurs in the others.
Number of words with more than 8 characters indicating Dutch as dutch has frequent long words.
The number of occurrences of the most frequent bigram “th” as part of the most frequent word
“the” is a good feature to pin down English. As is the bigram “ed” at the end of words for English in
the past tense as most Wikipedia articles are.
The raw feature counts are scaled to the length of the example to represent the fraction of the text
the feature occurs in giving me a feature list in the range of [0, 1].
Data gathering and data set construction.
I used the wikis for the three languages as instructed. But only about 25% of the data is Wikipedia.
The rest if from Project Gutenberg. I got three novels in text format for the three languages. A small
python script and some unix commands gave me my data set. The python script makes examples of
random word count between 10 and a 200 words.
I have a training file for each language. There is an example on each line. The test and validation
set were made in the same way. Each language thus has training, validation and testing sets.
Each training set has 600 examples for a total of 1800 examples for the three languages.
The validation set has 300 examples each for a total of 900 examples in the set for all three
languages.
The testing set is the same size as the validation set, but has different examples.
Training process:
All three training files are read and evaluated. The evaluation takes in a block of text and returns a
list of features that represent the text. This is used as input to the neural network. The output is a list
where the correct class is a 1 and the rest are 0s.
For instance, the text:
Led Zeppelin's next album, Houses of the Holy, was released in March 1973. It featured further
experimentation by the band, who expanded their use of synthesisers and mellotron orchestration.
The predominately orange album cover, designed by the London-based design group Hipgnosis,
depicts images of nude children climbing the Giant's Causeway in Northern Ireland. Although the
children are not shown from the front, the cover was controversial at the time of the album's
release. As with the band's fourth album, neither their name nor the album title was printed on the
sleeve
is evaluated as [0.2527, 0.0329, 0.0329, 0.4615] and the corresponding output classs for training is
[1, 0, 0] for English.
All the examples containing (input, output) are shuffled so the network is not trained on the classes sequentially.
The feed forward mechanism is implemented in a separate function so I can test novel examples
easily.
Backpropogation:
The number of epochs is given as input to the back_propogation method. In each epoch, the
network is trained on all examples and the error is recorded. After each epoch, the neural network is
validated on the validation set to keep track of how well is the network leaning. Ideally we'd like to
see a decrease in the errors after each epoch.
I use sum of squared error as the metric for error.
To avoid overfitting, I use a threshold value of validation error and stop training once it drops below
the threshold.
I tried a lot of other things. I let it run for different values of alpha, epochs and number of hidden
layers before settling for one set of values.
I did three random restarts and chose the best out of the three models based on error in the validtion
set and test set accuracy.
I used matplotlib to help me with these decisions.
Testing:
Once training is done, the user is prompted for the novel input through standard input. The user
must enter a piece of text and hit Enter. The neural network returns what language it thinks the text
belongs to.
The user can also enter “test” at the input followed by three test files in the order English, Italian
and Dutch. The network classifies all these and prints a confusion matrix for the given test set. The
user can also enter “default” to test with files named en_test, it_test and nl_test included with the
code.
I let my model classify the text of 15 randomly selected novels, 5 from each language to give a total
of more than 11,000 examples and got an accuracy of 98.87%.
That is a sufficiently high accuracy for a test set that is 6 times larger than the training set.
I have included these big test files. They have “_big” in their names.
The accuracy of the model over the test set is also printed.

About

Neural Network that can classify English, Italian and Dutch

Resources

Stars

5 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

For testing just run
python LanguageDetect.py
and then, enter an example
The model is stored in the .pickle file.
For training, change the read_from_file_and_test flag in the main function to False and run.
Network Design
I chose to make my neural network with objects as I felt it lead to a cleaner albeit slower
implementation.
The class Neuron has activation and a list of weights. There as many weights in a neuron as there
are neurons in the previous layer. The feed_forward method abstracts the weighted sum and
sigmoid at the neurons. Every neuron can thus compute its own activation.
The class NeuralNetwork has the learning rate alpha, and two lists for hidden and output neurons.
My network has 5 input neurons, 5 hidden neurons and 3 output neurons.
Feature Selection
I decided to use the following features for the three languages when I evaluate the training data.
The number of words that end in vowels is a very strong indicator for Italian.
Presence of the bigram “ij” for Dutch. “ij” has a heavier weight in my features as it is an alphabet in
Dutch and almost never occurs in the others.
Number of words with more than 8 characters indicating Dutch as dutch has frequent long words.
The number of occurrences of the most frequent bigram “th” as part of the most frequent word
“the” is a good feature to pin down English. As is the bigram “ed” at the end of words for English in
the past tense as most Wikipedia articles are.
The raw feature counts are scaled to the length of the example to represent the fraction of the text
the feature occurs in giving me a feature list in the range of [0, 1].
Data gathering and data set construction.
I used the wikis for the three languages as instructed. But only about 25% of the data is Wikipedia.
The rest if from Project Gutenberg. I got three novels in text format for the three languages. A small
python script and some unix commands gave me my data set. The python script makes examples of
random word count between 10 and a 200 words.
I have a training file for each language. There is an example on each line. The test and validation
set were made in the same way. Each language thus has training, validation and testing sets.
Each training set has 600 examples for a total of 1800 examples for the three languages.
The validation set has 300 examples each for a total of 900 examples in the set for all three
languages.
The testing set is the same size as the validation set, but has different examples.
Training process:
All three training files are read and evaluated. The evaluation takes in a block of text and returns a
list of features that represent the text. This is used as input to the neural network. The output is a list
where the correct class is a 1 and the rest are 0s.
For instance, the text:
Led Zeppelin's next album, Houses of the Holy, was released in March 1973. It featured further
experimentation by the band, who expanded their use of synthesisers and mellotron orchestration.
The predominately orange album cover, designed by the London-based design group Hipgnosis,
depicts images of nude children climbing the Giant's Causeway in Northern Ireland. Although the
children are not shown from the front, the cover was controversial at the time of the album's
release. As with the band's fourth album, neither their name nor the album title was printed on the
sleeve
is evaluated as [0.2527, 0.0329, 0.0329, 0.4615] and the corresponding output classs for training is
[1, 0, 0] for English.
All the examples containing (input, output) are shuffled so the network is not trained on the classes sequentially.
The feed forward mechanism is implemented in a separate function so I can test novel examples
easily.
Backpropogation:
The number of epochs is given as input to the back_propogation method. In each epoch, the
network is trained on all examples and the error is recorded. After each epoch, the neural network is
validated on the validation set to keep track of how well is the network leaning. Ideally we'd like to
see a decrease in the errors after each epoch.
I use sum of squared error as the metric for error.
To avoid overfitting, I use a threshold value of validation error and stop training once it drops below
the threshold.
I tried a lot of other things. I let it run for different values of alpha, epochs and number of hidden
layers before settling for one set of values.
I did three random restarts and chose the best out of the three models based on error in the validtion
set and test set accuracy.
I used matplotlib to help me with these decisions.
Testing:
Once training is done, the user is prompted for the novel input through standard input. The user
must enter a piece of text and hit Enter. The neural network returns what language it thinks the text
belongs to.
The user can also enter “test” at the input followed by three test files in the order English, Italian
and Dutch. The network classifies all these and prints a confusion matrix for the given test set. The
user can also enter “default” to test with files named en_test, it_test and nl_test included with the
code.
I let my model classify the text of 15 randomly selected novels, 5 from each language to give a total
of more than 11,000 examples and got an accuracy of 98.87%.
That is a sufficiently high accuracy for a test set that is 6 times larger than the training set.
I have included these big test files. They have “_big” in their names.
The accuracy of the model over the test set is also printed.

About

Neural Network that can classify English, Italian and Dutch

Resources

Stars

5 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

For testing just run
python LanguageDetect.py
and then, enter an example
The model is stored in the .pickle file.
For training, change the read_from_file_and_test flag in the main function to False and run.
Network Design
I chose to make my neural network with objects as I felt it lead to a cleaner albeit slower
implementation.
The class Neuron has activation and a list of weights. There as many weights in a neuron as there
are neurons in the previous layer. The feed_forward method abstracts the weighted sum and
sigmoid at the neurons. Every neuron can thus compute its own activation.
The class NeuralNetwork has the learning rate alpha, and two lists for hidden and output neurons.
My network has 5 input neurons, 5 hidden neurons and 3 output neurons.
Feature Selection
I decided to use the following features for the three languages when I evaluate the training data.
The number of words that end in vowels is a very strong indicator for Italian.
Presence of the bigram “ij” for Dutch. “ij” has a heavier weight in my features as it is an alphabet in
Dutch and almost never occurs in the others.
Number of words with more than 8 characters indicating Dutch as dutch has frequent long words.
The number of occurrences of the most frequent bigram “th” as part of the most frequent word
“the” is a good feature to pin down English. As is the bigram “ed” at the end of words for English in
the past tense as most Wikipedia articles are.
The raw feature counts are scaled to the length of the example to represent the fraction of the text
the feature occurs in giving me a feature list in the range of [0, 1].
Data gathering and data set construction.
I used the wikis for the three languages as instructed. But only about 25% of the data is Wikipedia.
The rest if from Project Gutenberg. I got three novels in text format for the three languages. A small
python script and some unix commands gave me my data set. The python script makes examples of
random word count between 10 and a 200 words.
I have a training file for each language. There is an example on each line. The test and validation
set were made in the same way. Each language thus has training, validation and testing sets.
Each training set has 600 examples for a total of 1800 examples for the three languages.
The validation set has 300 examples each for a total of 900 examples in the set for all three
languages.
The testing set is the same size as the validation set, but has different examples.
Training process:
All three training files are read and evaluated. The evaluation takes in a block of text and returns a
list of features that represent the text. This is used as input to the neural network. The output is a list
where the correct class is a 1 and the rest are 0s.
For instance, the text:
Led Zeppelin's next album, Houses of the Holy, was released in March 1973. It featured further
experimentation by the band, who expanded their use of synthesisers and mellotron orchestration.
The predominately orange album cover, designed by the London-based design group Hipgnosis,
depicts images of nude children climbing the Giant's Causeway in Northern Ireland. Although the
children are not shown from the front, the cover was controversial at the time of the album's
release. As with the band's fourth album, neither their name nor the album title was printed on the
sleeve
is evaluated as [0.2527, 0.0329, 0.0329, 0.4615] and the corresponding output classs for training is
[1, 0, 0] for English.
All the examples containing (input, output) are shuffled so the network is not trained on the classes sequentially.
The feed forward mechanism is implemented in a separate function so I can test novel examples
easily.
Backpropogation:
The number of epochs is given as input to the back_propogation method. In each epoch, the
network is trained on all examples and the error is recorded. After each epoch, the neural network is
validated on the validation set to keep track of how well is the network leaning. Ideally we'd like to
see a decrease in the errors after each epoch.
I use sum of squared error as the metric for error.
To avoid overfitting, I use a threshold value of validation error and stop training once it drops below
the threshold.
I tried a lot of other things. I let it run for different values of alpha, epochs and number of hidden
layers before settling for one set of values.
I did three random restarts and chose the best out of the three models based on error in the validtion
set and test set accuracy.
I used matplotlib to help me with these decisions.
Testing:
Once training is done, the user is prompted for the novel input through standard input. The user
must enter a piece of text and hit Enter. The neural network returns what language it thinks the text
belongs to.
The user can also enter “test” at the input followed by three test files in the order English, Italian
and Dutch. The network classifies all these and prints a confusion matrix for the given test set. The
user can also enter “default” to test with files named en_test, it_test and nl_test included with the
code.
I let my model classify the text of 15 randomly selected novels, 5 from each language to give a total
of more than 11,000 examples and got an accuracy of 98.87%.
That is a sufficiently high accuracy for a test set that is 6 times larger than the training set.
I have included these big test files. They have “_big” in their names.
The accuracy of the model over the test set is also printed.

About

Neural Network that can classify English, Italian and Dutch

Resources

Stars

5 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

For testing just run
python LanguageDetect.py
and then, enter an example
The model is stored in the .pickle file.
For training, change the read_from_file_and_test flag in the main function to False and run.
Network Design
I chose to make my neural network with objects as I felt it lead to a cleaner albeit slower
implementation.
The class Neuron has activation and a list of weights. There as many weights in a neuron as there
are neurons in the previous layer. The feed_forward method abstracts the weighted sum and
sigmoid at the neurons. Every neuron can thus compute its own activation.
The class NeuralNetwork has the learning rate alpha, and two lists for hidden and output neurons.
My network has 5 input neurons, 5 hidden neurons and 3 output neurons.
Feature Selection
I decided to use the following features for the three languages when I evaluate the training data.
The number of words that end in vowels is a very strong indicator for Italian.
Presence of the bigram “ij” for Dutch. “ij” has a heavier weight in my features as it is an alphabet in
Dutch and almost never occurs in the others.
Number of words with more than 8 characters indicating Dutch as dutch has frequent long words.
The number of occurrences of the most frequent bigram “th” as part of the most frequent word
“the” is a good feature to pin down English. As is the bigram “ed” at the end of words for English in
the past tense as most Wikipedia articles are.
The raw feature counts are scaled to the length of the example to represent the fraction of the text
the feature occurs in giving me a feature list in the range of [0, 1].
Data gathering and data set construction.
I used the wikis for the three languages as instructed. But only about 25% of the data is Wikipedia.
The rest if from Project Gutenberg. I got three novels in text format for the three languages. A small
python script and some unix commands gave me my data set. The python script makes examples of
random word count between 10 and a 200 words.
I have a training file for each language. There is an example on each line. The test and validation
set were made in the same way. Each language thus has training, validation and testing sets.
Each training set has 600 examples for a total of 1800 examples for the three languages.
The validation set has 300 examples each for a total of 900 examples in the set for all three
languages.
The testing set is the same size as the validation set, but has different examples.
Training process:
All three training files are read and evaluated. The evaluation takes in a block of text and returns a
list of features that represent the text. This is used as input to the neural network. The output is a list
where the correct class is a 1 and the rest are 0s.
For instance, the text:
Led Zeppelin's next album, Houses of the Holy, was released in March 1973. It featured further
experimentation by the band, who expanded their use of synthesisers and mellotron orchestration.
The predominately orange album cover, designed by the London-based design group Hipgnosis,
depicts images of nude children climbing the Giant's Causeway in Northern Ireland. Although the
children are not shown from the front, the cover was controversial at the time of the album's
release. As with the band's fourth album, neither their name nor the album title was printed on the
sleeve
is evaluated as [0.2527, 0.0329, 0.0329, 0.4615] and the corresponding output classs for training is
[1, 0, 0] for English.
All the examples containing (input, output) are shuffled so the network is not trained on the classes sequentially.
The feed forward mechanism is implemented in a separate function so I can test novel examples
easily.
Backpropogation:
The number of epochs is given as input to the back_propogation method. In each epoch, the
network is trained on all examples and the error is recorded. After each epoch, the neural network is
validated on the validation set to keep track of how well is the network leaning. Ideally we'd like to
see a decrease in the errors after each epoch.
I use sum of squared error as the metric for error.
To avoid overfitting, I use a threshold value of validation error and stop training once it drops below
the threshold.
I tried a lot of other things. I let it run for different values of alpha, epochs and number of hidden
layers before settling for one set of values.
I did three random restarts and chose the best out of the three models based on error in the validtion
set and test set accuracy.
I used matplotlib to help me with these decisions.
Testing:
Once training is done, the user is prompted for the novel input through standard input. The user
must enter a piece of text and hit Enter. The neural network returns what language it thinks the text
belongs to.
The user can also enter “test” at the input followed by three test files in the order English, Italian
and Dutch. The network classifies all these and prints a confusion matrix for the given test set. The
user can also enter “default” to test with files named en_test, it_test and nl_test included with the
code.
I let my model classify the text of 15 randomly selected novels, 5 from each language to give a total
of more than 11,000 examples and got an accuracy of 98.87%.
That is a sufficiently high accuracy for a test set that is 6 times larger than the training set.
I have included these big test files. They have “_big” in their names.
The accuracy of the model over the test set is also printed.

About

Neural Network that can classify English, Italian and Dutch

Resources

Stars

5 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages