Skip to content

Add perplexity loss algorithm - #10718

Closed
PoojanSmart wants to merge 0 commit into
TheAlgorithms:masterfrom
PoojanSmart:master
Closed

Add perplexity loss algorithm#10718
PoojanSmart wants to merge 0 commit into
TheAlgorithms:masterfrom
PoojanSmart:master

Conversation

@PoojanSmart

Copy link
Copy Markdown
Contributor

Describe your change:

  • Perplexity loss function which is used in NLP (Natural language processing) for finding accuracy of the model based on how much certain the model is on its predictions.
  • Add an algorithm?
  • Fix a bug or typo in an existing algorithm?
  • Add or change doctests? -- Note: Please avoid changing both code and tests in a single pull request.
  • Documentation change?

Checklist:

  • I have read CONTRIBUTING.md.
  • This pull request is all my own work -- I have not plagiarized.
  • I know that pull requests will not be merged if they fail the automated tests.
  • This PR only changes one algorithm file. To ease review, please open separate PRs for separate algorithms.
  • All new Python files are placed inside an existing directory.
  • All filenames are in all lowercase characters with no spaces or dashes.
  • All functions and variable names follow Python naming conventions.
  • All function parameters and return values are annotated with Python type hints.
  • All functions have doctests that pass the automated testing.
  • All new algorithms include at least one URL that points to Wikipedia or another similar explanation.
  • If this pull request resolves one or more open issues then the description above includes the issue number(s) with a closing keyword: "Fixes #ISSUE-NUMBER".

@algorithms-keeperalgorithms-keeperBot added awaiting reviews This PR is ready to be reviewed tests are failing Do not merge until tests pass labels Oct 20, 2023

@imSankoimSanko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Did the code passed the pre commit tests ??

@PoojanSmart

This comment was marked as outdated.

@PoojanSmart

Copy link
Copy Markdown
ContributorAuthor

Once #10713 is merged. This issue should get resolved.
@cclauss please review and merge #10713.

One more thing, in my case when I run pre-commit run --all-files --show-diff-on-failure, I am not getting the same types of errors as the log from git actions.

@algorithms-keeperalgorithms-keeperBot removed the tests are failing Do not merge until tests pass label Oct 20, 2023
Comment on lines +32 to +37
>>> y_pred = np.array( \
[[[0.28, 0.19, 0.21 , 0.15, 0.15], \
[0.24, 0.19, 0.09, 0.18, 0.27]], \
[[0.03, 0.26, 0.21, 0.18, 0.30], \
[0.28, 0.10, 0.33, 0.15, 0.12]]]\
)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

PEP8: Backslash line continuation should be avoided in Python.

Suggested change
>>>y_pred=np.array( \
[[[0.28, 0.19, 0.21 , 0.15, 0.15], \
[0.24, 0.19, 0.09, 0.18, 0.27]], \
[[0.03, 0.26, 0.21, 0.18, 0.30], \
[0.28, 0.10, 0.33, 0.15, 0.12]]]\
)
>>>y_pred=np.array(
...[[[0.28, 0.19, 0.21 , 0.15, 0.15],
...[0.24, 0.19, 0.09, 0.18, 0.27]],
...[[0.03, 0.26, 0.21, 0.18, 0.30],
... [0.28, 0.10, 0.33, 0.15, 0.12]]],
... )

@cclausscclauss self-assigned this Oct 20, 2023
@cclauss

Copy link
Copy Markdown
Member

pre-commit autoupdate

@cclausscclauss changed the title Adds perplexity loss algorithmAdd perplexity loss algorithmOct 20, 2023
@algorithms-keeperalgorithms-keeperBot added tests are failing Do not merge until tests pass and removed tests are failing Do not merge until tests pass labels Oct 20, 2023

@tianyizheng02tianyizheng02 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All loss function files were consolidated into machine_learning/loss_functions.py in #10737. Could you move your new code into that file?

Comment threadmachine_learning/loss_functions.py Outdated
Comment on lines +315 to +316
# Add small constant to avoid getting inf for log(0)
epsilon = 1e-7

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please make epsilon an optional function parameter so that users can change its value

Comment threadmachine_learning/loss_functions.py Outdated
Comment on lines +332 to +338
# Getting the matrix containing prediction for only true class
true_class_pred = np.sum(y_pred * filter_matrix, axis=2)

# Calculating perplexity for each sentence
perp_losses = np.exp(
np.negative(np.mean(np.log(true_class_pred + epsilon), axis=1))
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
# Getting the matrix containing prediction for only true class
true_class_pred=np.sum(y_pred*filter_matrix, axis=2)
# Calculating perplexity for each sentence
perp_losses=np.exp(
np.negative(np.mean(np.log(true_class_pred+epsilon), axis=1))
)
# Getting the matrix containing prediction for only true class
# Clip values to avoid log(0)
true_class_pred=np.sum(y_pred*filter_matrix, axis=2).clip(epsilon, 1)
# Calculating perplexity for each sentence
perp_losses=np.exp(np.negative(np.mean(np.log(true_class_pred), axis=1)))

You can use .clip() to restrict the range of the array's values instead of adding epsilon to every entry. This way, only the problematic values get changed while OK values remain the same. Note that this may change the value of your doctests.

@PoojanSmartPoojanSmart mentioned this pull request Oct 27, 2023
15 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting reviewsThis PR is ready to be reviewed

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@PoojanSmart@cclauss@tianyizheng02@imSanko
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Add perplexity loss algorithm by PoojanSmart · Pull Request #10718 · TheAlgorithms/Python · GitHub
Skip to content

Add perplexity loss algorithm - #10718

Closed
PoojanSmart wants to merge 0 commit into
TheAlgorithms:masterfrom
PoojanSmart:master
Closed

Add perplexity loss algorithm#10718
PoojanSmart wants to merge 0 commit into
TheAlgorithms:masterfrom
PoojanSmart:master

Conversation

@PoojanSmart

Copy link
Copy Markdown
Contributor

Describe your change:

  • Perplexity loss function which is used in NLP (Natural language processing) for finding accuracy of the model based on how much certain the model is on its predictions.
  • Add an algorithm?
  • Fix a bug or typo in an existing algorithm?
  • Add or change doctests? -- Note: Please avoid changing both code and tests in a single pull request.
  • Documentation change?

Checklist:

  • I have read CONTRIBUTING.md.
  • This pull request is all my own work -- I have not plagiarized.
  • I know that pull requests will not be merged if they fail the automated tests.
  • This PR only changes one algorithm file. To ease review, please open separate PRs for separate algorithms.
  • All new Python files are placed inside an existing directory.
  • All filenames are in all lowercase characters with no spaces or dashes.
  • All functions and variable names follow Python naming conventions.
  • All function parameters and return values are annotated with Python type hints.
  • All functions have doctests that pass the automated testing.
  • All new algorithms include at least one URL that points to Wikipedia or another similar explanation.
  • If this pull request resolves one or more open issues then the description above includes the issue number(s) with a closing keyword: "Fixes #ISSUE-NUMBER".

@algorithms-keeperalgorithms-keeperBot added awaiting reviews This PR is ready to be reviewed tests are failing Do not merge until tests pass labels Oct 20, 2023

@imSankoimSanko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Did the code passed the pre commit tests ??

@PoojanSmart

This comment was marked as outdated.

@PoojanSmart

Copy link
Copy Markdown
ContributorAuthor

Once #10713 is merged. This issue should get resolved.
@cclauss please review and merge #10713.

One more thing, in my case when I run pre-commit run --all-files --show-diff-on-failure, I am not getting the same types of errors as the log from git actions.

@algorithms-keeperalgorithms-keeperBot removed the tests are failing Do not merge until tests pass label Oct 20, 2023
Comment on lines +32 to +37
>>> y_pred = np.array( \
[[[0.28, 0.19, 0.21 , 0.15, 0.15], \
[0.24, 0.19, 0.09, 0.18, 0.27]], \
[[0.03, 0.26, 0.21, 0.18, 0.30], \
[0.28, 0.10, 0.33, 0.15, 0.12]]]\
)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

PEP8: Backslash line continuation should be avoided in Python.

Suggested change
>>>y_pred=np.array( \
[[[0.28, 0.19, 0.21 , 0.15, 0.15], \
[0.24, 0.19, 0.09, 0.18, 0.27]], \
[[0.03, 0.26, 0.21, 0.18, 0.30], \
[0.28, 0.10, 0.33, 0.15, 0.12]]]\
)
>>>y_pred=np.array(
...[[[0.28, 0.19, 0.21 , 0.15, 0.15],
...[0.24, 0.19, 0.09, 0.18, 0.27]],
...[[0.03, 0.26, 0.21, 0.18, 0.30],
... [0.28, 0.10, 0.33, 0.15, 0.12]]],
... )

@cclausscclauss self-assigned this Oct 20, 2023
@cclauss

Copy link
Copy Markdown
Member

pre-commit autoupdate

@cclausscclauss changed the title Adds perplexity loss algorithmAdd perplexity loss algorithmOct 20, 2023
@algorithms-keeperalgorithms-keeperBot added tests are failing Do not merge until tests pass and removed tests are failing Do not merge until tests pass labels Oct 20, 2023

@tianyizheng02tianyizheng02 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All loss function files were consolidated into machine_learning/loss_functions.py in #10737. Could you move your new code into that file?

Comment threadmachine_learning/loss_functions.py Outdated
Comment on lines +315 to +316
# Add small constant to avoid getting inf for log(0)
epsilon = 1e-7

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please make epsilon an optional function parameter so that users can change its value

Comment threadmachine_learning/loss_functions.py Outdated
Comment on lines +332 to +338
# Getting the matrix containing prediction for only true class
true_class_pred = np.sum(y_pred * filter_matrix, axis=2)

# Calculating perplexity for each sentence
perp_losses = np.exp(
np.negative(np.mean(np.log(true_class_pred + epsilon), axis=1))
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
# Getting the matrix containing prediction for only true class
true_class_pred=np.sum(y_pred*filter_matrix, axis=2)
# Calculating perplexity for each sentence
perp_losses=np.exp(
np.negative(np.mean(np.log(true_class_pred+epsilon), axis=1))
)
# Getting the matrix containing prediction for only true class
# Clip values to avoid log(0)
true_class_pred=np.sum(y_pred*filter_matrix, axis=2).clip(epsilon, 1)
# Calculating perplexity for each sentence
perp_losses=np.exp(np.negative(np.mean(np.log(true_class_pred), axis=1)))

You can use .clip() to restrict the range of the array's values instead of adding epsilon to every entry. This way, only the problematic values get changed while OK values remain the same. Note that this may change the value of your doctests.

@PoojanSmartPoojanSmart mentioned this pull request Oct 27, 2023
15 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting reviewsThis PR is ready to be reviewed

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@PoojanSmart@cclauss@tianyizheng02@imSanko
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Add perplexity loss algorithm by PoojanSmart · Pull Request #10718 · TheAlgorithms/Python · GitHub
Skip to content

Add perplexity loss algorithm - #10718

Closed
PoojanSmart wants to merge 0 commit into
TheAlgorithms:masterfrom
PoojanSmart:master
Closed

Add perplexity loss algorithm#10718
PoojanSmart wants to merge 0 commit into
TheAlgorithms:masterfrom
PoojanSmart:master

Conversation

@PoojanSmart

Copy link
Copy Markdown
Contributor

Describe your change:

  • Perplexity loss function which is used in NLP (Natural language processing) for finding accuracy of the model based on how much certain the model is on its predictions.
  • Add an algorithm?
  • Fix a bug or typo in an existing algorithm?
  • Add or change doctests? -- Note: Please avoid changing both code and tests in a single pull request.
  • Documentation change?

Checklist:

  • I have read CONTRIBUTING.md.
  • This pull request is all my own work -- I have not plagiarized.
  • I know that pull requests will not be merged if they fail the automated tests.
  • This PR only changes one algorithm file. To ease review, please open separate PRs for separate algorithms.
  • All new Python files are placed inside an existing directory.
  • All filenames are in all lowercase characters with no spaces or dashes.
  • All functions and variable names follow Python naming conventions.
  • All function parameters and return values are annotated with Python type hints.
  • All functions have doctests that pass the automated testing.
  • All new algorithms include at least one URL that points to Wikipedia or another similar explanation.
  • If this pull request resolves one or more open issues then the description above includes the issue number(s) with a closing keyword: "Fixes #ISSUE-NUMBER".

@algorithms-keeperalgorithms-keeperBot added awaiting reviews This PR is ready to be reviewed tests are failing Do not merge until tests pass labels Oct 20, 2023

@imSankoimSanko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Did the code passed the pre commit tests ??

@PoojanSmart

This comment was marked as outdated.

@PoojanSmart

Copy link
Copy Markdown
ContributorAuthor

Once #10713 is merged. This issue should get resolved.
@cclauss please review and merge #10713.

One more thing, in my case when I run pre-commit run --all-files --show-diff-on-failure, I am not getting the same types of errors as the log from git actions.

@algorithms-keeperalgorithms-keeperBot removed the tests are failing Do not merge until tests pass label Oct 20, 2023
Comment on lines +32 to +37
>>> y_pred = np.array( \
[[[0.28, 0.19, 0.21 , 0.15, 0.15], \
[0.24, 0.19, 0.09, 0.18, 0.27]], \
[[0.03, 0.26, 0.21, 0.18, 0.30], \
[0.28, 0.10, 0.33, 0.15, 0.12]]]\
)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

PEP8: Backslash line continuation should be avoided in Python.

Suggested change
>>>y_pred=np.array( \
[[[0.28, 0.19, 0.21 , 0.15, 0.15], \
[0.24, 0.19, 0.09, 0.18, 0.27]], \
[[0.03, 0.26, 0.21, 0.18, 0.30], \
[0.28, 0.10, 0.33, 0.15, 0.12]]]\
)
>>>y_pred=np.array(
...[[[0.28, 0.19, 0.21 , 0.15, 0.15],
...[0.24, 0.19, 0.09, 0.18, 0.27]],
...[[0.03, 0.26, 0.21, 0.18, 0.30],
... [0.28, 0.10, 0.33, 0.15, 0.12]]],
... )

@cclausscclauss self-assigned this Oct 20, 2023
@cclauss

Copy link
Copy Markdown
Member

pre-commit autoupdate

@cclausscclauss changed the title Adds perplexity loss algorithmAdd perplexity loss algorithmOct 20, 2023
@algorithms-keeperalgorithms-keeperBot added tests are failing Do not merge until tests pass and removed tests are failing Do not merge until tests pass labels Oct 20, 2023

@tianyizheng02tianyizheng02 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All loss function files were consolidated into machine_learning/loss_functions.py in #10737. Could you move your new code into that file?

Comment threadmachine_learning/loss_functions.py Outdated
Comment on lines +315 to +316
# Add small constant to avoid getting inf for log(0)
epsilon = 1e-7

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please make epsilon an optional function parameter so that users can change its value

Comment threadmachine_learning/loss_functions.py Outdated
Comment on lines +332 to +338
# Getting the matrix containing prediction for only true class
true_class_pred = np.sum(y_pred * filter_matrix, axis=2)

# Calculating perplexity for each sentence
perp_losses = np.exp(
np.negative(np.mean(np.log(true_class_pred + epsilon), axis=1))
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
# Getting the matrix containing prediction for only true class
true_class_pred=np.sum(y_pred*filter_matrix, axis=2)
# Calculating perplexity for each sentence
perp_losses=np.exp(
np.negative(np.mean(np.log(true_class_pred+epsilon), axis=1))
)
# Getting the matrix containing prediction for only true class
# Clip values to avoid log(0)
true_class_pred=np.sum(y_pred*filter_matrix, axis=2).clip(epsilon, 1)
# Calculating perplexity for each sentence
perp_losses=np.exp(np.negative(np.mean(np.log(true_class_pred), axis=1)))

You can use .clip() to restrict the range of the array's values instead of adding epsilon to every entry. This way, only the problematic values get changed while OK values remain the same. Note that this may change the value of your doctests.

@PoojanSmartPoojanSmart mentioned this pull request Oct 27, 2023
15 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting reviewsThis PR is ready to be reviewed

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@PoojanSmart@cclauss@tianyizheng02@imSanko
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Add perplexity loss algorithm by PoojanSmart · Pull Request #10718 · TheAlgorithms/Python · GitHub
Skip to content

Add perplexity loss algorithm - #10718

Closed
PoojanSmart wants to merge 0 commit into
TheAlgorithms:masterfrom
PoojanSmart:master
Closed

Add perplexity loss algorithm#10718
PoojanSmart wants to merge 0 commit into
TheAlgorithms:masterfrom
PoojanSmart:master

Conversation

@PoojanSmart

Copy link
Copy Markdown
Contributor

Describe your change:

  • Perplexity loss function which is used in NLP (Natural language processing) for finding accuracy of the model based on how much certain the model is on its predictions.
  • Add an algorithm?
  • Fix a bug or typo in an existing algorithm?
  • Add or change doctests? -- Note: Please avoid changing both code and tests in a single pull request.
  • Documentation change?

Checklist:

  • I have read CONTRIBUTING.md.
  • This pull request is all my own work -- I have not plagiarized.
  • I know that pull requests will not be merged if they fail the automated tests.
  • This PR only changes one algorithm file. To ease review, please open separate PRs for separate algorithms.
  • All new Python files are placed inside an existing directory.
  • All filenames are in all lowercase characters with no spaces or dashes.
  • All functions and variable names follow Python naming conventions.
  • All function parameters and return values are annotated with Python type hints.
  • All functions have doctests that pass the automated testing.
  • All new algorithms include at least one URL that points to Wikipedia or another similar explanation.
  • If this pull request resolves one or more open issues then the description above includes the issue number(s) with a closing keyword: "Fixes #ISSUE-NUMBER".

@algorithms-keeperalgorithms-keeperBot added awaiting reviews This PR is ready to be reviewed tests are failing Do not merge until tests pass labels Oct 20, 2023

@imSankoimSanko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Did the code passed the pre commit tests ??

@PoojanSmart

This comment was marked as outdated.

@PoojanSmart

Copy link
Copy Markdown
ContributorAuthor

Once #10713 is merged. This issue should get resolved.
@cclauss please review and merge #10713.

One more thing, in my case when I run pre-commit run --all-files --show-diff-on-failure, I am not getting the same types of errors as the log from git actions.

@algorithms-keeperalgorithms-keeperBot removed the tests are failing Do not merge until tests pass label Oct 20, 2023
Comment on lines +32 to +37
>>> y_pred = np.array( \
[[[0.28, 0.19, 0.21 , 0.15, 0.15], \
[0.24, 0.19, 0.09, 0.18, 0.27]], \
[[0.03, 0.26, 0.21, 0.18, 0.30], \
[0.28, 0.10, 0.33, 0.15, 0.12]]]\
)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

PEP8: Backslash line continuation should be avoided in Python.

Suggested change
>>>y_pred=np.array( \
[[[0.28, 0.19, 0.21 , 0.15, 0.15], \
[0.24, 0.19, 0.09, 0.18, 0.27]], \
[[0.03, 0.26, 0.21, 0.18, 0.30], \
[0.28, 0.10, 0.33, 0.15, 0.12]]]\
)
>>>y_pred=np.array(
...[[[0.28, 0.19, 0.21 , 0.15, 0.15],
...[0.24, 0.19, 0.09, 0.18, 0.27]],
...[[0.03, 0.26, 0.21, 0.18, 0.30],
... [0.28, 0.10, 0.33, 0.15, 0.12]]],
... )

@cclausscclauss self-assigned this Oct 20, 2023
@cclauss

Copy link
Copy Markdown
Member

pre-commit autoupdate

@cclausscclauss changed the title Adds perplexity loss algorithmAdd perplexity loss algorithmOct 20, 2023
@algorithms-keeperalgorithms-keeperBot added tests are failing Do not merge until tests pass and removed tests are failing Do not merge until tests pass labels Oct 20, 2023

@tianyizheng02tianyizheng02 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All loss function files were consolidated into machine_learning/loss_functions.py in #10737. Could you move your new code into that file?

Comment threadmachine_learning/loss_functions.py Outdated
Comment on lines +315 to +316
# Add small constant to avoid getting inf for log(0)
epsilon = 1e-7

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please make epsilon an optional function parameter so that users can change its value

Comment threadmachine_learning/loss_functions.py Outdated
Comment on lines +332 to +338
# Getting the matrix containing prediction for only true class
true_class_pred = np.sum(y_pred * filter_matrix, axis=2)

# Calculating perplexity for each sentence
perp_losses = np.exp(
np.negative(np.mean(np.log(true_class_pred + epsilon), axis=1))
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
# Getting the matrix containing prediction for only true class
true_class_pred=np.sum(y_pred*filter_matrix, axis=2)
# Calculating perplexity for each sentence
perp_losses=np.exp(
np.negative(np.mean(np.log(true_class_pred+epsilon), axis=1))
)
# Getting the matrix containing prediction for only true class
# Clip values to avoid log(0)
true_class_pred=np.sum(y_pred*filter_matrix, axis=2).clip(epsilon, 1)
# Calculating perplexity for each sentence
perp_losses=np.exp(np.negative(np.mean(np.log(true_class_pred), axis=1)))

You can use .clip() to restrict the range of the array's values instead of adding epsilon to every entry. This way, only the problematic values get changed while OK values remain the same. Note that this may change the value of your doctests.

@PoojanSmartPoojanSmart mentioned this pull request Oct 27, 2023
15 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting reviewsThis PR is ready to be reviewed

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@PoojanSmart@cclauss@tianyizheng02@imSanko
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' Add perplexity loss algorithm by PoojanSmart · Pull Request #10718 · TheAlgorithms/Python · GitHub
Skip to content

Add perplexity loss algorithm - #10718

Closed
PoojanSmart wants to merge 0 commit into
TheAlgorithms:masterfrom
PoojanSmart:master
Closed

Add perplexity loss algorithm#10718
PoojanSmart wants to merge 0 commit into
TheAlgorithms:masterfrom
PoojanSmart:master

Conversation

@PoojanSmart

Copy link
Copy Markdown
Contributor

Describe your change:

  • Perplexity loss function which is used in NLP (Natural language processing) for finding accuracy of the model based on how much certain the model is on its predictions.
  • Add an algorithm?
  • Fix a bug or typo in an existing algorithm?
  • Add or change doctests? -- Note: Please avoid changing both code and tests in a single pull request.
  • Documentation change?

Checklist:

  • I have read CONTRIBUTING.md.
  • This pull request is all my own work -- I have not plagiarized.
  • I know that pull requests will not be merged if they fail the automated tests.
  • This PR only changes one algorithm file. To ease review, please open separate PRs for separate algorithms.
  • All new Python files are placed inside an existing directory.
  • All filenames are in all lowercase characters with no spaces or dashes.
  • All functions and variable names follow Python naming conventions.
  • All function parameters and return values are annotated with Python type hints.
  • All functions have doctests that pass the automated testing.
  • All new algorithms include at least one URL that points to Wikipedia or another similar explanation.
  • If this pull request resolves one or more open issues then the description above includes the issue number(s) with a closing keyword: "Fixes #ISSUE-NUMBER".

@algorithms-keeperalgorithms-keeperBot added awaiting reviews This PR is ready to be reviewed tests are failing Do not merge until tests pass labels Oct 20, 2023

@imSankoimSanko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Did the code passed the pre commit tests ??

@PoojanSmart

This comment was marked as outdated.

@PoojanSmart

Copy link
Copy Markdown
ContributorAuthor

Once #10713 is merged. This issue should get resolved.
@cclauss please review and merge #10713.

One more thing, in my case when I run pre-commit run --all-files --show-diff-on-failure, I am not getting the same types of errors as the log from git actions.

@algorithms-keeperalgorithms-keeperBot removed the tests are failing Do not merge until tests pass label Oct 20, 2023
Comment on lines +32 to +37
>>> y_pred = np.array( \
[[[0.28, 0.19, 0.21 , 0.15, 0.15], \
[0.24, 0.19, 0.09, 0.18, 0.27]], \
[[0.03, 0.26, 0.21, 0.18, 0.30], \
[0.28, 0.10, 0.33, 0.15, 0.12]]]\
)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

PEP8: Backslash line continuation should be avoided in Python.

Suggested change
>>>y_pred=np.array( \
[[[0.28, 0.19, 0.21 , 0.15, 0.15], \
[0.24, 0.19, 0.09, 0.18, 0.27]], \
[[0.03, 0.26, 0.21, 0.18, 0.30], \
[0.28, 0.10, 0.33, 0.15, 0.12]]]\
)
>>>y_pred=np.array(
...[[[0.28, 0.19, 0.21 , 0.15, 0.15],
...[0.24, 0.19, 0.09, 0.18, 0.27]],
...[[0.03, 0.26, 0.21, 0.18, 0.30],
... [0.28, 0.10, 0.33, 0.15, 0.12]]],
... )

@cclausscclauss self-assigned this Oct 20, 2023
@cclauss

Copy link
Copy Markdown
Member

pre-commit autoupdate

@cclausscclauss changed the title Adds perplexity loss algorithmAdd perplexity loss algorithmOct 20, 2023
@algorithms-keeperalgorithms-keeperBot added tests are failing Do not merge until tests pass and removed tests are failing Do not merge until tests pass labels Oct 20, 2023

@tianyizheng02tianyizheng02 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All loss function files were consolidated into machine_learning/loss_functions.py in #10737. Could you move your new code into that file?

Comment threadmachine_learning/loss_functions.py Outdated
Comment on lines +315 to +316
# Add small constant to avoid getting inf for log(0)
epsilon = 1e-7

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please make epsilon an optional function parameter so that users can change its value

Comment threadmachine_learning/loss_functions.py Outdated
Comment on lines +332 to +338
# Getting the matrix containing prediction for only true class
true_class_pred = np.sum(y_pred * filter_matrix, axis=2)

# Calculating perplexity for each sentence
perp_losses = np.exp(
np.negative(np.mean(np.log(true_class_pred + epsilon), axis=1))
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
# Getting the matrix containing prediction for only true class
true_class_pred=np.sum(y_pred*filter_matrix, axis=2)
# Calculating perplexity for each sentence
perp_losses=np.exp(
np.negative(np.mean(np.log(true_class_pred+epsilon), axis=1))
)
# Getting the matrix containing prediction for only true class
# Clip values to avoid log(0)
true_class_pred=np.sum(y_pred*filter_matrix, axis=2).clip(epsilon, 1)
# Calculating perplexity for each sentence
perp_losses=np.exp(np.negative(np.mean(np.log(true_class_pred), axis=1)))

You can use .clip() to restrict the range of the array's values instead of adding epsilon to every entry. This way, only the problematic values get changed while OK values remain the same. Note that this may change the value of your doctests.

@PoojanSmartPoojanSmart mentioned this pull request Oct 27, 2023
15 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting reviewsThis PR is ready to be reviewed

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@PoojanSmart@cclauss@tianyizheng02@imSanko
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Add perplexity loss algorithm by PoojanSmart · Pull Request #10718 · TheAlgorithms/Python · GitHub
Skip to content

Add perplexity loss algorithm - #10718

Closed
PoojanSmart wants to merge 0 commit into
TheAlgorithms:masterfrom
PoojanSmart:master
Closed

Add perplexity loss algorithm#10718
PoojanSmart wants to merge 0 commit into
TheAlgorithms:masterfrom
PoojanSmart:master

Conversation

@PoojanSmart

Copy link
Copy Markdown
Contributor

Describe your change:

  • Perplexity loss function which is used in NLP (Natural language processing) for finding accuracy of the model based on how much certain the model is on its predictions.
  • Add an algorithm?
  • Fix a bug or typo in an existing algorithm?
  • Add or change doctests? -- Note: Please avoid changing both code and tests in a single pull request.
  • Documentation change?

Checklist:

  • I have read CONTRIBUTING.md.
  • This pull request is all my own work -- I have not plagiarized.
  • I know that pull requests will not be merged if they fail the automated tests.
  • This PR only changes one algorithm file. To ease review, please open separate PRs for separate algorithms.
  • All new Python files are placed inside an existing directory.
  • All filenames are in all lowercase characters with no spaces or dashes.
  • All functions and variable names follow Python naming conventions.
  • All function parameters and return values are annotated with Python type hints.
  • All functions have doctests that pass the automated testing.
  • All new algorithms include at least one URL that points to Wikipedia or another similar explanation.
  • If this pull request resolves one or more open issues then the description above includes the issue number(s) with a closing keyword: "Fixes #ISSUE-NUMBER".

@algorithms-keeperalgorithms-keeperBot added awaiting reviews This PR is ready to be reviewed tests are failing Do not merge until tests pass labels Oct 20, 2023

@imSankoimSanko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Did the code passed the pre commit tests ??

@PoojanSmart

This comment was marked as outdated.

@PoojanSmart

Copy link
Copy Markdown
ContributorAuthor

Once #10713 is merged. This issue should get resolved.
@cclauss please review and merge #10713.

One more thing, in my case when I run pre-commit run --all-files --show-diff-on-failure, I am not getting the same types of errors as the log from git actions.

@algorithms-keeperalgorithms-keeperBot removed the tests are failing Do not merge until tests pass label Oct 20, 2023
Comment on lines +32 to +37
>>> y_pred = np.array( \
[[[0.28, 0.19, 0.21 , 0.15, 0.15], \
[0.24, 0.19, 0.09, 0.18, 0.27]], \
[[0.03, 0.26, 0.21, 0.18, 0.30], \
[0.28, 0.10, 0.33, 0.15, 0.12]]]\
)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

PEP8: Backslash line continuation should be avoided in Python.

Suggested change
>>>y_pred=np.array( \
[[[0.28, 0.19, 0.21 , 0.15, 0.15], \
[0.24, 0.19, 0.09, 0.18, 0.27]], \
[[0.03, 0.26, 0.21, 0.18, 0.30], \
[0.28, 0.10, 0.33, 0.15, 0.12]]]\
)
>>>y_pred=np.array(
...[[[0.28, 0.19, 0.21 , 0.15, 0.15],
...[0.24, 0.19, 0.09, 0.18, 0.27]],
...[[0.03, 0.26, 0.21, 0.18, 0.30],
... [0.28, 0.10, 0.33, 0.15, 0.12]]],
... )

@cclausscclauss self-assigned this Oct 20, 2023
@cclauss

Copy link
Copy Markdown
Member

pre-commit autoupdate

@cclausscclauss changed the title Adds perplexity loss algorithmAdd perplexity loss algorithmOct 20, 2023
@algorithms-keeperalgorithms-keeperBot added tests are failing Do not merge until tests pass and removed tests are failing Do not merge until tests pass labels Oct 20, 2023

@tianyizheng02tianyizheng02 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All loss function files were consolidated into machine_learning/loss_functions.py in #10737. Could you move your new code into that file?

Comment threadmachine_learning/loss_functions.py Outdated
Comment on lines +315 to +316
# Add small constant to avoid getting inf for log(0)
epsilon = 1e-7

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please make epsilon an optional function parameter so that users can change its value

Comment threadmachine_learning/loss_functions.py Outdated
Comment on lines +332 to +338
# Getting the matrix containing prediction for only true class
true_class_pred = np.sum(y_pred * filter_matrix, axis=2)

# Calculating perplexity for each sentence
perp_losses = np.exp(
np.negative(np.mean(np.log(true_class_pred + epsilon), axis=1))
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
# Getting the matrix containing prediction for only true class
true_class_pred=np.sum(y_pred*filter_matrix, axis=2)
# Calculating perplexity for each sentence
perp_losses=np.exp(
np.negative(np.mean(np.log(true_class_pred+epsilon), axis=1))
)
# Getting the matrix containing prediction for only true class
# Clip values to avoid log(0)
true_class_pred=np.sum(y_pred*filter_matrix, axis=2).clip(epsilon, 1)
# Calculating perplexity for each sentence
perp_losses=np.exp(np.negative(np.mean(np.log(true_class_pred), axis=1)))

You can use .clip() to restrict the range of the array's values instead of adding epsilon to every entry. This way, only the problematic values get changed while OK values remain the same. Note that this may change the value of your doctests.

@PoojanSmartPoojanSmart mentioned this pull request Oct 27, 2023
15 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting reviewsThis PR is ready to be reviewed

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@PoojanSmart@cclauss@tianyizheng02@imSanko
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); })(); Add perplexity loss algorithm by PoojanSmart · Pull Request #10718 · TheAlgorithms/Python · GitHub
Skip to content

Add perplexity loss algorithm - #10718

Closed
PoojanSmart wants to merge 0 commit into
TheAlgorithms:masterfrom
PoojanSmart:master
Closed

Add perplexity loss algorithm#10718
PoojanSmart wants to merge 0 commit into
TheAlgorithms:masterfrom
PoojanSmart:master

Conversation

@PoojanSmart

Copy link
Copy Markdown
Contributor

Describe your change:

  • Perplexity loss function which is used in NLP (Natural language processing) for finding accuracy of the model based on how much certain the model is on its predictions.
  • Add an algorithm?
  • Fix a bug or typo in an existing algorithm?
  • Add or change doctests? -- Note: Please avoid changing both code and tests in a single pull request.
  • Documentation change?

Checklist:

  • I have read CONTRIBUTING.md.
  • This pull request is all my own work -- I have not plagiarized.
  • I know that pull requests will not be merged if they fail the automated tests.
  • This PR only changes one algorithm file. To ease review, please open separate PRs for separate algorithms.
  • All new Python files are placed inside an existing directory.
  • All filenames are in all lowercase characters with no spaces or dashes.
  • All functions and variable names follow Python naming conventions.
  • All function parameters and return values are annotated with Python type hints.
  • All functions have doctests that pass the automated testing.
  • All new algorithms include at least one URL that points to Wikipedia or another similar explanation.
  • If this pull request resolves one or more open issues then the description above includes the issue number(s) with a closing keyword: "Fixes #ISSUE-NUMBER".

@algorithms-keeperalgorithms-keeperBot added awaiting reviews This PR is ready to be reviewed tests are failing Do not merge until tests pass labels Oct 20, 2023

@imSankoimSanko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Did the code passed the pre commit tests ??

@PoojanSmart

This comment was marked as outdated.

@PoojanSmart

Copy link
Copy Markdown
ContributorAuthor

Once #10713 is merged. This issue should get resolved.
@cclauss please review and merge #10713.

One more thing, in my case when I run pre-commit run --all-files --show-diff-on-failure, I am not getting the same types of errors as the log from git actions.

@algorithms-keeperalgorithms-keeperBot removed the tests are failing Do not merge until tests pass label Oct 20, 2023
Comment on lines +32 to +37
>>> y_pred = np.array( \
[[[0.28, 0.19, 0.21 , 0.15, 0.15], \
[0.24, 0.19, 0.09, 0.18, 0.27]], \
[[0.03, 0.26, 0.21, 0.18, 0.30], \
[0.28, 0.10, 0.33, 0.15, 0.12]]]\
)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

PEP8: Backslash line continuation should be avoided in Python.

Suggested change
>>>y_pred=np.array( \
[[[0.28, 0.19, 0.21 , 0.15, 0.15], \
[0.24, 0.19, 0.09, 0.18, 0.27]], \
[[0.03, 0.26, 0.21, 0.18, 0.30], \
[0.28, 0.10, 0.33, 0.15, 0.12]]]\
)
>>>y_pred=np.array(
...[[[0.28, 0.19, 0.21 , 0.15, 0.15],
...[0.24, 0.19, 0.09, 0.18, 0.27]],
...[[0.03, 0.26, 0.21, 0.18, 0.30],
... [0.28, 0.10, 0.33, 0.15, 0.12]]],
... )

@cclausscclauss self-assigned this Oct 20, 2023
@cclauss

Copy link
Copy Markdown
Member

pre-commit autoupdate

@cclausscclauss changed the title Adds perplexity loss algorithmAdd perplexity loss algorithmOct 20, 2023
@algorithms-keeperalgorithms-keeperBot added tests are failing Do not merge until tests pass and removed tests are failing Do not merge until tests pass labels Oct 20, 2023

@tianyizheng02tianyizheng02 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All loss function files were consolidated into machine_learning/loss_functions.py in #10737. Could you move your new code into that file?

Comment threadmachine_learning/loss_functions.py Outdated
Comment on lines +315 to +316
# Add small constant to avoid getting inf for log(0)
epsilon = 1e-7

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please make epsilon an optional function parameter so that users can change its value

Comment threadmachine_learning/loss_functions.py Outdated
Comment on lines +332 to +338
# Getting the matrix containing prediction for only true class
true_class_pred = np.sum(y_pred * filter_matrix, axis=2)

# Calculating perplexity for each sentence
perp_losses = np.exp(
np.negative(np.mean(np.log(true_class_pred + epsilon), axis=1))
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
# Getting the matrix containing prediction for only true class
true_class_pred=np.sum(y_pred*filter_matrix, axis=2)
# Calculating perplexity for each sentence
perp_losses=np.exp(
np.negative(np.mean(np.log(true_class_pred+epsilon), axis=1))
)
# Getting the matrix containing prediction for only true class
# Clip values to avoid log(0)
true_class_pred=np.sum(y_pred*filter_matrix, axis=2).clip(epsilon, 1)
# Calculating perplexity for each sentence
perp_losses=np.exp(np.negative(np.mean(np.log(true_class_pred), axis=1)))

You can use .clip() to restrict the range of the array's values instead of adding epsilon to every entry. This way, only the problematic values get changed while OK values remain the same. Note that this may change the value of your doctests.

@PoojanSmartPoojanSmart mentioned this pull request Oct 27, 2023
15 tasks
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting reviewsThis PR is ready to be reviewed

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@PoojanSmart@cclauss@tianyizheng02@imSanko