Skip to content

Solving the Top k most frequent words problem using a max-heap - #8125

Closed
amos-syd wants to merge 0 commit into
TheAlgorithms:masterfrom
amos-syd:master
Closed

Solving the Top k most frequent words problem using a max-heap#8125
amos-syd wants to merge 0 commit into
TheAlgorithms:masterfrom
amos-syd:master

Conversation

@amos-syd

@amos-sydamos-syd commented Feb 7, 2023

Copy link
Copy Markdown
Contributor

Describe your change:

This PR aims to add an algorithm to identify the top k most frequent strings given a provided string list of elements.
To do this, the algorithm is using a max-heap implementation already existing in this repository (a generic type was introduced to allow the usage).

Time complexity is O(n), where n is the number of words:

  • O(n) for building the max-heap
  • k*O(logn) for extracting the k most frequent strings
  • Add an algorithm?
  • Fix a bug or typo in an existing algorithm?
  • Documentation change?

Checklist:

  • I have read CONTRIBUTING.md.
  • This pull request is all my own work -- I have not plagiarized.
  • I know that pull requests will not be merged if they fail the automated tests.
  • This PR only changes one algorithm file. To ease review, please open separate PRs for separate algorithms.
  • All new Python files are placed inside an existing directory.
  • All filenames are in all lowercase characters with no spaces or dashes.
  • All functions and variable names follow Python naming conventions.
  • All function parameters and return values are annotated with Python type hints.
  • All functions have doctests that pass the automated testing.
  • All new algorithms include at least one URL that points to Wikipedia or another similar explanation.
  • If this pull request resolves one or more open issues then the commit message contains Fixes: #{$ISSUE_NO}.

@algorithms-keeperalgorithms-keeperBot added awaiting reviews This PR is ready to be reviewed enhancement This PR modified some existing files require descriptive names This PR needs descriptive function and/or variable names require tests Tests [doctest/unittest/pytest] are required require type hints https://docs.python.org/3/library/typing.html labels Feb 7, 2023

@algorithms-keeperalgorithms-keeperBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Click here to look at the relevant links ⬇️

🔗 Relevant Links

Repository:

Python:

Automated review generated by algorithms-keeper. If there's any problem regarding this review, please open an issue about it.

algorithms-keeper commands and options

algorithms-keeper actions can be triggered by commenting on this PR:

  • @algorithms-keeper review to trigger the checks for only added pull request files
  • @algorithms-keeper review-all to trigger the checks for all the pull request files, including the modified files. As we cannot post review comments on lines not part of the diff, this command will post all the messages in one comment.

NOTE: Commands are in beta and so this feature is restricted only to a member or owner of the organization.

Comment threadstrings/top_k_frequent_words.py Outdated
Comment threadstrings/top_k_frequent_words.py Outdated
Comment threadstrings/top_k_frequent_words.py Outdated
Comment threadstrings/top_k_frequent_words.py Outdated
@algorithms-keeperalgorithms-keeperBot removed require descriptive names This PR needs descriptive function and/or variable names require tests Tests [doctest/unittest/pytest] are required require type hints https://docs.python.org/3/library/typing.html labels Feb 7, 2023

@CaedenPHCaedenPH left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very well written algorithm, great stuff!

@amos-syd

Copy link
Copy Markdown
ContributorAuthor

Hello @cclauss, I'm tagging you as I saw your involvement in several PRs.
Do you know who can I ask to for a review? Is there anything I should still change in the PR?
Thank you!

@cclauss

cclauss commented Apr 23, 2023

Copy link
Copy Markdown
Member

All of this seems like it could be done in just a few lines using only a collections.Counter so what are the advantages of this approach?

@amos-syd

Copy link
Copy Markdown
ContributorAuthor

All of this seems like it could be done in just a few lines using only a collections.Counter so what are the advantages of this approach?

Thanks for the feedback.
I had a look to collections.Counter and noticed that, indeed, most_common solves exactly this problem: https://github.com/python/cpython/blob/3.11/Lib/collections/__init__.py#L608

This is using heapq.nlargest behind the scenes: https://github.com/python/cpython/blob/3.11/Lib/heapq.py#L523

There is already a Heap class in this repository, I imagined it would be useful to show a typical usage of heaps (i.e. finding out order statistics). Do you think I should close this PR entirely, or do you have other suggestions?

Thank you

@cclauss

Copy link
Copy Markdown
Member
fromcollectionsimportCounterdeftop_k_frequent_words(words, k_value):
return [x[0] forxinCounter(words).most_common(k_value)]

@cclauss

cclauss commented Apr 23, 2023

Copy link
Copy Markdown
Member

Please put the text of our last two messages into the file’s docstring and then we can merge this. That will show why this work was done and that there also that the Python standard library provides a far more straightforward way to solve the same problem.

@amos-syd

Copy link
Copy Markdown
ContributorAuthor

Please put the text of our last two messages into the file’s docstring and then we can merge this. That will show why this work was done and that there also that the Python standard library provides a far more straightforward way to solve the same problem.

Thanks for the feedback.
I have updated the docstring mentioning the (preferable) Python standard library solution.

@algorithms-keeperalgorithms-keeperBot added the tests are failing Do not merge until tests pass label Apr 23, 2023

@cclausscclauss left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting reviewsThis PR is ready to be reviewedenhancementThis PR modified some existing filestests are failingDo not merge until tests pass

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@amos-syd@cclauss@CaedenPH
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Solving the `Top k most frequent words` problem using a max-heap by amos-syd · Pull Request #8125 · TheAlgorithms/Python · GitHub
Skip to content

Solving the Top k most frequent words problem using a max-heap - #8125

Closed
amos-syd wants to merge 0 commit into
TheAlgorithms:masterfrom
amos-syd:master
Closed

Solving the Top k most frequent words problem using a max-heap#8125
amos-syd wants to merge 0 commit into
TheAlgorithms:masterfrom
amos-syd:master

Conversation

@amos-syd

@amos-sydamos-syd commented Feb 7, 2023

Copy link
Copy Markdown
Contributor

Describe your change:

This PR aims to add an algorithm to identify the top k most frequent strings given a provided string list of elements.
To do this, the algorithm is using a max-heap implementation already existing in this repository (a generic type was introduced to allow the usage).

Time complexity is O(n), where n is the number of words:

  • O(n) for building the max-heap
  • k*O(logn) for extracting the k most frequent strings
  • Add an algorithm?
  • Fix a bug or typo in an existing algorithm?
  • Documentation change?

Checklist:

  • I have read CONTRIBUTING.md.
  • This pull request is all my own work -- I have not plagiarized.
  • I know that pull requests will not be merged if they fail the automated tests.
  • This PR only changes one algorithm file. To ease review, please open separate PRs for separate algorithms.
  • All new Python files are placed inside an existing directory.
  • All filenames are in all lowercase characters with no spaces or dashes.
  • All functions and variable names follow Python naming conventions.
  • All function parameters and return values are annotated with Python type hints.
  • All functions have doctests that pass the automated testing.
  • All new algorithms include at least one URL that points to Wikipedia or another similar explanation.
  • If this pull request resolves one or more open issues then the commit message contains Fixes: #{$ISSUE_NO}.

@algorithms-keeperalgorithms-keeperBot added awaiting reviews This PR is ready to be reviewed enhancement This PR modified some existing files require descriptive names This PR needs descriptive function and/or variable names require tests Tests [doctest/unittest/pytest] are required require type hints https://docs.python.org/3/library/typing.html labels Feb 7, 2023

@algorithms-keeperalgorithms-keeperBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Click here to look at the relevant links ⬇️

🔗 Relevant Links

Repository:

Python:

Automated review generated by algorithms-keeper. If there's any problem regarding this review, please open an issue about it.

algorithms-keeper commands and options

algorithms-keeper actions can be triggered by commenting on this PR:

  • @algorithms-keeper review to trigger the checks for only added pull request files
  • @algorithms-keeper review-all to trigger the checks for all the pull request files, including the modified files. As we cannot post review comments on lines not part of the diff, this command will post all the messages in one comment.

NOTE: Commands are in beta and so this feature is restricted only to a member or owner of the organization.

Comment threadstrings/top_k_frequent_words.py Outdated
Comment threadstrings/top_k_frequent_words.py Outdated
Comment threadstrings/top_k_frequent_words.py Outdated
Comment threadstrings/top_k_frequent_words.py Outdated
@algorithms-keeperalgorithms-keeperBot removed require descriptive names This PR needs descriptive function and/or variable names require tests Tests [doctest/unittest/pytest] are required require type hints https://docs.python.org/3/library/typing.html labels Feb 7, 2023

@CaedenPHCaedenPH left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very well written algorithm, great stuff!

@amos-syd

Copy link
Copy Markdown
ContributorAuthor

Hello @cclauss, I'm tagging you as I saw your involvement in several PRs.
Do you know who can I ask to for a review? Is there anything I should still change in the PR?
Thank you!

@cclauss

cclauss commented Apr 23, 2023

Copy link
Copy Markdown
Member

All of this seems like it could be done in just a few lines using only a collections.Counter so what are the advantages of this approach?

@amos-syd

Copy link
Copy Markdown
ContributorAuthor

All of this seems like it could be done in just a few lines using only a collections.Counter so what are the advantages of this approach?

Thanks for the feedback.
I had a look to collections.Counter and noticed that, indeed, most_common solves exactly this problem: https://github.com/python/cpython/blob/3.11/Lib/collections/__init__.py#L608

This is using heapq.nlargest behind the scenes: https://github.com/python/cpython/blob/3.11/Lib/heapq.py#L523

There is already a Heap class in this repository, I imagined it would be useful to show a typical usage of heaps (i.e. finding out order statistics). Do you think I should close this PR entirely, or do you have other suggestions?

Thank you

@cclauss

Copy link
Copy Markdown
Member
fromcollectionsimportCounterdeftop_k_frequent_words(words, k_value):
return [x[0] forxinCounter(words).most_common(k_value)]

@cclauss

cclauss commented Apr 23, 2023

Copy link
Copy Markdown
Member

Please put the text of our last two messages into the file’s docstring and then we can merge this. That will show why this work was done and that there also that the Python standard library provides a far more straightforward way to solve the same problem.

@amos-syd

Copy link
Copy Markdown
ContributorAuthor

Please put the text of our last two messages into the file’s docstring and then we can merge this. That will show why this work was done and that there also that the Python standard library provides a far more straightforward way to solve the same problem.

Thanks for the feedback.
I have updated the docstring mentioning the (preferable) Python standard library solution.

@algorithms-keeperalgorithms-keeperBot added the tests are failing Do not merge until tests pass label Apr 23, 2023

@cclausscclauss left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting reviewsThis PR is ready to be reviewedenhancementThis PR modified some existing filestests are failingDo not merge until tests pass

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@amos-syd@cclauss@CaedenPH
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Solving the `Top k most frequent words` problem using a max-heap by amos-syd · Pull Request #8125 · TheAlgorithms/Python · GitHub
Skip to content

Solving the Top k most frequent words problem using a max-heap - #8125

Closed
amos-syd wants to merge 0 commit into
TheAlgorithms:masterfrom
amos-syd:master
Closed

Solving the Top k most frequent words problem using a max-heap#8125
amos-syd wants to merge 0 commit into
TheAlgorithms:masterfrom
amos-syd:master

Conversation

@amos-syd

@amos-sydamos-syd commented Feb 7, 2023

Copy link
Copy Markdown
Contributor

Describe your change:

This PR aims to add an algorithm to identify the top k most frequent strings given a provided string list of elements.
To do this, the algorithm is using a max-heap implementation already existing in this repository (a generic type was introduced to allow the usage).

Time complexity is O(n), where n is the number of words:

  • O(n) for building the max-heap
  • k*O(logn) for extracting the k most frequent strings
  • Add an algorithm?
  • Fix a bug or typo in an existing algorithm?
  • Documentation change?

Checklist:

  • I have read CONTRIBUTING.md.
  • This pull request is all my own work -- I have not plagiarized.
  • I know that pull requests will not be merged if they fail the automated tests.
  • This PR only changes one algorithm file. To ease review, please open separate PRs for separate algorithms.
  • All new Python files are placed inside an existing directory.
  • All filenames are in all lowercase characters with no spaces or dashes.
  • All functions and variable names follow Python naming conventions.
  • All function parameters and return values are annotated with Python type hints.
  • All functions have doctests that pass the automated testing.
  • All new algorithms include at least one URL that points to Wikipedia or another similar explanation.
  • If this pull request resolves one or more open issues then the commit message contains Fixes: #{$ISSUE_NO}.

@algorithms-keeperalgorithms-keeperBot added awaiting reviews This PR is ready to be reviewed enhancement This PR modified some existing files require descriptive names This PR needs descriptive function and/or variable names require tests Tests [doctest/unittest/pytest] are required require type hints https://docs.python.org/3/library/typing.html labels Feb 7, 2023

@algorithms-keeperalgorithms-keeperBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Click here to look at the relevant links ⬇️

🔗 Relevant Links

Repository:

Python:

Automated review generated by algorithms-keeper. If there's any problem regarding this review, please open an issue about it.

algorithms-keeper commands and options

algorithms-keeper actions can be triggered by commenting on this PR:

  • @algorithms-keeper review to trigger the checks for only added pull request files
  • @algorithms-keeper review-all to trigger the checks for all the pull request files, including the modified files. As we cannot post review comments on lines not part of the diff, this command will post all the messages in one comment.

NOTE: Commands are in beta and so this feature is restricted only to a member or owner of the organization.

Comment threadstrings/top_k_frequent_words.py Outdated
Comment threadstrings/top_k_frequent_words.py Outdated
Comment threadstrings/top_k_frequent_words.py Outdated
Comment threadstrings/top_k_frequent_words.py Outdated
@algorithms-keeperalgorithms-keeperBot removed require descriptive names This PR needs descriptive function and/or variable names require tests Tests [doctest/unittest/pytest] are required require type hints https://docs.python.org/3/library/typing.html labels Feb 7, 2023

@CaedenPHCaedenPH left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very well written algorithm, great stuff!

@amos-syd

Copy link
Copy Markdown
ContributorAuthor

Hello @cclauss, I'm tagging you as I saw your involvement in several PRs.
Do you know who can I ask to for a review? Is there anything I should still change in the PR?
Thank you!

@cclauss

cclauss commented Apr 23, 2023

Copy link
Copy Markdown
Member

All of this seems like it could be done in just a few lines using only a collections.Counter so what are the advantages of this approach?

@amos-syd

Copy link
Copy Markdown
ContributorAuthor

All of this seems like it could be done in just a few lines using only a collections.Counter so what are the advantages of this approach?

Thanks for the feedback.
I had a look to collections.Counter and noticed that, indeed, most_common solves exactly this problem: https://github.com/python/cpython/blob/3.11/Lib/collections/__init__.py#L608

This is using heapq.nlargest behind the scenes: https://github.com/python/cpython/blob/3.11/Lib/heapq.py#L523

There is already a Heap class in this repository, I imagined it would be useful to show a typical usage of heaps (i.e. finding out order statistics). Do you think I should close this PR entirely, or do you have other suggestions?

Thank you

@cclauss

Copy link
Copy Markdown
Member
fromcollectionsimportCounterdeftop_k_frequent_words(words, k_value):
return [x[0] forxinCounter(words).most_common(k_value)]

@cclauss

cclauss commented Apr 23, 2023

Copy link
Copy Markdown
Member

Please put the text of our last two messages into the file’s docstring and then we can merge this. That will show why this work was done and that there also that the Python standard library provides a far more straightforward way to solve the same problem.

@amos-syd

Copy link
Copy Markdown
ContributorAuthor

Please put the text of our last two messages into the file’s docstring and then we can merge this. That will show why this work was done and that there also that the Python standard library provides a far more straightforward way to solve the same problem.

Thanks for the feedback.
I have updated the docstring mentioning the (preferable) Python standard library solution.

@algorithms-keeperalgorithms-keeperBot added the tests are failing Do not merge until tests pass label Apr 23, 2023

@cclausscclauss left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting reviewsThis PR is ready to be reviewedenhancementThis PR modified some existing filestests are failingDo not merge until tests pass

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@amos-syd@cclauss@CaedenPH
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Solving the `Top k most frequent words` problem using a max-heap by amos-syd · Pull Request #8125 · TheAlgorithms/Python · GitHub
Skip to content

Solving the Top k most frequent words problem using a max-heap - #8125

Closed
amos-syd wants to merge 0 commit into
TheAlgorithms:masterfrom
amos-syd:master
Closed

Solving the Top k most frequent words problem using a max-heap#8125
amos-syd wants to merge 0 commit into
TheAlgorithms:masterfrom
amos-syd:master

Conversation

@amos-syd

@amos-sydamos-syd commented Feb 7, 2023

Copy link
Copy Markdown
Contributor

Describe your change:

This PR aims to add an algorithm to identify the top k most frequent strings given a provided string list of elements.
To do this, the algorithm is using a max-heap implementation already existing in this repository (a generic type was introduced to allow the usage).

Time complexity is O(n), where n is the number of words:

  • O(n) for building the max-heap
  • k*O(logn) for extracting the k most frequent strings
  • Add an algorithm?
  • Fix a bug or typo in an existing algorithm?
  • Documentation change?

Checklist:

  • I have read CONTRIBUTING.md.
  • This pull request is all my own work -- I have not plagiarized.
  • I know that pull requests will not be merged if they fail the automated tests.
  • This PR only changes one algorithm file. To ease review, please open separate PRs for separate algorithms.
  • All new Python files are placed inside an existing directory.
  • All filenames are in all lowercase characters with no spaces or dashes.
  • All functions and variable names follow Python naming conventions.
  • All function parameters and return values are annotated with Python type hints.
  • All functions have doctests that pass the automated testing.
  • All new algorithms include at least one URL that points to Wikipedia or another similar explanation.
  • If this pull request resolves one or more open issues then the commit message contains Fixes: #{$ISSUE_NO}.

@algorithms-keeperalgorithms-keeperBot added awaiting reviews This PR is ready to be reviewed enhancement This PR modified some existing files require descriptive names This PR needs descriptive function and/or variable names require tests Tests [doctest/unittest/pytest] are required require type hints https://docs.python.org/3/library/typing.html labels Feb 7, 2023

@algorithms-keeperalgorithms-keeperBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Click here to look at the relevant links ⬇️

🔗 Relevant Links

Repository:

Python:

Automated review generated by algorithms-keeper. If there's any problem regarding this review, please open an issue about it.

algorithms-keeper commands and options

algorithms-keeper actions can be triggered by commenting on this PR:

  • @algorithms-keeper review to trigger the checks for only added pull request files
  • @algorithms-keeper review-all to trigger the checks for all the pull request files, including the modified files. As we cannot post review comments on lines not part of the diff, this command will post all the messages in one comment.

NOTE: Commands are in beta and so this feature is restricted only to a member or owner of the organization.

Comment threadstrings/top_k_frequent_words.py Outdated
Comment threadstrings/top_k_frequent_words.py Outdated
Comment threadstrings/top_k_frequent_words.py Outdated
Comment threadstrings/top_k_frequent_words.py Outdated
@algorithms-keeperalgorithms-keeperBot removed require descriptive names This PR needs descriptive function and/or variable names require tests Tests [doctest/unittest/pytest] are required require type hints https://docs.python.org/3/library/typing.html labels Feb 7, 2023

@CaedenPHCaedenPH left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very well written algorithm, great stuff!

@amos-syd

Copy link
Copy Markdown
ContributorAuthor

Hello @cclauss, I'm tagging you as I saw your involvement in several PRs.
Do you know who can I ask to for a review? Is there anything I should still change in the PR?
Thank you!

@cclauss

cclauss commented Apr 23, 2023

Copy link
Copy Markdown
Member

All of this seems like it could be done in just a few lines using only a collections.Counter so what are the advantages of this approach?

@amos-syd

Copy link
Copy Markdown
ContributorAuthor

All of this seems like it could be done in just a few lines using only a collections.Counter so what are the advantages of this approach?

Thanks for the feedback.
I had a look to collections.Counter and noticed that, indeed, most_common solves exactly this problem: https://github.com/python/cpython/blob/3.11/Lib/collections/__init__.py#L608

This is using heapq.nlargest behind the scenes: https://github.com/python/cpython/blob/3.11/Lib/heapq.py#L523

There is already a Heap class in this repository, I imagined it would be useful to show a typical usage of heaps (i.e. finding out order statistics). Do you think I should close this PR entirely, or do you have other suggestions?

Thank you

@cclauss

Copy link
Copy Markdown
Member
fromcollectionsimportCounterdeftop_k_frequent_words(words, k_value):
return [x[0] forxinCounter(words).most_common(k_value)]

@cclauss

cclauss commented Apr 23, 2023

Copy link
Copy Markdown
Member

Please put the text of our last two messages into the file’s docstring and then we can merge this. That will show why this work was done and that there also that the Python standard library provides a far more straightforward way to solve the same problem.

@amos-syd

Copy link
Copy Markdown
ContributorAuthor

Please put the text of our last two messages into the file’s docstring and then we can merge this. That will show why this work was done and that there also that the Python standard library provides a far more straightforward way to solve the same problem.

Thanks for the feedback.
I have updated the docstring mentioning the (preferable) Python standard library solution.

@algorithms-keeperalgorithms-keeperBot added the tests are failing Do not merge until tests pass label Apr 23, 2023

@cclausscclauss left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting reviewsThis PR is ready to be reviewedenhancementThis PR modified some existing filestests are failingDo not merge until tests pass

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@amos-syd@cclauss@CaedenPH
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' Solving the `Top k most frequent words` problem using a max-heap by amos-syd · Pull Request #8125 · TheAlgorithms/Python · GitHub
Skip to content

Solving the Top k most frequent words problem using a max-heap - #8125

Closed
amos-syd wants to merge 0 commit into
TheAlgorithms:masterfrom
amos-syd:master
Closed

Solving the Top k most frequent words problem using a max-heap#8125
amos-syd wants to merge 0 commit into
TheAlgorithms:masterfrom
amos-syd:master

Conversation

@amos-syd

@amos-sydamos-syd commented Feb 7, 2023

Copy link
Copy Markdown
Contributor

Describe your change:

This PR aims to add an algorithm to identify the top k most frequent strings given a provided string list of elements.
To do this, the algorithm is using a max-heap implementation already existing in this repository (a generic type was introduced to allow the usage).

Time complexity is O(n), where n is the number of words:

  • O(n) for building the max-heap
  • k*O(logn) for extracting the k most frequent strings
  • Add an algorithm?
  • Fix a bug or typo in an existing algorithm?
  • Documentation change?

Checklist:

  • I have read CONTRIBUTING.md.
  • This pull request is all my own work -- I have not plagiarized.
  • I know that pull requests will not be merged if they fail the automated tests.
  • This PR only changes one algorithm file. To ease review, please open separate PRs for separate algorithms.
  • All new Python files are placed inside an existing directory.
  • All filenames are in all lowercase characters with no spaces or dashes.
  • All functions and variable names follow Python naming conventions.
  • All function parameters and return values are annotated with Python type hints.
  • All functions have doctests that pass the automated testing.
  • All new algorithms include at least one URL that points to Wikipedia or another similar explanation.
  • If this pull request resolves one or more open issues then the commit message contains Fixes: #{$ISSUE_NO}.

@algorithms-keeperalgorithms-keeperBot added awaiting reviews This PR is ready to be reviewed enhancement This PR modified some existing files require descriptive names This PR needs descriptive function and/or variable names require tests Tests [doctest/unittest/pytest] are required require type hints https://docs.python.org/3/library/typing.html labels Feb 7, 2023

@algorithms-keeperalgorithms-keeperBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Click here to look at the relevant links ⬇️

🔗 Relevant Links

Repository:

Python:

Automated review generated by algorithms-keeper. If there's any problem regarding this review, please open an issue about it.

algorithms-keeper commands and options

algorithms-keeper actions can be triggered by commenting on this PR:

  • @algorithms-keeper review to trigger the checks for only added pull request files
  • @algorithms-keeper review-all to trigger the checks for all the pull request files, including the modified files. As we cannot post review comments on lines not part of the diff, this command will post all the messages in one comment.

NOTE: Commands are in beta and so this feature is restricted only to a member or owner of the organization.

Comment threadstrings/top_k_frequent_words.py Outdated
Comment threadstrings/top_k_frequent_words.py Outdated
Comment threadstrings/top_k_frequent_words.py Outdated
Comment threadstrings/top_k_frequent_words.py Outdated
@algorithms-keeperalgorithms-keeperBot removed require descriptive names This PR needs descriptive function and/or variable names require tests Tests [doctest/unittest/pytest] are required require type hints https://docs.python.org/3/library/typing.html labels Feb 7, 2023

@CaedenPHCaedenPH left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very well written algorithm, great stuff!

@amos-syd

Copy link
Copy Markdown
ContributorAuthor

Hello @cclauss, I'm tagging you as I saw your involvement in several PRs.
Do you know who can I ask to for a review? Is there anything I should still change in the PR?
Thank you!

@cclauss

cclauss commented Apr 23, 2023

Copy link
Copy Markdown
Member

All of this seems like it could be done in just a few lines using only a collections.Counter so what are the advantages of this approach?

@amos-syd

Copy link
Copy Markdown
ContributorAuthor

All of this seems like it could be done in just a few lines using only a collections.Counter so what are the advantages of this approach?

Thanks for the feedback.
I had a look to collections.Counter and noticed that, indeed, most_common solves exactly this problem: https://github.com/python/cpython/blob/3.11/Lib/collections/__init__.py#L608

This is using heapq.nlargest behind the scenes: https://github.com/python/cpython/blob/3.11/Lib/heapq.py#L523

There is already a Heap class in this repository, I imagined it would be useful to show a typical usage of heaps (i.e. finding out order statistics). Do you think I should close this PR entirely, or do you have other suggestions?

Thank you

@cclauss

Copy link
Copy Markdown
Member
fromcollectionsimportCounterdeftop_k_frequent_words(words, k_value):
return [x[0] forxinCounter(words).most_common(k_value)]

@cclauss

cclauss commented Apr 23, 2023

Copy link
Copy Markdown
Member

Please put the text of our last two messages into the file’s docstring and then we can merge this. That will show why this work was done and that there also that the Python standard library provides a far more straightforward way to solve the same problem.

@amos-syd

Copy link
Copy Markdown
ContributorAuthor

Please put the text of our last two messages into the file’s docstring and then we can merge this. That will show why this work was done and that there also that the Python standard library provides a far more straightforward way to solve the same problem.

Thanks for the feedback.
I have updated the docstring mentioning the (preferable) Python standard library solution.

@algorithms-keeperalgorithms-keeperBot added the tests are failing Do not merge until tests pass label Apr 23, 2023

@cclausscclauss left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting reviewsThis PR is ready to be reviewedenhancementThis PR modified some existing filestests are failingDo not merge until tests pass

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@amos-syd@cclauss@CaedenPH
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Solving the `Top k most frequent words` problem using a max-heap by amos-syd · Pull Request #8125 · TheAlgorithms/Python · GitHub
Skip to content

Solving the Top k most frequent words problem using a max-heap - #8125

Closed
amos-syd wants to merge 0 commit into
TheAlgorithms:masterfrom
amos-syd:master
Closed

Solving the Top k most frequent words problem using a max-heap#8125
amos-syd wants to merge 0 commit into
TheAlgorithms:masterfrom
amos-syd:master

Conversation

@amos-syd

@amos-sydamos-syd commented Feb 7, 2023

Copy link
Copy Markdown
Contributor

Describe your change:

This PR aims to add an algorithm to identify the top k most frequent strings given a provided string list of elements.
To do this, the algorithm is using a max-heap implementation already existing in this repository (a generic type was introduced to allow the usage).

Time complexity is O(n), where n is the number of words:

  • O(n) for building the max-heap
  • k*O(logn) for extracting the k most frequent strings
  • Add an algorithm?
  • Fix a bug or typo in an existing algorithm?
  • Documentation change?

Checklist:

  • I have read CONTRIBUTING.md.
  • This pull request is all my own work -- I have not plagiarized.
  • I know that pull requests will not be merged if they fail the automated tests.
  • This PR only changes one algorithm file. To ease review, please open separate PRs for separate algorithms.
  • All new Python files are placed inside an existing directory.
  • All filenames are in all lowercase characters with no spaces or dashes.
  • All functions and variable names follow Python naming conventions.
  • All function parameters and return values are annotated with Python type hints.
  • All functions have doctests that pass the automated testing.
  • All new algorithms include at least one URL that points to Wikipedia or another similar explanation.
  • If this pull request resolves one or more open issues then the commit message contains Fixes: #{$ISSUE_NO}.

@algorithms-keeperalgorithms-keeperBot added awaiting reviews This PR is ready to be reviewed enhancement This PR modified some existing files require descriptive names This PR needs descriptive function and/or variable names require tests Tests [doctest/unittest/pytest] are required require type hints https://docs.python.org/3/library/typing.html labels Feb 7, 2023

@algorithms-keeperalgorithms-keeperBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Click here to look at the relevant links ⬇️

🔗 Relevant Links

Repository:

Python:

Automated review generated by algorithms-keeper. If there's any problem regarding this review, please open an issue about it.

algorithms-keeper commands and options

algorithms-keeper actions can be triggered by commenting on this PR:

  • @algorithms-keeper review to trigger the checks for only added pull request files
  • @algorithms-keeper review-all to trigger the checks for all the pull request files, including the modified files. As we cannot post review comments on lines not part of the diff, this command will post all the messages in one comment.

NOTE: Commands are in beta and so this feature is restricted only to a member or owner of the organization.

Comment threadstrings/top_k_frequent_words.py Outdated
Comment threadstrings/top_k_frequent_words.py Outdated
Comment threadstrings/top_k_frequent_words.py Outdated
Comment threadstrings/top_k_frequent_words.py Outdated
@algorithms-keeperalgorithms-keeperBot removed require descriptive names This PR needs descriptive function and/or variable names require tests Tests [doctest/unittest/pytest] are required require type hints https://docs.python.org/3/library/typing.html labels Feb 7, 2023

@CaedenPHCaedenPH left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very well written algorithm, great stuff!

@amos-syd

Copy link
Copy Markdown
ContributorAuthor

Hello @cclauss, I'm tagging you as I saw your involvement in several PRs.
Do you know who can I ask to for a review? Is there anything I should still change in the PR?
Thank you!

@cclauss

cclauss commented Apr 23, 2023

Copy link
Copy Markdown
Member

All of this seems like it could be done in just a few lines using only a collections.Counter so what are the advantages of this approach?

@amos-syd

Copy link
Copy Markdown
ContributorAuthor

All of this seems like it could be done in just a few lines using only a collections.Counter so what are the advantages of this approach?

Thanks for the feedback.
I had a look to collections.Counter and noticed that, indeed, most_common solves exactly this problem: https://github.com/python/cpython/blob/3.11/Lib/collections/__init__.py#L608

This is using heapq.nlargest behind the scenes: https://github.com/python/cpython/blob/3.11/Lib/heapq.py#L523

There is already a Heap class in this repository, I imagined it would be useful to show a typical usage of heaps (i.e. finding out order statistics). Do you think I should close this PR entirely, or do you have other suggestions?

Thank you

@cclauss

Copy link
Copy Markdown
Member
fromcollectionsimportCounterdeftop_k_frequent_words(words, k_value):
return [x[0] forxinCounter(words).most_common(k_value)]

@cclauss

cclauss commented Apr 23, 2023

Copy link
Copy Markdown
Member

Please put the text of our last two messages into the file’s docstring and then we can merge this. That will show why this work was done and that there also that the Python standard library provides a far more straightforward way to solve the same problem.

@amos-syd

Copy link
Copy Markdown
ContributorAuthor

Please put the text of our last two messages into the file’s docstring and then we can merge this. That will show why this work was done and that there also that the Python standard library provides a far more straightforward way to solve the same problem.

Thanks for the feedback.
I have updated the docstring mentioning the (preferable) Python standard library solution.

@algorithms-keeperalgorithms-keeperBot added the tests are failing Do not merge until tests pass label Apr 23, 2023

@cclausscclauss left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting reviewsThis PR is ready to be reviewedenhancementThis PR modified some existing filestests are failingDo not merge until tests pass

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@amos-syd@cclauss@CaedenPH
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Solving the `Top k most frequent words` problem using a max-heap by amos-syd · Pull Request #8125 · TheAlgorithms/Python · GitHub
Skip to content

Solving the Top k most frequent words problem using a max-heap - #8125

Closed
amos-syd wants to merge 0 commit into
TheAlgorithms:masterfrom
amos-syd:master
Closed

Solving the Top k most frequent words problem using a max-heap#8125
amos-syd wants to merge 0 commit into
TheAlgorithms:masterfrom
amos-syd:master

Conversation

@amos-syd

@amos-sydamos-syd commented Feb 7, 2023

Copy link
Copy Markdown
Contributor

Describe your change:

This PR aims to add an algorithm to identify the top k most frequent strings given a provided string list of elements.
To do this, the algorithm is using a max-heap implementation already existing in this repository (a generic type was introduced to allow the usage).

Time complexity is O(n), where n is the number of words:

  • O(n) for building the max-heap
  • k*O(logn) for extracting the k most frequent strings
  • Add an algorithm?
  • Fix a bug or typo in an existing algorithm?
  • Documentation change?

Checklist:

  • I have read CONTRIBUTING.md.
  • This pull request is all my own work -- I have not plagiarized.
  • I know that pull requests will not be merged if they fail the automated tests.
  • This PR only changes one algorithm file. To ease review, please open separate PRs for separate algorithms.
  • All new Python files are placed inside an existing directory.
  • All filenames are in all lowercase characters with no spaces or dashes.
  • All functions and variable names follow Python naming conventions.
  • All function parameters and return values are annotated with Python type hints.
  • All functions have doctests that pass the automated testing.
  • All new algorithms include at least one URL that points to Wikipedia or another similar explanation.
  • If this pull request resolves one or more open issues then the commit message contains Fixes: #{$ISSUE_NO}.

@algorithms-keeperalgorithms-keeperBot added awaiting reviews This PR is ready to be reviewed enhancement This PR modified some existing files require descriptive names This PR needs descriptive function and/or variable names require tests Tests [doctest/unittest/pytest] are required require type hints https://docs.python.org/3/library/typing.html labels Feb 7, 2023

@algorithms-keeperalgorithms-keeperBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Click here to look at the relevant links ⬇️

🔗 Relevant Links

Repository:

Python:

Automated review generated by algorithms-keeper. If there's any problem regarding this review, please open an issue about it.

algorithms-keeper commands and options

algorithms-keeper actions can be triggered by commenting on this PR:

  • @algorithms-keeper review to trigger the checks for only added pull request files
  • @algorithms-keeper review-all to trigger the checks for all the pull request files, including the modified files. As we cannot post review comments on lines not part of the diff, this command will post all the messages in one comment.

NOTE: Commands are in beta and so this feature is restricted only to a member or owner of the organization.

Comment threadstrings/top_k_frequent_words.py Outdated
Comment threadstrings/top_k_frequent_words.py Outdated
Comment threadstrings/top_k_frequent_words.py Outdated
Comment threadstrings/top_k_frequent_words.py Outdated
@algorithms-keeperalgorithms-keeperBot removed require descriptive names This PR needs descriptive function and/or variable names require tests Tests [doctest/unittest/pytest] are required require type hints https://docs.python.org/3/library/typing.html labels Feb 7, 2023

@CaedenPHCaedenPH left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very well written algorithm, great stuff!

@amos-syd

Copy link
Copy Markdown
ContributorAuthor

Hello @cclauss, I'm tagging you as I saw your involvement in several PRs.
Do you know who can I ask to for a review? Is there anything I should still change in the PR?
Thank you!

@cclauss

cclauss commented Apr 23, 2023

Copy link
Copy Markdown
Member

All of this seems like it could be done in just a few lines using only a collections.Counter so what are the advantages of this approach?

@amos-syd

Copy link
Copy Markdown
ContributorAuthor

All of this seems like it could be done in just a few lines using only a collections.Counter so what are the advantages of this approach?

Thanks for the feedback.
I had a look to collections.Counter and noticed that, indeed, most_common solves exactly this problem: https://github.com/python/cpython/blob/3.11/Lib/collections/__init__.py#L608

This is using heapq.nlargest behind the scenes: https://github.com/python/cpython/blob/3.11/Lib/heapq.py#L523

There is already a Heap class in this repository, I imagined it would be useful to show a typical usage of heaps (i.e. finding out order statistics). Do you think I should close this PR entirely, or do you have other suggestions?

Thank you

@cclauss

Copy link
Copy Markdown
Member
fromcollectionsimportCounterdeftop_k_frequent_words(words, k_value):
return [x[0] forxinCounter(words).most_common(k_value)]

@cclauss

cclauss commented Apr 23, 2023

Copy link
Copy Markdown
Member

Please put the text of our last two messages into the file’s docstring and then we can merge this. That will show why this work was done and that there also that the Python standard library provides a far more straightforward way to solve the same problem.

@amos-syd

Copy link
Copy Markdown
ContributorAuthor

Please put the text of our last two messages into the file’s docstring and then we can merge this. That will show why this work was done and that there also that the Python standard library provides a far more straightforward way to solve the same problem.

Thanks for the feedback.
I have updated the docstring mentioning the (preferable) Python standard library solution.

@algorithms-keeperalgorithms-keeperBot added the tests are failing Do not merge until tests pass label Apr 23, 2023

@cclausscclauss left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting reviewsThis PR is ready to be reviewedenhancementThis PR modified some existing filestests are failingDo not merge until tests pass

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@amos-syd@cclauss@CaedenPH
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); Solving the `Top k most frequent words` problem using a max-heap by amos-syd · Pull Request #8125 · TheAlgorithms/Python · GitHub
Skip to content

Solving the Top k most frequent words problem using a max-heap - #8125

Closed
amos-syd wants to merge 0 commit into
TheAlgorithms:masterfrom
amos-syd:master
Closed

Solving the Top k most frequent words problem using a max-heap#8125
amos-syd wants to merge 0 commit into
TheAlgorithms:masterfrom
amos-syd:master

Conversation

@amos-syd

@amos-sydamos-syd commented Feb 7, 2023

Copy link
Copy Markdown
Contributor

Describe your change:

This PR aims to add an algorithm to identify the top k most frequent strings given a provided string list of elements.
To do this, the algorithm is using a max-heap implementation already existing in this repository (a generic type was introduced to allow the usage).

Time complexity is O(n), where n is the number of words:

  • O(n) for building the max-heap
  • k*O(logn) for extracting the k most frequent strings
  • Add an algorithm?
  • Fix a bug or typo in an existing algorithm?
  • Documentation change?

Checklist:

  • I have read CONTRIBUTING.md.
  • This pull request is all my own work -- I have not plagiarized.
  • I know that pull requests will not be merged if they fail the automated tests.
  • This PR only changes one algorithm file. To ease review, please open separate PRs for separate algorithms.
  • All new Python files are placed inside an existing directory.
  • All filenames are in all lowercase characters with no spaces or dashes.
  • All functions and variable names follow Python naming conventions.
  • All function parameters and return values are annotated with Python type hints.
  • All functions have doctests that pass the automated testing.
  • All new algorithms include at least one URL that points to Wikipedia or another similar explanation.
  • If this pull request resolves one or more open issues then the commit message contains Fixes: #{$ISSUE_NO}.

@algorithms-keeperalgorithms-keeperBot added awaiting reviews This PR is ready to be reviewed enhancement This PR modified some existing files require descriptive names This PR needs descriptive function and/or variable names require tests Tests [doctest/unittest/pytest] are required require type hints https://docs.python.org/3/library/typing.html labels Feb 7, 2023

@algorithms-keeperalgorithms-keeperBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Click here to look at the relevant links ⬇️

🔗 Relevant Links

Repository:

Python:

Automated review generated by algorithms-keeper. If there's any problem regarding this review, please open an issue about it.

algorithms-keeper commands and options

algorithms-keeper actions can be triggered by commenting on this PR:

  • @algorithms-keeper review to trigger the checks for only added pull request files
  • @algorithms-keeper review-all to trigger the checks for all the pull request files, including the modified files. As we cannot post review comments on lines not part of the diff, this command will post all the messages in one comment.

NOTE: Commands are in beta and so this feature is restricted only to a member or owner of the organization.

Comment threadstrings/top_k_frequent_words.py Outdated
Comment threadstrings/top_k_frequent_words.py Outdated
Comment threadstrings/top_k_frequent_words.py Outdated
Comment threadstrings/top_k_frequent_words.py Outdated
@algorithms-keeperalgorithms-keeperBot removed require descriptive names This PR needs descriptive function and/or variable names require tests Tests [doctest/unittest/pytest] are required require type hints https://docs.python.org/3/library/typing.html labels Feb 7, 2023

@CaedenPHCaedenPH left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very well written algorithm, great stuff!

@amos-syd

Copy link
Copy Markdown
ContributorAuthor

Hello @cclauss, I'm tagging you as I saw your involvement in several PRs.
Do you know who can I ask to for a review? Is there anything I should still change in the PR?
Thank you!

@cclauss

cclauss commented Apr 23, 2023

Copy link
Copy Markdown
Member

All of this seems like it could be done in just a few lines using only a collections.Counter so what are the advantages of this approach?

@amos-syd

Copy link
Copy Markdown
ContributorAuthor

All of this seems like it could be done in just a few lines using only a collections.Counter so what are the advantages of this approach?

Thanks for the feedback.
I had a look to collections.Counter and noticed that, indeed, most_common solves exactly this problem: https://github.com/python/cpython/blob/3.11/Lib/collections/__init__.py#L608

This is using heapq.nlargest behind the scenes: https://github.com/python/cpython/blob/3.11/Lib/heapq.py#L523

There is already a Heap class in this repository, I imagined it would be useful to show a typical usage of heaps (i.e. finding out order statistics). Do you think I should close this PR entirely, or do you have other suggestions?

Thank you

@cclauss

Copy link
Copy Markdown
Member
fromcollectionsimportCounterdeftop_k_frequent_words(words, k_value):
return [x[0] forxinCounter(words).most_common(k_value)]

@cclauss

cclauss commented Apr 23, 2023

Copy link
Copy Markdown
Member

Please put the text of our last two messages into the file’s docstring and then we can merge this. That will show why this work was done and that there also that the Python standard library provides a far more straightforward way to solve the same problem.

@amos-syd

Copy link
Copy Markdown
ContributorAuthor

Please put the text of our last two messages into the file’s docstring and then we can merge this. That will show why this work was done and that there also that the Python standard library provides a far more straightforward way to solve the same problem.

Thanks for the feedback.
I have updated the docstring mentioning the (preferable) Python standard library solution.

@algorithms-keeperalgorithms-keeperBot added the tests are failing Do not merge until tests pass label Apr 23, 2023

@cclausscclauss left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting reviewsThis PR is ready to be reviewedenhancementThis PR modified some existing filestests are failingDo not merge until tests pass

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@amos-syd@cclauss@CaedenPH