Cache channel metadata - #2333

Merged
rtibbles merged 16 commits into
learningequality:developfrom
jredrejo:delay_channel_size
Sep 30, 2020
Merged

Cache channel metadata#2333
rtibbles merged 16 commits into
learningequality:developfrom
jredrejo:delay_channel_size

Conversation

@jredrejo

@jredrejojredrejo commented Sep 28, 2020

Copy link
Copy Markdown
Contributor

Description

Removes AdminChannelViewSet annotations and defers them to be executed in an async task that will cache the results.

As operations are executed asynchronously, the values are not immediately available if they have not been previously cached. In this case the values are marked with a DEFERRED_FLAG so the frontend knows that it can request them later.

The structure of the created object can be seen in the contentcuration/contentcuration/utils/channel.py and it's done according to the https://www.notion.so/learningequality/2020-12-24-Team-Sonic-counting-descendants-meeting-7ea0f46d4a6648109a22eefbba124a92 decissions

NOTE: some small optimization have also been done, mainly removing unneded order_by requirements in some of the queries, and the use of CTE to avoid scanning the whole File table after paginating the results.

Issue Addressed (if applicable)

Current annotations in AdminChannelViewSet returns queries needing days to be executed (File table has 76 millions of rows and ContentNode has 10 millions and the queries where scanning and joining both tables)

Implementation Notes (optional)

  • Added an api endpoint for the frontend to retrieve cached metadata.

Example of use: http://localhost:8080/api/admin-channels/deferred_data?id__in=0598c68f00fc562486fc6616b63f267f,c1f2b7e6ac9f56a2bb44fa7a48b66dce

  • Two different tasks are implemented:

    • cache_channel_metadata_task(key, channel, tree_id) to cache one single channel
    • cache_multiple_channels_metadata_task(channels) to cache multiple channels

Cache invalidation

Cache keys will be stored forever
The timestamp stored in the key will be used to retrieve data, but will be marked as stale and recalculated after one hour (following a probabilistic cache invalidation algorithm to avoid a recalculation stampede)

@codecov

codecovBot commented Sep 28, 2020

Copy link
Copy Markdown

Codecov Report

Merging #2333 into develop will decrease coverage by 1.74%.
The diff coverage is 36.06%.

Impacted file tree graph

@@ Coverage Diff @@## develop #2333 +/- ##
===========================================
- Coverage 81.46% 79.72% -1.75% 
===========================================
Files 293 295 +2 Lines 14106 14587 +481 ===========================================
+ Hits 11491 11629 +138 - Misses 2615 2958 +343 
Impacted FilesCoverage Δ
...ontentcuration/contentcuration/viewsets/channel.py73.54% <25.58%> (-8.36%)⬇️
contentcuration/contentcuration/utils/channel.py34.21% <34.21%> (ø)
contentcuration/contentcuration/utils/cache.py43.90% <42.42%> (-19.74%)⬇️
contentcuration/contentcuration/tasks.py72.88% <66.66%> (-0.34%)⬇️
contentcuration/contentcuration/models.py86.99% <100.00%> (+0.16%)⬆️
...contentcuration/tests/viewsets/test_contentnode.py66.85% <0.00%> (-19.71%)⬇️
contentcuration/contentcuration/utils/nodes.py65.93% <0.00%> (-13.54%)⬇️
...uration/contentcuration/tests/test_contentnodes.py97.21% <0.00%> (-0.05%)⬇️
contentcuration/contentcuration/views/base.py55.27% <0.00%> (-0.05%)⬇️
contentcuration/contentcuration/test_settings.py100.00% <0.00%> (ø)
... and 8 more

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 1bf3f74...6a1a4e7. Read the comment docs.

@rtibblesrtibbles left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some questions, some ignorable suggestions, and one suggested extension.

Comment threadcontentcuration/contentcuration/utils/channel.py Outdated
Comment threadcontentcuration/contentcuration/viewsets/channel.py
Comment threadcontentcuration/contentcuration/viewsets/channel.py Outdated
Comment threadcontentcuration/contentcuration/viewsets/channel.py Outdated
Comment threadcontentcuration/contentcuration/utils/cache.py Outdated
Comment threadcontentcuration/contentcuration/viewsets/channel.py
ids = request.GET.get("id__in")
if not ids:
raise ValidationError("id__in GET parameter is required")
ids = ids.split(",")

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not necessary for the admin viewset, but if we generalize this we should probably do a quick filter here by the channels the user has view permissions for, so something like:

ids = self.get_queryset().filter(id__in=ids).values_list("id", flat=True)

Will then filter down the id list to the ones that the user has permissions for (in the admin use case, this is less important though).

@rtibbles

Copy link
Copy Markdown
Member

Just need to do a manual test, but this looks good to merge and build on!

@rtibblesrtibbles left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One comment, one question.



def cache_stampede(expire, beta=1):
"""Cache decorator with cache stampede protection.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we add a comment with a link to where we vendored it from?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I added a link to the pdf of the paper and its explanation. I'll add it to the diskcache code too, I stole some ideas from there (mainly the decorator usage) but for the main logic I used the paper. This implementation is simpler than diskcache's.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

OK, if you didn't lean on diskcache's code that much, this is fine then!

return result, delta

@functools.wraps(func)
def wrapper(*args, **kwargs):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we want to reuse this decorator, guessing we want to be able to parameterize the cache key template - but I think everything else could stay the same (except channel_id would be a more generic id variable name). Thoughts?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes, because the key is fetched from the decorated function, the change should be calling it with the final key, applying the parametrization before, it should be an easy change. Do you want me to do it now, so it's prepare for it?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sure, if it's no bother!

@rtibblesrtibbles left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's merge and iterate!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jredrejo@rtibbles
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Cache channel metadata - #2333

Merged
rtibbles merged 16 commits into
learningequality:developfrom
jredrejo:delay_channel_size
Sep 30, 2020
Merged

Cache channel metadata#2333
rtibbles merged 16 commits into
learningequality:developfrom
jredrejo:delay_channel_size

Conversation

@jredrejo

@jredrejojredrejo commented Sep 28, 2020

Copy link
Copy Markdown
Contributor

Description

Removes AdminChannelViewSet annotations and defers them to be executed in an async task that will cache the results.

As operations are executed asynchronously, the values are not immediately available if they have not been previously cached. In this case the values are marked with a DEFERRED_FLAG so the frontend knows that it can request them later.

The structure of the created object can be seen in the contentcuration/contentcuration/utils/channel.py and it's done according to the https://www.notion.so/learningequality/2020-12-24-Team-Sonic-counting-descendants-meeting-7ea0f46d4a6648109a22eefbba124a92 decissions

NOTE: some small optimization have also been done, mainly removing unneded order_by requirements in some of the queries, and the use of CTE to avoid scanning the whole File table after paginating the results.

Issue Addressed (if applicable)

Current annotations in AdminChannelViewSet returns queries needing days to be executed (File table has 76 millions of rows and ContentNode has 10 millions and the queries where scanning and joining both tables)

Implementation Notes (optional)

  • Added an api endpoint for the frontend to retrieve cached metadata.

Example of use: http://localhost:8080/api/admin-channels/deferred_data?id__in=0598c68f00fc562486fc6616b63f267f,c1f2b7e6ac9f56a2bb44fa7a48b66dce

  • Two different tasks are implemented:

    • cache_channel_metadata_task(key, channel, tree_id) to cache one single channel
    • cache_multiple_channels_metadata_task(channels) to cache multiple channels

Cache invalidation

Cache keys will be stored forever
The timestamp stored in the key will be used to retrieve data, but will be marked as stale and recalculated after one hour (following a probabilistic cache invalidation algorithm to avoid a recalculation stampede)

@codecov

codecovBot commented Sep 28, 2020

Copy link
Copy Markdown

Codecov Report

Merging #2333 into develop will decrease coverage by 1.74%.
The diff coverage is 36.06%.

Impacted file tree graph

@@ Coverage Diff @@## develop #2333 +/- ##
===========================================
- Coverage 81.46% 79.72% -1.75% 
===========================================
Files 293 295 +2 Lines 14106 14587 +481 ===========================================
+ Hits 11491 11629 +138 - Misses 2615 2958 +343 
Impacted FilesCoverage Δ
...ontentcuration/contentcuration/viewsets/channel.py73.54% <25.58%> (-8.36%)⬇️
contentcuration/contentcuration/utils/channel.py34.21% <34.21%> (ø)
contentcuration/contentcuration/utils/cache.py43.90% <42.42%> (-19.74%)⬇️
contentcuration/contentcuration/tasks.py72.88% <66.66%> (-0.34%)⬇️
contentcuration/contentcuration/models.py86.99% <100.00%> (+0.16%)⬆️
...contentcuration/tests/viewsets/test_contentnode.py66.85% <0.00%> (-19.71%)⬇️
contentcuration/contentcuration/utils/nodes.py65.93% <0.00%> (-13.54%)⬇️
...uration/contentcuration/tests/test_contentnodes.py97.21% <0.00%> (-0.05%)⬇️
contentcuration/contentcuration/views/base.py55.27% <0.00%> (-0.05%)⬇️
contentcuration/contentcuration/test_settings.py100.00% <0.00%> (ø)
... and 8 more

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 1bf3f74...6a1a4e7. Read the comment docs.

@rtibblesrtibbles left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some questions, some ignorable suggestions, and one suggested extension.

Comment threadcontentcuration/contentcuration/utils/channel.py Outdated
Comment threadcontentcuration/contentcuration/viewsets/channel.py
Comment threadcontentcuration/contentcuration/viewsets/channel.py Outdated
Comment threadcontentcuration/contentcuration/viewsets/channel.py Outdated
Comment threadcontentcuration/contentcuration/utils/cache.py Outdated
Comment threadcontentcuration/contentcuration/viewsets/channel.py
ids = request.GET.get("id__in")
if not ids:
raise ValidationError("id__in GET parameter is required")
ids = ids.split(",")

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not necessary for the admin viewset, but if we generalize this we should probably do a quick filter here by the channels the user has view permissions for, so something like:

ids = self.get_queryset().filter(id__in=ids).values_list("id", flat=True)

Will then filter down the id list to the ones that the user has permissions for (in the admin use case, this is less important though).

@rtibbles

Copy link
Copy Markdown
Member

Just need to do a manual test, but this looks good to merge and build on!

@rtibblesrtibbles left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One comment, one question.



def cache_stampede(expire, beta=1):
"""Cache decorator with cache stampede protection.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we add a comment with a link to where we vendored it from?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I added a link to the pdf of the paper and its explanation. I'll add it to the diskcache code too, I stole some ideas from there (mainly the decorator usage) but for the main logic I used the paper. This implementation is simpler than diskcache's.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

OK, if you didn't lean on diskcache's code that much, this is fine then!

return result, delta

@functools.wraps(func)
def wrapper(*args, **kwargs):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we want to reuse this decorator, guessing we want to be able to parameterize the cache key template - but I think everything else could stay the same (except channel_id would be a more generic id variable name). Thoughts?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes, because the key is fetched from the decorated function, the change should be calling it with the final key, applying the parametrization before, it should be an easy change. Do you want me to do it now, so it's prepare for it?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sure, if it's no bother!

@rtibblesrtibbles left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's merge and iterate!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jredrejo@rtibbles
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Cache channel metadata - #2333

Merged
rtibbles merged 16 commits into
learningequality:developfrom
jredrejo:delay_channel_size
Sep 30, 2020
Merged

Cache channel metadata#2333
rtibbles merged 16 commits into
learningequality:developfrom
jredrejo:delay_channel_size

Conversation

@jredrejo

@jredrejojredrejo commented Sep 28, 2020

Copy link
Copy Markdown
Contributor

Description

Removes AdminChannelViewSet annotations and defers them to be executed in an async task that will cache the results.

As operations are executed asynchronously, the values are not immediately available if they have not been previously cached. In this case the values are marked with a DEFERRED_FLAG so the frontend knows that it can request them later.

The structure of the created object can be seen in the contentcuration/contentcuration/utils/channel.py and it's done according to the https://www.notion.so/learningequality/2020-12-24-Team-Sonic-counting-descendants-meeting-7ea0f46d4a6648109a22eefbba124a92 decissions

NOTE: some small optimization have also been done, mainly removing unneded order_by requirements in some of the queries, and the use of CTE to avoid scanning the whole File table after paginating the results.

Issue Addressed (if applicable)

Current annotations in AdminChannelViewSet returns queries needing days to be executed (File table has 76 millions of rows and ContentNode has 10 millions and the queries where scanning and joining both tables)

Implementation Notes (optional)

  • Added an api endpoint for the frontend to retrieve cached metadata.

Example of use: http://localhost:8080/api/admin-channels/deferred_data?id__in=0598c68f00fc562486fc6616b63f267f,c1f2b7e6ac9f56a2bb44fa7a48b66dce

  • Two different tasks are implemented:

    • cache_channel_metadata_task(key, channel, tree_id) to cache one single channel
    • cache_multiple_channels_metadata_task(channels) to cache multiple channels

Cache invalidation

Cache keys will be stored forever
The timestamp stored in the key will be used to retrieve data, but will be marked as stale and recalculated after one hour (following a probabilistic cache invalidation algorithm to avoid a recalculation stampede)

@codecov

codecovBot commented Sep 28, 2020

Copy link
Copy Markdown

Codecov Report

Merging #2333 into develop will decrease coverage by 1.74%.
The diff coverage is 36.06%.

Impacted file tree graph

@@ Coverage Diff @@## develop #2333 +/- ##
===========================================
- Coverage 81.46% 79.72% -1.75% 
===========================================
Files 293 295 +2 Lines 14106 14587 +481 ===========================================
+ Hits 11491 11629 +138 - Misses 2615 2958 +343 
Impacted FilesCoverage Δ
...ontentcuration/contentcuration/viewsets/channel.py73.54% <25.58%> (-8.36%)⬇️
contentcuration/contentcuration/utils/channel.py34.21% <34.21%> (ø)
contentcuration/contentcuration/utils/cache.py43.90% <42.42%> (-19.74%)⬇️
contentcuration/contentcuration/tasks.py72.88% <66.66%> (-0.34%)⬇️
contentcuration/contentcuration/models.py86.99% <100.00%> (+0.16%)⬆️
...contentcuration/tests/viewsets/test_contentnode.py66.85% <0.00%> (-19.71%)⬇️
contentcuration/contentcuration/utils/nodes.py65.93% <0.00%> (-13.54%)⬇️
...uration/contentcuration/tests/test_contentnodes.py97.21% <0.00%> (-0.05%)⬇️
contentcuration/contentcuration/views/base.py55.27% <0.00%> (-0.05%)⬇️
contentcuration/contentcuration/test_settings.py100.00% <0.00%> (ø)
... and 8 more

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 1bf3f74...6a1a4e7. Read the comment docs.

@rtibblesrtibbles left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some questions, some ignorable suggestions, and one suggested extension.

Comment threadcontentcuration/contentcuration/utils/channel.py Outdated
Comment threadcontentcuration/contentcuration/viewsets/channel.py
Comment threadcontentcuration/contentcuration/viewsets/channel.py Outdated
Comment threadcontentcuration/contentcuration/viewsets/channel.py Outdated
Comment threadcontentcuration/contentcuration/utils/cache.py Outdated
Comment threadcontentcuration/contentcuration/viewsets/channel.py
ids = request.GET.get("id__in")
if not ids:
raise ValidationError("id__in GET parameter is required")
ids = ids.split(",")

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not necessary for the admin viewset, but if we generalize this we should probably do a quick filter here by the channels the user has view permissions for, so something like:

ids = self.get_queryset().filter(id__in=ids).values_list("id", flat=True)

Will then filter down the id list to the ones that the user has permissions for (in the admin use case, this is less important though).

@rtibbles

Copy link
Copy Markdown
Member

Just need to do a manual test, but this looks good to merge and build on!

@rtibblesrtibbles left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One comment, one question.



def cache_stampede(expire, beta=1):
"""Cache decorator with cache stampede protection.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we add a comment with a link to where we vendored it from?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I added a link to the pdf of the paper and its explanation. I'll add it to the diskcache code too, I stole some ideas from there (mainly the decorator usage) but for the main logic I used the paper. This implementation is simpler than diskcache's.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

OK, if you didn't lean on diskcache's code that much, this is fine then!

return result, delta

@functools.wraps(func)
def wrapper(*args, **kwargs):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we want to reuse this decorator, guessing we want to be able to parameterize the cache key template - but I think everything else could stay the same (except channel_id would be a more generic id variable name). Thoughts?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes, because the key is fetched from the decorated function, the change should be calling it with the final key, applying the parametrization before, it should be an easy change. Do you want me to do it now, so it's prepare for it?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sure, if it's no bother!

@rtibblesrtibbles left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's merge and iterate!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jredrejo@rtibbles
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Cache channel metadata - #2333

Merged
rtibbles merged 16 commits into
learningequality:developfrom
jredrejo:delay_channel_size
Sep 30, 2020
Merged

Cache channel metadata#2333
rtibbles merged 16 commits into
learningequality:developfrom
jredrejo:delay_channel_size

Conversation

@jredrejo

@jredrejojredrejo commented Sep 28, 2020

Copy link
Copy Markdown
Contributor

Description

Removes AdminChannelViewSet annotations and defers them to be executed in an async task that will cache the results.

As operations are executed asynchronously, the values are not immediately available if they have not been previously cached. In this case the values are marked with a DEFERRED_FLAG so the frontend knows that it can request them later.

The structure of the created object can be seen in the contentcuration/contentcuration/utils/channel.py and it's done according to the https://www.notion.so/learningequality/2020-12-24-Team-Sonic-counting-descendants-meeting-7ea0f46d4a6648109a22eefbba124a92 decissions

NOTE: some small optimization have also been done, mainly removing unneded order_by requirements in some of the queries, and the use of CTE to avoid scanning the whole File table after paginating the results.

Issue Addressed (if applicable)

Current annotations in AdminChannelViewSet returns queries needing days to be executed (File table has 76 millions of rows and ContentNode has 10 millions and the queries where scanning and joining both tables)

Implementation Notes (optional)

  • Added an api endpoint for the frontend to retrieve cached metadata.

Example of use: http://localhost:8080/api/admin-channels/deferred_data?id__in=0598c68f00fc562486fc6616b63f267f,c1f2b7e6ac9f56a2bb44fa7a48b66dce

  • Two different tasks are implemented:

    • cache_channel_metadata_task(key, channel, tree_id) to cache one single channel
    • cache_multiple_channels_metadata_task(channels) to cache multiple channels

Cache invalidation

Cache keys will be stored forever
The timestamp stored in the key will be used to retrieve data, but will be marked as stale and recalculated after one hour (following a probabilistic cache invalidation algorithm to avoid a recalculation stampede)

@codecov

codecovBot commented Sep 28, 2020

Copy link
Copy Markdown

Codecov Report

Merging #2333 into develop will decrease coverage by 1.74%.
The diff coverage is 36.06%.

Impacted file tree graph

@@ Coverage Diff @@## develop #2333 +/- ##
===========================================
- Coverage 81.46% 79.72% -1.75% 
===========================================
Files 293 295 +2 Lines 14106 14587 +481 ===========================================
+ Hits 11491 11629 +138 - Misses 2615 2958 +343 
Impacted FilesCoverage Δ
...ontentcuration/contentcuration/viewsets/channel.py73.54% <25.58%> (-8.36%)⬇️
contentcuration/contentcuration/utils/channel.py34.21% <34.21%> (ø)
contentcuration/contentcuration/utils/cache.py43.90% <42.42%> (-19.74%)⬇️
contentcuration/contentcuration/tasks.py72.88% <66.66%> (-0.34%)⬇️
contentcuration/contentcuration/models.py86.99% <100.00%> (+0.16%)⬆️
...contentcuration/tests/viewsets/test_contentnode.py66.85% <0.00%> (-19.71%)⬇️
contentcuration/contentcuration/utils/nodes.py65.93% <0.00%> (-13.54%)⬇️
...uration/contentcuration/tests/test_contentnodes.py97.21% <0.00%> (-0.05%)⬇️
contentcuration/contentcuration/views/base.py55.27% <0.00%> (-0.05%)⬇️
contentcuration/contentcuration/test_settings.py100.00% <0.00%> (ø)
... and 8 more

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 1bf3f74...6a1a4e7. Read the comment docs.

@rtibblesrtibbles left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some questions, some ignorable suggestions, and one suggested extension.

Comment threadcontentcuration/contentcuration/utils/channel.py Outdated
Comment threadcontentcuration/contentcuration/viewsets/channel.py
Comment threadcontentcuration/contentcuration/viewsets/channel.py Outdated
Comment threadcontentcuration/contentcuration/viewsets/channel.py Outdated
Comment threadcontentcuration/contentcuration/utils/cache.py Outdated
Comment threadcontentcuration/contentcuration/viewsets/channel.py
ids = request.GET.get("id__in")
if not ids:
raise ValidationError("id__in GET parameter is required")
ids = ids.split(",")

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not necessary for the admin viewset, but if we generalize this we should probably do a quick filter here by the channels the user has view permissions for, so something like:

ids = self.get_queryset().filter(id__in=ids).values_list("id", flat=True)

Will then filter down the id list to the ones that the user has permissions for (in the admin use case, this is less important though).

@rtibbles

Copy link
Copy Markdown
Member

Just need to do a manual test, but this looks good to merge and build on!

@rtibblesrtibbles left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One comment, one question.



def cache_stampede(expire, beta=1):
"""Cache decorator with cache stampede protection.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we add a comment with a link to where we vendored it from?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I added a link to the pdf of the paper and its explanation. I'll add it to the diskcache code too, I stole some ideas from there (mainly the decorator usage) but for the main logic I used the paper. This implementation is simpler than diskcache's.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

OK, if you didn't lean on diskcache's code that much, this is fine then!

return result, delta

@functools.wraps(func)
def wrapper(*args, **kwargs):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we want to reuse this decorator, guessing we want to be able to parameterize the cache key template - but I think everything else could stay the same (except channel_id would be a more generic id variable name). Thoughts?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes, because the key is fetched from the decorated function, the change should be calling it with the final key, applying the parametrization before, it should be an easy change. Do you want me to do it now, so it's prepare for it?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sure, if it's no bother!

@rtibblesrtibbles left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's merge and iterate!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jredrejo@rtibbles
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Cache channel metadata - #2333

Merged
rtibbles merged 16 commits into
learningequality:developfrom
jredrejo:delay_channel_size
Sep 30, 2020
Merged

Cache channel metadata#2333
rtibbles merged 16 commits into
learningequality:developfrom
jredrejo:delay_channel_size

Conversation

@jredrejo

@jredrejojredrejo commented Sep 28, 2020

Copy link
Copy Markdown
Contributor

Description

Removes AdminChannelViewSet annotations and defers them to be executed in an async task that will cache the results.

As operations are executed asynchronously, the values are not immediately available if they have not been previously cached. In this case the values are marked with a DEFERRED_FLAG so the frontend knows that it can request them later.

The structure of the created object can be seen in the contentcuration/contentcuration/utils/channel.py and it's done according to the https://www.notion.so/learningequality/2020-12-24-Team-Sonic-counting-descendants-meeting-7ea0f46d4a6648109a22eefbba124a92 decissions

NOTE: some small optimization have also been done, mainly removing unneded order_by requirements in some of the queries, and the use of CTE to avoid scanning the whole File table after paginating the results.

Issue Addressed (if applicable)

Current annotations in AdminChannelViewSet returns queries needing days to be executed (File table has 76 millions of rows and ContentNode has 10 millions and the queries where scanning and joining both tables)

Implementation Notes (optional)

  • Added an api endpoint for the frontend to retrieve cached metadata.

Example of use: http://localhost:8080/api/admin-channels/deferred_data?id__in=0598c68f00fc562486fc6616b63f267f,c1f2b7e6ac9f56a2bb44fa7a48b66dce

  • Two different tasks are implemented:

    • cache_channel_metadata_task(key, channel, tree_id) to cache one single channel
    • cache_multiple_channels_metadata_task(channels) to cache multiple channels

Cache invalidation

Cache keys will be stored forever
The timestamp stored in the key will be used to retrieve data, but will be marked as stale and recalculated after one hour (following a probabilistic cache invalidation algorithm to avoid a recalculation stampede)

@codecov

codecovBot commented Sep 28, 2020

Copy link
Copy Markdown

Codecov Report

Merging #2333 into develop will decrease coverage by 1.74%.
The diff coverage is 36.06%.

Impacted file tree graph

@@ Coverage Diff @@## develop #2333 +/- ##
===========================================
- Coverage 81.46% 79.72% -1.75% 
===========================================
Files 293 295 +2 Lines 14106 14587 +481 ===========================================
+ Hits 11491 11629 +138 - Misses 2615 2958 +343 
Impacted FilesCoverage Δ
...ontentcuration/contentcuration/viewsets/channel.py73.54% <25.58%> (-8.36%)⬇️
contentcuration/contentcuration/utils/channel.py34.21% <34.21%> (ø)
contentcuration/contentcuration/utils/cache.py43.90% <42.42%> (-19.74%)⬇️
contentcuration/contentcuration/tasks.py72.88% <66.66%> (-0.34%)⬇️
contentcuration/contentcuration/models.py86.99% <100.00%> (+0.16%)⬆️
...contentcuration/tests/viewsets/test_contentnode.py66.85% <0.00%> (-19.71%)⬇️
contentcuration/contentcuration/utils/nodes.py65.93% <0.00%> (-13.54%)⬇️
...uration/contentcuration/tests/test_contentnodes.py97.21% <0.00%> (-0.05%)⬇️
contentcuration/contentcuration/views/base.py55.27% <0.00%> (-0.05%)⬇️
contentcuration/contentcuration/test_settings.py100.00% <0.00%> (ø)
... and 8 more

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 1bf3f74...6a1a4e7. Read the comment docs.

@rtibblesrtibbles left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some questions, some ignorable suggestions, and one suggested extension.

Comment threadcontentcuration/contentcuration/utils/channel.py Outdated
Comment threadcontentcuration/contentcuration/viewsets/channel.py
Comment threadcontentcuration/contentcuration/viewsets/channel.py Outdated
Comment threadcontentcuration/contentcuration/viewsets/channel.py Outdated
Comment threadcontentcuration/contentcuration/utils/cache.py Outdated
Comment threadcontentcuration/contentcuration/viewsets/channel.py
ids = request.GET.get("id__in")
if not ids:
raise ValidationError("id__in GET parameter is required")
ids = ids.split(",")

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not necessary for the admin viewset, but if we generalize this we should probably do a quick filter here by the channels the user has view permissions for, so something like:

ids = self.get_queryset().filter(id__in=ids).values_list("id", flat=True)

Will then filter down the id list to the ones that the user has permissions for (in the admin use case, this is less important though).

@rtibbles

Copy link
Copy Markdown
Member

Just need to do a manual test, but this looks good to merge and build on!

@rtibblesrtibbles left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One comment, one question.



def cache_stampede(expire, beta=1):
"""Cache decorator with cache stampede protection.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we add a comment with a link to where we vendored it from?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I added a link to the pdf of the paper and its explanation. I'll add it to the diskcache code too, I stole some ideas from there (mainly the decorator usage) but for the main logic I used the paper. This implementation is simpler than diskcache's.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

OK, if you didn't lean on diskcache's code that much, this is fine then!

return result, delta

@functools.wraps(func)
def wrapper(*args, **kwargs):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we want to reuse this decorator, guessing we want to be able to parameterize the cache key template - but I think everything else could stay the same (except channel_id would be a more generic id variable name). Thoughts?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes, because the key is fetched from the decorated function, the change should be calling it with the final key, applying the parametrization before, it should be an easy change. Do you want me to do it now, so it's prepare for it?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sure, if it's no bother!

@rtibblesrtibbles left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's merge and iterate!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jredrejo@rtibbles
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Cache channel metadata - #2333

Merged
rtibbles merged 16 commits into
learningequality:developfrom
jredrejo:delay_channel_size
Sep 30, 2020
Merged

Cache channel metadata#2333
rtibbles merged 16 commits into
learningequality:developfrom
jredrejo:delay_channel_size

Conversation

@jredrejo

@jredrejojredrejo commented Sep 28, 2020

Copy link
Copy Markdown
Contributor

Description

Removes AdminChannelViewSet annotations and defers them to be executed in an async task that will cache the results.

As operations are executed asynchronously, the values are not immediately available if they have not been previously cached. In this case the values are marked with a DEFERRED_FLAG so the frontend knows that it can request them later.

The structure of the created object can be seen in the contentcuration/contentcuration/utils/channel.py and it's done according to the https://www.notion.so/learningequality/2020-12-24-Team-Sonic-counting-descendants-meeting-7ea0f46d4a6648109a22eefbba124a92 decissions

NOTE: some small optimization have also been done, mainly removing unneded order_by requirements in some of the queries, and the use of CTE to avoid scanning the whole File table after paginating the results.

Issue Addressed (if applicable)

Current annotations in AdminChannelViewSet returns queries needing days to be executed (File table has 76 millions of rows and ContentNode has 10 millions and the queries where scanning and joining both tables)

Implementation Notes (optional)

  • Added an api endpoint for the frontend to retrieve cached metadata.

Example of use: http://localhost:8080/api/admin-channels/deferred_data?id__in=0598c68f00fc562486fc6616b63f267f,c1f2b7e6ac9f56a2bb44fa7a48b66dce

  • Two different tasks are implemented:

    • cache_channel_metadata_task(key, channel, tree_id) to cache one single channel
    • cache_multiple_channels_metadata_task(channels) to cache multiple channels

Cache invalidation

Cache keys will be stored forever
The timestamp stored in the key will be used to retrieve data, but will be marked as stale and recalculated after one hour (following a probabilistic cache invalidation algorithm to avoid a recalculation stampede)

@codecov

codecovBot commented Sep 28, 2020

Copy link
Copy Markdown

Codecov Report

Merging #2333 into develop will decrease coverage by 1.74%.
The diff coverage is 36.06%.

Impacted file tree graph

@@ Coverage Diff @@## develop #2333 +/- ##
===========================================
- Coverage 81.46% 79.72% -1.75% 
===========================================
Files 293 295 +2 Lines 14106 14587 +481 ===========================================
+ Hits 11491 11629 +138 - Misses 2615 2958 +343 
Impacted FilesCoverage Δ
...ontentcuration/contentcuration/viewsets/channel.py73.54% <25.58%> (-8.36%)⬇️
contentcuration/contentcuration/utils/channel.py34.21% <34.21%> (ø)
contentcuration/contentcuration/utils/cache.py43.90% <42.42%> (-19.74%)⬇️
contentcuration/contentcuration/tasks.py72.88% <66.66%> (-0.34%)⬇️
contentcuration/contentcuration/models.py86.99% <100.00%> (+0.16%)⬆️
...contentcuration/tests/viewsets/test_contentnode.py66.85% <0.00%> (-19.71%)⬇️
contentcuration/contentcuration/utils/nodes.py65.93% <0.00%> (-13.54%)⬇️
...uration/contentcuration/tests/test_contentnodes.py97.21% <0.00%> (-0.05%)⬇️
contentcuration/contentcuration/views/base.py55.27% <0.00%> (-0.05%)⬇️
contentcuration/contentcuration/test_settings.py100.00% <0.00%> (ø)
... and 8 more

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 1bf3f74...6a1a4e7. Read the comment docs.

@rtibblesrtibbles left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some questions, some ignorable suggestions, and one suggested extension.

Comment threadcontentcuration/contentcuration/utils/channel.py Outdated
Comment threadcontentcuration/contentcuration/viewsets/channel.py
Comment threadcontentcuration/contentcuration/viewsets/channel.py Outdated
Comment threadcontentcuration/contentcuration/viewsets/channel.py Outdated
Comment threadcontentcuration/contentcuration/utils/cache.py Outdated
Comment threadcontentcuration/contentcuration/viewsets/channel.py
ids = request.GET.get("id__in")
if not ids:
raise ValidationError("id__in GET parameter is required")
ids = ids.split(",")

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not necessary for the admin viewset, but if we generalize this we should probably do a quick filter here by the channels the user has view permissions for, so something like:

ids = self.get_queryset().filter(id__in=ids).values_list("id", flat=True)

Will then filter down the id list to the ones that the user has permissions for (in the admin use case, this is less important though).

@rtibbles

Copy link
Copy Markdown
Member

Just need to do a manual test, but this looks good to merge and build on!

@rtibblesrtibbles left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One comment, one question.



def cache_stampede(expire, beta=1):
"""Cache decorator with cache stampede protection.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we add a comment with a link to where we vendored it from?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I added a link to the pdf of the paper and its explanation. I'll add it to the diskcache code too, I stole some ideas from there (mainly the decorator usage) but for the main logic I used the paper. This implementation is simpler than diskcache's.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

OK, if you didn't lean on diskcache's code that much, this is fine then!

return result, delta

@functools.wraps(func)
def wrapper(*args, **kwargs):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we want to reuse this decorator, guessing we want to be able to parameterize the cache key template - but I think everything else could stay the same (except channel_id would be a more generic id variable name). Thoughts?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes, because the key is fetched from the decorated function, the change should be calling it with the final key, applying the parametrization before, it should be an easy change. Do you want me to do it now, so it's prepare for it?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sure, if it's no bother!

@rtibblesrtibbles left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's merge and iterate!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jredrejo@rtibbles
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Cache channel metadata - #2333

Merged
rtibbles merged 16 commits into
learningequality:developfrom
jredrejo:delay_channel_size
Sep 30, 2020
Merged

Cache channel metadata#2333
rtibbles merged 16 commits into
learningequality:developfrom
jredrejo:delay_channel_size

Conversation

@jredrejo

@jredrejojredrejo commented Sep 28, 2020

Copy link
Copy Markdown
Contributor

Description

Removes AdminChannelViewSet annotations and defers them to be executed in an async task that will cache the results.

As operations are executed asynchronously, the values are not immediately available if they have not been previously cached. In this case the values are marked with a DEFERRED_FLAG so the frontend knows that it can request them later.

The structure of the created object can be seen in the contentcuration/contentcuration/utils/channel.py and it's done according to the https://www.notion.so/learningequality/2020-12-24-Team-Sonic-counting-descendants-meeting-7ea0f46d4a6648109a22eefbba124a92 decissions

NOTE: some small optimization have also been done, mainly removing unneded order_by requirements in some of the queries, and the use of CTE to avoid scanning the whole File table after paginating the results.

Issue Addressed (if applicable)

Current annotations in AdminChannelViewSet returns queries needing days to be executed (File table has 76 millions of rows and ContentNode has 10 millions and the queries where scanning and joining both tables)

Implementation Notes (optional)

  • Added an api endpoint for the frontend to retrieve cached metadata.

Example of use: http://localhost:8080/api/admin-channels/deferred_data?id__in=0598c68f00fc562486fc6616b63f267f,c1f2b7e6ac9f56a2bb44fa7a48b66dce

  • Two different tasks are implemented:

    • cache_channel_metadata_task(key, channel, tree_id) to cache one single channel
    • cache_multiple_channels_metadata_task(channels) to cache multiple channels

Cache invalidation

Cache keys will be stored forever
The timestamp stored in the key will be used to retrieve data, but will be marked as stale and recalculated after one hour (following a probabilistic cache invalidation algorithm to avoid a recalculation stampede)

@codecov

codecovBot commented Sep 28, 2020

Copy link
Copy Markdown

Codecov Report

Merging #2333 into develop will decrease coverage by 1.74%.
The diff coverage is 36.06%.

Impacted file tree graph

@@ Coverage Diff @@## develop #2333 +/- ##
===========================================
- Coverage 81.46% 79.72% -1.75% 
===========================================
Files 293 295 +2 Lines 14106 14587 +481 ===========================================
+ Hits 11491 11629 +138 - Misses 2615 2958 +343 
Impacted FilesCoverage Δ
...ontentcuration/contentcuration/viewsets/channel.py73.54% <25.58%> (-8.36%)⬇️
contentcuration/contentcuration/utils/channel.py34.21% <34.21%> (ø)
contentcuration/contentcuration/utils/cache.py43.90% <42.42%> (-19.74%)⬇️
contentcuration/contentcuration/tasks.py72.88% <66.66%> (-0.34%)⬇️
contentcuration/contentcuration/models.py86.99% <100.00%> (+0.16%)⬆️
...contentcuration/tests/viewsets/test_contentnode.py66.85% <0.00%> (-19.71%)⬇️
contentcuration/contentcuration/utils/nodes.py65.93% <0.00%> (-13.54%)⬇️
...uration/contentcuration/tests/test_contentnodes.py97.21% <0.00%> (-0.05%)⬇️
contentcuration/contentcuration/views/base.py55.27% <0.00%> (-0.05%)⬇️
contentcuration/contentcuration/test_settings.py100.00% <0.00%> (ø)
... and 8 more

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 1bf3f74...6a1a4e7. Read the comment docs.

@rtibblesrtibbles left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some questions, some ignorable suggestions, and one suggested extension.

Comment threadcontentcuration/contentcuration/utils/channel.py Outdated
Comment threadcontentcuration/contentcuration/viewsets/channel.py
Comment threadcontentcuration/contentcuration/viewsets/channel.py Outdated
Comment threadcontentcuration/contentcuration/viewsets/channel.py Outdated
Comment threadcontentcuration/contentcuration/utils/cache.py Outdated
Comment threadcontentcuration/contentcuration/viewsets/channel.py
ids = request.GET.get("id__in")
if not ids:
raise ValidationError("id__in GET parameter is required")
ids = ids.split(",")

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not necessary for the admin viewset, but if we generalize this we should probably do a quick filter here by the channels the user has view permissions for, so something like:

ids = self.get_queryset().filter(id__in=ids).values_list("id", flat=True)

Will then filter down the id list to the ones that the user has permissions for (in the admin use case, this is less important though).

@rtibbles

Copy link
Copy Markdown
Member

Just need to do a manual test, but this looks good to merge and build on!

@rtibblesrtibbles left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One comment, one question.



def cache_stampede(expire, beta=1):
"""Cache decorator with cache stampede protection.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we add a comment with a link to where we vendored it from?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I added a link to the pdf of the paper and its explanation. I'll add it to the diskcache code too, I stole some ideas from there (mainly the decorator usage) but for the main logic I used the paper. This implementation is simpler than diskcache's.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

OK, if you didn't lean on diskcache's code that much, this is fine then!

return result, delta

@functools.wraps(func)
def wrapper(*args, **kwargs):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we want to reuse this decorator, guessing we want to be able to parameterize the cache key template - but I think everything else could stay the same (except channel_id would be a more generic id variable name). Thoughts?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes, because the key is fetched from the decorated function, the change should be calling it with the final key, applying the parametrization before, it should be an easy change. Do you want me to do it now, so it's prepare for it?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sure, if it's no bother!

@rtibblesrtibbles left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's merge and iterate!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jredrejo@rtibbles
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Cache channel metadata - #2333

Merged
rtibbles merged 16 commits into
learningequality:developfrom
jredrejo:delay_channel_size
Sep 30, 2020
Merged

Cache channel metadata#2333
rtibbles merged 16 commits into
learningequality:developfrom
jredrejo:delay_channel_size

Conversation

@jredrejo

@jredrejojredrejo commented Sep 28, 2020

Copy link
Copy Markdown
Contributor

Description

Removes AdminChannelViewSet annotations and defers them to be executed in an async task that will cache the results.

As operations are executed asynchronously, the values are not immediately available if they have not been previously cached. In this case the values are marked with a DEFERRED_FLAG so the frontend knows that it can request them later.

The structure of the created object can be seen in the contentcuration/contentcuration/utils/channel.py and it's done according to the https://www.notion.so/learningequality/2020-12-24-Team-Sonic-counting-descendants-meeting-7ea0f46d4a6648109a22eefbba124a92 decissions

NOTE: some small optimization have also been done, mainly removing unneded order_by requirements in some of the queries, and the use of CTE to avoid scanning the whole File table after paginating the results.

Issue Addressed (if applicable)

Current annotations in AdminChannelViewSet returns queries needing days to be executed (File table has 76 millions of rows and ContentNode has 10 millions and the queries where scanning and joining both tables)

Implementation Notes (optional)

  • Added an api endpoint for the frontend to retrieve cached metadata.

Example of use: http://localhost:8080/api/admin-channels/deferred_data?id__in=0598c68f00fc562486fc6616b63f267f,c1f2b7e6ac9f56a2bb44fa7a48b66dce

  • Two different tasks are implemented:

    • cache_channel_metadata_task(key, channel, tree_id) to cache one single channel
    • cache_multiple_channels_metadata_task(channels) to cache multiple channels

Cache invalidation

Cache keys will be stored forever
The timestamp stored in the key will be used to retrieve data, but will be marked as stale and recalculated after one hour (following a probabilistic cache invalidation algorithm to avoid a recalculation stampede)

@codecov

codecovBot commented Sep 28, 2020

Copy link
Copy Markdown

Codecov Report

Merging #2333 into develop will decrease coverage by 1.74%.
The diff coverage is 36.06%.

Impacted file tree graph

@@ Coverage Diff @@## develop #2333 +/- ##
===========================================
- Coverage 81.46% 79.72% -1.75% 
===========================================
Files 293 295 +2 Lines 14106 14587 +481 ===========================================
+ Hits 11491 11629 +138 - Misses 2615 2958 +343 
Impacted FilesCoverage Δ
...ontentcuration/contentcuration/viewsets/channel.py73.54% <25.58%> (-8.36%)⬇️
contentcuration/contentcuration/utils/channel.py34.21% <34.21%> (ø)
contentcuration/contentcuration/utils/cache.py43.90% <42.42%> (-19.74%)⬇️
contentcuration/contentcuration/tasks.py72.88% <66.66%> (-0.34%)⬇️
contentcuration/contentcuration/models.py86.99% <100.00%> (+0.16%)⬆️
...contentcuration/tests/viewsets/test_contentnode.py66.85% <0.00%> (-19.71%)⬇️
contentcuration/contentcuration/utils/nodes.py65.93% <0.00%> (-13.54%)⬇️
...uration/contentcuration/tests/test_contentnodes.py97.21% <0.00%> (-0.05%)⬇️
contentcuration/contentcuration/views/base.py55.27% <0.00%> (-0.05%)⬇️
contentcuration/contentcuration/test_settings.py100.00% <0.00%> (ø)
... and 8 more

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 1bf3f74...6a1a4e7. Read the comment docs.

@rtibblesrtibbles left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some questions, some ignorable suggestions, and one suggested extension.

Comment threadcontentcuration/contentcuration/utils/channel.py Outdated
Comment threadcontentcuration/contentcuration/viewsets/channel.py
Comment threadcontentcuration/contentcuration/viewsets/channel.py Outdated
Comment threadcontentcuration/contentcuration/viewsets/channel.py Outdated
Comment threadcontentcuration/contentcuration/utils/cache.py Outdated
Comment threadcontentcuration/contentcuration/viewsets/channel.py
ids = request.GET.get("id__in")
if not ids:
raise ValidationError("id__in GET parameter is required")
ids = ids.split(",")

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not necessary for the admin viewset, but if we generalize this we should probably do a quick filter here by the channels the user has view permissions for, so something like:

ids = self.get_queryset().filter(id__in=ids).values_list("id", flat=True)

Will then filter down the id list to the ones that the user has permissions for (in the admin use case, this is less important though).

@rtibbles

Copy link
Copy Markdown
Member

Just need to do a manual test, but this looks good to merge and build on!

@rtibblesrtibbles left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One comment, one question.



def cache_stampede(expire, beta=1):
"""Cache decorator with cache stampede protection.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we add a comment with a link to where we vendored it from?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I added a link to the pdf of the paper and its explanation. I'll add it to the diskcache code too, I stole some ideas from there (mainly the decorator usage) but for the main logic I used the paper. This implementation is simpler than diskcache's.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

OK, if you didn't lean on diskcache's code that much, this is fine then!

return result, delta

@functools.wraps(func)
def wrapper(*args, **kwargs):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we want to reuse this decorator, guessing we want to be able to parameterize the cache key template - but I think everything else could stay the same (except channel_id would be a more generic id variable name). Thoughts?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes, because the key is fetched from the decorated function, the change should be calling it with the final key, applying the parametrization before, it should be an easy change. Do you want me to do it now, so it's prepare for it?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sure, if it's no bother!

@rtibblesrtibbles left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's merge and iterate!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jredrejo@rtibbles