[WIP] Attempt to continue LRU cache for decoded chunks - #1214

Closed
croth1 wants to merge 58 commits into
zarr-developers:mainfrom
croth1:chunk_cache
Closed

[WIP] Attempt to continue LRU cache for decoded chunks#1214
croth1 wants to merge 58 commits into
zarr-developers:mainfrom
croth1:chunk_cache

Conversation

@croth1

@croth1croth1 commented Oct 23, 2022

Copy link
Copy Markdown

Continuation attempt of #306 - do not merge yet - very likely still not correct behaviour when chunks get deleted from store by #738.

Still have very limited knowledge of the code base - will require me to dig a bit deeper to gain bit more understanding.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

shikharsgand others added 30 commits August 29, 2018 13:17
… decode round trip for object arrays(with tests)
factoring out mapping code from LRUStoreCache and LRUChunkCache
@croth1

croth1 commented Oct 23, 2022

Copy link
Copy Markdown
Author

Hmh, ok - the diff of the merge commit looks still a bit insane - some code from other PRs somehow leaked in/got duplicated - or I am just bad at interpreting the diff. Will investigate tomorrow what happened there.

@@ -65,10 +117,6 @@ def create_store(self, **kwargs): # pragma: no cover
# implement in sub-class

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

IIRC, I needed to move this one, because LRUChunkCache doesn't implement the contextmanager interface - maybe implement it instead?

Comment threadzarr/core.py
@@ -87,6 +87,15 @@ class Array:
read and decompressed when possible.

.. versionadded:: 2.7

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this will need to get updated.

Comment threadzarr/creation.py
non-fill-value data are stored, at the expense of overhead associated
with checking the data of each chunk.

.. versionadded:: 2.7

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

needs to get updated

@joshmoore

Copy link
Copy Markdown
Member

Kicked off the GHA workflows.

@codecov

codecovBot commented Oct 24, 2022

Copy link
Copy Markdown

Codecov Report

All modified and coverable lines are covered by tests ✅

Project coverage is 99.99%. Comparing base (f361631) to head (16793f5).
Report is 720 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #1214 +/- ##
========================================
Coverage 99.99% 99.99% ========================================
Files 35 35 Lines 14136 14366 +230 ========================================
+ Hits 14135 14365 +230 
Misses 1 1 
Files with missing linesCoverage Δ
zarr/__init__.py100.00% <ø> (ø)
zarr/core.py100.00% <100.00%> (ø)
zarr/creation.py100.00% <ø> (ø)
zarr/hierarchy.py99.78% <100.00%> (+<0.01%)⬆️
zarr/storage.py100.00% <100.00%> (ø)
zarr/tests/test_core.py100.00% <100.00%> (ø)
zarr/tests/test_storage.py100.00% <100.00%> (ø)

@joshmoore

Copy link
Copy Markdown
Member

Green. Are there any other commits expected from your side, @croth1?

@croth1

Copy link
Copy Markdown
Author

I am still not super sure whether this is the final form or I would rather re-implement it as a caching store ontop of other stores, just like the LRUStoreCache - it feels a bit more intuitive to me. Also I still want to check whether I can do write caching. If I remember correctly, this implementation is write-through. Not sure whether in combination with LRUStoreCache cached writes can be achieved. Need to do more research, but I'm a bit busy right now.

@FarzanT

FarzanT commented May 20, 2023

Copy link
Copy Markdown

Hi, any news on when this feature will get pushed to main? It seems like a very useful feature!
I want to reduce the overhead of sampling from the same chunk which is expected to be in RAM. If I'm not mistaken, right now, every time you pass an index to the zarr array, it has to decompress that index. I want to have an option to keep the entire chunk decompressed in the cache until it's discarded. Specifically, I'm using a large zarr array in my pytorch dataset module, and need to randomly iterate through samples in the same chunk before randomly picking the next chunk. Avoiding the decompression overhead is quite useful, and unfortunately I can't leave the data fully decompressed on disk, as it would be too large.

@croth1

Copy link
Copy Markdown
Author

@FarzanT@joshmoore, I have in the meantime changed my approach and would not need this anymore - hence the long silence. If that's useful for other people, we can try finishing it up as is. Would need some rebasing and performance testing, though.

IIRC, last time I checked all tests were green, although this would need a very thorough review because frequently I was not 100% sure what I was doing during the quite substantial rebase.

@joshmoore

Copy link
Copy Markdown
Member

I'll leave @FarzanT and others to say how pressing their need is. Happy to help how I can @croth1 if you'd like to pursue this.

@FarzanT

Copy link
Copy Markdown

Thank you @croth1 and @joshmoore, my primary use case has also been addressed by #278 (comment), so it's not a pressing issue for me at the moment. But I'd say it would be much nicer to just flip a switch and have zarr handle this internally. If this pull request can address this, then it shouldn't be abandoned IMO!

@jhamman

Copy link
Copy Markdown
Member

I'm going to close this as stale. Folks should feel free to reopen if there is interest in continuing this work.

@jhammanjhamman closed this Oct 11, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@croth1@joshmoore@FarzanT@jhamman@shikharsg@jakirkham
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

[WIP] Attempt to continue LRU cache for decoded chunks - #1214

Closed
croth1 wants to merge 58 commits into
zarr-developers:mainfrom
croth1:chunk_cache
Closed

[WIP] Attempt to continue LRU cache for decoded chunks#1214
croth1 wants to merge 58 commits into
zarr-developers:mainfrom
croth1:chunk_cache

Conversation

@croth1

@croth1croth1 commented Oct 23, 2022

Copy link
Copy Markdown

Continuation attempt of #306 - do not merge yet - very likely still not correct behaviour when chunks get deleted from store by #738.

Still have very limited knowledge of the code base - will require me to dig a bit deeper to gain bit more understanding.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

shikharsgand others added 30 commits August 29, 2018 13:17
… decode round trip for object arrays(with tests)
factoring out mapping code from LRUStoreCache and LRUChunkCache
@croth1

croth1 commented Oct 23, 2022

Copy link
Copy Markdown
Author

Hmh, ok - the diff of the merge commit looks still a bit insane - some code from other PRs somehow leaked in/got duplicated - or I am just bad at interpreting the diff. Will investigate tomorrow what happened there.

@@ -65,10 +117,6 @@ def create_store(self, **kwargs): # pragma: no cover
# implement in sub-class

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

IIRC, I needed to move this one, because LRUChunkCache doesn't implement the contextmanager interface - maybe implement it instead?

Comment threadzarr/core.py
@@ -87,6 +87,15 @@ class Array:
read and decompressed when possible.

.. versionadded:: 2.7

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this will need to get updated.

Comment threadzarr/creation.py
non-fill-value data are stored, at the expense of overhead associated
with checking the data of each chunk.

.. versionadded:: 2.7

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

needs to get updated

@joshmoore

Copy link
Copy Markdown
Member

Kicked off the GHA workflows.

@codecov

codecovBot commented Oct 24, 2022

Copy link
Copy Markdown

Codecov Report

All modified and coverable lines are covered by tests ✅

Project coverage is 99.99%. Comparing base (f361631) to head (16793f5).
Report is 720 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #1214 +/- ##
========================================
Coverage 99.99% 99.99% ========================================
Files 35 35 Lines 14136 14366 +230 ========================================
+ Hits 14135 14365 +230 
Misses 1 1 
Files with missing linesCoverage Δ
zarr/__init__.py100.00% <ø> (ø)
zarr/core.py100.00% <100.00%> (ø)
zarr/creation.py100.00% <ø> (ø)
zarr/hierarchy.py99.78% <100.00%> (+<0.01%)⬆️
zarr/storage.py100.00% <100.00%> (ø)
zarr/tests/test_core.py100.00% <100.00%> (ø)
zarr/tests/test_storage.py100.00% <100.00%> (ø)

@joshmoore

Copy link
Copy Markdown
Member

Green. Are there any other commits expected from your side, @croth1?

@croth1

Copy link
Copy Markdown
Author

I am still not super sure whether this is the final form or I would rather re-implement it as a caching store ontop of other stores, just like the LRUStoreCache - it feels a bit more intuitive to me. Also I still want to check whether I can do write caching. If I remember correctly, this implementation is write-through. Not sure whether in combination with LRUStoreCache cached writes can be achieved. Need to do more research, but I'm a bit busy right now.

@FarzanT

FarzanT commented May 20, 2023

Copy link
Copy Markdown

Hi, any news on when this feature will get pushed to main? It seems like a very useful feature!
I want to reduce the overhead of sampling from the same chunk which is expected to be in RAM. If I'm not mistaken, right now, every time you pass an index to the zarr array, it has to decompress that index. I want to have an option to keep the entire chunk decompressed in the cache until it's discarded. Specifically, I'm using a large zarr array in my pytorch dataset module, and need to randomly iterate through samples in the same chunk before randomly picking the next chunk. Avoiding the decompression overhead is quite useful, and unfortunately I can't leave the data fully decompressed on disk, as it would be too large.

@croth1

Copy link
Copy Markdown
Author

@FarzanT@joshmoore, I have in the meantime changed my approach and would not need this anymore - hence the long silence. If that's useful for other people, we can try finishing it up as is. Would need some rebasing and performance testing, though.

IIRC, last time I checked all tests were green, although this would need a very thorough review because frequently I was not 100% sure what I was doing during the quite substantial rebase.

@joshmoore

Copy link
Copy Markdown
Member

I'll leave @FarzanT and others to say how pressing their need is. Happy to help how I can @croth1 if you'd like to pursue this.

@FarzanT

Copy link
Copy Markdown

Thank you @croth1 and @joshmoore, my primary use case has also been addressed by #278 (comment), so it's not a pressing issue for me at the moment. But I'd say it would be much nicer to just flip a switch and have zarr handle this internally. If this pull request can address this, then it shouldn't be abandoned IMO!

@jhamman

Copy link
Copy Markdown
Member

I'm going to close this as stale. Folks should feel free to reopen if there is interest in continuing this work.

@jhammanjhamman closed this Oct 11, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@croth1@joshmoore@FarzanT@jhamman@shikharsg@jakirkham
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[WIP] Attempt to continue LRU cache for decoded chunks - #1214

Closed
croth1 wants to merge 58 commits into
zarr-developers:mainfrom
croth1:chunk_cache
Closed

[WIP] Attempt to continue LRU cache for decoded chunks#1214
croth1 wants to merge 58 commits into
zarr-developers:mainfrom
croth1:chunk_cache

Conversation

@croth1

@croth1croth1 commented Oct 23, 2022

Copy link
Copy Markdown

Continuation attempt of #306 - do not merge yet - very likely still not correct behaviour when chunks get deleted from store by #738.

Still have very limited knowledge of the code base - will require me to dig a bit deeper to gain bit more understanding.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

shikharsgand others added 30 commits August 29, 2018 13:17
… decode round trip for object arrays(with tests)
factoring out mapping code from LRUStoreCache and LRUChunkCache
@croth1

croth1 commented Oct 23, 2022

Copy link
Copy Markdown
Author

Hmh, ok - the diff of the merge commit looks still a bit insane - some code from other PRs somehow leaked in/got duplicated - or I am just bad at interpreting the diff. Will investigate tomorrow what happened there.

@@ -65,10 +117,6 @@ def create_store(self, **kwargs): # pragma: no cover
# implement in sub-class

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

IIRC, I needed to move this one, because LRUChunkCache doesn't implement the contextmanager interface - maybe implement it instead?

Comment threadzarr/core.py
@@ -87,6 +87,15 @@ class Array:
read and decompressed when possible.

.. versionadded:: 2.7

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this will need to get updated.

Comment threadzarr/creation.py
non-fill-value data are stored, at the expense of overhead associated
with checking the data of each chunk.

.. versionadded:: 2.7

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

needs to get updated

@joshmoore

Copy link
Copy Markdown
Member

Kicked off the GHA workflows.

@codecov

codecovBot commented Oct 24, 2022

Copy link
Copy Markdown

Codecov Report

All modified and coverable lines are covered by tests ✅

Project coverage is 99.99%. Comparing base (f361631) to head (16793f5).
Report is 720 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #1214 +/- ##
========================================
Coverage 99.99% 99.99% ========================================
Files 35 35 Lines 14136 14366 +230 ========================================
+ Hits 14135 14365 +230 
Misses 1 1 
Files with missing linesCoverage Δ
zarr/__init__.py100.00% <ø> (ø)
zarr/core.py100.00% <100.00%> (ø)
zarr/creation.py100.00% <ø> (ø)
zarr/hierarchy.py99.78% <100.00%> (+<0.01%)⬆️
zarr/storage.py100.00% <100.00%> (ø)
zarr/tests/test_core.py100.00% <100.00%> (ø)
zarr/tests/test_storage.py100.00% <100.00%> (ø)

@joshmoore

Copy link
Copy Markdown
Member

Green. Are there any other commits expected from your side, @croth1?

@croth1

Copy link
Copy Markdown
Author

I am still not super sure whether this is the final form or I would rather re-implement it as a caching store ontop of other stores, just like the LRUStoreCache - it feels a bit more intuitive to me. Also I still want to check whether I can do write caching. If I remember correctly, this implementation is write-through. Not sure whether in combination with LRUStoreCache cached writes can be achieved. Need to do more research, but I'm a bit busy right now.

@FarzanT

FarzanT commented May 20, 2023

Copy link
Copy Markdown

Hi, any news on when this feature will get pushed to main? It seems like a very useful feature!
I want to reduce the overhead of sampling from the same chunk which is expected to be in RAM. If I'm not mistaken, right now, every time you pass an index to the zarr array, it has to decompress that index. I want to have an option to keep the entire chunk decompressed in the cache until it's discarded. Specifically, I'm using a large zarr array in my pytorch dataset module, and need to randomly iterate through samples in the same chunk before randomly picking the next chunk. Avoiding the decompression overhead is quite useful, and unfortunately I can't leave the data fully decompressed on disk, as it would be too large.

@croth1

Copy link
Copy Markdown
Author

@FarzanT@joshmoore, I have in the meantime changed my approach and would not need this anymore - hence the long silence. If that's useful for other people, we can try finishing it up as is. Would need some rebasing and performance testing, though.

IIRC, last time I checked all tests were green, although this would need a very thorough review because frequently I was not 100% sure what I was doing during the quite substantial rebase.

@joshmoore

Copy link
Copy Markdown
Member

I'll leave @FarzanT and others to say how pressing their need is. Happy to help how I can @croth1 if you'd like to pursue this.

@FarzanT

Copy link
Copy Markdown

Thank you @croth1 and @joshmoore, my primary use case has also been addressed by #278 (comment), so it's not a pressing issue for me at the moment. But I'd say it would be much nicer to just flip a switch and have zarr handle this internally. If this pull request can address this, then it shouldn't be abandoned IMO!

@jhamman

Copy link
Copy Markdown
Member

I'm going to close this as stale. Folks should feel free to reopen if there is interest in continuing this work.

@jhammanjhamman closed this Oct 11, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@croth1@joshmoore@FarzanT@jhamman@shikharsg@jakirkham
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[WIP] Attempt to continue LRU cache for decoded chunks - #1214

Closed
croth1 wants to merge 58 commits into
zarr-developers:mainfrom
croth1:chunk_cache
Closed

[WIP] Attempt to continue LRU cache for decoded chunks#1214
croth1 wants to merge 58 commits into
zarr-developers:mainfrom
croth1:chunk_cache

Conversation

@croth1

@croth1croth1 commented Oct 23, 2022

Copy link
Copy Markdown

Continuation attempt of #306 - do not merge yet - very likely still not correct behaviour when chunks get deleted from store by #738.

Still have very limited knowledge of the code base - will require me to dig a bit deeper to gain bit more understanding.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

shikharsgand others added 30 commits August 29, 2018 13:17
… decode round trip for object arrays(with tests)
factoring out mapping code from LRUStoreCache and LRUChunkCache
@croth1

croth1 commented Oct 23, 2022

Copy link
Copy Markdown
Author

Hmh, ok - the diff of the merge commit looks still a bit insane - some code from other PRs somehow leaked in/got duplicated - or I am just bad at interpreting the diff. Will investigate tomorrow what happened there.

@@ -65,10 +117,6 @@ def create_store(self, **kwargs): # pragma: no cover
# implement in sub-class

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

IIRC, I needed to move this one, because LRUChunkCache doesn't implement the contextmanager interface - maybe implement it instead?

Comment threadzarr/core.py
@@ -87,6 +87,15 @@ class Array:
read and decompressed when possible.

.. versionadded:: 2.7

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this will need to get updated.

Comment threadzarr/creation.py
non-fill-value data are stored, at the expense of overhead associated
with checking the data of each chunk.

.. versionadded:: 2.7

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

needs to get updated

@joshmoore

Copy link
Copy Markdown
Member

Kicked off the GHA workflows.

@codecov

codecovBot commented Oct 24, 2022

Copy link
Copy Markdown

Codecov Report

All modified and coverable lines are covered by tests ✅

Project coverage is 99.99%. Comparing base (f361631) to head (16793f5).
Report is 720 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #1214 +/- ##
========================================
Coverage 99.99% 99.99% ========================================
Files 35 35 Lines 14136 14366 +230 ========================================
+ Hits 14135 14365 +230 
Misses 1 1 
Files with missing linesCoverage Δ
zarr/__init__.py100.00% <ø> (ø)
zarr/core.py100.00% <100.00%> (ø)
zarr/creation.py100.00% <ø> (ø)
zarr/hierarchy.py99.78% <100.00%> (+<0.01%)⬆️
zarr/storage.py100.00% <100.00%> (ø)
zarr/tests/test_core.py100.00% <100.00%> (ø)
zarr/tests/test_storage.py100.00% <100.00%> (ø)

@joshmoore

Copy link
Copy Markdown
Member

Green. Are there any other commits expected from your side, @croth1?

@croth1

Copy link
Copy Markdown
Author

I am still not super sure whether this is the final form or I would rather re-implement it as a caching store ontop of other stores, just like the LRUStoreCache - it feels a bit more intuitive to me. Also I still want to check whether I can do write caching. If I remember correctly, this implementation is write-through. Not sure whether in combination with LRUStoreCache cached writes can be achieved. Need to do more research, but I'm a bit busy right now.

@FarzanT

FarzanT commented May 20, 2023

Copy link
Copy Markdown

Hi, any news on when this feature will get pushed to main? It seems like a very useful feature!
I want to reduce the overhead of sampling from the same chunk which is expected to be in RAM. If I'm not mistaken, right now, every time you pass an index to the zarr array, it has to decompress that index. I want to have an option to keep the entire chunk decompressed in the cache until it's discarded. Specifically, I'm using a large zarr array in my pytorch dataset module, and need to randomly iterate through samples in the same chunk before randomly picking the next chunk. Avoiding the decompression overhead is quite useful, and unfortunately I can't leave the data fully decompressed on disk, as it would be too large.

@croth1

Copy link
Copy Markdown
Author

@FarzanT@joshmoore, I have in the meantime changed my approach and would not need this anymore - hence the long silence. If that's useful for other people, we can try finishing it up as is. Would need some rebasing and performance testing, though.

IIRC, last time I checked all tests were green, although this would need a very thorough review because frequently I was not 100% sure what I was doing during the quite substantial rebase.

@joshmoore

Copy link
Copy Markdown
Member

I'll leave @FarzanT and others to say how pressing their need is. Happy to help how I can @croth1 if you'd like to pursue this.

@FarzanT

Copy link
Copy Markdown

Thank you @croth1 and @joshmoore, my primary use case has also been addressed by #278 (comment), so it's not a pressing issue for me at the moment. But I'd say it would be much nicer to just flip a switch and have zarr handle this internally. If this pull request can address this, then it shouldn't be abandoned IMO!

@jhamman

Copy link
Copy Markdown
Member

I'm going to close this as stale. Folks should feel free to reopen if there is interest in continuing this work.

@jhammanjhamman closed this Oct 11, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@croth1@joshmoore@FarzanT@jhamman@shikharsg@jakirkham
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

[WIP] Attempt to continue LRU cache for decoded chunks - #1214

Closed
croth1 wants to merge 58 commits into
zarr-developers:mainfrom
croth1:chunk_cache
Closed

[WIP] Attempt to continue LRU cache for decoded chunks#1214
croth1 wants to merge 58 commits into
zarr-developers:mainfrom
croth1:chunk_cache

Conversation

@croth1

@croth1croth1 commented Oct 23, 2022

Copy link
Copy Markdown

Continuation attempt of #306 - do not merge yet - very likely still not correct behaviour when chunks get deleted from store by #738.

Still have very limited knowledge of the code base - will require me to dig a bit deeper to gain bit more understanding.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

shikharsgand others added 30 commits August 29, 2018 13:17
… decode round trip for object arrays(with tests)
factoring out mapping code from LRUStoreCache and LRUChunkCache
@croth1

croth1 commented Oct 23, 2022

Copy link
Copy Markdown
Author

Hmh, ok - the diff of the merge commit looks still a bit insane - some code from other PRs somehow leaked in/got duplicated - or I am just bad at interpreting the diff. Will investigate tomorrow what happened there.

@@ -65,10 +117,6 @@ def create_store(self, **kwargs): # pragma: no cover
# implement in sub-class

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

IIRC, I needed to move this one, because LRUChunkCache doesn't implement the contextmanager interface - maybe implement it instead?

Comment threadzarr/core.py
@@ -87,6 +87,15 @@ class Array:
read and decompressed when possible.

.. versionadded:: 2.7

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this will need to get updated.

Comment threadzarr/creation.py
non-fill-value data are stored, at the expense of overhead associated
with checking the data of each chunk.

.. versionadded:: 2.7

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

needs to get updated

@joshmoore

Copy link
Copy Markdown
Member

Kicked off the GHA workflows.

@codecov

codecovBot commented Oct 24, 2022

Copy link
Copy Markdown

Codecov Report

All modified and coverable lines are covered by tests ✅

Project coverage is 99.99%. Comparing base (f361631) to head (16793f5).
Report is 720 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #1214 +/- ##
========================================
Coverage 99.99% 99.99% ========================================
Files 35 35 Lines 14136 14366 +230 ========================================
+ Hits 14135 14365 +230 
Misses 1 1 
Files with missing linesCoverage Δ
zarr/__init__.py100.00% <ø> (ø)
zarr/core.py100.00% <100.00%> (ø)
zarr/creation.py100.00% <ø> (ø)
zarr/hierarchy.py99.78% <100.00%> (+<0.01%)⬆️
zarr/storage.py100.00% <100.00%> (ø)
zarr/tests/test_core.py100.00% <100.00%> (ø)
zarr/tests/test_storage.py100.00% <100.00%> (ø)

@joshmoore

Copy link
Copy Markdown
Member

Green. Are there any other commits expected from your side, @croth1?

@croth1

Copy link
Copy Markdown
Author

I am still not super sure whether this is the final form or I would rather re-implement it as a caching store ontop of other stores, just like the LRUStoreCache - it feels a bit more intuitive to me. Also I still want to check whether I can do write caching. If I remember correctly, this implementation is write-through. Not sure whether in combination with LRUStoreCache cached writes can be achieved. Need to do more research, but I'm a bit busy right now.

@FarzanT

FarzanT commented May 20, 2023

Copy link
Copy Markdown

Hi, any news on when this feature will get pushed to main? It seems like a very useful feature!
I want to reduce the overhead of sampling from the same chunk which is expected to be in RAM. If I'm not mistaken, right now, every time you pass an index to the zarr array, it has to decompress that index. I want to have an option to keep the entire chunk decompressed in the cache until it's discarded. Specifically, I'm using a large zarr array in my pytorch dataset module, and need to randomly iterate through samples in the same chunk before randomly picking the next chunk. Avoiding the decompression overhead is quite useful, and unfortunately I can't leave the data fully decompressed on disk, as it would be too large.

@croth1

Copy link
Copy Markdown
Author

@FarzanT@joshmoore, I have in the meantime changed my approach and would not need this anymore - hence the long silence. If that's useful for other people, we can try finishing it up as is. Would need some rebasing and performance testing, though.

IIRC, last time I checked all tests were green, although this would need a very thorough review because frequently I was not 100% sure what I was doing during the quite substantial rebase.

@joshmoore

Copy link
Copy Markdown
Member

I'll leave @FarzanT and others to say how pressing their need is. Happy to help how I can @croth1 if you'd like to pursue this.

@FarzanT

Copy link
Copy Markdown

Thank you @croth1 and @joshmoore, my primary use case has also been addressed by #278 (comment), so it's not a pressing issue for me at the moment. But I'd say it would be much nicer to just flip a switch and have zarr handle this internally. If this pull request can address this, then it shouldn't be abandoned IMO!

@jhamman

Copy link
Copy Markdown
Member

I'm going to close this as stale. Folks should feel free to reopen if there is interest in continuing this work.

@jhammanjhamman closed this Oct 11, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@croth1@joshmoore@FarzanT@jhamman@shikharsg@jakirkham
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[WIP] Attempt to continue LRU cache for decoded chunks - #1214

Closed
croth1 wants to merge 58 commits into
zarr-developers:mainfrom
croth1:chunk_cache
Closed

[WIP] Attempt to continue LRU cache for decoded chunks#1214
croth1 wants to merge 58 commits into
zarr-developers:mainfrom
croth1:chunk_cache

Conversation

@croth1

@croth1croth1 commented Oct 23, 2022

Copy link
Copy Markdown

Continuation attempt of #306 - do not merge yet - very likely still not correct behaviour when chunks get deleted from store by #738.

Still have very limited knowledge of the code base - will require me to dig a bit deeper to gain bit more understanding.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

shikharsgand others added 30 commits August 29, 2018 13:17
… decode round trip for object arrays(with tests)
factoring out mapping code from LRUStoreCache and LRUChunkCache
@croth1

croth1 commented Oct 23, 2022

Copy link
Copy Markdown
Author

Hmh, ok - the diff of the merge commit looks still a bit insane - some code from other PRs somehow leaked in/got duplicated - or I am just bad at interpreting the diff. Will investigate tomorrow what happened there.

@@ -65,10 +117,6 @@ def create_store(self, **kwargs): # pragma: no cover
# implement in sub-class

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

IIRC, I needed to move this one, because LRUChunkCache doesn't implement the contextmanager interface - maybe implement it instead?

Comment threadzarr/core.py
@@ -87,6 +87,15 @@ class Array:
read and decompressed when possible.

.. versionadded:: 2.7

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this will need to get updated.

Comment threadzarr/creation.py
non-fill-value data are stored, at the expense of overhead associated
with checking the data of each chunk.

.. versionadded:: 2.7

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

needs to get updated

@joshmoore

Copy link
Copy Markdown
Member

Kicked off the GHA workflows.

@codecov

codecovBot commented Oct 24, 2022

Copy link
Copy Markdown

Codecov Report

All modified and coverable lines are covered by tests ✅

Project coverage is 99.99%. Comparing base (f361631) to head (16793f5).
Report is 720 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #1214 +/- ##
========================================
Coverage 99.99% 99.99% ========================================
Files 35 35 Lines 14136 14366 +230 ========================================
+ Hits 14135 14365 +230 
Misses 1 1 
Files with missing linesCoverage Δ
zarr/__init__.py100.00% <ø> (ø)
zarr/core.py100.00% <100.00%> (ø)
zarr/creation.py100.00% <ø> (ø)
zarr/hierarchy.py99.78% <100.00%> (+<0.01%)⬆️
zarr/storage.py100.00% <100.00%> (ø)
zarr/tests/test_core.py100.00% <100.00%> (ø)
zarr/tests/test_storage.py100.00% <100.00%> (ø)

@joshmoore

Copy link
Copy Markdown
Member

Green. Are there any other commits expected from your side, @croth1?

@croth1

Copy link
Copy Markdown
Author

I am still not super sure whether this is the final form or I would rather re-implement it as a caching store ontop of other stores, just like the LRUStoreCache - it feels a bit more intuitive to me. Also I still want to check whether I can do write caching. If I remember correctly, this implementation is write-through. Not sure whether in combination with LRUStoreCache cached writes can be achieved. Need to do more research, but I'm a bit busy right now.

@FarzanT

FarzanT commented May 20, 2023

Copy link
Copy Markdown

Hi, any news on when this feature will get pushed to main? It seems like a very useful feature!
I want to reduce the overhead of sampling from the same chunk which is expected to be in RAM. If I'm not mistaken, right now, every time you pass an index to the zarr array, it has to decompress that index. I want to have an option to keep the entire chunk decompressed in the cache until it's discarded. Specifically, I'm using a large zarr array in my pytorch dataset module, and need to randomly iterate through samples in the same chunk before randomly picking the next chunk. Avoiding the decompression overhead is quite useful, and unfortunately I can't leave the data fully decompressed on disk, as it would be too large.

@croth1

Copy link
Copy Markdown
Author

@FarzanT@joshmoore, I have in the meantime changed my approach and would not need this anymore - hence the long silence. If that's useful for other people, we can try finishing it up as is. Would need some rebasing and performance testing, though.

IIRC, last time I checked all tests were green, although this would need a very thorough review because frequently I was not 100% sure what I was doing during the quite substantial rebase.

@joshmoore

Copy link
Copy Markdown
Member

I'll leave @FarzanT and others to say how pressing their need is. Happy to help how I can @croth1 if you'd like to pursue this.

@FarzanT

Copy link
Copy Markdown

Thank you @croth1 and @joshmoore, my primary use case has also been addressed by #278 (comment), so it's not a pressing issue for me at the moment. But I'd say it would be much nicer to just flip a switch and have zarr handle this internally. If this pull request can address this, then it shouldn't be abandoned IMO!

@jhamman

Copy link
Copy Markdown
Member

I'm going to close this as stale. Folks should feel free to reopen if there is interest in continuing this work.

@jhammanjhamman closed this Oct 11, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@croth1@joshmoore@FarzanT@jhamman@shikharsg@jakirkham
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[WIP] Attempt to continue LRU cache for decoded chunks - #1214

Closed
croth1 wants to merge 58 commits into
zarr-developers:mainfrom
croth1:chunk_cache
Closed

[WIP] Attempt to continue LRU cache for decoded chunks#1214
croth1 wants to merge 58 commits into
zarr-developers:mainfrom
croth1:chunk_cache

Conversation

@croth1

@croth1croth1 commented Oct 23, 2022

Copy link
Copy Markdown

Continuation attempt of #306 - do not merge yet - very likely still not correct behaviour when chunks get deleted from store by #738.

Still have very limited knowledge of the code base - will require me to dig a bit deeper to gain bit more understanding.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

shikharsgand others added 30 commits August 29, 2018 13:17
… decode round trip for object arrays(with tests)
factoring out mapping code from LRUStoreCache and LRUChunkCache
@croth1

croth1 commented Oct 23, 2022

Copy link
Copy Markdown
Author

Hmh, ok - the diff of the merge commit looks still a bit insane - some code from other PRs somehow leaked in/got duplicated - or I am just bad at interpreting the diff. Will investigate tomorrow what happened there.

@@ -65,10 +117,6 @@ def create_store(self, **kwargs): # pragma: no cover
# implement in sub-class

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

IIRC, I needed to move this one, because LRUChunkCache doesn't implement the contextmanager interface - maybe implement it instead?

Comment threadzarr/core.py
@@ -87,6 +87,15 @@ class Array:
read and decompressed when possible.

.. versionadded:: 2.7

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this will need to get updated.

Comment threadzarr/creation.py
non-fill-value data are stored, at the expense of overhead associated
with checking the data of each chunk.

.. versionadded:: 2.7

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

needs to get updated

@joshmoore

Copy link
Copy Markdown
Member

Kicked off the GHA workflows.

@codecov

codecovBot commented Oct 24, 2022

Copy link
Copy Markdown

Codecov Report

All modified and coverable lines are covered by tests ✅

Project coverage is 99.99%. Comparing base (f361631) to head (16793f5).
Report is 720 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #1214 +/- ##
========================================
Coverage 99.99% 99.99% ========================================
Files 35 35 Lines 14136 14366 +230 ========================================
+ Hits 14135 14365 +230 
Misses 1 1 
Files with missing linesCoverage Δ
zarr/__init__.py100.00% <ø> (ø)
zarr/core.py100.00% <100.00%> (ø)
zarr/creation.py100.00% <ø> (ø)
zarr/hierarchy.py99.78% <100.00%> (+<0.01%)⬆️
zarr/storage.py100.00% <100.00%> (ø)
zarr/tests/test_core.py100.00% <100.00%> (ø)
zarr/tests/test_storage.py100.00% <100.00%> (ø)

@joshmoore

Copy link
Copy Markdown
Member

Green. Are there any other commits expected from your side, @croth1?

@croth1

Copy link
Copy Markdown
Author

I am still not super sure whether this is the final form or I would rather re-implement it as a caching store ontop of other stores, just like the LRUStoreCache - it feels a bit more intuitive to me. Also I still want to check whether I can do write caching. If I remember correctly, this implementation is write-through. Not sure whether in combination with LRUStoreCache cached writes can be achieved. Need to do more research, but I'm a bit busy right now.

@FarzanT

FarzanT commented May 20, 2023

Copy link
Copy Markdown

Hi, any news on when this feature will get pushed to main? It seems like a very useful feature!
I want to reduce the overhead of sampling from the same chunk which is expected to be in RAM. If I'm not mistaken, right now, every time you pass an index to the zarr array, it has to decompress that index. I want to have an option to keep the entire chunk decompressed in the cache until it's discarded. Specifically, I'm using a large zarr array in my pytorch dataset module, and need to randomly iterate through samples in the same chunk before randomly picking the next chunk. Avoiding the decompression overhead is quite useful, and unfortunately I can't leave the data fully decompressed on disk, as it would be too large.

@croth1

Copy link
Copy Markdown
Author

@FarzanT@joshmoore, I have in the meantime changed my approach and would not need this anymore - hence the long silence. If that's useful for other people, we can try finishing it up as is. Would need some rebasing and performance testing, though.

IIRC, last time I checked all tests were green, although this would need a very thorough review because frequently I was not 100% sure what I was doing during the quite substantial rebase.

@joshmoore

Copy link
Copy Markdown
Member

I'll leave @FarzanT and others to say how pressing their need is. Happy to help how I can @croth1 if you'd like to pursue this.

@FarzanT

Copy link
Copy Markdown

Thank you @croth1 and @joshmoore, my primary use case has also been addressed by #278 (comment), so it's not a pressing issue for me at the moment. But I'd say it would be much nicer to just flip a switch and have zarr handle this internally. If this pull request can address this, then it shouldn't be abandoned IMO!

@jhamman

Copy link
Copy Markdown
Member

I'm going to close this as stale. Folks should feel free to reopen if there is interest in continuing this work.

@jhammanjhamman closed this Oct 11, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@croth1@joshmoore@FarzanT@jhamman@shikharsg@jakirkham
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

[WIP] Attempt to continue LRU cache for decoded chunks - #1214

Closed
croth1 wants to merge 58 commits into
zarr-developers:mainfrom
croth1:chunk_cache
Closed

[WIP] Attempt to continue LRU cache for decoded chunks#1214
croth1 wants to merge 58 commits into
zarr-developers:mainfrom
croth1:chunk_cache

Conversation

@croth1

@croth1croth1 commented Oct 23, 2022

Copy link
Copy Markdown

Continuation attempt of #306 - do not merge yet - very likely still not correct behaviour when chunks get deleted from store by #738.

Still have very limited knowledge of the code base - will require me to dig a bit deeper to gain bit more understanding.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

shikharsgand others added 30 commits August 29, 2018 13:17
… decode round trip for object arrays(with tests)
factoring out mapping code from LRUStoreCache and LRUChunkCache
@croth1

croth1 commented Oct 23, 2022

Copy link
Copy Markdown
Author

Hmh, ok - the diff of the merge commit looks still a bit insane - some code from other PRs somehow leaked in/got duplicated - or I am just bad at interpreting the diff. Will investigate tomorrow what happened there.

@@ -65,10 +117,6 @@ def create_store(self, **kwargs): # pragma: no cover
# implement in sub-class

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

IIRC, I needed to move this one, because LRUChunkCache doesn't implement the contextmanager interface - maybe implement it instead?

Comment threadzarr/core.py
@@ -87,6 +87,15 @@ class Array:
read and decompressed when possible.

.. versionadded:: 2.7

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this will need to get updated.

Comment threadzarr/creation.py
non-fill-value data are stored, at the expense of overhead associated
with checking the data of each chunk.

.. versionadded:: 2.7

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

needs to get updated

@joshmoore

Copy link
Copy Markdown
Member

Kicked off the GHA workflows.

@codecov

codecovBot commented Oct 24, 2022

Copy link
Copy Markdown

Codecov Report

All modified and coverable lines are covered by tests ✅

Project coverage is 99.99%. Comparing base (f361631) to head (16793f5).
Report is 720 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #1214 +/- ##
========================================
Coverage 99.99% 99.99% ========================================
Files 35 35 Lines 14136 14366 +230 ========================================
+ Hits 14135 14365 +230 
Misses 1 1 
Files with missing linesCoverage Δ
zarr/__init__.py100.00% <ø> (ø)
zarr/core.py100.00% <100.00%> (ø)
zarr/creation.py100.00% <ø> (ø)
zarr/hierarchy.py99.78% <100.00%> (+<0.01%)⬆️
zarr/storage.py100.00% <100.00%> (ø)
zarr/tests/test_core.py100.00% <100.00%> (ø)
zarr/tests/test_storage.py100.00% <100.00%> (ø)

@joshmoore

Copy link
Copy Markdown
Member

Green. Are there any other commits expected from your side, @croth1?

@croth1

Copy link
Copy Markdown
Author

I am still not super sure whether this is the final form or I would rather re-implement it as a caching store ontop of other stores, just like the LRUStoreCache - it feels a bit more intuitive to me. Also I still want to check whether I can do write caching. If I remember correctly, this implementation is write-through. Not sure whether in combination with LRUStoreCache cached writes can be achieved. Need to do more research, but I'm a bit busy right now.

@FarzanT

FarzanT commented May 20, 2023

Copy link
Copy Markdown

Hi, any news on when this feature will get pushed to main? It seems like a very useful feature!
I want to reduce the overhead of sampling from the same chunk which is expected to be in RAM. If I'm not mistaken, right now, every time you pass an index to the zarr array, it has to decompress that index. I want to have an option to keep the entire chunk decompressed in the cache until it's discarded. Specifically, I'm using a large zarr array in my pytorch dataset module, and need to randomly iterate through samples in the same chunk before randomly picking the next chunk. Avoiding the decompression overhead is quite useful, and unfortunately I can't leave the data fully decompressed on disk, as it would be too large.

@croth1

Copy link
Copy Markdown
Author

@FarzanT@joshmoore, I have in the meantime changed my approach and would not need this anymore - hence the long silence. If that's useful for other people, we can try finishing it up as is. Would need some rebasing and performance testing, though.

IIRC, last time I checked all tests were green, although this would need a very thorough review because frequently I was not 100% sure what I was doing during the quite substantial rebase.

@joshmoore

Copy link
Copy Markdown
Member

I'll leave @FarzanT and others to say how pressing their need is. Happy to help how I can @croth1 if you'd like to pursue this.

@FarzanT

Copy link
Copy Markdown

Thank you @croth1 and @joshmoore, my primary use case has also been addressed by #278 (comment), so it's not a pressing issue for me at the moment. But I'd say it would be much nicer to just flip a switch and have zarr handle this internally. If this pull request can address this, then it shouldn't be abandoned IMO!

@jhamman

Copy link
Copy Markdown
Member

I'm going to close this as stale. Folks should feel free to reopen if there is interest in continuing this work.

@jhammanjhamman closed this Oct 11, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@croth1@joshmoore@FarzanT@jhamman@shikharsg@jakirkham