Open zip store from path with zip suffix - #2856

Open
aulemahal wants to merge 14 commits into
zarr-developers:mainfrom
aulemahal:zipstore-from-path
Open

Open zip store from path with zip suffix#2856
aulemahal wants to merge 14 commits into
zarr-developers:mainfrom
aulemahal:zipstore-from-path

Conversation

@aulemahal

@aulemahalaulemahal commented Feb 21, 2025

Copy link
Copy Markdown
  • Opening a string or Path path is redirected to ZipStore if the path has a .zip suffix.

Fixes#2831

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/user-guide/*.rst
  • Changes documented as a new file in changes/
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

The added test is not as elegant as the other from the suite. Is there a way to get an existing zipped zarr dataset within the test suite ? Should I create another fixture instead of writing my own zarr in the test ?

@github-actionsgithub-actionsBot added the needs release notes Automatically applied to PRs which haven't added release notes label Feb 21, 2025
@github-actionsgithub-actionsBot removed the needs release notes Automatically applied to PRs which haven't added release notes label Mar 3, 2025
@aulemahal

Copy link
Copy Markdown
Author

I did not add anything in the user-guide/*.rst files as I didn't feel it was necessary. The docstring of open( already makes mention of this feature.

@dstansbydstansby left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good! Just one minor comment, but I think worth fixing 👍

Comment threadsrc/zarr/storage/_common.py Outdated
@aulemahal

Copy link
Copy Markdown
Author

Hi @dstansby , I'm not sure, but it seems your suggested change that worked in march now breaks the tests. Did some assumptions change about mode=None ?

@dstansbydstansby added this to the 3.1.2 milestone Jul 31, 2025
@codecov

codecovBot commented May 22, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 93.47%. Comparing base (ed60e13) to head (8c234da).

Additional details and impacted files
@@ Coverage Diff @@## main #2856 +/- ##
=======================================
Coverage 93.47% 93.47% =======================================
Files 90 90 Lines 11967 11970 +3 =======================================
+ Hits 11186 11189 +3 
Misses 781 781 
Files with missing linesCoverage Δ
src/zarr/storage/_common.py89.67% <100.00%> (+0.14%)⬆️
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@aulemahal

Copy link
Copy Markdown
Author

Hi @dstansby , I updated the PR with the current main.

To justify coming back to this PR after almost one year:

In between the moment I opened this PR and now, I saw there was the objective of ZEP8 which would have been another way of implementing this convenience. However, this project seems to have been stalled for the last 6 months.

My organisation has a lot of zipped zarrs in its database. We are running a lot of code on a traditionnal HPC which has a relatively low limit on inodes. Zipped zarr offer much better performance than netCDFs and a single node. Local zarrs are more performant, but yield insane amounts of inodes. Xarray not being able to read zipped zarr directly is an issue that blocks us from upgrading to Zarr 3.

We have a few solutions:

  • Having zarr do it natively : the easiest one, this PR
  • A custom xarray backend + changes in intake-esm : more effort, less elegant, needs development in a low-activity repo.
  • Monkeypatch zarr : extra ugly, much less nice than contributing to open source software
  • Flip to netCDF / unzipped zarr : lots of work to rewrite everything, great loss in performance on the HPC

I would understand if my patch here is not the most elegant way forward in the grand scheme of Zarr. But, I still think this could be useful for more than us. It could be seen as a "temporary solution" before ZEP8 is implemented fully.

@d-v-b

Copy link
Copy Markdown
Contributor

Zipped zarr offer much better performance than netCDFs and a single node. Local zarrs are more performant, but yield insane amounts of inodes.

I will look at this PR but in the meantime have you considered sharding? that gives you most of the benefits of zipping your data

@aulemahal

Copy link
Copy Markdown
Author

Thanks!

Yes, we did, but not that much yet, we were a bit slow to begin to look into new Zarr v3 features. For a typical gridded climate datasets with a single data variable, you still have at least 3 inodes per variable (folder, data and json), which can easily yield 20x more inodes considering you have a few auxiliary coords, even if writing the data in a single shard (not sure yet if this is good or not for multi-GB dataseets). So sharding could be a compromise indeed, but this is still a lot of inodes.

@ospinajulian

Copy link
Copy Markdown

Hi, is there any update on this PR?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Can't conveniently open zip store from path with zarr v3

4 participants

@aulemahal@d-v-b@ospinajulian@dstansby
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Open zip store from path with zip suffix - #2856

Open
aulemahal wants to merge 14 commits into
zarr-developers:mainfrom
aulemahal:zipstore-from-path
Open

Open zip store from path with zip suffix#2856
aulemahal wants to merge 14 commits into
zarr-developers:mainfrom
aulemahal:zipstore-from-path

Conversation

@aulemahal

@aulemahalaulemahal commented Feb 21, 2025

Copy link
Copy Markdown
  • Opening a string or Path path is redirected to ZipStore if the path has a .zip suffix.

Fixes#2831

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/user-guide/*.rst
  • Changes documented as a new file in changes/
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

The added test is not as elegant as the other from the suite. Is there a way to get an existing zipped zarr dataset within the test suite ? Should I create another fixture instead of writing my own zarr in the test ?

@github-actionsgithub-actionsBot added the needs release notes Automatically applied to PRs which haven't added release notes label Feb 21, 2025
@github-actionsgithub-actionsBot removed the needs release notes Automatically applied to PRs which haven't added release notes label Mar 3, 2025
@aulemahal

Copy link
Copy Markdown
Author

I did not add anything in the user-guide/*.rst files as I didn't feel it was necessary. The docstring of open( already makes mention of this feature.

@dstansbydstansby left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good! Just one minor comment, but I think worth fixing 👍

Comment threadsrc/zarr/storage/_common.py Outdated
@aulemahal

Copy link
Copy Markdown
Author

Hi @dstansby , I'm not sure, but it seems your suggested change that worked in march now breaks the tests. Did some assumptions change about mode=None ?

@dstansbydstansby added this to the 3.1.2 milestone Jul 31, 2025
@codecov

codecovBot commented May 22, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 93.47%. Comparing base (ed60e13) to head (8c234da).

Additional details and impacted files
@@ Coverage Diff @@## main #2856 +/- ##
=======================================
Coverage 93.47% 93.47% =======================================
Files 90 90 Lines 11967 11970 +3 =======================================
+ Hits 11186 11189 +3 
Misses 781 781 
Files with missing linesCoverage Δ
src/zarr/storage/_common.py89.67% <100.00%> (+0.14%)⬆️
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@aulemahal

Copy link
Copy Markdown
Author

Hi @dstansby , I updated the PR with the current main.

To justify coming back to this PR after almost one year:

In between the moment I opened this PR and now, I saw there was the objective of ZEP8 which would have been another way of implementing this convenience. However, this project seems to have been stalled for the last 6 months.

My organisation has a lot of zipped zarrs in its database. We are running a lot of code on a traditionnal HPC which has a relatively low limit on inodes. Zipped zarr offer much better performance than netCDFs and a single node. Local zarrs are more performant, but yield insane amounts of inodes. Xarray not being able to read zipped zarr directly is an issue that blocks us from upgrading to Zarr 3.

We have a few solutions:

  • Having zarr do it natively : the easiest one, this PR
  • A custom xarray backend + changes in intake-esm : more effort, less elegant, needs development in a low-activity repo.
  • Monkeypatch zarr : extra ugly, much less nice than contributing to open source software
  • Flip to netCDF / unzipped zarr : lots of work to rewrite everything, great loss in performance on the HPC

I would understand if my patch here is not the most elegant way forward in the grand scheme of Zarr. But, I still think this could be useful for more than us. It could be seen as a "temporary solution" before ZEP8 is implemented fully.

@d-v-b

Copy link
Copy Markdown
Contributor

Zipped zarr offer much better performance than netCDFs and a single node. Local zarrs are more performant, but yield insane amounts of inodes.

I will look at this PR but in the meantime have you considered sharding? that gives you most of the benefits of zipping your data

@aulemahal

Copy link
Copy Markdown
Author

Thanks!

Yes, we did, but not that much yet, we were a bit slow to begin to look into new Zarr v3 features. For a typical gridded climate datasets with a single data variable, you still have at least 3 inodes per variable (folder, data and json), which can easily yield 20x more inodes considering you have a few auxiliary coords, even if writing the data in a single shard (not sure yet if this is good or not for multi-GB dataseets). So sharding could be a compromise indeed, but this is still a lot of inodes.

@ospinajulian

Copy link
Copy Markdown

Hi, is there any update on this PR?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Can't conveniently open zip store from path with zarr v3

4 participants

@aulemahal@d-v-b@ospinajulian@dstansby
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Open zip store from path with zip suffix - #2856

Open
aulemahal wants to merge 14 commits into
zarr-developers:mainfrom
aulemahal:zipstore-from-path
Open

Open zip store from path with zip suffix#2856
aulemahal wants to merge 14 commits into
zarr-developers:mainfrom
aulemahal:zipstore-from-path

Conversation

@aulemahal

@aulemahalaulemahal commented Feb 21, 2025

Copy link
Copy Markdown
  • Opening a string or Path path is redirected to ZipStore if the path has a .zip suffix.

Fixes#2831

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/user-guide/*.rst
  • Changes documented as a new file in changes/
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

The added test is not as elegant as the other from the suite. Is there a way to get an existing zipped zarr dataset within the test suite ? Should I create another fixture instead of writing my own zarr in the test ?

@github-actionsgithub-actionsBot added the needs release notes Automatically applied to PRs which haven't added release notes label Feb 21, 2025
@github-actionsgithub-actionsBot removed the needs release notes Automatically applied to PRs which haven't added release notes label Mar 3, 2025
@aulemahal

Copy link
Copy Markdown
Author

I did not add anything in the user-guide/*.rst files as I didn't feel it was necessary. The docstring of open( already makes mention of this feature.

@dstansbydstansby left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good! Just one minor comment, but I think worth fixing 👍

Comment threadsrc/zarr/storage/_common.py Outdated
@aulemahal

Copy link
Copy Markdown
Author

Hi @dstansby , I'm not sure, but it seems your suggested change that worked in march now breaks the tests. Did some assumptions change about mode=None ?

@dstansbydstansby added this to the 3.1.2 milestone Jul 31, 2025
@codecov

codecovBot commented May 22, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 93.47%. Comparing base (ed60e13) to head (8c234da).

Additional details and impacted files
@@ Coverage Diff @@## main #2856 +/- ##
=======================================
Coverage 93.47% 93.47% =======================================
Files 90 90 Lines 11967 11970 +3 =======================================
+ Hits 11186 11189 +3 
Misses 781 781 
Files with missing linesCoverage Δ
src/zarr/storage/_common.py89.67% <100.00%> (+0.14%)⬆️
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@aulemahal

Copy link
Copy Markdown
Author

Hi @dstansby , I updated the PR with the current main.

To justify coming back to this PR after almost one year:

In between the moment I opened this PR and now, I saw there was the objective of ZEP8 which would have been another way of implementing this convenience. However, this project seems to have been stalled for the last 6 months.

My organisation has a lot of zipped zarrs in its database. We are running a lot of code on a traditionnal HPC which has a relatively low limit on inodes. Zipped zarr offer much better performance than netCDFs and a single node. Local zarrs are more performant, but yield insane amounts of inodes. Xarray not being able to read zipped zarr directly is an issue that blocks us from upgrading to Zarr 3.

We have a few solutions:

  • Having zarr do it natively : the easiest one, this PR
  • A custom xarray backend + changes in intake-esm : more effort, less elegant, needs development in a low-activity repo.
  • Monkeypatch zarr : extra ugly, much less nice than contributing to open source software
  • Flip to netCDF / unzipped zarr : lots of work to rewrite everything, great loss in performance on the HPC

I would understand if my patch here is not the most elegant way forward in the grand scheme of Zarr. But, I still think this could be useful for more than us. It could be seen as a "temporary solution" before ZEP8 is implemented fully.

@d-v-b

Copy link
Copy Markdown
Contributor

Zipped zarr offer much better performance than netCDFs and a single node. Local zarrs are more performant, but yield insane amounts of inodes.

I will look at this PR but in the meantime have you considered sharding? that gives you most of the benefits of zipping your data

@aulemahal

Copy link
Copy Markdown
Author

Thanks!

Yes, we did, but not that much yet, we were a bit slow to begin to look into new Zarr v3 features. For a typical gridded climate datasets with a single data variable, you still have at least 3 inodes per variable (folder, data and json), which can easily yield 20x more inodes considering you have a few auxiliary coords, even if writing the data in a single shard (not sure yet if this is good or not for multi-GB dataseets). So sharding could be a compromise indeed, but this is still a lot of inodes.

@ospinajulian

Copy link
Copy Markdown

Hi, is there any update on this PR?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Can't conveniently open zip store from path with zarr v3

4 participants

@aulemahal@d-v-b@ospinajulian@dstansby
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Open zip store from path with zip suffix - #2856

Open
aulemahal wants to merge 14 commits into
zarr-developers:mainfrom
aulemahal:zipstore-from-path
Open

Open zip store from path with zip suffix#2856
aulemahal wants to merge 14 commits into
zarr-developers:mainfrom
aulemahal:zipstore-from-path

Conversation

@aulemahal

@aulemahalaulemahal commented Feb 21, 2025

Copy link
Copy Markdown
  • Opening a string or Path path is redirected to ZipStore if the path has a .zip suffix.

Fixes#2831

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/user-guide/*.rst
  • Changes documented as a new file in changes/
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

The added test is not as elegant as the other from the suite. Is there a way to get an existing zipped zarr dataset within the test suite ? Should I create another fixture instead of writing my own zarr in the test ?

@github-actionsgithub-actionsBot added the needs release notes Automatically applied to PRs which haven't added release notes label Feb 21, 2025
@github-actionsgithub-actionsBot removed the needs release notes Automatically applied to PRs which haven't added release notes label Mar 3, 2025
@aulemahal

Copy link
Copy Markdown
Author

I did not add anything in the user-guide/*.rst files as I didn't feel it was necessary. The docstring of open( already makes mention of this feature.

@dstansbydstansby left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good! Just one minor comment, but I think worth fixing 👍

Comment threadsrc/zarr/storage/_common.py Outdated
@aulemahal

Copy link
Copy Markdown
Author

Hi @dstansby , I'm not sure, but it seems your suggested change that worked in march now breaks the tests. Did some assumptions change about mode=None ?

@dstansbydstansby added this to the 3.1.2 milestone Jul 31, 2025
@codecov

codecovBot commented May 22, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 93.47%. Comparing base (ed60e13) to head (8c234da).

Additional details and impacted files
@@ Coverage Diff @@## main #2856 +/- ##
=======================================
Coverage 93.47% 93.47% =======================================
Files 90 90 Lines 11967 11970 +3 =======================================
+ Hits 11186 11189 +3 
Misses 781 781 
Files with missing linesCoverage Δ
src/zarr/storage/_common.py89.67% <100.00%> (+0.14%)⬆️
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@aulemahal

Copy link
Copy Markdown
Author

Hi @dstansby , I updated the PR with the current main.

To justify coming back to this PR after almost one year:

In between the moment I opened this PR and now, I saw there was the objective of ZEP8 which would have been another way of implementing this convenience. However, this project seems to have been stalled for the last 6 months.

My organisation has a lot of zipped zarrs in its database. We are running a lot of code on a traditionnal HPC which has a relatively low limit on inodes. Zipped zarr offer much better performance than netCDFs and a single node. Local zarrs are more performant, but yield insane amounts of inodes. Xarray not being able to read zipped zarr directly is an issue that blocks us from upgrading to Zarr 3.

We have a few solutions:

  • Having zarr do it natively : the easiest one, this PR
  • A custom xarray backend + changes in intake-esm : more effort, less elegant, needs development in a low-activity repo.
  • Monkeypatch zarr : extra ugly, much less nice than contributing to open source software
  • Flip to netCDF / unzipped zarr : lots of work to rewrite everything, great loss in performance on the HPC

I would understand if my patch here is not the most elegant way forward in the grand scheme of Zarr. But, I still think this could be useful for more than us. It could be seen as a "temporary solution" before ZEP8 is implemented fully.

@d-v-b

Copy link
Copy Markdown
Contributor

Zipped zarr offer much better performance than netCDFs and a single node. Local zarrs are more performant, but yield insane amounts of inodes.

I will look at this PR but in the meantime have you considered sharding? that gives you most of the benefits of zipping your data

@aulemahal

Copy link
Copy Markdown
Author

Thanks!

Yes, we did, but not that much yet, we were a bit slow to begin to look into new Zarr v3 features. For a typical gridded climate datasets with a single data variable, you still have at least 3 inodes per variable (folder, data and json), which can easily yield 20x more inodes considering you have a few auxiliary coords, even if writing the data in a single shard (not sure yet if this is good or not for multi-GB dataseets). So sharding could be a compromise indeed, but this is still a lot of inodes.

@ospinajulian

Copy link
Copy Markdown

Hi, is there any update on this PR?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Can't conveniently open zip store from path with zarr v3

4 participants

@aulemahal@d-v-b@ospinajulian@dstansby
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Open zip store from path with zip suffix - #2856

Open
aulemahal wants to merge 14 commits into
zarr-developers:mainfrom
aulemahal:zipstore-from-path
Open

Open zip store from path with zip suffix#2856
aulemahal wants to merge 14 commits into
zarr-developers:mainfrom
aulemahal:zipstore-from-path

Conversation

@aulemahal

@aulemahalaulemahal commented Feb 21, 2025

Copy link
Copy Markdown
  • Opening a string or Path path is redirected to ZipStore if the path has a .zip suffix.

Fixes#2831

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/user-guide/*.rst
  • Changes documented as a new file in changes/
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

The added test is not as elegant as the other from the suite. Is there a way to get an existing zipped zarr dataset within the test suite ? Should I create another fixture instead of writing my own zarr in the test ?

@github-actionsgithub-actionsBot added the needs release notes Automatically applied to PRs which haven't added release notes label Feb 21, 2025
@github-actionsgithub-actionsBot removed the needs release notes Automatically applied to PRs which haven't added release notes label Mar 3, 2025
@aulemahal

Copy link
Copy Markdown
Author

I did not add anything in the user-guide/*.rst files as I didn't feel it was necessary. The docstring of open( already makes mention of this feature.

@dstansbydstansby left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good! Just one minor comment, but I think worth fixing 👍

Comment threadsrc/zarr/storage/_common.py Outdated
@aulemahal

Copy link
Copy Markdown
Author

Hi @dstansby , I'm not sure, but it seems your suggested change that worked in march now breaks the tests. Did some assumptions change about mode=None ?

@dstansbydstansby added this to the 3.1.2 milestone Jul 31, 2025
@codecov

codecovBot commented May 22, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 93.47%. Comparing base (ed60e13) to head (8c234da).

Additional details and impacted files
@@ Coverage Diff @@## main #2856 +/- ##
=======================================
Coverage 93.47% 93.47% =======================================
Files 90 90 Lines 11967 11970 +3 =======================================
+ Hits 11186 11189 +3 
Misses 781 781 
Files with missing linesCoverage Δ
src/zarr/storage/_common.py89.67% <100.00%> (+0.14%)⬆️
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@aulemahal

Copy link
Copy Markdown
Author

Hi @dstansby , I updated the PR with the current main.

To justify coming back to this PR after almost one year:

In between the moment I opened this PR and now, I saw there was the objective of ZEP8 which would have been another way of implementing this convenience. However, this project seems to have been stalled for the last 6 months.

My organisation has a lot of zipped zarrs in its database. We are running a lot of code on a traditionnal HPC which has a relatively low limit on inodes. Zipped zarr offer much better performance than netCDFs and a single node. Local zarrs are more performant, but yield insane amounts of inodes. Xarray not being able to read zipped zarr directly is an issue that blocks us from upgrading to Zarr 3.

We have a few solutions:

  • Having zarr do it natively : the easiest one, this PR
  • A custom xarray backend + changes in intake-esm : more effort, less elegant, needs development in a low-activity repo.
  • Monkeypatch zarr : extra ugly, much less nice than contributing to open source software
  • Flip to netCDF / unzipped zarr : lots of work to rewrite everything, great loss in performance on the HPC

I would understand if my patch here is not the most elegant way forward in the grand scheme of Zarr. But, I still think this could be useful for more than us. It could be seen as a "temporary solution" before ZEP8 is implemented fully.

@d-v-b

Copy link
Copy Markdown
Contributor

Zipped zarr offer much better performance than netCDFs and a single node. Local zarrs are more performant, but yield insane amounts of inodes.

I will look at this PR but in the meantime have you considered sharding? that gives you most of the benefits of zipping your data

@aulemahal

Copy link
Copy Markdown
Author

Thanks!

Yes, we did, but not that much yet, we were a bit slow to begin to look into new Zarr v3 features. For a typical gridded climate datasets with a single data variable, you still have at least 3 inodes per variable (folder, data and json), which can easily yield 20x more inodes considering you have a few auxiliary coords, even if writing the data in a single shard (not sure yet if this is good or not for multi-GB dataseets). So sharding could be a compromise indeed, but this is still a lot of inodes.

@ospinajulian

Copy link
Copy Markdown

Hi, is there any update on this PR?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Can't conveniently open zip store from path with zarr v3

4 participants

@aulemahal@d-v-b@ospinajulian@dstansby
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Open zip store from path with zip suffix - #2856

Open
aulemahal wants to merge 14 commits into
zarr-developers:mainfrom
aulemahal:zipstore-from-path
Open

Open zip store from path with zip suffix#2856
aulemahal wants to merge 14 commits into
zarr-developers:mainfrom
aulemahal:zipstore-from-path

Conversation

@aulemahal

@aulemahalaulemahal commented Feb 21, 2025

Copy link
Copy Markdown
  • Opening a string or Path path is redirected to ZipStore if the path has a .zip suffix.

Fixes#2831

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/user-guide/*.rst
  • Changes documented as a new file in changes/
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

The added test is not as elegant as the other from the suite. Is there a way to get an existing zipped zarr dataset within the test suite ? Should I create another fixture instead of writing my own zarr in the test ?

@github-actionsgithub-actionsBot added the needs release notes Automatically applied to PRs which haven't added release notes label Feb 21, 2025
@github-actionsgithub-actionsBot removed the needs release notes Automatically applied to PRs which haven't added release notes label Mar 3, 2025
@aulemahal

Copy link
Copy Markdown
Author

I did not add anything in the user-guide/*.rst files as I didn't feel it was necessary. The docstring of open( already makes mention of this feature.

@dstansbydstansby left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good! Just one minor comment, but I think worth fixing 👍

Comment threadsrc/zarr/storage/_common.py Outdated
@aulemahal

Copy link
Copy Markdown
Author

Hi @dstansby , I'm not sure, but it seems your suggested change that worked in march now breaks the tests. Did some assumptions change about mode=None ?

@dstansbydstansby added this to the 3.1.2 milestone Jul 31, 2025
@codecov

codecovBot commented May 22, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 93.47%. Comparing base (ed60e13) to head (8c234da).

Additional details and impacted files
@@ Coverage Diff @@## main #2856 +/- ##
=======================================
Coverage 93.47% 93.47% =======================================
Files 90 90 Lines 11967 11970 +3 =======================================
+ Hits 11186 11189 +3 
Misses 781 781 
Files with missing linesCoverage Δ
src/zarr/storage/_common.py89.67% <100.00%> (+0.14%)⬆️
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@aulemahal

Copy link
Copy Markdown
Author

Hi @dstansby , I updated the PR with the current main.

To justify coming back to this PR after almost one year:

In between the moment I opened this PR and now, I saw there was the objective of ZEP8 which would have been another way of implementing this convenience. However, this project seems to have been stalled for the last 6 months.

My organisation has a lot of zipped zarrs in its database. We are running a lot of code on a traditionnal HPC which has a relatively low limit on inodes. Zipped zarr offer much better performance than netCDFs and a single node. Local zarrs are more performant, but yield insane amounts of inodes. Xarray not being able to read zipped zarr directly is an issue that blocks us from upgrading to Zarr 3.

We have a few solutions:

  • Having zarr do it natively : the easiest one, this PR
  • A custom xarray backend + changes in intake-esm : more effort, less elegant, needs development in a low-activity repo.
  • Monkeypatch zarr : extra ugly, much less nice than contributing to open source software
  • Flip to netCDF / unzipped zarr : lots of work to rewrite everything, great loss in performance on the HPC

I would understand if my patch here is not the most elegant way forward in the grand scheme of Zarr. But, I still think this could be useful for more than us. It could be seen as a "temporary solution" before ZEP8 is implemented fully.

@d-v-b

Copy link
Copy Markdown
Contributor

Zipped zarr offer much better performance than netCDFs and a single node. Local zarrs are more performant, but yield insane amounts of inodes.

I will look at this PR but in the meantime have you considered sharding? that gives you most of the benefits of zipping your data

@aulemahal

Copy link
Copy Markdown
Author

Thanks!

Yes, we did, but not that much yet, we were a bit slow to begin to look into new Zarr v3 features. For a typical gridded climate datasets with a single data variable, you still have at least 3 inodes per variable (folder, data and json), which can easily yield 20x more inodes considering you have a few auxiliary coords, even if writing the data in a single shard (not sure yet if this is good or not for multi-GB dataseets). So sharding could be a compromise indeed, but this is still a lot of inodes.

@ospinajulian

Copy link
Copy Markdown

Hi, is there any update on this PR?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Can't conveniently open zip store from path with zarr v3

4 participants

@aulemahal@d-v-b@ospinajulian@dstansby
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Open zip store from path with zip suffix - #2856

Open
aulemahal wants to merge 14 commits into
zarr-developers:mainfrom
aulemahal:zipstore-from-path
Open

Open zip store from path with zip suffix#2856
aulemahal wants to merge 14 commits into
zarr-developers:mainfrom
aulemahal:zipstore-from-path

Conversation

@aulemahal

@aulemahalaulemahal commented Feb 21, 2025

Copy link
Copy Markdown
  • Opening a string or Path path is redirected to ZipStore if the path has a .zip suffix.

Fixes#2831

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/user-guide/*.rst
  • Changes documented as a new file in changes/
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

The added test is not as elegant as the other from the suite. Is there a way to get an existing zipped zarr dataset within the test suite ? Should I create another fixture instead of writing my own zarr in the test ?

@github-actionsgithub-actionsBot added the needs release notes Automatically applied to PRs which haven't added release notes label Feb 21, 2025
@github-actionsgithub-actionsBot removed the needs release notes Automatically applied to PRs which haven't added release notes label Mar 3, 2025
@aulemahal

Copy link
Copy Markdown
Author

I did not add anything in the user-guide/*.rst files as I didn't feel it was necessary. The docstring of open( already makes mention of this feature.

@dstansbydstansby left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good! Just one minor comment, but I think worth fixing 👍

Comment threadsrc/zarr/storage/_common.py Outdated
@aulemahal

Copy link
Copy Markdown
Author

Hi @dstansby , I'm not sure, but it seems your suggested change that worked in march now breaks the tests. Did some assumptions change about mode=None ?

@dstansbydstansby added this to the 3.1.2 milestone Jul 31, 2025
@codecov

codecovBot commented May 22, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 93.47%. Comparing base (ed60e13) to head (8c234da).

Additional details and impacted files
@@ Coverage Diff @@## main #2856 +/- ##
=======================================
Coverage 93.47% 93.47% =======================================
Files 90 90 Lines 11967 11970 +3 =======================================
+ Hits 11186 11189 +3 
Misses 781 781 
Files with missing linesCoverage Δ
src/zarr/storage/_common.py89.67% <100.00%> (+0.14%)⬆️
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@aulemahal

Copy link
Copy Markdown
Author

Hi @dstansby , I updated the PR with the current main.

To justify coming back to this PR after almost one year:

In between the moment I opened this PR and now, I saw there was the objective of ZEP8 which would have been another way of implementing this convenience. However, this project seems to have been stalled for the last 6 months.

My organisation has a lot of zipped zarrs in its database. We are running a lot of code on a traditionnal HPC which has a relatively low limit on inodes. Zipped zarr offer much better performance than netCDFs and a single node. Local zarrs are more performant, but yield insane amounts of inodes. Xarray not being able to read zipped zarr directly is an issue that blocks us from upgrading to Zarr 3.

We have a few solutions:

  • Having zarr do it natively : the easiest one, this PR
  • A custom xarray backend + changes in intake-esm : more effort, less elegant, needs development in a low-activity repo.
  • Monkeypatch zarr : extra ugly, much less nice than contributing to open source software
  • Flip to netCDF / unzipped zarr : lots of work to rewrite everything, great loss in performance on the HPC

I would understand if my patch here is not the most elegant way forward in the grand scheme of Zarr. But, I still think this could be useful for more than us. It could be seen as a "temporary solution" before ZEP8 is implemented fully.

@d-v-b

Copy link
Copy Markdown
Contributor

Zipped zarr offer much better performance than netCDFs and a single node. Local zarrs are more performant, but yield insane amounts of inodes.

I will look at this PR but in the meantime have you considered sharding? that gives you most of the benefits of zipping your data

@aulemahal

Copy link
Copy Markdown
Author

Thanks!

Yes, we did, but not that much yet, we were a bit slow to begin to look into new Zarr v3 features. For a typical gridded climate datasets with a single data variable, you still have at least 3 inodes per variable (folder, data and json), which can easily yield 20x more inodes considering you have a few auxiliary coords, even if writing the data in a single shard (not sure yet if this is good or not for multi-GB dataseets). So sharding could be a compromise indeed, but this is still a lot of inodes.

@ospinajulian

Copy link
Copy Markdown

Hi, is there any update on this PR?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Can't conveniently open zip store from path with zarr v3

4 participants

@aulemahal@d-v-b@ospinajulian@dstansby
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Open zip store from path with zip suffix - #2856

Open
aulemahal wants to merge 14 commits into
zarr-developers:mainfrom
aulemahal:zipstore-from-path
Open

Open zip store from path with zip suffix#2856
aulemahal wants to merge 14 commits into
zarr-developers:mainfrom
aulemahal:zipstore-from-path

Conversation

@aulemahal

@aulemahalaulemahal commented Feb 21, 2025

Copy link
Copy Markdown
  • Opening a string or Path path is redirected to ZipStore if the path has a .zip suffix.

Fixes#2831

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/user-guide/*.rst
  • Changes documented as a new file in changes/
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

The added test is not as elegant as the other from the suite. Is there a way to get an existing zipped zarr dataset within the test suite ? Should I create another fixture instead of writing my own zarr in the test ?

@github-actionsgithub-actionsBot added the needs release notes Automatically applied to PRs which haven't added release notes label Feb 21, 2025
@github-actionsgithub-actionsBot removed the needs release notes Automatically applied to PRs which haven't added release notes label Mar 3, 2025
@aulemahal

Copy link
Copy Markdown
Author

I did not add anything in the user-guide/*.rst files as I didn't feel it was necessary. The docstring of open( already makes mention of this feature.

@dstansbydstansby left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good! Just one minor comment, but I think worth fixing 👍

Comment threadsrc/zarr/storage/_common.py Outdated
@aulemahal

Copy link
Copy Markdown
Author

Hi @dstansby , I'm not sure, but it seems your suggested change that worked in march now breaks the tests. Did some assumptions change about mode=None ?

@dstansbydstansby added this to the 3.1.2 milestone Jul 31, 2025
@codecov

codecovBot commented May 22, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 93.47%. Comparing base (ed60e13) to head (8c234da).

Additional details and impacted files
@@ Coverage Diff @@## main #2856 +/- ##
=======================================
Coverage 93.47% 93.47% =======================================
Files 90 90 Lines 11967 11970 +3 =======================================
+ Hits 11186 11189 +3 
Misses 781 781 
Files with missing linesCoverage Δ
src/zarr/storage/_common.py89.67% <100.00%> (+0.14%)⬆️
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@aulemahal

Copy link
Copy Markdown
Author

Hi @dstansby , I updated the PR with the current main.

To justify coming back to this PR after almost one year:

In between the moment I opened this PR and now, I saw there was the objective of ZEP8 which would have been another way of implementing this convenience. However, this project seems to have been stalled for the last 6 months.

My organisation has a lot of zipped zarrs in its database. We are running a lot of code on a traditionnal HPC which has a relatively low limit on inodes. Zipped zarr offer much better performance than netCDFs and a single node. Local zarrs are more performant, but yield insane amounts of inodes. Xarray not being able to read zipped zarr directly is an issue that blocks us from upgrading to Zarr 3.

We have a few solutions:

  • Having zarr do it natively : the easiest one, this PR
  • A custom xarray backend + changes in intake-esm : more effort, less elegant, needs development in a low-activity repo.
  • Monkeypatch zarr : extra ugly, much less nice than contributing to open source software
  • Flip to netCDF / unzipped zarr : lots of work to rewrite everything, great loss in performance on the HPC

I would understand if my patch here is not the most elegant way forward in the grand scheme of Zarr. But, I still think this could be useful for more than us. It could be seen as a "temporary solution" before ZEP8 is implemented fully.

@d-v-b

Copy link
Copy Markdown
Contributor

Zipped zarr offer much better performance than netCDFs and a single node. Local zarrs are more performant, but yield insane amounts of inodes.

I will look at this PR but in the meantime have you considered sharding? that gives you most of the benefits of zipping your data

@aulemahal

Copy link
Copy Markdown
Author

Thanks!

Yes, we did, but not that much yet, we were a bit slow to begin to look into new Zarr v3 features. For a typical gridded climate datasets with a single data variable, you still have at least 3 inodes per variable (folder, data and json), which can easily yield 20x more inodes considering you have a few auxiliary coords, even if writing the data in a single shard (not sure yet if this is good or not for multi-GB dataseets). So sharding could be a compromise indeed, but this is still a lot of inodes.

@ospinajulian

Copy link
Copy Markdown

Hi, is there any update on this PR?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Can't conveniently open zip store from path with zarr v3

4 participants

@aulemahal@d-v-b@ospinajulian@dstansby