Zarr consolidated - #2559

Merged
shoyer merged 16 commits into
pydata:masterfrom
rabernat:zarr_consolidated
Dec 4, 2018
Merged

Zarr consolidated#2559
shoyer merged 16 commits into
pydata:masterfrom
rabernat:zarr_consolidated

Conversation

@rabernat

@rabernatrabernat commented Nov 20, 2018

Copy link
Copy Markdown
Contributor

This PR adds support for reading and writing of consolidated metadata in zarr stores.

  • Closeshow to incorporate zarr's new open_consolidated method? #2558 (remove if there is no corresponding issue, which should only be the case for minor changes)
  • Tests added (for all bug fixes or enhancements)
  • Fully documented, including whats-new.rst for all changes and api.rst for new API (remove if this change should not be visible to users, e.g., if it is an internal clean-up, or if this is part of a larger project that will be documented later)

@pep8speaks

pep8speaks commented Nov 20, 2018

Copy link
Copy Markdown

Hello @rabernat! Thanks for updating the PR.

Line 240:80: E501 line too long (82 > 79 characters)

Comment last updated on December 04, 2018 at 19:34 Hours UTC

Comment threadxarray/backends/api.py Outdated
@rabernat

Copy link
Copy Markdown
ContributorAuthor

Ping @lilyminium for a review.

Comment threadxarray/backends/api.py Outdated
if consolidate:
import zarr
zarr.consolidate_metadata(store)
# do we need to reload ztore now that we have consolidated?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

would it make sense for zarr to handle this?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What do you mean?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I meant reloading the zarr store automatically.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think that would be hard to achieve. And I'm not sure it's necessary. Frankly I don't know why we return a store object from to_zarr at all.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

zarr.consolidate_metadata returns the output of open_consolidated on the same store, so this is already happening

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Also need to add some version checks...this will only work with zarr > 2.2.

Comment threaddoc/io.rst
Comment threadxarray/tests/test_backends.py Outdated
Comment threadxarray/backends/zarr.py Outdated

def __init__(self, zarr_group):
if consolidated or consolidate_on_close:
if LooseVersion(zarr.__version__) <= '2.2': # pragma: no cover

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

reminder to update this version check too.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Being more explicit about the version seems to fix this issue here. In the tests I have used the importorskip approach.

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Not sure I understand why there are tests failing now. The failing function is test_basic_compute.

https://travis-ci.org/pydata/xarray/jobs/460873430#L7489

At first glance, this does not appear to have anything to do with my PR. The relevant error is:


______________________________ test_basic_compute ______________________________
def test_basic_compute():
ds = Dataset({'foo': ('x', range(5)),
'bar': ('x', range(5))}).chunk({'x': 2})
for get in [dask.threaded.get,
dask.multiprocessing.get,
dask.local.get_sync,
None]:
with (dask.config.set(scheduler=get)
if LooseVersion(dask.__version__) >= LooseVersion('0.19.4')
else dask.config.set(scheduler=get)
if LooseVersion(dask.__version__) >= LooseVersion('0.18.0')
else dask.set_options(get=get)):
> ds.compute()
xarray/tests/test_dask.py:843: _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ xarray/core/dataset.py:597: in compute
return new.load(**kwargs)
xarray/core/dataset.py:494: in load
evaluated_data = da.compute(*lazy_data.values(), **kwargs)
../../../miniconda/envs/test_env/lib/python3.6/site-packages/dask/base.py:390: in compute
collections=collections)
../../../miniconda/envs/test_env/lib/python3.6/site-packages/dask/base.py:865: in get_scheduler
return get_scheduler(scheduler=config.get('scheduler', None))
_ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ get = None, scheduler = <function get at 0x7fc31d9ae048>, collections = None
cls = None
def get_scheduler(get=None, scheduler=None, collections=None, cls=None):
""" Get scheduler function
There are various ways to specify the scheduler to use:
1. Passing in get= parameters (deprecated)
2. Passing in scheduler= parameters
3. Passing these into global confiuration
4. Using defaults of a dask collection
This function centralizes the logic to determine the right scheduler to use
from those many options
"""
if get is not None:
if scheduler is not None:
raise ValueError("Both get= and scheduler= provided. Choose one")
warn_on_get(get)
return get
if scheduler is not None:
> if scheduler.lower() in named_schedulers:
E AttributeError: 'function' object has no attribute 'lower'
../../../miniconda/envs/test_env/lib/python3.6/site-packages/dask/base.py:854: AttributeError

@shoyer

Copy link
Copy Markdown
Member

I bet this is due to the latest dask release (1.0). We can fix this in another PR.

@lilyminium

Copy link
Copy Markdown
Contributor

I remember dealing with this in my pull request -- if I recall correctly scheduler was pointing to the scheduler.get function instead. It was a minor bug that was either fixed in the next release of xarray (0.11.0) or Dask (0.20.1).

@rabernat

Copy link
Copy Markdown
ContributorAuthor

So if the test issues can be considered resolved, the only decision we need to make is about the API.

Do we prefer (the current way):

ds.to_zarr(fname, consolidate=True)
xr.open_zarr(fname, consolidated=True)

or @shoyer's suggestion

ds.to_zarr(fname, consolidated=True)
xr.open_zarr(fname, consolidated=True)

???

@martindurant

Copy link
Copy Markdown
Contributor

Will the default for both options be False for the time being?

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Will the default for both options be False for the time being?

Yes

@martindurant

Copy link
Copy Markdown
Contributor

Glad to see this happening, by the way. Once in, catalogs using intake-xarray can be updated and I don't thin the code will need to change.

@alimanfoo

Copy link
Copy Markdown
Contributor

Great to see this. On the API, FWIW I'd vote for using the same keyword (consolidated) in both, less burden on the user to remember what to use.

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Keywords are now all consolidated and all tests are go.

Ready to merge?

@jhammanjhamman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is basically ready. I had a few small questions/comments but this looks safe for a merge here soon.

Comment threaddoc/whats-new.rst
~~~~~~~~~~~~

- Ability to read and write consolidated metadata in zarr stores.
By `Ryan Abernathey <https://github.com/rabernat>`_.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you reference the issue this is attached to: (:issue:`2558`).


open_kwargs = dict(mode=mode, synchronizer=synchronizer, path=group)
if consolidated:
# TODO: an option to pass the metadata_key keyword

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

do we need to consider this TODO here?


open_kwargs = dict(mode=mode, synchronizer=synchronizer, path=group)
if consolidated:
# TODO: an option to pass the metadata_key keyword

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Anything to do here now?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we feel that it's important to expose this functionality from within xarray? I don't.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I also don't.
I think it's ok for xarray to have an opinion on what the special key is called.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I propose we just leave these TODO's here as is. If anyone ever needs this feature from the xarray side, this will help guide them on how to implement it.

@martindurant

Copy link
Copy Markdown
Contributor

LGTM

Do you think there should be more explicit text of how to add consolidation to existing zarr/xarray data-sets, rather than creating them with consolidation turned on?

We may also need some text around updating consolidated data-sets, but that can maybe wait to see what kind of usage people try.

@rabernat

Copy link
Copy Markdown
ContributorAuthor

We may also need some text around updating consolidated data-sets, but that can maybe wait to see what kind of usage people try.

Since xarray cannot append or modify in-place existing zarr stores, this seems outside the scope of xarray for now. But maybe it is worth mentioning in the docs.

@jhamman

Copy link
Copy Markdown
Member

I'm happy here. ...but Appveyor is not.

@shoyer

shoyer commented Dec 4, 2018

Copy link
Copy Markdown
Member

@rabernat if you're ready, let's merge this.

The failures on Appveyor are unrelated (an issue with int32 and cftime)

@rabernat

rabernat commented Dec 4, 2018 via email

Copy link
Copy Markdown
ContributorAuthor

@shoyer
shoyer merged commit 3ae93ac into pydata:masterDec 4, 2018
@rabernat

Copy link
Copy Markdown
ContributorAuthor

If anyone wants to see how awesome consolidated metadata is, you can try it in this binder:
https://github.com/rabernat/pangeo_ecco_examples/

I did a bit of lazy profiling here:
https://gist.github.com/rabernat/ce1fb414cf53541afe2245363b06c49d

Things that used to take ~40s now take ~1s. Especially since loading the data is one of the first steps in any pangeo notebook, this is a huge improvement in usability.

Thanks to everyone who helped make it happen!

@martindurant

Copy link
Copy Markdown
Contributor

I like those timings.

dcherian pushed a commit to yohai/xarray that referenced this pull request Dec 16, 2018
* upstream/master:
Feature: N-dimensional auto_combine (pydata#2553)
Support HighLevelGraphs (pydata#2603)
Bump cftime version in doc environment (pydata#2604)
use keep_attrs in binary operations II (pydata#2590)
Temporarily mark dask-dev build as an allowed failure (pydata#2602)
Fix wrong error message in interp() (pydata#2598)
Add dayofyear and dayofweek accessors (pydata#2599)
Fix h5netcdf saving scalars with filters or chunks (pydata#2591)
Minor update to PR template (pydata#2596)
Zarr consolidated (pydata#2559)
fix examples (pydata#2581)
Fix typo (pydata#2578)
Concat docstring typo (pydata#2577)
DOC: remove example using Dataset.T (pydata#2572)
python setup.py test now works by default (pydata#2573)
Return slices when possible from CFTimeIndex.get_loc() (pydata#2569)
DOC: fix computation.rst (pydata#2567)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

how to incorporate zarr's new open_consolidated method?

7 participants

@rabernat@pep8speaks@shoyer@lilyminium@martindurant@alimanfoo@jhamman
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Zarr consolidated - #2559

Merged
shoyer merged 16 commits into
pydata:masterfrom
rabernat:zarr_consolidated
Dec 4, 2018
Merged

Zarr consolidated#2559
shoyer merged 16 commits into
pydata:masterfrom
rabernat:zarr_consolidated

Conversation

@rabernat

@rabernatrabernat commented Nov 20, 2018

Copy link
Copy Markdown
Contributor

This PR adds support for reading and writing of consolidated metadata in zarr stores.

  • Closeshow to incorporate zarr's new open_consolidated method? #2558 (remove if there is no corresponding issue, which should only be the case for minor changes)
  • Tests added (for all bug fixes or enhancements)
  • Fully documented, including whats-new.rst for all changes and api.rst for new API (remove if this change should not be visible to users, e.g., if it is an internal clean-up, or if this is part of a larger project that will be documented later)

@pep8speaks

pep8speaks commented Nov 20, 2018

Copy link
Copy Markdown

Hello @rabernat! Thanks for updating the PR.

Line 240:80: E501 line too long (82 > 79 characters)

Comment last updated on December 04, 2018 at 19:34 Hours UTC

Comment threadxarray/backends/api.py Outdated
@rabernat

Copy link
Copy Markdown
ContributorAuthor

Ping @lilyminium for a review.

Comment threadxarray/backends/api.py Outdated
if consolidate:
import zarr
zarr.consolidate_metadata(store)
# do we need to reload ztore now that we have consolidated?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

would it make sense for zarr to handle this?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What do you mean?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I meant reloading the zarr store automatically.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think that would be hard to achieve. And I'm not sure it's necessary. Frankly I don't know why we return a store object from to_zarr at all.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

zarr.consolidate_metadata returns the output of open_consolidated on the same store, so this is already happening

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Also need to add some version checks...this will only work with zarr > 2.2.

Comment threaddoc/io.rst
Comment threadxarray/tests/test_backends.py Outdated
Comment threadxarray/backends/zarr.py Outdated

def __init__(self, zarr_group):
if consolidated or consolidate_on_close:
if LooseVersion(zarr.__version__) <= '2.2': # pragma: no cover

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

reminder to update this version check too.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Being more explicit about the version seems to fix this issue here. In the tests I have used the importorskip approach.

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Not sure I understand why there are tests failing now. The failing function is test_basic_compute.

https://travis-ci.org/pydata/xarray/jobs/460873430#L7489

At first glance, this does not appear to have anything to do with my PR. The relevant error is:


______________________________ test_basic_compute ______________________________
def test_basic_compute():
ds = Dataset({'foo': ('x', range(5)),
'bar': ('x', range(5))}).chunk({'x': 2})
for get in [dask.threaded.get,
dask.multiprocessing.get,
dask.local.get_sync,
None]:
with (dask.config.set(scheduler=get)
if LooseVersion(dask.__version__) >= LooseVersion('0.19.4')
else dask.config.set(scheduler=get)
if LooseVersion(dask.__version__) >= LooseVersion('0.18.0')
else dask.set_options(get=get)):
> ds.compute()
xarray/tests/test_dask.py:843: _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ xarray/core/dataset.py:597: in compute
return new.load(**kwargs)
xarray/core/dataset.py:494: in load
evaluated_data = da.compute(*lazy_data.values(), **kwargs)
../../../miniconda/envs/test_env/lib/python3.6/site-packages/dask/base.py:390: in compute
collections=collections)
../../../miniconda/envs/test_env/lib/python3.6/site-packages/dask/base.py:865: in get_scheduler
return get_scheduler(scheduler=config.get('scheduler', None))
_ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ get = None, scheduler = <function get at 0x7fc31d9ae048>, collections = None
cls = None
def get_scheduler(get=None, scheduler=None, collections=None, cls=None):
""" Get scheduler function
There are various ways to specify the scheduler to use:
1. Passing in get= parameters (deprecated)
2. Passing in scheduler= parameters
3. Passing these into global confiuration
4. Using defaults of a dask collection
This function centralizes the logic to determine the right scheduler to use
from those many options
"""
if get is not None:
if scheduler is not None:
raise ValueError("Both get= and scheduler= provided. Choose one")
warn_on_get(get)
return get
if scheduler is not None:
> if scheduler.lower() in named_schedulers:
E AttributeError: 'function' object has no attribute 'lower'
../../../miniconda/envs/test_env/lib/python3.6/site-packages/dask/base.py:854: AttributeError

@shoyer

Copy link
Copy Markdown
Member

I bet this is due to the latest dask release (1.0). We can fix this in another PR.

@lilyminium

Copy link
Copy Markdown
Contributor

I remember dealing with this in my pull request -- if I recall correctly scheduler was pointing to the scheduler.get function instead. It was a minor bug that was either fixed in the next release of xarray (0.11.0) or Dask (0.20.1).

@rabernat

Copy link
Copy Markdown
ContributorAuthor

So if the test issues can be considered resolved, the only decision we need to make is about the API.

Do we prefer (the current way):

ds.to_zarr(fname, consolidate=True)
xr.open_zarr(fname, consolidated=True)

or @shoyer's suggestion

ds.to_zarr(fname, consolidated=True)
xr.open_zarr(fname, consolidated=True)

???

@martindurant

Copy link
Copy Markdown
Contributor

Will the default for both options be False for the time being?

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Will the default for both options be False for the time being?

Yes

@martindurant

Copy link
Copy Markdown
Contributor

Glad to see this happening, by the way. Once in, catalogs using intake-xarray can be updated and I don't thin the code will need to change.

@alimanfoo

Copy link
Copy Markdown
Contributor

Great to see this. On the API, FWIW I'd vote for using the same keyword (consolidated) in both, less burden on the user to remember what to use.

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Keywords are now all consolidated and all tests are go.

Ready to merge?

@jhammanjhamman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is basically ready. I had a few small questions/comments but this looks safe for a merge here soon.

Comment threaddoc/whats-new.rst
~~~~~~~~~~~~

- Ability to read and write consolidated metadata in zarr stores.
By `Ryan Abernathey <https://github.com/rabernat>`_.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you reference the issue this is attached to: (:issue:`2558`).


open_kwargs = dict(mode=mode, synchronizer=synchronizer, path=group)
if consolidated:
# TODO: an option to pass the metadata_key keyword

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

do we need to consider this TODO here?


open_kwargs = dict(mode=mode, synchronizer=synchronizer, path=group)
if consolidated:
# TODO: an option to pass the metadata_key keyword

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Anything to do here now?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we feel that it's important to expose this functionality from within xarray? I don't.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I also don't.
I think it's ok for xarray to have an opinion on what the special key is called.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I propose we just leave these TODO's here as is. If anyone ever needs this feature from the xarray side, this will help guide them on how to implement it.

@martindurant

Copy link
Copy Markdown
Contributor

LGTM

Do you think there should be more explicit text of how to add consolidation to existing zarr/xarray data-sets, rather than creating them with consolidation turned on?

We may also need some text around updating consolidated data-sets, but that can maybe wait to see what kind of usage people try.

@rabernat

Copy link
Copy Markdown
ContributorAuthor

We may also need some text around updating consolidated data-sets, but that can maybe wait to see what kind of usage people try.

Since xarray cannot append or modify in-place existing zarr stores, this seems outside the scope of xarray for now. But maybe it is worth mentioning in the docs.

@jhamman

Copy link
Copy Markdown
Member

I'm happy here. ...but Appveyor is not.

@shoyer

shoyer commented Dec 4, 2018

Copy link
Copy Markdown
Member

@rabernat if you're ready, let's merge this.

The failures on Appveyor are unrelated (an issue with int32 and cftime)

@rabernat

rabernat commented Dec 4, 2018 via email

Copy link
Copy Markdown
ContributorAuthor

@shoyer
shoyer merged commit 3ae93ac into pydata:masterDec 4, 2018
@rabernat

Copy link
Copy Markdown
ContributorAuthor

If anyone wants to see how awesome consolidated metadata is, you can try it in this binder:
https://github.com/rabernat/pangeo_ecco_examples/

I did a bit of lazy profiling here:
https://gist.github.com/rabernat/ce1fb414cf53541afe2245363b06c49d

Things that used to take ~40s now take ~1s. Especially since loading the data is one of the first steps in any pangeo notebook, this is a huge improvement in usability.

Thanks to everyone who helped make it happen!

@martindurant

Copy link
Copy Markdown
Contributor

I like those timings.

dcherian pushed a commit to yohai/xarray that referenced this pull request Dec 16, 2018
* upstream/master:
Feature: N-dimensional auto_combine (pydata#2553)
Support HighLevelGraphs (pydata#2603)
Bump cftime version in doc environment (pydata#2604)
use keep_attrs in binary operations II (pydata#2590)
Temporarily mark dask-dev build as an allowed failure (pydata#2602)
Fix wrong error message in interp() (pydata#2598)
Add dayofyear and dayofweek accessors (pydata#2599)
Fix h5netcdf saving scalars with filters or chunks (pydata#2591)
Minor update to PR template (pydata#2596)
Zarr consolidated (pydata#2559)
fix examples (pydata#2581)
Fix typo (pydata#2578)
Concat docstring typo (pydata#2577)
DOC: remove example using Dataset.T (pydata#2572)
python setup.py test now works by default (pydata#2573)
Return slices when possible from CFTimeIndex.get_loc() (pydata#2569)
DOC: fix computation.rst (pydata#2567)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

how to incorporate zarr's new open_consolidated method?

7 participants

@rabernat@pep8speaks@shoyer@lilyminium@martindurant@alimanfoo@jhamman
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Zarr consolidated - #2559

Merged
shoyer merged 16 commits into
pydata:masterfrom
rabernat:zarr_consolidated
Dec 4, 2018
Merged

Zarr consolidated#2559
shoyer merged 16 commits into
pydata:masterfrom
rabernat:zarr_consolidated

Conversation

@rabernat

@rabernatrabernat commented Nov 20, 2018

Copy link
Copy Markdown
Contributor

This PR adds support for reading and writing of consolidated metadata in zarr stores.

  • Closeshow to incorporate zarr's new open_consolidated method? #2558 (remove if there is no corresponding issue, which should only be the case for minor changes)
  • Tests added (for all bug fixes or enhancements)
  • Fully documented, including whats-new.rst for all changes and api.rst for new API (remove if this change should not be visible to users, e.g., if it is an internal clean-up, or if this is part of a larger project that will be documented later)

@pep8speaks

pep8speaks commented Nov 20, 2018

Copy link
Copy Markdown

Hello @rabernat! Thanks for updating the PR.

Line 240:80: E501 line too long (82 > 79 characters)

Comment last updated on December 04, 2018 at 19:34 Hours UTC

Comment threadxarray/backends/api.py Outdated
@rabernat

Copy link
Copy Markdown
ContributorAuthor

Ping @lilyminium for a review.

Comment threadxarray/backends/api.py Outdated
if consolidate:
import zarr
zarr.consolidate_metadata(store)
# do we need to reload ztore now that we have consolidated?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

would it make sense for zarr to handle this?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What do you mean?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I meant reloading the zarr store automatically.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think that would be hard to achieve. And I'm not sure it's necessary. Frankly I don't know why we return a store object from to_zarr at all.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

zarr.consolidate_metadata returns the output of open_consolidated on the same store, so this is already happening

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Also need to add some version checks...this will only work with zarr > 2.2.

Comment threaddoc/io.rst
Comment threadxarray/tests/test_backends.py Outdated
Comment threadxarray/backends/zarr.py Outdated

def __init__(self, zarr_group):
if consolidated or consolidate_on_close:
if LooseVersion(zarr.__version__) <= '2.2': # pragma: no cover

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

reminder to update this version check too.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Being more explicit about the version seems to fix this issue here. In the tests I have used the importorskip approach.

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Not sure I understand why there are tests failing now. The failing function is test_basic_compute.

https://travis-ci.org/pydata/xarray/jobs/460873430#L7489

At first glance, this does not appear to have anything to do with my PR. The relevant error is:


______________________________ test_basic_compute ______________________________
def test_basic_compute():
ds = Dataset({'foo': ('x', range(5)),
'bar': ('x', range(5))}).chunk({'x': 2})
for get in [dask.threaded.get,
dask.multiprocessing.get,
dask.local.get_sync,
None]:
with (dask.config.set(scheduler=get)
if LooseVersion(dask.__version__) >= LooseVersion('0.19.4')
else dask.config.set(scheduler=get)
if LooseVersion(dask.__version__) >= LooseVersion('0.18.0')
else dask.set_options(get=get)):
> ds.compute()
xarray/tests/test_dask.py:843: _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ xarray/core/dataset.py:597: in compute
return new.load(**kwargs)
xarray/core/dataset.py:494: in load
evaluated_data = da.compute(*lazy_data.values(), **kwargs)
../../../miniconda/envs/test_env/lib/python3.6/site-packages/dask/base.py:390: in compute
collections=collections)
../../../miniconda/envs/test_env/lib/python3.6/site-packages/dask/base.py:865: in get_scheduler
return get_scheduler(scheduler=config.get('scheduler', None))
_ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ get = None, scheduler = <function get at 0x7fc31d9ae048>, collections = None
cls = None
def get_scheduler(get=None, scheduler=None, collections=None, cls=None):
""" Get scheduler function
There are various ways to specify the scheduler to use:
1. Passing in get= parameters (deprecated)
2. Passing in scheduler= parameters
3. Passing these into global confiuration
4. Using defaults of a dask collection
This function centralizes the logic to determine the right scheduler to use
from those many options
"""
if get is not None:
if scheduler is not None:
raise ValueError("Both get= and scheduler= provided. Choose one")
warn_on_get(get)
return get
if scheduler is not None:
> if scheduler.lower() in named_schedulers:
E AttributeError: 'function' object has no attribute 'lower'
../../../miniconda/envs/test_env/lib/python3.6/site-packages/dask/base.py:854: AttributeError

@shoyer

Copy link
Copy Markdown
Member

I bet this is due to the latest dask release (1.0). We can fix this in another PR.

@lilyminium

Copy link
Copy Markdown
Contributor

I remember dealing with this in my pull request -- if I recall correctly scheduler was pointing to the scheduler.get function instead. It was a minor bug that was either fixed in the next release of xarray (0.11.0) or Dask (0.20.1).

@rabernat

Copy link
Copy Markdown
ContributorAuthor

So if the test issues can be considered resolved, the only decision we need to make is about the API.

Do we prefer (the current way):

ds.to_zarr(fname, consolidate=True)
xr.open_zarr(fname, consolidated=True)

or @shoyer's suggestion

ds.to_zarr(fname, consolidated=True)
xr.open_zarr(fname, consolidated=True)

???

@martindurant

Copy link
Copy Markdown
Contributor

Will the default for both options be False for the time being?

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Will the default for both options be False for the time being?

Yes

@martindurant

Copy link
Copy Markdown
Contributor

Glad to see this happening, by the way. Once in, catalogs using intake-xarray can be updated and I don't thin the code will need to change.

@alimanfoo

Copy link
Copy Markdown
Contributor

Great to see this. On the API, FWIW I'd vote for using the same keyword (consolidated) in both, less burden on the user to remember what to use.

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Keywords are now all consolidated and all tests are go.

Ready to merge?

@jhammanjhamman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is basically ready. I had a few small questions/comments but this looks safe for a merge here soon.

Comment threaddoc/whats-new.rst
~~~~~~~~~~~~

- Ability to read and write consolidated metadata in zarr stores.
By `Ryan Abernathey <https://github.com/rabernat>`_.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you reference the issue this is attached to: (:issue:`2558`).


open_kwargs = dict(mode=mode, synchronizer=synchronizer, path=group)
if consolidated:
# TODO: an option to pass the metadata_key keyword

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

do we need to consider this TODO here?


open_kwargs = dict(mode=mode, synchronizer=synchronizer, path=group)
if consolidated:
# TODO: an option to pass the metadata_key keyword

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Anything to do here now?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we feel that it's important to expose this functionality from within xarray? I don't.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I also don't.
I think it's ok for xarray to have an opinion on what the special key is called.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I propose we just leave these TODO's here as is. If anyone ever needs this feature from the xarray side, this will help guide them on how to implement it.

@martindurant

Copy link
Copy Markdown
Contributor

LGTM

Do you think there should be more explicit text of how to add consolidation to existing zarr/xarray data-sets, rather than creating them with consolidation turned on?

We may also need some text around updating consolidated data-sets, but that can maybe wait to see what kind of usage people try.

@rabernat

Copy link
Copy Markdown
ContributorAuthor

We may also need some text around updating consolidated data-sets, but that can maybe wait to see what kind of usage people try.

Since xarray cannot append or modify in-place existing zarr stores, this seems outside the scope of xarray for now. But maybe it is worth mentioning in the docs.

@jhamman

Copy link
Copy Markdown
Member

I'm happy here. ...but Appveyor is not.

@shoyer

shoyer commented Dec 4, 2018

Copy link
Copy Markdown
Member

@rabernat if you're ready, let's merge this.

The failures on Appveyor are unrelated (an issue with int32 and cftime)

@rabernat

rabernat commented Dec 4, 2018 via email

Copy link
Copy Markdown
ContributorAuthor

@shoyer
shoyer merged commit 3ae93ac into pydata:masterDec 4, 2018
@rabernat

Copy link
Copy Markdown
ContributorAuthor

If anyone wants to see how awesome consolidated metadata is, you can try it in this binder:
https://github.com/rabernat/pangeo_ecco_examples/

I did a bit of lazy profiling here:
https://gist.github.com/rabernat/ce1fb414cf53541afe2245363b06c49d

Things that used to take ~40s now take ~1s. Especially since loading the data is one of the first steps in any pangeo notebook, this is a huge improvement in usability.

Thanks to everyone who helped make it happen!

@martindurant

Copy link
Copy Markdown
Contributor

I like those timings.

dcherian pushed a commit to yohai/xarray that referenced this pull request Dec 16, 2018
* upstream/master:
Feature: N-dimensional auto_combine (pydata#2553)
Support HighLevelGraphs (pydata#2603)
Bump cftime version in doc environment (pydata#2604)
use keep_attrs in binary operations II (pydata#2590)
Temporarily mark dask-dev build as an allowed failure (pydata#2602)
Fix wrong error message in interp() (pydata#2598)
Add dayofyear and dayofweek accessors (pydata#2599)
Fix h5netcdf saving scalars with filters or chunks (pydata#2591)
Minor update to PR template (pydata#2596)
Zarr consolidated (pydata#2559)
fix examples (pydata#2581)
Fix typo (pydata#2578)
Concat docstring typo (pydata#2577)
DOC: remove example using Dataset.T (pydata#2572)
python setup.py test now works by default (pydata#2573)
Return slices when possible from CFTimeIndex.get_loc() (pydata#2569)
DOC: fix computation.rst (pydata#2567)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

how to incorporate zarr's new open_consolidated method?

7 participants

@rabernat@pep8speaks@shoyer@lilyminium@martindurant@alimanfoo@jhamman
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Zarr consolidated - #2559

Merged
shoyer merged 16 commits into
pydata:masterfrom
rabernat:zarr_consolidated
Dec 4, 2018
Merged

Zarr consolidated#2559
shoyer merged 16 commits into
pydata:masterfrom
rabernat:zarr_consolidated

Conversation

@rabernat

@rabernatrabernat commented Nov 20, 2018

Copy link
Copy Markdown
Contributor

This PR adds support for reading and writing of consolidated metadata in zarr stores.

  • Closeshow to incorporate zarr's new open_consolidated method? #2558 (remove if there is no corresponding issue, which should only be the case for minor changes)
  • Tests added (for all bug fixes or enhancements)
  • Fully documented, including whats-new.rst for all changes and api.rst for new API (remove if this change should not be visible to users, e.g., if it is an internal clean-up, or if this is part of a larger project that will be documented later)

@pep8speaks

pep8speaks commented Nov 20, 2018

Copy link
Copy Markdown

Hello @rabernat! Thanks for updating the PR.

Line 240:80: E501 line too long (82 > 79 characters)

Comment last updated on December 04, 2018 at 19:34 Hours UTC

Comment threadxarray/backends/api.py Outdated
@rabernat

Copy link
Copy Markdown
ContributorAuthor

Ping @lilyminium for a review.

Comment threadxarray/backends/api.py Outdated
if consolidate:
import zarr
zarr.consolidate_metadata(store)
# do we need to reload ztore now that we have consolidated?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

would it make sense for zarr to handle this?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What do you mean?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I meant reloading the zarr store automatically.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think that would be hard to achieve. And I'm not sure it's necessary. Frankly I don't know why we return a store object from to_zarr at all.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

zarr.consolidate_metadata returns the output of open_consolidated on the same store, so this is already happening

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Also need to add some version checks...this will only work with zarr > 2.2.

Comment threaddoc/io.rst
Comment threadxarray/tests/test_backends.py Outdated
Comment threadxarray/backends/zarr.py Outdated

def __init__(self, zarr_group):
if consolidated or consolidate_on_close:
if LooseVersion(zarr.__version__) <= '2.2': # pragma: no cover

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

reminder to update this version check too.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Being more explicit about the version seems to fix this issue here. In the tests I have used the importorskip approach.

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Not sure I understand why there are tests failing now. The failing function is test_basic_compute.

https://travis-ci.org/pydata/xarray/jobs/460873430#L7489

At first glance, this does not appear to have anything to do with my PR. The relevant error is:


______________________________ test_basic_compute ______________________________
def test_basic_compute():
ds = Dataset({'foo': ('x', range(5)),
'bar': ('x', range(5))}).chunk({'x': 2})
for get in [dask.threaded.get,
dask.multiprocessing.get,
dask.local.get_sync,
None]:
with (dask.config.set(scheduler=get)
if LooseVersion(dask.__version__) >= LooseVersion('0.19.4')
else dask.config.set(scheduler=get)
if LooseVersion(dask.__version__) >= LooseVersion('0.18.0')
else dask.set_options(get=get)):
> ds.compute()
xarray/tests/test_dask.py:843: _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ xarray/core/dataset.py:597: in compute
return new.load(**kwargs)
xarray/core/dataset.py:494: in load
evaluated_data = da.compute(*lazy_data.values(), **kwargs)
../../../miniconda/envs/test_env/lib/python3.6/site-packages/dask/base.py:390: in compute
collections=collections)
../../../miniconda/envs/test_env/lib/python3.6/site-packages/dask/base.py:865: in get_scheduler
return get_scheduler(scheduler=config.get('scheduler', None))
_ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ get = None, scheduler = <function get at 0x7fc31d9ae048>, collections = None
cls = None
def get_scheduler(get=None, scheduler=None, collections=None, cls=None):
""" Get scheduler function
There are various ways to specify the scheduler to use:
1. Passing in get= parameters (deprecated)
2. Passing in scheduler= parameters
3. Passing these into global confiuration
4. Using defaults of a dask collection
This function centralizes the logic to determine the right scheduler to use
from those many options
"""
if get is not None:
if scheduler is not None:
raise ValueError("Both get= and scheduler= provided. Choose one")
warn_on_get(get)
return get
if scheduler is not None:
> if scheduler.lower() in named_schedulers:
E AttributeError: 'function' object has no attribute 'lower'
../../../miniconda/envs/test_env/lib/python3.6/site-packages/dask/base.py:854: AttributeError

@shoyer

Copy link
Copy Markdown
Member

I bet this is due to the latest dask release (1.0). We can fix this in another PR.

@lilyminium

Copy link
Copy Markdown
Contributor

I remember dealing with this in my pull request -- if I recall correctly scheduler was pointing to the scheduler.get function instead. It was a minor bug that was either fixed in the next release of xarray (0.11.0) or Dask (0.20.1).

@rabernat

Copy link
Copy Markdown
ContributorAuthor

So if the test issues can be considered resolved, the only decision we need to make is about the API.

Do we prefer (the current way):

ds.to_zarr(fname, consolidate=True)
xr.open_zarr(fname, consolidated=True)

or @shoyer's suggestion

ds.to_zarr(fname, consolidated=True)
xr.open_zarr(fname, consolidated=True)

???

@martindurant

Copy link
Copy Markdown
Contributor

Will the default for both options be False for the time being?

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Will the default for both options be False for the time being?

Yes

@martindurant

Copy link
Copy Markdown
Contributor

Glad to see this happening, by the way. Once in, catalogs using intake-xarray can be updated and I don't thin the code will need to change.

@alimanfoo

Copy link
Copy Markdown
Contributor

Great to see this. On the API, FWIW I'd vote for using the same keyword (consolidated) in both, less burden on the user to remember what to use.

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Keywords are now all consolidated and all tests are go.

Ready to merge?

@jhammanjhamman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is basically ready. I had a few small questions/comments but this looks safe for a merge here soon.

Comment threaddoc/whats-new.rst
~~~~~~~~~~~~

- Ability to read and write consolidated metadata in zarr stores.
By `Ryan Abernathey <https://github.com/rabernat>`_.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you reference the issue this is attached to: (:issue:`2558`).


open_kwargs = dict(mode=mode, synchronizer=synchronizer, path=group)
if consolidated:
# TODO: an option to pass the metadata_key keyword

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

do we need to consider this TODO here?


open_kwargs = dict(mode=mode, synchronizer=synchronizer, path=group)
if consolidated:
# TODO: an option to pass the metadata_key keyword

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Anything to do here now?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we feel that it's important to expose this functionality from within xarray? I don't.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I also don't.
I think it's ok for xarray to have an opinion on what the special key is called.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I propose we just leave these TODO's here as is. If anyone ever needs this feature from the xarray side, this will help guide them on how to implement it.

@martindurant

Copy link
Copy Markdown
Contributor

LGTM

Do you think there should be more explicit text of how to add consolidation to existing zarr/xarray data-sets, rather than creating them with consolidation turned on?

We may also need some text around updating consolidated data-sets, but that can maybe wait to see what kind of usage people try.

@rabernat

Copy link
Copy Markdown
ContributorAuthor

We may also need some text around updating consolidated data-sets, but that can maybe wait to see what kind of usage people try.

Since xarray cannot append or modify in-place existing zarr stores, this seems outside the scope of xarray for now. But maybe it is worth mentioning in the docs.

@jhamman

Copy link
Copy Markdown
Member

I'm happy here. ...but Appveyor is not.

@shoyer

shoyer commented Dec 4, 2018

Copy link
Copy Markdown
Member

@rabernat if you're ready, let's merge this.

The failures on Appveyor are unrelated (an issue with int32 and cftime)

@rabernat

rabernat commented Dec 4, 2018 via email

Copy link
Copy Markdown
ContributorAuthor

@shoyer
shoyer merged commit 3ae93ac into pydata:masterDec 4, 2018
@rabernat

Copy link
Copy Markdown
ContributorAuthor

If anyone wants to see how awesome consolidated metadata is, you can try it in this binder:
https://github.com/rabernat/pangeo_ecco_examples/

I did a bit of lazy profiling here:
https://gist.github.com/rabernat/ce1fb414cf53541afe2245363b06c49d

Things that used to take ~40s now take ~1s. Especially since loading the data is one of the first steps in any pangeo notebook, this is a huge improvement in usability.

Thanks to everyone who helped make it happen!

@martindurant

Copy link
Copy Markdown
Contributor

I like those timings.

dcherian pushed a commit to yohai/xarray that referenced this pull request Dec 16, 2018
* upstream/master:
Feature: N-dimensional auto_combine (pydata#2553)
Support HighLevelGraphs (pydata#2603)
Bump cftime version in doc environment (pydata#2604)
use keep_attrs in binary operations II (pydata#2590)
Temporarily mark dask-dev build as an allowed failure (pydata#2602)
Fix wrong error message in interp() (pydata#2598)
Add dayofyear and dayofweek accessors (pydata#2599)
Fix h5netcdf saving scalars with filters or chunks (pydata#2591)
Minor update to PR template (pydata#2596)
Zarr consolidated (pydata#2559)
fix examples (pydata#2581)
Fix typo (pydata#2578)
Concat docstring typo (pydata#2577)
DOC: remove example using Dataset.T (pydata#2572)
python setup.py test now works by default (pydata#2573)
Return slices when possible from CFTimeIndex.get_loc() (pydata#2569)
DOC: fix computation.rst (pydata#2567)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

how to incorporate zarr's new open_consolidated method?

7 participants

@rabernat@pep8speaks@shoyer@lilyminium@martindurant@alimanfoo@jhamman
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Zarr consolidated - #2559

Merged
shoyer merged 16 commits into
pydata:masterfrom
rabernat:zarr_consolidated
Dec 4, 2018
Merged

Zarr consolidated#2559
shoyer merged 16 commits into
pydata:masterfrom
rabernat:zarr_consolidated

Conversation

@rabernat

@rabernatrabernat commented Nov 20, 2018

Copy link
Copy Markdown
Contributor

This PR adds support for reading and writing of consolidated metadata in zarr stores.

  • Closeshow to incorporate zarr's new open_consolidated method? #2558 (remove if there is no corresponding issue, which should only be the case for minor changes)
  • Tests added (for all bug fixes or enhancements)
  • Fully documented, including whats-new.rst for all changes and api.rst for new API (remove if this change should not be visible to users, e.g., if it is an internal clean-up, or if this is part of a larger project that will be documented later)

@pep8speaks

pep8speaks commented Nov 20, 2018

Copy link
Copy Markdown

Hello @rabernat! Thanks for updating the PR.

Line 240:80: E501 line too long (82 > 79 characters)

Comment last updated on December 04, 2018 at 19:34 Hours UTC

Comment threadxarray/backends/api.py Outdated
@rabernat

Copy link
Copy Markdown
ContributorAuthor

Ping @lilyminium for a review.

Comment threadxarray/backends/api.py Outdated
if consolidate:
import zarr
zarr.consolidate_metadata(store)
# do we need to reload ztore now that we have consolidated?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

would it make sense for zarr to handle this?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What do you mean?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I meant reloading the zarr store automatically.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think that would be hard to achieve. And I'm not sure it's necessary. Frankly I don't know why we return a store object from to_zarr at all.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

zarr.consolidate_metadata returns the output of open_consolidated on the same store, so this is already happening

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Also need to add some version checks...this will only work with zarr > 2.2.

Comment threaddoc/io.rst
Comment threadxarray/tests/test_backends.py Outdated
Comment threadxarray/backends/zarr.py Outdated

def __init__(self, zarr_group):
if consolidated or consolidate_on_close:
if LooseVersion(zarr.__version__) <= '2.2': # pragma: no cover

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

reminder to update this version check too.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Being more explicit about the version seems to fix this issue here. In the tests I have used the importorskip approach.

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Not sure I understand why there are tests failing now. The failing function is test_basic_compute.

https://travis-ci.org/pydata/xarray/jobs/460873430#L7489

At first glance, this does not appear to have anything to do with my PR. The relevant error is:


______________________________ test_basic_compute ______________________________
def test_basic_compute():
ds = Dataset({'foo': ('x', range(5)),
'bar': ('x', range(5))}).chunk({'x': 2})
for get in [dask.threaded.get,
dask.multiprocessing.get,
dask.local.get_sync,
None]:
with (dask.config.set(scheduler=get)
if LooseVersion(dask.__version__) >= LooseVersion('0.19.4')
else dask.config.set(scheduler=get)
if LooseVersion(dask.__version__) >= LooseVersion('0.18.0')
else dask.set_options(get=get)):
> ds.compute()
xarray/tests/test_dask.py:843: _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ xarray/core/dataset.py:597: in compute
return new.load(**kwargs)
xarray/core/dataset.py:494: in load
evaluated_data = da.compute(*lazy_data.values(), **kwargs)
../../../miniconda/envs/test_env/lib/python3.6/site-packages/dask/base.py:390: in compute
collections=collections)
../../../miniconda/envs/test_env/lib/python3.6/site-packages/dask/base.py:865: in get_scheduler
return get_scheduler(scheduler=config.get('scheduler', None))
_ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ get = None, scheduler = <function get at 0x7fc31d9ae048>, collections = None
cls = None
def get_scheduler(get=None, scheduler=None, collections=None, cls=None):
""" Get scheduler function
There are various ways to specify the scheduler to use:
1. Passing in get= parameters (deprecated)
2. Passing in scheduler= parameters
3. Passing these into global confiuration
4. Using defaults of a dask collection
This function centralizes the logic to determine the right scheduler to use
from those many options
"""
if get is not None:
if scheduler is not None:
raise ValueError("Both get= and scheduler= provided. Choose one")
warn_on_get(get)
return get
if scheduler is not None:
> if scheduler.lower() in named_schedulers:
E AttributeError: 'function' object has no attribute 'lower'
../../../miniconda/envs/test_env/lib/python3.6/site-packages/dask/base.py:854: AttributeError

@shoyer

Copy link
Copy Markdown
Member

I bet this is due to the latest dask release (1.0). We can fix this in another PR.

@lilyminium

Copy link
Copy Markdown
Contributor

I remember dealing with this in my pull request -- if I recall correctly scheduler was pointing to the scheduler.get function instead. It was a minor bug that was either fixed in the next release of xarray (0.11.0) or Dask (0.20.1).

@rabernat

Copy link
Copy Markdown
ContributorAuthor

So if the test issues can be considered resolved, the only decision we need to make is about the API.

Do we prefer (the current way):

ds.to_zarr(fname, consolidate=True)
xr.open_zarr(fname, consolidated=True)

or @shoyer's suggestion

ds.to_zarr(fname, consolidated=True)
xr.open_zarr(fname, consolidated=True)

???

@martindurant

Copy link
Copy Markdown
Contributor

Will the default for both options be False for the time being?

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Will the default for both options be False for the time being?

Yes

@martindurant

Copy link
Copy Markdown
Contributor

Glad to see this happening, by the way. Once in, catalogs using intake-xarray can be updated and I don't thin the code will need to change.

@alimanfoo

Copy link
Copy Markdown
Contributor

Great to see this. On the API, FWIW I'd vote for using the same keyword (consolidated) in both, less burden on the user to remember what to use.

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Keywords are now all consolidated and all tests are go.

Ready to merge?

@jhammanjhamman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is basically ready. I had a few small questions/comments but this looks safe for a merge here soon.

Comment threaddoc/whats-new.rst
~~~~~~~~~~~~

- Ability to read and write consolidated metadata in zarr stores.
By `Ryan Abernathey <https://github.com/rabernat>`_.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you reference the issue this is attached to: (:issue:`2558`).


open_kwargs = dict(mode=mode, synchronizer=synchronizer, path=group)
if consolidated:
# TODO: an option to pass the metadata_key keyword

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

do we need to consider this TODO here?


open_kwargs = dict(mode=mode, synchronizer=synchronizer, path=group)
if consolidated:
# TODO: an option to pass the metadata_key keyword

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Anything to do here now?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we feel that it's important to expose this functionality from within xarray? I don't.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I also don't.
I think it's ok for xarray to have an opinion on what the special key is called.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I propose we just leave these TODO's here as is. If anyone ever needs this feature from the xarray side, this will help guide them on how to implement it.

@martindurant

Copy link
Copy Markdown
Contributor

LGTM

Do you think there should be more explicit text of how to add consolidation to existing zarr/xarray data-sets, rather than creating them with consolidation turned on?

We may also need some text around updating consolidated data-sets, but that can maybe wait to see what kind of usage people try.

@rabernat

Copy link
Copy Markdown
ContributorAuthor

We may also need some text around updating consolidated data-sets, but that can maybe wait to see what kind of usage people try.

Since xarray cannot append or modify in-place existing zarr stores, this seems outside the scope of xarray for now. But maybe it is worth mentioning in the docs.

@jhamman

Copy link
Copy Markdown
Member

I'm happy here. ...but Appveyor is not.

@shoyer

shoyer commented Dec 4, 2018

Copy link
Copy Markdown
Member

@rabernat if you're ready, let's merge this.

The failures on Appveyor are unrelated (an issue with int32 and cftime)

@rabernat

rabernat commented Dec 4, 2018 via email

Copy link
Copy Markdown
ContributorAuthor

@shoyer
shoyer merged commit 3ae93ac into pydata:masterDec 4, 2018
@rabernat

Copy link
Copy Markdown
ContributorAuthor

If anyone wants to see how awesome consolidated metadata is, you can try it in this binder:
https://github.com/rabernat/pangeo_ecco_examples/

I did a bit of lazy profiling here:
https://gist.github.com/rabernat/ce1fb414cf53541afe2245363b06c49d

Things that used to take ~40s now take ~1s. Especially since loading the data is one of the first steps in any pangeo notebook, this is a huge improvement in usability.

Thanks to everyone who helped make it happen!

@martindurant

Copy link
Copy Markdown
Contributor

I like those timings.

dcherian pushed a commit to yohai/xarray that referenced this pull request Dec 16, 2018
* upstream/master:
Feature: N-dimensional auto_combine (pydata#2553)
Support HighLevelGraphs (pydata#2603)
Bump cftime version in doc environment (pydata#2604)
use keep_attrs in binary operations II (pydata#2590)
Temporarily mark dask-dev build as an allowed failure (pydata#2602)
Fix wrong error message in interp() (pydata#2598)
Add dayofyear and dayofweek accessors (pydata#2599)
Fix h5netcdf saving scalars with filters or chunks (pydata#2591)
Minor update to PR template (pydata#2596)
Zarr consolidated (pydata#2559)
fix examples (pydata#2581)
Fix typo (pydata#2578)
Concat docstring typo (pydata#2577)
DOC: remove example using Dataset.T (pydata#2572)
python setup.py test now works by default (pydata#2573)
Return slices when possible from CFTimeIndex.get_loc() (pydata#2569)
DOC: fix computation.rst (pydata#2567)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

how to incorporate zarr's new open_consolidated method?

7 participants

@rabernat@pep8speaks@shoyer@lilyminium@martindurant@alimanfoo@jhamman
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Zarr consolidated - #2559

Merged
shoyer merged 16 commits into
pydata:masterfrom
rabernat:zarr_consolidated
Dec 4, 2018
Merged

Zarr consolidated#2559
shoyer merged 16 commits into
pydata:masterfrom
rabernat:zarr_consolidated

Conversation

@rabernat

@rabernatrabernat commented Nov 20, 2018

Copy link
Copy Markdown
Contributor

This PR adds support for reading and writing of consolidated metadata in zarr stores.

  • Closeshow to incorporate zarr's new open_consolidated method? #2558 (remove if there is no corresponding issue, which should only be the case for minor changes)
  • Tests added (for all bug fixes or enhancements)
  • Fully documented, including whats-new.rst for all changes and api.rst for new API (remove if this change should not be visible to users, e.g., if it is an internal clean-up, or if this is part of a larger project that will be documented later)

@pep8speaks

pep8speaks commented Nov 20, 2018

Copy link
Copy Markdown

Hello @rabernat! Thanks for updating the PR.

Line 240:80: E501 line too long (82 > 79 characters)

Comment last updated on December 04, 2018 at 19:34 Hours UTC

Comment threadxarray/backends/api.py Outdated
@rabernat

Copy link
Copy Markdown
ContributorAuthor

Ping @lilyminium for a review.

Comment threadxarray/backends/api.py Outdated
if consolidate:
import zarr
zarr.consolidate_metadata(store)
# do we need to reload ztore now that we have consolidated?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

would it make sense for zarr to handle this?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What do you mean?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I meant reloading the zarr store automatically.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think that would be hard to achieve. And I'm not sure it's necessary. Frankly I don't know why we return a store object from to_zarr at all.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

zarr.consolidate_metadata returns the output of open_consolidated on the same store, so this is already happening

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Also need to add some version checks...this will only work with zarr > 2.2.

Comment threaddoc/io.rst
Comment threadxarray/tests/test_backends.py Outdated
Comment threadxarray/backends/zarr.py Outdated

def __init__(self, zarr_group):
if consolidated or consolidate_on_close:
if LooseVersion(zarr.__version__) <= '2.2': # pragma: no cover

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

reminder to update this version check too.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Being more explicit about the version seems to fix this issue here. In the tests I have used the importorskip approach.

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Not sure I understand why there are tests failing now. The failing function is test_basic_compute.

https://travis-ci.org/pydata/xarray/jobs/460873430#L7489

At first glance, this does not appear to have anything to do with my PR. The relevant error is:


______________________________ test_basic_compute ______________________________
def test_basic_compute():
ds = Dataset({'foo': ('x', range(5)),
'bar': ('x', range(5))}).chunk({'x': 2})
for get in [dask.threaded.get,
dask.multiprocessing.get,
dask.local.get_sync,
None]:
with (dask.config.set(scheduler=get)
if LooseVersion(dask.__version__) >= LooseVersion('0.19.4')
else dask.config.set(scheduler=get)
if LooseVersion(dask.__version__) >= LooseVersion('0.18.0')
else dask.set_options(get=get)):
> ds.compute()
xarray/tests/test_dask.py:843: _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ xarray/core/dataset.py:597: in compute
return new.load(**kwargs)
xarray/core/dataset.py:494: in load
evaluated_data = da.compute(*lazy_data.values(), **kwargs)
../../../miniconda/envs/test_env/lib/python3.6/site-packages/dask/base.py:390: in compute
collections=collections)
../../../miniconda/envs/test_env/lib/python3.6/site-packages/dask/base.py:865: in get_scheduler
return get_scheduler(scheduler=config.get('scheduler', None))
_ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ get = None, scheduler = <function get at 0x7fc31d9ae048>, collections = None
cls = None
def get_scheduler(get=None, scheduler=None, collections=None, cls=None):
""" Get scheduler function
There are various ways to specify the scheduler to use:
1. Passing in get= parameters (deprecated)
2. Passing in scheduler= parameters
3. Passing these into global confiuration
4. Using defaults of a dask collection
This function centralizes the logic to determine the right scheduler to use
from those many options
"""
if get is not None:
if scheduler is not None:
raise ValueError("Both get= and scheduler= provided. Choose one")
warn_on_get(get)
return get
if scheduler is not None:
> if scheduler.lower() in named_schedulers:
E AttributeError: 'function' object has no attribute 'lower'
../../../miniconda/envs/test_env/lib/python3.6/site-packages/dask/base.py:854: AttributeError

@shoyer

Copy link
Copy Markdown
Member

I bet this is due to the latest dask release (1.0). We can fix this in another PR.

@lilyminium

Copy link
Copy Markdown
Contributor

I remember dealing with this in my pull request -- if I recall correctly scheduler was pointing to the scheduler.get function instead. It was a minor bug that was either fixed in the next release of xarray (0.11.0) or Dask (0.20.1).

@rabernat

Copy link
Copy Markdown
ContributorAuthor

So if the test issues can be considered resolved, the only decision we need to make is about the API.

Do we prefer (the current way):

ds.to_zarr(fname, consolidate=True)
xr.open_zarr(fname, consolidated=True)

or @shoyer's suggestion

ds.to_zarr(fname, consolidated=True)
xr.open_zarr(fname, consolidated=True)

???

@martindurant

Copy link
Copy Markdown
Contributor

Will the default for both options be False for the time being?

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Will the default for both options be False for the time being?

Yes

@martindurant

Copy link
Copy Markdown
Contributor

Glad to see this happening, by the way. Once in, catalogs using intake-xarray can be updated and I don't thin the code will need to change.

@alimanfoo

Copy link
Copy Markdown
Contributor

Great to see this. On the API, FWIW I'd vote for using the same keyword (consolidated) in both, less burden on the user to remember what to use.

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Keywords are now all consolidated and all tests are go.

Ready to merge?

@jhammanjhamman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is basically ready. I had a few small questions/comments but this looks safe for a merge here soon.

Comment threaddoc/whats-new.rst
~~~~~~~~~~~~

- Ability to read and write consolidated metadata in zarr stores.
By `Ryan Abernathey <https://github.com/rabernat>`_.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you reference the issue this is attached to: (:issue:`2558`).


open_kwargs = dict(mode=mode, synchronizer=synchronizer, path=group)
if consolidated:
# TODO: an option to pass the metadata_key keyword

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

do we need to consider this TODO here?


open_kwargs = dict(mode=mode, synchronizer=synchronizer, path=group)
if consolidated:
# TODO: an option to pass the metadata_key keyword

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Anything to do here now?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we feel that it's important to expose this functionality from within xarray? I don't.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I also don't.
I think it's ok for xarray to have an opinion on what the special key is called.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I propose we just leave these TODO's here as is. If anyone ever needs this feature from the xarray side, this will help guide them on how to implement it.

@martindurant

Copy link
Copy Markdown
Contributor

LGTM

Do you think there should be more explicit text of how to add consolidation to existing zarr/xarray data-sets, rather than creating them with consolidation turned on?

We may also need some text around updating consolidated data-sets, but that can maybe wait to see what kind of usage people try.

@rabernat

Copy link
Copy Markdown
ContributorAuthor

We may also need some text around updating consolidated data-sets, but that can maybe wait to see what kind of usage people try.

Since xarray cannot append or modify in-place existing zarr stores, this seems outside the scope of xarray for now. But maybe it is worth mentioning in the docs.

@jhamman

Copy link
Copy Markdown
Member

I'm happy here. ...but Appveyor is not.

@shoyer

shoyer commented Dec 4, 2018

Copy link
Copy Markdown
Member

@rabernat if you're ready, let's merge this.

The failures on Appveyor are unrelated (an issue with int32 and cftime)

@rabernat

rabernat commented Dec 4, 2018 via email

Copy link
Copy Markdown
ContributorAuthor

@shoyer
shoyer merged commit 3ae93ac into pydata:masterDec 4, 2018
@rabernat

Copy link
Copy Markdown
ContributorAuthor

If anyone wants to see how awesome consolidated metadata is, you can try it in this binder:
https://github.com/rabernat/pangeo_ecco_examples/

I did a bit of lazy profiling here:
https://gist.github.com/rabernat/ce1fb414cf53541afe2245363b06c49d

Things that used to take ~40s now take ~1s. Especially since loading the data is one of the first steps in any pangeo notebook, this is a huge improvement in usability.

Thanks to everyone who helped make it happen!

@martindurant

Copy link
Copy Markdown
Contributor

I like those timings.

dcherian pushed a commit to yohai/xarray that referenced this pull request Dec 16, 2018
* upstream/master:
Feature: N-dimensional auto_combine (pydata#2553)
Support HighLevelGraphs (pydata#2603)
Bump cftime version in doc environment (pydata#2604)
use keep_attrs in binary operations II (pydata#2590)
Temporarily mark dask-dev build as an allowed failure (pydata#2602)
Fix wrong error message in interp() (pydata#2598)
Add dayofyear and dayofweek accessors (pydata#2599)
Fix h5netcdf saving scalars with filters or chunks (pydata#2591)
Minor update to PR template (pydata#2596)
Zarr consolidated (pydata#2559)
fix examples (pydata#2581)
Fix typo (pydata#2578)
Concat docstring typo (pydata#2577)
DOC: remove example using Dataset.T (pydata#2572)
python setup.py test now works by default (pydata#2573)
Return slices when possible from CFTimeIndex.get_loc() (pydata#2569)
DOC: fix computation.rst (pydata#2567)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

how to incorporate zarr's new open_consolidated method?

7 participants

@rabernat@pep8speaks@shoyer@lilyminium@martindurant@alimanfoo@jhamman
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Zarr consolidated - #2559

Merged
shoyer merged 16 commits into
pydata:masterfrom
rabernat:zarr_consolidated
Dec 4, 2018
Merged

Zarr consolidated#2559
shoyer merged 16 commits into
pydata:masterfrom
rabernat:zarr_consolidated

Conversation

@rabernat

@rabernatrabernat commented Nov 20, 2018

Copy link
Copy Markdown
Contributor

This PR adds support for reading and writing of consolidated metadata in zarr stores.

  • Closeshow to incorporate zarr's new open_consolidated method? #2558 (remove if there is no corresponding issue, which should only be the case for minor changes)
  • Tests added (for all bug fixes or enhancements)
  • Fully documented, including whats-new.rst for all changes and api.rst for new API (remove if this change should not be visible to users, e.g., if it is an internal clean-up, or if this is part of a larger project that will be documented later)

@pep8speaks

pep8speaks commented Nov 20, 2018

Copy link
Copy Markdown

Hello @rabernat! Thanks for updating the PR.

Line 240:80: E501 line too long (82 > 79 characters)

Comment last updated on December 04, 2018 at 19:34 Hours UTC

Comment threadxarray/backends/api.py Outdated
@rabernat

Copy link
Copy Markdown
ContributorAuthor

Ping @lilyminium for a review.

Comment threadxarray/backends/api.py Outdated
if consolidate:
import zarr
zarr.consolidate_metadata(store)
# do we need to reload ztore now that we have consolidated?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

would it make sense for zarr to handle this?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What do you mean?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I meant reloading the zarr store automatically.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think that would be hard to achieve. And I'm not sure it's necessary. Frankly I don't know why we return a store object from to_zarr at all.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

zarr.consolidate_metadata returns the output of open_consolidated on the same store, so this is already happening

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Also need to add some version checks...this will only work with zarr > 2.2.

Comment threaddoc/io.rst
Comment threadxarray/tests/test_backends.py Outdated
Comment threadxarray/backends/zarr.py Outdated

def __init__(self, zarr_group):
if consolidated or consolidate_on_close:
if LooseVersion(zarr.__version__) <= '2.2': # pragma: no cover

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

reminder to update this version check too.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Being more explicit about the version seems to fix this issue here. In the tests I have used the importorskip approach.

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Not sure I understand why there are tests failing now. The failing function is test_basic_compute.

https://travis-ci.org/pydata/xarray/jobs/460873430#L7489

At first glance, this does not appear to have anything to do with my PR. The relevant error is:


______________________________ test_basic_compute ______________________________
def test_basic_compute():
ds = Dataset({'foo': ('x', range(5)),
'bar': ('x', range(5))}).chunk({'x': 2})
for get in [dask.threaded.get,
dask.multiprocessing.get,
dask.local.get_sync,
None]:
with (dask.config.set(scheduler=get)
if LooseVersion(dask.__version__) >= LooseVersion('0.19.4')
else dask.config.set(scheduler=get)
if LooseVersion(dask.__version__) >= LooseVersion('0.18.0')
else dask.set_options(get=get)):
> ds.compute()
xarray/tests/test_dask.py:843: _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ xarray/core/dataset.py:597: in compute
return new.load(**kwargs)
xarray/core/dataset.py:494: in load
evaluated_data = da.compute(*lazy_data.values(), **kwargs)
../../../miniconda/envs/test_env/lib/python3.6/site-packages/dask/base.py:390: in compute
collections=collections)
../../../miniconda/envs/test_env/lib/python3.6/site-packages/dask/base.py:865: in get_scheduler
return get_scheduler(scheduler=config.get('scheduler', None))
_ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ get = None, scheduler = <function get at 0x7fc31d9ae048>, collections = None
cls = None
def get_scheduler(get=None, scheduler=None, collections=None, cls=None):
""" Get scheduler function
There are various ways to specify the scheduler to use:
1. Passing in get= parameters (deprecated)
2. Passing in scheduler= parameters
3. Passing these into global confiuration
4. Using defaults of a dask collection
This function centralizes the logic to determine the right scheduler to use
from those many options
"""
if get is not None:
if scheduler is not None:
raise ValueError("Both get= and scheduler= provided. Choose one")
warn_on_get(get)
return get
if scheduler is not None:
> if scheduler.lower() in named_schedulers:
E AttributeError: 'function' object has no attribute 'lower'
../../../miniconda/envs/test_env/lib/python3.6/site-packages/dask/base.py:854: AttributeError

@shoyer

Copy link
Copy Markdown
Member

I bet this is due to the latest dask release (1.0). We can fix this in another PR.

@lilyminium

Copy link
Copy Markdown
Contributor

I remember dealing with this in my pull request -- if I recall correctly scheduler was pointing to the scheduler.get function instead. It was a minor bug that was either fixed in the next release of xarray (0.11.0) or Dask (0.20.1).

@rabernat

Copy link
Copy Markdown
ContributorAuthor

So if the test issues can be considered resolved, the only decision we need to make is about the API.

Do we prefer (the current way):

ds.to_zarr(fname, consolidate=True)
xr.open_zarr(fname, consolidated=True)

or @shoyer's suggestion

ds.to_zarr(fname, consolidated=True)
xr.open_zarr(fname, consolidated=True)

???

@martindurant

Copy link
Copy Markdown
Contributor

Will the default for both options be False for the time being?

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Will the default for both options be False for the time being?

Yes

@martindurant

Copy link
Copy Markdown
Contributor

Glad to see this happening, by the way. Once in, catalogs using intake-xarray can be updated and I don't thin the code will need to change.

@alimanfoo

Copy link
Copy Markdown
Contributor

Great to see this. On the API, FWIW I'd vote for using the same keyword (consolidated) in both, less burden on the user to remember what to use.

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Keywords are now all consolidated and all tests are go.

Ready to merge?

@jhammanjhamman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is basically ready. I had a few small questions/comments but this looks safe for a merge here soon.

Comment threaddoc/whats-new.rst
~~~~~~~~~~~~

- Ability to read and write consolidated metadata in zarr stores.
By `Ryan Abernathey <https://github.com/rabernat>`_.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you reference the issue this is attached to: (:issue:`2558`).


open_kwargs = dict(mode=mode, synchronizer=synchronizer, path=group)
if consolidated:
# TODO: an option to pass the metadata_key keyword

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

do we need to consider this TODO here?


open_kwargs = dict(mode=mode, synchronizer=synchronizer, path=group)
if consolidated:
# TODO: an option to pass the metadata_key keyword

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Anything to do here now?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we feel that it's important to expose this functionality from within xarray? I don't.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I also don't.
I think it's ok for xarray to have an opinion on what the special key is called.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I propose we just leave these TODO's here as is. If anyone ever needs this feature from the xarray side, this will help guide them on how to implement it.

@martindurant

Copy link
Copy Markdown
Contributor

LGTM

Do you think there should be more explicit text of how to add consolidation to existing zarr/xarray data-sets, rather than creating them with consolidation turned on?

We may also need some text around updating consolidated data-sets, but that can maybe wait to see what kind of usage people try.

@rabernat

Copy link
Copy Markdown
ContributorAuthor

We may also need some text around updating consolidated data-sets, but that can maybe wait to see what kind of usage people try.

Since xarray cannot append or modify in-place existing zarr stores, this seems outside the scope of xarray for now. But maybe it is worth mentioning in the docs.

@jhamman

Copy link
Copy Markdown
Member

I'm happy here. ...but Appveyor is not.

@shoyer

shoyer commented Dec 4, 2018

Copy link
Copy Markdown
Member

@rabernat if you're ready, let's merge this.

The failures on Appveyor are unrelated (an issue with int32 and cftime)

@rabernat

rabernat commented Dec 4, 2018 via email

Copy link
Copy Markdown
ContributorAuthor

@shoyer
shoyer merged commit 3ae93ac into pydata:masterDec 4, 2018
@rabernat

Copy link
Copy Markdown
ContributorAuthor

If anyone wants to see how awesome consolidated metadata is, you can try it in this binder:
https://github.com/rabernat/pangeo_ecco_examples/

I did a bit of lazy profiling here:
https://gist.github.com/rabernat/ce1fb414cf53541afe2245363b06c49d

Things that used to take ~40s now take ~1s. Especially since loading the data is one of the first steps in any pangeo notebook, this is a huge improvement in usability.

Thanks to everyone who helped make it happen!

@martindurant

Copy link
Copy Markdown
Contributor

I like those timings.

dcherian pushed a commit to yohai/xarray that referenced this pull request Dec 16, 2018
* upstream/master:
Feature: N-dimensional auto_combine (pydata#2553)
Support HighLevelGraphs (pydata#2603)
Bump cftime version in doc environment (pydata#2604)
use keep_attrs in binary operations II (pydata#2590)
Temporarily mark dask-dev build as an allowed failure (pydata#2602)
Fix wrong error message in interp() (pydata#2598)
Add dayofyear and dayofweek accessors (pydata#2599)
Fix h5netcdf saving scalars with filters or chunks (pydata#2591)
Minor update to PR template (pydata#2596)
Zarr consolidated (pydata#2559)
fix examples (pydata#2581)
Fix typo (pydata#2578)
Concat docstring typo (pydata#2577)
DOC: remove example using Dataset.T (pydata#2572)
python setup.py test now works by default (pydata#2573)
Return slices when possible from CFTimeIndex.get_loc() (pydata#2569)
DOC: fix computation.rst (pydata#2567)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

how to incorporate zarr's new open_consolidated method?

7 participants

@rabernat@pep8speaks@shoyer@lilyminium@martindurant@alimanfoo@jhamman
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Zarr consolidated - #2559

Merged
shoyer merged 16 commits into
pydata:masterfrom
rabernat:zarr_consolidated
Dec 4, 2018
Merged

Zarr consolidated#2559
shoyer merged 16 commits into
pydata:masterfrom
rabernat:zarr_consolidated

Conversation

@rabernat

@rabernatrabernat commented Nov 20, 2018

Copy link
Copy Markdown
Contributor

This PR adds support for reading and writing of consolidated metadata in zarr stores.

  • Closeshow to incorporate zarr's new open_consolidated method? #2558 (remove if there is no corresponding issue, which should only be the case for minor changes)
  • Tests added (for all bug fixes or enhancements)
  • Fully documented, including whats-new.rst for all changes and api.rst for new API (remove if this change should not be visible to users, e.g., if it is an internal clean-up, or if this is part of a larger project that will be documented later)

@pep8speaks

pep8speaks commented Nov 20, 2018

Copy link
Copy Markdown

Hello @rabernat! Thanks for updating the PR.

Line 240:80: E501 line too long (82 > 79 characters)

Comment last updated on December 04, 2018 at 19:34 Hours UTC

Comment threadxarray/backends/api.py Outdated
@rabernat

Copy link
Copy Markdown
ContributorAuthor

Ping @lilyminium for a review.

Comment threadxarray/backends/api.py Outdated
if consolidate:
import zarr
zarr.consolidate_metadata(store)
# do we need to reload ztore now that we have consolidated?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

would it make sense for zarr to handle this?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What do you mean?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I meant reloading the zarr store automatically.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think that would be hard to achieve. And I'm not sure it's necessary. Frankly I don't know why we return a store object from to_zarr at all.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

zarr.consolidate_metadata returns the output of open_consolidated on the same store, so this is already happening

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Also need to add some version checks...this will only work with zarr > 2.2.

Comment threaddoc/io.rst
Comment threadxarray/tests/test_backends.py Outdated
Comment threadxarray/backends/zarr.py Outdated

def __init__(self, zarr_group):
if consolidated or consolidate_on_close:
if LooseVersion(zarr.__version__) <= '2.2': # pragma: no cover

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

reminder to update this version check too.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Being more explicit about the version seems to fix this issue here. In the tests I have used the importorskip approach.

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Not sure I understand why there are tests failing now. The failing function is test_basic_compute.

https://travis-ci.org/pydata/xarray/jobs/460873430#L7489

At first glance, this does not appear to have anything to do with my PR. The relevant error is:


______________________________ test_basic_compute ______________________________
def test_basic_compute():
ds = Dataset({'foo': ('x', range(5)),
'bar': ('x', range(5))}).chunk({'x': 2})
for get in [dask.threaded.get,
dask.multiprocessing.get,
dask.local.get_sync,
None]:
with (dask.config.set(scheduler=get)
if LooseVersion(dask.__version__) >= LooseVersion('0.19.4')
else dask.config.set(scheduler=get)
if LooseVersion(dask.__version__) >= LooseVersion('0.18.0')
else dask.set_options(get=get)):
> ds.compute()
xarray/tests/test_dask.py:843: _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ xarray/core/dataset.py:597: in compute
return new.load(**kwargs)
xarray/core/dataset.py:494: in load
evaluated_data = da.compute(*lazy_data.values(), **kwargs)
../../../miniconda/envs/test_env/lib/python3.6/site-packages/dask/base.py:390: in compute
collections=collections)
../../../miniconda/envs/test_env/lib/python3.6/site-packages/dask/base.py:865: in get_scheduler
return get_scheduler(scheduler=config.get('scheduler', None))
_ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ get = None, scheduler = <function get at 0x7fc31d9ae048>, collections = None
cls = None
def get_scheduler(get=None, scheduler=None, collections=None, cls=None):
""" Get scheduler function
There are various ways to specify the scheduler to use:
1. Passing in get= parameters (deprecated)
2. Passing in scheduler= parameters
3. Passing these into global confiuration
4. Using defaults of a dask collection
This function centralizes the logic to determine the right scheduler to use
from those many options
"""
if get is not None:
if scheduler is not None:
raise ValueError("Both get= and scheduler= provided. Choose one")
warn_on_get(get)
return get
if scheduler is not None:
> if scheduler.lower() in named_schedulers:
E AttributeError: 'function' object has no attribute 'lower'
../../../miniconda/envs/test_env/lib/python3.6/site-packages/dask/base.py:854: AttributeError

@shoyer

Copy link
Copy Markdown
Member

I bet this is due to the latest dask release (1.0). We can fix this in another PR.

@lilyminium

Copy link
Copy Markdown
Contributor

I remember dealing with this in my pull request -- if I recall correctly scheduler was pointing to the scheduler.get function instead. It was a minor bug that was either fixed in the next release of xarray (0.11.0) or Dask (0.20.1).

@rabernat

Copy link
Copy Markdown
ContributorAuthor

So if the test issues can be considered resolved, the only decision we need to make is about the API.

Do we prefer (the current way):

ds.to_zarr(fname, consolidate=True)
xr.open_zarr(fname, consolidated=True)

or @shoyer's suggestion

ds.to_zarr(fname, consolidated=True)
xr.open_zarr(fname, consolidated=True)

???

@martindurant

Copy link
Copy Markdown
Contributor

Will the default for both options be False for the time being?

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Will the default for both options be False for the time being?

Yes

@martindurant

Copy link
Copy Markdown
Contributor

Glad to see this happening, by the way. Once in, catalogs using intake-xarray can be updated and I don't thin the code will need to change.

@alimanfoo

Copy link
Copy Markdown
Contributor

Great to see this. On the API, FWIW I'd vote for using the same keyword (consolidated) in both, less burden on the user to remember what to use.

@rabernat

Copy link
Copy Markdown
ContributorAuthor

Keywords are now all consolidated and all tests are go.

Ready to merge?

@jhammanjhamman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is basically ready. I had a few small questions/comments but this looks safe for a merge here soon.

Comment threaddoc/whats-new.rst
~~~~~~~~~~~~

- Ability to read and write consolidated metadata in zarr stores.
By `Ryan Abernathey <https://github.com/rabernat>`_.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you reference the issue this is attached to: (:issue:`2558`).


open_kwargs = dict(mode=mode, synchronizer=synchronizer, path=group)
if consolidated:
# TODO: an option to pass the metadata_key keyword

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

do we need to consider this TODO here?


open_kwargs = dict(mode=mode, synchronizer=synchronizer, path=group)
if consolidated:
# TODO: an option to pass the metadata_key keyword

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Anything to do here now?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we feel that it's important to expose this functionality from within xarray? I don't.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I also don't.
I think it's ok for xarray to have an opinion on what the special key is called.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I propose we just leave these TODO's here as is. If anyone ever needs this feature from the xarray side, this will help guide them on how to implement it.

@martindurant

Copy link
Copy Markdown
Contributor

LGTM

Do you think there should be more explicit text of how to add consolidation to existing zarr/xarray data-sets, rather than creating them with consolidation turned on?

We may also need some text around updating consolidated data-sets, but that can maybe wait to see what kind of usage people try.

@rabernat

Copy link
Copy Markdown
ContributorAuthor

We may also need some text around updating consolidated data-sets, but that can maybe wait to see what kind of usage people try.

Since xarray cannot append or modify in-place existing zarr stores, this seems outside the scope of xarray for now. But maybe it is worth mentioning in the docs.

@jhamman

Copy link
Copy Markdown
Member

I'm happy here. ...but Appveyor is not.

@shoyer

shoyer commented Dec 4, 2018

Copy link
Copy Markdown
Member

@rabernat if you're ready, let's merge this.

The failures on Appveyor are unrelated (an issue with int32 and cftime)

@rabernat

rabernat commented Dec 4, 2018 via email

Copy link
Copy Markdown
ContributorAuthor

@shoyer
shoyer merged commit 3ae93ac into pydata:masterDec 4, 2018
@rabernat

Copy link
Copy Markdown
ContributorAuthor

If anyone wants to see how awesome consolidated metadata is, you can try it in this binder:
https://github.com/rabernat/pangeo_ecco_examples/

I did a bit of lazy profiling here:
https://gist.github.com/rabernat/ce1fb414cf53541afe2245363b06c49d

Things that used to take ~40s now take ~1s. Especially since loading the data is one of the first steps in any pangeo notebook, this is a huge improvement in usability.

Thanks to everyone who helped make it happen!

@martindurant

Copy link
Copy Markdown
Contributor

I like those timings.

dcherian pushed a commit to yohai/xarray that referenced this pull request Dec 16, 2018
* upstream/master:
Feature: N-dimensional auto_combine (pydata#2553)
Support HighLevelGraphs (pydata#2603)
Bump cftime version in doc environment (pydata#2604)
use keep_attrs in binary operations II (pydata#2590)
Temporarily mark dask-dev build as an allowed failure (pydata#2602)
Fix wrong error message in interp() (pydata#2598)
Add dayofyear and dayofweek accessors (pydata#2599)
Fix h5netcdf saving scalars with filters or chunks (pydata#2591)
Minor update to PR template (pydata#2596)
Zarr consolidated (pydata#2559)
fix examples (pydata#2581)
Fix typo (pydata#2578)
Concat docstring typo (pydata#2577)
DOC: remove example using Dataset.T (pydata#2572)
python setup.py test now works by default (pydata#2573)
Return slices when possible from CFTimeIndex.get_loc() (pydata#2569)
DOC: fix computation.rst (pydata#2567)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

how to incorporate zarr's new open_consolidated method?

7 participants

@rabernat@pep8speaks@shoyer@lilyminium@martindurant@alimanfoo@jhamman