Bump Numcodecs requirement to 0.6.1 - #347

Closed
jakirkham wants to merge 21 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.1
Closed

Bump Numcodecs requirement to 0.6.1#347
jakirkham wants to merge 21 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.1

Conversation

@jakirkham

Copy link
Copy Markdown
Member

As there are some critical fixes in the latest Numcodecs, this bumps our lower bound to the latest version.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

@alimanfooalimanfoo added this to the v2.3 milestone Nov 30, 2018

@alimanfooalimanfoo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @jakirkham. Probably worth adding a brief release note.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

FTR saw some test failures locally that it looks like CI also encounters. Figured we'd want to look at these and determine how we want to fix them.

@alimanfoo

Copy link
Copy Markdown
Member

Ah, yes, sorry, forgot about that. I expect some test failures due to #324. May be other consequences of the upgrade too. Suggest dealing with them in this PR.

@jakirkham

jakirkham commented Nov 30, 2018

Copy link
Copy Markdown
MemberAuthor

SGTM. How about this one (MsgPack)?

# create an object array using picklez=self.create_array(shape=10, chunks=3, dtype=object, object_codec=Pickle())
z[0] ='foo'>assertz[0] =='foo'

ref: https://travis-ci.org/zarr-developers/zarr/jobs/461883537#L2717-L2720

...or this one (Categorize)?

withpytest.raises(RuntimeError):
# noinspection PyStatementEffect>v[:]
EFailed: DIDNOTRAISE<type'exceptions.RuntimeError'>

ref: https://travis-ci.org/zarr-developers/zarr/jobs/461883537#L2621-L2624

Previously MsgPack was turning bytes objects to unicode objects when
round-tripping them. However this has been fixed in the latest version
of Numcodecs. So correct this test now that MsgPack is working
correctly.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Also there is another failure. Appears Pickle wasn't able to handle general buffer protocol conforming types in decode. Though every other codec worked fine with these (which seems to be a consequence of them using the new utility functions 😉). Apparently Pickle was not using these. Have put together PR ( zarr-developers/numcodecs#143 ) with a fix for this. Would be good to do another patch release to get that out (and anything else these failures might need).

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Based on testing locally, it looks like we are down to the RuntimeError not being raised by an object array after removing the filters.

forcompressorinZlib(1), Blosc():
z=self.create_array(shape=len(data), chunks=30, dtype=object,
object_codec=Categorize(greetings,
dtype=object),
compressor=compressor)
z[:] =datav=z.view(filters=[])
withpytest.raises(RuntimeError):
# noinspection PyStatementEffect>v[:]
EFailed: DIDNOTRAISE<type'exceptions.RuntimeError'>

@alimanfoo

Copy link
Copy Markdown
Member

Hi @jakirkham, took the liberty to push a commit that should resolve the test failure due to RuntimeError not being raised for object arrays.

Other issues should be resolved with the fixes you've merged into numcodecs, so I'll cut a numcodecs 0.6.2 release then come back and update the version requirement here.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Thanks @alimanfoo!

As we already ensured the `chunk` is an `ndarray` viewing the original
data, there is no need for us to do that here as well. Plus the checks
performed by `ensure_contiguous_ndarray` are not needed for our use case
here. Particularly as we have already handled the unusual type cases
above. We also don't need to constrain the buffer size. As such the only
thing we really need is to flatten the array and make it contiguous,
which is what we handle here directly.
As both the expected `object` case and the non-`object` case perform a
`reshape` to flatten the data, go ahead and refactor that out of both
cases and handle it generally. Simplifies the code a bit.
As refactoring of the `reshape` step has effectively dropped the
expected `object` type case, the checks for different types is a little
more complicated than needed. To fix this, basically invert and swap the
case ordering. This way we can handle all generally expected types first
and simply cast them. Then we can raise if an `object` type shows up and
is unexpected.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Did a little bit of refactoring on your change. Hope that is ok. Please let me know your thoughts on that.

The `DictStore` is pretty reliant on the fact that values are immutable
and can be easily compared. For example `__eq__` assumes that all
contents can be compared easily. This works fine if the data is `bytes`.
However it doesn't really work for `ndarray`s for example. Previously we
would have stored whatever the user gave us here. This means comparisons
could falldown in those cases as well (much as the example in the
tutorial has highlighted on CI). Now we effectively require that the
data be something that can either be coerced to `bytes` (e.g. via the
new/old buffer protocol) or is `bytes` to begin with. Make sure not to
force this requirement when nesting one `MutableMapping` within another.
This test case seems to be ill-posed. Anytime we store `object`s to
`Array`s we require an `object_codec` to be specified. Otherwise we have
no clean way to serialize the data. However this `DictStore` test breaks
that assumption by explicitly storing an `object` type in it even though
this would never work for the other stores (particularly when working
with `Array`s). This includes in-memory Zarr `Array`s, which would be
backed by `DictStore`. Given this, we go ahead and drop this test case.
Instead of using a Python `dict` as the `default` store for a Zarr
`Array`, use the `DictStore`. This ensures that all blobs will be
represented as `bytes` regardless of what the user provided as data.
Thus things like comparisons of stores will work well in the default
case.
@jakirkham

jakirkham commented Dec 2, 2018

Copy link
Copy Markdown
MemberAuthor

On a different point, it looks like __eq__ in DictStore is assuming all items in it are comparable. This assumption runs into problems however when anything that is not easily comparable is stored in the DictStore. For instance, ndarray is comparable, but does not reduce to a bool as expected, which causes the tutorial's comparison to fail CI. There are likely other ill-posed situations we can imagine.

Looking at this example before the recent Numcodecs release, it seems the contents of DictStore were always being filled with bytes likely as an indirect consequence of applying filters to the data first. Have added a commit, which enforces this behavior intentionally using ensure_bytes from Numcodecs. Thus making this behavior something we can rely upon for things like __eq__ and such.

However there is a wrinkle with this approach. It appears we had a test that was assuming that objects could be placed in DictStores. Though this behavior doesn't really work for any of the other stores and we usually require an object_codec to be specified to even serialize an object type (specifically with Array). Given this, I've dropped that test and am contending it doesn't conform to our expectations about how stores typically should work with object types.

Also it appears that the Array's default store was a dict, which we cannot apply these fixes to. So have updated the default to DictStore. I'm not sure why this wasn't already the case, which means I'm probably missing something.

Happy to hear thoughts or receive pushback on this approach as I may very well be missing something. If you know of a better approach, would also be interested in hearing about that as well.

Also raised issue ( #348 ) to provide a nice simple example of this problem.

@jakirkham

jakirkham commented Dec 2, 2018

Copy link
Copy Markdown
MemberAuthor

Seems the tutorial test is now hung up on some whitespace issue. Not totally sure how that got introduced.

NVM this was more fallout from the switch to DictStore as the default store for Array.

Edit: This has since been fixed.

As we are now using `DictStore` to back the `Array`, we can correctly
measure how much memory it is using. So update the examples in `info`
and the tutorial to show how much memory is being used. Also update the
store type listed in info as well.
As Numcodecs now includes a very versatile and effective `ensure_bytes`
function, there is no need to define our own in `zarr.storage` as well.
So go ahead and drop it.
As this is no longer being used by `ensure_bytes` as that function was
dropped, go ahead and drop `binary_type` as well.
Make use of Numcodecs' `ensure_ndarray` to take `ndarray` views onto
buffers to be stored in a few cases so as to reshape them and avoid a
copy (thanks to the buffer protocol).
Rewrite `buffer_size` to just use Numcodecs' `ensure_ndarray` to get an
`ndarray` that views the data. Once the `ndarray` is gotten, all that is
needed is to access its `nbytes` member, which returns the number of
bytes that it takes up.
If the data is already a `str` instance, turn `ensure_str` into a no-op.
For all other cases, make use of Numcodecs' `ensure_bytes` to aid
`ensure_str` in coercing data through the buffer protocol. If we are on
Python 3, then decode the `bytes` object to a `str`.
As `DictStore` now must only store `bytes` or types coercible to bytes
via the buffer protocol, there is no possibility for it to have unknown
sizes as `bytes` always have a known size. So drop these cases where the
size can be `-1`.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Simplifies some of the code by making use of Numcodecs utility functions. This avoids a few copies when working with some of the stores. Also generally makes the rest easier to understand.

Make sure that datetime/timedelta arrays are cast to a type that
supports the buffer protocol. Ensure this is a type that can handle all
of the datetime/timedelta values and has the same itemsize.
Instead of using `ensure_ndarray`, use `ensure_contiguous_ndarray` with
the stores. This ensures that datetime/timedeltas are handled by
default. Also catches things like object arrays. Finally this handles
flattening the array if needed.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Replacing with PR ( #352 ), which goes to Numcodecs 0.6.2. It also consolidates the changes from here and drops a few that have been pulled into other PRs. Namely ensuring DictStore contains only bytes for data ( #350 ) and changing the default backend of Array to DictStore ( #351 ).

@jakirkhamjakirkham removed this from the v2.3 milestone Dec 4, 2018
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jakirkham@alimanfoo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Bump Numcodecs requirement to 0.6.1 - #347

Closed
jakirkham wants to merge 21 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.1
Closed

Bump Numcodecs requirement to 0.6.1#347
jakirkham wants to merge 21 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.1

Conversation

@jakirkham

Copy link
Copy Markdown
Member

As there are some critical fixes in the latest Numcodecs, this bumps our lower bound to the latest version.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

@alimanfooalimanfoo added this to the v2.3 milestone Nov 30, 2018

@alimanfooalimanfoo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @jakirkham. Probably worth adding a brief release note.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

FTR saw some test failures locally that it looks like CI also encounters. Figured we'd want to look at these and determine how we want to fix them.

@alimanfoo

Copy link
Copy Markdown
Member

Ah, yes, sorry, forgot about that. I expect some test failures due to #324. May be other consequences of the upgrade too. Suggest dealing with them in this PR.

@jakirkham

jakirkham commented Nov 30, 2018

Copy link
Copy Markdown
MemberAuthor

SGTM. How about this one (MsgPack)?

# create an object array using picklez=self.create_array(shape=10, chunks=3, dtype=object, object_codec=Pickle())
z[0] ='foo'>assertz[0] =='foo'

ref: https://travis-ci.org/zarr-developers/zarr/jobs/461883537#L2717-L2720

...or this one (Categorize)?

withpytest.raises(RuntimeError):
# noinspection PyStatementEffect>v[:]
EFailed: DIDNOTRAISE<type'exceptions.RuntimeError'>

ref: https://travis-ci.org/zarr-developers/zarr/jobs/461883537#L2621-L2624

Previously MsgPack was turning bytes objects to unicode objects when
round-tripping them. However this has been fixed in the latest version
of Numcodecs. So correct this test now that MsgPack is working
correctly.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Also there is another failure. Appears Pickle wasn't able to handle general buffer protocol conforming types in decode. Though every other codec worked fine with these (which seems to be a consequence of them using the new utility functions 😉). Apparently Pickle was not using these. Have put together PR ( zarr-developers/numcodecs#143 ) with a fix for this. Would be good to do another patch release to get that out (and anything else these failures might need).

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Based on testing locally, it looks like we are down to the RuntimeError not being raised by an object array after removing the filters.

forcompressorinZlib(1), Blosc():
z=self.create_array(shape=len(data), chunks=30, dtype=object,
object_codec=Categorize(greetings,
dtype=object),
compressor=compressor)
z[:] =datav=z.view(filters=[])
withpytest.raises(RuntimeError):
# noinspection PyStatementEffect>v[:]
EFailed: DIDNOTRAISE<type'exceptions.RuntimeError'>

@alimanfoo

Copy link
Copy Markdown
Member

Hi @jakirkham, took the liberty to push a commit that should resolve the test failure due to RuntimeError not being raised for object arrays.

Other issues should be resolved with the fixes you've merged into numcodecs, so I'll cut a numcodecs 0.6.2 release then come back and update the version requirement here.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Thanks @alimanfoo!

As we already ensured the `chunk` is an `ndarray` viewing the original
data, there is no need for us to do that here as well. Plus the checks
performed by `ensure_contiguous_ndarray` are not needed for our use case
here. Particularly as we have already handled the unusual type cases
above. We also don't need to constrain the buffer size. As such the only
thing we really need is to flatten the array and make it contiguous,
which is what we handle here directly.
As both the expected `object` case and the non-`object` case perform a
`reshape` to flatten the data, go ahead and refactor that out of both
cases and handle it generally. Simplifies the code a bit.
As refactoring of the `reshape` step has effectively dropped the
expected `object` type case, the checks for different types is a little
more complicated than needed. To fix this, basically invert and swap the
case ordering. This way we can handle all generally expected types first
and simply cast them. Then we can raise if an `object` type shows up and
is unexpected.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Did a little bit of refactoring on your change. Hope that is ok. Please let me know your thoughts on that.

The `DictStore` is pretty reliant on the fact that values are immutable
and can be easily compared. For example `__eq__` assumes that all
contents can be compared easily. This works fine if the data is `bytes`.
However it doesn't really work for `ndarray`s for example. Previously we
would have stored whatever the user gave us here. This means comparisons
could falldown in those cases as well (much as the example in the
tutorial has highlighted on CI). Now we effectively require that the
data be something that can either be coerced to `bytes` (e.g. via the
new/old buffer protocol) or is `bytes` to begin with. Make sure not to
force this requirement when nesting one `MutableMapping` within another.
This test case seems to be ill-posed. Anytime we store `object`s to
`Array`s we require an `object_codec` to be specified. Otherwise we have
no clean way to serialize the data. However this `DictStore` test breaks
that assumption by explicitly storing an `object` type in it even though
this would never work for the other stores (particularly when working
with `Array`s). This includes in-memory Zarr `Array`s, which would be
backed by `DictStore`. Given this, we go ahead and drop this test case.
Instead of using a Python `dict` as the `default` store for a Zarr
`Array`, use the `DictStore`. This ensures that all blobs will be
represented as `bytes` regardless of what the user provided as data.
Thus things like comparisons of stores will work well in the default
case.
@jakirkham

jakirkham commented Dec 2, 2018

Copy link
Copy Markdown
MemberAuthor

On a different point, it looks like __eq__ in DictStore is assuming all items in it are comparable. This assumption runs into problems however when anything that is not easily comparable is stored in the DictStore. For instance, ndarray is comparable, but does not reduce to a bool as expected, which causes the tutorial's comparison to fail CI. There are likely other ill-posed situations we can imagine.

Looking at this example before the recent Numcodecs release, it seems the contents of DictStore were always being filled with bytes likely as an indirect consequence of applying filters to the data first. Have added a commit, which enforces this behavior intentionally using ensure_bytes from Numcodecs. Thus making this behavior something we can rely upon for things like __eq__ and such.

However there is a wrinkle with this approach. It appears we had a test that was assuming that objects could be placed in DictStores. Though this behavior doesn't really work for any of the other stores and we usually require an object_codec to be specified to even serialize an object type (specifically with Array). Given this, I've dropped that test and am contending it doesn't conform to our expectations about how stores typically should work with object types.

Also it appears that the Array's default store was a dict, which we cannot apply these fixes to. So have updated the default to DictStore. I'm not sure why this wasn't already the case, which means I'm probably missing something.

Happy to hear thoughts or receive pushback on this approach as I may very well be missing something. If you know of a better approach, would also be interested in hearing about that as well.

Also raised issue ( #348 ) to provide a nice simple example of this problem.

@jakirkham

jakirkham commented Dec 2, 2018

Copy link
Copy Markdown
MemberAuthor

Seems the tutorial test is now hung up on some whitespace issue. Not totally sure how that got introduced.

NVM this was more fallout from the switch to DictStore as the default store for Array.

Edit: This has since been fixed.

As we are now using `DictStore` to back the `Array`, we can correctly
measure how much memory it is using. So update the examples in `info`
and the tutorial to show how much memory is being used. Also update the
store type listed in info as well.
As Numcodecs now includes a very versatile and effective `ensure_bytes`
function, there is no need to define our own in `zarr.storage` as well.
So go ahead and drop it.
As this is no longer being used by `ensure_bytes` as that function was
dropped, go ahead and drop `binary_type` as well.
Make use of Numcodecs' `ensure_ndarray` to take `ndarray` views onto
buffers to be stored in a few cases so as to reshape them and avoid a
copy (thanks to the buffer protocol).
Rewrite `buffer_size` to just use Numcodecs' `ensure_ndarray` to get an
`ndarray` that views the data. Once the `ndarray` is gotten, all that is
needed is to access its `nbytes` member, which returns the number of
bytes that it takes up.
If the data is already a `str` instance, turn `ensure_str` into a no-op.
For all other cases, make use of Numcodecs' `ensure_bytes` to aid
`ensure_str` in coercing data through the buffer protocol. If we are on
Python 3, then decode the `bytes` object to a `str`.
As `DictStore` now must only store `bytes` or types coercible to bytes
via the buffer protocol, there is no possibility for it to have unknown
sizes as `bytes` always have a known size. So drop these cases where the
size can be `-1`.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Simplifies some of the code by making use of Numcodecs utility functions. This avoids a few copies when working with some of the stores. Also generally makes the rest easier to understand.

Make sure that datetime/timedelta arrays are cast to a type that
supports the buffer protocol. Ensure this is a type that can handle all
of the datetime/timedelta values and has the same itemsize.
Instead of using `ensure_ndarray`, use `ensure_contiguous_ndarray` with
the stores. This ensures that datetime/timedeltas are handled by
default. Also catches things like object arrays. Finally this handles
flattening the array if needed.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Replacing with PR ( #352 ), which goes to Numcodecs 0.6.2. It also consolidates the changes from here and drops a few that have been pulled into other PRs. Namely ensuring DictStore contains only bytes for data ( #350 ) and changing the default backend of Array to DictStore ( #351 ).

@jakirkhamjakirkham removed this from the v2.3 milestone Dec 4, 2018
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jakirkham@alimanfoo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Bump Numcodecs requirement to 0.6.1 - #347

Closed
jakirkham wants to merge 21 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.1
Closed

Bump Numcodecs requirement to 0.6.1#347
jakirkham wants to merge 21 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.1

Conversation

@jakirkham

Copy link
Copy Markdown
Member

As there are some critical fixes in the latest Numcodecs, this bumps our lower bound to the latest version.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

@alimanfooalimanfoo added this to the v2.3 milestone Nov 30, 2018

@alimanfooalimanfoo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @jakirkham. Probably worth adding a brief release note.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

FTR saw some test failures locally that it looks like CI also encounters. Figured we'd want to look at these and determine how we want to fix them.

@alimanfoo

Copy link
Copy Markdown
Member

Ah, yes, sorry, forgot about that. I expect some test failures due to #324. May be other consequences of the upgrade too. Suggest dealing with them in this PR.

@jakirkham

jakirkham commented Nov 30, 2018

Copy link
Copy Markdown
MemberAuthor

SGTM. How about this one (MsgPack)?

# create an object array using picklez=self.create_array(shape=10, chunks=3, dtype=object, object_codec=Pickle())
z[0] ='foo'>assertz[0] =='foo'

ref: https://travis-ci.org/zarr-developers/zarr/jobs/461883537#L2717-L2720

...or this one (Categorize)?

withpytest.raises(RuntimeError):
# noinspection PyStatementEffect>v[:]
EFailed: DIDNOTRAISE<type'exceptions.RuntimeError'>

ref: https://travis-ci.org/zarr-developers/zarr/jobs/461883537#L2621-L2624

Previously MsgPack was turning bytes objects to unicode objects when
round-tripping them. However this has been fixed in the latest version
of Numcodecs. So correct this test now that MsgPack is working
correctly.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Also there is another failure. Appears Pickle wasn't able to handle general buffer protocol conforming types in decode. Though every other codec worked fine with these (which seems to be a consequence of them using the new utility functions 😉). Apparently Pickle was not using these. Have put together PR ( zarr-developers/numcodecs#143 ) with a fix for this. Would be good to do another patch release to get that out (and anything else these failures might need).

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Based on testing locally, it looks like we are down to the RuntimeError not being raised by an object array after removing the filters.

forcompressorinZlib(1), Blosc():
z=self.create_array(shape=len(data), chunks=30, dtype=object,
object_codec=Categorize(greetings,
dtype=object),
compressor=compressor)
z[:] =datav=z.view(filters=[])
withpytest.raises(RuntimeError):
# noinspection PyStatementEffect>v[:]
EFailed: DIDNOTRAISE<type'exceptions.RuntimeError'>

@alimanfoo

Copy link
Copy Markdown
Member

Hi @jakirkham, took the liberty to push a commit that should resolve the test failure due to RuntimeError not being raised for object arrays.

Other issues should be resolved with the fixes you've merged into numcodecs, so I'll cut a numcodecs 0.6.2 release then come back and update the version requirement here.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Thanks @alimanfoo!

As we already ensured the `chunk` is an `ndarray` viewing the original
data, there is no need for us to do that here as well. Plus the checks
performed by `ensure_contiguous_ndarray` are not needed for our use case
here. Particularly as we have already handled the unusual type cases
above. We also don't need to constrain the buffer size. As such the only
thing we really need is to flatten the array and make it contiguous,
which is what we handle here directly.
As both the expected `object` case and the non-`object` case perform a
`reshape` to flatten the data, go ahead and refactor that out of both
cases and handle it generally. Simplifies the code a bit.
As refactoring of the `reshape` step has effectively dropped the
expected `object` type case, the checks for different types is a little
more complicated than needed. To fix this, basically invert and swap the
case ordering. This way we can handle all generally expected types first
and simply cast them. Then we can raise if an `object` type shows up and
is unexpected.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Did a little bit of refactoring on your change. Hope that is ok. Please let me know your thoughts on that.

The `DictStore` is pretty reliant on the fact that values are immutable
and can be easily compared. For example `__eq__` assumes that all
contents can be compared easily. This works fine if the data is `bytes`.
However it doesn't really work for `ndarray`s for example. Previously we
would have stored whatever the user gave us here. This means comparisons
could falldown in those cases as well (much as the example in the
tutorial has highlighted on CI). Now we effectively require that the
data be something that can either be coerced to `bytes` (e.g. via the
new/old buffer protocol) or is `bytes` to begin with. Make sure not to
force this requirement when nesting one `MutableMapping` within another.
This test case seems to be ill-posed. Anytime we store `object`s to
`Array`s we require an `object_codec` to be specified. Otherwise we have
no clean way to serialize the data. However this `DictStore` test breaks
that assumption by explicitly storing an `object` type in it even though
this would never work for the other stores (particularly when working
with `Array`s). This includes in-memory Zarr `Array`s, which would be
backed by `DictStore`. Given this, we go ahead and drop this test case.
Instead of using a Python `dict` as the `default` store for a Zarr
`Array`, use the `DictStore`. This ensures that all blobs will be
represented as `bytes` regardless of what the user provided as data.
Thus things like comparisons of stores will work well in the default
case.
@jakirkham

jakirkham commented Dec 2, 2018

Copy link
Copy Markdown
MemberAuthor

On a different point, it looks like __eq__ in DictStore is assuming all items in it are comparable. This assumption runs into problems however when anything that is not easily comparable is stored in the DictStore. For instance, ndarray is comparable, but does not reduce to a bool as expected, which causes the tutorial's comparison to fail CI. There are likely other ill-posed situations we can imagine.

Looking at this example before the recent Numcodecs release, it seems the contents of DictStore were always being filled with bytes likely as an indirect consequence of applying filters to the data first. Have added a commit, which enforces this behavior intentionally using ensure_bytes from Numcodecs. Thus making this behavior something we can rely upon for things like __eq__ and such.

However there is a wrinkle with this approach. It appears we had a test that was assuming that objects could be placed in DictStores. Though this behavior doesn't really work for any of the other stores and we usually require an object_codec to be specified to even serialize an object type (specifically with Array). Given this, I've dropped that test and am contending it doesn't conform to our expectations about how stores typically should work with object types.

Also it appears that the Array's default store was a dict, which we cannot apply these fixes to. So have updated the default to DictStore. I'm not sure why this wasn't already the case, which means I'm probably missing something.

Happy to hear thoughts or receive pushback on this approach as I may very well be missing something. If you know of a better approach, would also be interested in hearing about that as well.

Also raised issue ( #348 ) to provide a nice simple example of this problem.

@jakirkham

jakirkham commented Dec 2, 2018

Copy link
Copy Markdown
MemberAuthor

Seems the tutorial test is now hung up on some whitespace issue. Not totally sure how that got introduced.

NVM this was more fallout from the switch to DictStore as the default store for Array.

Edit: This has since been fixed.

As we are now using `DictStore` to back the `Array`, we can correctly
measure how much memory it is using. So update the examples in `info`
and the tutorial to show how much memory is being used. Also update the
store type listed in info as well.
As Numcodecs now includes a very versatile and effective `ensure_bytes`
function, there is no need to define our own in `zarr.storage` as well.
So go ahead and drop it.
As this is no longer being used by `ensure_bytes` as that function was
dropped, go ahead and drop `binary_type` as well.
Make use of Numcodecs' `ensure_ndarray` to take `ndarray` views onto
buffers to be stored in a few cases so as to reshape them and avoid a
copy (thanks to the buffer protocol).
Rewrite `buffer_size` to just use Numcodecs' `ensure_ndarray` to get an
`ndarray` that views the data. Once the `ndarray` is gotten, all that is
needed is to access its `nbytes` member, which returns the number of
bytes that it takes up.
If the data is already a `str` instance, turn `ensure_str` into a no-op.
For all other cases, make use of Numcodecs' `ensure_bytes` to aid
`ensure_str` in coercing data through the buffer protocol. If we are on
Python 3, then decode the `bytes` object to a `str`.
As `DictStore` now must only store `bytes` or types coercible to bytes
via the buffer protocol, there is no possibility for it to have unknown
sizes as `bytes` always have a known size. So drop these cases where the
size can be `-1`.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Simplifies some of the code by making use of Numcodecs utility functions. This avoids a few copies when working with some of the stores. Also generally makes the rest easier to understand.

Make sure that datetime/timedelta arrays are cast to a type that
supports the buffer protocol. Ensure this is a type that can handle all
of the datetime/timedelta values and has the same itemsize.
Instead of using `ensure_ndarray`, use `ensure_contiguous_ndarray` with
the stores. This ensures that datetime/timedeltas are handled by
default. Also catches things like object arrays. Finally this handles
flattening the array if needed.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Replacing with PR ( #352 ), which goes to Numcodecs 0.6.2. It also consolidates the changes from here and drops a few that have been pulled into other PRs. Namely ensuring DictStore contains only bytes for data ( #350 ) and changing the default backend of Array to DictStore ( #351 ).

@jakirkhamjakirkham removed this from the v2.3 milestone Dec 4, 2018
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jakirkham@alimanfoo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Bump Numcodecs requirement to 0.6.1 - #347

Closed
jakirkham wants to merge 21 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.1
Closed

Bump Numcodecs requirement to 0.6.1#347
jakirkham wants to merge 21 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.1

Conversation

@jakirkham

Copy link
Copy Markdown
Member

As there are some critical fixes in the latest Numcodecs, this bumps our lower bound to the latest version.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

@alimanfooalimanfoo added this to the v2.3 milestone Nov 30, 2018

@alimanfooalimanfoo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @jakirkham. Probably worth adding a brief release note.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

FTR saw some test failures locally that it looks like CI also encounters. Figured we'd want to look at these and determine how we want to fix them.

@alimanfoo

Copy link
Copy Markdown
Member

Ah, yes, sorry, forgot about that. I expect some test failures due to #324. May be other consequences of the upgrade too. Suggest dealing with them in this PR.

@jakirkham

jakirkham commented Nov 30, 2018

Copy link
Copy Markdown
MemberAuthor

SGTM. How about this one (MsgPack)?

# create an object array using picklez=self.create_array(shape=10, chunks=3, dtype=object, object_codec=Pickle())
z[0] ='foo'>assertz[0] =='foo'

ref: https://travis-ci.org/zarr-developers/zarr/jobs/461883537#L2717-L2720

...or this one (Categorize)?

withpytest.raises(RuntimeError):
# noinspection PyStatementEffect>v[:]
EFailed: DIDNOTRAISE<type'exceptions.RuntimeError'>

ref: https://travis-ci.org/zarr-developers/zarr/jobs/461883537#L2621-L2624

Previously MsgPack was turning bytes objects to unicode objects when
round-tripping them. However this has been fixed in the latest version
of Numcodecs. So correct this test now that MsgPack is working
correctly.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Also there is another failure. Appears Pickle wasn't able to handle general buffer protocol conforming types in decode. Though every other codec worked fine with these (which seems to be a consequence of them using the new utility functions 😉). Apparently Pickle was not using these. Have put together PR ( zarr-developers/numcodecs#143 ) with a fix for this. Would be good to do another patch release to get that out (and anything else these failures might need).

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Based on testing locally, it looks like we are down to the RuntimeError not being raised by an object array after removing the filters.

forcompressorinZlib(1), Blosc():
z=self.create_array(shape=len(data), chunks=30, dtype=object,
object_codec=Categorize(greetings,
dtype=object),
compressor=compressor)
z[:] =datav=z.view(filters=[])
withpytest.raises(RuntimeError):
# noinspection PyStatementEffect>v[:]
EFailed: DIDNOTRAISE<type'exceptions.RuntimeError'>

@alimanfoo

Copy link
Copy Markdown
Member

Hi @jakirkham, took the liberty to push a commit that should resolve the test failure due to RuntimeError not being raised for object arrays.

Other issues should be resolved with the fixes you've merged into numcodecs, so I'll cut a numcodecs 0.6.2 release then come back and update the version requirement here.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Thanks @alimanfoo!

As we already ensured the `chunk` is an `ndarray` viewing the original
data, there is no need for us to do that here as well. Plus the checks
performed by `ensure_contiguous_ndarray` are not needed for our use case
here. Particularly as we have already handled the unusual type cases
above. We also don't need to constrain the buffer size. As such the only
thing we really need is to flatten the array and make it contiguous,
which is what we handle here directly.
As both the expected `object` case and the non-`object` case perform a
`reshape` to flatten the data, go ahead and refactor that out of both
cases and handle it generally. Simplifies the code a bit.
As refactoring of the `reshape` step has effectively dropped the
expected `object` type case, the checks for different types is a little
more complicated than needed. To fix this, basically invert and swap the
case ordering. This way we can handle all generally expected types first
and simply cast them. Then we can raise if an `object` type shows up and
is unexpected.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Did a little bit of refactoring on your change. Hope that is ok. Please let me know your thoughts on that.

The `DictStore` is pretty reliant on the fact that values are immutable
and can be easily compared. For example `__eq__` assumes that all
contents can be compared easily. This works fine if the data is `bytes`.
However it doesn't really work for `ndarray`s for example. Previously we
would have stored whatever the user gave us here. This means comparisons
could falldown in those cases as well (much as the example in the
tutorial has highlighted on CI). Now we effectively require that the
data be something that can either be coerced to `bytes` (e.g. via the
new/old buffer protocol) or is `bytes` to begin with. Make sure not to
force this requirement when nesting one `MutableMapping` within another.
This test case seems to be ill-posed. Anytime we store `object`s to
`Array`s we require an `object_codec` to be specified. Otherwise we have
no clean way to serialize the data. However this `DictStore` test breaks
that assumption by explicitly storing an `object` type in it even though
this would never work for the other stores (particularly when working
with `Array`s). This includes in-memory Zarr `Array`s, which would be
backed by `DictStore`. Given this, we go ahead and drop this test case.
Instead of using a Python `dict` as the `default` store for a Zarr
`Array`, use the `DictStore`. This ensures that all blobs will be
represented as `bytes` regardless of what the user provided as data.
Thus things like comparisons of stores will work well in the default
case.
@jakirkham

jakirkham commented Dec 2, 2018

Copy link
Copy Markdown
MemberAuthor

On a different point, it looks like __eq__ in DictStore is assuming all items in it are comparable. This assumption runs into problems however when anything that is not easily comparable is stored in the DictStore. For instance, ndarray is comparable, but does not reduce to a bool as expected, which causes the tutorial's comparison to fail CI. There are likely other ill-posed situations we can imagine.

Looking at this example before the recent Numcodecs release, it seems the contents of DictStore were always being filled with bytes likely as an indirect consequence of applying filters to the data first. Have added a commit, which enforces this behavior intentionally using ensure_bytes from Numcodecs. Thus making this behavior something we can rely upon for things like __eq__ and such.

However there is a wrinkle with this approach. It appears we had a test that was assuming that objects could be placed in DictStores. Though this behavior doesn't really work for any of the other stores and we usually require an object_codec to be specified to even serialize an object type (specifically with Array). Given this, I've dropped that test and am contending it doesn't conform to our expectations about how stores typically should work with object types.

Also it appears that the Array's default store was a dict, which we cannot apply these fixes to. So have updated the default to DictStore. I'm not sure why this wasn't already the case, which means I'm probably missing something.

Happy to hear thoughts or receive pushback on this approach as I may very well be missing something. If you know of a better approach, would also be interested in hearing about that as well.

Also raised issue ( #348 ) to provide a nice simple example of this problem.

@jakirkham

jakirkham commented Dec 2, 2018

Copy link
Copy Markdown
MemberAuthor

Seems the tutorial test is now hung up on some whitespace issue. Not totally sure how that got introduced.

NVM this was more fallout from the switch to DictStore as the default store for Array.

Edit: This has since been fixed.

As we are now using `DictStore` to back the `Array`, we can correctly
measure how much memory it is using. So update the examples in `info`
and the tutorial to show how much memory is being used. Also update the
store type listed in info as well.
As Numcodecs now includes a very versatile and effective `ensure_bytes`
function, there is no need to define our own in `zarr.storage` as well.
So go ahead and drop it.
As this is no longer being used by `ensure_bytes` as that function was
dropped, go ahead and drop `binary_type` as well.
Make use of Numcodecs' `ensure_ndarray` to take `ndarray` views onto
buffers to be stored in a few cases so as to reshape them and avoid a
copy (thanks to the buffer protocol).
Rewrite `buffer_size` to just use Numcodecs' `ensure_ndarray` to get an
`ndarray` that views the data. Once the `ndarray` is gotten, all that is
needed is to access its `nbytes` member, which returns the number of
bytes that it takes up.
If the data is already a `str` instance, turn `ensure_str` into a no-op.
For all other cases, make use of Numcodecs' `ensure_bytes` to aid
`ensure_str` in coercing data through the buffer protocol. If we are on
Python 3, then decode the `bytes` object to a `str`.
As `DictStore` now must only store `bytes` or types coercible to bytes
via the buffer protocol, there is no possibility for it to have unknown
sizes as `bytes` always have a known size. So drop these cases where the
size can be `-1`.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Simplifies some of the code by making use of Numcodecs utility functions. This avoids a few copies when working with some of the stores. Also generally makes the rest easier to understand.

Make sure that datetime/timedelta arrays are cast to a type that
supports the buffer protocol. Ensure this is a type that can handle all
of the datetime/timedelta values and has the same itemsize.
Instead of using `ensure_ndarray`, use `ensure_contiguous_ndarray` with
the stores. This ensures that datetime/timedeltas are handled by
default. Also catches things like object arrays. Finally this handles
flattening the array if needed.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Replacing with PR ( #352 ), which goes to Numcodecs 0.6.2. It also consolidates the changes from here and drops a few that have been pulled into other PRs. Namely ensuring DictStore contains only bytes for data ( #350 ) and changing the default backend of Array to DictStore ( #351 ).

@jakirkhamjakirkham removed this from the v2.3 milestone Dec 4, 2018
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jakirkham@alimanfoo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Bump Numcodecs requirement to 0.6.1 - #347

Closed
jakirkham wants to merge 21 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.1
Closed

Bump Numcodecs requirement to 0.6.1#347
jakirkham wants to merge 21 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.1

Conversation

@jakirkham

Copy link
Copy Markdown
Member

As there are some critical fixes in the latest Numcodecs, this bumps our lower bound to the latest version.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

@alimanfooalimanfoo added this to the v2.3 milestone Nov 30, 2018

@alimanfooalimanfoo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @jakirkham. Probably worth adding a brief release note.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

FTR saw some test failures locally that it looks like CI also encounters. Figured we'd want to look at these and determine how we want to fix them.

@alimanfoo

Copy link
Copy Markdown
Member

Ah, yes, sorry, forgot about that. I expect some test failures due to #324. May be other consequences of the upgrade too. Suggest dealing with them in this PR.

@jakirkham

jakirkham commented Nov 30, 2018

Copy link
Copy Markdown
MemberAuthor

SGTM. How about this one (MsgPack)?

# create an object array using picklez=self.create_array(shape=10, chunks=3, dtype=object, object_codec=Pickle())
z[0] ='foo'>assertz[0] =='foo'

ref: https://travis-ci.org/zarr-developers/zarr/jobs/461883537#L2717-L2720

...or this one (Categorize)?

withpytest.raises(RuntimeError):
# noinspection PyStatementEffect>v[:]
EFailed: DIDNOTRAISE<type'exceptions.RuntimeError'>

ref: https://travis-ci.org/zarr-developers/zarr/jobs/461883537#L2621-L2624

Previously MsgPack was turning bytes objects to unicode objects when
round-tripping them. However this has been fixed in the latest version
of Numcodecs. So correct this test now that MsgPack is working
correctly.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Also there is another failure. Appears Pickle wasn't able to handle general buffer protocol conforming types in decode. Though every other codec worked fine with these (which seems to be a consequence of them using the new utility functions 😉). Apparently Pickle was not using these. Have put together PR ( zarr-developers/numcodecs#143 ) with a fix for this. Would be good to do another patch release to get that out (and anything else these failures might need).

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Based on testing locally, it looks like we are down to the RuntimeError not being raised by an object array after removing the filters.

forcompressorinZlib(1), Blosc():
z=self.create_array(shape=len(data), chunks=30, dtype=object,
object_codec=Categorize(greetings,
dtype=object),
compressor=compressor)
z[:] =datav=z.view(filters=[])
withpytest.raises(RuntimeError):
# noinspection PyStatementEffect>v[:]
EFailed: DIDNOTRAISE<type'exceptions.RuntimeError'>

@alimanfoo

Copy link
Copy Markdown
Member

Hi @jakirkham, took the liberty to push a commit that should resolve the test failure due to RuntimeError not being raised for object arrays.

Other issues should be resolved with the fixes you've merged into numcodecs, so I'll cut a numcodecs 0.6.2 release then come back and update the version requirement here.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Thanks @alimanfoo!

As we already ensured the `chunk` is an `ndarray` viewing the original
data, there is no need for us to do that here as well. Plus the checks
performed by `ensure_contiguous_ndarray` are not needed for our use case
here. Particularly as we have already handled the unusual type cases
above. We also don't need to constrain the buffer size. As such the only
thing we really need is to flatten the array and make it contiguous,
which is what we handle here directly.
As both the expected `object` case and the non-`object` case perform a
`reshape` to flatten the data, go ahead and refactor that out of both
cases and handle it generally. Simplifies the code a bit.
As refactoring of the `reshape` step has effectively dropped the
expected `object` type case, the checks for different types is a little
more complicated than needed. To fix this, basically invert and swap the
case ordering. This way we can handle all generally expected types first
and simply cast them. Then we can raise if an `object` type shows up and
is unexpected.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Did a little bit of refactoring on your change. Hope that is ok. Please let me know your thoughts on that.

The `DictStore` is pretty reliant on the fact that values are immutable
and can be easily compared. For example `__eq__` assumes that all
contents can be compared easily. This works fine if the data is `bytes`.
However it doesn't really work for `ndarray`s for example. Previously we
would have stored whatever the user gave us here. This means comparisons
could falldown in those cases as well (much as the example in the
tutorial has highlighted on CI). Now we effectively require that the
data be something that can either be coerced to `bytes` (e.g. via the
new/old buffer protocol) or is `bytes` to begin with. Make sure not to
force this requirement when nesting one `MutableMapping` within another.
This test case seems to be ill-posed. Anytime we store `object`s to
`Array`s we require an `object_codec` to be specified. Otherwise we have
no clean way to serialize the data. However this `DictStore` test breaks
that assumption by explicitly storing an `object` type in it even though
this would never work for the other stores (particularly when working
with `Array`s). This includes in-memory Zarr `Array`s, which would be
backed by `DictStore`. Given this, we go ahead and drop this test case.
Instead of using a Python `dict` as the `default` store for a Zarr
`Array`, use the `DictStore`. This ensures that all blobs will be
represented as `bytes` regardless of what the user provided as data.
Thus things like comparisons of stores will work well in the default
case.
@jakirkham

jakirkham commented Dec 2, 2018

Copy link
Copy Markdown
MemberAuthor

On a different point, it looks like __eq__ in DictStore is assuming all items in it are comparable. This assumption runs into problems however when anything that is not easily comparable is stored in the DictStore. For instance, ndarray is comparable, but does not reduce to a bool as expected, which causes the tutorial's comparison to fail CI. There are likely other ill-posed situations we can imagine.

Looking at this example before the recent Numcodecs release, it seems the contents of DictStore were always being filled with bytes likely as an indirect consequence of applying filters to the data first. Have added a commit, which enforces this behavior intentionally using ensure_bytes from Numcodecs. Thus making this behavior something we can rely upon for things like __eq__ and such.

However there is a wrinkle with this approach. It appears we had a test that was assuming that objects could be placed in DictStores. Though this behavior doesn't really work for any of the other stores and we usually require an object_codec to be specified to even serialize an object type (specifically with Array). Given this, I've dropped that test and am contending it doesn't conform to our expectations about how stores typically should work with object types.

Also it appears that the Array's default store was a dict, which we cannot apply these fixes to. So have updated the default to DictStore. I'm not sure why this wasn't already the case, which means I'm probably missing something.

Happy to hear thoughts or receive pushback on this approach as I may very well be missing something. If you know of a better approach, would also be interested in hearing about that as well.

Also raised issue ( #348 ) to provide a nice simple example of this problem.

@jakirkham

jakirkham commented Dec 2, 2018

Copy link
Copy Markdown
MemberAuthor

Seems the tutorial test is now hung up on some whitespace issue. Not totally sure how that got introduced.

NVM this was more fallout from the switch to DictStore as the default store for Array.

Edit: This has since been fixed.

As we are now using `DictStore` to back the `Array`, we can correctly
measure how much memory it is using. So update the examples in `info`
and the tutorial to show how much memory is being used. Also update the
store type listed in info as well.
As Numcodecs now includes a very versatile and effective `ensure_bytes`
function, there is no need to define our own in `zarr.storage` as well.
So go ahead and drop it.
As this is no longer being used by `ensure_bytes` as that function was
dropped, go ahead and drop `binary_type` as well.
Make use of Numcodecs' `ensure_ndarray` to take `ndarray` views onto
buffers to be stored in a few cases so as to reshape them and avoid a
copy (thanks to the buffer protocol).
Rewrite `buffer_size` to just use Numcodecs' `ensure_ndarray` to get an
`ndarray` that views the data. Once the `ndarray` is gotten, all that is
needed is to access its `nbytes` member, which returns the number of
bytes that it takes up.
If the data is already a `str` instance, turn `ensure_str` into a no-op.
For all other cases, make use of Numcodecs' `ensure_bytes` to aid
`ensure_str` in coercing data through the buffer protocol. If we are on
Python 3, then decode the `bytes` object to a `str`.
As `DictStore` now must only store `bytes` or types coercible to bytes
via the buffer protocol, there is no possibility for it to have unknown
sizes as `bytes` always have a known size. So drop these cases where the
size can be `-1`.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Simplifies some of the code by making use of Numcodecs utility functions. This avoids a few copies when working with some of the stores. Also generally makes the rest easier to understand.

Make sure that datetime/timedelta arrays are cast to a type that
supports the buffer protocol. Ensure this is a type that can handle all
of the datetime/timedelta values and has the same itemsize.
Instead of using `ensure_ndarray`, use `ensure_contiguous_ndarray` with
the stores. This ensures that datetime/timedeltas are handled by
default. Also catches things like object arrays. Finally this handles
flattening the array if needed.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Replacing with PR ( #352 ), which goes to Numcodecs 0.6.2. It also consolidates the changes from here and drops a few that have been pulled into other PRs. Namely ensuring DictStore contains only bytes for data ( #350 ) and changing the default backend of Array to DictStore ( #351 ).

@jakirkhamjakirkham removed this from the v2.3 milestone Dec 4, 2018
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jakirkham@alimanfoo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Bump Numcodecs requirement to 0.6.1 - #347

Closed
jakirkham wants to merge 21 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.1
Closed

Bump Numcodecs requirement to 0.6.1#347
jakirkham wants to merge 21 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.1

Conversation

@jakirkham

Copy link
Copy Markdown
Member

As there are some critical fixes in the latest Numcodecs, this bumps our lower bound to the latest version.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

@alimanfooalimanfoo added this to the v2.3 milestone Nov 30, 2018

@alimanfooalimanfoo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @jakirkham. Probably worth adding a brief release note.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

FTR saw some test failures locally that it looks like CI also encounters. Figured we'd want to look at these and determine how we want to fix them.

@alimanfoo

Copy link
Copy Markdown
Member

Ah, yes, sorry, forgot about that. I expect some test failures due to #324. May be other consequences of the upgrade too. Suggest dealing with them in this PR.

@jakirkham

jakirkham commented Nov 30, 2018

Copy link
Copy Markdown
MemberAuthor

SGTM. How about this one (MsgPack)?

# create an object array using picklez=self.create_array(shape=10, chunks=3, dtype=object, object_codec=Pickle())
z[0] ='foo'>assertz[0] =='foo'

ref: https://travis-ci.org/zarr-developers/zarr/jobs/461883537#L2717-L2720

...or this one (Categorize)?

withpytest.raises(RuntimeError):
# noinspection PyStatementEffect>v[:]
EFailed: DIDNOTRAISE<type'exceptions.RuntimeError'>

ref: https://travis-ci.org/zarr-developers/zarr/jobs/461883537#L2621-L2624

Previously MsgPack was turning bytes objects to unicode objects when
round-tripping them. However this has been fixed in the latest version
of Numcodecs. So correct this test now that MsgPack is working
correctly.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Also there is another failure. Appears Pickle wasn't able to handle general buffer protocol conforming types in decode. Though every other codec worked fine with these (which seems to be a consequence of them using the new utility functions 😉). Apparently Pickle was not using these. Have put together PR ( zarr-developers/numcodecs#143 ) with a fix for this. Would be good to do another patch release to get that out (and anything else these failures might need).

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Based on testing locally, it looks like we are down to the RuntimeError not being raised by an object array after removing the filters.

forcompressorinZlib(1), Blosc():
z=self.create_array(shape=len(data), chunks=30, dtype=object,
object_codec=Categorize(greetings,
dtype=object),
compressor=compressor)
z[:] =datav=z.view(filters=[])
withpytest.raises(RuntimeError):
# noinspection PyStatementEffect>v[:]
EFailed: DIDNOTRAISE<type'exceptions.RuntimeError'>

@alimanfoo

Copy link
Copy Markdown
Member

Hi @jakirkham, took the liberty to push a commit that should resolve the test failure due to RuntimeError not being raised for object arrays.

Other issues should be resolved with the fixes you've merged into numcodecs, so I'll cut a numcodecs 0.6.2 release then come back and update the version requirement here.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Thanks @alimanfoo!

As we already ensured the `chunk` is an `ndarray` viewing the original
data, there is no need for us to do that here as well. Plus the checks
performed by `ensure_contiguous_ndarray` are not needed for our use case
here. Particularly as we have already handled the unusual type cases
above. We also don't need to constrain the buffer size. As such the only
thing we really need is to flatten the array and make it contiguous,
which is what we handle here directly.
As both the expected `object` case and the non-`object` case perform a
`reshape` to flatten the data, go ahead and refactor that out of both
cases and handle it generally. Simplifies the code a bit.
As refactoring of the `reshape` step has effectively dropped the
expected `object` type case, the checks for different types is a little
more complicated than needed. To fix this, basically invert and swap the
case ordering. This way we can handle all generally expected types first
and simply cast them. Then we can raise if an `object` type shows up and
is unexpected.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Did a little bit of refactoring on your change. Hope that is ok. Please let me know your thoughts on that.

The `DictStore` is pretty reliant on the fact that values are immutable
and can be easily compared. For example `__eq__` assumes that all
contents can be compared easily. This works fine if the data is `bytes`.
However it doesn't really work for `ndarray`s for example. Previously we
would have stored whatever the user gave us here. This means comparisons
could falldown in those cases as well (much as the example in the
tutorial has highlighted on CI). Now we effectively require that the
data be something that can either be coerced to `bytes` (e.g. via the
new/old buffer protocol) or is `bytes` to begin with. Make sure not to
force this requirement when nesting one `MutableMapping` within another.
This test case seems to be ill-posed. Anytime we store `object`s to
`Array`s we require an `object_codec` to be specified. Otherwise we have
no clean way to serialize the data. However this `DictStore` test breaks
that assumption by explicitly storing an `object` type in it even though
this would never work for the other stores (particularly when working
with `Array`s). This includes in-memory Zarr `Array`s, which would be
backed by `DictStore`. Given this, we go ahead and drop this test case.
Instead of using a Python `dict` as the `default` store for a Zarr
`Array`, use the `DictStore`. This ensures that all blobs will be
represented as `bytes` regardless of what the user provided as data.
Thus things like comparisons of stores will work well in the default
case.
@jakirkham

jakirkham commented Dec 2, 2018

Copy link
Copy Markdown
MemberAuthor

On a different point, it looks like __eq__ in DictStore is assuming all items in it are comparable. This assumption runs into problems however when anything that is not easily comparable is stored in the DictStore. For instance, ndarray is comparable, but does not reduce to a bool as expected, which causes the tutorial's comparison to fail CI. There are likely other ill-posed situations we can imagine.

Looking at this example before the recent Numcodecs release, it seems the contents of DictStore were always being filled with bytes likely as an indirect consequence of applying filters to the data first. Have added a commit, which enforces this behavior intentionally using ensure_bytes from Numcodecs. Thus making this behavior something we can rely upon for things like __eq__ and such.

However there is a wrinkle with this approach. It appears we had a test that was assuming that objects could be placed in DictStores. Though this behavior doesn't really work for any of the other stores and we usually require an object_codec to be specified to even serialize an object type (specifically with Array). Given this, I've dropped that test and am contending it doesn't conform to our expectations about how stores typically should work with object types.

Also it appears that the Array's default store was a dict, which we cannot apply these fixes to. So have updated the default to DictStore. I'm not sure why this wasn't already the case, which means I'm probably missing something.

Happy to hear thoughts or receive pushback on this approach as I may very well be missing something. If you know of a better approach, would also be interested in hearing about that as well.

Also raised issue ( #348 ) to provide a nice simple example of this problem.

@jakirkham

jakirkham commented Dec 2, 2018

Copy link
Copy Markdown
MemberAuthor

Seems the tutorial test is now hung up on some whitespace issue. Not totally sure how that got introduced.

NVM this was more fallout from the switch to DictStore as the default store for Array.

Edit: This has since been fixed.

As we are now using `DictStore` to back the `Array`, we can correctly
measure how much memory it is using. So update the examples in `info`
and the tutorial to show how much memory is being used. Also update the
store type listed in info as well.
As Numcodecs now includes a very versatile and effective `ensure_bytes`
function, there is no need to define our own in `zarr.storage` as well.
So go ahead and drop it.
As this is no longer being used by `ensure_bytes` as that function was
dropped, go ahead and drop `binary_type` as well.
Make use of Numcodecs' `ensure_ndarray` to take `ndarray` views onto
buffers to be stored in a few cases so as to reshape them and avoid a
copy (thanks to the buffer protocol).
Rewrite `buffer_size` to just use Numcodecs' `ensure_ndarray` to get an
`ndarray` that views the data. Once the `ndarray` is gotten, all that is
needed is to access its `nbytes` member, which returns the number of
bytes that it takes up.
If the data is already a `str` instance, turn `ensure_str` into a no-op.
For all other cases, make use of Numcodecs' `ensure_bytes` to aid
`ensure_str` in coercing data through the buffer protocol. If we are on
Python 3, then decode the `bytes` object to a `str`.
As `DictStore` now must only store `bytes` or types coercible to bytes
via the buffer protocol, there is no possibility for it to have unknown
sizes as `bytes` always have a known size. So drop these cases where the
size can be `-1`.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Simplifies some of the code by making use of Numcodecs utility functions. This avoids a few copies when working with some of the stores. Also generally makes the rest easier to understand.

Make sure that datetime/timedelta arrays are cast to a type that
supports the buffer protocol. Ensure this is a type that can handle all
of the datetime/timedelta values and has the same itemsize.
Instead of using `ensure_ndarray`, use `ensure_contiguous_ndarray` with
the stores. This ensures that datetime/timedeltas are handled by
default. Also catches things like object arrays. Finally this handles
flattening the array if needed.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Replacing with PR ( #352 ), which goes to Numcodecs 0.6.2. It also consolidates the changes from here and drops a few that have been pulled into other PRs. Namely ensuring DictStore contains only bytes for data ( #350 ) and changing the default backend of Array to DictStore ( #351 ).

@jakirkhamjakirkham removed this from the v2.3 milestone Dec 4, 2018
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jakirkham@alimanfoo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Bump Numcodecs requirement to 0.6.1 - #347

Closed
jakirkham wants to merge 21 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.1
Closed

Bump Numcodecs requirement to 0.6.1#347
jakirkham wants to merge 21 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.1

Conversation

@jakirkham

Copy link
Copy Markdown
Member

As there are some critical fixes in the latest Numcodecs, this bumps our lower bound to the latest version.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

@alimanfooalimanfoo added this to the v2.3 milestone Nov 30, 2018

@alimanfooalimanfoo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @jakirkham. Probably worth adding a brief release note.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

FTR saw some test failures locally that it looks like CI also encounters. Figured we'd want to look at these and determine how we want to fix them.

@alimanfoo

Copy link
Copy Markdown
Member

Ah, yes, sorry, forgot about that. I expect some test failures due to #324. May be other consequences of the upgrade too. Suggest dealing with them in this PR.

@jakirkham

jakirkham commented Nov 30, 2018

Copy link
Copy Markdown
MemberAuthor

SGTM. How about this one (MsgPack)?

# create an object array using picklez=self.create_array(shape=10, chunks=3, dtype=object, object_codec=Pickle())
z[0] ='foo'>assertz[0] =='foo'

ref: https://travis-ci.org/zarr-developers/zarr/jobs/461883537#L2717-L2720

...or this one (Categorize)?

withpytest.raises(RuntimeError):
# noinspection PyStatementEffect>v[:]
EFailed: DIDNOTRAISE<type'exceptions.RuntimeError'>

ref: https://travis-ci.org/zarr-developers/zarr/jobs/461883537#L2621-L2624

Previously MsgPack was turning bytes objects to unicode objects when
round-tripping them. However this has been fixed in the latest version
of Numcodecs. So correct this test now that MsgPack is working
correctly.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Also there is another failure. Appears Pickle wasn't able to handle general buffer protocol conforming types in decode. Though every other codec worked fine with these (which seems to be a consequence of them using the new utility functions 😉). Apparently Pickle was not using these. Have put together PR ( zarr-developers/numcodecs#143 ) with a fix for this. Would be good to do another patch release to get that out (and anything else these failures might need).

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Based on testing locally, it looks like we are down to the RuntimeError not being raised by an object array after removing the filters.

forcompressorinZlib(1), Blosc():
z=self.create_array(shape=len(data), chunks=30, dtype=object,
object_codec=Categorize(greetings,
dtype=object),
compressor=compressor)
z[:] =datav=z.view(filters=[])
withpytest.raises(RuntimeError):
# noinspection PyStatementEffect>v[:]
EFailed: DIDNOTRAISE<type'exceptions.RuntimeError'>

@alimanfoo

Copy link
Copy Markdown
Member

Hi @jakirkham, took the liberty to push a commit that should resolve the test failure due to RuntimeError not being raised for object arrays.

Other issues should be resolved with the fixes you've merged into numcodecs, so I'll cut a numcodecs 0.6.2 release then come back and update the version requirement here.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Thanks @alimanfoo!

As we already ensured the `chunk` is an `ndarray` viewing the original
data, there is no need for us to do that here as well. Plus the checks
performed by `ensure_contiguous_ndarray` are not needed for our use case
here. Particularly as we have already handled the unusual type cases
above. We also don't need to constrain the buffer size. As such the only
thing we really need is to flatten the array and make it contiguous,
which is what we handle here directly.
As both the expected `object` case and the non-`object` case perform a
`reshape` to flatten the data, go ahead and refactor that out of both
cases and handle it generally. Simplifies the code a bit.
As refactoring of the `reshape` step has effectively dropped the
expected `object` type case, the checks for different types is a little
more complicated than needed. To fix this, basically invert and swap the
case ordering. This way we can handle all generally expected types first
and simply cast them. Then we can raise if an `object` type shows up and
is unexpected.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Did a little bit of refactoring on your change. Hope that is ok. Please let me know your thoughts on that.

The `DictStore` is pretty reliant on the fact that values are immutable
and can be easily compared. For example `__eq__` assumes that all
contents can be compared easily. This works fine if the data is `bytes`.
However it doesn't really work for `ndarray`s for example. Previously we
would have stored whatever the user gave us here. This means comparisons
could falldown in those cases as well (much as the example in the
tutorial has highlighted on CI). Now we effectively require that the
data be something that can either be coerced to `bytes` (e.g. via the
new/old buffer protocol) or is `bytes` to begin with. Make sure not to
force this requirement when nesting one `MutableMapping` within another.
This test case seems to be ill-posed. Anytime we store `object`s to
`Array`s we require an `object_codec` to be specified. Otherwise we have
no clean way to serialize the data. However this `DictStore` test breaks
that assumption by explicitly storing an `object` type in it even though
this would never work for the other stores (particularly when working
with `Array`s). This includes in-memory Zarr `Array`s, which would be
backed by `DictStore`. Given this, we go ahead and drop this test case.
Instead of using a Python `dict` as the `default` store for a Zarr
`Array`, use the `DictStore`. This ensures that all blobs will be
represented as `bytes` regardless of what the user provided as data.
Thus things like comparisons of stores will work well in the default
case.
@jakirkham

jakirkham commented Dec 2, 2018

Copy link
Copy Markdown
MemberAuthor

On a different point, it looks like __eq__ in DictStore is assuming all items in it are comparable. This assumption runs into problems however when anything that is not easily comparable is stored in the DictStore. For instance, ndarray is comparable, but does not reduce to a bool as expected, which causes the tutorial's comparison to fail CI. There are likely other ill-posed situations we can imagine.

Looking at this example before the recent Numcodecs release, it seems the contents of DictStore were always being filled with bytes likely as an indirect consequence of applying filters to the data first. Have added a commit, which enforces this behavior intentionally using ensure_bytes from Numcodecs. Thus making this behavior something we can rely upon for things like __eq__ and such.

However there is a wrinkle with this approach. It appears we had a test that was assuming that objects could be placed in DictStores. Though this behavior doesn't really work for any of the other stores and we usually require an object_codec to be specified to even serialize an object type (specifically with Array). Given this, I've dropped that test and am contending it doesn't conform to our expectations about how stores typically should work with object types.

Also it appears that the Array's default store was a dict, which we cannot apply these fixes to. So have updated the default to DictStore. I'm not sure why this wasn't already the case, which means I'm probably missing something.

Happy to hear thoughts or receive pushback on this approach as I may very well be missing something. If you know of a better approach, would also be interested in hearing about that as well.

Also raised issue ( #348 ) to provide a nice simple example of this problem.

@jakirkham

jakirkham commented Dec 2, 2018

Copy link
Copy Markdown
MemberAuthor

Seems the tutorial test is now hung up on some whitespace issue. Not totally sure how that got introduced.

NVM this was more fallout from the switch to DictStore as the default store for Array.

Edit: This has since been fixed.

As we are now using `DictStore` to back the `Array`, we can correctly
measure how much memory it is using. So update the examples in `info`
and the tutorial to show how much memory is being used. Also update the
store type listed in info as well.
As Numcodecs now includes a very versatile and effective `ensure_bytes`
function, there is no need to define our own in `zarr.storage` as well.
So go ahead and drop it.
As this is no longer being used by `ensure_bytes` as that function was
dropped, go ahead and drop `binary_type` as well.
Make use of Numcodecs' `ensure_ndarray` to take `ndarray` views onto
buffers to be stored in a few cases so as to reshape them and avoid a
copy (thanks to the buffer protocol).
Rewrite `buffer_size` to just use Numcodecs' `ensure_ndarray` to get an
`ndarray` that views the data. Once the `ndarray` is gotten, all that is
needed is to access its `nbytes` member, which returns the number of
bytes that it takes up.
If the data is already a `str` instance, turn `ensure_str` into a no-op.
For all other cases, make use of Numcodecs' `ensure_bytes` to aid
`ensure_str` in coercing data through the buffer protocol. If we are on
Python 3, then decode the `bytes` object to a `str`.
As `DictStore` now must only store `bytes` or types coercible to bytes
via the buffer protocol, there is no possibility for it to have unknown
sizes as `bytes` always have a known size. So drop these cases where the
size can be `-1`.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Simplifies some of the code by making use of Numcodecs utility functions. This avoids a few copies when working with some of the stores. Also generally makes the rest easier to understand.

Make sure that datetime/timedelta arrays are cast to a type that
supports the buffer protocol. Ensure this is a type that can handle all
of the datetime/timedelta values and has the same itemsize.
Instead of using `ensure_ndarray`, use `ensure_contiguous_ndarray` with
the stores. This ensures that datetime/timedeltas are handled by
default. Also catches things like object arrays. Finally this handles
flattening the array if needed.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Replacing with PR ( #352 ), which goes to Numcodecs 0.6.2. It also consolidates the changes from here and drops a few that have been pulled into other PRs. Namely ensuring DictStore contains only bytes for data ( #350 ) and changing the default backend of Array to DictStore ( #351 ).

@jakirkhamjakirkham removed this from the v2.3 milestone Dec 4, 2018
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jakirkham@alimanfoo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Bump Numcodecs requirement to 0.6.1 - #347

Closed
jakirkham wants to merge 21 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.1
Closed

Bump Numcodecs requirement to 0.6.1#347
jakirkham wants to merge 21 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.1

Conversation

@jakirkham

Copy link
Copy Markdown
Member

As there are some critical fixes in the latest Numcodecs, this bumps our lower bound to the latest version.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

@alimanfooalimanfoo added this to the v2.3 milestone Nov 30, 2018

@alimanfooalimanfoo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @jakirkham. Probably worth adding a brief release note.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

FTR saw some test failures locally that it looks like CI also encounters. Figured we'd want to look at these and determine how we want to fix them.

@alimanfoo

Copy link
Copy Markdown
Member

Ah, yes, sorry, forgot about that. I expect some test failures due to #324. May be other consequences of the upgrade too. Suggest dealing with them in this PR.

@jakirkham

jakirkham commented Nov 30, 2018

Copy link
Copy Markdown
MemberAuthor

SGTM. How about this one (MsgPack)?

# create an object array using picklez=self.create_array(shape=10, chunks=3, dtype=object, object_codec=Pickle())
z[0] ='foo'>assertz[0] =='foo'

ref: https://travis-ci.org/zarr-developers/zarr/jobs/461883537#L2717-L2720

...or this one (Categorize)?

withpytest.raises(RuntimeError):
# noinspection PyStatementEffect>v[:]
EFailed: DIDNOTRAISE<type'exceptions.RuntimeError'>

ref: https://travis-ci.org/zarr-developers/zarr/jobs/461883537#L2621-L2624

Previously MsgPack was turning bytes objects to unicode objects when
round-tripping them. However this has been fixed in the latest version
of Numcodecs. So correct this test now that MsgPack is working
correctly.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Also there is another failure. Appears Pickle wasn't able to handle general buffer protocol conforming types in decode. Though every other codec worked fine with these (which seems to be a consequence of them using the new utility functions 😉). Apparently Pickle was not using these. Have put together PR ( zarr-developers/numcodecs#143 ) with a fix for this. Would be good to do another patch release to get that out (and anything else these failures might need).

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Based on testing locally, it looks like we are down to the RuntimeError not being raised by an object array after removing the filters.

forcompressorinZlib(1), Blosc():
z=self.create_array(shape=len(data), chunks=30, dtype=object,
object_codec=Categorize(greetings,
dtype=object),
compressor=compressor)
z[:] =datav=z.view(filters=[])
withpytest.raises(RuntimeError):
# noinspection PyStatementEffect>v[:]
EFailed: DIDNOTRAISE<type'exceptions.RuntimeError'>

@alimanfoo

Copy link
Copy Markdown
Member

Hi @jakirkham, took the liberty to push a commit that should resolve the test failure due to RuntimeError not being raised for object arrays.

Other issues should be resolved with the fixes you've merged into numcodecs, so I'll cut a numcodecs 0.6.2 release then come back and update the version requirement here.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Thanks @alimanfoo!

As we already ensured the `chunk` is an `ndarray` viewing the original
data, there is no need for us to do that here as well. Plus the checks
performed by `ensure_contiguous_ndarray` are not needed for our use case
here. Particularly as we have already handled the unusual type cases
above. We also don't need to constrain the buffer size. As such the only
thing we really need is to flatten the array and make it contiguous,
which is what we handle here directly.
As both the expected `object` case and the non-`object` case perform a
`reshape` to flatten the data, go ahead and refactor that out of both
cases and handle it generally. Simplifies the code a bit.
As refactoring of the `reshape` step has effectively dropped the
expected `object` type case, the checks for different types is a little
more complicated than needed. To fix this, basically invert and swap the
case ordering. This way we can handle all generally expected types first
and simply cast them. Then we can raise if an `object` type shows up and
is unexpected.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Did a little bit of refactoring on your change. Hope that is ok. Please let me know your thoughts on that.

The `DictStore` is pretty reliant on the fact that values are immutable
and can be easily compared. For example `__eq__` assumes that all
contents can be compared easily. This works fine if the data is `bytes`.
However it doesn't really work for `ndarray`s for example. Previously we
would have stored whatever the user gave us here. This means comparisons
could falldown in those cases as well (much as the example in the
tutorial has highlighted on CI). Now we effectively require that the
data be something that can either be coerced to `bytes` (e.g. via the
new/old buffer protocol) or is `bytes` to begin with. Make sure not to
force this requirement when nesting one `MutableMapping` within another.
This test case seems to be ill-posed. Anytime we store `object`s to
`Array`s we require an `object_codec` to be specified. Otherwise we have
no clean way to serialize the data. However this `DictStore` test breaks
that assumption by explicitly storing an `object` type in it even though
this would never work for the other stores (particularly when working
with `Array`s). This includes in-memory Zarr `Array`s, which would be
backed by `DictStore`. Given this, we go ahead and drop this test case.
Instead of using a Python `dict` as the `default` store for a Zarr
`Array`, use the `DictStore`. This ensures that all blobs will be
represented as `bytes` regardless of what the user provided as data.
Thus things like comparisons of stores will work well in the default
case.
@jakirkham

jakirkham commented Dec 2, 2018

Copy link
Copy Markdown
MemberAuthor

On a different point, it looks like __eq__ in DictStore is assuming all items in it are comparable. This assumption runs into problems however when anything that is not easily comparable is stored in the DictStore. For instance, ndarray is comparable, but does not reduce to a bool as expected, which causes the tutorial's comparison to fail CI. There are likely other ill-posed situations we can imagine.

Looking at this example before the recent Numcodecs release, it seems the contents of DictStore were always being filled with bytes likely as an indirect consequence of applying filters to the data first. Have added a commit, which enforces this behavior intentionally using ensure_bytes from Numcodecs. Thus making this behavior something we can rely upon for things like __eq__ and such.

However there is a wrinkle with this approach. It appears we had a test that was assuming that objects could be placed in DictStores. Though this behavior doesn't really work for any of the other stores and we usually require an object_codec to be specified to even serialize an object type (specifically with Array). Given this, I've dropped that test and am contending it doesn't conform to our expectations about how stores typically should work with object types.

Also it appears that the Array's default store was a dict, which we cannot apply these fixes to. So have updated the default to DictStore. I'm not sure why this wasn't already the case, which means I'm probably missing something.

Happy to hear thoughts or receive pushback on this approach as I may very well be missing something. If you know of a better approach, would also be interested in hearing about that as well.

Also raised issue ( #348 ) to provide a nice simple example of this problem.

@jakirkham

jakirkham commented Dec 2, 2018

Copy link
Copy Markdown
MemberAuthor

Seems the tutorial test is now hung up on some whitespace issue. Not totally sure how that got introduced.

NVM this was more fallout from the switch to DictStore as the default store for Array.

Edit: This has since been fixed.

As we are now using `DictStore` to back the `Array`, we can correctly
measure how much memory it is using. So update the examples in `info`
and the tutorial to show how much memory is being used. Also update the
store type listed in info as well.
As Numcodecs now includes a very versatile and effective `ensure_bytes`
function, there is no need to define our own in `zarr.storage` as well.
So go ahead and drop it.
As this is no longer being used by `ensure_bytes` as that function was
dropped, go ahead and drop `binary_type` as well.
Make use of Numcodecs' `ensure_ndarray` to take `ndarray` views onto
buffers to be stored in a few cases so as to reshape them and avoid a
copy (thanks to the buffer protocol).
Rewrite `buffer_size` to just use Numcodecs' `ensure_ndarray` to get an
`ndarray` that views the data. Once the `ndarray` is gotten, all that is
needed is to access its `nbytes` member, which returns the number of
bytes that it takes up.
If the data is already a `str` instance, turn `ensure_str` into a no-op.
For all other cases, make use of Numcodecs' `ensure_bytes` to aid
`ensure_str` in coercing data through the buffer protocol. If we are on
Python 3, then decode the `bytes` object to a `str`.
As `DictStore` now must only store `bytes` or types coercible to bytes
via the buffer protocol, there is no possibility for it to have unknown
sizes as `bytes` always have a known size. So drop these cases where the
size can be `-1`.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Simplifies some of the code by making use of Numcodecs utility functions. This avoids a few copies when working with some of the stores. Also generally makes the rest easier to understand.

Make sure that datetime/timedelta arrays are cast to a type that
supports the buffer protocol. Ensure this is a type that can handle all
of the datetime/timedelta values and has the same itemsize.
Instead of using `ensure_ndarray`, use `ensure_contiguous_ndarray` with
the stores. This ensures that datetime/timedeltas are handled by
default. Also catches things like object arrays. Finally this handles
flattening the array if needed.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Replacing with PR ( #352 ), which goes to Numcodecs 0.6.2. It also consolidates the changes from here and drops a few that have been pulled into other PRs. Namely ensuring DictStore contains only bytes for data ( #350 ) and changing the default backend of Array to DictStore ( #351 ).

@jakirkhamjakirkham removed this from the v2.3 milestone Dec 4, 2018
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jakirkham@alimanfoo