Use ensure_ndarray in a few more places - #506

Merged
jakirkham merged 5 commits into
zarr-developers:masterfrom
jakirkham:use_ensure_ndarray_more
Nov 12, 2019
Merged

Use ensure_ndarray in a few more places#506
jakirkham merged 5 commits into
zarr-developers:masterfrom
jakirkham:use_ensure_ndarray_more

Conversation

@jakirkham

Copy link
Copy Markdown
Member

As we can coerce any object that is array-like to an ndarray, there should be no need for these explicit ndarray checks. So eliminate them by using ensure_ndarray to get an ndarray object instead.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

Instead of checking to see if `chunk` is an `ndarray`, use
`ensure_ndarray` to get an `ndarray` viewing the underlying data in
`chunk`. This way we know attributes like `dtype` are available and can
be checked easily. Also makes this a bit more friendly with other
array-like types.
Since we are interested in getting an `ndarray` representing the buffer
within `out` for writing into, go ahead and use `ensure_ndarray` to
coerce the underlying buffer into an `ndarray`. This way we can avoid a
needless check and just write into any array-like value for `out` that
is provided.
Appears that `out` can also be a Zarr `Array` or any other array-like
that does not expose a buffer, but does allow writing into. In these
cases `ensure_ndarray` will fail as there is not an underlying buffer
that can be used with an `ndarray`. To also handle this case, catch the
`TypeError` that `ensure_ndarray` will raise in this case and use that
to indicate whether `out` is now an `ndarray` or not. This allows us to
continue to write into arbitrary buffers, but also correctly handle
objects that do not expose buffers.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

@jrbourbeau, if you are around, this could use a review 🙂

Comment threadzarr/core.py

out_is_ndarray = True
try:
out = ensure_ndarray(out)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note that this is already checked by test_get_selection_out. Without this try/except, we get the following TypeError because out is a Zarr Array, which cannot be coerced to an ndarray. This is ok and intentional. So we just carry on without coercing that case and note that we do not have an ndarray when checking later.

Details
_________________________test_get_selection_out____________________________deftest_get_selection_out():
# basic selectionsa=np.arange(1050)
z=zarr.create(shape=1050, chunks=100, dtype=a.dtype)
z[:] =aselections= [
slice(50, 150),
slice(0, 1050),
slice(1, 2),
]
forselectioninselections:
expect=a[selection]
out=zarr.create(shape=expect.shape, chunks=10, dtype=expect.dtype, fill_value=0)
>z.get_basic_selection(selection, out=out)
zarr/tests/test_indexing.py:1036: _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ zarr/core.py:698: inget_basic_selectionfields=fields)
zarr/core.py:740: in_get_basic_selection_ndreturnself._get_selection(indexer=indexer, out=out, fields=fields)
zarr/core.py:1028: in_get_selectiondrop_axes=indexer.drop_axes, fields=fields)
zarr/core.py:1573: in_chunk_getitemout=ensure_ndarray(out)
__ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ buf=<zarr.core.Array (100,) int64>defensure_ndarray(buf):
"""Convenience function to coerce `buf` to a numpy array, if it is not already a numpy array. Parameters ---------- buf : array-like or bytes-like A numpy array or any object exporting a buffer interface. Returns ------- arr : ndarray A numpy array, sharing memory with `buf`. Notes ----- This function will not create a copy under any circumstances, it is guaranteed to return a view on memory exported by `buf`. """ifisinstance(buf, np.ndarray):
# already a numpy arrayarr=bufelifisinstance(buf, array.array) andbuf.typecodein'cu':
# Guard condition, do not support array.array with unicode type, this is# problematic because numpy does not support it on all platforms. Also do not# support char as it was removed in Python 3.raiseTypeError('array.array with char or unicode type is not supported')
else:
# N.B., first take a memoryview to make sure that we subsequently create a# numpy array from a memory buffer with no copyifPY2: # pragma: py3 no covertry:
mem=memoryview(buf)
exceptTypeError:
# on PY2 also check if object exports old-style buffer interfacemem=np.getbuffer(buf)
else: # pragma: py2 no cover>mem=memoryview(buf)
ETypeError: memoryview: abytes-likeobjectisrequired, not'Array'
.tox/py36/lib/python3.6/site-packages/numcodecs/compat.py:74: TypeError

ref: https://travis-ci.org/zarr-developers/zarr-python/jobs/610533967#L2316

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Planning on merging EOD tomorrow if no comments.

@jrbourbeaujrbourbeau left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the PR @jakirkham! Generally these changes seem fine by me. I have one question about when ensure_ndarray raises an error that I've left below.

Comment threadzarr/core.py

# check object encoding
if isinstance(chunk, np.ndarray) and chunk.dtype == object:
if ensure_ndarray(chunk).dtype == object:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is it possible that some non-ndarrays would previously skip over this if block but now fail if ensure_ndarray raises an error? For example:

In [28]: fromnumcodecs.compatimportensure_ndarrayIn [29]: importarrayIn [30]: chunk=array.array('u', 'hello \u2641')
In [31]: ifisinstance(chunk, np.ndarray) andchunk.dtype==object:
...: pass
...:
In [32]: ifensure_ndarray(chunk).dtype==object:
...: pass
...:
---------------------------------------------------------------------------TypeErrorTraceback (mostrecentcalllast)
<ipython-input-32-e1e15cbf81c2>in<module>---->1ifensure_ndarray(chunk).dtype==object:
2pass3~/miniconda/envs/zarr-python-dev/lib/python3.7/site-packages/numcodecs/compat.pyinensure_ndarray(buf)
63# problematic because numpy does not support it on all platforms. Also do not64# support char as it was removed in Python 3.--->65raiseTypeError('array.array with char or unicode type is not supported')
6667else:
TypeError: array.arraywithcharorunicodetypeisnotsupported

To be clear, I'm not sure how likely this is (or if it's even possible) to come up in practice. Do you have a sense for this?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That exception would be expected as we have decided not to work with Python builtin arrays that use character or unicode types. There is some more detailed discussion in this thread.

@jakirkhamjakirkhamNov 12, 2019

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Basically the only exception that we need to worry about is the one coming from memoryview, which will be a TypeError if the object doesn't support the buffer protocol. After that point we have a memoryview. So we know the rest of the code will work (unless NumPy gets a bug ;).

Edit: Sorry was looking at the other change for a second.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For context this function is trying to turn some data into bytes that can be serialized (say to a file on disk).

The filters before this step should be returning something that either is an ndarray or could be coerced to one. This is required by the compressors that follow and the storage layer afterwards. So if ensure_ndarray fails here, then we have invalid data and raising an exception to the user would be appropriate.

@jrbourbeaujrbourbeau left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great, thank you for providing all the additional context surrounding this section. I pushed a commit with a changelog entry. Feel free to merge on green. Thanks @jakirkham!

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Thanks @jrbourbeau for the review and fix! 😄

@jakirkham
jakirkham merged commit 58b1786 into zarr-developers:masterNov 12, 2019
@jakirkham
jakirkham deleted the use_ensure_ndarray_more branch November 12, 2019 18:56
@CarreauCarreau added this to the v2.4 milestone Sep 9, 2020
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jakirkham@jrbourbeau@Carreau
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Use ensure_ndarray in a few more places - #506

Merged
jakirkham merged 5 commits into
zarr-developers:masterfrom
jakirkham:use_ensure_ndarray_more
Nov 12, 2019
Merged

Use ensure_ndarray in a few more places#506
jakirkham merged 5 commits into
zarr-developers:masterfrom
jakirkham:use_ensure_ndarray_more

Conversation

@jakirkham

Copy link
Copy Markdown
Member

As we can coerce any object that is array-like to an ndarray, there should be no need for these explicit ndarray checks. So eliminate them by using ensure_ndarray to get an ndarray object instead.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

Instead of checking to see if `chunk` is an `ndarray`, use
`ensure_ndarray` to get an `ndarray` viewing the underlying data in
`chunk`. This way we know attributes like `dtype` are available and can
be checked easily. Also makes this a bit more friendly with other
array-like types.
Since we are interested in getting an `ndarray` representing the buffer
within `out` for writing into, go ahead and use `ensure_ndarray` to
coerce the underlying buffer into an `ndarray`. This way we can avoid a
needless check and just write into any array-like value for `out` that
is provided.
Appears that `out` can also be a Zarr `Array` or any other array-like
that does not expose a buffer, but does allow writing into. In these
cases `ensure_ndarray` will fail as there is not an underlying buffer
that can be used with an `ndarray`. To also handle this case, catch the
`TypeError` that `ensure_ndarray` will raise in this case and use that
to indicate whether `out` is now an `ndarray` or not. This allows us to
continue to write into arbitrary buffers, but also correctly handle
objects that do not expose buffers.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

@jrbourbeau, if you are around, this could use a review 🙂

Comment threadzarr/core.py

out_is_ndarray = True
try:
out = ensure_ndarray(out)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note that this is already checked by test_get_selection_out. Without this try/except, we get the following TypeError because out is a Zarr Array, which cannot be coerced to an ndarray. This is ok and intentional. So we just carry on without coercing that case and note that we do not have an ndarray when checking later.

Details
_________________________test_get_selection_out____________________________deftest_get_selection_out():
# basic selectionsa=np.arange(1050)
z=zarr.create(shape=1050, chunks=100, dtype=a.dtype)
z[:] =aselections= [
slice(50, 150),
slice(0, 1050),
slice(1, 2),
]
forselectioninselections:
expect=a[selection]
out=zarr.create(shape=expect.shape, chunks=10, dtype=expect.dtype, fill_value=0)
>z.get_basic_selection(selection, out=out)
zarr/tests/test_indexing.py:1036: _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ zarr/core.py:698: inget_basic_selectionfields=fields)
zarr/core.py:740: in_get_basic_selection_ndreturnself._get_selection(indexer=indexer, out=out, fields=fields)
zarr/core.py:1028: in_get_selectiondrop_axes=indexer.drop_axes, fields=fields)
zarr/core.py:1573: in_chunk_getitemout=ensure_ndarray(out)
__ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ buf=<zarr.core.Array (100,) int64>defensure_ndarray(buf):
"""Convenience function to coerce `buf` to a numpy array, if it is not already a numpy array. Parameters ---------- buf : array-like or bytes-like A numpy array or any object exporting a buffer interface. Returns ------- arr : ndarray A numpy array, sharing memory with `buf`. Notes ----- This function will not create a copy under any circumstances, it is guaranteed to return a view on memory exported by `buf`. """ifisinstance(buf, np.ndarray):
# already a numpy arrayarr=bufelifisinstance(buf, array.array) andbuf.typecodein'cu':
# Guard condition, do not support array.array with unicode type, this is# problematic because numpy does not support it on all platforms. Also do not# support char as it was removed in Python 3.raiseTypeError('array.array with char or unicode type is not supported')
else:
# N.B., first take a memoryview to make sure that we subsequently create a# numpy array from a memory buffer with no copyifPY2: # pragma: py3 no covertry:
mem=memoryview(buf)
exceptTypeError:
# on PY2 also check if object exports old-style buffer interfacemem=np.getbuffer(buf)
else: # pragma: py2 no cover>mem=memoryview(buf)
ETypeError: memoryview: abytes-likeobjectisrequired, not'Array'
.tox/py36/lib/python3.6/site-packages/numcodecs/compat.py:74: TypeError

ref: https://travis-ci.org/zarr-developers/zarr-python/jobs/610533967#L2316

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Planning on merging EOD tomorrow if no comments.

@jrbourbeaujrbourbeau left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the PR @jakirkham! Generally these changes seem fine by me. I have one question about when ensure_ndarray raises an error that I've left below.

Comment threadzarr/core.py

# check object encoding
if isinstance(chunk, np.ndarray) and chunk.dtype == object:
if ensure_ndarray(chunk).dtype == object:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is it possible that some non-ndarrays would previously skip over this if block but now fail if ensure_ndarray raises an error? For example:

In [28]: fromnumcodecs.compatimportensure_ndarrayIn [29]: importarrayIn [30]: chunk=array.array('u', 'hello \u2641')
In [31]: ifisinstance(chunk, np.ndarray) andchunk.dtype==object:
...: pass
...:
In [32]: ifensure_ndarray(chunk).dtype==object:
...: pass
...:
---------------------------------------------------------------------------TypeErrorTraceback (mostrecentcalllast)
<ipython-input-32-e1e15cbf81c2>in<module>---->1ifensure_ndarray(chunk).dtype==object:
2pass3~/miniconda/envs/zarr-python-dev/lib/python3.7/site-packages/numcodecs/compat.pyinensure_ndarray(buf)
63# problematic because numpy does not support it on all platforms. Also do not64# support char as it was removed in Python 3.--->65raiseTypeError('array.array with char or unicode type is not supported')
6667else:
TypeError: array.arraywithcharorunicodetypeisnotsupported

To be clear, I'm not sure how likely this is (or if it's even possible) to come up in practice. Do you have a sense for this?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That exception would be expected as we have decided not to work with Python builtin arrays that use character or unicode types. There is some more detailed discussion in this thread.

@jakirkhamjakirkhamNov 12, 2019

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Basically the only exception that we need to worry about is the one coming from memoryview, which will be a TypeError if the object doesn't support the buffer protocol. After that point we have a memoryview. So we know the rest of the code will work (unless NumPy gets a bug ;).

Edit: Sorry was looking at the other change for a second.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For context this function is trying to turn some data into bytes that can be serialized (say to a file on disk).

The filters before this step should be returning something that either is an ndarray or could be coerced to one. This is required by the compressors that follow and the storage layer afterwards. So if ensure_ndarray fails here, then we have invalid data and raising an exception to the user would be appropriate.

@jrbourbeaujrbourbeau left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great, thank you for providing all the additional context surrounding this section. I pushed a commit with a changelog entry. Feel free to merge on green. Thanks @jakirkham!

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Thanks @jrbourbeau for the review and fix! 😄

@jakirkham
jakirkham merged commit 58b1786 into zarr-developers:masterNov 12, 2019
@jakirkham
jakirkham deleted the use_ensure_ndarray_more branch November 12, 2019 18:56
@CarreauCarreau added this to the v2.4 milestone Sep 9, 2020
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jakirkham@jrbourbeau@Carreau
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use ensure_ndarray in a few more places - #506

Merged
jakirkham merged 5 commits into
zarr-developers:masterfrom
jakirkham:use_ensure_ndarray_more
Nov 12, 2019
Merged

Use ensure_ndarray in a few more places#506
jakirkham merged 5 commits into
zarr-developers:masterfrom
jakirkham:use_ensure_ndarray_more

Conversation

@jakirkham

Copy link
Copy Markdown
Member

As we can coerce any object that is array-like to an ndarray, there should be no need for these explicit ndarray checks. So eliminate them by using ensure_ndarray to get an ndarray object instead.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

Instead of checking to see if `chunk` is an `ndarray`, use
`ensure_ndarray` to get an `ndarray` viewing the underlying data in
`chunk`. This way we know attributes like `dtype` are available and can
be checked easily. Also makes this a bit more friendly with other
array-like types.
Since we are interested in getting an `ndarray` representing the buffer
within `out` for writing into, go ahead and use `ensure_ndarray` to
coerce the underlying buffer into an `ndarray`. This way we can avoid a
needless check and just write into any array-like value for `out` that
is provided.
Appears that `out` can also be a Zarr `Array` or any other array-like
that does not expose a buffer, but does allow writing into. In these
cases `ensure_ndarray` will fail as there is not an underlying buffer
that can be used with an `ndarray`. To also handle this case, catch the
`TypeError` that `ensure_ndarray` will raise in this case and use that
to indicate whether `out` is now an `ndarray` or not. This allows us to
continue to write into arbitrary buffers, but also correctly handle
objects that do not expose buffers.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

@jrbourbeau, if you are around, this could use a review 🙂

Comment threadzarr/core.py

out_is_ndarray = True
try:
out = ensure_ndarray(out)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note that this is already checked by test_get_selection_out. Without this try/except, we get the following TypeError because out is a Zarr Array, which cannot be coerced to an ndarray. This is ok and intentional. So we just carry on without coercing that case and note that we do not have an ndarray when checking later.

Details
_________________________test_get_selection_out____________________________deftest_get_selection_out():
# basic selectionsa=np.arange(1050)
z=zarr.create(shape=1050, chunks=100, dtype=a.dtype)
z[:] =aselections= [
slice(50, 150),
slice(0, 1050),
slice(1, 2),
]
forselectioninselections:
expect=a[selection]
out=zarr.create(shape=expect.shape, chunks=10, dtype=expect.dtype, fill_value=0)
>z.get_basic_selection(selection, out=out)
zarr/tests/test_indexing.py:1036: _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ zarr/core.py:698: inget_basic_selectionfields=fields)
zarr/core.py:740: in_get_basic_selection_ndreturnself._get_selection(indexer=indexer, out=out, fields=fields)
zarr/core.py:1028: in_get_selectiondrop_axes=indexer.drop_axes, fields=fields)
zarr/core.py:1573: in_chunk_getitemout=ensure_ndarray(out)
__ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ buf=<zarr.core.Array (100,) int64>defensure_ndarray(buf):
"""Convenience function to coerce `buf` to a numpy array, if it is not already a numpy array. Parameters ---------- buf : array-like or bytes-like A numpy array or any object exporting a buffer interface. Returns ------- arr : ndarray A numpy array, sharing memory with `buf`. Notes ----- This function will not create a copy under any circumstances, it is guaranteed to return a view on memory exported by `buf`. """ifisinstance(buf, np.ndarray):
# already a numpy arrayarr=bufelifisinstance(buf, array.array) andbuf.typecodein'cu':
# Guard condition, do not support array.array with unicode type, this is# problematic because numpy does not support it on all platforms. Also do not# support char as it was removed in Python 3.raiseTypeError('array.array with char or unicode type is not supported')
else:
# N.B., first take a memoryview to make sure that we subsequently create a# numpy array from a memory buffer with no copyifPY2: # pragma: py3 no covertry:
mem=memoryview(buf)
exceptTypeError:
# on PY2 also check if object exports old-style buffer interfacemem=np.getbuffer(buf)
else: # pragma: py2 no cover>mem=memoryview(buf)
ETypeError: memoryview: abytes-likeobjectisrequired, not'Array'
.tox/py36/lib/python3.6/site-packages/numcodecs/compat.py:74: TypeError

ref: https://travis-ci.org/zarr-developers/zarr-python/jobs/610533967#L2316

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Planning on merging EOD tomorrow if no comments.

@jrbourbeaujrbourbeau left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the PR @jakirkham! Generally these changes seem fine by me. I have one question about when ensure_ndarray raises an error that I've left below.

Comment threadzarr/core.py

# check object encoding
if isinstance(chunk, np.ndarray) and chunk.dtype == object:
if ensure_ndarray(chunk).dtype == object:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is it possible that some non-ndarrays would previously skip over this if block but now fail if ensure_ndarray raises an error? For example:

In [28]: fromnumcodecs.compatimportensure_ndarrayIn [29]: importarrayIn [30]: chunk=array.array('u', 'hello \u2641')
In [31]: ifisinstance(chunk, np.ndarray) andchunk.dtype==object:
...: pass
...:
In [32]: ifensure_ndarray(chunk).dtype==object:
...: pass
...:
---------------------------------------------------------------------------TypeErrorTraceback (mostrecentcalllast)
<ipython-input-32-e1e15cbf81c2>in<module>---->1ifensure_ndarray(chunk).dtype==object:
2pass3~/miniconda/envs/zarr-python-dev/lib/python3.7/site-packages/numcodecs/compat.pyinensure_ndarray(buf)
63# problematic because numpy does not support it on all platforms. Also do not64# support char as it was removed in Python 3.--->65raiseTypeError('array.array with char or unicode type is not supported')
6667else:
TypeError: array.arraywithcharorunicodetypeisnotsupported

To be clear, I'm not sure how likely this is (or if it's even possible) to come up in practice. Do you have a sense for this?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That exception would be expected as we have decided not to work with Python builtin arrays that use character or unicode types. There is some more detailed discussion in this thread.

@jakirkhamjakirkhamNov 12, 2019

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Basically the only exception that we need to worry about is the one coming from memoryview, which will be a TypeError if the object doesn't support the buffer protocol. After that point we have a memoryview. So we know the rest of the code will work (unless NumPy gets a bug ;).

Edit: Sorry was looking at the other change for a second.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For context this function is trying to turn some data into bytes that can be serialized (say to a file on disk).

The filters before this step should be returning something that either is an ndarray or could be coerced to one. This is required by the compressors that follow and the storage layer afterwards. So if ensure_ndarray fails here, then we have invalid data and raising an exception to the user would be appropriate.

@jrbourbeaujrbourbeau left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great, thank you for providing all the additional context surrounding this section. I pushed a commit with a changelog entry. Feel free to merge on green. Thanks @jakirkham!

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Thanks @jrbourbeau for the review and fix! 😄

@jakirkham
jakirkham merged commit 58b1786 into zarr-developers:masterNov 12, 2019
@jakirkham
jakirkham deleted the use_ensure_ndarray_more branch November 12, 2019 18:56
@CarreauCarreau added this to the v2.4 milestone Sep 9, 2020
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jakirkham@jrbourbeau@Carreau
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use ensure_ndarray in a few more places - #506

Merged
jakirkham merged 5 commits into
zarr-developers:masterfrom
jakirkham:use_ensure_ndarray_more
Nov 12, 2019
Merged

Use ensure_ndarray in a few more places#506
jakirkham merged 5 commits into
zarr-developers:masterfrom
jakirkham:use_ensure_ndarray_more

Conversation

@jakirkham

Copy link
Copy Markdown
Member

As we can coerce any object that is array-like to an ndarray, there should be no need for these explicit ndarray checks. So eliminate them by using ensure_ndarray to get an ndarray object instead.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

Instead of checking to see if `chunk` is an `ndarray`, use
`ensure_ndarray` to get an `ndarray` viewing the underlying data in
`chunk`. This way we know attributes like `dtype` are available and can
be checked easily. Also makes this a bit more friendly with other
array-like types.
Since we are interested in getting an `ndarray` representing the buffer
within `out` for writing into, go ahead and use `ensure_ndarray` to
coerce the underlying buffer into an `ndarray`. This way we can avoid a
needless check and just write into any array-like value for `out` that
is provided.
Appears that `out` can also be a Zarr `Array` or any other array-like
that does not expose a buffer, but does allow writing into. In these
cases `ensure_ndarray` will fail as there is not an underlying buffer
that can be used with an `ndarray`. To also handle this case, catch the
`TypeError` that `ensure_ndarray` will raise in this case and use that
to indicate whether `out` is now an `ndarray` or not. This allows us to
continue to write into arbitrary buffers, but also correctly handle
objects that do not expose buffers.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

@jrbourbeau, if you are around, this could use a review 🙂

Comment threadzarr/core.py

out_is_ndarray = True
try:
out = ensure_ndarray(out)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note that this is already checked by test_get_selection_out. Without this try/except, we get the following TypeError because out is a Zarr Array, which cannot be coerced to an ndarray. This is ok and intentional. So we just carry on without coercing that case and note that we do not have an ndarray when checking later.

Details
_________________________test_get_selection_out____________________________deftest_get_selection_out():
# basic selectionsa=np.arange(1050)
z=zarr.create(shape=1050, chunks=100, dtype=a.dtype)
z[:] =aselections= [
slice(50, 150),
slice(0, 1050),
slice(1, 2),
]
forselectioninselections:
expect=a[selection]
out=zarr.create(shape=expect.shape, chunks=10, dtype=expect.dtype, fill_value=0)
>z.get_basic_selection(selection, out=out)
zarr/tests/test_indexing.py:1036: _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ zarr/core.py:698: inget_basic_selectionfields=fields)
zarr/core.py:740: in_get_basic_selection_ndreturnself._get_selection(indexer=indexer, out=out, fields=fields)
zarr/core.py:1028: in_get_selectiondrop_axes=indexer.drop_axes, fields=fields)
zarr/core.py:1573: in_chunk_getitemout=ensure_ndarray(out)
__ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ buf=<zarr.core.Array (100,) int64>defensure_ndarray(buf):
"""Convenience function to coerce `buf` to a numpy array, if it is not already a numpy array. Parameters ---------- buf : array-like or bytes-like A numpy array or any object exporting a buffer interface. Returns ------- arr : ndarray A numpy array, sharing memory with `buf`. Notes ----- This function will not create a copy under any circumstances, it is guaranteed to return a view on memory exported by `buf`. """ifisinstance(buf, np.ndarray):
# already a numpy arrayarr=bufelifisinstance(buf, array.array) andbuf.typecodein'cu':
# Guard condition, do not support array.array with unicode type, this is# problematic because numpy does not support it on all platforms. Also do not# support char as it was removed in Python 3.raiseTypeError('array.array with char or unicode type is not supported')
else:
# N.B., first take a memoryview to make sure that we subsequently create a# numpy array from a memory buffer with no copyifPY2: # pragma: py3 no covertry:
mem=memoryview(buf)
exceptTypeError:
# on PY2 also check if object exports old-style buffer interfacemem=np.getbuffer(buf)
else: # pragma: py2 no cover>mem=memoryview(buf)
ETypeError: memoryview: abytes-likeobjectisrequired, not'Array'
.tox/py36/lib/python3.6/site-packages/numcodecs/compat.py:74: TypeError

ref: https://travis-ci.org/zarr-developers/zarr-python/jobs/610533967#L2316

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Planning on merging EOD tomorrow if no comments.

@jrbourbeaujrbourbeau left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the PR @jakirkham! Generally these changes seem fine by me. I have one question about when ensure_ndarray raises an error that I've left below.

Comment threadzarr/core.py

# check object encoding
if isinstance(chunk, np.ndarray) and chunk.dtype == object:
if ensure_ndarray(chunk).dtype == object:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is it possible that some non-ndarrays would previously skip over this if block but now fail if ensure_ndarray raises an error? For example:

In [28]: fromnumcodecs.compatimportensure_ndarrayIn [29]: importarrayIn [30]: chunk=array.array('u', 'hello \u2641')
In [31]: ifisinstance(chunk, np.ndarray) andchunk.dtype==object:
...: pass
...:
In [32]: ifensure_ndarray(chunk).dtype==object:
...: pass
...:
---------------------------------------------------------------------------TypeErrorTraceback (mostrecentcalllast)
<ipython-input-32-e1e15cbf81c2>in<module>---->1ifensure_ndarray(chunk).dtype==object:
2pass3~/miniconda/envs/zarr-python-dev/lib/python3.7/site-packages/numcodecs/compat.pyinensure_ndarray(buf)
63# problematic because numpy does not support it on all platforms. Also do not64# support char as it was removed in Python 3.--->65raiseTypeError('array.array with char or unicode type is not supported')
6667else:
TypeError: array.arraywithcharorunicodetypeisnotsupported

To be clear, I'm not sure how likely this is (or if it's even possible) to come up in practice. Do you have a sense for this?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That exception would be expected as we have decided not to work with Python builtin arrays that use character or unicode types. There is some more detailed discussion in this thread.

@jakirkhamjakirkhamNov 12, 2019

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Basically the only exception that we need to worry about is the one coming from memoryview, which will be a TypeError if the object doesn't support the buffer protocol. After that point we have a memoryview. So we know the rest of the code will work (unless NumPy gets a bug ;).

Edit: Sorry was looking at the other change for a second.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For context this function is trying to turn some data into bytes that can be serialized (say to a file on disk).

The filters before this step should be returning something that either is an ndarray or could be coerced to one. This is required by the compressors that follow and the storage layer afterwards. So if ensure_ndarray fails here, then we have invalid data and raising an exception to the user would be appropriate.

@jrbourbeaujrbourbeau left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great, thank you for providing all the additional context surrounding this section. I pushed a commit with a changelog entry. Feel free to merge on green. Thanks @jakirkham!

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Thanks @jrbourbeau for the review and fix! 😄

@jakirkham
jakirkham merged commit 58b1786 into zarr-developers:masterNov 12, 2019
@jakirkham
jakirkham deleted the use_ensure_ndarray_more branch November 12, 2019 18:56
@CarreauCarreau added this to the v2.4 milestone Sep 9, 2020
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jakirkham@jrbourbeau@Carreau
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Use ensure_ndarray in a few more places - #506

Merged
jakirkham merged 5 commits into
zarr-developers:masterfrom
jakirkham:use_ensure_ndarray_more
Nov 12, 2019
Merged

Use ensure_ndarray in a few more places#506
jakirkham merged 5 commits into
zarr-developers:masterfrom
jakirkham:use_ensure_ndarray_more

Conversation

@jakirkham

Copy link
Copy Markdown
Member

As we can coerce any object that is array-like to an ndarray, there should be no need for these explicit ndarray checks. So eliminate them by using ensure_ndarray to get an ndarray object instead.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

Instead of checking to see if `chunk` is an `ndarray`, use
`ensure_ndarray` to get an `ndarray` viewing the underlying data in
`chunk`. This way we know attributes like `dtype` are available and can
be checked easily. Also makes this a bit more friendly with other
array-like types.
Since we are interested in getting an `ndarray` representing the buffer
within `out` for writing into, go ahead and use `ensure_ndarray` to
coerce the underlying buffer into an `ndarray`. This way we can avoid a
needless check and just write into any array-like value for `out` that
is provided.
Appears that `out` can also be a Zarr `Array` or any other array-like
that does not expose a buffer, but does allow writing into. In these
cases `ensure_ndarray` will fail as there is not an underlying buffer
that can be used with an `ndarray`. To also handle this case, catch the
`TypeError` that `ensure_ndarray` will raise in this case and use that
to indicate whether `out` is now an `ndarray` or not. This allows us to
continue to write into arbitrary buffers, but also correctly handle
objects that do not expose buffers.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

@jrbourbeau, if you are around, this could use a review 🙂

Comment threadzarr/core.py

out_is_ndarray = True
try:
out = ensure_ndarray(out)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note that this is already checked by test_get_selection_out. Without this try/except, we get the following TypeError because out is a Zarr Array, which cannot be coerced to an ndarray. This is ok and intentional. So we just carry on without coercing that case and note that we do not have an ndarray when checking later.

Details
_________________________test_get_selection_out____________________________deftest_get_selection_out():
# basic selectionsa=np.arange(1050)
z=zarr.create(shape=1050, chunks=100, dtype=a.dtype)
z[:] =aselections= [
slice(50, 150),
slice(0, 1050),
slice(1, 2),
]
forselectioninselections:
expect=a[selection]
out=zarr.create(shape=expect.shape, chunks=10, dtype=expect.dtype, fill_value=0)
>z.get_basic_selection(selection, out=out)
zarr/tests/test_indexing.py:1036: _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ zarr/core.py:698: inget_basic_selectionfields=fields)
zarr/core.py:740: in_get_basic_selection_ndreturnself._get_selection(indexer=indexer, out=out, fields=fields)
zarr/core.py:1028: in_get_selectiondrop_axes=indexer.drop_axes, fields=fields)
zarr/core.py:1573: in_chunk_getitemout=ensure_ndarray(out)
__ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ buf=<zarr.core.Array (100,) int64>defensure_ndarray(buf):
"""Convenience function to coerce `buf` to a numpy array, if it is not already a numpy array. Parameters ---------- buf : array-like or bytes-like A numpy array or any object exporting a buffer interface. Returns ------- arr : ndarray A numpy array, sharing memory with `buf`. Notes ----- This function will not create a copy under any circumstances, it is guaranteed to return a view on memory exported by `buf`. """ifisinstance(buf, np.ndarray):
# already a numpy arrayarr=bufelifisinstance(buf, array.array) andbuf.typecodein'cu':
# Guard condition, do not support array.array with unicode type, this is# problematic because numpy does not support it on all platforms. Also do not# support char as it was removed in Python 3.raiseTypeError('array.array with char or unicode type is not supported')
else:
# N.B., first take a memoryview to make sure that we subsequently create a# numpy array from a memory buffer with no copyifPY2: # pragma: py3 no covertry:
mem=memoryview(buf)
exceptTypeError:
# on PY2 also check if object exports old-style buffer interfacemem=np.getbuffer(buf)
else: # pragma: py2 no cover>mem=memoryview(buf)
ETypeError: memoryview: abytes-likeobjectisrequired, not'Array'
.tox/py36/lib/python3.6/site-packages/numcodecs/compat.py:74: TypeError

ref: https://travis-ci.org/zarr-developers/zarr-python/jobs/610533967#L2316

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Planning on merging EOD tomorrow if no comments.

@jrbourbeaujrbourbeau left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the PR @jakirkham! Generally these changes seem fine by me. I have one question about when ensure_ndarray raises an error that I've left below.

Comment threadzarr/core.py

# check object encoding
if isinstance(chunk, np.ndarray) and chunk.dtype == object:
if ensure_ndarray(chunk).dtype == object:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is it possible that some non-ndarrays would previously skip over this if block but now fail if ensure_ndarray raises an error? For example:

In [28]: fromnumcodecs.compatimportensure_ndarrayIn [29]: importarrayIn [30]: chunk=array.array('u', 'hello \u2641')
In [31]: ifisinstance(chunk, np.ndarray) andchunk.dtype==object:
...: pass
...:
In [32]: ifensure_ndarray(chunk).dtype==object:
...: pass
...:
---------------------------------------------------------------------------TypeErrorTraceback (mostrecentcalllast)
<ipython-input-32-e1e15cbf81c2>in<module>---->1ifensure_ndarray(chunk).dtype==object:
2pass3~/miniconda/envs/zarr-python-dev/lib/python3.7/site-packages/numcodecs/compat.pyinensure_ndarray(buf)
63# problematic because numpy does not support it on all platforms. Also do not64# support char as it was removed in Python 3.--->65raiseTypeError('array.array with char or unicode type is not supported')
6667else:
TypeError: array.arraywithcharorunicodetypeisnotsupported

To be clear, I'm not sure how likely this is (or if it's even possible) to come up in practice. Do you have a sense for this?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That exception would be expected as we have decided not to work with Python builtin arrays that use character or unicode types. There is some more detailed discussion in this thread.

@jakirkhamjakirkhamNov 12, 2019

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Basically the only exception that we need to worry about is the one coming from memoryview, which will be a TypeError if the object doesn't support the buffer protocol. After that point we have a memoryview. So we know the rest of the code will work (unless NumPy gets a bug ;).

Edit: Sorry was looking at the other change for a second.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For context this function is trying to turn some data into bytes that can be serialized (say to a file on disk).

The filters before this step should be returning something that either is an ndarray or could be coerced to one. This is required by the compressors that follow and the storage layer afterwards. So if ensure_ndarray fails here, then we have invalid data and raising an exception to the user would be appropriate.

@jrbourbeaujrbourbeau left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great, thank you for providing all the additional context surrounding this section. I pushed a commit with a changelog entry. Feel free to merge on green. Thanks @jakirkham!

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Thanks @jrbourbeau for the review and fix! 😄

@jakirkham
jakirkham merged commit 58b1786 into zarr-developers:masterNov 12, 2019
@jakirkham
jakirkham deleted the use_ensure_ndarray_more branch November 12, 2019 18:56
@CarreauCarreau added this to the v2.4 milestone Sep 9, 2020
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jakirkham@jrbourbeau@Carreau
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use ensure_ndarray in a few more places - #506

Merged
jakirkham merged 5 commits into
zarr-developers:masterfrom
jakirkham:use_ensure_ndarray_more
Nov 12, 2019
Merged

Use ensure_ndarray in a few more places#506
jakirkham merged 5 commits into
zarr-developers:masterfrom
jakirkham:use_ensure_ndarray_more

Conversation

@jakirkham

Copy link
Copy Markdown
Member

As we can coerce any object that is array-like to an ndarray, there should be no need for these explicit ndarray checks. So eliminate them by using ensure_ndarray to get an ndarray object instead.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

Instead of checking to see if `chunk` is an `ndarray`, use
`ensure_ndarray` to get an `ndarray` viewing the underlying data in
`chunk`. This way we know attributes like `dtype` are available and can
be checked easily. Also makes this a bit more friendly with other
array-like types.
Since we are interested in getting an `ndarray` representing the buffer
within `out` for writing into, go ahead and use `ensure_ndarray` to
coerce the underlying buffer into an `ndarray`. This way we can avoid a
needless check and just write into any array-like value for `out` that
is provided.
Appears that `out` can also be a Zarr `Array` or any other array-like
that does not expose a buffer, but does allow writing into. In these
cases `ensure_ndarray` will fail as there is not an underlying buffer
that can be used with an `ndarray`. To also handle this case, catch the
`TypeError` that `ensure_ndarray` will raise in this case and use that
to indicate whether `out` is now an `ndarray` or not. This allows us to
continue to write into arbitrary buffers, but also correctly handle
objects that do not expose buffers.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

@jrbourbeau, if you are around, this could use a review 🙂

Comment threadzarr/core.py

out_is_ndarray = True
try:
out = ensure_ndarray(out)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note that this is already checked by test_get_selection_out. Without this try/except, we get the following TypeError because out is a Zarr Array, which cannot be coerced to an ndarray. This is ok and intentional. So we just carry on without coercing that case and note that we do not have an ndarray when checking later.

Details
_________________________test_get_selection_out____________________________deftest_get_selection_out():
# basic selectionsa=np.arange(1050)
z=zarr.create(shape=1050, chunks=100, dtype=a.dtype)
z[:] =aselections= [
slice(50, 150),
slice(0, 1050),
slice(1, 2),
]
forselectioninselections:
expect=a[selection]
out=zarr.create(shape=expect.shape, chunks=10, dtype=expect.dtype, fill_value=0)
>z.get_basic_selection(selection, out=out)
zarr/tests/test_indexing.py:1036: _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ zarr/core.py:698: inget_basic_selectionfields=fields)
zarr/core.py:740: in_get_basic_selection_ndreturnself._get_selection(indexer=indexer, out=out, fields=fields)
zarr/core.py:1028: in_get_selectiondrop_axes=indexer.drop_axes, fields=fields)
zarr/core.py:1573: in_chunk_getitemout=ensure_ndarray(out)
__ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ buf=<zarr.core.Array (100,) int64>defensure_ndarray(buf):
"""Convenience function to coerce `buf` to a numpy array, if it is not already a numpy array. Parameters ---------- buf : array-like or bytes-like A numpy array or any object exporting a buffer interface. Returns ------- arr : ndarray A numpy array, sharing memory with `buf`. Notes ----- This function will not create a copy under any circumstances, it is guaranteed to return a view on memory exported by `buf`. """ifisinstance(buf, np.ndarray):
# already a numpy arrayarr=bufelifisinstance(buf, array.array) andbuf.typecodein'cu':
# Guard condition, do not support array.array with unicode type, this is# problematic because numpy does not support it on all platforms. Also do not# support char as it was removed in Python 3.raiseTypeError('array.array with char or unicode type is not supported')
else:
# N.B., first take a memoryview to make sure that we subsequently create a# numpy array from a memory buffer with no copyifPY2: # pragma: py3 no covertry:
mem=memoryview(buf)
exceptTypeError:
# on PY2 also check if object exports old-style buffer interfacemem=np.getbuffer(buf)
else: # pragma: py2 no cover>mem=memoryview(buf)
ETypeError: memoryview: abytes-likeobjectisrequired, not'Array'
.tox/py36/lib/python3.6/site-packages/numcodecs/compat.py:74: TypeError

ref: https://travis-ci.org/zarr-developers/zarr-python/jobs/610533967#L2316

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Planning on merging EOD tomorrow if no comments.

@jrbourbeaujrbourbeau left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the PR @jakirkham! Generally these changes seem fine by me. I have one question about when ensure_ndarray raises an error that I've left below.

Comment threadzarr/core.py

# check object encoding
if isinstance(chunk, np.ndarray) and chunk.dtype == object:
if ensure_ndarray(chunk).dtype == object:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is it possible that some non-ndarrays would previously skip over this if block but now fail if ensure_ndarray raises an error? For example:

In [28]: fromnumcodecs.compatimportensure_ndarrayIn [29]: importarrayIn [30]: chunk=array.array('u', 'hello \u2641')
In [31]: ifisinstance(chunk, np.ndarray) andchunk.dtype==object:
...: pass
...:
In [32]: ifensure_ndarray(chunk).dtype==object:
...: pass
...:
---------------------------------------------------------------------------TypeErrorTraceback (mostrecentcalllast)
<ipython-input-32-e1e15cbf81c2>in<module>---->1ifensure_ndarray(chunk).dtype==object:
2pass3~/miniconda/envs/zarr-python-dev/lib/python3.7/site-packages/numcodecs/compat.pyinensure_ndarray(buf)
63# problematic because numpy does not support it on all platforms. Also do not64# support char as it was removed in Python 3.--->65raiseTypeError('array.array with char or unicode type is not supported')
6667else:
TypeError: array.arraywithcharorunicodetypeisnotsupported

To be clear, I'm not sure how likely this is (or if it's even possible) to come up in practice. Do you have a sense for this?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That exception would be expected as we have decided not to work with Python builtin arrays that use character or unicode types. There is some more detailed discussion in this thread.

@jakirkhamjakirkhamNov 12, 2019

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Basically the only exception that we need to worry about is the one coming from memoryview, which will be a TypeError if the object doesn't support the buffer protocol. After that point we have a memoryview. So we know the rest of the code will work (unless NumPy gets a bug ;).

Edit: Sorry was looking at the other change for a second.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For context this function is trying to turn some data into bytes that can be serialized (say to a file on disk).

The filters before this step should be returning something that either is an ndarray or could be coerced to one. This is required by the compressors that follow and the storage layer afterwards. So if ensure_ndarray fails here, then we have invalid data and raising an exception to the user would be appropriate.

@jrbourbeaujrbourbeau left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great, thank you for providing all the additional context surrounding this section. I pushed a commit with a changelog entry. Feel free to merge on green. Thanks @jakirkham!

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Thanks @jrbourbeau for the review and fix! 😄

@jakirkham
jakirkham merged commit 58b1786 into zarr-developers:masterNov 12, 2019
@jakirkham
jakirkham deleted the use_ensure_ndarray_more branch November 12, 2019 18:56
@CarreauCarreau added this to the v2.4 milestone Sep 9, 2020
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jakirkham@jrbourbeau@Carreau
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use ensure_ndarray in a few more places - #506

Merged
jakirkham merged 5 commits into
zarr-developers:masterfrom
jakirkham:use_ensure_ndarray_more
Nov 12, 2019
Merged

Use ensure_ndarray in a few more places#506
jakirkham merged 5 commits into
zarr-developers:masterfrom
jakirkham:use_ensure_ndarray_more

Conversation

@jakirkham

Copy link
Copy Markdown
Member

As we can coerce any object that is array-like to an ndarray, there should be no need for these explicit ndarray checks. So eliminate them by using ensure_ndarray to get an ndarray object instead.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

Instead of checking to see if `chunk` is an `ndarray`, use
`ensure_ndarray` to get an `ndarray` viewing the underlying data in
`chunk`. This way we know attributes like `dtype` are available and can
be checked easily. Also makes this a bit more friendly with other
array-like types.
Since we are interested in getting an `ndarray` representing the buffer
within `out` for writing into, go ahead and use `ensure_ndarray` to
coerce the underlying buffer into an `ndarray`. This way we can avoid a
needless check and just write into any array-like value for `out` that
is provided.
Appears that `out` can also be a Zarr `Array` or any other array-like
that does not expose a buffer, but does allow writing into. In these
cases `ensure_ndarray` will fail as there is not an underlying buffer
that can be used with an `ndarray`. To also handle this case, catch the
`TypeError` that `ensure_ndarray` will raise in this case and use that
to indicate whether `out` is now an `ndarray` or not. This allows us to
continue to write into arbitrary buffers, but also correctly handle
objects that do not expose buffers.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

@jrbourbeau, if you are around, this could use a review 🙂

Comment threadzarr/core.py

out_is_ndarray = True
try:
out = ensure_ndarray(out)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note that this is already checked by test_get_selection_out. Without this try/except, we get the following TypeError because out is a Zarr Array, which cannot be coerced to an ndarray. This is ok and intentional. So we just carry on without coercing that case and note that we do not have an ndarray when checking later.

Details
_________________________test_get_selection_out____________________________deftest_get_selection_out():
# basic selectionsa=np.arange(1050)
z=zarr.create(shape=1050, chunks=100, dtype=a.dtype)
z[:] =aselections= [
slice(50, 150),
slice(0, 1050),
slice(1, 2),
]
forselectioninselections:
expect=a[selection]
out=zarr.create(shape=expect.shape, chunks=10, dtype=expect.dtype, fill_value=0)
>z.get_basic_selection(selection, out=out)
zarr/tests/test_indexing.py:1036: _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ zarr/core.py:698: inget_basic_selectionfields=fields)
zarr/core.py:740: in_get_basic_selection_ndreturnself._get_selection(indexer=indexer, out=out, fields=fields)
zarr/core.py:1028: in_get_selectiondrop_axes=indexer.drop_axes, fields=fields)
zarr/core.py:1573: in_chunk_getitemout=ensure_ndarray(out)
__ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ buf=<zarr.core.Array (100,) int64>defensure_ndarray(buf):
"""Convenience function to coerce `buf` to a numpy array, if it is not already a numpy array. Parameters ---------- buf : array-like or bytes-like A numpy array or any object exporting a buffer interface. Returns ------- arr : ndarray A numpy array, sharing memory with `buf`. Notes ----- This function will not create a copy under any circumstances, it is guaranteed to return a view on memory exported by `buf`. """ifisinstance(buf, np.ndarray):
# already a numpy arrayarr=bufelifisinstance(buf, array.array) andbuf.typecodein'cu':
# Guard condition, do not support array.array with unicode type, this is# problematic because numpy does not support it on all platforms. Also do not# support char as it was removed in Python 3.raiseTypeError('array.array with char or unicode type is not supported')
else:
# N.B., first take a memoryview to make sure that we subsequently create a# numpy array from a memory buffer with no copyifPY2: # pragma: py3 no covertry:
mem=memoryview(buf)
exceptTypeError:
# on PY2 also check if object exports old-style buffer interfacemem=np.getbuffer(buf)
else: # pragma: py2 no cover>mem=memoryview(buf)
ETypeError: memoryview: abytes-likeobjectisrequired, not'Array'
.tox/py36/lib/python3.6/site-packages/numcodecs/compat.py:74: TypeError

ref: https://travis-ci.org/zarr-developers/zarr-python/jobs/610533967#L2316

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Planning on merging EOD tomorrow if no comments.

@jrbourbeaujrbourbeau left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the PR @jakirkham! Generally these changes seem fine by me. I have one question about when ensure_ndarray raises an error that I've left below.

Comment threadzarr/core.py

# check object encoding
if isinstance(chunk, np.ndarray) and chunk.dtype == object:
if ensure_ndarray(chunk).dtype == object:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is it possible that some non-ndarrays would previously skip over this if block but now fail if ensure_ndarray raises an error? For example:

In [28]: fromnumcodecs.compatimportensure_ndarrayIn [29]: importarrayIn [30]: chunk=array.array('u', 'hello \u2641')
In [31]: ifisinstance(chunk, np.ndarray) andchunk.dtype==object:
...: pass
...:
In [32]: ifensure_ndarray(chunk).dtype==object:
...: pass
...:
---------------------------------------------------------------------------TypeErrorTraceback (mostrecentcalllast)
<ipython-input-32-e1e15cbf81c2>in<module>---->1ifensure_ndarray(chunk).dtype==object:
2pass3~/miniconda/envs/zarr-python-dev/lib/python3.7/site-packages/numcodecs/compat.pyinensure_ndarray(buf)
63# problematic because numpy does not support it on all platforms. Also do not64# support char as it was removed in Python 3.--->65raiseTypeError('array.array with char or unicode type is not supported')
6667else:
TypeError: array.arraywithcharorunicodetypeisnotsupported

To be clear, I'm not sure how likely this is (or if it's even possible) to come up in practice. Do you have a sense for this?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That exception would be expected as we have decided not to work with Python builtin arrays that use character or unicode types. There is some more detailed discussion in this thread.

@jakirkhamjakirkhamNov 12, 2019

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Basically the only exception that we need to worry about is the one coming from memoryview, which will be a TypeError if the object doesn't support the buffer protocol. After that point we have a memoryview. So we know the rest of the code will work (unless NumPy gets a bug ;).

Edit: Sorry was looking at the other change for a second.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For context this function is trying to turn some data into bytes that can be serialized (say to a file on disk).

The filters before this step should be returning something that either is an ndarray or could be coerced to one. This is required by the compressors that follow and the storage layer afterwards. So if ensure_ndarray fails here, then we have invalid data and raising an exception to the user would be appropriate.

@jrbourbeaujrbourbeau left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great, thank you for providing all the additional context surrounding this section. I pushed a commit with a changelog entry. Feel free to merge on green. Thanks @jakirkham!

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Thanks @jrbourbeau for the review and fix! 😄

@jakirkham
jakirkham merged commit 58b1786 into zarr-developers:masterNov 12, 2019
@jakirkham
jakirkham deleted the use_ensure_ndarray_more branch November 12, 2019 18:56
@CarreauCarreau added this to the v2.4 milestone Sep 9, 2020
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jakirkham@jrbourbeau@Carreau
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Use ensure_ndarray in a few more places - #506

Merged
jakirkham merged 5 commits into
zarr-developers:masterfrom
jakirkham:use_ensure_ndarray_more
Nov 12, 2019
Merged

Use ensure_ndarray in a few more places#506
jakirkham merged 5 commits into
zarr-developers:masterfrom
jakirkham:use_ensure_ndarray_more

Conversation

@jakirkham

Copy link
Copy Markdown
Member

As we can coerce any object that is array-like to an ndarray, there should be no need for these explicit ndarray checks. So eliminate them by using ensure_ndarray to get an ndarray object instead.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

Instead of checking to see if `chunk` is an `ndarray`, use
`ensure_ndarray` to get an `ndarray` viewing the underlying data in
`chunk`. This way we know attributes like `dtype` are available and can
be checked easily. Also makes this a bit more friendly with other
array-like types.
Since we are interested in getting an `ndarray` representing the buffer
within `out` for writing into, go ahead and use `ensure_ndarray` to
coerce the underlying buffer into an `ndarray`. This way we can avoid a
needless check and just write into any array-like value for `out` that
is provided.
Appears that `out` can also be a Zarr `Array` or any other array-like
that does not expose a buffer, but does allow writing into. In these
cases `ensure_ndarray` will fail as there is not an underlying buffer
that can be used with an `ndarray`. To also handle this case, catch the
`TypeError` that `ensure_ndarray` will raise in this case and use that
to indicate whether `out` is now an `ndarray` or not. This allows us to
continue to write into arbitrary buffers, but also correctly handle
objects that do not expose buffers.
@jakirkham

Copy link
Copy Markdown
MemberAuthor

@jrbourbeau, if you are around, this could use a review 🙂

Comment threadzarr/core.py

out_is_ndarray = True
try:
out = ensure_ndarray(out)

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note that this is already checked by test_get_selection_out. Without this try/except, we get the following TypeError because out is a Zarr Array, which cannot be coerced to an ndarray. This is ok and intentional. So we just carry on without coercing that case and note that we do not have an ndarray when checking later.

Details
_________________________test_get_selection_out____________________________deftest_get_selection_out():
# basic selectionsa=np.arange(1050)
z=zarr.create(shape=1050, chunks=100, dtype=a.dtype)
z[:] =aselections= [
slice(50, 150),
slice(0, 1050),
slice(1, 2),
]
forselectioninselections:
expect=a[selection]
out=zarr.create(shape=expect.shape, chunks=10, dtype=expect.dtype, fill_value=0)
>z.get_basic_selection(selection, out=out)
zarr/tests/test_indexing.py:1036: _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ zarr/core.py:698: inget_basic_selectionfields=fields)
zarr/core.py:740: in_get_basic_selection_ndreturnself._get_selection(indexer=indexer, out=out, fields=fields)
zarr/core.py:1028: in_get_selectiondrop_axes=indexer.drop_axes, fields=fields)
zarr/core.py:1573: in_chunk_getitemout=ensure_ndarray(out)
__ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ buf=<zarr.core.Array (100,) int64>defensure_ndarray(buf):
"""Convenience function to coerce `buf` to a numpy array, if it is not already a numpy array. Parameters ---------- buf : array-like or bytes-like A numpy array or any object exporting a buffer interface. Returns ------- arr : ndarray A numpy array, sharing memory with `buf`. Notes ----- This function will not create a copy under any circumstances, it is guaranteed to return a view on memory exported by `buf`. """ifisinstance(buf, np.ndarray):
# already a numpy arrayarr=bufelifisinstance(buf, array.array) andbuf.typecodein'cu':
# Guard condition, do not support array.array with unicode type, this is# problematic because numpy does not support it on all platforms. Also do not# support char as it was removed in Python 3.raiseTypeError('array.array with char or unicode type is not supported')
else:
# N.B., first take a memoryview to make sure that we subsequently create a# numpy array from a memory buffer with no copyifPY2: # pragma: py3 no covertry:
mem=memoryview(buf)
exceptTypeError:
# on PY2 also check if object exports old-style buffer interfacemem=np.getbuffer(buf)
else: # pragma: py2 no cover>mem=memoryview(buf)
ETypeError: memoryview: abytes-likeobjectisrequired, not'Array'
.tox/py36/lib/python3.6/site-packages/numcodecs/compat.py:74: TypeError

ref: https://travis-ci.org/zarr-developers/zarr-python/jobs/610533967#L2316

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Planning on merging EOD tomorrow if no comments.

@jrbourbeaujrbourbeau left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the PR @jakirkham! Generally these changes seem fine by me. I have one question about when ensure_ndarray raises an error that I've left below.

Comment threadzarr/core.py

# check object encoding
if isinstance(chunk, np.ndarray) and chunk.dtype == object:
if ensure_ndarray(chunk).dtype == object:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is it possible that some non-ndarrays would previously skip over this if block but now fail if ensure_ndarray raises an error? For example:

In [28]: fromnumcodecs.compatimportensure_ndarrayIn [29]: importarrayIn [30]: chunk=array.array('u', 'hello \u2641')
In [31]: ifisinstance(chunk, np.ndarray) andchunk.dtype==object:
...: pass
...:
In [32]: ifensure_ndarray(chunk).dtype==object:
...: pass
...:
---------------------------------------------------------------------------TypeErrorTraceback (mostrecentcalllast)
<ipython-input-32-e1e15cbf81c2>in<module>---->1ifensure_ndarray(chunk).dtype==object:
2pass3~/miniconda/envs/zarr-python-dev/lib/python3.7/site-packages/numcodecs/compat.pyinensure_ndarray(buf)
63# problematic because numpy does not support it on all platforms. Also do not64# support char as it was removed in Python 3.--->65raiseTypeError('array.array with char or unicode type is not supported')
6667else:
TypeError: array.arraywithcharorunicodetypeisnotsupported

To be clear, I'm not sure how likely this is (or if it's even possible) to come up in practice. Do you have a sense for this?

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That exception would be expected as we have decided not to work with Python builtin arrays that use character or unicode types. There is some more detailed discussion in this thread.

@jakirkhamjakirkhamNov 12, 2019

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Basically the only exception that we need to worry about is the one coming from memoryview, which will be a TypeError if the object doesn't support the buffer protocol. After that point we have a memoryview. So we know the rest of the code will work (unless NumPy gets a bug ;).

Edit: Sorry was looking at the other change for a second.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For context this function is trying to turn some data into bytes that can be serialized (say to a file on disk).

The filters before this step should be returning something that either is an ndarray or could be coerced to one. This is required by the compressors that follow and the storage layer afterwards. So if ensure_ndarray fails here, then we have invalid data and raising an exception to the user would be appropriate.

@jrbourbeaujrbourbeau left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great, thank you for providing all the additional context surrounding this section. I pushed a commit with a changelog entry. Feel free to merge on green. Thanks @jakirkham!

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Thanks @jrbourbeau for the review and fix! 😄

@jakirkham
jakirkham merged commit 58b1786 into zarr-developers:masterNov 12, 2019
@jakirkham
jakirkham deleted the use_ensure_ndarray_more branch November 12, 2019 18:56
@CarreauCarreau added this to the v2.4 milestone Sep 9, 2020
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jakirkham@jrbourbeau@Carreau