Bump Numcodecs requirement to 0.6.2 - #352

Merged
alimanfoo merged 15 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.2
Dec 4, 2018
Merged

Bump Numcodecs requirement to 0.6.2#352
alimanfoo merged 15 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.2

Conversation

@jakirkham

@jakirkhamjakirkham commented Dec 4, 2018

Copy link
Copy Markdown
Member

Follow-up to PR ( #347 )

Fixes#324

As there are some critical fixes and needed features in the latest Numcodecs, this bumps our lower bound to the latest version. Pulls most of the content from the aforementioned PR. Drops some commits that have been broken out into subsequent PRs to keep this focused on the Numcodecs upgrade. Does some refactoring and cleanup of internal functions thanks to utility functions now included in Numcodecs.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

jakirkhamand others added 12 commits November 30, 2018 12:18
Previously MsgPack was turning bytes objects to unicode objects when
round-tripping them. However this has been fixed in the latest version
of Numcodecs. So correct this test now that MsgPack is working
correctly.
As we already ensured the `chunk` is an `ndarray` viewing the original
data, there is no need for us to do that here as well. Plus the checks
performed by `ensure_contiguous_ndarray` are not needed for our use case
here. Particularly as we have already handled the unusual type cases
above. We also don't need to constrain the buffer size. As such the only
thing we really need is to flatten the array and make it contiguous,
which is what we handle here directly.
As both the expected `object` case and the non-`object` case perform a
`reshape` to flatten the data, go ahead and refactor that out of both
cases and handle it generally. Simplifies the code a bit.
As refactoring of the `reshape` step has effectively dropped the
expected `object` type case, the checks for different types is a little
more complicated than needed. To fix this, basically invert and swap the
case ordering. This way we can handle all generally expected types first
and simply cast them. Then we can raise if an `object` type shows up and
is unexpected.
As Numcodecs now includes a very versatile and effective `ensure_bytes`
function, there is no need to define our own in `zarr.storage` as well.
So go ahead and drop it.
Make use of Numcodecs' `ensure_contiguous_ndarray` to take `ndarray`
views onto buffers to be stored in a few cases so as to reshape them and
avoid a copy (thanks to the buffer protocol). This ensures that
datetime/timedeltas are handled by default. Also catches things like
object arrays. Finally this handles flattening the array if needed.
All-in-all this gets as close to a `bytes` object as possible while not
copying and doing its best to preserve type information while
constructing something that fits the buffer protocol.
Rewrite `buffer_size` to just use Numcodecs' `ensure_ndarray` to get an
`ndarray` that views the data. Once the `ndarray` is gotten, all that is
needed is to access its `nbytes` member, which returns the number of
bytes that it takes up.
If the data is already a `str` instance, turn `ensure_str` into a no-op.
For all other cases, make use of Numcodecs' `ensure_bytes` to aid
`ensure_str` in coercing data through the buffer protocol. If we are on
Python 3, then decode the `bytes` object to a `str`.
@jakirkhamjakirkham mentioned this pull request Dec 4, 2018
7 tasks
As Blosc got upgraded and it contained an upgrade of Zstd, the results
changed a little bit for this example. So update them accordingly.
Should fix the doctest failure.
@jakirkhamjakirkham added this to the v2.3 milestone Dec 4, 2018

@alimanfooalimanfoo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @jakirkham, all looks good.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Looks like the line not covered is a fallback for when file removal fails. Given PR ( #327 ) obviates that, maybe we should merge that PR and drop that fallback. Thoughts?

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Made PR ( #355 ) to simply ignore coverage on that group of lines.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Alright, think this is ready now. 😄

@alimanfoo

Copy link
Copy Markdown
Member

Awesome, merging...

@alimanfoo
alimanfoo merged commit c4427a4 into zarr-developers:masterDec 4, 2018
@alimanfooalimanfoo mentioned this pull request Dec 4, 2018
7 tasks
@jakirkham
jakirkham deleted the use_numcodecs_0.6.2 branch December 5, 2018 02:48
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Missed a bit of code that could be simplified with ensure_ndarray. Addressing that in PR ( #360 ).

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Also dropping a workaround for older Numcodecs. ( #361 ) Please let me know if there are more of these.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

MsgPack codec is broken for array of bytes objects

2 participants

@jakirkham@alimanfoo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Bump Numcodecs requirement to 0.6.2 - #352

Merged
alimanfoo merged 15 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.2
Dec 4, 2018
Merged

Bump Numcodecs requirement to 0.6.2#352
alimanfoo merged 15 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.2

Conversation

@jakirkham

@jakirkhamjakirkham commented Dec 4, 2018

Copy link
Copy Markdown
Member

Follow-up to PR ( #347 )

Fixes#324

As there are some critical fixes and needed features in the latest Numcodecs, this bumps our lower bound to the latest version. Pulls most of the content from the aforementioned PR. Drops some commits that have been broken out into subsequent PRs to keep this focused on the Numcodecs upgrade. Does some refactoring and cleanup of internal functions thanks to utility functions now included in Numcodecs.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

jakirkhamand others added 12 commits November 30, 2018 12:18
Previously MsgPack was turning bytes objects to unicode objects when
round-tripping them. However this has been fixed in the latest version
of Numcodecs. So correct this test now that MsgPack is working
correctly.
As we already ensured the `chunk` is an `ndarray` viewing the original
data, there is no need for us to do that here as well. Plus the checks
performed by `ensure_contiguous_ndarray` are not needed for our use case
here. Particularly as we have already handled the unusual type cases
above. We also don't need to constrain the buffer size. As such the only
thing we really need is to flatten the array and make it contiguous,
which is what we handle here directly.
As both the expected `object` case and the non-`object` case perform a
`reshape` to flatten the data, go ahead and refactor that out of both
cases and handle it generally. Simplifies the code a bit.
As refactoring of the `reshape` step has effectively dropped the
expected `object` type case, the checks for different types is a little
more complicated than needed. To fix this, basically invert and swap the
case ordering. This way we can handle all generally expected types first
and simply cast them. Then we can raise if an `object` type shows up and
is unexpected.
As Numcodecs now includes a very versatile and effective `ensure_bytes`
function, there is no need to define our own in `zarr.storage` as well.
So go ahead and drop it.
Make use of Numcodecs' `ensure_contiguous_ndarray` to take `ndarray`
views onto buffers to be stored in a few cases so as to reshape them and
avoid a copy (thanks to the buffer protocol). This ensures that
datetime/timedeltas are handled by default. Also catches things like
object arrays. Finally this handles flattening the array if needed.
All-in-all this gets as close to a `bytes` object as possible while not
copying and doing its best to preserve type information while
constructing something that fits the buffer protocol.
Rewrite `buffer_size` to just use Numcodecs' `ensure_ndarray` to get an
`ndarray` that views the data. Once the `ndarray` is gotten, all that is
needed is to access its `nbytes` member, which returns the number of
bytes that it takes up.
If the data is already a `str` instance, turn `ensure_str` into a no-op.
For all other cases, make use of Numcodecs' `ensure_bytes` to aid
`ensure_str` in coercing data through the buffer protocol. If we are on
Python 3, then decode the `bytes` object to a `str`.
@jakirkhamjakirkham mentioned this pull request Dec 4, 2018
7 tasks
As Blosc got upgraded and it contained an upgrade of Zstd, the results
changed a little bit for this example. So update them accordingly.
Should fix the doctest failure.
@jakirkhamjakirkham added this to the v2.3 milestone Dec 4, 2018

@alimanfooalimanfoo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @jakirkham, all looks good.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Looks like the line not covered is a fallback for when file removal fails. Given PR ( #327 ) obviates that, maybe we should merge that PR and drop that fallback. Thoughts?

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Made PR ( #355 ) to simply ignore coverage on that group of lines.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Alright, think this is ready now. 😄

@alimanfoo

Copy link
Copy Markdown
Member

Awesome, merging...

@alimanfoo
alimanfoo merged commit c4427a4 into zarr-developers:masterDec 4, 2018
@alimanfooalimanfoo mentioned this pull request Dec 4, 2018
7 tasks
@jakirkham
jakirkham deleted the use_numcodecs_0.6.2 branch December 5, 2018 02:48
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Missed a bit of code that could be simplified with ensure_ndarray. Addressing that in PR ( #360 ).

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Also dropping a workaround for older Numcodecs. ( #361 ) Please let me know if there are more of these.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

MsgPack codec is broken for array of bytes objects

2 participants

@jakirkham@alimanfoo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Bump Numcodecs requirement to 0.6.2 - #352

Merged
alimanfoo merged 15 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.2
Dec 4, 2018
Merged

Bump Numcodecs requirement to 0.6.2#352
alimanfoo merged 15 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.2

Conversation

@jakirkham

@jakirkhamjakirkham commented Dec 4, 2018

Copy link
Copy Markdown
Member

Follow-up to PR ( #347 )

Fixes#324

As there are some critical fixes and needed features in the latest Numcodecs, this bumps our lower bound to the latest version. Pulls most of the content from the aforementioned PR. Drops some commits that have been broken out into subsequent PRs to keep this focused on the Numcodecs upgrade. Does some refactoring and cleanup of internal functions thanks to utility functions now included in Numcodecs.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

jakirkhamand others added 12 commits November 30, 2018 12:18
Previously MsgPack was turning bytes objects to unicode objects when
round-tripping them. However this has been fixed in the latest version
of Numcodecs. So correct this test now that MsgPack is working
correctly.
As we already ensured the `chunk` is an `ndarray` viewing the original
data, there is no need for us to do that here as well. Plus the checks
performed by `ensure_contiguous_ndarray` are not needed for our use case
here. Particularly as we have already handled the unusual type cases
above. We also don't need to constrain the buffer size. As such the only
thing we really need is to flatten the array and make it contiguous,
which is what we handle here directly.
As both the expected `object` case and the non-`object` case perform a
`reshape` to flatten the data, go ahead and refactor that out of both
cases and handle it generally. Simplifies the code a bit.
As refactoring of the `reshape` step has effectively dropped the
expected `object` type case, the checks for different types is a little
more complicated than needed. To fix this, basically invert and swap the
case ordering. This way we can handle all generally expected types first
and simply cast them. Then we can raise if an `object` type shows up and
is unexpected.
As Numcodecs now includes a very versatile and effective `ensure_bytes`
function, there is no need to define our own in `zarr.storage` as well.
So go ahead and drop it.
Make use of Numcodecs' `ensure_contiguous_ndarray` to take `ndarray`
views onto buffers to be stored in a few cases so as to reshape them and
avoid a copy (thanks to the buffer protocol). This ensures that
datetime/timedeltas are handled by default. Also catches things like
object arrays. Finally this handles flattening the array if needed.
All-in-all this gets as close to a `bytes` object as possible while not
copying and doing its best to preserve type information while
constructing something that fits the buffer protocol.
Rewrite `buffer_size` to just use Numcodecs' `ensure_ndarray` to get an
`ndarray` that views the data. Once the `ndarray` is gotten, all that is
needed is to access its `nbytes` member, which returns the number of
bytes that it takes up.
If the data is already a `str` instance, turn `ensure_str` into a no-op.
For all other cases, make use of Numcodecs' `ensure_bytes` to aid
`ensure_str` in coercing data through the buffer protocol. If we are on
Python 3, then decode the `bytes` object to a `str`.
@jakirkhamjakirkham mentioned this pull request Dec 4, 2018
7 tasks
As Blosc got upgraded and it contained an upgrade of Zstd, the results
changed a little bit for this example. So update them accordingly.
Should fix the doctest failure.
@jakirkhamjakirkham added this to the v2.3 milestone Dec 4, 2018

@alimanfooalimanfoo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @jakirkham, all looks good.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Looks like the line not covered is a fallback for when file removal fails. Given PR ( #327 ) obviates that, maybe we should merge that PR and drop that fallback. Thoughts?

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Made PR ( #355 ) to simply ignore coverage on that group of lines.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Alright, think this is ready now. 😄

@alimanfoo

Copy link
Copy Markdown
Member

Awesome, merging...

@alimanfoo
alimanfoo merged commit c4427a4 into zarr-developers:masterDec 4, 2018
@alimanfooalimanfoo mentioned this pull request Dec 4, 2018
7 tasks
@jakirkham
jakirkham deleted the use_numcodecs_0.6.2 branch December 5, 2018 02:48
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Missed a bit of code that could be simplified with ensure_ndarray. Addressing that in PR ( #360 ).

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Also dropping a workaround for older Numcodecs. ( #361 ) Please let me know if there are more of these.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

MsgPack codec is broken for array of bytes objects

2 participants

@jakirkham@alimanfoo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Bump Numcodecs requirement to 0.6.2 - #352

Merged
alimanfoo merged 15 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.2
Dec 4, 2018
Merged

Bump Numcodecs requirement to 0.6.2#352
alimanfoo merged 15 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.2

Conversation

@jakirkham

@jakirkhamjakirkham commented Dec 4, 2018

Copy link
Copy Markdown
Member

Follow-up to PR ( #347 )

Fixes#324

As there are some critical fixes and needed features in the latest Numcodecs, this bumps our lower bound to the latest version. Pulls most of the content from the aforementioned PR. Drops some commits that have been broken out into subsequent PRs to keep this focused on the Numcodecs upgrade. Does some refactoring and cleanup of internal functions thanks to utility functions now included in Numcodecs.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

jakirkhamand others added 12 commits November 30, 2018 12:18
Previously MsgPack was turning bytes objects to unicode objects when
round-tripping them. However this has been fixed in the latest version
of Numcodecs. So correct this test now that MsgPack is working
correctly.
As we already ensured the `chunk` is an `ndarray` viewing the original
data, there is no need for us to do that here as well. Plus the checks
performed by `ensure_contiguous_ndarray` are not needed for our use case
here. Particularly as we have already handled the unusual type cases
above. We also don't need to constrain the buffer size. As such the only
thing we really need is to flatten the array and make it contiguous,
which is what we handle here directly.
As both the expected `object` case and the non-`object` case perform a
`reshape` to flatten the data, go ahead and refactor that out of both
cases and handle it generally. Simplifies the code a bit.
As refactoring of the `reshape` step has effectively dropped the
expected `object` type case, the checks for different types is a little
more complicated than needed. To fix this, basically invert and swap the
case ordering. This way we can handle all generally expected types first
and simply cast them. Then we can raise if an `object` type shows up and
is unexpected.
As Numcodecs now includes a very versatile and effective `ensure_bytes`
function, there is no need to define our own in `zarr.storage` as well.
So go ahead and drop it.
Make use of Numcodecs' `ensure_contiguous_ndarray` to take `ndarray`
views onto buffers to be stored in a few cases so as to reshape them and
avoid a copy (thanks to the buffer protocol). This ensures that
datetime/timedeltas are handled by default. Also catches things like
object arrays. Finally this handles flattening the array if needed.
All-in-all this gets as close to a `bytes` object as possible while not
copying and doing its best to preserve type information while
constructing something that fits the buffer protocol.
Rewrite `buffer_size` to just use Numcodecs' `ensure_ndarray` to get an
`ndarray` that views the data. Once the `ndarray` is gotten, all that is
needed is to access its `nbytes` member, which returns the number of
bytes that it takes up.
If the data is already a `str` instance, turn `ensure_str` into a no-op.
For all other cases, make use of Numcodecs' `ensure_bytes` to aid
`ensure_str` in coercing data through the buffer protocol. If we are on
Python 3, then decode the `bytes` object to a `str`.
@jakirkhamjakirkham mentioned this pull request Dec 4, 2018
7 tasks
As Blosc got upgraded and it contained an upgrade of Zstd, the results
changed a little bit for this example. So update them accordingly.
Should fix the doctest failure.
@jakirkhamjakirkham added this to the v2.3 milestone Dec 4, 2018

@alimanfooalimanfoo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @jakirkham, all looks good.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Looks like the line not covered is a fallback for when file removal fails. Given PR ( #327 ) obviates that, maybe we should merge that PR and drop that fallback. Thoughts?

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Made PR ( #355 ) to simply ignore coverage on that group of lines.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Alright, think this is ready now. 😄

@alimanfoo

Copy link
Copy Markdown
Member

Awesome, merging...

@alimanfoo
alimanfoo merged commit c4427a4 into zarr-developers:masterDec 4, 2018
@alimanfooalimanfoo mentioned this pull request Dec 4, 2018
7 tasks
@jakirkham
jakirkham deleted the use_numcodecs_0.6.2 branch December 5, 2018 02:48
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Missed a bit of code that could be simplified with ensure_ndarray. Addressing that in PR ( #360 ).

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Also dropping a workaround for older Numcodecs. ( #361 ) Please let me know if there are more of these.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

MsgPack codec is broken for array of bytes objects

2 participants

@jakirkham@alimanfoo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Bump Numcodecs requirement to 0.6.2 - #352

Merged
alimanfoo merged 15 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.2
Dec 4, 2018
Merged

Bump Numcodecs requirement to 0.6.2#352
alimanfoo merged 15 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.2

Conversation

@jakirkham

@jakirkhamjakirkham commented Dec 4, 2018

Copy link
Copy Markdown
Member

Follow-up to PR ( #347 )

Fixes#324

As there are some critical fixes and needed features in the latest Numcodecs, this bumps our lower bound to the latest version. Pulls most of the content from the aforementioned PR. Drops some commits that have been broken out into subsequent PRs to keep this focused on the Numcodecs upgrade. Does some refactoring and cleanup of internal functions thanks to utility functions now included in Numcodecs.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

jakirkhamand others added 12 commits November 30, 2018 12:18
Previously MsgPack was turning bytes objects to unicode objects when
round-tripping them. However this has been fixed in the latest version
of Numcodecs. So correct this test now that MsgPack is working
correctly.
As we already ensured the `chunk` is an `ndarray` viewing the original
data, there is no need for us to do that here as well. Plus the checks
performed by `ensure_contiguous_ndarray` are not needed for our use case
here. Particularly as we have already handled the unusual type cases
above. We also don't need to constrain the buffer size. As such the only
thing we really need is to flatten the array and make it contiguous,
which is what we handle here directly.
As both the expected `object` case and the non-`object` case perform a
`reshape` to flatten the data, go ahead and refactor that out of both
cases and handle it generally. Simplifies the code a bit.
As refactoring of the `reshape` step has effectively dropped the
expected `object` type case, the checks for different types is a little
more complicated than needed. To fix this, basically invert and swap the
case ordering. This way we can handle all generally expected types first
and simply cast them. Then we can raise if an `object` type shows up and
is unexpected.
As Numcodecs now includes a very versatile and effective `ensure_bytes`
function, there is no need to define our own in `zarr.storage` as well.
So go ahead and drop it.
Make use of Numcodecs' `ensure_contiguous_ndarray` to take `ndarray`
views onto buffers to be stored in a few cases so as to reshape them and
avoid a copy (thanks to the buffer protocol). This ensures that
datetime/timedeltas are handled by default. Also catches things like
object arrays. Finally this handles flattening the array if needed.
All-in-all this gets as close to a `bytes` object as possible while not
copying and doing its best to preserve type information while
constructing something that fits the buffer protocol.
Rewrite `buffer_size` to just use Numcodecs' `ensure_ndarray` to get an
`ndarray` that views the data. Once the `ndarray` is gotten, all that is
needed is to access its `nbytes` member, which returns the number of
bytes that it takes up.
If the data is already a `str` instance, turn `ensure_str` into a no-op.
For all other cases, make use of Numcodecs' `ensure_bytes` to aid
`ensure_str` in coercing data through the buffer protocol. If we are on
Python 3, then decode the `bytes` object to a `str`.
@jakirkhamjakirkham mentioned this pull request Dec 4, 2018
7 tasks
As Blosc got upgraded and it contained an upgrade of Zstd, the results
changed a little bit for this example. So update them accordingly.
Should fix the doctest failure.
@jakirkhamjakirkham added this to the v2.3 milestone Dec 4, 2018

@alimanfooalimanfoo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @jakirkham, all looks good.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Looks like the line not covered is a fallback for when file removal fails. Given PR ( #327 ) obviates that, maybe we should merge that PR and drop that fallback. Thoughts?

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Made PR ( #355 ) to simply ignore coverage on that group of lines.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Alright, think this is ready now. 😄

@alimanfoo

Copy link
Copy Markdown
Member

Awesome, merging...

@alimanfoo
alimanfoo merged commit c4427a4 into zarr-developers:masterDec 4, 2018
@alimanfooalimanfoo mentioned this pull request Dec 4, 2018
7 tasks
@jakirkham
jakirkham deleted the use_numcodecs_0.6.2 branch December 5, 2018 02:48
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Missed a bit of code that could be simplified with ensure_ndarray. Addressing that in PR ( #360 ).

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Also dropping a workaround for older Numcodecs. ( #361 ) Please let me know if there are more of these.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

MsgPack codec is broken for array of bytes objects

2 participants

@jakirkham@alimanfoo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Bump Numcodecs requirement to 0.6.2 - #352

Merged
alimanfoo merged 15 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.2
Dec 4, 2018
Merged

Bump Numcodecs requirement to 0.6.2#352
alimanfoo merged 15 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.2

Conversation

@jakirkham

@jakirkhamjakirkham commented Dec 4, 2018

Copy link
Copy Markdown
Member

Follow-up to PR ( #347 )

Fixes#324

As there are some critical fixes and needed features in the latest Numcodecs, this bumps our lower bound to the latest version. Pulls most of the content from the aforementioned PR. Drops some commits that have been broken out into subsequent PRs to keep this focused on the Numcodecs upgrade. Does some refactoring and cleanup of internal functions thanks to utility functions now included in Numcodecs.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

jakirkhamand others added 12 commits November 30, 2018 12:18
Previously MsgPack was turning bytes objects to unicode objects when
round-tripping them. However this has been fixed in the latest version
of Numcodecs. So correct this test now that MsgPack is working
correctly.
As we already ensured the `chunk` is an `ndarray` viewing the original
data, there is no need for us to do that here as well. Plus the checks
performed by `ensure_contiguous_ndarray` are not needed for our use case
here. Particularly as we have already handled the unusual type cases
above. We also don't need to constrain the buffer size. As such the only
thing we really need is to flatten the array and make it contiguous,
which is what we handle here directly.
As both the expected `object` case and the non-`object` case perform a
`reshape` to flatten the data, go ahead and refactor that out of both
cases and handle it generally. Simplifies the code a bit.
As refactoring of the `reshape` step has effectively dropped the
expected `object` type case, the checks for different types is a little
more complicated than needed. To fix this, basically invert and swap the
case ordering. This way we can handle all generally expected types first
and simply cast them. Then we can raise if an `object` type shows up and
is unexpected.
As Numcodecs now includes a very versatile and effective `ensure_bytes`
function, there is no need to define our own in `zarr.storage` as well.
So go ahead and drop it.
Make use of Numcodecs' `ensure_contiguous_ndarray` to take `ndarray`
views onto buffers to be stored in a few cases so as to reshape them and
avoid a copy (thanks to the buffer protocol). This ensures that
datetime/timedeltas are handled by default. Also catches things like
object arrays. Finally this handles flattening the array if needed.
All-in-all this gets as close to a `bytes` object as possible while not
copying and doing its best to preserve type information while
constructing something that fits the buffer protocol.
Rewrite `buffer_size` to just use Numcodecs' `ensure_ndarray` to get an
`ndarray` that views the data. Once the `ndarray` is gotten, all that is
needed is to access its `nbytes` member, which returns the number of
bytes that it takes up.
If the data is already a `str` instance, turn `ensure_str` into a no-op.
For all other cases, make use of Numcodecs' `ensure_bytes` to aid
`ensure_str` in coercing data through the buffer protocol. If we are on
Python 3, then decode the `bytes` object to a `str`.
@jakirkhamjakirkham mentioned this pull request Dec 4, 2018
7 tasks
As Blosc got upgraded and it contained an upgrade of Zstd, the results
changed a little bit for this example. So update them accordingly.
Should fix the doctest failure.
@jakirkhamjakirkham added this to the v2.3 milestone Dec 4, 2018

@alimanfooalimanfoo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @jakirkham, all looks good.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Looks like the line not covered is a fallback for when file removal fails. Given PR ( #327 ) obviates that, maybe we should merge that PR and drop that fallback. Thoughts?

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Made PR ( #355 ) to simply ignore coverage on that group of lines.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Alright, think this is ready now. 😄

@alimanfoo

Copy link
Copy Markdown
Member

Awesome, merging...

@alimanfoo
alimanfoo merged commit c4427a4 into zarr-developers:masterDec 4, 2018
@alimanfooalimanfoo mentioned this pull request Dec 4, 2018
7 tasks
@jakirkham
jakirkham deleted the use_numcodecs_0.6.2 branch December 5, 2018 02:48
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Missed a bit of code that could be simplified with ensure_ndarray. Addressing that in PR ( #360 ).

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Also dropping a workaround for older Numcodecs. ( #361 ) Please let me know if there are more of these.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

MsgPack codec is broken for array of bytes objects

2 participants

@jakirkham@alimanfoo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Bump Numcodecs requirement to 0.6.2 - #352

Merged
alimanfoo merged 15 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.2
Dec 4, 2018
Merged

Bump Numcodecs requirement to 0.6.2#352
alimanfoo merged 15 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.2

Conversation

@jakirkham

@jakirkhamjakirkham commented Dec 4, 2018

Copy link
Copy Markdown
Member

Follow-up to PR ( #347 )

Fixes#324

As there are some critical fixes and needed features in the latest Numcodecs, this bumps our lower bound to the latest version. Pulls most of the content from the aforementioned PR. Drops some commits that have been broken out into subsequent PRs to keep this focused on the Numcodecs upgrade. Does some refactoring and cleanup of internal functions thanks to utility functions now included in Numcodecs.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

jakirkhamand others added 12 commits November 30, 2018 12:18
Previously MsgPack was turning bytes objects to unicode objects when
round-tripping them. However this has been fixed in the latest version
of Numcodecs. So correct this test now that MsgPack is working
correctly.
As we already ensured the `chunk` is an `ndarray` viewing the original
data, there is no need for us to do that here as well. Plus the checks
performed by `ensure_contiguous_ndarray` are not needed for our use case
here. Particularly as we have already handled the unusual type cases
above. We also don't need to constrain the buffer size. As such the only
thing we really need is to flatten the array and make it contiguous,
which is what we handle here directly.
As both the expected `object` case and the non-`object` case perform a
`reshape` to flatten the data, go ahead and refactor that out of both
cases and handle it generally. Simplifies the code a bit.
As refactoring of the `reshape` step has effectively dropped the
expected `object` type case, the checks for different types is a little
more complicated than needed. To fix this, basically invert and swap the
case ordering. This way we can handle all generally expected types first
and simply cast them. Then we can raise if an `object` type shows up and
is unexpected.
As Numcodecs now includes a very versatile and effective `ensure_bytes`
function, there is no need to define our own in `zarr.storage` as well.
So go ahead and drop it.
Make use of Numcodecs' `ensure_contiguous_ndarray` to take `ndarray`
views onto buffers to be stored in a few cases so as to reshape them and
avoid a copy (thanks to the buffer protocol). This ensures that
datetime/timedeltas are handled by default. Also catches things like
object arrays. Finally this handles flattening the array if needed.
All-in-all this gets as close to a `bytes` object as possible while not
copying and doing its best to preserve type information while
constructing something that fits the buffer protocol.
Rewrite `buffer_size` to just use Numcodecs' `ensure_ndarray` to get an
`ndarray` that views the data. Once the `ndarray` is gotten, all that is
needed is to access its `nbytes` member, which returns the number of
bytes that it takes up.
If the data is already a `str` instance, turn `ensure_str` into a no-op.
For all other cases, make use of Numcodecs' `ensure_bytes` to aid
`ensure_str` in coercing data through the buffer protocol. If we are on
Python 3, then decode the `bytes` object to a `str`.
@jakirkhamjakirkham mentioned this pull request Dec 4, 2018
7 tasks
As Blosc got upgraded and it contained an upgrade of Zstd, the results
changed a little bit for this example. So update them accordingly.
Should fix the doctest failure.
@jakirkhamjakirkham added this to the v2.3 milestone Dec 4, 2018

@alimanfooalimanfoo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @jakirkham, all looks good.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Looks like the line not covered is a fallback for when file removal fails. Given PR ( #327 ) obviates that, maybe we should merge that PR and drop that fallback. Thoughts?

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Made PR ( #355 ) to simply ignore coverage on that group of lines.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Alright, think this is ready now. 😄

@alimanfoo

Copy link
Copy Markdown
Member

Awesome, merging...

@alimanfoo
alimanfoo merged commit c4427a4 into zarr-developers:masterDec 4, 2018
@alimanfooalimanfoo mentioned this pull request Dec 4, 2018
7 tasks
@jakirkham
jakirkham deleted the use_numcodecs_0.6.2 branch December 5, 2018 02:48
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Missed a bit of code that could be simplified with ensure_ndarray. Addressing that in PR ( #360 ).

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Also dropping a workaround for older Numcodecs. ( #361 ) Please let me know if there are more of these.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

MsgPack codec is broken for array of bytes objects

2 participants

@jakirkham@alimanfoo
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Bump Numcodecs requirement to 0.6.2 - #352

Merged
alimanfoo merged 15 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.2
Dec 4, 2018
Merged

Bump Numcodecs requirement to 0.6.2#352
alimanfoo merged 15 commits into
zarr-developers:masterfrom
jakirkham:use_numcodecs_0.6.2

Conversation

@jakirkham

@jakirkhamjakirkham commented Dec 4, 2018

Copy link
Copy Markdown
Member

Follow-up to PR ( #347 )

Fixes#324

As there are some critical fixes and needed features in the latest Numcodecs, this bumps our lower bound to the latest version. Pulls most of the content from the aforementioned PR. Drops some commits that have been broken out into subsequent PRs to keep this focused on the Numcodecs upgrade. Does some refactoring and cleanup of internal functions thanks to utility functions now included in Numcodecs.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • Docs build locally (e.g., run tox -e docs)
  • AppVeyor and Travis CI passes
  • Test coverage is 100% (Coveralls passes)

jakirkhamand others added 12 commits November 30, 2018 12:18
Previously MsgPack was turning bytes objects to unicode objects when
round-tripping them. However this has been fixed in the latest version
of Numcodecs. So correct this test now that MsgPack is working
correctly.
As we already ensured the `chunk` is an `ndarray` viewing the original
data, there is no need for us to do that here as well. Plus the checks
performed by `ensure_contiguous_ndarray` are not needed for our use case
here. Particularly as we have already handled the unusual type cases
above. We also don't need to constrain the buffer size. As such the only
thing we really need is to flatten the array and make it contiguous,
which is what we handle here directly.
As both the expected `object` case and the non-`object` case perform a
`reshape` to flatten the data, go ahead and refactor that out of both
cases and handle it generally. Simplifies the code a bit.
As refactoring of the `reshape` step has effectively dropped the
expected `object` type case, the checks for different types is a little
more complicated than needed. To fix this, basically invert and swap the
case ordering. This way we can handle all generally expected types first
and simply cast them. Then we can raise if an `object` type shows up and
is unexpected.
As Numcodecs now includes a very versatile and effective `ensure_bytes`
function, there is no need to define our own in `zarr.storage` as well.
So go ahead and drop it.
Make use of Numcodecs' `ensure_contiguous_ndarray` to take `ndarray`
views onto buffers to be stored in a few cases so as to reshape them and
avoid a copy (thanks to the buffer protocol). This ensures that
datetime/timedeltas are handled by default. Also catches things like
object arrays. Finally this handles flattening the array if needed.
All-in-all this gets as close to a `bytes` object as possible while not
copying and doing its best to preserve type information while
constructing something that fits the buffer protocol.
Rewrite `buffer_size` to just use Numcodecs' `ensure_ndarray` to get an
`ndarray` that views the data. Once the `ndarray` is gotten, all that is
needed is to access its `nbytes` member, which returns the number of
bytes that it takes up.
If the data is already a `str` instance, turn `ensure_str` into a no-op.
For all other cases, make use of Numcodecs' `ensure_bytes` to aid
`ensure_str` in coercing data through the buffer protocol. If we are on
Python 3, then decode the `bytes` object to a `str`.
@jakirkhamjakirkham mentioned this pull request Dec 4, 2018
7 tasks
As Blosc got upgraded and it contained an upgrade of Zstd, the results
changed a little bit for this example. So update them accordingly.
Should fix the doctest failure.
@jakirkhamjakirkham added this to the v2.3 milestone Dec 4, 2018

@alimanfooalimanfoo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @jakirkham, all looks good.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Looks like the line not covered is a fallback for when file removal fails. Given PR ( #327 ) obviates that, maybe we should merge that PR and drop that fallback. Thoughts?

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Made PR ( #355 ) to simply ignore coverage on that group of lines.

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Alright, think this is ready now. 😄

@alimanfoo

Copy link
Copy Markdown
Member

Awesome, merging...

@alimanfoo
alimanfoo merged commit c4427a4 into zarr-developers:masterDec 4, 2018
@alimanfooalimanfoo mentioned this pull request Dec 4, 2018
7 tasks
@jakirkham
jakirkham deleted the use_numcodecs_0.6.2 branch December 5, 2018 02:48
@jakirkham

Copy link
Copy Markdown
MemberAuthor

Missed a bit of code that could be simplified with ensure_ndarray. Addressing that in PR ( #360 ).

@jakirkham

Copy link
Copy Markdown
MemberAuthor

Also dropping a workaround for older Numcodecs. ( #361 ) Please let me know if there are more of these.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

MsgPack codec is broken for array of bytes objects

2 participants

@jakirkham@alimanfoo