Support all indexing variants - #1917

Merged
normanrz merged 16 commits into
v3from
v3-indexing
Jun 3, 2024
Merged

Support all indexing variants#1917
normanrz merged 16 commits into
v3from
v3-indexing

Conversation

@normanrz

Copy link
Copy Markdown
Member

In this PR I ported over all indexing variants from the v2 codebase into v3. The Array class exposes the sames methods as in the v2 code base (e.g. get_basic_selection, set_orthogonal_selection, oindex[...]) whereas the AsyncArray only has _get_selection(indexer, ...) and _set_selection(indexer, ...).

The current status of this PR is: the tests are green but typing still causes some headaches.

There are a few breaking changes of the v3 code that we should discuss:

  • 0-dimensional arrays (i.e. single values) are not supported
  • selecting fields (object dtype) is shaky
  • The out kwarg for getitem doesn't work and can be complicated to implement with the new NDBuffer abstraction

Please let me know your thoughts on these limitations @jhamman@d-v-b.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

@normanrznormanrz added the V3 label May 27, 2024
@normanrznormanrz self-assigned this May 27, 2024
@d-v-b

Copy link
Copy Markdown
Contributor

Thanks for this effort! I will give it a closer look later today.

regarding these two concerns:

  • 0-dimensional arrays (i.e. single values) are not supported

  • selecting fields (object dtype) is shaky

I don't have a lot of experience with either of these features from v2, so I am probably not the best judge of how we want this to look. Are there any big time 0-dimensional array or object dtype users we can ping to have them look it over?

The out kwarg for getitem doesn't work and can be complicated to implement with the new NDBuffer abstraction

Similarly, I never use out, but maybe @madsbk has some thoughts here?

@madsbk

Copy link
Copy Markdown
Contributor

Regarding out, I think it should be all or nothing. That is, either we implement an output-by-caller policy where all components takes an output argument, or we don't use output arguments at all.

As @akshaysubr point out in #1751 (comment), it might be diffecult to support an output-by-caller policy so I am leaning to remove the out argument and make it clear that the Zarr stack will copy data in most cases.

@normanrz

Copy link
Copy Markdown
MemberAuthor

While I agree that providing an out kwarg is pointless if zarr-python is doing copies anyways, we have it in the existing API. I wonder if we should continue to provide the out arg but issue a (deprecation) warning?

@madsbk

madsbk commented May 29, 2024

Copy link
Copy Markdown
Contributor

Good point, a deprecation warning is a good idea :)

@rabernat

Copy link
Copy Markdown
Contributor

Are there any big time 0-dimensional array or object dtype users we can ping to have them look it over?

#1874 recently revealed that there are definitely users of 0-dimensional arrays. (Also python objects dtypes.)

@normanrznormanrz added this to the 3.0.0.alpha milestone May 31, 2024
Comment threadsrc/zarr/codecs/pipeline.py Outdated
Comment threadsrc/zarr/indexing.py Outdated
BlockSelection = BlockSelector | tuple[BlockSelector, ...]
BlockSelectionNormalized = tuple[BlockSelector, ...]
MaskSelection = npt.NDArray[np.bool_]
OrthogonalSelector = int | slice | npt.NDArray[np.intp | np.bool_]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we should look for a way to simplify this as much as we can.OrthogonalSelector vs OrthogonalSelection, where the latter contains the former, is a recipe for confusion.

Comment threadsrc/zarr/indexing.py
CoordinateSelection = npt.NDArray[np.intp]
BlockSelector = int | slice
BlockSelection = BlockSelector | tuple[BlockSelector, ...]
BlockSelectionNormalized = tuple[BlockSelector, ...]

@d-v-bd-v-bJun 1, 2024

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think when we use a type alias like BlockSelectionNormalized for tuple[BlockSelector,...] we lose some readability. To know that BlockSelectionNormalized is tuple[int | slice, ...] I have to do 2 lookups, and we aren't even saving lines of code because the type aliases are longer than the types. Was there a particular problem caused by the plain type annotations?

Comment threadsrc/zarr/indexing.py
MaskSelection = npt.NDArray[np.bool_]
OrthogonalSelector = int | slice | npt.NDArray[np.intp | np.bool_]
OrthogonalSelection = OrthogonalSelector | tuple[OrthogonalSelector, ...]
OrthogonalSelectionNormalized = tuple[OrthogonalSelector, ...]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

From the name OrthogonalSelectionNormalized I would expect that this type takes OrthogonalSelection, but it takes OrthogonalSelecTOR, which is a bit surprising / confusing

@normanrz

Copy link
Copy Markdown
MemberAuthor

I got the types to a state where mypy doesn't complain anymore. I know that the types are far from ideal. However, it will take more work to iron that out given the variety of indexing methods that zarr-python supports and I don't want to hold this off from the alpha release.

@normanrz
normanrz marked this pull request as ready for review June 1, 2024 20:27
@d-v-b

d-v-b commented Jun 1, 2024

Copy link
Copy Markdown
Contributor

This is great, thanks @normanrz! Agreed that we can sort out the types in a later effort.

@jhammanjhamman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looking great @normanrz -- a few comments to help tidy this up.

Comment threadtests/v3/util.py Outdated
Comment threadtests/v3/util.py Outdated
Comment threadtests/v3/util.py Outdated
@normanrz
normanrz merged commit 24e855c into v3Jun 3, 2024
@normanrz
normanrz deleted the v3-indexing branch June 3, 2024 11:57
d-v-b pushed a commit to d-v-b/zarr-python that referenced this pull request Jun 4, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

5 participants

@normanrz@d-v-b@madsbk@rabernat@jhamman
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Support all indexing variants - #1917

Merged
normanrz merged 16 commits into
v3from
v3-indexing
Jun 3, 2024
Merged

Support all indexing variants#1917
normanrz merged 16 commits into
v3from
v3-indexing

Conversation

@normanrz

Copy link
Copy Markdown
Member

In this PR I ported over all indexing variants from the v2 codebase into v3. The Array class exposes the sames methods as in the v2 code base (e.g. get_basic_selection, set_orthogonal_selection, oindex[...]) whereas the AsyncArray only has _get_selection(indexer, ...) and _set_selection(indexer, ...).

The current status of this PR is: the tests are green but typing still causes some headaches.

There are a few breaking changes of the v3 code that we should discuss:

  • 0-dimensional arrays (i.e. single values) are not supported
  • selecting fields (object dtype) is shaky
  • The out kwarg for getitem doesn't work and can be complicated to implement with the new NDBuffer abstraction

Please let me know your thoughts on these limitations @jhamman@d-v-b.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

@normanrznormanrz added the V3 label May 27, 2024
@normanrznormanrz self-assigned this May 27, 2024
@d-v-b

Copy link
Copy Markdown
Contributor

Thanks for this effort! I will give it a closer look later today.

regarding these two concerns:

  • 0-dimensional arrays (i.e. single values) are not supported

  • selecting fields (object dtype) is shaky

I don't have a lot of experience with either of these features from v2, so I am probably not the best judge of how we want this to look. Are there any big time 0-dimensional array or object dtype users we can ping to have them look it over?

The out kwarg for getitem doesn't work and can be complicated to implement with the new NDBuffer abstraction

Similarly, I never use out, but maybe @madsbk has some thoughts here?

@madsbk

Copy link
Copy Markdown
Contributor

Regarding out, I think it should be all or nothing. That is, either we implement an output-by-caller policy where all components takes an output argument, or we don't use output arguments at all.

As @akshaysubr point out in #1751 (comment), it might be diffecult to support an output-by-caller policy so I am leaning to remove the out argument and make it clear that the Zarr stack will copy data in most cases.

@normanrz

Copy link
Copy Markdown
MemberAuthor

While I agree that providing an out kwarg is pointless if zarr-python is doing copies anyways, we have it in the existing API. I wonder if we should continue to provide the out arg but issue a (deprecation) warning?

@madsbk

madsbk commented May 29, 2024

Copy link
Copy Markdown
Contributor

Good point, a deprecation warning is a good idea :)

@rabernat

Copy link
Copy Markdown
Contributor

Are there any big time 0-dimensional array or object dtype users we can ping to have them look it over?

#1874 recently revealed that there are definitely users of 0-dimensional arrays. (Also python objects dtypes.)

@normanrznormanrz added this to the 3.0.0.alpha milestone May 31, 2024
Comment threadsrc/zarr/codecs/pipeline.py Outdated
Comment threadsrc/zarr/indexing.py Outdated
BlockSelection = BlockSelector | tuple[BlockSelector, ...]
BlockSelectionNormalized = tuple[BlockSelector, ...]
MaskSelection = npt.NDArray[np.bool_]
OrthogonalSelector = int | slice | npt.NDArray[np.intp | np.bool_]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we should look for a way to simplify this as much as we can.OrthogonalSelector vs OrthogonalSelection, where the latter contains the former, is a recipe for confusion.

Comment threadsrc/zarr/indexing.py
CoordinateSelection = npt.NDArray[np.intp]
BlockSelector = int | slice
BlockSelection = BlockSelector | tuple[BlockSelector, ...]
BlockSelectionNormalized = tuple[BlockSelector, ...]

@d-v-bd-v-bJun 1, 2024

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think when we use a type alias like BlockSelectionNormalized for tuple[BlockSelector,...] we lose some readability. To know that BlockSelectionNormalized is tuple[int | slice, ...] I have to do 2 lookups, and we aren't even saving lines of code because the type aliases are longer than the types. Was there a particular problem caused by the plain type annotations?

Comment threadsrc/zarr/indexing.py
MaskSelection = npt.NDArray[np.bool_]
OrthogonalSelector = int | slice | npt.NDArray[np.intp | np.bool_]
OrthogonalSelection = OrthogonalSelector | tuple[OrthogonalSelector, ...]
OrthogonalSelectionNormalized = tuple[OrthogonalSelector, ...]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

From the name OrthogonalSelectionNormalized I would expect that this type takes OrthogonalSelection, but it takes OrthogonalSelecTOR, which is a bit surprising / confusing

@normanrz

Copy link
Copy Markdown
MemberAuthor

I got the types to a state where mypy doesn't complain anymore. I know that the types are far from ideal. However, it will take more work to iron that out given the variety of indexing methods that zarr-python supports and I don't want to hold this off from the alpha release.

@normanrz
normanrz marked this pull request as ready for review June 1, 2024 20:27
@d-v-b

d-v-b commented Jun 1, 2024

Copy link
Copy Markdown
Contributor

This is great, thanks @normanrz! Agreed that we can sort out the types in a later effort.

@jhammanjhamman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looking great @normanrz -- a few comments to help tidy this up.

Comment threadtests/v3/util.py Outdated
Comment threadtests/v3/util.py Outdated
Comment threadtests/v3/util.py Outdated
@normanrz
normanrz merged commit 24e855c into v3Jun 3, 2024
@normanrz
normanrz deleted the v3-indexing branch June 3, 2024 11:57
d-v-b pushed a commit to d-v-b/zarr-python that referenced this pull request Jun 4, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

5 participants

@normanrz@d-v-b@madsbk@rabernat@jhamman
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Support all indexing variants - #1917

Merged
normanrz merged 16 commits into
v3from
v3-indexing
Jun 3, 2024
Merged

Support all indexing variants#1917
normanrz merged 16 commits into
v3from
v3-indexing

Conversation

@normanrz

Copy link
Copy Markdown
Member

In this PR I ported over all indexing variants from the v2 codebase into v3. The Array class exposes the sames methods as in the v2 code base (e.g. get_basic_selection, set_orthogonal_selection, oindex[...]) whereas the AsyncArray only has _get_selection(indexer, ...) and _set_selection(indexer, ...).

The current status of this PR is: the tests are green but typing still causes some headaches.

There are a few breaking changes of the v3 code that we should discuss:

  • 0-dimensional arrays (i.e. single values) are not supported
  • selecting fields (object dtype) is shaky
  • The out kwarg for getitem doesn't work and can be complicated to implement with the new NDBuffer abstraction

Please let me know your thoughts on these limitations @jhamman@d-v-b.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

@normanrznormanrz added the V3 label May 27, 2024
@normanrznormanrz self-assigned this May 27, 2024
@d-v-b

Copy link
Copy Markdown
Contributor

Thanks for this effort! I will give it a closer look later today.

regarding these two concerns:

  • 0-dimensional arrays (i.e. single values) are not supported

  • selecting fields (object dtype) is shaky

I don't have a lot of experience with either of these features from v2, so I am probably not the best judge of how we want this to look. Are there any big time 0-dimensional array or object dtype users we can ping to have them look it over?

The out kwarg for getitem doesn't work and can be complicated to implement with the new NDBuffer abstraction

Similarly, I never use out, but maybe @madsbk has some thoughts here?

@madsbk

Copy link
Copy Markdown
Contributor

Regarding out, I think it should be all or nothing. That is, either we implement an output-by-caller policy where all components takes an output argument, or we don't use output arguments at all.

As @akshaysubr point out in #1751 (comment), it might be diffecult to support an output-by-caller policy so I am leaning to remove the out argument and make it clear that the Zarr stack will copy data in most cases.

@normanrz

Copy link
Copy Markdown
MemberAuthor

While I agree that providing an out kwarg is pointless if zarr-python is doing copies anyways, we have it in the existing API. I wonder if we should continue to provide the out arg but issue a (deprecation) warning?

@madsbk

madsbk commented May 29, 2024

Copy link
Copy Markdown
Contributor

Good point, a deprecation warning is a good idea :)

@rabernat

Copy link
Copy Markdown
Contributor

Are there any big time 0-dimensional array or object dtype users we can ping to have them look it over?

#1874 recently revealed that there are definitely users of 0-dimensional arrays. (Also python objects dtypes.)

@normanrznormanrz added this to the 3.0.0.alpha milestone May 31, 2024
Comment threadsrc/zarr/codecs/pipeline.py Outdated
Comment threadsrc/zarr/indexing.py Outdated
BlockSelection = BlockSelector | tuple[BlockSelector, ...]
BlockSelectionNormalized = tuple[BlockSelector, ...]
MaskSelection = npt.NDArray[np.bool_]
OrthogonalSelector = int | slice | npt.NDArray[np.intp | np.bool_]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we should look for a way to simplify this as much as we can.OrthogonalSelector vs OrthogonalSelection, where the latter contains the former, is a recipe for confusion.

Comment threadsrc/zarr/indexing.py
CoordinateSelection = npt.NDArray[np.intp]
BlockSelector = int | slice
BlockSelection = BlockSelector | tuple[BlockSelector, ...]
BlockSelectionNormalized = tuple[BlockSelector, ...]

@d-v-bd-v-bJun 1, 2024

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think when we use a type alias like BlockSelectionNormalized for tuple[BlockSelector,...] we lose some readability. To know that BlockSelectionNormalized is tuple[int | slice, ...] I have to do 2 lookups, and we aren't even saving lines of code because the type aliases are longer than the types. Was there a particular problem caused by the plain type annotations?

Comment threadsrc/zarr/indexing.py
MaskSelection = npt.NDArray[np.bool_]
OrthogonalSelector = int | slice | npt.NDArray[np.intp | np.bool_]
OrthogonalSelection = OrthogonalSelector | tuple[OrthogonalSelector, ...]
OrthogonalSelectionNormalized = tuple[OrthogonalSelector, ...]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

From the name OrthogonalSelectionNormalized I would expect that this type takes OrthogonalSelection, but it takes OrthogonalSelecTOR, which is a bit surprising / confusing

@normanrz

Copy link
Copy Markdown
MemberAuthor

I got the types to a state where mypy doesn't complain anymore. I know that the types are far from ideal. However, it will take more work to iron that out given the variety of indexing methods that zarr-python supports and I don't want to hold this off from the alpha release.

@normanrz
normanrz marked this pull request as ready for review June 1, 2024 20:27
@d-v-b

d-v-b commented Jun 1, 2024

Copy link
Copy Markdown
Contributor

This is great, thanks @normanrz! Agreed that we can sort out the types in a later effort.

@jhammanjhamman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looking great @normanrz -- a few comments to help tidy this up.

Comment threadtests/v3/util.py Outdated
Comment threadtests/v3/util.py Outdated
Comment threadtests/v3/util.py Outdated
@normanrz
normanrz merged commit 24e855c into v3Jun 3, 2024
@normanrz
normanrz deleted the v3-indexing branch June 3, 2024 11:57
d-v-b pushed a commit to d-v-b/zarr-python that referenced this pull request Jun 4, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

5 participants

@normanrz@d-v-b@madsbk@rabernat@jhamman
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Support all indexing variants - #1917

Merged
normanrz merged 16 commits into
v3from
v3-indexing
Jun 3, 2024
Merged

Support all indexing variants#1917
normanrz merged 16 commits into
v3from
v3-indexing

Conversation

@normanrz

Copy link
Copy Markdown
Member

In this PR I ported over all indexing variants from the v2 codebase into v3. The Array class exposes the sames methods as in the v2 code base (e.g. get_basic_selection, set_orthogonal_selection, oindex[...]) whereas the AsyncArray only has _get_selection(indexer, ...) and _set_selection(indexer, ...).

The current status of this PR is: the tests are green but typing still causes some headaches.

There are a few breaking changes of the v3 code that we should discuss:

  • 0-dimensional arrays (i.e. single values) are not supported
  • selecting fields (object dtype) is shaky
  • The out kwarg for getitem doesn't work and can be complicated to implement with the new NDBuffer abstraction

Please let me know your thoughts on these limitations @jhamman@d-v-b.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

@normanrznormanrz added the V3 label May 27, 2024
@normanrznormanrz self-assigned this May 27, 2024
@d-v-b

Copy link
Copy Markdown
Contributor

Thanks for this effort! I will give it a closer look later today.

regarding these two concerns:

  • 0-dimensional arrays (i.e. single values) are not supported

  • selecting fields (object dtype) is shaky

I don't have a lot of experience with either of these features from v2, so I am probably not the best judge of how we want this to look. Are there any big time 0-dimensional array or object dtype users we can ping to have them look it over?

The out kwarg for getitem doesn't work and can be complicated to implement with the new NDBuffer abstraction

Similarly, I never use out, but maybe @madsbk has some thoughts here?

@madsbk

Copy link
Copy Markdown
Contributor

Regarding out, I think it should be all or nothing. That is, either we implement an output-by-caller policy where all components takes an output argument, or we don't use output arguments at all.

As @akshaysubr point out in #1751 (comment), it might be diffecult to support an output-by-caller policy so I am leaning to remove the out argument and make it clear that the Zarr stack will copy data in most cases.

@normanrz

Copy link
Copy Markdown
MemberAuthor

While I agree that providing an out kwarg is pointless if zarr-python is doing copies anyways, we have it in the existing API. I wonder if we should continue to provide the out arg but issue a (deprecation) warning?

@madsbk

madsbk commented May 29, 2024

Copy link
Copy Markdown
Contributor

Good point, a deprecation warning is a good idea :)

@rabernat

Copy link
Copy Markdown
Contributor

Are there any big time 0-dimensional array or object dtype users we can ping to have them look it over?

#1874 recently revealed that there are definitely users of 0-dimensional arrays. (Also python objects dtypes.)

@normanrznormanrz added this to the 3.0.0.alpha milestone May 31, 2024
Comment threadsrc/zarr/codecs/pipeline.py Outdated
Comment threadsrc/zarr/indexing.py Outdated
BlockSelection = BlockSelector | tuple[BlockSelector, ...]
BlockSelectionNormalized = tuple[BlockSelector, ...]
MaskSelection = npt.NDArray[np.bool_]
OrthogonalSelector = int | slice | npt.NDArray[np.intp | np.bool_]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we should look for a way to simplify this as much as we can.OrthogonalSelector vs OrthogonalSelection, where the latter contains the former, is a recipe for confusion.

Comment threadsrc/zarr/indexing.py
CoordinateSelection = npt.NDArray[np.intp]
BlockSelector = int | slice
BlockSelection = BlockSelector | tuple[BlockSelector, ...]
BlockSelectionNormalized = tuple[BlockSelector, ...]

@d-v-bd-v-bJun 1, 2024

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think when we use a type alias like BlockSelectionNormalized for tuple[BlockSelector,...] we lose some readability. To know that BlockSelectionNormalized is tuple[int | slice, ...] I have to do 2 lookups, and we aren't even saving lines of code because the type aliases are longer than the types. Was there a particular problem caused by the plain type annotations?

Comment threadsrc/zarr/indexing.py
MaskSelection = npt.NDArray[np.bool_]
OrthogonalSelector = int | slice | npt.NDArray[np.intp | np.bool_]
OrthogonalSelection = OrthogonalSelector | tuple[OrthogonalSelector, ...]
OrthogonalSelectionNormalized = tuple[OrthogonalSelector, ...]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

From the name OrthogonalSelectionNormalized I would expect that this type takes OrthogonalSelection, but it takes OrthogonalSelecTOR, which is a bit surprising / confusing

@normanrz

Copy link
Copy Markdown
MemberAuthor

I got the types to a state where mypy doesn't complain anymore. I know that the types are far from ideal. However, it will take more work to iron that out given the variety of indexing methods that zarr-python supports and I don't want to hold this off from the alpha release.

@normanrz
normanrz marked this pull request as ready for review June 1, 2024 20:27
@d-v-b

d-v-b commented Jun 1, 2024

Copy link
Copy Markdown
Contributor

This is great, thanks @normanrz! Agreed that we can sort out the types in a later effort.

@jhammanjhamman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looking great @normanrz -- a few comments to help tidy this up.

Comment threadtests/v3/util.py Outdated
Comment threadtests/v3/util.py Outdated
Comment threadtests/v3/util.py Outdated
@normanrz
normanrz merged commit 24e855c into v3Jun 3, 2024
@normanrz
normanrz deleted the v3-indexing branch June 3, 2024 11:57
d-v-b pushed a commit to d-v-b/zarr-python that referenced this pull request Jun 4, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

5 participants

@normanrz@d-v-b@madsbk@rabernat@jhamman
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Support all indexing variants - #1917

Merged
normanrz merged 16 commits into
v3from
v3-indexing
Jun 3, 2024
Merged

Support all indexing variants#1917
normanrz merged 16 commits into
v3from
v3-indexing

Conversation

@normanrz

Copy link
Copy Markdown
Member

In this PR I ported over all indexing variants from the v2 codebase into v3. The Array class exposes the sames methods as in the v2 code base (e.g. get_basic_selection, set_orthogonal_selection, oindex[...]) whereas the AsyncArray only has _get_selection(indexer, ...) and _set_selection(indexer, ...).

The current status of this PR is: the tests are green but typing still causes some headaches.

There are a few breaking changes of the v3 code that we should discuss:

  • 0-dimensional arrays (i.e. single values) are not supported
  • selecting fields (object dtype) is shaky
  • The out kwarg for getitem doesn't work and can be complicated to implement with the new NDBuffer abstraction

Please let me know your thoughts on these limitations @jhamman@d-v-b.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

@normanrznormanrz added the V3 label May 27, 2024
@normanrznormanrz self-assigned this May 27, 2024
@d-v-b

Copy link
Copy Markdown
Contributor

Thanks for this effort! I will give it a closer look later today.

regarding these two concerns:

  • 0-dimensional arrays (i.e. single values) are not supported

  • selecting fields (object dtype) is shaky

I don't have a lot of experience with either of these features from v2, so I am probably not the best judge of how we want this to look. Are there any big time 0-dimensional array or object dtype users we can ping to have them look it over?

The out kwarg for getitem doesn't work and can be complicated to implement with the new NDBuffer abstraction

Similarly, I never use out, but maybe @madsbk has some thoughts here?

@madsbk

Copy link
Copy Markdown
Contributor

Regarding out, I think it should be all or nothing. That is, either we implement an output-by-caller policy where all components takes an output argument, or we don't use output arguments at all.

As @akshaysubr point out in #1751 (comment), it might be diffecult to support an output-by-caller policy so I am leaning to remove the out argument and make it clear that the Zarr stack will copy data in most cases.

@normanrz

Copy link
Copy Markdown
MemberAuthor

While I agree that providing an out kwarg is pointless if zarr-python is doing copies anyways, we have it in the existing API. I wonder if we should continue to provide the out arg but issue a (deprecation) warning?

@madsbk

madsbk commented May 29, 2024

Copy link
Copy Markdown
Contributor

Good point, a deprecation warning is a good idea :)

@rabernat

Copy link
Copy Markdown
Contributor

Are there any big time 0-dimensional array or object dtype users we can ping to have them look it over?

#1874 recently revealed that there are definitely users of 0-dimensional arrays. (Also python objects dtypes.)

@normanrznormanrz added this to the 3.0.0.alpha milestone May 31, 2024
Comment threadsrc/zarr/codecs/pipeline.py Outdated
Comment threadsrc/zarr/indexing.py Outdated
BlockSelection = BlockSelector | tuple[BlockSelector, ...]
BlockSelectionNormalized = tuple[BlockSelector, ...]
MaskSelection = npt.NDArray[np.bool_]
OrthogonalSelector = int | slice | npt.NDArray[np.intp | np.bool_]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we should look for a way to simplify this as much as we can.OrthogonalSelector vs OrthogonalSelection, where the latter contains the former, is a recipe for confusion.

Comment threadsrc/zarr/indexing.py
CoordinateSelection = npt.NDArray[np.intp]
BlockSelector = int | slice
BlockSelection = BlockSelector | tuple[BlockSelector, ...]
BlockSelectionNormalized = tuple[BlockSelector, ...]

@d-v-bd-v-bJun 1, 2024

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think when we use a type alias like BlockSelectionNormalized for tuple[BlockSelector,...] we lose some readability. To know that BlockSelectionNormalized is tuple[int | slice, ...] I have to do 2 lookups, and we aren't even saving lines of code because the type aliases are longer than the types. Was there a particular problem caused by the plain type annotations?

Comment threadsrc/zarr/indexing.py
MaskSelection = npt.NDArray[np.bool_]
OrthogonalSelector = int | slice | npt.NDArray[np.intp | np.bool_]
OrthogonalSelection = OrthogonalSelector | tuple[OrthogonalSelector, ...]
OrthogonalSelectionNormalized = tuple[OrthogonalSelector, ...]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

From the name OrthogonalSelectionNormalized I would expect that this type takes OrthogonalSelection, but it takes OrthogonalSelecTOR, which is a bit surprising / confusing

@normanrz

Copy link
Copy Markdown
MemberAuthor

I got the types to a state where mypy doesn't complain anymore. I know that the types are far from ideal. However, it will take more work to iron that out given the variety of indexing methods that zarr-python supports and I don't want to hold this off from the alpha release.

@normanrz
normanrz marked this pull request as ready for review June 1, 2024 20:27
@d-v-b

d-v-b commented Jun 1, 2024

Copy link
Copy Markdown
Contributor

This is great, thanks @normanrz! Agreed that we can sort out the types in a later effort.

@jhammanjhamman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looking great @normanrz -- a few comments to help tidy this up.

Comment threadtests/v3/util.py Outdated
Comment threadtests/v3/util.py Outdated
Comment threadtests/v3/util.py Outdated
@normanrz
normanrz merged commit 24e855c into v3Jun 3, 2024
@normanrz
normanrz deleted the v3-indexing branch June 3, 2024 11:57
d-v-b pushed a commit to d-v-b/zarr-python that referenced this pull request Jun 4, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

5 participants

@normanrz@d-v-b@madsbk@rabernat@jhamman
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Support all indexing variants - #1917

Merged
normanrz merged 16 commits into
v3from
v3-indexing
Jun 3, 2024
Merged

Support all indexing variants#1917
normanrz merged 16 commits into
v3from
v3-indexing

Conversation

@normanrz

Copy link
Copy Markdown
Member

In this PR I ported over all indexing variants from the v2 codebase into v3. The Array class exposes the sames methods as in the v2 code base (e.g. get_basic_selection, set_orthogonal_selection, oindex[...]) whereas the AsyncArray only has _get_selection(indexer, ...) and _set_selection(indexer, ...).

The current status of this PR is: the tests are green but typing still causes some headaches.

There are a few breaking changes of the v3 code that we should discuss:

  • 0-dimensional arrays (i.e. single values) are not supported
  • selecting fields (object dtype) is shaky
  • The out kwarg for getitem doesn't work and can be complicated to implement with the new NDBuffer abstraction

Please let me know your thoughts on these limitations @jhamman@d-v-b.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

@normanrznormanrz added the V3 label May 27, 2024
@normanrznormanrz self-assigned this May 27, 2024
@d-v-b

Copy link
Copy Markdown
Contributor

Thanks for this effort! I will give it a closer look later today.

regarding these two concerns:

  • 0-dimensional arrays (i.e. single values) are not supported

  • selecting fields (object dtype) is shaky

I don't have a lot of experience with either of these features from v2, so I am probably not the best judge of how we want this to look. Are there any big time 0-dimensional array or object dtype users we can ping to have them look it over?

The out kwarg for getitem doesn't work and can be complicated to implement with the new NDBuffer abstraction

Similarly, I never use out, but maybe @madsbk has some thoughts here?

@madsbk

Copy link
Copy Markdown
Contributor

Regarding out, I think it should be all or nothing. That is, either we implement an output-by-caller policy where all components takes an output argument, or we don't use output arguments at all.

As @akshaysubr point out in #1751 (comment), it might be diffecult to support an output-by-caller policy so I am leaning to remove the out argument and make it clear that the Zarr stack will copy data in most cases.

@normanrz

Copy link
Copy Markdown
MemberAuthor

While I agree that providing an out kwarg is pointless if zarr-python is doing copies anyways, we have it in the existing API. I wonder if we should continue to provide the out arg but issue a (deprecation) warning?

@madsbk

madsbk commented May 29, 2024

Copy link
Copy Markdown
Contributor

Good point, a deprecation warning is a good idea :)

@rabernat

Copy link
Copy Markdown
Contributor

Are there any big time 0-dimensional array or object dtype users we can ping to have them look it over?

#1874 recently revealed that there are definitely users of 0-dimensional arrays. (Also python objects dtypes.)

@normanrznormanrz added this to the 3.0.0.alpha milestone May 31, 2024
Comment threadsrc/zarr/codecs/pipeline.py Outdated
Comment threadsrc/zarr/indexing.py Outdated
BlockSelection = BlockSelector | tuple[BlockSelector, ...]
BlockSelectionNormalized = tuple[BlockSelector, ...]
MaskSelection = npt.NDArray[np.bool_]
OrthogonalSelector = int | slice | npt.NDArray[np.intp | np.bool_]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we should look for a way to simplify this as much as we can.OrthogonalSelector vs OrthogonalSelection, where the latter contains the former, is a recipe for confusion.

Comment threadsrc/zarr/indexing.py
CoordinateSelection = npt.NDArray[np.intp]
BlockSelector = int | slice
BlockSelection = BlockSelector | tuple[BlockSelector, ...]
BlockSelectionNormalized = tuple[BlockSelector, ...]

@d-v-bd-v-bJun 1, 2024

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think when we use a type alias like BlockSelectionNormalized for tuple[BlockSelector,...] we lose some readability. To know that BlockSelectionNormalized is tuple[int | slice, ...] I have to do 2 lookups, and we aren't even saving lines of code because the type aliases are longer than the types. Was there a particular problem caused by the plain type annotations?

Comment threadsrc/zarr/indexing.py
MaskSelection = npt.NDArray[np.bool_]
OrthogonalSelector = int | slice | npt.NDArray[np.intp | np.bool_]
OrthogonalSelection = OrthogonalSelector | tuple[OrthogonalSelector, ...]
OrthogonalSelectionNormalized = tuple[OrthogonalSelector, ...]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

From the name OrthogonalSelectionNormalized I would expect that this type takes OrthogonalSelection, but it takes OrthogonalSelecTOR, which is a bit surprising / confusing

@normanrz

Copy link
Copy Markdown
MemberAuthor

I got the types to a state where mypy doesn't complain anymore. I know that the types are far from ideal. However, it will take more work to iron that out given the variety of indexing methods that zarr-python supports and I don't want to hold this off from the alpha release.

@normanrz
normanrz marked this pull request as ready for review June 1, 2024 20:27
@d-v-b

d-v-b commented Jun 1, 2024

Copy link
Copy Markdown
Contributor

This is great, thanks @normanrz! Agreed that we can sort out the types in a later effort.

@jhammanjhamman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looking great @normanrz -- a few comments to help tidy this up.

Comment threadtests/v3/util.py Outdated
Comment threadtests/v3/util.py Outdated
Comment threadtests/v3/util.py Outdated
@normanrz
normanrz merged commit 24e855c into v3Jun 3, 2024
@normanrz
normanrz deleted the v3-indexing branch June 3, 2024 11:57
d-v-b pushed a commit to d-v-b/zarr-python that referenced this pull request Jun 4, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

5 participants

@normanrz@d-v-b@madsbk@rabernat@jhamman
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Support all indexing variants - #1917

Merged
normanrz merged 16 commits into
v3from
v3-indexing
Jun 3, 2024
Merged

Support all indexing variants#1917
normanrz merged 16 commits into
v3from
v3-indexing

Conversation

@normanrz

Copy link
Copy Markdown
Member

In this PR I ported over all indexing variants from the v2 codebase into v3. The Array class exposes the sames methods as in the v2 code base (e.g. get_basic_selection, set_orthogonal_selection, oindex[...]) whereas the AsyncArray only has _get_selection(indexer, ...) and _set_selection(indexer, ...).

The current status of this PR is: the tests are green but typing still causes some headaches.

There are a few breaking changes of the v3 code that we should discuss:

  • 0-dimensional arrays (i.e. single values) are not supported
  • selecting fields (object dtype) is shaky
  • The out kwarg for getitem doesn't work and can be complicated to implement with the new NDBuffer abstraction

Please let me know your thoughts on these limitations @jhamman@d-v-b.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

@normanrznormanrz added the V3 label May 27, 2024
@normanrznormanrz self-assigned this May 27, 2024
@d-v-b

Copy link
Copy Markdown
Contributor

Thanks for this effort! I will give it a closer look later today.

regarding these two concerns:

  • 0-dimensional arrays (i.e. single values) are not supported

  • selecting fields (object dtype) is shaky

I don't have a lot of experience with either of these features from v2, so I am probably not the best judge of how we want this to look. Are there any big time 0-dimensional array or object dtype users we can ping to have them look it over?

The out kwarg for getitem doesn't work and can be complicated to implement with the new NDBuffer abstraction

Similarly, I never use out, but maybe @madsbk has some thoughts here?

@madsbk

Copy link
Copy Markdown
Contributor

Regarding out, I think it should be all or nothing. That is, either we implement an output-by-caller policy where all components takes an output argument, or we don't use output arguments at all.

As @akshaysubr point out in #1751 (comment), it might be diffecult to support an output-by-caller policy so I am leaning to remove the out argument and make it clear that the Zarr stack will copy data in most cases.

@normanrz

Copy link
Copy Markdown
MemberAuthor

While I agree that providing an out kwarg is pointless if zarr-python is doing copies anyways, we have it in the existing API. I wonder if we should continue to provide the out arg but issue a (deprecation) warning?

@madsbk

madsbk commented May 29, 2024

Copy link
Copy Markdown
Contributor

Good point, a deprecation warning is a good idea :)

@rabernat

Copy link
Copy Markdown
Contributor

Are there any big time 0-dimensional array or object dtype users we can ping to have them look it over?

#1874 recently revealed that there are definitely users of 0-dimensional arrays. (Also python objects dtypes.)

@normanrznormanrz added this to the 3.0.0.alpha milestone May 31, 2024
Comment threadsrc/zarr/codecs/pipeline.py Outdated
Comment threadsrc/zarr/indexing.py Outdated
BlockSelection = BlockSelector | tuple[BlockSelector, ...]
BlockSelectionNormalized = tuple[BlockSelector, ...]
MaskSelection = npt.NDArray[np.bool_]
OrthogonalSelector = int | slice | npt.NDArray[np.intp | np.bool_]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we should look for a way to simplify this as much as we can.OrthogonalSelector vs OrthogonalSelection, where the latter contains the former, is a recipe for confusion.

Comment threadsrc/zarr/indexing.py
CoordinateSelection = npt.NDArray[np.intp]
BlockSelector = int | slice
BlockSelection = BlockSelector | tuple[BlockSelector, ...]
BlockSelectionNormalized = tuple[BlockSelector, ...]

@d-v-bd-v-bJun 1, 2024

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think when we use a type alias like BlockSelectionNormalized for tuple[BlockSelector,...] we lose some readability. To know that BlockSelectionNormalized is tuple[int | slice, ...] I have to do 2 lookups, and we aren't even saving lines of code because the type aliases are longer than the types. Was there a particular problem caused by the plain type annotations?

Comment threadsrc/zarr/indexing.py
MaskSelection = npt.NDArray[np.bool_]
OrthogonalSelector = int | slice | npt.NDArray[np.intp | np.bool_]
OrthogonalSelection = OrthogonalSelector | tuple[OrthogonalSelector, ...]
OrthogonalSelectionNormalized = tuple[OrthogonalSelector, ...]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

From the name OrthogonalSelectionNormalized I would expect that this type takes OrthogonalSelection, but it takes OrthogonalSelecTOR, which is a bit surprising / confusing

@normanrz

Copy link
Copy Markdown
MemberAuthor

I got the types to a state where mypy doesn't complain anymore. I know that the types are far from ideal. However, it will take more work to iron that out given the variety of indexing methods that zarr-python supports and I don't want to hold this off from the alpha release.

@normanrz
normanrz marked this pull request as ready for review June 1, 2024 20:27
@d-v-b

d-v-b commented Jun 1, 2024

Copy link
Copy Markdown
Contributor

This is great, thanks @normanrz! Agreed that we can sort out the types in a later effort.

@jhammanjhamman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looking great @normanrz -- a few comments to help tidy this up.

Comment threadtests/v3/util.py Outdated
Comment threadtests/v3/util.py Outdated
Comment threadtests/v3/util.py Outdated
@normanrz
normanrz merged commit 24e855c into v3Jun 3, 2024
@normanrz
normanrz deleted the v3-indexing branch June 3, 2024 11:57
d-v-b pushed a commit to d-v-b/zarr-python that referenced this pull request Jun 4, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

5 participants

@normanrz@d-v-b@madsbk@rabernat@jhamman
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Support all indexing variants - #1917

Merged
normanrz merged 16 commits into
v3from
v3-indexing
Jun 3, 2024
Merged

Support all indexing variants#1917
normanrz merged 16 commits into
v3from
v3-indexing

Conversation

@normanrz

Copy link
Copy Markdown
Member

In this PR I ported over all indexing variants from the v2 codebase into v3. The Array class exposes the sames methods as in the v2 code base (e.g. get_basic_selection, set_orthogonal_selection, oindex[...]) whereas the AsyncArray only has _get_selection(indexer, ...) and _set_selection(indexer, ...).

The current status of this PR is: the tests are green but typing still causes some headaches.

There are a few breaking changes of the v3 code that we should discuss:

  • 0-dimensional arrays (i.e. single values) are not supported
  • selecting fields (object dtype) is shaky
  • The out kwarg for getitem doesn't work and can be complicated to implement with the new NDBuffer abstraction

Please let me know your thoughts on these limitations @jhamman@d-v-b.

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

@normanrznormanrz added the V3 label May 27, 2024
@normanrznormanrz self-assigned this May 27, 2024
@d-v-b

Copy link
Copy Markdown
Contributor

Thanks for this effort! I will give it a closer look later today.

regarding these two concerns:

  • 0-dimensional arrays (i.e. single values) are not supported

  • selecting fields (object dtype) is shaky

I don't have a lot of experience with either of these features from v2, so I am probably not the best judge of how we want this to look. Are there any big time 0-dimensional array or object dtype users we can ping to have them look it over?

The out kwarg for getitem doesn't work and can be complicated to implement with the new NDBuffer abstraction

Similarly, I never use out, but maybe @madsbk has some thoughts here?

@madsbk

Copy link
Copy Markdown
Contributor

Regarding out, I think it should be all or nothing. That is, either we implement an output-by-caller policy where all components takes an output argument, or we don't use output arguments at all.

As @akshaysubr point out in #1751 (comment), it might be diffecult to support an output-by-caller policy so I am leaning to remove the out argument and make it clear that the Zarr stack will copy data in most cases.

@normanrz

Copy link
Copy Markdown
MemberAuthor

While I agree that providing an out kwarg is pointless if zarr-python is doing copies anyways, we have it in the existing API. I wonder if we should continue to provide the out arg but issue a (deprecation) warning?

@madsbk

madsbk commented May 29, 2024

Copy link
Copy Markdown
Contributor

Good point, a deprecation warning is a good idea :)

@rabernat

Copy link
Copy Markdown
Contributor

Are there any big time 0-dimensional array or object dtype users we can ping to have them look it over?

#1874 recently revealed that there are definitely users of 0-dimensional arrays. (Also python objects dtypes.)

@normanrznormanrz added this to the 3.0.0.alpha milestone May 31, 2024
Comment threadsrc/zarr/codecs/pipeline.py Outdated
Comment threadsrc/zarr/indexing.py Outdated
BlockSelection = BlockSelector | tuple[BlockSelector, ...]
BlockSelectionNormalized = tuple[BlockSelector, ...]
MaskSelection = npt.NDArray[np.bool_]
OrthogonalSelector = int | slice | npt.NDArray[np.intp | np.bool_]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we should look for a way to simplify this as much as we can.OrthogonalSelector vs OrthogonalSelection, where the latter contains the former, is a recipe for confusion.

Comment threadsrc/zarr/indexing.py
CoordinateSelection = npt.NDArray[np.intp]
BlockSelector = int | slice
BlockSelection = BlockSelector | tuple[BlockSelector, ...]
BlockSelectionNormalized = tuple[BlockSelector, ...]

@d-v-bd-v-bJun 1, 2024

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think when we use a type alias like BlockSelectionNormalized for tuple[BlockSelector,...] we lose some readability. To know that BlockSelectionNormalized is tuple[int | slice, ...] I have to do 2 lookups, and we aren't even saving lines of code because the type aliases are longer than the types. Was there a particular problem caused by the plain type annotations?

Comment threadsrc/zarr/indexing.py
MaskSelection = npt.NDArray[np.bool_]
OrthogonalSelector = int | slice | npt.NDArray[np.intp | np.bool_]
OrthogonalSelection = OrthogonalSelector | tuple[OrthogonalSelector, ...]
OrthogonalSelectionNormalized = tuple[OrthogonalSelector, ...]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

From the name OrthogonalSelectionNormalized I would expect that this type takes OrthogonalSelection, but it takes OrthogonalSelecTOR, which is a bit surprising / confusing

@normanrz

Copy link
Copy Markdown
MemberAuthor

I got the types to a state where mypy doesn't complain anymore. I know that the types are far from ideal. However, it will take more work to iron that out given the variety of indexing methods that zarr-python supports and I don't want to hold this off from the alpha release.

@normanrz
normanrz marked this pull request as ready for review June 1, 2024 20:27
@d-v-b

d-v-b commented Jun 1, 2024

Copy link
Copy Markdown
Contributor

This is great, thanks @normanrz! Agreed that we can sort out the types in a later effort.

@jhammanjhamman left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looking great @normanrz -- a few comments to help tidy this up.

Comment threadtests/v3/util.py Outdated
Comment threadtests/v3/util.py Outdated
Comment threadtests/v3/util.py Outdated
@normanrz
normanrz merged commit 24e855c into v3Jun 3, 2024
@normanrz
normanrz deleted the v3-indexing branch June 3, 2024 11:57
d-v-b pushed a commit to d-v-b/zarr-python that referenced this pull request Jun 4, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

5 participants

@normanrz@d-v-b@madsbk@rabernat@jhamman