implement store.list_prefix and store._set_many - #2064

Merged
d-v-b merged 9 commits into
v3from
store-list-prefix
Sep 19, 2024
Merged

implement store.list_prefix and store._set_many#2064
d-v-b merged 9 commits into
v3from
store-list-prefix

Conversation

@d-v-b

@d-v-bd-v-b commented Aug 3, 2024

Copy link
Copy Markdown
Contributor
  • fixes / implements list_prefix for stores. list_prefix(prefix=foo) now consistently returns keys with the shared prefix foo stripped. I'd be fine altering this to return absolute keys instead.
  • implements tests for list_prefix on the StoreTests base class
  • adds a _set_dict(Mapping[str, Buffer]) method to stores that allows declaring a collection of key: value pairs to write to storage. The primary usage is to make store tests simpler via a declarative API. I also suspect that making tests simpler anticipates making other code simpler. The default implementation of _set_dict simply wraps store.set, but it's easy to imagine fancier batching / transaction implementations.
  • alters some of the store tests to use _set_dict.
  • adds some missing type annotations in the remotestore tests

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

@d-v-b
d-v-b requested review from jhamman and martindurant and removed request for martindurantAugust 3, 2024 09:59
@d-v-b

d-v-b commented Aug 3, 2024

Copy link
Copy Markdown
ContributorAuthor

@martindurant i'm on flaky train internet and the github web API is giving me inconsistent signals about whether you have been requested for review; to be clear, I would appreciate your review :)

@d-v-bd-v-b mentioned this pull request Aug 3, 2024
6 tasks
@martindurant

Copy link
Copy Markdown
Member

list_prefix(prefix=foo) now consistently returns keys with the shared prefix foo stripped. I'd be fine altering this to return absolute keys instead.

What matters is how we intend to call it! fsspec likes to provide full paths (but has explicit prefix interfaces) - of course it's fine to tailor the situation to what we need, but I don't really know what that is.

I would appreciate your review

I may have got a notification or a few.

Comment threadsrc/zarr/abc/store.py Outdated
"""
for p in (self.root / prefix).rglob("*"):
if p.is_file():
yield str(p)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Were we getting duplicates?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

i think we were. this code path was not tested until this PR

Comment threadsrc/zarr/store/local.py Outdated
for p in (self.root / prefix).rglob("*"):
if p.is_file():
yield str(p).replace(to_strip, "")
yield str(p).removeprefix(to_strip)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1

Comment threadsrc/zarr/store/remote.py Outdated
for onefile in await self._fs._ls(prefix, detail=False):
yield onefile
find_str = "/".join([self.path, prefix])
for onefile in await self._fs._find(find_str):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The defaults for find are: maxdepth=None, withdirs=False, detail=False; maybe good to be specific.

Why is find() better than ls()? The former will return all child files, not just one level deep - is that the intent? If not, ls() ought to be generally more efficient.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

using find here is merely due to my ignorance of fsspec. I will implement ls as you suggest

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It depends on whether you want one directory level or everything below it. When I wrote the original, I didn't know the intent.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

i believe the intent here is to list everything below prefix (at least, that's how I'm using it)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I misunderstood your first comment. since the intent is to use the behavior of _find, I'm keeping it, but adding explicit kwargs as you suggested.

Comment threadsrc/zarr/sync.py
"""
result = []
async for x in data:
result.append(x)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

asyncio.gather? Like above, not much point in having coroutines if we serially wait for them.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think we can use asyncio.gather here, because AsyncGenerator is not iterable. Happy to be corrected, since I don't really know asyncio very well.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and for clarification, _collect_aiterator largely exists for convenience in testing, because I need some way to collect async generators when debugging with pdb. This function is not intended for use in anything performance sensitive.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should probably re-examine the use of async-iterators, though. If we can't gather() on them (seems to be true?), then they are the wrong abstraction since gather() is probably always what we actually want.

@martindurantmartindurantAug 3, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On second thoughts, maybe I'm wrong - does async for schedule all the coroutines at once?? Should be easy to test.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think it schedules them all at once. in [x async for x in async_generator], x is not an awaitable; it's already awaited. since the basic model of the generator is that it's a resumable, stateful iterator, I don't think we can schedule all the tasks at once.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the idea with the generators is to a) support seamless pagination and b) support pipelining (del_prefix will be able to take advantage of this at some point).

store_dict = dict(zip(keys, data_buf, strict=True))
await store._set_dict(store_dict)
for k, v in store_dict.items():
assert self.get(store, k).to_bytes() == v.to_bytes()

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if x.to_bytes() == y.to_bytes(), does x== y?

Isn't there a multiple get? Maybe not important here.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if x.to_bytes() == y.to_bytes(), does x== y?

no, and I suspect this might be deliberate since in principle Buffer instances can have identical bytes but different devices (e.g., gpu memory vs host memory); thus x == y might only be true if two buffers are bytes-equal and device-equal, but I'm speculating here. @madsbk would have a better answer I think.

Isn't there a multiple get? Maybe not important here.

there is no multiple get (nor a multiple set, nor a multiple delete).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@pytest.fixture(scope="function")
def store(self, store_kwargs: dict[str, str | bool]) -> RemoteStore:
url = store_kwargs["url"]
async def store(self, store_kwargs: dict[str, str | bool | UPath]) -> RemoteStore:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This isn't actually async

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

correct, but the class we are inheriting from defines this as an async method

Comment threadsrc/zarr/abc/store.py Outdated
self._is_open = False
pass

async def _set_dict(self, dict: Mapping[str, Buffer]) -> None:

@jhammanjhammanAug 9, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

set_many() (analogous to insert_many)?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I moved away from set_dict and switched to _set_many

@jhammanjhamman added the V3 label Aug 9, 2024
@dcheriandcherian mentioned this pull request Aug 15, 2024
6 tasks
@d-v-b
d-v-b requested a review from jhammanSeptember 19, 2024 16:13
@d-v-b

Copy link
Copy Markdown
ContributorAuthor

we should try to get this in, because it fixes problematic list_prefix behavior that has been noted elsewhere.

@d-v-bd-v-b changed the title implement store.list_prefix and store._set_dictimplement store.list_prefix and store._set_manySep 19, 2024
@d-v-b
d-v-b merged commit 06e3215 into v3Sep 19, 2024
@d-v-b
d-v-b deleted the store-list-prefix branch September 19, 2024 16:26
dcherian added a commit to dcherian/zarr-python that referenced this pull request Sep 24, 2024
* v3:
chore: update pre-commit hooks (zarr-developers#2222)
fix: validate v3 dtypes when loading/creating v3 metadata (zarr-developers#2209)
fix typo in store integration test (zarr-developers#2223)
Basic Zarr-python 2.x compatibility changes (zarr-developers#2098)
Make Group.arrays, groups compatible with v2 (zarr-developers#2213)
Typing fixes to test_indexing (zarr-developers#2193)
Default to RemoteStore for fsspec URIs (zarr-developers#2198)
Make MemoryStore serialiazable (zarr-developers#2204)
[v3] Implement Group methods for empty, full, ones, and zeros (zarr-developers#2210)
implement `store.list_prefix` and `store._set_many` (zarr-developers#2064)
Fixed codec for v2 data with no fill value (zarr-developers#2207)
@jhammanjhamman added this to the 3.0.0.beta milestone Oct 17, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants

@d-v-b@martindurant@jhamman
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

implement store.list_prefix and store._set_many - #2064

Merged
d-v-b merged 9 commits into
v3from
store-list-prefix
Sep 19, 2024
Merged

implement store.list_prefix and store._set_many#2064
d-v-b merged 9 commits into
v3from
store-list-prefix

Conversation

@d-v-b

@d-v-bd-v-b commented Aug 3, 2024

Copy link
Copy Markdown
Contributor
  • fixes / implements list_prefix for stores. list_prefix(prefix=foo) now consistently returns keys with the shared prefix foo stripped. I'd be fine altering this to return absolute keys instead.
  • implements tests for list_prefix on the StoreTests base class
  • adds a _set_dict(Mapping[str, Buffer]) method to stores that allows declaring a collection of key: value pairs to write to storage. The primary usage is to make store tests simpler via a declarative API. I also suspect that making tests simpler anticipates making other code simpler. The default implementation of _set_dict simply wraps store.set, but it's easy to imagine fancier batching / transaction implementations.
  • alters some of the store tests to use _set_dict.
  • adds some missing type annotations in the remotestore tests

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

@d-v-b
d-v-b requested review from jhamman and martindurant and removed request for martindurantAugust 3, 2024 09:59
@d-v-b

d-v-b commented Aug 3, 2024

Copy link
Copy Markdown
ContributorAuthor

@martindurant i'm on flaky train internet and the github web API is giving me inconsistent signals about whether you have been requested for review; to be clear, I would appreciate your review :)

@d-v-bd-v-b mentioned this pull request Aug 3, 2024
6 tasks
@martindurant

Copy link
Copy Markdown
Member

list_prefix(prefix=foo) now consistently returns keys with the shared prefix foo stripped. I'd be fine altering this to return absolute keys instead.

What matters is how we intend to call it! fsspec likes to provide full paths (but has explicit prefix interfaces) - of course it's fine to tailor the situation to what we need, but I don't really know what that is.

I would appreciate your review

I may have got a notification or a few.

Comment threadsrc/zarr/abc/store.py Outdated
"""
for p in (self.root / prefix).rglob("*"):
if p.is_file():
yield str(p)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Were we getting duplicates?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

i think we were. this code path was not tested until this PR

Comment threadsrc/zarr/store/local.py Outdated
for p in (self.root / prefix).rglob("*"):
if p.is_file():
yield str(p).replace(to_strip, "")
yield str(p).removeprefix(to_strip)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1

Comment threadsrc/zarr/store/remote.py Outdated
for onefile in await self._fs._ls(prefix, detail=False):
yield onefile
find_str = "/".join([self.path, prefix])
for onefile in await self._fs._find(find_str):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The defaults for find are: maxdepth=None, withdirs=False, detail=False; maybe good to be specific.

Why is find() better than ls()? The former will return all child files, not just one level deep - is that the intent? If not, ls() ought to be generally more efficient.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

using find here is merely due to my ignorance of fsspec. I will implement ls as you suggest

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It depends on whether you want one directory level or everything below it. When I wrote the original, I didn't know the intent.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

i believe the intent here is to list everything below prefix (at least, that's how I'm using it)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I misunderstood your first comment. since the intent is to use the behavior of _find, I'm keeping it, but adding explicit kwargs as you suggested.

Comment threadsrc/zarr/sync.py
"""
result = []
async for x in data:
result.append(x)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

asyncio.gather? Like above, not much point in having coroutines if we serially wait for them.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think we can use asyncio.gather here, because AsyncGenerator is not iterable. Happy to be corrected, since I don't really know asyncio very well.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and for clarification, _collect_aiterator largely exists for convenience in testing, because I need some way to collect async generators when debugging with pdb. This function is not intended for use in anything performance sensitive.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should probably re-examine the use of async-iterators, though. If we can't gather() on them (seems to be true?), then they are the wrong abstraction since gather() is probably always what we actually want.

@martindurantmartindurantAug 3, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On second thoughts, maybe I'm wrong - does async for schedule all the coroutines at once?? Should be easy to test.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think it schedules them all at once. in [x async for x in async_generator], x is not an awaitable; it's already awaited. since the basic model of the generator is that it's a resumable, stateful iterator, I don't think we can schedule all the tasks at once.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the idea with the generators is to a) support seamless pagination and b) support pipelining (del_prefix will be able to take advantage of this at some point).

store_dict = dict(zip(keys, data_buf, strict=True))
await store._set_dict(store_dict)
for k, v in store_dict.items():
assert self.get(store, k).to_bytes() == v.to_bytes()

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if x.to_bytes() == y.to_bytes(), does x== y?

Isn't there a multiple get? Maybe not important here.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if x.to_bytes() == y.to_bytes(), does x== y?

no, and I suspect this might be deliberate since in principle Buffer instances can have identical bytes but different devices (e.g., gpu memory vs host memory); thus x == y might only be true if two buffers are bytes-equal and device-equal, but I'm speculating here. @madsbk would have a better answer I think.

Isn't there a multiple get? Maybe not important here.

there is no multiple get (nor a multiple set, nor a multiple delete).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@pytest.fixture(scope="function")
def store(self, store_kwargs: dict[str, str | bool]) -> RemoteStore:
url = store_kwargs["url"]
async def store(self, store_kwargs: dict[str, str | bool | UPath]) -> RemoteStore:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This isn't actually async

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

correct, but the class we are inheriting from defines this as an async method

Comment threadsrc/zarr/abc/store.py Outdated
self._is_open = False
pass

async def _set_dict(self, dict: Mapping[str, Buffer]) -> None:

@jhammanjhammanAug 9, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

set_many() (analogous to insert_many)?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I moved away from set_dict and switched to _set_many

@jhammanjhamman added the V3 label Aug 9, 2024
@dcheriandcherian mentioned this pull request Aug 15, 2024
6 tasks
@d-v-b
d-v-b requested a review from jhammanSeptember 19, 2024 16:13
@d-v-b

Copy link
Copy Markdown
ContributorAuthor

we should try to get this in, because it fixes problematic list_prefix behavior that has been noted elsewhere.

@d-v-bd-v-b changed the title implement store.list_prefix and store._set_dictimplement store.list_prefix and store._set_manySep 19, 2024
@d-v-b
d-v-b merged commit 06e3215 into v3Sep 19, 2024
@d-v-b
d-v-b deleted the store-list-prefix branch September 19, 2024 16:26
dcherian added a commit to dcherian/zarr-python that referenced this pull request Sep 24, 2024
* v3:
chore: update pre-commit hooks (zarr-developers#2222)
fix: validate v3 dtypes when loading/creating v3 metadata (zarr-developers#2209)
fix typo in store integration test (zarr-developers#2223)
Basic Zarr-python 2.x compatibility changes (zarr-developers#2098)
Make Group.arrays, groups compatible with v2 (zarr-developers#2213)
Typing fixes to test_indexing (zarr-developers#2193)
Default to RemoteStore for fsspec URIs (zarr-developers#2198)
Make MemoryStore serialiazable (zarr-developers#2204)
[v3] Implement Group methods for empty, full, ones, and zeros (zarr-developers#2210)
implement `store.list_prefix` and `store._set_many` (zarr-developers#2064)
Fixed codec for v2 data with no fill value (zarr-developers#2207)
@jhammanjhamman added this to the 3.0.0.beta milestone Oct 17, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants

@d-v-b@martindurant@jhamman
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

implement store.list_prefix and store._set_many - #2064

Merged
d-v-b merged 9 commits into
v3from
store-list-prefix
Sep 19, 2024
Merged

implement store.list_prefix and store._set_many#2064
d-v-b merged 9 commits into
v3from
store-list-prefix

Conversation

@d-v-b

@d-v-bd-v-b commented Aug 3, 2024

Copy link
Copy Markdown
Contributor
  • fixes / implements list_prefix for stores. list_prefix(prefix=foo) now consistently returns keys with the shared prefix foo stripped. I'd be fine altering this to return absolute keys instead.
  • implements tests for list_prefix on the StoreTests base class
  • adds a _set_dict(Mapping[str, Buffer]) method to stores that allows declaring a collection of key: value pairs to write to storage. The primary usage is to make store tests simpler via a declarative API. I also suspect that making tests simpler anticipates making other code simpler. The default implementation of _set_dict simply wraps store.set, but it's easy to imagine fancier batching / transaction implementations.
  • alters some of the store tests to use _set_dict.
  • adds some missing type annotations in the remotestore tests

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

@d-v-b
d-v-b requested review from jhamman and martindurant and removed request for martindurantAugust 3, 2024 09:59
@d-v-b

d-v-b commented Aug 3, 2024

Copy link
Copy Markdown
ContributorAuthor

@martindurant i'm on flaky train internet and the github web API is giving me inconsistent signals about whether you have been requested for review; to be clear, I would appreciate your review :)

@d-v-bd-v-b mentioned this pull request Aug 3, 2024
6 tasks
@martindurant

Copy link
Copy Markdown
Member

list_prefix(prefix=foo) now consistently returns keys with the shared prefix foo stripped. I'd be fine altering this to return absolute keys instead.

What matters is how we intend to call it! fsspec likes to provide full paths (but has explicit prefix interfaces) - of course it's fine to tailor the situation to what we need, but I don't really know what that is.

I would appreciate your review

I may have got a notification or a few.

Comment threadsrc/zarr/abc/store.py Outdated
"""
for p in (self.root / prefix).rglob("*"):
if p.is_file():
yield str(p)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Were we getting duplicates?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

i think we were. this code path was not tested until this PR

Comment threadsrc/zarr/store/local.py Outdated
for p in (self.root / prefix).rglob("*"):
if p.is_file():
yield str(p).replace(to_strip, "")
yield str(p).removeprefix(to_strip)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1

Comment threadsrc/zarr/store/remote.py Outdated
for onefile in await self._fs._ls(prefix, detail=False):
yield onefile
find_str = "/".join([self.path, prefix])
for onefile in await self._fs._find(find_str):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The defaults for find are: maxdepth=None, withdirs=False, detail=False; maybe good to be specific.

Why is find() better than ls()? The former will return all child files, not just one level deep - is that the intent? If not, ls() ought to be generally more efficient.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

using find here is merely due to my ignorance of fsspec. I will implement ls as you suggest

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It depends on whether you want one directory level or everything below it. When I wrote the original, I didn't know the intent.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

i believe the intent here is to list everything below prefix (at least, that's how I'm using it)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I misunderstood your first comment. since the intent is to use the behavior of _find, I'm keeping it, but adding explicit kwargs as you suggested.

Comment threadsrc/zarr/sync.py
"""
result = []
async for x in data:
result.append(x)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

asyncio.gather? Like above, not much point in having coroutines if we serially wait for them.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think we can use asyncio.gather here, because AsyncGenerator is not iterable. Happy to be corrected, since I don't really know asyncio very well.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and for clarification, _collect_aiterator largely exists for convenience in testing, because I need some way to collect async generators when debugging with pdb. This function is not intended for use in anything performance sensitive.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should probably re-examine the use of async-iterators, though. If we can't gather() on them (seems to be true?), then they are the wrong abstraction since gather() is probably always what we actually want.

@martindurantmartindurantAug 3, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On second thoughts, maybe I'm wrong - does async for schedule all the coroutines at once?? Should be easy to test.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think it schedules them all at once. in [x async for x in async_generator], x is not an awaitable; it's already awaited. since the basic model of the generator is that it's a resumable, stateful iterator, I don't think we can schedule all the tasks at once.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the idea with the generators is to a) support seamless pagination and b) support pipelining (del_prefix will be able to take advantage of this at some point).

store_dict = dict(zip(keys, data_buf, strict=True))
await store._set_dict(store_dict)
for k, v in store_dict.items():
assert self.get(store, k).to_bytes() == v.to_bytes()

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if x.to_bytes() == y.to_bytes(), does x== y?

Isn't there a multiple get? Maybe not important here.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if x.to_bytes() == y.to_bytes(), does x== y?

no, and I suspect this might be deliberate since in principle Buffer instances can have identical bytes but different devices (e.g., gpu memory vs host memory); thus x == y might only be true if two buffers are bytes-equal and device-equal, but I'm speculating here. @madsbk would have a better answer I think.

Isn't there a multiple get? Maybe not important here.

there is no multiple get (nor a multiple set, nor a multiple delete).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@pytest.fixture(scope="function")
def store(self, store_kwargs: dict[str, str | bool]) -> RemoteStore:
url = store_kwargs["url"]
async def store(self, store_kwargs: dict[str, str | bool | UPath]) -> RemoteStore:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This isn't actually async

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

correct, but the class we are inheriting from defines this as an async method

Comment threadsrc/zarr/abc/store.py Outdated
self._is_open = False
pass

async def _set_dict(self, dict: Mapping[str, Buffer]) -> None:

@jhammanjhammanAug 9, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

set_many() (analogous to insert_many)?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I moved away from set_dict and switched to _set_many

@jhammanjhamman added the V3 label Aug 9, 2024
@dcheriandcherian mentioned this pull request Aug 15, 2024
6 tasks
@d-v-b
d-v-b requested a review from jhammanSeptember 19, 2024 16:13
@d-v-b

Copy link
Copy Markdown
ContributorAuthor

we should try to get this in, because it fixes problematic list_prefix behavior that has been noted elsewhere.

@d-v-bd-v-b changed the title implement store.list_prefix and store._set_dictimplement store.list_prefix and store._set_manySep 19, 2024
@d-v-b
d-v-b merged commit 06e3215 into v3Sep 19, 2024
@d-v-b
d-v-b deleted the store-list-prefix branch September 19, 2024 16:26
dcherian added a commit to dcherian/zarr-python that referenced this pull request Sep 24, 2024
* v3:
chore: update pre-commit hooks (zarr-developers#2222)
fix: validate v3 dtypes when loading/creating v3 metadata (zarr-developers#2209)
fix typo in store integration test (zarr-developers#2223)
Basic Zarr-python 2.x compatibility changes (zarr-developers#2098)
Make Group.arrays, groups compatible with v2 (zarr-developers#2213)
Typing fixes to test_indexing (zarr-developers#2193)
Default to RemoteStore for fsspec URIs (zarr-developers#2198)
Make MemoryStore serialiazable (zarr-developers#2204)
[v3] Implement Group methods for empty, full, ones, and zeros (zarr-developers#2210)
implement `store.list_prefix` and `store._set_many` (zarr-developers#2064)
Fixed codec for v2 data with no fill value (zarr-developers#2207)
@jhammanjhamman added this to the 3.0.0.beta milestone Oct 17, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants

@d-v-b@martindurant@jhamman
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

implement store.list_prefix and store._set_many - #2064

Merged
d-v-b merged 9 commits into
v3from
store-list-prefix
Sep 19, 2024
Merged

implement store.list_prefix and store._set_many#2064
d-v-b merged 9 commits into
v3from
store-list-prefix

Conversation

@d-v-b

@d-v-bd-v-b commented Aug 3, 2024

Copy link
Copy Markdown
Contributor
  • fixes / implements list_prefix for stores. list_prefix(prefix=foo) now consistently returns keys with the shared prefix foo stripped. I'd be fine altering this to return absolute keys instead.
  • implements tests for list_prefix on the StoreTests base class
  • adds a _set_dict(Mapping[str, Buffer]) method to stores that allows declaring a collection of key: value pairs to write to storage. The primary usage is to make store tests simpler via a declarative API. I also suspect that making tests simpler anticipates making other code simpler. The default implementation of _set_dict simply wraps store.set, but it's easy to imagine fancier batching / transaction implementations.
  • alters some of the store tests to use _set_dict.
  • adds some missing type annotations in the remotestore tests

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

@d-v-b
d-v-b requested review from jhamman and martindurant and removed request for martindurantAugust 3, 2024 09:59
@d-v-b

d-v-b commented Aug 3, 2024

Copy link
Copy Markdown
ContributorAuthor

@martindurant i'm on flaky train internet and the github web API is giving me inconsistent signals about whether you have been requested for review; to be clear, I would appreciate your review :)

@d-v-bd-v-b mentioned this pull request Aug 3, 2024
6 tasks
@martindurant

Copy link
Copy Markdown
Member

list_prefix(prefix=foo) now consistently returns keys with the shared prefix foo stripped. I'd be fine altering this to return absolute keys instead.

What matters is how we intend to call it! fsspec likes to provide full paths (but has explicit prefix interfaces) - of course it's fine to tailor the situation to what we need, but I don't really know what that is.

I would appreciate your review

I may have got a notification or a few.

Comment threadsrc/zarr/abc/store.py Outdated
"""
for p in (self.root / prefix).rglob("*"):
if p.is_file():
yield str(p)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Were we getting duplicates?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

i think we were. this code path was not tested until this PR

Comment threadsrc/zarr/store/local.py Outdated
for p in (self.root / prefix).rglob("*"):
if p.is_file():
yield str(p).replace(to_strip, "")
yield str(p).removeprefix(to_strip)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1

Comment threadsrc/zarr/store/remote.py Outdated
for onefile in await self._fs._ls(prefix, detail=False):
yield onefile
find_str = "/".join([self.path, prefix])
for onefile in await self._fs._find(find_str):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The defaults for find are: maxdepth=None, withdirs=False, detail=False; maybe good to be specific.

Why is find() better than ls()? The former will return all child files, not just one level deep - is that the intent? If not, ls() ought to be generally more efficient.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

using find here is merely due to my ignorance of fsspec. I will implement ls as you suggest

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It depends on whether you want one directory level or everything below it. When I wrote the original, I didn't know the intent.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

i believe the intent here is to list everything below prefix (at least, that's how I'm using it)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I misunderstood your first comment. since the intent is to use the behavior of _find, I'm keeping it, but adding explicit kwargs as you suggested.

Comment threadsrc/zarr/sync.py
"""
result = []
async for x in data:
result.append(x)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

asyncio.gather? Like above, not much point in having coroutines if we serially wait for them.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think we can use asyncio.gather here, because AsyncGenerator is not iterable. Happy to be corrected, since I don't really know asyncio very well.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and for clarification, _collect_aiterator largely exists for convenience in testing, because I need some way to collect async generators when debugging with pdb. This function is not intended for use in anything performance sensitive.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should probably re-examine the use of async-iterators, though. If we can't gather() on them (seems to be true?), then they are the wrong abstraction since gather() is probably always what we actually want.

@martindurantmartindurantAug 3, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On second thoughts, maybe I'm wrong - does async for schedule all the coroutines at once?? Should be easy to test.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think it schedules them all at once. in [x async for x in async_generator], x is not an awaitable; it's already awaited. since the basic model of the generator is that it's a resumable, stateful iterator, I don't think we can schedule all the tasks at once.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the idea with the generators is to a) support seamless pagination and b) support pipelining (del_prefix will be able to take advantage of this at some point).

store_dict = dict(zip(keys, data_buf, strict=True))
await store._set_dict(store_dict)
for k, v in store_dict.items():
assert self.get(store, k).to_bytes() == v.to_bytes()

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if x.to_bytes() == y.to_bytes(), does x== y?

Isn't there a multiple get? Maybe not important here.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if x.to_bytes() == y.to_bytes(), does x== y?

no, and I suspect this might be deliberate since in principle Buffer instances can have identical bytes but different devices (e.g., gpu memory vs host memory); thus x == y might only be true if two buffers are bytes-equal and device-equal, but I'm speculating here. @madsbk would have a better answer I think.

Isn't there a multiple get? Maybe not important here.

there is no multiple get (nor a multiple set, nor a multiple delete).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@pytest.fixture(scope="function")
def store(self, store_kwargs: dict[str, str | bool]) -> RemoteStore:
url = store_kwargs["url"]
async def store(self, store_kwargs: dict[str, str | bool | UPath]) -> RemoteStore:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This isn't actually async

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

correct, but the class we are inheriting from defines this as an async method

Comment threadsrc/zarr/abc/store.py Outdated
self._is_open = False
pass

async def _set_dict(self, dict: Mapping[str, Buffer]) -> None:

@jhammanjhammanAug 9, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

set_many() (analogous to insert_many)?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I moved away from set_dict and switched to _set_many

@jhammanjhamman added the V3 label Aug 9, 2024
@dcheriandcherian mentioned this pull request Aug 15, 2024
6 tasks
@d-v-b
d-v-b requested a review from jhammanSeptember 19, 2024 16:13
@d-v-b

Copy link
Copy Markdown
ContributorAuthor

we should try to get this in, because it fixes problematic list_prefix behavior that has been noted elsewhere.

@d-v-bd-v-b changed the title implement store.list_prefix and store._set_dictimplement store.list_prefix and store._set_manySep 19, 2024
@d-v-b
d-v-b merged commit 06e3215 into v3Sep 19, 2024
@d-v-b
d-v-b deleted the store-list-prefix branch September 19, 2024 16:26
dcherian added a commit to dcherian/zarr-python that referenced this pull request Sep 24, 2024
* v3:
chore: update pre-commit hooks (zarr-developers#2222)
fix: validate v3 dtypes when loading/creating v3 metadata (zarr-developers#2209)
fix typo in store integration test (zarr-developers#2223)
Basic Zarr-python 2.x compatibility changes (zarr-developers#2098)
Make Group.arrays, groups compatible with v2 (zarr-developers#2213)
Typing fixes to test_indexing (zarr-developers#2193)
Default to RemoteStore for fsspec URIs (zarr-developers#2198)
Make MemoryStore serialiazable (zarr-developers#2204)
[v3] Implement Group methods for empty, full, ones, and zeros (zarr-developers#2210)
implement `store.list_prefix` and `store._set_many` (zarr-developers#2064)
Fixed codec for v2 data with no fill value (zarr-developers#2207)
@jhammanjhamman added this to the 3.0.0.beta milestone Oct 17, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants

@d-v-b@martindurant@jhamman
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

implement store.list_prefix and store._set_many - #2064

Merged
d-v-b merged 9 commits into
v3from
store-list-prefix
Sep 19, 2024
Merged

implement store.list_prefix and store._set_many#2064
d-v-b merged 9 commits into
v3from
store-list-prefix

Conversation

@d-v-b

@d-v-bd-v-b commented Aug 3, 2024

Copy link
Copy Markdown
Contributor
  • fixes / implements list_prefix for stores. list_prefix(prefix=foo) now consistently returns keys with the shared prefix foo stripped. I'd be fine altering this to return absolute keys instead.
  • implements tests for list_prefix on the StoreTests base class
  • adds a _set_dict(Mapping[str, Buffer]) method to stores that allows declaring a collection of key: value pairs to write to storage. The primary usage is to make store tests simpler via a declarative API. I also suspect that making tests simpler anticipates making other code simpler. The default implementation of _set_dict simply wraps store.set, but it's easy to imagine fancier batching / transaction implementations.
  • alters some of the store tests to use _set_dict.
  • adds some missing type annotations in the remotestore tests

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

@d-v-b
d-v-b requested review from jhamman and martindurant and removed request for martindurantAugust 3, 2024 09:59
@d-v-b

d-v-b commented Aug 3, 2024

Copy link
Copy Markdown
ContributorAuthor

@martindurant i'm on flaky train internet and the github web API is giving me inconsistent signals about whether you have been requested for review; to be clear, I would appreciate your review :)

@d-v-bd-v-b mentioned this pull request Aug 3, 2024
6 tasks
@martindurant

Copy link
Copy Markdown
Member

list_prefix(prefix=foo) now consistently returns keys with the shared prefix foo stripped. I'd be fine altering this to return absolute keys instead.

What matters is how we intend to call it! fsspec likes to provide full paths (but has explicit prefix interfaces) - of course it's fine to tailor the situation to what we need, but I don't really know what that is.

I would appreciate your review

I may have got a notification or a few.

Comment threadsrc/zarr/abc/store.py Outdated
"""
for p in (self.root / prefix).rglob("*"):
if p.is_file():
yield str(p)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Were we getting duplicates?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

i think we were. this code path was not tested until this PR

Comment threadsrc/zarr/store/local.py Outdated
for p in (self.root / prefix).rglob("*"):
if p.is_file():
yield str(p).replace(to_strip, "")
yield str(p).removeprefix(to_strip)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1

Comment threadsrc/zarr/store/remote.py Outdated
for onefile in await self._fs._ls(prefix, detail=False):
yield onefile
find_str = "/".join([self.path, prefix])
for onefile in await self._fs._find(find_str):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The defaults for find are: maxdepth=None, withdirs=False, detail=False; maybe good to be specific.

Why is find() better than ls()? The former will return all child files, not just one level deep - is that the intent? If not, ls() ought to be generally more efficient.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

using find here is merely due to my ignorance of fsspec. I will implement ls as you suggest

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It depends on whether you want one directory level or everything below it. When I wrote the original, I didn't know the intent.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

i believe the intent here is to list everything below prefix (at least, that's how I'm using it)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I misunderstood your first comment. since the intent is to use the behavior of _find, I'm keeping it, but adding explicit kwargs as you suggested.

Comment threadsrc/zarr/sync.py
"""
result = []
async for x in data:
result.append(x)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

asyncio.gather? Like above, not much point in having coroutines if we serially wait for them.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think we can use asyncio.gather here, because AsyncGenerator is not iterable. Happy to be corrected, since I don't really know asyncio very well.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and for clarification, _collect_aiterator largely exists for convenience in testing, because I need some way to collect async generators when debugging with pdb. This function is not intended for use in anything performance sensitive.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should probably re-examine the use of async-iterators, though. If we can't gather() on them (seems to be true?), then they are the wrong abstraction since gather() is probably always what we actually want.

@martindurantmartindurantAug 3, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On second thoughts, maybe I'm wrong - does async for schedule all the coroutines at once?? Should be easy to test.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think it schedules them all at once. in [x async for x in async_generator], x is not an awaitable; it's already awaited. since the basic model of the generator is that it's a resumable, stateful iterator, I don't think we can schedule all the tasks at once.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the idea with the generators is to a) support seamless pagination and b) support pipelining (del_prefix will be able to take advantage of this at some point).

store_dict = dict(zip(keys, data_buf, strict=True))
await store._set_dict(store_dict)
for k, v in store_dict.items():
assert self.get(store, k).to_bytes() == v.to_bytes()

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if x.to_bytes() == y.to_bytes(), does x== y?

Isn't there a multiple get? Maybe not important here.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if x.to_bytes() == y.to_bytes(), does x== y?

no, and I suspect this might be deliberate since in principle Buffer instances can have identical bytes but different devices (e.g., gpu memory vs host memory); thus x == y might only be true if two buffers are bytes-equal and device-equal, but I'm speculating here. @madsbk would have a better answer I think.

Isn't there a multiple get? Maybe not important here.

there is no multiple get (nor a multiple set, nor a multiple delete).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@pytest.fixture(scope="function")
def store(self, store_kwargs: dict[str, str | bool]) -> RemoteStore:
url = store_kwargs["url"]
async def store(self, store_kwargs: dict[str, str | bool | UPath]) -> RemoteStore:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This isn't actually async

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

correct, but the class we are inheriting from defines this as an async method

Comment threadsrc/zarr/abc/store.py Outdated
self._is_open = False
pass

async def _set_dict(self, dict: Mapping[str, Buffer]) -> None:

@jhammanjhammanAug 9, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

set_many() (analogous to insert_many)?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I moved away from set_dict and switched to _set_many

@jhammanjhamman added the V3 label Aug 9, 2024
@dcheriandcherian mentioned this pull request Aug 15, 2024
6 tasks
@d-v-b
d-v-b requested a review from jhammanSeptember 19, 2024 16:13
@d-v-b

Copy link
Copy Markdown
ContributorAuthor

we should try to get this in, because it fixes problematic list_prefix behavior that has been noted elsewhere.

@d-v-bd-v-b changed the title implement store.list_prefix and store._set_dictimplement store.list_prefix and store._set_manySep 19, 2024
@d-v-b
d-v-b merged commit 06e3215 into v3Sep 19, 2024
@d-v-b
d-v-b deleted the store-list-prefix branch September 19, 2024 16:26
dcherian added a commit to dcherian/zarr-python that referenced this pull request Sep 24, 2024
* v3:
chore: update pre-commit hooks (zarr-developers#2222)
fix: validate v3 dtypes when loading/creating v3 metadata (zarr-developers#2209)
fix typo in store integration test (zarr-developers#2223)
Basic Zarr-python 2.x compatibility changes (zarr-developers#2098)
Make Group.arrays, groups compatible with v2 (zarr-developers#2213)
Typing fixes to test_indexing (zarr-developers#2193)
Default to RemoteStore for fsspec URIs (zarr-developers#2198)
Make MemoryStore serialiazable (zarr-developers#2204)
[v3] Implement Group methods for empty, full, ones, and zeros (zarr-developers#2210)
implement `store.list_prefix` and `store._set_many` (zarr-developers#2064)
Fixed codec for v2 data with no fill value (zarr-developers#2207)
@jhammanjhamman added this to the 3.0.0.beta milestone Oct 17, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants

@d-v-b@martindurant@jhamman
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

implement store.list_prefix and store._set_many - #2064

Merged
d-v-b merged 9 commits into
v3from
store-list-prefix
Sep 19, 2024
Merged

implement store.list_prefix and store._set_many#2064
d-v-b merged 9 commits into
v3from
store-list-prefix

Conversation

@d-v-b

@d-v-bd-v-b commented Aug 3, 2024

Copy link
Copy Markdown
Contributor
  • fixes / implements list_prefix for stores. list_prefix(prefix=foo) now consistently returns keys with the shared prefix foo stripped. I'd be fine altering this to return absolute keys instead.
  • implements tests for list_prefix on the StoreTests base class
  • adds a _set_dict(Mapping[str, Buffer]) method to stores that allows declaring a collection of key: value pairs to write to storage. The primary usage is to make store tests simpler via a declarative API. I also suspect that making tests simpler anticipates making other code simpler. The default implementation of _set_dict simply wraps store.set, but it's easy to imagine fancier batching / transaction implementations.
  • alters some of the store tests to use _set_dict.
  • adds some missing type annotations in the remotestore tests

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

@d-v-b
d-v-b requested review from jhamman and martindurant and removed request for martindurantAugust 3, 2024 09:59
@d-v-b

d-v-b commented Aug 3, 2024

Copy link
Copy Markdown
ContributorAuthor

@martindurant i'm on flaky train internet and the github web API is giving me inconsistent signals about whether you have been requested for review; to be clear, I would appreciate your review :)

@d-v-bd-v-b mentioned this pull request Aug 3, 2024
6 tasks
@martindurant

Copy link
Copy Markdown
Member

list_prefix(prefix=foo) now consistently returns keys with the shared prefix foo stripped. I'd be fine altering this to return absolute keys instead.

What matters is how we intend to call it! fsspec likes to provide full paths (but has explicit prefix interfaces) - of course it's fine to tailor the situation to what we need, but I don't really know what that is.

I would appreciate your review

I may have got a notification or a few.

Comment threadsrc/zarr/abc/store.py Outdated
"""
for p in (self.root / prefix).rglob("*"):
if p.is_file():
yield str(p)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Were we getting duplicates?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

i think we were. this code path was not tested until this PR

Comment threadsrc/zarr/store/local.py Outdated
for p in (self.root / prefix).rglob("*"):
if p.is_file():
yield str(p).replace(to_strip, "")
yield str(p).removeprefix(to_strip)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1

Comment threadsrc/zarr/store/remote.py Outdated
for onefile in await self._fs._ls(prefix, detail=False):
yield onefile
find_str = "/".join([self.path, prefix])
for onefile in await self._fs._find(find_str):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The defaults for find are: maxdepth=None, withdirs=False, detail=False; maybe good to be specific.

Why is find() better than ls()? The former will return all child files, not just one level deep - is that the intent? If not, ls() ought to be generally more efficient.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

using find here is merely due to my ignorance of fsspec. I will implement ls as you suggest

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It depends on whether you want one directory level or everything below it. When I wrote the original, I didn't know the intent.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

i believe the intent here is to list everything below prefix (at least, that's how I'm using it)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I misunderstood your first comment. since the intent is to use the behavior of _find, I'm keeping it, but adding explicit kwargs as you suggested.

Comment threadsrc/zarr/sync.py
"""
result = []
async for x in data:
result.append(x)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

asyncio.gather? Like above, not much point in having coroutines if we serially wait for them.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think we can use asyncio.gather here, because AsyncGenerator is not iterable. Happy to be corrected, since I don't really know asyncio very well.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and for clarification, _collect_aiterator largely exists for convenience in testing, because I need some way to collect async generators when debugging with pdb. This function is not intended for use in anything performance sensitive.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should probably re-examine the use of async-iterators, though. If we can't gather() on them (seems to be true?), then they are the wrong abstraction since gather() is probably always what we actually want.

@martindurantmartindurantAug 3, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On second thoughts, maybe I'm wrong - does async for schedule all the coroutines at once?? Should be easy to test.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think it schedules them all at once. in [x async for x in async_generator], x is not an awaitable; it's already awaited. since the basic model of the generator is that it's a resumable, stateful iterator, I don't think we can schedule all the tasks at once.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the idea with the generators is to a) support seamless pagination and b) support pipelining (del_prefix will be able to take advantage of this at some point).

store_dict = dict(zip(keys, data_buf, strict=True))
await store._set_dict(store_dict)
for k, v in store_dict.items():
assert self.get(store, k).to_bytes() == v.to_bytes()

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if x.to_bytes() == y.to_bytes(), does x== y?

Isn't there a multiple get? Maybe not important here.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if x.to_bytes() == y.to_bytes(), does x== y?

no, and I suspect this might be deliberate since in principle Buffer instances can have identical bytes but different devices (e.g., gpu memory vs host memory); thus x == y might only be true if two buffers are bytes-equal and device-equal, but I'm speculating here. @madsbk would have a better answer I think.

Isn't there a multiple get? Maybe not important here.

there is no multiple get (nor a multiple set, nor a multiple delete).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@pytest.fixture(scope="function")
def store(self, store_kwargs: dict[str, str | bool]) -> RemoteStore:
url = store_kwargs["url"]
async def store(self, store_kwargs: dict[str, str | bool | UPath]) -> RemoteStore:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This isn't actually async

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

correct, but the class we are inheriting from defines this as an async method

Comment threadsrc/zarr/abc/store.py Outdated
self._is_open = False
pass

async def _set_dict(self, dict: Mapping[str, Buffer]) -> None:

@jhammanjhammanAug 9, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

set_many() (analogous to insert_many)?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I moved away from set_dict and switched to _set_many

@jhammanjhamman added the V3 label Aug 9, 2024
@dcheriandcherian mentioned this pull request Aug 15, 2024
6 tasks
@d-v-b
d-v-b requested a review from jhammanSeptember 19, 2024 16:13
@d-v-b

Copy link
Copy Markdown
ContributorAuthor

we should try to get this in, because it fixes problematic list_prefix behavior that has been noted elsewhere.

@d-v-bd-v-b changed the title implement store.list_prefix and store._set_dictimplement store.list_prefix and store._set_manySep 19, 2024
@d-v-b
d-v-b merged commit 06e3215 into v3Sep 19, 2024
@d-v-b
d-v-b deleted the store-list-prefix branch September 19, 2024 16:26
dcherian added a commit to dcherian/zarr-python that referenced this pull request Sep 24, 2024
* v3:
chore: update pre-commit hooks (zarr-developers#2222)
fix: validate v3 dtypes when loading/creating v3 metadata (zarr-developers#2209)
fix typo in store integration test (zarr-developers#2223)
Basic Zarr-python 2.x compatibility changes (zarr-developers#2098)
Make Group.arrays, groups compatible with v2 (zarr-developers#2213)
Typing fixes to test_indexing (zarr-developers#2193)
Default to RemoteStore for fsspec URIs (zarr-developers#2198)
Make MemoryStore serialiazable (zarr-developers#2204)
[v3] Implement Group methods for empty, full, ones, and zeros (zarr-developers#2210)
implement `store.list_prefix` and `store._set_many` (zarr-developers#2064)
Fixed codec for v2 data with no fill value (zarr-developers#2207)
@jhammanjhamman added this to the 3.0.0.beta milestone Oct 17, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants

@d-v-b@martindurant@jhamman
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

implement store.list_prefix and store._set_many - #2064

Merged
d-v-b merged 9 commits into
v3from
store-list-prefix
Sep 19, 2024
Merged

implement store.list_prefix and store._set_many#2064
d-v-b merged 9 commits into
v3from
store-list-prefix

Conversation

@d-v-b

@d-v-bd-v-b commented Aug 3, 2024

Copy link
Copy Markdown
Contributor
  • fixes / implements list_prefix for stores. list_prefix(prefix=foo) now consistently returns keys with the shared prefix foo stripped. I'd be fine altering this to return absolute keys instead.
  • implements tests for list_prefix on the StoreTests base class
  • adds a _set_dict(Mapping[str, Buffer]) method to stores that allows declaring a collection of key: value pairs to write to storage. The primary usage is to make store tests simpler via a declarative API. I also suspect that making tests simpler anticipates making other code simpler. The default implementation of _set_dict simply wraps store.set, but it's easy to imagine fancier batching / transaction implementations.
  • alters some of the store tests to use _set_dict.
  • adds some missing type annotations in the remotestore tests

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

@d-v-b
d-v-b requested review from jhamman and martindurant and removed request for martindurantAugust 3, 2024 09:59
@d-v-b

d-v-b commented Aug 3, 2024

Copy link
Copy Markdown
ContributorAuthor

@martindurant i'm on flaky train internet and the github web API is giving me inconsistent signals about whether you have been requested for review; to be clear, I would appreciate your review :)

@d-v-bd-v-b mentioned this pull request Aug 3, 2024
6 tasks
@martindurant

Copy link
Copy Markdown
Member

list_prefix(prefix=foo) now consistently returns keys with the shared prefix foo stripped. I'd be fine altering this to return absolute keys instead.

What matters is how we intend to call it! fsspec likes to provide full paths (but has explicit prefix interfaces) - of course it's fine to tailor the situation to what we need, but I don't really know what that is.

I would appreciate your review

I may have got a notification or a few.

Comment threadsrc/zarr/abc/store.py Outdated
"""
for p in (self.root / prefix).rglob("*"):
if p.is_file():
yield str(p)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Were we getting duplicates?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

i think we were. this code path was not tested until this PR

Comment threadsrc/zarr/store/local.py Outdated
for p in (self.root / prefix).rglob("*"):
if p.is_file():
yield str(p).replace(to_strip, "")
yield str(p).removeprefix(to_strip)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1

Comment threadsrc/zarr/store/remote.py Outdated
for onefile in await self._fs._ls(prefix, detail=False):
yield onefile
find_str = "/".join([self.path, prefix])
for onefile in await self._fs._find(find_str):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The defaults for find are: maxdepth=None, withdirs=False, detail=False; maybe good to be specific.

Why is find() better than ls()? The former will return all child files, not just one level deep - is that the intent? If not, ls() ought to be generally more efficient.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

using find here is merely due to my ignorance of fsspec. I will implement ls as you suggest

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It depends on whether you want one directory level or everything below it. When I wrote the original, I didn't know the intent.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

i believe the intent here is to list everything below prefix (at least, that's how I'm using it)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I misunderstood your first comment. since the intent is to use the behavior of _find, I'm keeping it, but adding explicit kwargs as you suggested.

Comment threadsrc/zarr/sync.py
"""
result = []
async for x in data:
result.append(x)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

asyncio.gather? Like above, not much point in having coroutines if we serially wait for them.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think we can use asyncio.gather here, because AsyncGenerator is not iterable. Happy to be corrected, since I don't really know asyncio very well.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and for clarification, _collect_aiterator largely exists for convenience in testing, because I need some way to collect async generators when debugging with pdb. This function is not intended for use in anything performance sensitive.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should probably re-examine the use of async-iterators, though. If we can't gather() on them (seems to be true?), then they are the wrong abstraction since gather() is probably always what we actually want.

@martindurantmartindurantAug 3, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On second thoughts, maybe I'm wrong - does async for schedule all the coroutines at once?? Should be easy to test.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think it schedules them all at once. in [x async for x in async_generator], x is not an awaitable; it's already awaited. since the basic model of the generator is that it's a resumable, stateful iterator, I don't think we can schedule all the tasks at once.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the idea with the generators is to a) support seamless pagination and b) support pipelining (del_prefix will be able to take advantage of this at some point).

store_dict = dict(zip(keys, data_buf, strict=True))
await store._set_dict(store_dict)
for k, v in store_dict.items():
assert self.get(store, k).to_bytes() == v.to_bytes()

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if x.to_bytes() == y.to_bytes(), does x== y?

Isn't there a multiple get? Maybe not important here.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if x.to_bytes() == y.to_bytes(), does x== y?

no, and I suspect this might be deliberate since in principle Buffer instances can have identical bytes but different devices (e.g., gpu memory vs host memory); thus x == y might only be true if two buffers are bytes-equal and device-equal, but I'm speculating here. @madsbk would have a better answer I think.

Isn't there a multiple get? Maybe not important here.

there is no multiple get (nor a multiple set, nor a multiple delete).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@pytest.fixture(scope="function")
def store(self, store_kwargs: dict[str, str | bool]) -> RemoteStore:
url = store_kwargs["url"]
async def store(self, store_kwargs: dict[str, str | bool | UPath]) -> RemoteStore:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This isn't actually async

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

correct, but the class we are inheriting from defines this as an async method

Comment threadsrc/zarr/abc/store.py Outdated
self._is_open = False
pass

async def _set_dict(self, dict: Mapping[str, Buffer]) -> None:

@jhammanjhammanAug 9, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

set_many() (analogous to insert_many)?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I moved away from set_dict and switched to _set_many

@jhammanjhamman added the V3 label Aug 9, 2024
@dcheriandcherian mentioned this pull request Aug 15, 2024
6 tasks
@d-v-b
d-v-b requested a review from jhammanSeptember 19, 2024 16:13
@d-v-b

Copy link
Copy Markdown
ContributorAuthor

we should try to get this in, because it fixes problematic list_prefix behavior that has been noted elsewhere.

@d-v-bd-v-b changed the title implement store.list_prefix and store._set_dictimplement store.list_prefix and store._set_manySep 19, 2024
@d-v-b
d-v-b merged commit 06e3215 into v3Sep 19, 2024
@d-v-b
d-v-b deleted the store-list-prefix branch September 19, 2024 16:26
dcherian added a commit to dcherian/zarr-python that referenced this pull request Sep 24, 2024
* v3:
chore: update pre-commit hooks (zarr-developers#2222)
fix: validate v3 dtypes when loading/creating v3 metadata (zarr-developers#2209)
fix typo in store integration test (zarr-developers#2223)
Basic Zarr-python 2.x compatibility changes (zarr-developers#2098)
Make Group.arrays, groups compatible with v2 (zarr-developers#2213)
Typing fixes to test_indexing (zarr-developers#2193)
Default to RemoteStore for fsspec URIs (zarr-developers#2198)
Make MemoryStore serialiazable (zarr-developers#2204)
[v3] Implement Group methods for empty, full, ones, and zeros (zarr-developers#2210)
implement `store.list_prefix` and `store._set_many` (zarr-developers#2064)
Fixed codec for v2 data with no fill value (zarr-developers#2207)
@jhammanjhamman added this to the 3.0.0.beta milestone Oct 17, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants

@d-v-b@martindurant@jhamman
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

implement store.list_prefix and store._set_many - #2064

Merged
d-v-b merged 9 commits into
v3from
store-list-prefix
Sep 19, 2024
Merged

implement store.list_prefix and store._set_many#2064
d-v-b merged 9 commits into
v3from
store-list-prefix

Conversation

@d-v-b

@d-v-bd-v-b commented Aug 3, 2024

Copy link
Copy Markdown
Contributor
  • fixes / implements list_prefix for stores. list_prefix(prefix=foo) now consistently returns keys with the shared prefix foo stripped. I'd be fine altering this to return absolute keys instead.
  • implements tests for list_prefix on the StoreTests base class
  • adds a _set_dict(Mapping[str, Buffer]) method to stores that allows declaring a collection of key: value pairs to write to storage. The primary usage is to make store tests simpler via a declarative API. I also suspect that making tests simpler anticipates making other code simpler. The default implementation of _set_dict simply wraps store.set, but it's easy to imagine fancier batching / transaction implementations.
  • alters some of the store tests to use _set_dict.
  • adds some missing type annotations in the remotestore tests

TODO:

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/tutorial.rst
  • Changes documented in docs/release.rst
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

@d-v-b
d-v-b requested review from jhamman and martindurant and removed request for martindurantAugust 3, 2024 09:59
@d-v-b

d-v-b commented Aug 3, 2024

Copy link
Copy Markdown
ContributorAuthor

@martindurant i'm on flaky train internet and the github web API is giving me inconsistent signals about whether you have been requested for review; to be clear, I would appreciate your review :)

@d-v-bd-v-b mentioned this pull request Aug 3, 2024
6 tasks
@martindurant

Copy link
Copy Markdown
Member

list_prefix(prefix=foo) now consistently returns keys with the shared prefix foo stripped. I'd be fine altering this to return absolute keys instead.

What matters is how we intend to call it! fsspec likes to provide full paths (but has explicit prefix interfaces) - of course it's fine to tailor the situation to what we need, but I don't really know what that is.

I would appreciate your review

I may have got a notification or a few.

Comment threadsrc/zarr/abc/store.py Outdated
"""
for p in (self.root / prefix).rglob("*"):
if p.is_file():
yield str(p)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Were we getting duplicates?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

i think we were. this code path was not tested until this PR

Comment threadsrc/zarr/store/local.py Outdated
for p in (self.root / prefix).rglob("*"):
if p.is_file():
yield str(p).replace(to_strip, "")
yield str(p).removeprefix(to_strip)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1

Comment threadsrc/zarr/store/remote.py Outdated
for onefile in await self._fs._ls(prefix, detail=False):
yield onefile
find_str = "/".join([self.path, prefix])
for onefile in await self._fs._find(find_str):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The defaults for find are: maxdepth=None, withdirs=False, detail=False; maybe good to be specific.

Why is find() better than ls()? The former will return all child files, not just one level deep - is that the intent? If not, ls() ought to be generally more efficient.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

using find here is merely due to my ignorance of fsspec. I will implement ls as you suggest

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It depends on whether you want one directory level or everything below it. When I wrote the original, I didn't know the intent.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

i believe the intent here is to list everything below prefix (at least, that's how I'm using it)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I misunderstood your first comment. since the intent is to use the behavior of _find, I'm keeping it, but adding explicit kwargs as you suggested.

Comment threadsrc/zarr/sync.py
"""
result = []
async for x in data:
result.append(x)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

asyncio.gather? Like above, not much point in having coroutines if we serially wait for them.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think we can use asyncio.gather here, because AsyncGenerator is not iterable. Happy to be corrected, since I don't really know asyncio very well.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and for clarification, _collect_aiterator largely exists for convenience in testing, because I need some way to collect async generators when debugging with pdb. This function is not intended for use in anything performance sensitive.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should probably re-examine the use of async-iterators, though. If we can't gather() on them (seems to be true?), then they are the wrong abstraction since gather() is probably always what we actually want.

@martindurantmartindurantAug 3, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On second thoughts, maybe I'm wrong - does async for schedule all the coroutines at once?? Should be easy to test.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think it schedules them all at once. in [x async for x in async_generator], x is not an awaitable; it's already awaited. since the basic model of the generator is that it's a resumable, stateful iterator, I don't think we can schedule all the tasks at once.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the idea with the generators is to a) support seamless pagination and b) support pipelining (del_prefix will be able to take advantage of this at some point).

store_dict = dict(zip(keys, data_buf, strict=True))
await store._set_dict(store_dict)
for k, v in store_dict.items():
assert self.get(store, k).to_bytes() == v.to_bytes()

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if x.to_bytes() == y.to_bytes(), does x== y?

Isn't there a multiple get? Maybe not important here.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if x.to_bytes() == y.to_bytes(), does x== y?

no, and I suspect this might be deliberate since in principle Buffer instances can have identical bytes but different devices (e.g., gpu memory vs host memory); thus x == y might only be true if two buffers are bytes-equal and device-equal, but I'm speculating here. @madsbk would have a better answer I think.

Isn't there a multiple get? Maybe not important here.

there is no multiple get (nor a multiple set, nor a multiple delete).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@pytest.fixture(scope="function")
def store(self, store_kwargs: dict[str, str | bool]) -> RemoteStore:
url = store_kwargs["url"]
async def store(self, store_kwargs: dict[str, str | bool | UPath]) -> RemoteStore:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This isn't actually async

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

correct, but the class we are inheriting from defines this as an async method

Comment threadsrc/zarr/abc/store.py Outdated
self._is_open = False
pass

async def _set_dict(self, dict: Mapping[str, Buffer]) -> None:

@jhammanjhammanAug 9, 2024

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

set_many() (analogous to insert_many)?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I moved away from set_dict and switched to _set_many

@jhammanjhamman added the V3 label Aug 9, 2024
@dcheriandcherian mentioned this pull request Aug 15, 2024
6 tasks
@d-v-b
d-v-b requested a review from jhammanSeptember 19, 2024 16:13
@d-v-b

Copy link
Copy Markdown
ContributorAuthor

we should try to get this in, because it fixes problematic list_prefix behavior that has been noted elsewhere.

@d-v-bd-v-b changed the title implement store.list_prefix and store._set_dictimplement store.list_prefix and store._set_manySep 19, 2024
@d-v-b
d-v-b merged commit 06e3215 into v3Sep 19, 2024
@d-v-b
d-v-b deleted the store-list-prefix branch September 19, 2024 16:26
dcherian added a commit to dcherian/zarr-python that referenced this pull request Sep 24, 2024
* v3:
chore: update pre-commit hooks (zarr-developers#2222)
fix: validate v3 dtypes when loading/creating v3 metadata (zarr-developers#2209)
fix typo in store integration test (zarr-developers#2223)
Basic Zarr-python 2.x compatibility changes (zarr-developers#2098)
Make Group.arrays, groups compatible with v2 (zarr-developers#2213)
Typing fixes to test_indexing (zarr-developers#2193)
Default to RemoteStore for fsspec URIs (zarr-developers#2198)
Make MemoryStore serialiazable (zarr-developers#2204)
[v3] Implement Group methods for empty, full, ones, and zeros (zarr-developers#2210)
implement `store.list_prefix` and `store._set_many` (zarr-developers#2064)
Fixed codec for v2 data with no fill value (zarr-developers#2207)
@jhammanjhamman added this to the 3.0.0.beta milestone Oct 17, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants

@d-v-b@martindurant@jhamman