Skip to content

Ensure file is created when it doesnt exist - #5

Merged
auxym merged 13 commits into
mainfrom
ensure-file-created
Jul 24, 2026
Merged

Ensure file is created when it doesnt exist#5
auxym merged 13 commits into
mainfrom
ensure-file-created

Conversation

@auxym

Copy link
Copy Markdown
Owner

Resolves#3

  • Change mode=rw to mode=rwc
  • Add unit test

- Change mode=rw to mode=rwc
- Add unit test
@auxym

Copy link
Copy Markdown
OwnerAuthor

@newville can you please confirm this fixes the issue for you also?

@newville

Copy link
Copy Markdown

@auxym Thanks -- This also needs the if TYPE_CHECKiNG test removed, and those imports always done.

I think a basic test, with no async code, such as

importzarrfromzarr_sqliteimportSQLiteStorefname='test1.zarrdb'store=SQLiteStore(fname, read_only=False)
root=zarr.open(store=store, mode='a')
group1=root.create_group('group1')
i=np.arange(80000)/60.0x=i+np.random.normal(size=len(i), scale=1.5)
y=np.sin(x/13) +0.7*np.cos(x/47) +np.random.normal(size=len(i), scale=0.05)
x.shape= (400, 200)
y.shape= (200, 400)
group1.create_array(data=x, name='xdat')
group1.create_array(data=y, name='ydat')
print(group1, list(group1.keys()))
##print("wrote data, now reading....")
time.sleep(0.25)
read_root=zarr.open(store=SQLiteStore(fname, read_only=True), mode='r')
ytest=read_root['group1/ydat'][()]
print(f"## DONE {read_root=}{ytest=}")

Something like this to create, write, and read should probably be included in the tests.

I find the use of async sort of odd here. I effectively never use async.

I would even say that the

read_root=zarr.open(store=SQLiteStore(fname, read_only=True), mode='r')

is sort of a wart here. The more obvious (to me anyway)

read_root=zarr.open(store=fname, mode='r')

fails.

@auxym

Copy link
Copy Markdown
OwnerAuthor

I agree on most things, will look into including your proposed test and the import fix. I also agree on async, I rarely use it, but I believe it is required to satisfy the zarr-python store abstract base class: https://github.com/zarr-developers/zarr-python/blob/main/src/zarr/abc/store.py

read_root = zarr.open(store=fname, mode='r')

I don't understand how this should work, how would zarr know it should use the SQLiteStore class to create the store based on only a file name?

@auxym

auxym commented Jul 17, 2026

Copy link
Copy Markdown
OwnerAuthor

OK, I have implemented your suggested test, see latest commit. It passes (even without the sleep).

For the imports, looking at it, I actually don't think it's necessary to remove the type_checking check, and I assume it could save time during module imports to leave it in place. You can check yourself: you can completely delete the if TYPE_CHECKING block and run the tests (uv run pytest) and every test still passes. Can you confirm that works for you also?

@newville

Copy link
Copy Markdown

TYPE_CHECKING is False at runtime (https://docs.python.org/3/library/typing.html#typing.TYPE_CHECKING).

So the code is not importing collections.abc.Iterable at runtime. But that is needed for more than type checking.
SQLiteStore.list() and SQLiteStore.list_prefix() cannot work without it.

I'll admit that I have little experience with (or love for) async, that I don't know why (from def list())

@overrideasyncdeflist(self) ->AsyncIterator[str]:
cur=awaitself._execute("SELECT k FROM zarr")
forrowincast(Iterable[tuple[str]], cur):
yieldrow[0]

is not (I do not mean "just"):

deflist(self):
return [row[0] forrowinself._execute("SELECT k FROM zarr")]

I can believe there is a reason for the extra code, and that cast(Iterable[tuple[str]], cur) is doing something useful, but I really do not know what that is. (And, like, the name is literally"list".... "returns a list" seems so obvious to me that I must be profoundly missing something and too old for all this fancy.... stuff???).

But, if you're going to use cast(Iterable[tuple[str]], cur), then, you have to import Iterable.

I would just remove the check for TYPE_CHECKING. You can time the imports -- you will find they are insignificant.

read_root = zarr.open(store=fname, mode='r')

I don't understand how this should work, how would zarr know it should use the SQLiteStore class to create the store based on only a file name?

Yeah, I guess this is a complaint about Zarr, not SQLiteStore.

root=zarr.open('store.zarr', mode='r')

works for a FileStore, but that wouldn't work for a ZipStore either. It seems like a registry of file types would be needed.

@auxym

auxym commented Jul 19, 2026

Copy link
Copy Markdown
OwnerAuthor

Oh yeah, missed those cast calls. The existing minimal tests did not test that code path. I added a whole bunch of tests (with LLM assistance) and I completely removed the if TYPE_CHECKING as you suggested.

Regarding the async stuff, my understanding is that zarr-python uses async "under the hood", but wraps a sync API around it. Thus the user of zarr can use either sync or async calls.

@newville

Copy link
Copy Markdown

@auxym Thanks.

It seems to me that the point of cast is to tell the type checker not to check types. ;) What happens if you replace for row in cast(Iterable[tuple[str]], cur): with for row in cur:? Does some IDE complain?

Regarding the async stuff, my understanding is that zarr-python uses async "under the hood", but wraps a sync API around it. Thus the user of zarr can use either sync or async calls.

I can believe all of that.

I admit a bias against async (I support a lot of Python projects for data collection and analysis, and I just never use async, even when asynchronous callbacks from a distributed control system are a vital part of the software. Threads for specific tasks, multiprocessing, yes definitely).

Async adds a bunch of keywords all over Python to bolt on coroutines that are really about asynchronous I/O (and not anything else), and then requires adding a bunch more keywords ("await"s) all over Python code to say "don't be asynchronous. You can ignore this as a rant, but I think hat

deflist(self):
return [row[0] forrowinself._execute("select k from zarr")]

has merit that

@overrideasyncdeflist(self) ->AsyncIterator[str]:
cur=awaitself._execute("SELECT k FROM zarr")
forrowincast(Iterable[tuple[str]], cur):
yieldrow[0]

does not necessarily improve. There is a bunch of code to signal non-runtime type-checkers to ignore type checking. There is code to "await" the (allegedly) expensive database lookup, and then yields over a list of values that has already been synchronously fetched. [If "cast key to str" is needed, then I think AsyncIterator[str] is not guaranteed. Perhaps the yielded value should be str(row[0])?].

When I look in detail here, it seems that all of the calls to self._execute() and self._execute_write() are awaited. Doesn't that mean these functions could be synchronous, and sort of really are synchronous?
I'll also guess that for SQLite with WAL, these database queries (simple select in a two-column table) will be fast.
I'm comfortable saying that any claim about speed needs real benchmarks, but I'll say that goes for the claim that "using async makes stuff faster".

Similarly, where the code has await self._ensure_open(), couldn't that be if not self.is_open: open()?

Feel free to ignore all this.

@auxym

Copy link
Copy Markdown
OwnerAuthor

Yes, I believe the casts were necessary to stop mypy/pyright from complaining.

That said,

Perhaps the yielded value should be str(row[0])?].

I do like that suggestion. Latest commit got rid of all casts :)

When I look in detail here, it seems that all of the calls to self._execute() and self._execute_write()

Yeah, I wrapped all db transactions in those methods. Yes, they are currently blocking (non-async), since the stdlib sqlite3 is blocking. The goal of putting all transactions in a central place is that I aim, in the future, to implement real async transactions, either using aiosqlite3 (3rd party lib) or implementing our own background thread for db work (probably the former). I also personally do not use async, but zarr itself offers async everywhere, so it would be a better match for that pattern.

Similarly, where the code has await self._ensure_open(), couldn't that be if not self.is_open: open()?

_ensure_open is from zarr's store base class, once again I'm trying to follow zarr-python's existing patterns.

https://github.com/zarr-developers/zarr-python/blob/50b7e016590d76011382771fd52843949782c9c0/src/zarr/abc/store.py#L140

@auxym
auxym merged commit ef46459 into mainJul 24, 2026
@auxym
auxym deleted the ensure-file-created branch July 24, 2026 01:03
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

trouble using with synchronous access

2 participants

@auxym@newville
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Ensure file is created when it doesnt exist by auxym · Pull Request #5 · auxym/zarr-sqlite-python · GitHub
Skip to content

Ensure file is created when it doesnt exist - #5

Merged
auxym merged 13 commits into
mainfrom
ensure-file-created
Jul 24, 2026
Merged

Ensure file is created when it doesnt exist#5
auxym merged 13 commits into
mainfrom
ensure-file-created

Conversation

@auxym

Copy link
Copy Markdown
Owner

Resolves#3

  • Change mode=rw to mode=rwc
  • Add unit test

- Change mode=rw to mode=rwc
- Add unit test
@auxym

Copy link
Copy Markdown
OwnerAuthor

@newville can you please confirm this fixes the issue for you also?

@newville

Copy link
Copy Markdown

@auxym Thanks -- This also needs the if TYPE_CHECKiNG test removed, and those imports always done.

I think a basic test, with no async code, such as

importzarrfromzarr_sqliteimportSQLiteStorefname='test1.zarrdb'store=SQLiteStore(fname, read_only=False)
root=zarr.open(store=store, mode='a')
group1=root.create_group('group1')
i=np.arange(80000)/60.0x=i+np.random.normal(size=len(i), scale=1.5)
y=np.sin(x/13) +0.7*np.cos(x/47) +np.random.normal(size=len(i), scale=0.05)
x.shape= (400, 200)
y.shape= (200, 400)
group1.create_array(data=x, name='xdat')
group1.create_array(data=y, name='ydat')
print(group1, list(group1.keys()))
##print("wrote data, now reading....")
time.sleep(0.25)
read_root=zarr.open(store=SQLiteStore(fname, read_only=True), mode='r')
ytest=read_root['group1/ydat'][()]
print(f"## DONE {read_root=}{ytest=}")

Something like this to create, write, and read should probably be included in the tests.

I find the use of async sort of odd here. I effectively never use async.

I would even say that the

read_root=zarr.open(store=SQLiteStore(fname, read_only=True), mode='r')

is sort of a wart here. The more obvious (to me anyway)

read_root=zarr.open(store=fname, mode='r')

fails.

@auxym

Copy link
Copy Markdown
OwnerAuthor

I agree on most things, will look into including your proposed test and the import fix. I also agree on async, I rarely use it, but I believe it is required to satisfy the zarr-python store abstract base class: https://github.com/zarr-developers/zarr-python/blob/main/src/zarr/abc/store.py

read_root = zarr.open(store=fname, mode='r')

I don't understand how this should work, how would zarr know it should use the SQLiteStore class to create the store based on only a file name?

@auxym

auxym commented Jul 17, 2026

Copy link
Copy Markdown
OwnerAuthor

OK, I have implemented your suggested test, see latest commit. It passes (even without the sleep).

For the imports, looking at it, I actually don't think it's necessary to remove the type_checking check, and I assume it could save time during module imports to leave it in place. You can check yourself: you can completely delete the if TYPE_CHECKING block and run the tests (uv run pytest) and every test still passes. Can you confirm that works for you also?

@newville

Copy link
Copy Markdown

TYPE_CHECKING is False at runtime (https://docs.python.org/3/library/typing.html#typing.TYPE_CHECKING).

So the code is not importing collections.abc.Iterable at runtime. But that is needed for more than type checking.
SQLiteStore.list() and SQLiteStore.list_prefix() cannot work without it.

I'll admit that I have little experience with (or love for) async, that I don't know why (from def list())

@overrideasyncdeflist(self) ->AsyncIterator[str]:
cur=awaitself._execute("SELECT k FROM zarr")
forrowincast(Iterable[tuple[str]], cur):
yieldrow[0]

is not (I do not mean "just"):

deflist(self):
return [row[0] forrowinself._execute("SELECT k FROM zarr")]

I can believe there is a reason for the extra code, and that cast(Iterable[tuple[str]], cur) is doing something useful, but I really do not know what that is. (And, like, the name is literally"list".... "returns a list" seems so obvious to me that I must be profoundly missing something and too old for all this fancy.... stuff???).

But, if you're going to use cast(Iterable[tuple[str]], cur), then, you have to import Iterable.

I would just remove the check for TYPE_CHECKING. You can time the imports -- you will find they are insignificant.

read_root = zarr.open(store=fname, mode='r')

I don't understand how this should work, how would zarr know it should use the SQLiteStore class to create the store based on only a file name?

Yeah, I guess this is a complaint about Zarr, not SQLiteStore.

root=zarr.open('store.zarr', mode='r')

works for a FileStore, but that wouldn't work for a ZipStore either. It seems like a registry of file types would be needed.

@auxym

auxym commented Jul 19, 2026

Copy link
Copy Markdown
OwnerAuthor

Oh yeah, missed those cast calls. The existing minimal tests did not test that code path. I added a whole bunch of tests (with LLM assistance) and I completely removed the if TYPE_CHECKING as you suggested.

Regarding the async stuff, my understanding is that zarr-python uses async "under the hood", but wraps a sync API around it. Thus the user of zarr can use either sync or async calls.

@newville

Copy link
Copy Markdown

@auxym Thanks.

It seems to me that the point of cast is to tell the type checker not to check types. ;) What happens if you replace for row in cast(Iterable[tuple[str]], cur): with for row in cur:? Does some IDE complain?

Regarding the async stuff, my understanding is that zarr-python uses async "under the hood", but wraps a sync API around it. Thus the user of zarr can use either sync or async calls.

I can believe all of that.

I admit a bias against async (I support a lot of Python projects for data collection and analysis, and I just never use async, even when asynchronous callbacks from a distributed control system are a vital part of the software. Threads for specific tasks, multiprocessing, yes definitely).

Async adds a bunch of keywords all over Python to bolt on coroutines that are really about asynchronous I/O (and not anything else), and then requires adding a bunch more keywords ("await"s) all over Python code to say "don't be asynchronous. You can ignore this as a rant, but I think hat

deflist(self):
return [row[0] forrowinself._execute("select k from zarr")]

has merit that

@overrideasyncdeflist(self) ->AsyncIterator[str]:
cur=awaitself._execute("SELECT k FROM zarr")
forrowincast(Iterable[tuple[str]], cur):
yieldrow[0]

does not necessarily improve. There is a bunch of code to signal non-runtime type-checkers to ignore type checking. There is code to "await" the (allegedly) expensive database lookup, and then yields over a list of values that has already been synchronously fetched. [If "cast key to str" is needed, then I think AsyncIterator[str] is not guaranteed. Perhaps the yielded value should be str(row[0])?].

When I look in detail here, it seems that all of the calls to self._execute() and self._execute_write() are awaited. Doesn't that mean these functions could be synchronous, and sort of really are synchronous?
I'll also guess that for SQLite with WAL, these database queries (simple select in a two-column table) will be fast.
I'm comfortable saying that any claim about speed needs real benchmarks, but I'll say that goes for the claim that "using async makes stuff faster".

Similarly, where the code has await self._ensure_open(), couldn't that be if not self.is_open: open()?

Feel free to ignore all this.

@auxym

Copy link
Copy Markdown
OwnerAuthor

Yes, I believe the casts were necessary to stop mypy/pyright from complaining.

That said,

Perhaps the yielded value should be str(row[0])?].

I do like that suggestion. Latest commit got rid of all casts :)

When I look in detail here, it seems that all of the calls to self._execute() and self._execute_write()

Yeah, I wrapped all db transactions in those methods. Yes, they are currently blocking (non-async), since the stdlib sqlite3 is blocking. The goal of putting all transactions in a central place is that I aim, in the future, to implement real async transactions, either using aiosqlite3 (3rd party lib) or implementing our own background thread for db work (probably the former). I also personally do not use async, but zarr itself offers async everywhere, so it would be a better match for that pattern.

Similarly, where the code has await self._ensure_open(), couldn't that be if not self.is_open: open()?

_ensure_open is from zarr's store base class, once again I'm trying to follow zarr-python's existing patterns.

https://github.com/zarr-developers/zarr-python/blob/50b7e016590d76011382771fd52843949782c9c0/src/zarr/abc/store.py#L140

@auxym
auxym merged commit ef46459 into mainJul 24, 2026
@auxym
auxym deleted the ensure-file-created branch July 24, 2026 01:03
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

trouble using with synchronous access

2 participants

@auxym@newville
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Ensure file is created when it doesnt exist by auxym · Pull Request #5 · auxym/zarr-sqlite-python · GitHub
Skip to content

Ensure file is created when it doesnt exist - #5

Merged
auxym merged 13 commits into
mainfrom
ensure-file-created
Jul 24, 2026
Merged

Ensure file is created when it doesnt exist#5
auxym merged 13 commits into
mainfrom
ensure-file-created

Conversation

@auxym

Copy link
Copy Markdown
Owner

Resolves#3

  • Change mode=rw to mode=rwc
  • Add unit test

- Change mode=rw to mode=rwc
- Add unit test
@auxym

Copy link
Copy Markdown
OwnerAuthor

@newville can you please confirm this fixes the issue for you also?

@newville

Copy link
Copy Markdown

@auxym Thanks -- This also needs the if TYPE_CHECKiNG test removed, and those imports always done.

I think a basic test, with no async code, such as

importzarrfromzarr_sqliteimportSQLiteStorefname='test1.zarrdb'store=SQLiteStore(fname, read_only=False)
root=zarr.open(store=store, mode='a')
group1=root.create_group('group1')
i=np.arange(80000)/60.0x=i+np.random.normal(size=len(i), scale=1.5)
y=np.sin(x/13) +0.7*np.cos(x/47) +np.random.normal(size=len(i), scale=0.05)
x.shape= (400, 200)
y.shape= (200, 400)
group1.create_array(data=x, name='xdat')
group1.create_array(data=y, name='ydat')
print(group1, list(group1.keys()))
##print("wrote data, now reading....")
time.sleep(0.25)
read_root=zarr.open(store=SQLiteStore(fname, read_only=True), mode='r')
ytest=read_root['group1/ydat'][()]
print(f"## DONE {read_root=}{ytest=}")

Something like this to create, write, and read should probably be included in the tests.

I find the use of async sort of odd here. I effectively never use async.

I would even say that the

read_root=zarr.open(store=SQLiteStore(fname, read_only=True), mode='r')

is sort of a wart here. The more obvious (to me anyway)

read_root=zarr.open(store=fname, mode='r')

fails.

@auxym

Copy link
Copy Markdown
OwnerAuthor

I agree on most things, will look into including your proposed test and the import fix. I also agree on async, I rarely use it, but I believe it is required to satisfy the zarr-python store abstract base class: https://github.com/zarr-developers/zarr-python/blob/main/src/zarr/abc/store.py

read_root = zarr.open(store=fname, mode='r')

I don't understand how this should work, how would zarr know it should use the SQLiteStore class to create the store based on only a file name?

@auxym

auxym commented Jul 17, 2026

Copy link
Copy Markdown
OwnerAuthor

OK, I have implemented your suggested test, see latest commit. It passes (even without the sleep).

For the imports, looking at it, I actually don't think it's necessary to remove the type_checking check, and I assume it could save time during module imports to leave it in place. You can check yourself: you can completely delete the if TYPE_CHECKING block and run the tests (uv run pytest) and every test still passes. Can you confirm that works for you also?

@newville

Copy link
Copy Markdown

TYPE_CHECKING is False at runtime (https://docs.python.org/3/library/typing.html#typing.TYPE_CHECKING).

So the code is not importing collections.abc.Iterable at runtime. But that is needed for more than type checking.
SQLiteStore.list() and SQLiteStore.list_prefix() cannot work without it.

I'll admit that I have little experience with (or love for) async, that I don't know why (from def list())

@overrideasyncdeflist(self) ->AsyncIterator[str]:
cur=awaitself._execute("SELECT k FROM zarr")
forrowincast(Iterable[tuple[str]], cur):
yieldrow[0]

is not (I do not mean "just"):

deflist(self):
return [row[0] forrowinself._execute("SELECT k FROM zarr")]

I can believe there is a reason for the extra code, and that cast(Iterable[tuple[str]], cur) is doing something useful, but I really do not know what that is. (And, like, the name is literally"list".... "returns a list" seems so obvious to me that I must be profoundly missing something and too old for all this fancy.... stuff???).

But, if you're going to use cast(Iterable[tuple[str]], cur), then, you have to import Iterable.

I would just remove the check for TYPE_CHECKING. You can time the imports -- you will find they are insignificant.

read_root = zarr.open(store=fname, mode='r')

I don't understand how this should work, how would zarr know it should use the SQLiteStore class to create the store based on only a file name?

Yeah, I guess this is a complaint about Zarr, not SQLiteStore.

root=zarr.open('store.zarr', mode='r')

works for a FileStore, but that wouldn't work for a ZipStore either. It seems like a registry of file types would be needed.

@auxym

auxym commented Jul 19, 2026

Copy link
Copy Markdown
OwnerAuthor

Oh yeah, missed those cast calls. The existing minimal tests did not test that code path. I added a whole bunch of tests (with LLM assistance) and I completely removed the if TYPE_CHECKING as you suggested.

Regarding the async stuff, my understanding is that zarr-python uses async "under the hood", but wraps a sync API around it. Thus the user of zarr can use either sync or async calls.

@newville

Copy link
Copy Markdown

@auxym Thanks.

It seems to me that the point of cast is to tell the type checker not to check types. ;) What happens if you replace for row in cast(Iterable[tuple[str]], cur): with for row in cur:? Does some IDE complain?

Regarding the async stuff, my understanding is that zarr-python uses async "under the hood", but wraps a sync API around it. Thus the user of zarr can use either sync or async calls.

I can believe all of that.

I admit a bias against async (I support a lot of Python projects for data collection and analysis, and I just never use async, even when asynchronous callbacks from a distributed control system are a vital part of the software. Threads for specific tasks, multiprocessing, yes definitely).

Async adds a bunch of keywords all over Python to bolt on coroutines that are really about asynchronous I/O (and not anything else), and then requires adding a bunch more keywords ("await"s) all over Python code to say "don't be asynchronous. You can ignore this as a rant, but I think hat

deflist(self):
return [row[0] forrowinself._execute("select k from zarr")]

has merit that

@overrideasyncdeflist(self) ->AsyncIterator[str]:
cur=awaitself._execute("SELECT k FROM zarr")
forrowincast(Iterable[tuple[str]], cur):
yieldrow[0]

does not necessarily improve. There is a bunch of code to signal non-runtime type-checkers to ignore type checking. There is code to "await" the (allegedly) expensive database lookup, and then yields over a list of values that has already been synchronously fetched. [If "cast key to str" is needed, then I think AsyncIterator[str] is not guaranteed. Perhaps the yielded value should be str(row[0])?].

When I look in detail here, it seems that all of the calls to self._execute() and self._execute_write() are awaited. Doesn't that mean these functions could be synchronous, and sort of really are synchronous?
I'll also guess that for SQLite with WAL, these database queries (simple select in a two-column table) will be fast.
I'm comfortable saying that any claim about speed needs real benchmarks, but I'll say that goes for the claim that "using async makes stuff faster".

Similarly, where the code has await self._ensure_open(), couldn't that be if not self.is_open: open()?

Feel free to ignore all this.

@auxym

Copy link
Copy Markdown
OwnerAuthor

Yes, I believe the casts were necessary to stop mypy/pyright from complaining.

That said,

Perhaps the yielded value should be str(row[0])?].

I do like that suggestion. Latest commit got rid of all casts :)

When I look in detail here, it seems that all of the calls to self._execute() and self._execute_write()

Yeah, I wrapped all db transactions in those methods. Yes, they are currently blocking (non-async), since the stdlib sqlite3 is blocking. The goal of putting all transactions in a central place is that I aim, in the future, to implement real async transactions, either using aiosqlite3 (3rd party lib) or implementing our own background thread for db work (probably the former). I also personally do not use async, but zarr itself offers async everywhere, so it would be a better match for that pattern.

Similarly, where the code has await self._ensure_open(), couldn't that be if not self.is_open: open()?

_ensure_open is from zarr's store base class, once again I'm trying to follow zarr-python's existing patterns.

https://github.com/zarr-developers/zarr-python/blob/50b7e016590d76011382771fd52843949782c9c0/src/zarr/abc/store.py#L140

@auxym
auxym merged commit ef46459 into mainJul 24, 2026
@auxym
auxym deleted the ensure-file-created branch July 24, 2026 01:03
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

trouble using with synchronous access

2 participants

@auxym@newville
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Ensure file is created when it doesnt exist by auxym · Pull Request #5 · auxym/zarr-sqlite-python · GitHub
Skip to content

Ensure file is created when it doesnt exist - #5

Merged
auxym merged 13 commits into
mainfrom
ensure-file-created
Jul 24, 2026
Merged

Ensure file is created when it doesnt exist#5
auxym merged 13 commits into
mainfrom
ensure-file-created

Conversation

@auxym

Copy link
Copy Markdown
Owner

Resolves#3

  • Change mode=rw to mode=rwc
  • Add unit test

- Change mode=rw to mode=rwc
- Add unit test
@auxym

Copy link
Copy Markdown
OwnerAuthor

@newville can you please confirm this fixes the issue for you also?

@newville

Copy link
Copy Markdown

@auxym Thanks -- This also needs the if TYPE_CHECKiNG test removed, and those imports always done.

I think a basic test, with no async code, such as

importzarrfromzarr_sqliteimportSQLiteStorefname='test1.zarrdb'store=SQLiteStore(fname, read_only=False)
root=zarr.open(store=store, mode='a')
group1=root.create_group('group1')
i=np.arange(80000)/60.0x=i+np.random.normal(size=len(i), scale=1.5)
y=np.sin(x/13) +0.7*np.cos(x/47) +np.random.normal(size=len(i), scale=0.05)
x.shape= (400, 200)
y.shape= (200, 400)
group1.create_array(data=x, name='xdat')
group1.create_array(data=y, name='ydat')
print(group1, list(group1.keys()))
##print("wrote data, now reading....")
time.sleep(0.25)
read_root=zarr.open(store=SQLiteStore(fname, read_only=True), mode='r')
ytest=read_root['group1/ydat'][()]
print(f"## DONE {read_root=}{ytest=}")

Something like this to create, write, and read should probably be included in the tests.

I find the use of async sort of odd here. I effectively never use async.

I would even say that the

read_root=zarr.open(store=SQLiteStore(fname, read_only=True), mode='r')

is sort of a wart here. The more obvious (to me anyway)

read_root=zarr.open(store=fname, mode='r')

fails.

@auxym

Copy link
Copy Markdown
OwnerAuthor

I agree on most things, will look into including your proposed test and the import fix. I also agree on async, I rarely use it, but I believe it is required to satisfy the zarr-python store abstract base class: https://github.com/zarr-developers/zarr-python/blob/main/src/zarr/abc/store.py

read_root = zarr.open(store=fname, mode='r')

I don't understand how this should work, how would zarr know it should use the SQLiteStore class to create the store based on only a file name?

@auxym

auxym commented Jul 17, 2026

Copy link
Copy Markdown
OwnerAuthor

OK, I have implemented your suggested test, see latest commit. It passes (even without the sleep).

For the imports, looking at it, I actually don't think it's necessary to remove the type_checking check, and I assume it could save time during module imports to leave it in place. You can check yourself: you can completely delete the if TYPE_CHECKING block and run the tests (uv run pytest) and every test still passes. Can you confirm that works for you also?

@newville

Copy link
Copy Markdown

TYPE_CHECKING is False at runtime (https://docs.python.org/3/library/typing.html#typing.TYPE_CHECKING).

So the code is not importing collections.abc.Iterable at runtime. But that is needed for more than type checking.
SQLiteStore.list() and SQLiteStore.list_prefix() cannot work without it.

I'll admit that I have little experience with (or love for) async, that I don't know why (from def list())

@overrideasyncdeflist(self) ->AsyncIterator[str]:
cur=awaitself._execute("SELECT k FROM zarr")
forrowincast(Iterable[tuple[str]], cur):
yieldrow[0]

is not (I do not mean "just"):

deflist(self):
return [row[0] forrowinself._execute("SELECT k FROM zarr")]

I can believe there is a reason for the extra code, and that cast(Iterable[tuple[str]], cur) is doing something useful, but I really do not know what that is. (And, like, the name is literally"list".... "returns a list" seems so obvious to me that I must be profoundly missing something and too old for all this fancy.... stuff???).

But, if you're going to use cast(Iterable[tuple[str]], cur), then, you have to import Iterable.

I would just remove the check for TYPE_CHECKING. You can time the imports -- you will find they are insignificant.

read_root = zarr.open(store=fname, mode='r')

I don't understand how this should work, how would zarr know it should use the SQLiteStore class to create the store based on only a file name?

Yeah, I guess this is a complaint about Zarr, not SQLiteStore.

root=zarr.open('store.zarr', mode='r')

works for a FileStore, but that wouldn't work for a ZipStore either. It seems like a registry of file types would be needed.

@auxym

auxym commented Jul 19, 2026

Copy link
Copy Markdown
OwnerAuthor

Oh yeah, missed those cast calls. The existing minimal tests did not test that code path. I added a whole bunch of tests (with LLM assistance) and I completely removed the if TYPE_CHECKING as you suggested.

Regarding the async stuff, my understanding is that zarr-python uses async "under the hood", but wraps a sync API around it. Thus the user of zarr can use either sync or async calls.

@newville

Copy link
Copy Markdown

@auxym Thanks.

It seems to me that the point of cast is to tell the type checker not to check types. ;) What happens if you replace for row in cast(Iterable[tuple[str]], cur): with for row in cur:? Does some IDE complain?

Regarding the async stuff, my understanding is that zarr-python uses async "under the hood", but wraps a sync API around it. Thus the user of zarr can use either sync or async calls.

I can believe all of that.

I admit a bias against async (I support a lot of Python projects for data collection and analysis, and I just never use async, even when asynchronous callbacks from a distributed control system are a vital part of the software. Threads for specific tasks, multiprocessing, yes definitely).

Async adds a bunch of keywords all over Python to bolt on coroutines that are really about asynchronous I/O (and not anything else), and then requires adding a bunch more keywords ("await"s) all over Python code to say "don't be asynchronous. You can ignore this as a rant, but I think hat

deflist(self):
return [row[0] forrowinself._execute("select k from zarr")]

has merit that

@overrideasyncdeflist(self) ->AsyncIterator[str]:
cur=awaitself._execute("SELECT k FROM zarr")
forrowincast(Iterable[tuple[str]], cur):
yieldrow[0]

does not necessarily improve. There is a bunch of code to signal non-runtime type-checkers to ignore type checking. There is code to "await" the (allegedly) expensive database lookup, and then yields over a list of values that has already been synchronously fetched. [If "cast key to str" is needed, then I think AsyncIterator[str] is not guaranteed. Perhaps the yielded value should be str(row[0])?].

When I look in detail here, it seems that all of the calls to self._execute() and self._execute_write() are awaited. Doesn't that mean these functions could be synchronous, and sort of really are synchronous?
I'll also guess that for SQLite with WAL, these database queries (simple select in a two-column table) will be fast.
I'm comfortable saying that any claim about speed needs real benchmarks, but I'll say that goes for the claim that "using async makes stuff faster".

Similarly, where the code has await self._ensure_open(), couldn't that be if not self.is_open: open()?

Feel free to ignore all this.

@auxym

Copy link
Copy Markdown
OwnerAuthor

Yes, I believe the casts were necessary to stop mypy/pyright from complaining.

That said,

Perhaps the yielded value should be str(row[0])?].

I do like that suggestion. Latest commit got rid of all casts :)

When I look in detail here, it seems that all of the calls to self._execute() and self._execute_write()

Yeah, I wrapped all db transactions in those methods. Yes, they are currently blocking (non-async), since the stdlib sqlite3 is blocking. The goal of putting all transactions in a central place is that I aim, in the future, to implement real async transactions, either using aiosqlite3 (3rd party lib) or implementing our own background thread for db work (probably the former). I also personally do not use async, but zarr itself offers async everywhere, so it would be a better match for that pattern.

Similarly, where the code has await self._ensure_open(), couldn't that be if not self.is_open: open()?

_ensure_open is from zarr's store base class, once again I'm trying to follow zarr-python's existing patterns.

https://github.com/zarr-developers/zarr-python/blob/50b7e016590d76011382771fd52843949782c9c0/src/zarr/abc/store.py#L140

@auxym
auxym merged commit ef46459 into mainJul 24, 2026
@auxym
auxym deleted the ensure-file-created branch July 24, 2026 01:03
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

trouble using with synchronous access

2 participants

@auxym@newville
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' Ensure file is created when it doesnt exist by auxym · Pull Request #5 · auxym/zarr-sqlite-python · GitHub
Skip to content

Ensure file is created when it doesnt exist - #5

Merged
auxym merged 13 commits into
mainfrom
ensure-file-created
Jul 24, 2026
Merged

Ensure file is created when it doesnt exist#5
auxym merged 13 commits into
mainfrom
ensure-file-created

Conversation

@auxym

Copy link
Copy Markdown
Owner

Resolves#3

  • Change mode=rw to mode=rwc
  • Add unit test

- Change mode=rw to mode=rwc
- Add unit test
@auxym

Copy link
Copy Markdown
OwnerAuthor

@newville can you please confirm this fixes the issue for you also?

@newville

Copy link
Copy Markdown

@auxym Thanks -- This also needs the if TYPE_CHECKiNG test removed, and those imports always done.

I think a basic test, with no async code, such as

importzarrfromzarr_sqliteimportSQLiteStorefname='test1.zarrdb'store=SQLiteStore(fname, read_only=False)
root=zarr.open(store=store, mode='a')
group1=root.create_group('group1')
i=np.arange(80000)/60.0x=i+np.random.normal(size=len(i), scale=1.5)
y=np.sin(x/13) +0.7*np.cos(x/47) +np.random.normal(size=len(i), scale=0.05)
x.shape= (400, 200)
y.shape= (200, 400)
group1.create_array(data=x, name='xdat')
group1.create_array(data=y, name='ydat')
print(group1, list(group1.keys()))
##print("wrote data, now reading....")
time.sleep(0.25)
read_root=zarr.open(store=SQLiteStore(fname, read_only=True), mode='r')
ytest=read_root['group1/ydat'][()]
print(f"## DONE {read_root=}{ytest=}")

Something like this to create, write, and read should probably be included in the tests.

I find the use of async sort of odd here. I effectively never use async.

I would even say that the

read_root=zarr.open(store=SQLiteStore(fname, read_only=True), mode='r')

is sort of a wart here. The more obvious (to me anyway)

read_root=zarr.open(store=fname, mode='r')

fails.

@auxym

Copy link
Copy Markdown
OwnerAuthor

I agree on most things, will look into including your proposed test and the import fix. I also agree on async, I rarely use it, but I believe it is required to satisfy the zarr-python store abstract base class: https://github.com/zarr-developers/zarr-python/blob/main/src/zarr/abc/store.py

read_root = zarr.open(store=fname, mode='r')

I don't understand how this should work, how would zarr know it should use the SQLiteStore class to create the store based on only a file name?

@auxym

auxym commented Jul 17, 2026

Copy link
Copy Markdown
OwnerAuthor

OK, I have implemented your suggested test, see latest commit. It passes (even without the sleep).

For the imports, looking at it, I actually don't think it's necessary to remove the type_checking check, and I assume it could save time during module imports to leave it in place. You can check yourself: you can completely delete the if TYPE_CHECKING block and run the tests (uv run pytest) and every test still passes. Can you confirm that works for you also?

@newville

Copy link
Copy Markdown

TYPE_CHECKING is False at runtime (https://docs.python.org/3/library/typing.html#typing.TYPE_CHECKING).

So the code is not importing collections.abc.Iterable at runtime. But that is needed for more than type checking.
SQLiteStore.list() and SQLiteStore.list_prefix() cannot work without it.

I'll admit that I have little experience with (or love for) async, that I don't know why (from def list())

@overrideasyncdeflist(self) ->AsyncIterator[str]:
cur=awaitself._execute("SELECT k FROM zarr")
forrowincast(Iterable[tuple[str]], cur):
yieldrow[0]

is not (I do not mean "just"):

deflist(self):
return [row[0] forrowinself._execute("SELECT k FROM zarr")]

I can believe there is a reason for the extra code, and that cast(Iterable[tuple[str]], cur) is doing something useful, but I really do not know what that is. (And, like, the name is literally"list".... "returns a list" seems so obvious to me that I must be profoundly missing something and too old for all this fancy.... stuff???).

But, if you're going to use cast(Iterable[tuple[str]], cur), then, you have to import Iterable.

I would just remove the check for TYPE_CHECKING. You can time the imports -- you will find they are insignificant.

read_root = zarr.open(store=fname, mode='r')

I don't understand how this should work, how would zarr know it should use the SQLiteStore class to create the store based on only a file name?

Yeah, I guess this is a complaint about Zarr, not SQLiteStore.

root=zarr.open('store.zarr', mode='r')

works for a FileStore, but that wouldn't work for a ZipStore either. It seems like a registry of file types would be needed.

@auxym

auxym commented Jul 19, 2026

Copy link
Copy Markdown
OwnerAuthor

Oh yeah, missed those cast calls. The existing minimal tests did not test that code path. I added a whole bunch of tests (with LLM assistance) and I completely removed the if TYPE_CHECKING as you suggested.

Regarding the async stuff, my understanding is that zarr-python uses async "under the hood", but wraps a sync API around it. Thus the user of zarr can use either sync or async calls.

@newville

Copy link
Copy Markdown

@auxym Thanks.

It seems to me that the point of cast is to tell the type checker not to check types. ;) What happens if you replace for row in cast(Iterable[tuple[str]], cur): with for row in cur:? Does some IDE complain?

Regarding the async stuff, my understanding is that zarr-python uses async "under the hood", but wraps a sync API around it. Thus the user of zarr can use either sync or async calls.

I can believe all of that.

I admit a bias against async (I support a lot of Python projects for data collection and analysis, and I just never use async, even when asynchronous callbacks from a distributed control system are a vital part of the software. Threads for specific tasks, multiprocessing, yes definitely).

Async adds a bunch of keywords all over Python to bolt on coroutines that are really about asynchronous I/O (and not anything else), and then requires adding a bunch more keywords ("await"s) all over Python code to say "don't be asynchronous. You can ignore this as a rant, but I think hat

deflist(self):
return [row[0] forrowinself._execute("select k from zarr")]

has merit that

@overrideasyncdeflist(self) ->AsyncIterator[str]:
cur=awaitself._execute("SELECT k FROM zarr")
forrowincast(Iterable[tuple[str]], cur):
yieldrow[0]

does not necessarily improve. There is a bunch of code to signal non-runtime type-checkers to ignore type checking. There is code to "await" the (allegedly) expensive database lookup, and then yields over a list of values that has already been synchronously fetched. [If "cast key to str" is needed, then I think AsyncIterator[str] is not guaranteed. Perhaps the yielded value should be str(row[0])?].

When I look in detail here, it seems that all of the calls to self._execute() and self._execute_write() are awaited. Doesn't that mean these functions could be synchronous, and sort of really are synchronous?
I'll also guess that for SQLite with WAL, these database queries (simple select in a two-column table) will be fast.
I'm comfortable saying that any claim about speed needs real benchmarks, but I'll say that goes for the claim that "using async makes stuff faster".

Similarly, where the code has await self._ensure_open(), couldn't that be if not self.is_open: open()?

Feel free to ignore all this.

@auxym

Copy link
Copy Markdown
OwnerAuthor

Yes, I believe the casts were necessary to stop mypy/pyright from complaining.

That said,

Perhaps the yielded value should be str(row[0])?].

I do like that suggestion. Latest commit got rid of all casts :)

When I look in detail here, it seems that all of the calls to self._execute() and self._execute_write()

Yeah, I wrapped all db transactions in those methods. Yes, they are currently blocking (non-async), since the stdlib sqlite3 is blocking. The goal of putting all transactions in a central place is that I aim, in the future, to implement real async transactions, either using aiosqlite3 (3rd party lib) or implementing our own background thread for db work (probably the former). I also personally do not use async, but zarr itself offers async everywhere, so it would be a better match for that pattern.

Similarly, where the code has await self._ensure_open(), couldn't that be if not self.is_open: open()?

_ensure_open is from zarr's store base class, once again I'm trying to follow zarr-python's existing patterns.

https://github.com/zarr-developers/zarr-python/blob/50b7e016590d76011382771fd52843949782c9c0/src/zarr/abc/store.py#L140

@auxym
auxym merged commit ef46459 into mainJul 24, 2026
@auxym
auxym deleted the ensure-file-created branch July 24, 2026 01:03
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

trouble using with synchronous access

2 participants

@auxym@newville
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Ensure file is created when it doesnt exist by auxym · Pull Request #5 · auxym/zarr-sqlite-python · GitHub
Skip to content

Ensure file is created when it doesnt exist - #5

Merged
auxym merged 13 commits into
mainfrom
ensure-file-created
Jul 24, 2026
Merged

Ensure file is created when it doesnt exist#5
auxym merged 13 commits into
mainfrom
ensure-file-created

Conversation

@auxym

Copy link
Copy Markdown
Owner

Resolves#3

  • Change mode=rw to mode=rwc
  • Add unit test

- Change mode=rw to mode=rwc
- Add unit test
@auxym

Copy link
Copy Markdown
OwnerAuthor

@newville can you please confirm this fixes the issue for you also?

@newville

Copy link
Copy Markdown

@auxym Thanks -- This also needs the if TYPE_CHECKiNG test removed, and those imports always done.

I think a basic test, with no async code, such as

importzarrfromzarr_sqliteimportSQLiteStorefname='test1.zarrdb'store=SQLiteStore(fname, read_only=False)
root=zarr.open(store=store, mode='a')
group1=root.create_group('group1')
i=np.arange(80000)/60.0x=i+np.random.normal(size=len(i), scale=1.5)
y=np.sin(x/13) +0.7*np.cos(x/47) +np.random.normal(size=len(i), scale=0.05)
x.shape= (400, 200)
y.shape= (200, 400)
group1.create_array(data=x, name='xdat')
group1.create_array(data=y, name='ydat')
print(group1, list(group1.keys()))
##print("wrote data, now reading....")
time.sleep(0.25)
read_root=zarr.open(store=SQLiteStore(fname, read_only=True), mode='r')
ytest=read_root['group1/ydat'][()]
print(f"## DONE {read_root=}{ytest=}")

Something like this to create, write, and read should probably be included in the tests.

I find the use of async sort of odd here. I effectively never use async.

I would even say that the

read_root=zarr.open(store=SQLiteStore(fname, read_only=True), mode='r')

is sort of a wart here. The more obvious (to me anyway)

read_root=zarr.open(store=fname, mode='r')

fails.

@auxym

Copy link
Copy Markdown
OwnerAuthor

I agree on most things, will look into including your proposed test and the import fix. I also agree on async, I rarely use it, but I believe it is required to satisfy the zarr-python store abstract base class: https://github.com/zarr-developers/zarr-python/blob/main/src/zarr/abc/store.py

read_root = zarr.open(store=fname, mode='r')

I don't understand how this should work, how would zarr know it should use the SQLiteStore class to create the store based on only a file name?

@auxym

auxym commented Jul 17, 2026

Copy link
Copy Markdown
OwnerAuthor

OK, I have implemented your suggested test, see latest commit. It passes (even without the sleep).

For the imports, looking at it, I actually don't think it's necessary to remove the type_checking check, and I assume it could save time during module imports to leave it in place. You can check yourself: you can completely delete the if TYPE_CHECKING block and run the tests (uv run pytest) and every test still passes. Can you confirm that works for you also?

@newville

Copy link
Copy Markdown

TYPE_CHECKING is False at runtime (https://docs.python.org/3/library/typing.html#typing.TYPE_CHECKING).

So the code is not importing collections.abc.Iterable at runtime. But that is needed for more than type checking.
SQLiteStore.list() and SQLiteStore.list_prefix() cannot work without it.

I'll admit that I have little experience with (or love for) async, that I don't know why (from def list())

@overrideasyncdeflist(self) ->AsyncIterator[str]:
cur=awaitself._execute("SELECT k FROM zarr")
forrowincast(Iterable[tuple[str]], cur):
yieldrow[0]

is not (I do not mean "just"):

deflist(self):
return [row[0] forrowinself._execute("SELECT k FROM zarr")]

I can believe there is a reason for the extra code, and that cast(Iterable[tuple[str]], cur) is doing something useful, but I really do not know what that is. (And, like, the name is literally"list".... "returns a list" seems so obvious to me that I must be profoundly missing something and too old for all this fancy.... stuff???).

But, if you're going to use cast(Iterable[tuple[str]], cur), then, you have to import Iterable.

I would just remove the check for TYPE_CHECKING. You can time the imports -- you will find they are insignificant.

read_root = zarr.open(store=fname, mode='r')

I don't understand how this should work, how would zarr know it should use the SQLiteStore class to create the store based on only a file name?

Yeah, I guess this is a complaint about Zarr, not SQLiteStore.

root=zarr.open('store.zarr', mode='r')

works for a FileStore, but that wouldn't work for a ZipStore either. It seems like a registry of file types would be needed.

@auxym

auxym commented Jul 19, 2026

Copy link
Copy Markdown
OwnerAuthor

Oh yeah, missed those cast calls. The existing minimal tests did not test that code path. I added a whole bunch of tests (with LLM assistance) and I completely removed the if TYPE_CHECKING as you suggested.

Regarding the async stuff, my understanding is that zarr-python uses async "under the hood", but wraps a sync API around it. Thus the user of zarr can use either sync or async calls.

@newville

Copy link
Copy Markdown

@auxym Thanks.

It seems to me that the point of cast is to tell the type checker not to check types. ;) What happens if you replace for row in cast(Iterable[tuple[str]], cur): with for row in cur:? Does some IDE complain?

Regarding the async stuff, my understanding is that zarr-python uses async "under the hood", but wraps a sync API around it. Thus the user of zarr can use either sync or async calls.

I can believe all of that.

I admit a bias against async (I support a lot of Python projects for data collection and analysis, and I just never use async, even when asynchronous callbacks from a distributed control system are a vital part of the software. Threads for specific tasks, multiprocessing, yes definitely).

Async adds a bunch of keywords all over Python to bolt on coroutines that are really about asynchronous I/O (and not anything else), and then requires adding a bunch more keywords ("await"s) all over Python code to say "don't be asynchronous. You can ignore this as a rant, but I think hat

deflist(self):
return [row[0] forrowinself._execute("select k from zarr")]

has merit that

@overrideasyncdeflist(self) ->AsyncIterator[str]:
cur=awaitself._execute("SELECT k FROM zarr")
forrowincast(Iterable[tuple[str]], cur):
yieldrow[0]

does not necessarily improve. There is a bunch of code to signal non-runtime type-checkers to ignore type checking. There is code to "await" the (allegedly) expensive database lookup, and then yields over a list of values that has already been synchronously fetched. [If "cast key to str" is needed, then I think AsyncIterator[str] is not guaranteed. Perhaps the yielded value should be str(row[0])?].

When I look in detail here, it seems that all of the calls to self._execute() and self._execute_write() are awaited. Doesn't that mean these functions could be synchronous, and sort of really are synchronous?
I'll also guess that for SQLite with WAL, these database queries (simple select in a two-column table) will be fast.
I'm comfortable saying that any claim about speed needs real benchmarks, but I'll say that goes for the claim that "using async makes stuff faster".

Similarly, where the code has await self._ensure_open(), couldn't that be if not self.is_open: open()?

Feel free to ignore all this.

@auxym

Copy link
Copy Markdown
OwnerAuthor

Yes, I believe the casts were necessary to stop mypy/pyright from complaining.

That said,

Perhaps the yielded value should be str(row[0])?].

I do like that suggestion. Latest commit got rid of all casts :)

When I look in detail here, it seems that all of the calls to self._execute() and self._execute_write()

Yeah, I wrapped all db transactions in those methods. Yes, they are currently blocking (non-async), since the stdlib sqlite3 is blocking. The goal of putting all transactions in a central place is that I aim, in the future, to implement real async transactions, either using aiosqlite3 (3rd party lib) or implementing our own background thread for db work (probably the former). I also personally do not use async, but zarr itself offers async everywhere, so it would be a better match for that pattern.

Similarly, where the code has await self._ensure_open(), couldn't that be if not self.is_open: open()?

_ensure_open is from zarr's store base class, once again I'm trying to follow zarr-python's existing patterns.

https://github.com/zarr-developers/zarr-python/blob/50b7e016590d76011382771fd52843949782c9c0/src/zarr/abc/store.py#L140

@auxym
auxym merged commit ef46459 into mainJul 24, 2026
@auxym
auxym deleted the ensure-file-created branch July 24, 2026 01:03
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

trouble using with synchronous access

2 participants

@auxym@newville
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Ensure file is created when it doesnt exist by auxym · Pull Request #5 · auxym/zarr-sqlite-python · GitHub
Skip to content

Ensure file is created when it doesnt exist - #5

Merged
auxym merged 13 commits into
mainfrom
ensure-file-created
Jul 24, 2026
Merged

Ensure file is created when it doesnt exist#5
auxym merged 13 commits into
mainfrom
ensure-file-created

Conversation

@auxym

Copy link
Copy Markdown
Owner

Resolves#3

  • Change mode=rw to mode=rwc
  • Add unit test

- Change mode=rw to mode=rwc
- Add unit test
@auxym

Copy link
Copy Markdown
OwnerAuthor

@newville can you please confirm this fixes the issue for you also?

@newville

Copy link
Copy Markdown

@auxym Thanks -- This also needs the if TYPE_CHECKiNG test removed, and those imports always done.

I think a basic test, with no async code, such as

importzarrfromzarr_sqliteimportSQLiteStorefname='test1.zarrdb'store=SQLiteStore(fname, read_only=False)
root=zarr.open(store=store, mode='a')
group1=root.create_group('group1')
i=np.arange(80000)/60.0x=i+np.random.normal(size=len(i), scale=1.5)
y=np.sin(x/13) +0.7*np.cos(x/47) +np.random.normal(size=len(i), scale=0.05)
x.shape= (400, 200)
y.shape= (200, 400)
group1.create_array(data=x, name='xdat')
group1.create_array(data=y, name='ydat')
print(group1, list(group1.keys()))
##print("wrote data, now reading....")
time.sleep(0.25)
read_root=zarr.open(store=SQLiteStore(fname, read_only=True), mode='r')
ytest=read_root['group1/ydat'][()]
print(f"## DONE {read_root=}{ytest=}")

Something like this to create, write, and read should probably be included in the tests.

I find the use of async sort of odd here. I effectively never use async.

I would even say that the

read_root=zarr.open(store=SQLiteStore(fname, read_only=True), mode='r')

is sort of a wart here. The more obvious (to me anyway)

read_root=zarr.open(store=fname, mode='r')

fails.

@auxym

Copy link
Copy Markdown
OwnerAuthor

I agree on most things, will look into including your proposed test and the import fix. I also agree on async, I rarely use it, but I believe it is required to satisfy the zarr-python store abstract base class: https://github.com/zarr-developers/zarr-python/blob/main/src/zarr/abc/store.py

read_root = zarr.open(store=fname, mode='r')

I don't understand how this should work, how would zarr know it should use the SQLiteStore class to create the store based on only a file name?

@auxym

auxym commented Jul 17, 2026

Copy link
Copy Markdown
OwnerAuthor

OK, I have implemented your suggested test, see latest commit. It passes (even without the sleep).

For the imports, looking at it, I actually don't think it's necessary to remove the type_checking check, and I assume it could save time during module imports to leave it in place. You can check yourself: you can completely delete the if TYPE_CHECKING block and run the tests (uv run pytest) and every test still passes. Can you confirm that works for you also?

@newville

Copy link
Copy Markdown

TYPE_CHECKING is False at runtime (https://docs.python.org/3/library/typing.html#typing.TYPE_CHECKING).

So the code is not importing collections.abc.Iterable at runtime. But that is needed for more than type checking.
SQLiteStore.list() and SQLiteStore.list_prefix() cannot work without it.

I'll admit that I have little experience with (or love for) async, that I don't know why (from def list())

@overrideasyncdeflist(self) ->AsyncIterator[str]:
cur=awaitself._execute("SELECT k FROM zarr")
forrowincast(Iterable[tuple[str]], cur):
yieldrow[0]

is not (I do not mean "just"):

deflist(self):
return [row[0] forrowinself._execute("SELECT k FROM zarr")]

I can believe there is a reason for the extra code, and that cast(Iterable[tuple[str]], cur) is doing something useful, but I really do not know what that is. (And, like, the name is literally"list".... "returns a list" seems so obvious to me that I must be profoundly missing something and too old for all this fancy.... stuff???).

But, if you're going to use cast(Iterable[tuple[str]], cur), then, you have to import Iterable.

I would just remove the check for TYPE_CHECKING. You can time the imports -- you will find they are insignificant.

read_root = zarr.open(store=fname, mode='r')

I don't understand how this should work, how would zarr know it should use the SQLiteStore class to create the store based on only a file name?

Yeah, I guess this is a complaint about Zarr, not SQLiteStore.

root=zarr.open('store.zarr', mode='r')

works for a FileStore, but that wouldn't work for a ZipStore either. It seems like a registry of file types would be needed.

@auxym

auxym commented Jul 19, 2026

Copy link
Copy Markdown
OwnerAuthor

Oh yeah, missed those cast calls. The existing minimal tests did not test that code path. I added a whole bunch of tests (with LLM assistance) and I completely removed the if TYPE_CHECKING as you suggested.

Regarding the async stuff, my understanding is that zarr-python uses async "under the hood", but wraps a sync API around it. Thus the user of zarr can use either sync or async calls.

@newville

Copy link
Copy Markdown

@auxym Thanks.

It seems to me that the point of cast is to tell the type checker not to check types. ;) What happens if you replace for row in cast(Iterable[tuple[str]], cur): with for row in cur:? Does some IDE complain?

Regarding the async stuff, my understanding is that zarr-python uses async "under the hood", but wraps a sync API around it. Thus the user of zarr can use either sync or async calls.

I can believe all of that.

I admit a bias against async (I support a lot of Python projects for data collection and analysis, and I just never use async, even when asynchronous callbacks from a distributed control system are a vital part of the software. Threads for specific tasks, multiprocessing, yes definitely).

Async adds a bunch of keywords all over Python to bolt on coroutines that are really about asynchronous I/O (and not anything else), and then requires adding a bunch more keywords ("await"s) all over Python code to say "don't be asynchronous. You can ignore this as a rant, but I think hat

deflist(self):
return [row[0] forrowinself._execute("select k from zarr")]

has merit that

@overrideasyncdeflist(self) ->AsyncIterator[str]:
cur=awaitself._execute("SELECT k FROM zarr")
forrowincast(Iterable[tuple[str]], cur):
yieldrow[0]

does not necessarily improve. There is a bunch of code to signal non-runtime type-checkers to ignore type checking. There is code to "await" the (allegedly) expensive database lookup, and then yields over a list of values that has already been synchronously fetched. [If "cast key to str" is needed, then I think AsyncIterator[str] is not guaranteed. Perhaps the yielded value should be str(row[0])?].

When I look in detail here, it seems that all of the calls to self._execute() and self._execute_write() are awaited. Doesn't that mean these functions could be synchronous, and sort of really are synchronous?
I'll also guess that for SQLite with WAL, these database queries (simple select in a two-column table) will be fast.
I'm comfortable saying that any claim about speed needs real benchmarks, but I'll say that goes for the claim that "using async makes stuff faster".

Similarly, where the code has await self._ensure_open(), couldn't that be if not self.is_open: open()?

Feel free to ignore all this.

@auxym

Copy link
Copy Markdown
OwnerAuthor

Yes, I believe the casts were necessary to stop mypy/pyright from complaining.

That said,

Perhaps the yielded value should be str(row[0])?].

I do like that suggestion. Latest commit got rid of all casts :)

When I look in detail here, it seems that all of the calls to self._execute() and self._execute_write()

Yeah, I wrapped all db transactions in those methods. Yes, they are currently blocking (non-async), since the stdlib sqlite3 is blocking. The goal of putting all transactions in a central place is that I aim, in the future, to implement real async transactions, either using aiosqlite3 (3rd party lib) or implementing our own background thread for db work (probably the former). I also personally do not use async, but zarr itself offers async everywhere, so it would be a better match for that pattern.

Similarly, where the code has await self._ensure_open(), couldn't that be if not self.is_open: open()?

_ensure_open is from zarr's store base class, once again I'm trying to follow zarr-python's existing patterns.

https://github.com/zarr-developers/zarr-python/blob/50b7e016590d76011382771fd52843949782c9c0/src/zarr/abc/store.py#L140

@auxym
auxym merged commit ef46459 into mainJul 24, 2026
@auxym
auxym deleted the ensure-file-created branch July 24, 2026 01:03
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

trouble using with synchronous access

2 participants

@auxym@newville
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); Ensure file is created when it doesnt exist by auxym · Pull Request #5 · auxym/zarr-sqlite-python · GitHub
Skip to content

Ensure file is created when it doesnt exist - #5

Merged
auxym merged 13 commits into
mainfrom
ensure-file-created
Jul 24, 2026
Merged

Ensure file is created when it doesnt exist#5
auxym merged 13 commits into
mainfrom
ensure-file-created

Conversation

@auxym

Copy link
Copy Markdown
Owner

Resolves#3

  • Change mode=rw to mode=rwc
  • Add unit test

- Change mode=rw to mode=rwc
- Add unit test
@auxym

Copy link
Copy Markdown
OwnerAuthor

@newville can you please confirm this fixes the issue for you also?

@newville

Copy link
Copy Markdown

@auxym Thanks -- This also needs the if TYPE_CHECKiNG test removed, and those imports always done.

I think a basic test, with no async code, such as

importzarrfromzarr_sqliteimportSQLiteStorefname='test1.zarrdb'store=SQLiteStore(fname, read_only=False)
root=zarr.open(store=store, mode='a')
group1=root.create_group('group1')
i=np.arange(80000)/60.0x=i+np.random.normal(size=len(i), scale=1.5)
y=np.sin(x/13) +0.7*np.cos(x/47) +np.random.normal(size=len(i), scale=0.05)
x.shape= (400, 200)
y.shape= (200, 400)
group1.create_array(data=x, name='xdat')
group1.create_array(data=y, name='ydat')
print(group1, list(group1.keys()))
##print("wrote data, now reading....")
time.sleep(0.25)
read_root=zarr.open(store=SQLiteStore(fname, read_only=True), mode='r')
ytest=read_root['group1/ydat'][()]
print(f"## DONE {read_root=}{ytest=}")

Something like this to create, write, and read should probably be included in the tests.

I find the use of async sort of odd here. I effectively never use async.

I would even say that the

read_root=zarr.open(store=SQLiteStore(fname, read_only=True), mode='r')

is sort of a wart here. The more obvious (to me anyway)

read_root=zarr.open(store=fname, mode='r')

fails.

@auxym

Copy link
Copy Markdown
OwnerAuthor

I agree on most things, will look into including your proposed test and the import fix. I also agree on async, I rarely use it, but I believe it is required to satisfy the zarr-python store abstract base class: https://github.com/zarr-developers/zarr-python/blob/main/src/zarr/abc/store.py

read_root = zarr.open(store=fname, mode='r')

I don't understand how this should work, how would zarr know it should use the SQLiteStore class to create the store based on only a file name?

@auxym

auxym commented Jul 17, 2026

Copy link
Copy Markdown
OwnerAuthor

OK, I have implemented your suggested test, see latest commit. It passes (even without the sleep).

For the imports, looking at it, I actually don't think it's necessary to remove the type_checking check, and I assume it could save time during module imports to leave it in place. You can check yourself: you can completely delete the if TYPE_CHECKING block and run the tests (uv run pytest) and every test still passes. Can you confirm that works for you also?

@newville

Copy link
Copy Markdown

TYPE_CHECKING is False at runtime (https://docs.python.org/3/library/typing.html#typing.TYPE_CHECKING).

So the code is not importing collections.abc.Iterable at runtime. But that is needed for more than type checking.
SQLiteStore.list() and SQLiteStore.list_prefix() cannot work without it.

I'll admit that I have little experience with (or love for) async, that I don't know why (from def list())

@overrideasyncdeflist(self) ->AsyncIterator[str]:
cur=awaitself._execute("SELECT k FROM zarr")
forrowincast(Iterable[tuple[str]], cur):
yieldrow[0]

is not (I do not mean "just"):

deflist(self):
return [row[0] forrowinself._execute("SELECT k FROM zarr")]

I can believe there is a reason for the extra code, and that cast(Iterable[tuple[str]], cur) is doing something useful, but I really do not know what that is. (And, like, the name is literally"list".... "returns a list" seems so obvious to me that I must be profoundly missing something and too old for all this fancy.... stuff???).

But, if you're going to use cast(Iterable[tuple[str]], cur), then, you have to import Iterable.

I would just remove the check for TYPE_CHECKING. You can time the imports -- you will find they are insignificant.

read_root = zarr.open(store=fname, mode='r')

I don't understand how this should work, how would zarr know it should use the SQLiteStore class to create the store based on only a file name?

Yeah, I guess this is a complaint about Zarr, not SQLiteStore.

root=zarr.open('store.zarr', mode='r')

works for a FileStore, but that wouldn't work for a ZipStore either. It seems like a registry of file types would be needed.

@auxym

auxym commented Jul 19, 2026

Copy link
Copy Markdown
OwnerAuthor

Oh yeah, missed those cast calls. The existing minimal tests did not test that code path. I added a whole bunch of tests (with LLM assistance) and I completely removed the if TYPE_CHECKING as you suggested.

Regarding the async stuff, my understanding is that zarr-python uses async "under the hood", but wraps a sync API around it. Thus the user of zarr can use either sync or async calls.

@newville

Copy link
Copy Markdown

@auxym Thanks.

It seems to me that the point of cast is to tell the type checker not to check types. ;) What happens if you replace for row in cast(Iterable[tuple[str]], cur): with for row in cur:? Does some IDE complain?

Regarding the async stuff, my understanding is that zarr-python uses async "under the hood", but wraps a sync API around it. Thus the user of zarr can use either sync or async calls.

I can believe all of that.

I admit a bias against async (I support a lot of Python projects for data collection and analysis, and I just never use async, even when asynchronous callbacks from a distributed control system are a vital part of the software. Threads for specific tasks, multiprocessing, yes definitely).

Async adds a bunch of keywords all over Python to bolt on coroutines that are really about asynchronous I/O (and not anything else), and then requires adding a bunch more keywords ("await"s) all over Python code to say "don't be asynchronous. You can ignore this as a rant, but I think hat

deflist(self):
return [row[0] forrowinself._execute("select k from zarr")]

has merit that

@overrideasyncdeflist(self) ->AsyncIterator[str]:
cur=awaitself._execute("SELECT k FROM zarr")
forrowincast(Iterable[tuple[str]], cur):
yieldrow[0]

does not necessarily improve. There is a bunch of code to signal non-runtime type-checkers to ignore type checking. There is code to "await" the (allegedly) expensive database lookup, and then yields over a list of values that has already been synchronously fetched. [If "cast key to str" is needed, then I think AsyncIterator[str] is not guaranteed. Perhaps the yielded value should be str(row[0])?].

When I look in detail here, it seems that all of the calls to self._execute() and self._execute_write() are awaited. Doesn't that mean these functions could be synchronous, and sort of really are synchronous?
I'll also guess that for SQLite with WAL, these database queries (simple select in a two-column table) will be fast.
I'm comfortable saying that any claim about speed needs real benchmarks, but I'll say that goes for the claim that "using async makes stuff faster".

Similarly, where the code has await self._ensure_open(), couldn't that be if not self.is_open: open()?

Feel free to ignore all this.

@auxym

Copy link
Copy Markdown
OwnerAuthor

Yes, I believe the casts were necessary to stop mypy/pyright from complaining.

That said,

Perhaps the yielded value should be str(row[0])?].

I do like that suggestion. Latest commit got rid of all casts :)

When I look in detail here, it seems that all of the calls to self._execute() and self._execute_write()

Yeah, I wrapped all db transactions in those methods. Yes, they are currently blocking (non-async), since the stdlib sqlite3 is blocking. The goal of putting all transactions in a central place is that I aim, in the future, to implement real async transactions, either using aiosqlite3 (3rd party lib) or implementing our own background thread for db work (probably the former). I also personally do not use async, but zarr itself offers async everywhere, so it would be a better match for that pattern.

Similarly, where the code has await self._ensure_open(), couldn't that be if not self.is_open: open()?

_ensure_open is from zarr's store base class, once again I'm trying to follow zarr-python's existing patterns.

https://github.com/zarr-developers/zarr-python/blob/50b7e016590d76011382771fd52843949782c9c0/src/zarr/abc/store.py#L140

@auxym
auxym merged commit ef46459 into mainJul 24, 2026
@auxym
auxym deleted the ensure-file-created branch July 24, 2026 01:03
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

trouble using with synchronous access

2 participants

@auxym@newville