feat(datasets): reusable dataset registry + downloader - #40

Merged
flying-sheep merged 26 commits into
scverse:mainfrom
timtreis:feat/datasets
Jun 19, 2026
Merged

feat(datasets): reusable dataset registry + downloader#40
flying-sheep merged 26 commits into
scverse:mainfrom
timtreis:feat/datasets

Conversation

@timtreis

@timtreistimtreis commented Jun 15, 2026

Copy link
Copy Markdown
Member

Motivation

Several scverse packages each reimplement the same thing: a registry of downloadable datasets + a pooch-based downloader with hash verification. squidpy, scanpy and pertpy all roll their own; there is no shared, reusable building block. This PR adds one to scverse-misc so packages can drop their bespoke infra and register only their domain-specific loaders.

fromscverse_misc.datasetsimportFetcher, register_loader@register_loader("spatialdata")defload_sd(ctx):
zip_path=ctx.download(ctx.entry.file(suffix=".zip"))
ctx.extract_archive(zip_path)
importspatialdataassdreturnsd.read_zarr(ctx.target_dir/f"{ctx.entry.name}.zarr")
sdata=Fetcher("datasets.yaml").fetch("cells")

timtreisand others added 2 commits June 15, 2026 12:44
Add a `datasets` subpackage (behind the `datasets` extra) that packages can
share instead of each reimplementing pooch-based dataset downloading:
- DatasetRegistry / DatasetEntry / FileEntry: declarative YAML registry,
supporting both full URLs (e.g. Zenodo) and base_url + s3_key.
- Fetcher: pooch download with SHA-256 verification, URL fallback, caching,
archive extraction (via FetchContext helpers).
- register_loader: pluggable loader registry keyed by the free-form dataset
`type` string, so domain loaders (image, spatialdata, visium, ...) are
registered by the consuming package. Ships a built-in `anndata` loader.
Tests cover registry parsing, URL building, loader dispatch and extraction
(no network). Verified end-to-end against a real S3-hosted SpatialData zip.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@codecov

codecovBot commented Jun 15, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 93.60%. Comparing base (e084b6a) to head (9fa4c8d).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #40 +/- ##
==========================================
+ Coverage 91.36% 93.60% +2.24% 
==========================================
Files 8 11 +3 Lines 440 532 +92 ==========================================
+ Hits 402 498 +96 + Misses 38 34 -4 
Files with missing linesCoverage Δ
src/scverse_misc/datasets/__init__.py100.00% <100.00%> (ø)
src/scverse_misc/datasets/_fetcher.py100.00% <100.00%> (ø)
src/scverse_misc/datasets/_registry.py100.00% <100.00%> (ø)

... and 2 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

timtreisand others added 2 commits June 15, 2026 12:57
- type register_loader via overloads so the decorator preserves loader types
- narrow file() suffix matching for mypy
- positional-only ctx in the built-in anndata loader to match the Loader protocol
- ignore_missing_imports for stubless optional deps (pooch, anndata, yaml) via a
single mypy override
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Multi-file datasets (e.g. a per-sample Visium layout) need to place files in a
subdirectory rather than the shared type cache dir.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- add a built-in spatialdata loader (zip -> .zarr -> read_zarr), behind the new
'spatialdata' extra
- promote anndata to a core dependency so the anndata loader works out of the box
- [datasets] extra is now just the download machinery (pooch, pyyaml, tqdm)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
timtreisand others added 4 commits June 15, 2026 13:37
Matches the Loader protocol; the pre-commit mypy hook type-checks tests too
(local runs over src/ alone missed it).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…r, retries)
- replace the hand-rolled url-fallback loop with pooch.create(...).fetch(),
gaining retry_if_failed for free
- replace manual shutil.unpack_archive with pooch's Unzip/Untar processors
(passed via FetchContext.download(processor=...))
- drop the now-redundant extract_archive helper and module logger
- Fetcher gains a 'retries' arg (pooch retry_if_failed, default 3)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- FileEntry.urls() returned a candidate list but only [0] was ever used and
pooch.create takes one url per key -> collapse to resolve_url() -> str
- remove unused FetchContext.download_all
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Address review feedback (over-engineered): remove DatasetRegistry, Fetcher and
FetchContext. Keep the typed FileEntry/DatasetEntry dataclasses; the registry is
now a plain dict[str, DatasetEntry] from parse_registry(), and downloading is a
fetch() function. Loaders are (entry, target, download, **kwargs) callables.
parse_registry folds every YAML key except type/files into entry.metadata, so it
no longer hardcodes domain fields (shape/library_id). ~216 -> ~96 lines, 6 classes
-> 2 (both pure-data).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Comment thread.gitignore Outdated
Comment threadsrc/scverse_misc/datasets/_fetcher.py Outdated
Comment threadsrc/scverse_misc/datasets/_fetcher.py Outdated
Co-authored-by: Philipp A. <flying-sheep@web.de>
Comment threadsrc/scverse_misc/datasets/_fetcher.py Outdated

@flying-sheepflying-sheep left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@flying-sheepflying-sheep left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Apart from the above, also needs param docs.

timtreisand others added 2 commits June 18, 2026 17:43
- sphinx_ext: read __scverse_misc_canonical_instance_name__ (the attr the
namespace decorator sets); the old __scverse_misc_namespace_name__ never
existed, so namespace-decorator docstrings were silently never rendered
- datasets: parse_registry drops unknown per-file YAML keys so extras
(e.g. `description`) no longer crash FileEntry(**fd)
- datasets: _load_spatialdata extracts into a per-dataset dir and finds the
store by glob("*.zarr") instead of hardcoding <name>.zarr — decouples from
zip layout, avoids collisions in the shared target
- tests: cover the download closure, both built-in loaders, file(name=...),
extra-key tolerance, and the new glob/error paths (datasets module now 100%)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@flying-sheep
flying-sheep self-requested a review June 18, 2026 16:23

@flying-sheepflying-sheep left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

looks great! one small docs problem:

Comment threadsrc/scverse_misc/datasets/_registry.py Outdated
Comment threadtests/test_datasets.py Outdated
Comment threaddocs/api.md Outdated
flying-sheepand others added 4 commits June 19, 2026 15:53
Co-authored-by: Philipp A. <flying-sheep@web.de>
- parse_registry now warns on (and still drops) unrecognised per-file keys
so typos surface, via a small `_file_entry` helper
- remove unused `calls["processor"]` capture in test_download_drives_pooch
(the FakePup.fetch method itself backs the real pup.fetch call and stays)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@flying-sheep
flying-sheep merged commit 3926c37 into scverse:mainJun 19, 2026
10 checks passed
@timtreis
timtreis deleted the feat/datasets branch June 19, 2026 14:11
timtreis added a commit to timtreis/squidpy that referenced this pull request Jun 23, 2026
Replace squidpy's internal pooch-based registry/downloader with the shared
scverse_misc.datasets system (scverse-misc[datasets]):
- _registry.py: build a scverse_misc DatasetRegistry from datasets.yaml,
folding squidpy-specific shape/library_id into the generic metadata mapping.
Drops squidpy's duplicated FileEntry/DatasetEntry/DatasetRegistry/DatasetType.
- _downloader.py: register squidpy's domain loaders (image -> ImageContainer,
visium_10x -> read.visium, spatialdata -> read_zarr) via register_loader and
override the built-in anndata loader for the shape warning. The pooch
download/verify/extract machinery now lives in scverse-misc.
- _datasets.py: public API unchanged; type dispatch uses plain strings.
- pyproject: drop direct pooch dep (now via scverse-misc[datasets]).
Net ~750 lines deleted. Public API (sq.datasets.*) is unchanged.
Depends on scverse/scverse-misc#40.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
timtreis added a commit to scverse/squidpy that referenced this pull request Aug 11, 2026
…wnloader (#1213)
* refactor(datasets): use scverse-misc dataset registry + downloader
Replace squidpy's internal pooch-based registry/downloader with the shared
scverse_misc.datasets system (scverse-misc[datasets]):
- _registry.py: build a scverse_misc DatasetRegistry from datasets.yaml,
folding squidpy-specific shape/library_id into the generic metadata mapping.
Drops squidpy's duplicated FileEntry/DatasetEntry/DatasetRegistry/DatasetType.
- _downloader.py: register squidpy's domain loaders (image -> ImageContainer,
visium_10x -> read.visium, spatialdata -> read_zarr) via register_loader and
override the built-in anndata loader for the shape warning. The pooch
download/verify/extract machinery now lives in scverse-misc.
- _datasets.py: public API unchanged; type dispatch uses plain strings.
- pyproject: drop direct pooch dep (now via scverse-misc[datasets]).
Net ~750 lines deleted. Public API (sq.datasets.*) is unchanged.
Depends on scverse/scverse-misc#40.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(datasets): use scverse-misc's built-in spatialdata loader
scverse-misc now ships a generic spatialdata loader, so squidpy no longer needs
its own; it registers only its domain loaders (image, visium_10x) plus the
anndata shape-warning override.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(datasets): consistent <datasetdir>/<type>/ cache layout for visium
Drop the redundant 'visium' prefix in visium() so downloads land in
<datasetdir>/visium_10x/<sample>/ like every other type (was doubly nested
under visium/visium_10x). Update the hires-image path assertion accordingly.
Verified: all @internet datasets tests pass (8 passed).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(ci): update prefetch script for the scverse-misc registry
registry.{anndata,image,spatialdata}_datasets were squidpy's old registry
properties, removed in the migration. Use dataset_names(type) instead. Visium
samples now cache to <datasetdir>/visium_10x/<sample>/ via the public API.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(datasets): track scverse-misc slim API (parse_registry + fetch)
scverse-misc dropped its DatasetRegistry/Fetcher/FetchContext classes for a typed
data model + functions. Adapt:
- _registry: parse_registry() -> (base_url, dict[str, DatasetEntry]); get_registry()
returns the dict, get_base_url() the base. shape/library_id/doc_header now live in
entry.metadata.
- _downloader: loaders are (entry, target, download, **kwargs); DatasetDownloader wraps
fetch(); visium uses pooch.Untar instead of a manual tarfile loop.
- _datasets/tests updated to the dict + metadata shape.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* refactor(datasets): collapse downloader wrapper, fix pooch/pyyaml deps
- replace DatasetDownloader class + get_downloader singleton + module
wrapper with a single download() over scverse_misc.datasets.fetch
- restore pooch as a direct dependency (used via pooch.Untar in the
visium loader); drop now-unused direct pyyaml dependency
- drop redundant visium() name validation, keeping its specific message
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* fix(datasets): address review feedback on scverse-misc migration
- visium(): guard against the visium_10x names specifically so a valid-but-
wrong-type name (e.g. "imc") fails with a clear error instead of dying deep
in the anndata loader with an "unexpected keyword argument" TypeError
- get_registry(): return a read-only MappingProxyType so the lru_cached
singleton can't be mutated by callers; rename _parsed() -> _load()
- cap scverse-misc[datasets] to <0.2 (0.1.x API still churning)
- document the <datasetdir>/<type>/ cache-path change (visium/->visium_10x/,
images/->image/) in the release notes
- restore the include_hires_tiff metadata-toggle test; add offline regression
tests for the visium guard; comment the intentional global anndata override
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* chore(datasets): drop notes-dev.md changelog entry (no longer used)
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Selman Özleyen <32667648+selmanozleyen@users.noreply.github.com>
@flying-sheep

Copy link
Copy Markdown
Member

Because I was interested: the pyyaml dependency is pulled in both by dask and zarr, so it costs nothing:

pipdeptree --python=(^hatch env find hatch-test.stable | path join bin/python) -r -p pyyaml,scanpy[dask]PyYAML==6.0.3┣━━ dask==2026.7.1 [requires: PyYAML>=5.4.1]┃ ┗━━ scanpy==1.14.0.dev57+g41165bfc9 [requires: dask>=2024.10, extra: dask]┗━━ donfig==0.8.1.post1 [requires: PyYAML] ┗━━ zarr==3.2.1 [requires: donfig>=0.8] ┣━━ anndata==0.13.2 [requires: zarr>=3.1] ┃ ┗━━ scanpy==1.14.0.dev57+g41165bfc9 [requires: anndata>=0.12.14, extra: dask] ┗━━ scanpy==1.14.0.dev57+g41165bfc9 [requires: zarr>=3.2]

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@timtreis@flying-sheep
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat(datasets): reusable dataset registry + downloader - #40

Merged
flying-sheep merged 26 commits into
scverse:mainfrom
timtreis:feat/datasets
Jun 19, 2026
Merged

feat(datasets): reusable dataset registry + downloader#40
flying-sheep merged 26 commits into
scverse:mainfrom
timtreis:feat/datasets

Conversation

@timtreis

@timtreistimtreis commented Jun 15, 2026

Copy link
Copy Markdown
Member

Motivation

Several scverse packages each reimplement the same thing: a registry of downloadable datasets + a pooch-based downloader with hash verification. squidpy, scanpy and pertpy all roll their own; there is no shared, reusable building block. This PR adds one to scverse-misc so packages can drop their bespoke infra and register only their domain-specific loaders.

fromscverse_misc.datasetsimportFetcher, register_loader@register_loader("spatialdata")defload_sd(ctx):
zip_path=ctx.download(ctx.entry.file(suffix=".zip"))
ctx.extract_archive(zip_path)
importspatialdataassdreturnsd.read_zarr(ctx.target_dir/f"{ctx.entry.name}.zarr")
sdata=Fetcher("datasets.yaml").fetch("cells")

timtreisand others added 2 commits June 15, 2026 12:44
Add a `datasets` subpackage (behind the `datasets` extra) that packages can
share instead of each reimplementing pooch-based dataset downloading:
- DatasetRegistry / DatasetEntry / FileEntry: declarative YAML registry,
supporting both full URLs (e.g. Zenodo) and base_url + s3_key.
- Fetcher: pooch download with SHA-256 verification, URL fallback, caching,
archive extraction (via FetchContext helpers).
- register_loader: pluggable loader registry keyed by the free-form dataset
`type` string, so domain loaders (image, spatialdata, visium, ...) are
registered by the consuming package. Ships a built-in `anndata` loader.
Tests cover registry parsing, URL building, loader dispatch and extraction
(no network). Verified end-to-end against a real S3-hosted SpatialData zip.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@codecov

codecovBot commented Jun 15, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 93.60%. Comparing base (e084b6a) to head (9fa4c8d).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #40 +/- ##
==========================================
+ Coverage 91.36% 93.60% +2.24% 
==========================================
Files 8 11 +3 Lines 440 532 +92 ==========================================
+ Hits 402 498 +96 + Misses 38 34 -4 
Files with missing linesCoverage Δ
src/scverse_misc/datasets/__init__.py100.00% <100.00%> (ø)
src/scverse_misc/datasets/_fetcher.py100.00% <100.00%> (ø)
src/scverse_misc/datasets/_registry.py100.00% <100.00%> (ø)

... and 2 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

timtreisand others added 2 commits June 15, 2026 12:57
- type register_loader via overloads so the decorator preserves loader types
- narrow file() suffix matching for mypy
- positional-only ctx in the built-in anndata loader to match the Loader protocol
- ignore_missing_imports for stubless optional deps (pooch, anndata, yaml) via a
single mypy override
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Multi-file datasets (e.g. a per-sample Visium layout) need to place files in a
subdirectory rather than the shared type cache dir.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- add a built-in spatialdata loader (zip -> .zarr -> read_zarr), behind the new
'spatialdata' extra
- promote anndata to a core dependency so the anndata loader works out of the box
- [datasets] extra is now just the download machinery (pooch, pyyaml, tqdm)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
timtreisand others added 4 commits June 15, 2026 13:37
Matches the Loader protocol; the pre-commit mypy hook type-checks tests too
(local runs over src/ alone missed it).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…r, retries)
- replace the hand-rolled url-fallback loop with pooch.create(...).fetch(),
gaining retry_if_failed for free
- replace manual shutil.unpack_archive with pooch's Unzip/Untar processors
(passed via FetchContext.download(processor=...))
- drop the now-redundant extract_archive helper and module logger
- Fetcher gains a 'retries' arg (pooch retry_if_failed, default 3)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- FileEntry.urls() returned a candidate list but only [0] was ever used and
pooch.create takes one url per key -> collapse to resolve_url() -> str
- remove unused FetchContext.download_all
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Address review feedback (over-engineered): remove DatasetRegistry, Fetcher and
FetchContext. Keep the typed FileEntry/DatasetEntry dataclasses; the registry is
now a plain dict[str, DatasetEntry] from parse_registry(), and downloading is a
fetch() function. Loaders are (entry, target, download, **kwargs) callables.
parse_registry folds every YAML key except type/files into entry.metadata, so it
no longer hardcodes domain fields (shape/library_id). ~216 -> ~96 lines, 6 classes
-> 2 (both pure-data).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Comment thread.gitignore Outdated
Comment threadsrc/scverse_misc/datasets/_fetcher.py Outdated
Comment threadsrc/scverse_misc/datasets/_fetcher.py Outdated
Co-authored-by: Philipp A. <flying-sheep@web.de>
Comment threadsrc/scverse_misc/datasets/_fetcher.py Outdated

@flying-sheepflying-sheep left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@flying-sheepflying-sheep left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Apart from the above, also needs param docs.

timtreisand others added 2 commits June 18, 2026 17:43
- sphinx_ext: read __scverse_misc_canonical_instance_name__ (the attr the
namespace decorator sets); the old __scverse_misc_namespace_name__ never
existed, so namespace-decorator docstrings were silently never rendered
- datasets: parse_registry drops unknown per-file YAML keys so extras
(e.g. `description`) no longer crash FileEntry(**fd)
- datasets: _load_spatialdata extracts into a per-dataset dir and finds the
store by glob("*.zarr") instead of hardcoding <name>.zarr — decouples from
zip layout, avoids collisions in the shared target
- tests: cover the download closure, both built-in loaders, file(name=...),
extra-key tolerance, and the new glob/error paths (datasets module now 100%)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@flying-sheep
flying-sheep self-requested a review June 18, 2026 16:23

@flying-sheepflying-sheep left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

looks great! one small docs problem:

Comment threadsrc/scverse_misc/datasets/_registry.py Outdated
Comment threadtests/test_datasets.py Outdated
Comment threaddocs/api.md Outdated
flying-sheepand others added 4 commits June 19, 2026 15:53
Co-authored-by: Philipp A. <flying-sheep@web.de>
- parse_registry now warns on (and still drops) unrecognised per-file keys
so typos surface, via a small `_file_entry` helper
- remove unused `calls["processor"]` capture in test_download_drives_pooch
(the FakePup.fetch method itself backs the real pup.fetch call and stays)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@flying-sheep
flying-sheep merged commit 3926c37 into scverse:mainJun 19, 2026
10 checks passed
@timtreis
timtreis deleted the feat/datasets branch June 19, 2026 14:11
timtreis added a commit to timtreis/squidpy that referenced this pull request Jun 23, 2026
Replace squidpy's internal pooch-based registry/downloader with the shared
scverse_misc.datasets system (scverse-misc[datasets]):
- _registry.py: build a scverse_misc DatasetRegistry from datasets.yaml,
folding squidpy-specific shape/library_id into the generic metadata mapping.
Drops squidpy's duplicated FileEntry/DatasetEntry/DatasetRegistry/DatasetType.
- _downloader.py: register squidpy's domain loaders (image -> ImageContainer,
visium_10x -> read.visium, spatialdata -> read_zarr) via register_loader and
override the built-in anndata loader for the shape warning. The pooch
download/verify/extract machinery now lives in scverse-misc.
- _datasets.py: public API unchanged; type dispatch uses plain strings.
- pyproject: drop direct pooch dep (now via scverse-misc[datasets]).
Net ~750 lines deleted. Public API (sq.datasets.*) is unchanged.
Depends on scverse/scverse-misc#40.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
timtreis added a commit to scverse/squidpy that referenced this pull request Aug 11, 2026
…wnloader (#1213)
* refactor(datasets): use scverse-misc dataset registry + downloader
Replace squidpy's internal pooch-based registry/downloader with the shared
scverse_misc.datasets system (scverse-misc[datasets]):
- _registry.py: build a scverse_misc DatasetRegistry from datasets.yaml,
folding squidpy-specific shape/library_id into the generic metadata mapping.
Drops squidpy's duplicated FileEntry/DatasetEntry/DatasetRegistry/DatasetType.
- _downloader.py: register squidpy's domain loaders (image -> ImageContainer,
visium_10x -> read.visium, spatialdata -> read_zarr) via register_loader and
override the built-in anndata loader for the shape warning. The pooch
download/verify/extract machinery now lives in scverse-misc.
- _datasets.py: public API unchanged; type dispatch uses plain strings.
- pyproject: drop direct pooch dep (now via scverse-misc[datasets]).
Net ~750 lines deleted. Public API (sq.datasets.*) is unchanged.
Depends on scverse/scverse-misc#40.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(datasets): use scverse-misc's built-in spatialdata loader
scverse-misc now ships a generic spatialdata loader, so squidpy no longer needs
its own; it registers only its domain loaders (image, visium_10x) plus the
anndata shape-warning override.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(datasets): consistent <datasetdir>/<type>/ cache layout for visium
Drop the redundant 'visium' prefix in visium() so downloads land in
<datasetdir>/visium_10x/<sample>/ like every other type (was doubly nested
under visium/visium_10x). Update the hires-image path assertion accordingly.
Verified: all @internet datasets tests pass (8 passed).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(ci): update prefetch script for the scverse-misc registry
registry.{anndata,image,spatialdata}_datasets were squidpy's old registry
properties, removed in the migration. Use dataset_names(type) instead. Visium
samples now cache to <datasetdir>/visium_10x/<sample>/ via the public API.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(datasets): track scverse-misc slim API (parse_registry + fetch)
scverse-misc dropped its DatasetRegistry/Fetcher/FetchContext classes for a typed
data model + functions. Adapt:
- _registry: parse_registry() -> (base_url, dict[str, DatasetEntry]); get_registry()
returns the dict, get_base_url() the base. shape/library_id/doc_header now live in
entry.metadata.
- _downloader: loaders are (entry, target, download, **kwargs); DatasetDownloader wraps
fetch(); visium uses pooch.Untar instead of a manual tarfile loop.
- _datasets/tests updated to the dict + metadata shape.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* refactor(datasets): collapse downloader wrapper, fix pooch/pyyaml deps
- replace DatasetDownloader class + get_downloader singleton + module
wrapper with a single download() over scverse_misc.datasets.fetch
- restore pooch as a direct dependency (used via pooch.Untar in the
visium loader); drop now-unused direct pyyaml dependency
- drop redundant visium() name validation, keeping its specific message
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* fix(datasets): address review feedback on scverse-misc migration
- visium(): guard against the visium_10x names specifically so a valid-but-
wrong-type name (e.g. "imc") fails with a clear error instead of dying deep
in the anndata loader with an "unexpected keyword argument" TypeError
- get_registry(): return a read-only MappingProxyType so the lru_cached
singleton can't be mutated by callers; rename _parsed() -> _load()
- cap scverse-misc[datasets] to <0.2 (0.1.x API still churning)
- document the <datasetdir>/<type>/ cache-path change (visium/->visium_10x/,
images/->image/) in the release notes
- restore the include_hires_tiff metadata-toggle test; add offline regression
tests for the visium guard; comment the intentional global anndata override
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* chore(datasets): drop notes-dev.md changelog entry (no longer used)
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Selman Özleyen <32667648+selmanozleyen@users.noreply.github.com>
@flying-sheep

Copy link
Copy Markdown
Member

Because I was interested: the pyyaml dependency is pulled in both by dask and zarr, so it costs nothing:

pipdeptree --python=(^hatch env find hatch-test.stable | path join bin/python) -r -p pyyaml,scanpy[dask]PyYAML==6.0.3┣━━ dask==2026.7.1 [requires: PyYAML>=5.4.1]┃ ┗━━ scanpy==1.14.0.dev57+g41165bfc9 [requires: dask>=2024.10, extra: dask]┗━━ donfig==0.8.1.post1 [requires: PyYAML] ┗━━ zarr==3.2.1 [requires: donfig>=0.8] ┣━━ anndata==0.13.2 [requires: zarr>=3.1] ┃ ┗━━ scanpy==1.14.0.dev57+g41165bfc9 [requires: anndata>=0.12.14, extra: dask] ┗━━ scanpy==1.14.0.dev57+g41165bfc9 [requires: zarr>=3.2]

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@timtreis@flying-sheep
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(datasets): reusable dataset registry + downloader - #40

Merged
flying-sheep merged 26 commits into
scverse:mainfrom
timtreis:feat/datasets
Jun 19, 2026
Merged

feat(datasets): reusable dataset registry + downloader#40
flying-sheep merged 26 commits into
scverse:mainfrom
timtreis:feat/datasets

Conversation

@timtreis

@timtreistimtreis commented Jun 15, 2026

Copy link
Copy Markdown
Member

Motivation

Several scverse packages each reimplement the same thing: a registry of downloadable datasets + a pooch-based downloader with hash verification. squidpy, scanpy and pertpy all roll their own; there is no shared, reusable building block. This PR adds one to scverse-misc so packages can drop their bespoke infra and register only their domain-specific loaders.

fromscverse_misc.datasetsimportFetcher, register_loader@register_loader("spatialdata")defload_sd(ctx):
zip_path=ctx.download(ctx.entry.file(suffix=".zip"))
ctx.extract_archive(zip_path)
importspatialdataassdreturnsd.read_zarr(ctx.target_dir/f"{ctx.entry.name}.zarr")
sdata=Fetcher("datasets.yaml").fetch("cells")

timtreisand others added 2 commits June 15, 2026 12:44
Add a `datasets` subpackage (behind the `datasets` extra) that packages can
share instead of each reimplementing pooch-based dataset downloading:
- DatasetRegistry / DatasetEntry / FileEntry: declarative YAML registry,
supporting both full URLs (e.g. Zenodo) and base_url + s3_key.
- Fetcher: pooch download with SHA-256 verification, URL fallback, caching,
archive extraction (via FetchContext helpers).
- register_loader: pluggable loader registry keyed by the free-form dataset
`type` string, so domain loaders (image, spatialdata, visium, ...) are
registered by the consuming package. Ships a built-in `anndata` loader.
Tests cover registry parsing, URL building, loader dispatch and extraction
(no network). Verified end-to-end against a real S3-hosted SpatialData zip.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@codecov

codecovBot commented Jun 15, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 93.60%. Comparing base (e084b6a) to head (9fa4c8d).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #40 +/- ##
==========================================
+ Coverage 91.36% 93.60% +2.24% 
==========================================
Files 8 11 +3 Lines 440 532 +92 ==========================================
+ Hits 402 498 +96 + Misses 38 34 -4 
Files with missing linesCoverage Δ
src/scverse_misc/datasets/__init__.py100.00% <100.00%> (ø)
src/scverse_misc/datasets/_fetcher.py100.00% <100.00%> (ø)
src/scverse_misc/datasets/_registry.py100.00% <100.00%> (ø)

... and 2 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

timtreisand others added 2 commits June 15, 2026 12:57
- type register_loader via overloads so the decorator preserves loader types
- narrow file() suffix matching for mypy
- positional-only ctx in the built-in anndata loader to match the Loader protocol
- ignore_missing_imports for stubless optional deps (pooch, anndata, yaml) via a
single mypy override
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Multi-file datasets (e.g. a per-sample Visium layout) need to place files in a
subdirectory rather than the shared type cache dir.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- add a built-in spatialdata loader (zip -> .zarr -> read_zarr), behind the new
'spatialdata' extra
- promote anndata to a core dependency so the anndata loader works out of the box
- [datasets] extra is now just the download machinery (pooch, pyyaml, tqdm)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
timtreisand others added 4 commits June 15, 2026 13:37
Matches the Loader protocol; the pre-commit mypy hook type-checks tests too
(local runs over src/ alone missed it).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…r, retries)
- replace the hand-rolled url-fallback loop with pooch.create(...).fetch(),
gaining retry_if_failed for free
- replace manual shutil.unpack_archive with pooch's Unzip/Untar processors
(passed via FetchContext.download(processor=...))
- drop the now-redundant extract_archive helper and module logger
- Fetcher gains a 'retries' arg (pooch retry_if_failed, default 3)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- FileEntry.urls() returned a candidate list but only [0] was ever used and
pooch.create takes one url per key -> collapse to resolve_url() -> str
- remove unused FetchContext.download_all
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Address review feedback (over-engineered): remove DatasetRegistry, Fetcher and
FetchContext. Keep the typed FileEntry/DatasetEntry dataclasses; the registry is
now a plain dict[str, DatasetEntry] from parse_registry(), and downloading is a
fetch() function. Loaders are (entry, target, download, **kwargs) callables.
parse_registry folds every YAML key except type/files into entry.metadata, so it
no longer hardcodes domain fields (shape/library_id). ~216 -> ~96 lines, 6 classes
-> 2 (both pure-data).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Comment thread.gitignore Outdated
Comment threadsrc/scverse_misc/datasets/_fetcher.py Outdated
Comment threadsrc/scverse_misc/datasets/_fetcher.py Outdated
Co-authored-by: Philipp A. <flying-sheep@web.de>
Comment threadsrc/scverse_misc/datasets/_fetcher.py Outdated

@flying-sheepflying-sheep left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@flying-sheepflying-sheep left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Apart from the above, also needs param docs.

timtreisand others added 2 commits June 18, 2026 17:43
- sphinx_ext: read __scverse_misc_canonical_instance_name__ (the attr the
namespace decorator sets); the old __scverse_misc_namespace_name__ never
existed, so namespace-decorator docstrings were silently never rendered
- datasets: parse_registry drops unknown per-file YAML keys so extras
(e.g. `description`) no longer crash FileEntry(**fd)
- datasets: _load_spatialdata extracts into a per-dataset dir and finds the
store by glob("*.zarr") instead of hardcoding <name>.zarr — decouples from
zip layout, avoids collisions in the shared target
- tests: cover the download closure, both built-in loaders, file(name=...),
extra-key tolerance, and the new glob/error paths (datasets module now 100%)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@flying-sheep
flying-sheep self-requested a review June 18, 2026 16:23

@flying-sheepflying-sheep left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

looks great! one small docs problem:

Comment threadsrc/scverse_misc/datasets/_registry.py Outdated
Comment threadtests/test_datasets.py Outdated
Comment threaddocs/api.md Outdated
flying-sheepand others added 4 commits June 19, 2026 15:53
Co-authored-by: Philipp A. <flying-sheep@web.de>
- parse_registry now warns on (and still drops) unrecognised per-file keys
so typos surface, via a small `_file_entry` helper
- remove unused `calls["processor"]` capture in test_download_drives_pooch
(the FakePup.fetch method itself backs the real pup.fetch call and stays)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@flying-sheep
flying-sheep merged commit 3926c37 into scverse:mainJun 19, 2026
10 checks passed
@timtreis
timtreis deleted the feat/datasets branch June 19, 2026 14:11
timtreis added a commit to timtreis/squidpy that referenced this pull request Jun 23, 2026
Replace squidpy's internal pooch-based registry/downloader with the shared
scverse_misc.datasets system (scverse-misc[datasets]):
- _registry.py: build a scverse_misc DatasetRegistry from datasets.yaml,
folding squidpy-specific shape/library_id into the generic metadata mapping.
Drops squidpy's duplicated FileEntry/DatasetEntry/DatasetRegistry/DatasetType.
- _downloader.py: register squidpy's domain loaders (image -> ImageContainer,
visium_10x -> read.visium, spatialdata -> read_zarr) via register_loader and
override the built-in anndata loader for the shape warning. The pooch
download/verify/extract machinery now lives in scverse-misc.
- _datasets.py: public API unchanged; type dispatch uses plain strings.
- pyproject: drop direct pooch dep (now via scverse-misc[datasets]).
Net ~750 lines deleted. Public API (sq.datasets.*) is unchanged.
Depends on scverse/scverse-misc#40.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
timtreis added a commit to scverse/squidpy that referenced this pull request Aug 11, 2026
…wnloader (#1213)
* refactor(datasets): use scverse-misc dataset registry + downloader
Replace squidpy's internal pooch-based registry/downloader with the shared
scverse_misc.datasets system (scverse-misc[datasets]):
- _registry.py: build a scverse_misc DatasetRegistry from datasets.yaml,
folding squidpy-specific shape/library_id into the generic metadata mapping.
Drops squidpy's duplicated FileEntry/DatasetEntry/DatasetRegistry/DatasetType.
- _downloader.py: register squidpy's domain loaders (image -> ImageContainer,
visium_10x -> read.visium, spatialdata -> read_zarr) via register_loader and
override the built-in anndata loader for the shape warning. The pooch
download/verify/extract machinery now lives in scverse-misc.
- _datasets.py: public API unchanged; type dispatch uses plain strings.
- pyproject: drop direct pooch dep (now via scverse-misc[datasets]).
Net ~750 lines deleted. Public API (sq.datasets.*) is unchanged.
Depends on scverse/scverse-misc#40.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(datasets): use scverse-misc's built-in spatialdata loader
scverse-misc now ships a generic spatialdata loader, so squidpy no longer needs
its own; it registers only its domain loaders (image, visium_10x) plus the
anndata shape-warning override.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(datasets): consistent <datasetdir>/<type>/ cache layout for visium
Drop the redundant 'visium' prefix in visium() so downloads land in
<datasetdir>/visium_10x/<sample>/ like every other type (was doubly nested
under visium/visium_10x). Update the hires-image path assertion accordingly.
Verified: all @internet datasets tests pass (8 passed).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(ci): update prefetch script for the scverse-misc registry
registry.{anndata,image,spatialdata}_datasets were squidpy's old registry
properties, removed in the migration. Use dataset_names(type) instead. Visium
samples now cache to <datasetdir>/visium_10x/<sample>/ via the public API.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(datasets): track scverse-misc slim API (parse_registry + fetch)
scverse-misc dropped its DatasetRegistry/Fetcher/FetchContext classes for a typed
data model + functions. Adapt:
- _registry: parse_registry() -> (base_url, dict[str, DatasetEntry]); get_registry()
returns the dict, get_base_url() the base. shape/library_id/doc_header now live in
entry.metadata.
- _downloader: loaders are (entry, target, download, **kwargs); DatasetDownloader wraps
fetch(); visium uses pooch.Untar instead of a manual tarfile loop.
- _datasets/tests updated to the dict + metadata shape.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* refactor(datasets): collapse downloader wrapper, fix pooch/pyyaml deps
- replace DatasetDownloader class + get_downloader singleton + module
wrapper with a single download() over scverse_misc.datasets.fetch
- restore pooch as a direct dependency (used via pooch.Untar in the
visium loader); drop now-unused direct pyyaml dependency
- drop redundant visium() name validation, keeping its specific message
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* fix(datasets): address review feedback on scverse-misc migration
- visium(): guard against the visium_10x names specifically so a valid-but-
wrong-type name (e.g. "imc") fails with a clear error instead of dying deep
in the anndata loader with an "unexpected keyword argument" TypeError
- get_registry(): return a read-only MappingProxyType so the lru_cached
singleton can't be mutated by callers; rename _parsed() -> _load()
- cap scverse-misc[datasets] to <0.2 (0.1.x API still churning)
- document the <datasetdir>/<type>/ cache-path change (visium/->visium_10x/,
images/->image/) in the release notes
- restore the include_hires_tiff metadata-toggle test; add offline regression
tests for the visium guard; comment the intentional global anndata override
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* chore(datasets): drop notes-dev.md changelog entry (no longer used)
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Selman Özleyen <32667648+selmanozleyen@users.noreply.github.com>
@flying-sheep

Copy link
Copy Markdown
Member

Because I was interested: the pyyaml dependency is pulled in both by dask and zarr, so it costs nothing:

pipdeptree --python=(^hatch env find hatch-test.stable | path join bin/python) -r -p pyyaml,scanpy[dask]PyYAML==6.0.3┣━━ dask==2026.7.1 [requires: PyYAML>=5.4.1]┃ ┗━━ scanpy==1.14.0.dev57+g41165bfc9 [requires: dask>=2024.10, extra: dask]┗━━ donfig==0.8.1.post1 [requires: PyYAML] ┗━━ zarr==3.2.1 [requires: donfig>=0.8] ┣━━ anndata==0.13.2 [requires: zarr>=3.1] ┃ ┗━━ scanpy==1.14.0.dev57+g41165bfc9 [requires: anndata>=0.12.14, extra: dask] ┗━━ scanpy==1.14.0.dev57+g41165bfc9 [requires: zarr>=3.2]

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@timtreis@flying-sheep
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(datasets): reusable dataset registry + downloader - #40

Merged
flying-sheep merged 26 commits into
scverse:mainfrom
timtreis:feat/datasets
Jun 19, 2026
Merged

feat(datasets): reusable dataset registry + downloader#40
flying-sheep merged 26 commits into
scverse:mainfrom
timtreis:feat/datasets

Conversation

@timtreis

@timtreistimtreis commented Jun 15, 2026

Copy link
Copy Markdown
Member

Motivation

Several scverse packages each reimplement the same thing: a registry of downloadable datasets + a pooch-based downloader with hash verification. squidpy, scanpy and pertpy all roll their own; there is no shared, reusable building block. This PR adds one to scverse-misc so packages can drop their bespoke infra and register only their domain-specific loaders.

fromscverse_misc.datasetsimportFetcher, register_loader@register_loader("spatialdata")defload_sd(ctx):
zip_path=ctx.download(ctx.entry.file(suffix=".zip"))
ctx.extract_archive(zip_path)
importspatialdataassdreturnsd.read_zarr(ctx.target_dir/f"{ctx.entry.name}.zarr")
sdata=Fetcher("datasets.yaml").fetch("cells")

timtreisand others added 2 commits June 15, 2026 12:44
Add a `datasets` subpackage (behind the `datasets` extra) that packages can
share instead of each reimplementing pooch-based dataset downloading:
- DatasetRegistry / DatasetEntry / FileEntry: declarative YAML registry,
supporting both full URLs (e.g. Zenodo) and base_url + s3_key.
- Fetcher: pooch download with SHA-256 verification, URL fallback, caching,
archive extraction (via FetchContext helpers).
- register_loader: pluggable loader registry keyed by the free-form dataset
`type` string, so domain loaders (image, spatialdata, visium, ...) are
registered by the consuming package. Ships a built-in `anndata` loader.
Tests cover registry parsing, URL building, loader dispatch and extraction
(no network). Verified end-to-end against a real S3-hosted SpatialData zip.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@codecov

codecovBot commented Jun 15, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 93.60%. Comparing base (e084b6a) to head (9fa4c8d).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #40 +/- ##
==========================================
+ Coverage 91.36% 93.60% +2.24% 
==========================================
Files 8 11 +3 Lines 440 532 +92 ==========================================
+ Hits 402 498 +96 + Misses 38 34 -4 
Files with missing linesCoverage Δ
src/scverse_misc/datasets/__init__.py100.00% <100.00%> (ø)
src/scverse_misc/datasets/_fetcher.py100.00% <100.00%> (ø)
src/scverse_misc/datasets/_registry.py100.00% <100.00%> (ø)

... and 2 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

timtreisand others added 2 commits June 15, 2026 12:57
- type register_loader via overloads so the decorator preserves loader types
- narrow file() suffix matching for mypy
- positional-only ctx in the built-in anndata loader to match the Loader protocol
- ignore_missing_imports for stubless optional deps (pooch, anndata, yaml) via a
single mypy override
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Multi-file datasets (e.g. a per-sample Visium layout) need to place files in a
subdirectory rather than the shared type cache dir.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- add a built-in spatialdata loader (zip -> .zarr -> read_zarr), behind the new
'spatialdata' extra
- promote anndata to a core dependency so the anndata loader works out of the box
- [datasets] extra is now just the download machinery (pooch, pyyaml, tqdm)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
timtreisand others added 4 commits June 15, 2026 13:37
Matches the Loader protocol; the pre-commit mypy hook type-checks tests too
(local runs over src/ alone missed it).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…r, retries)
- replace the hand-rolled url-fallback loop with pooch.create(...).fetch(),
gaining retry_if_failed for free
- replace manual shutil.unpack_archive with pooch's Unzip/Untar processors
(passed via FetchContext.download(processor=...))
- drop the now-redundant extract_archive helper and module logger
- Fetcher gains a 'retries' arg (pooch retry_if_failed, default 3)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- FileEntry.urls() returned a candidate list but only [0] was ever used and
pooch.create takes one url per key -> collapse to resolve_url() -> str
- remove unused FetchContext.download_all
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Address review feedback (over-engineered): remove DatasetRegistry, Fetcher and
FetchContext. Keep the typed FileEntry/DatasetEntry dataclasses; the registry is
now a plain dict[str, DatasetEntry] from parse_registry(), and downloading is a
fetch() function. Loaders are (entry, target, download, **kwargs) callables.
parse_registry folds every YAML key except type/files into entry.metadata, so it
no longer hardcodes domain fields (shape/library_id). ~216 -> ~96 lines, 6 classes
-> 2 (both pure-data).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Comment thread.gitignore Outdated
Comment threadsrc/scverse_misc/datasets/_fetcher.py Outdated
Comment threadsrc/scverse_misc/datasets/_fetcher.py Outdated
Co-authored-by: Philipp A. <flying-sheep@web.de>
Comment threadsrc/scverse_misc/datasets/_fetcher.py Outdated

@flying-sheepflying-sheep left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@flying-sheepflying-sheep left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Apart from the above, also needs param docs.

timtreisand others added 2 commits June 18, 2026 17:43
- sphinx_ext: read __scverse_misc_canonical_instance_name__ (the attr the
namespace decorator sets); the old __scverse_misc_namespace_name__ never
existed, so namespace-decorator docstrings were silently never rendered
- datasets: parse_registry drops unknown per-file YAML keys so extras
(e.g. `description`) no longer crash FileEntry(**fd)
- datasets: _load_spatialdata extracts into a per-dataset dir and finds the
store by glob("*.zarr") instead of hardcoding <name>.zarr — decouples from
zip layout, avoids collisions in the shared target
- tests: cover the download closure, both built-in loaders, file(name=...),
extra-key tolerance, and the new glob/error paths (datasets module now 100%)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@flying-sheep
flying-sheep self-requested a review June 18, 2026 16:23

@flying-sheepflying-sheep left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

looks great! one small docs problem:

Comment threadsrc/scverse_misc/datasets/_registry.py Outdated
Comment threadtests/test_datasets.py Outdated
Comment threaddocs/api.md Outdated
flying-sheepand others added 4 commits June 19, 2026 15:53
Co-authored-by: Philipp A. <flying-sheep@web.de>
- parse_registry now warns on (and still drops) unrecognised per-file keys
so typos surface, via a small `_file_entry` helper
- remove unused `calls["processor"]` capture in test_download_drives_pooch
(the FakePup.fetch method itself backs the real pup.fetch call and stays)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@flying-sheep
flying-sheep merged commit 3926c37 into scverse:mainJun 19, 2026
10 checks passed
@timtreis
timtreis deleted the feat/datasets branch June 19, 2026 14:11
timtreis added a commit to timtreis/squidpy that referenced this pull request Jun 23, 2026
Replace squidpy's internal pooch-based registry/downloader with the shared
scverse_misc.datasets system (scverse-misc[datasets]):
- _registry.py: build a scverse_misc DatasetRegistry from datasets.yaml,
folding squidpy-specific shape/library_id into the generic metadata mapping.
Drops squidpy's duplicated FileEntry/DatasetEntry/DatasetRegistry/DatasetType.
- _downloader.py: register squidpy's domain loaders (image -> ImageContainer,
visium_10x -> read.visium, spatialdata -> read_zarr) via register_loader and
override the built-in anndata loader for the shape warning. The pooch
download/verify/extract machinery now lives in scverse-misc.
- _datasets.py: public API unchanged; type dispatch uses plain strings.
- pyproject: drop direct pooch dep (now via scverse-misc[datasets]).
Net ~750 lines deleted. Public API (sq.datasets.*) is unchanged.
Depends on scverse/scverse-misc#40.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
timtreis added a commit to scverse/squidpy that referenced this pull request Aug 11, 2026
…wnloader (#1213)
* refactor(datasets): use scverse-misc dataset registry + downloader
Replace squidpy's internal pooch-based registry/downloader with the shared
scverse_misc.datasets system (scverse-misc[datasets]):
- _registry.py: build a scverse_misc DatasetRegistry from datasets.yaml,
folding squidpy-specific shape/library_id into the generic metadata mapping.
Drops squidpy's duplicated FileEntry/DatasetEntry/DatasetRegistry/DatasetType.
- _downloader.py: register squidpy's domain loaders (image -> ImageContainer,
visium_10x -> read.visium, spatialdata -> read_zarr) via register_loader and
override the built-in anndata loader for the shape warning. The pooch
download/verify/extract machinery now lives in scverse-misc.
- _datasets.py: public API unchanged; type dispatch uses plain strings.
- pyproject: drop direct pooch dep (now via scverse-misc[datasets]).
Net ~750 lines deleted. Public API (sq.datasets.*) is unchanged.
Depends on scverse/scverse-misc#40.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(datasets): use scverse-misc's built-in spatialdata loader
scverse-misc now ships a generic spatialdata loader, so squidpy no longer needs
its own; it registers only its domain loaders (image, visium_10x) plus the
anndata shape-warning override.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(datasets): consistent <datasetdir>/<type>/ cache layout for visium
Drop the redundant 'visium' prefix in visium() so downloads land in
<datasetdir>/visium_10x/<sample>/ like every other type (was doubly nested
under visium/visium_10x). Update the hires-image path assertion accordingly.
Verified: all @internet datasets tests pass (8 passed).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(ci): update prefetch script for the scverse-misc registry
registry.{anndata,image,spatialdata}_datasets were squidpy's old registry
properties, removed in the migration. Use dataset_names(type) instead. Visium
samples now cache to <datasetdir>/visium_10x/<sample>/ via the public API.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(datasets): track scverse-misc slim API (parse_registry + fetch)
scverse-misc dropped its DatasetRegistry/Fetcher/FetchContext classes for a typed
data model + functions. Adapt:
- _registry: parse_registry() -> (base_url, dict[str, DatasetEntry]); get_registry()
returns the dict, get_base_url() the base. shape/library_id/doc_header now live in
entry.metadata.
- _downloader: loaders are (entry, target, download, **kwargs); DatasetDownloader wraps
fetch(); visium uses pooch.Untar instead of a manual tarfile loop.
- _datasets/tests updated to the dict + metadata shape.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* refactor(datasets): collapse downloader wrapper, fix pooch/pyyaml deps
- replace DatasetDownloader class + get_downloader singleton + module
wrapper with a single download() over scverse_misc.datasets.fetch
- restore pooch as a direct dependency (used via pooch.Untar in the
visium loader); drop now-unused direct pyyaml dependency
- drop redundant visium() name validation, keeping its specific message
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* fix(datasets): address review feedback on scverse-misc migration
- visium(): guard against the visium_10x names specifically so a valid-but-
wrong-type name (e.g. "imc") fails with a clear error instead of dying deep
in the anndata loader with an "unexpected keyword argument" TypeError
- get_registry(): return a read-only MappingProxyType so the lru_cached
singleton can't be mutated by callers; rename _parsed() -> _load()
- cap scverse-misc[datasets] to <0.2 (0.1.x API still churning)
- document the <datasetdir>/<type>/ cache-path change (visium/->visium_10x/,
images/->image/) in the release notes
- restore the include_hires_tiff metadata-toggle test; add offline regression
tests for the visium guard; comment the intentional global anndata override
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* chore(datasets): drop notes-dev.md changelog entry (no longer used)
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Selman Özleyen <32667648+selmanozleyen@users.noreply.github.com>
@flying-sheep

Copy link
Copy Markdown
Member

Because I was interested: the pyyaml dependency is pulled in both by dask and zarr, so it costs nothing:

pipdeptree --python=(^hatch env find hatch-test.stable | path join bin/python) -r -p pyyaml,scanpy[dask]PyYAML==6.0.3┣━━ dask==2026.7.1 [requires: PyYAML>=5.4.1]┃ ┗━━ scanpy==1.14.0.dev57+g41165bfc9 [requires: dask>=2024.10, extra: dask]┗━━ donfig==0.8.1.post1 [requires: PyYAML] ┗━━ zarr==3.2.1 [requires: donfig>=0.8] ┣━━ anndata==0.13.2 [requires: zarr>=3.1] ┃ ┗━━ scanpy==1.14.0.dev57+g41165bfc9 [requires: anndata>=0.12.14, extra: dask] ┗━━ scanpy==1.14.0.dev57+g41165bfc9 [requires: zarr>=3.2]

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@timtreis@flying-sheep
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat(datasets): reusable dataset registry + downloader - #40

Merged
flying-sheep merged 26 commits into
scverse:mainfrom
timtreis:feat/datasets
Jun 19, 2026
Merged

feat(datasets): reusable dataset registry + downloader#40
flying-sheep merged 26 commits into
scverse:mainfrom
timtreis:feat/datasets

Conversation

@timtreis

@timtreistimtreis commented Jun 15, 2026

Copy link
Copy Markdown
Member

Motivation

Several scverse packages each reimplement the same thing: a registry of downloadable datasets + a pooch-based downloader with hash verification. squidpy, scanpy and pertpy all roll their own; there is no shared, reusable building block. This PR adds one to scverse-misc so packages can drop their bespoke infra and register only their domain-specific loaders.

fromscverse_misc.datasetsimportFetcher, register_loader@register_loader("spatialdata")defload_sd(ctx):
zip_path=ctx.download(ctx.entry.file(suffix=".zip"))
ctx.extract_archive(zip_path)
importspatialdataassdreturnsd.read_zarr(ctx.target_dir/f"{ctx.entry.name}.zarr")
sdata=Fetcher("datasets.yaml").fetch("cells")

timtreisand others added 2 commits June 15, 2026 12:44
Add a `datasets` subpackage (behind the `datasets` extra) that packages can
share instead of each reimplementing pooch-based dataset downloading:
- DatasetRegistry / DatasetEntry / FileEntry: declarative YAML registry,
supporting both full URLs (e.g. Zenodo) and base_url + s3_key.
- Fetcher: pooch download with SHA-256 verification, URL fallback, caching,
archive extraction (via FetchContext helpers).
- register_loader: pluggable loader registry keyed by the free-form dataset
`type` string, so domain loaders (image, spatialdata, visium, ...) are
registered by the consuming package. Ships a built-in `anndata` loader.
Tests cover registry parsing, URL building, loader dispatch and extraction
(no network). Verified end-to-end against a real S3-hosted SpatialData zip.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@codecov

codecovBot commented Jun 15, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 93.60%. Comparing base (e084b6a) to head (9fa4c8d).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #40 +/- ##
==========================================
+ Coverage 91.36% 93.60% +2.24% 
==========================================
Files 8 11 +3 Lines 440 532 +92 ==========================================
+ Hits 402 498 +96 + Misses 38 34 -4 
Files with missing linesCoverage Δ
src/scverse_misc/datasets/__init__.py100.00% <100.00%> (ø)
src/scverse_misc/datasets/_fetcher.py100.00% <100.00%> (ø)
src/scverse_misc/datasets/_registry.py100.00% <100.00%> (ø)

... and 2 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

timtreisand others added 2 commits June 15, 2026 12:57
- type register_loader via overloads so the decorator preserves loader types
- narrow file() suffix matching for mypy
- positional-only ctx in the built-in anndata loader to match the Loader protocol
- ignore_missing_imports for stubless optional deps (pooch, anndata, yaml) via a
single mypy override
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Multi-file datasets (e.g. a per-sample Visium layout) need to place files in a
subdirectory rather than the shared type cache dir.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- add a built-in spatialdata loader (zip -> .zarr -> read_zarr), behind the new
'spatialdata' extra
- promote anndata to a core dependency so the anndata loader works out of the box
- [datasets] extra is now just the download machinery (pooch, pyyaml, tqdm)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
timtreisand others added 4 commits June 15, 2026 13:37
Matches the Loader protocol; the pre-commit mypy hook type-checks tests too
(local runs over src/ alone missed it).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…r, retries)
- replace the hand-rolled url-fallback loop with pooch.create(...).fetch(),
gaining retry_if_failed for free
- replace manual shutil.unpack_archive with pooch's Unzip/Untar processors
(passed via FetchContext.download(processor=...))
- drop the now-redundant extract_archive helper and module logger
- Fetcher gains a 'retries' arg (pooch retry_if_failed, default 3)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- FileEntry.urls() returned a candidate list but only [0] was ever used and
pooch.create takes one url per key -> collapse to resolve_url() -> str
- remove unused FetchContext.download_all
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Address review feedback (over-engineered): remove DatasetRegistry, Fetcher and
FetchContext. Keep the typed FileEntry/DatasetEntry dataclasses; the registry is
now a plain dict[str, DatasetEntry] from parse_registry(), and downloading is a
fetch() function. Loaders are (entry, target, download, **kwargs) callables.
parse_registry folds every YAML key except type/files into entry.metadata, so it
no longer hardcodes domain fields (shape/library_id). ~216 -> ~96 lines, 6 classes
-> 2 (both pure-data).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Comment thread.gitignore Outdated
Comment threadsrc/scverse_misc/datasets/_fetcher.py Outdated
Comment threadsrc/scverse_misc/datasets/_fetcher.py Outdated
Co-authored-by: Philipp A. <flying-sheep@web.de>
Comment threadsrc/scverse_misc/datasets/_fetcher.py Outdated

@flying-sheepflying-sheep left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@flying-sheepflying-sheep left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Apart from the above, also needs param docs.

timtreisand others added 2 commits June 18, 2026 17:43
- sphinx_ext: read __scverse_misc_canonical_instance_name__ (the attr the
namespace decorator sets); the old __scverse_misc_namespace_name__ never
existed, so namespace-decorator docstrings were silently never rendered
- datasets: parse_registry drops unknown per-file YAML keys so extras
(e.g. `description`) no longer crash FileEntry(**fd)
- datasets: _load_spatialdata extracts into a per-dataset dir and finds the
store by glob("*.zarr") instead of hardcoding <name>.zarr — decouples from
zip layout, avoids collisions in the shared target
- tests: cover the download closure, both built-in loaders, file(name=...),
extra-key tolerance, and the new glob/error paths (datasets module now 100%)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@flying-sheep
flying-sheep self-requested a review June 18, 2026 16:23

@flying-sheepflying-sheep left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

looks great! one small docs problem:

Comment threadsrc/scverse_misc/datasets/_registry.py Outdated
Comment threadtests/test_datasets.py Outdated
Comment threaddocs/api.md Outdated
flying-sheepand others added 4 commits June 19, 2026 15:53
Co-authored-by: Philipp A. <flying-sheep@web.de>
- parse_registry now warns on (and still drops) unrecognised per-file keys
so typos surface, via a small `_file_entry` helper
- remove unused `calls["processor"]` capture in test_download_drives_pooch
(the FakePup.fetch method itself backs the real pup.fetch call and stays)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@flying-sheep
flying-sheep merged commit 3926c37 into scverse:mainJun 19, 2026
10 checks passed
@timtreis
timtreis deleted the feat/datasets branch June 19, 2026 14:11
timtreis added a commit to timtreis/squidpy that referenced this pull request Jun 23, 2026
Replace squidpy's internal pooch-based registry/downloader with the shared
scverse_misc.datasets system (scverse-misc[datasets]):
- _registry.py: build a scverse_misc DatasetRegistry from datasets.yaml,
folding squidpy-specific shape/library_id into the generic metadata mapping.
Drops squidpy's duplicated FileEntry/DatasetEntry/DatasetRegistry/DatasetType.
- _downloader.py: register squidpy's domain loaders (image -> ImageContainer,
visium_10x -> read.visium, spatialdata -> read_zarr) via register_loader and
override the built-in anndata loader for the shape warning. The pooch
download/verify/extract machinery now lives in scverse-misc.
- _datasets.py: public API unchanged; type dispatch uses plain strings.
- pyproject: drop direct pooch dep (now via scverse-misc[datasets]).
Net ~750 lines deleted. Public API (sq.datasets.*) is unchanged.
Depends on scverse/scverse-misc#40.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
timtreis added a commit to scverse/squidpy that referenced this pull request Aug 11, 2026
…wnloader (#1213)
* refactor(datasets): use scverse-misc dataset registry + downloader
Replace squidpy's internal pooch-based registry/downloader with the shared
scverse_misc.datasets system (scverse-misc[datasets]):
- _registry.py: build a scverse_misc DatasetRegistry from datasets.yaml,
folding squidpy-specific shape/library_id into the generic metadata mapping.
Drops squidpy's duplicated FileEntry/DatasetEntry/DatasetRegistry/DatasetType.
- _downloader.py: register squidpy's domain loaders (image -> ImageContainer,
visium_10x -> read.visium, spatialdata -> read_zarr) via register_loader and
override the built-in anndata loader for the shape warning. The pooch
download/verify/extract machinery now lives in scverse-misc.
- _datasets.py: public API unchanged; type dispatch uses plain strings.
- pyproject: drop direct pooch dep (now via scverse-misc[datasets]).
Net ~750 lines deleted. Public API (sq.datasets.*) is unchanged.
Depends on scverse/scverse-misc#40.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(datasets): use scverse-misc's built-in spatialdata loader
scverse-misc now ships a generic spatialdata loader, so squidpy no longer needs
its own; it registers only its domain loaders (image, visium_10x) plus the
anndata shape-warning override.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(datasets): consistent <datasetdir>/<type>/ cache layout for visium
Drop the redundant 'visium' prefix in visium() so downloads land in
<datasetdir>/visium_10x/<sample>/ like every other type (was doubly nested
under visium/visium_10x). Update the hires-image path assertion accordingly.
Verified: all @internet datasets tests pass (8 passed).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(ci): update prefetch script for the scverse-misc registry
registry.{anndata,image,spatialdata}_datasets were squidpy's old registry
properties, removed in the migration. Use dataset_names(type) instead. Visium
samples now cache to <datasetdir>/visium_10x/<sample>/ via the public API.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(datasets): track scverse-misc slim API (parse_registry + fetch)
scverse-misc dropped its DatasetRegistry/Fetcher/FetchContext classes for a typed
data model + functions. Adapt:
- _registry: parse_registry() -> (base_url, dict[str, DatasetEntry]); get_registry()
returns the dict, get_base_url() the base. shape/library_id/doc_header now live in
entry.metadata.
- _downloader: loaders are (entry, target, download, **kwargs); DatasetDownloader wraps
fetch(); visium uses pooch.Untar instead of a manual tarfile loop.
- _datasets/tests updated to the dict + metadata shape.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* refactor(datasets): collapse downloader wrapper, fix pooch/pyyaml deps
- replace DatasetDownloader class + get_downloader singleton + module
wrapper with a single download() over scverse_misc.datasets.fetch
- restore pooch as a direct dependency (used via pooch.Untar in the
visium loader); drop now-unused direct pyyaml dependency
- drop redundant visium() name validation, keeping its specific message
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* fix(datasets): address review feedback on scverse-misc migration
- visium(): guard against the visium_10x names specifically so a valid-but-
wrong-type name (e.g. "imc") fails with a clear error instead of dying deep
in the anndata loader with an "unexpected keyword argument" TypeError
- get_registry(): return a read-only MappingProxyType so the lru_cached
singleton can't be mutated by callers; rename _parsed() -> _load()
- cap scverse-misc[datasets] to <0.2 (0.1.x API still churning)
- document the <datasetdir>/<type>/ cache-path change (visium/->visium_10x/,
images/->image/) in the release notes
- restore the include_hires_tiff metadata-toggle test; add offline regression
tests for the visium guard; comment the intentional global anndata override
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* chore(datasets): drop notes-dev.md changelog entry (no longer used)
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Selman Özleyen <32667648+selmanozleyen@users.noreply.github.com>
@flying-sheep

Copy link
Copy Markdown
Member

Because I was interested: the pyyaml dependency is pulled in both by dask and zarr, so it costs nothing:

pipdeptree --python=(^hatch env find hatch-test.stable | path join bin/python) -r -p pyyaml,scanpy[dask]PyYAML==6.0.3┣━━ dask==2026.7.1 [requires: PyYAML>=5.4.1]┃ ┗━━ scanpy==1.14.0.dev57+g41165bfc9 [requires: dask>=2024.10, extra: dask]┗━━ donfig==0.8.1.post1 [requires: PyYAML] ┗━━ zarr==3.2.1 [requires: donfig>=0.8] ┣━━ anndata==0.13.2 [requires: zarr>=3.1] ┃ ┗━━ scanpy==1.14.0.dev57+g41165bfc9 [requires: anndata>=0.12.14, extra: dask] ┗━━ scanpy==1.14.0.dev57+g41165bfc9 [requires: zarr>=3.2]

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@timtreis@flying-sheep
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(datasets): reusable dataset registry + downloader - #40

Merged
flying-sheep merged 26 commits into
scverse:mainfrom
timtreis:feat/datasets
Jun 19, 2026
Merged

feat(datasets): reusable dataset registry + downloader#40
flying-sheep merged 26 commits into
scverse:mainfrom
timtreis:feat/datasets

Conversation

@timtreis

@timtreistimtreis commented Jun 15, 2026

Copy link
Copy Markdown
Member

Motivation

Several scverse packages each reimplement the same thing: a registry of downloadable datasets + a pooch-based downloader with hash verification. squidpy, scanpy and pertpy all roll their own; there is no shared, reusable building block. This PR adds one to scverse-misc so packages can drop their bespoke infra and register only their domain-specific loaders.

fromscverse_misc.datasetsimportFetcher, register_loader@register_loader("spatialdata")defload_sd(ctx):
zip_path=ctx.download(ctx.entry.file(suffix=".zip"))
ctx.extract_archive(zip_path)
importspatialdataassdreturnsd.read_zarr(ctx.target_dir/f"{ctx.entry.name}.zarr")
sdata=Fetcher("datasets.yaml").fetch("cells")

timtreisand others added 2 commits June 15, 2026 12:44
Add a `datasets` subpackage (behind the `datasets` extra) that packages can
share instead of each reimplementing pooch-based dataset downloading:
- DatasetRegistry / DatasetEntry / FileEntry: declarative YAML registry,
supporting both full URLs (e.g. Zenodo) and base_url + s3_key.
- Fetcher: pooch download with SHA-256 verification, URL fallback, caching,
archive extraction (via FetchContext helpers).
- register_loader: pluggable loader registry keyed by the free-form dataset
`type` string, so domain loaders (image, spatialdata, visium, ...) are
registered by the consuming package. Ships a built-in `anndata` loader.
Tests cover registry parsing, URL building, loader dispatch and extraction
(no network). Verified end-to-end against a real S3-hosted SpatialData zip.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@codecov

codecovBot commented Jun 15, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 93.60%. Comparing base (e084b6a) to head (9fa4c8d).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #40 +/- ##
==========================================
+ Coverage 91.36% 93.60% +2.24% 
==========================================
Files 8 11 +3 Lines 440 532 +92 ==========================================
+ Hits 402 498 +96 + Misses 38 34 -4 
Files with missing linesCoverage Δ
src/scverse_misc/datasets/__init__.py100.00% <100.00%> (ø)
src/scverse_misc/datasets/_fetcher.py100.00% <100.00%> (ø)
src/scverse_misc/datasets/_registry.py100.00% <100.00%> (ø)

... and 2 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

timtreisand others added 2 commits June 15, 2026 12:57
- type register_loader via overloads so the decorator preserves loader types
- narrow file() suffix matching for mypy
- positional-only ctx in the built-in anndata loader to match the Loader protocol
- ignore_missing_imports for stubless optional deps (pooch, anndata, yaml) via a
single mypy override
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Multi-file datasets (e.g. a per-sample Visium layout) need to place files in a
subdirectory rather than the shared type cache dir.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- add a built-in spatialdata loader (zip -> .zarr -> read_zarr), behind the new
'spatialdata' extra
- promote anndata to a core dependency so the anndata loader works out of the box
- [datasets] extra is now just the download machinery (pooch, pyyaml, tqdm)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
timtreisand others added 4 commits June 15, 2026 13:37
Matches the Loader protocol; the pre-commit mypy hook type-checks tests too
(local runs over src/ alone missed it).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…r, retries)
- replace the hand-rolled url-fallback loop with pooch.create(...).fetch(),
gaining retry_if_failed for free
- replace manual shutil.unpack_archive with pooch's Unzip/Untar processors
(passed via FetchContext.download(processor=...))
- drop the now-redundant extract_archive helper and module logger
- Fetcher gains a 'retries' arg (pooch retry_if_failed, default 3)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- FileEntry.urls() returned a candidate list but only [0] was ever used and
pooch.create takes one url per key -> collapse to resolve_url() -> str
- remove unused FetchContext.download_all
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Address review feedback (over-engineered): remove DatasetRegistry, Fetcher and
FetchContext. Keep the typed FileEntry/DatasetEntry dataclasses; the registry is
now a plain dict[str, DatasetEntry] from parse_registry(), and downloading is a
fetch() function. Loaders are (entry, target, download, **kwargs) callables.
parse_registry folds every YAML key except type/files into entry.metadata, so it
no longer hardcodes domain fields (shape/library_id). ~216 -> ~96 lines, 6 classes
-> 2 (both pure-data).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Comment thread.gitignore Outdated
Comment threadsrc/scverse_misc/datasets/_fetcher.py Outdated
Comment threadsrc/scverse_misc/datasets/_fetcher.py Outdated
Co-authored-by: Philipp A. <flying-sheep@web.de>
Comment threadsrc/scverse_misc/datasets/_fetcher.py Outdated

@flying-sheepflying-sheep left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@flying-sheepflying-sheep left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Apart from the above, also needs param docs.

timtreisand others added 2 commits June 18, 2026 17:43
- sphinx_ext: read __scverse_misc_canonical_instance_name__ (the attr the
namespace decorator sets); the old __scverse_misc_namespace_name__ never
existed, so namespace-decorator docstrings were silently never rendered
- datasets: parse_registry drops unknown per-file YAML keys so extras
(e.g. `description`) no longer crash FileEntry(**fd)
- datasets: _load_spatialdata extracts into a per-dataset dir and finds the
store by glob("*.zarr") instead of hardcoding <name>.zarr — decouples from
zip layout, avoids collisions in the shared target
- tests: cover the download closure, both built-in loaders, file(name=...),
extra-key tolerance, and the new glob/error paths (datasets module now 100%)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@flying-sheep
flying-sheep self-requested a review June 18, 2026 16:23

@flying-sheepflying-sheep left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

looks great! one small docs problem:

Comment threadsrc/scverse_misc/datasets/_registry.py Outdated
Comment threadtests/test_datasets.py Outdated
Comment threaddocs/api.md Outdated
flying-sheepand others added 4 commits June 19, 2026 15:53
Co-authored-by: Philipp A. <flying-sheep@web.de>
- parse_registry now warns on (and still drops) unrecognised per-file keys
so typos surface, via a small `_file_entry` helper
- remove unused `calls["processor"]` capture in test_download_drives_pooch
(the FakePup.fetch method itself backs the real pup.fetch call and stays)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@flying-sheep
flying-sheep merged commit 3926c37 into scverse:mainJun 19, 2026
10 checks passed
@timtreis
timtreis deleted the feat/datasets branch June 19, 2026 14:11
timtreis added a commit to timtreis/squidpy that referenced this pull request Jun 23, 2026
Replace squidpy's internal pooch-based registry/downloader with the shared
scverse_misc.datasets system (scverse-misc[datasets]):
- _registry.py: build a scverse_misc DatasetRegistry from datasets.yaml,
folding squidpy-specific shape/library_id into the generic metadata mapping.
Drops squidpy's duplicated FileEntry/DatasetEntry/DatasetRegistry/DatasetType.
- _downloader.py: register squidpy's domain loaders (image -> ImageContainer,
visium_10x -> read.visium, spatialdata -> read_zarr) via register_loader and
override the built-in anndata loader for the shape warning. The pooch
download/verify/extract machinery now lives in scverse-misc.
- _datasets.py: public API unchanged; type dispatch uses plain strings.
- pyproject: drop direct pooch dep (now via scverse-misc[datasets]).
Net ~750 lines deleted. Public API (sq.datasets.*) is unchanged.
Depends on scverse/scverse-misc#40.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
timtreis added a commit to scverse/squidpy that referenced this pull request Aug 11, 2026
…wnloader (#1213)
* refactor(datasets): use scverse-misc dataset registry + downloader
Replace squidpy's internal pooch-based registry/downloader with the shared
scverse_misc.datasets system (scverse-misc[datasets]):
- _registry.py: build a scverse_misc DatasetRegistry from datasets.yaml,
folding squidpy-specific shape/library_id into the generic metadata mapping.
Drops squidpy's duplicated FileEntry/DatasetEntry/DatasetRegistry/DatasetType.
- _downloader.py: register squidpy's domain loaders (image -> ImageContainer,
visium_10x -> read.visium, spatialdata -> read_zarr) via register_loader and
override the built-in anndata loader for the shape warning. The pooch
download/verify/extract machinery now lives in scverse-misc.
- _datasets.py: public API unchanged; type dispatch uses plain strings.
- pyproject: drop direct pooch dep (now via scverse-misc[datasets]).
Net ~750 lines deleted. Public API (sq.datasets.*) is unchanged.
Depends on scverse/scverse-misc#40.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(datasets): use scverse-misc's built-in spatialdata loader
scverse-misc now ships a generic spatialdata loader, so squidpy no longer needs
its own; it registers only its domain loaders (image, visium_10x) plus the
anndata shape-warning override.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(datasets): consistent <datasetdir>/<type>/ cache layout for visium
Drop the redundant 'visium' prefix in visium() so downloads land in
<datasetdir>/visium_10x/<sample>/ like every other type (was doubly nested
under visium/visium_10x). Update the hires-image path assertion accordingly.
Verified: all @internet datasets tests pass (8 passed).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(ci): update prefetch script for the scverse-misc registry
registry.{anndata,image,spatialdata}_datasets were squidpy's old registry
properties, removed in the migration. Use dataset_names(type) instead. Visium
samples now cache to <datasetdir>/visium_10x/<sample>/ via the public API.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(datasets): track scverse-misc slim API (parse_registry + fetch)
scverse-misc dropped its DatasetRegistry/Fetcher/FetchContext classes for a typed
data model + functions. Adapt:
- _registry: parse_registry() -> (base_url, dict[str, DatasetEntry]); get_registry()
returns the dict, get_base_url() the base. shape/library_id/doc_header now live in
entry.metadata.
- _downloader: loaders are (entry, target, download, **kwargs); DatasetDownloader wraps
fetch(); visium uses pooch.Untar instead of a manual tarfile loop.
- _datasets/tests updated to the dict + metadata shape.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* refactor(datasets): collapse downloader wrapper, fix pooch/pyyaml deps
- replace DatasetDownloader class + get_downloader singleton + module
wrapper with a single download() over scverse_misc.datasets.fetch
- restore pooch as a direct dependency (used via pooch.Untar in the
visium loader); drop now-unused direct pyyaml dependency
- drop redundant visium() name validation, keeping its specific message
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* fix(datasets): address review feedback on scverse-misc migration
- visium(): guard against the visium_10x names specifically so a valid-but-
wrong-type name (e.g. "imc") fails with a clear error instead of dying deep
in the anndata loader with an "unexpected keyword argument" TypeError
- get_registry(): return a read-only MappingProxyType so the lru_cached
singleton can't be mutated by callers; rename _parsed() -> _load()
- cap scverse-misc[datasets] to <0.2 (0.1.x API still churning)
- document the <datasetdir>/<type>/ cache-path change (visium/->visium_10x/,
images/->image/) in the release notes
- restore the include_hires_tiff metadata-toggle test; add offline regression
tests for the visium guard; comment the intentional global anndata override
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* chore(datasets): drop notes-dev.md changelog entry (no longer used)
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Selman Özleyen <32667648+selmanozleyen@users.noreply.github.com>
@flying-sheep

Copy link
Copy Markdown
Member

Because I was interested: the pyyaml dependency is pulled in both by dask and zarr, so it costs nothing:

pipdeptree --python=(^hatch env find hatch-test.stable | path join bin/python) -r -p pyyaml,scanpy[dask]PyYAML==6.0.3┣━━ dask==2026.7.1 [requires: PyYAML>=5.4.1]┃ ┗━━ scanpy==1.14.0.dev57+g41165bfc9 [requires: dask>=2024.10, extra: dask]┗━━ donfig==0.8.1.post1 [requires: PyYAML] ┗━━ zarr==3.2.1 [requires: donfig>=0.8] ┣━━ anndata==0.13.2 [requires: zarr>=3.1] ┃ ┗━━ scanpy==1.14.0.dev57+g41165bfc9 [requires: anndata>=0.12.14, extra: dask] ┗━━ scanpy==1.14.0.dev57+g41165bfc9 [requires: zarr>=3.2]

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@timtreis@flying-sheep
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(datasets): reusable dataset registry + downloader - #40

Merged
flying-sheep merged 26 commits into
scverse:mainfrom
timtreis:feat/datasets
Jun 19, 2026
Merged

feat(datasets): reusable dataset registry + downloader#40
flying-sheep merged 26 commits into
scverse:mainfrom
timtreis:feat/datasets

Conversation

@timtreis

@timtreistimtreis commented Jun 15, 2026

Copy link
Copy Markdown
Member

Motivation

Several scverse packages each reimplement the same thing: a registry of downloadable datasets + a pooch-based downloader with hash verification. squidpy, scanpy and pertpy all roll their own; there is no shared, reusable building block. This PR adds one to scverse-misc so packages can drop their bespoke infra and register only their domain-specific loaders.

fromscverse_misc.datasetsimportFetcher, register_loader@register_loader("spatialdata")defload_sd(ctx):
zip_path=ctx.download(ctx.entry.file(suffix=".zip"))
ctx.extract_archive(zip_path)
importspatialdataassdreturnsd.read_zarr(ctx.target_dir/f"{ctx.entry.name}.zarr")
sdata=Fetcher("datasets.yaml").fetch("cells")

timtreisand others added 2 commits June 15, 2026 12:44
Add a `datasets` subpackage (behind the `datasets` extra) that packages can
share instead of each reimplementing pooch-based dataset downloading:
- DatasetRegistry / DatasetEntry / FileEntry: declarative YAML registry,
supporting both full URLs (e.g. Zenodo) and base_url + s3_key.
- Fetcher: pooch download with SHA-256 verification, URL fallback, caching,
archive extraction (via FetchContext helpers).
- register_loader: pluggable loader registry keyed by the free-form dataset
`type` string, so domain loaders (image, spatialdata, visium, ...) are
registered by the consuming package. Ships a built-in `anndata` loader.
Tests cover registry parsing, URL building, loader dispatch and extraction
(no network). Verified end-to-end against a real S3-hosted SpatialData zip.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@codecov

codecovBot commented Jun 15, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 93.60%. Comparing base (e084b6a) to head (9fa4c8d).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #40 +/- ##
==========================================
+ Coverage 91.36% 93.60% +2.24% 
==========================================
Files 8 11 +3 Lines 440 532 +92 ==========================================
+ Hits 402 498 +96 + Misses 38 34 -4 
Files with missing linesCoverage Δ
src/scverse_misc/datasets/__init__.py100.00% <100.00%> (ø)
src/scverse_misc/datasets/_fetcher.py100.00% <100.00%> (ø)
src/scverse_misc/datasets/_registry.py100.00% <100.00%> (ø)

... and 2 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

timtreisand others added 2 commits June 15, 2026 12:57
- type register_loader via overloads so the decorator preserves loader types
- narrow file() suffix matching for mypy
- positional-only ctx in the built-in anndata loader to match the Loader protocol
- ignore_missing_imports for stubless optional deps (pooch, anndata, yaml) via a
single mypy override
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Multi-file datasets (e.g. a per-sample Visium layout) need to place files in a
subdirectory rather than the shared type cache dir.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- add a built-in spatialdata loader (zip -> .zarr -> read_zarr), behind the new
'spatialdata' extra
- promote anndata to a core dependency so the anndata loader works out of the box
- [datasets] extra is now just the download machinery (pooch, pyyaml, tqdm)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
timtreisand others added 4 commits June 15, 2026 13:37
Matches the Loader protocol; the pre-commit mypy hook type-checks tests too
(local runs over src/ alone missed it).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…r, retries)
- replace the hand-rolled url-fallback loop with pooch.create(...).fetch(),
gaining retry_if_failed for free
- replace manual shutil.unpack_archive with pooch's Unzip/Untar processors
(passed via FetchContext.download(processor=...))
- drop the now-redundant extract_archive helper and module logger
- Fetcher gains a 'retries' arg (pooch retry_if_failed, default 3)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- FileEntry.urls() returned a candidate list but only [0] was ever used and
pooch.create takes one url per key -> collapse to resolve_url() -> str
- remove unused FetchContext.download_all
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Address review feedback (over-engineered): remove DatasetRegistry, Fetcher and
FetchContext. Keep the typed FileEntry/DatasetEntry dataclasses; the registry is
now a plain dict[str, DatasetEntry] from parse_registry(), and downloading is a
fetch() function. Loaders are (entry, target, download, **kwargs) callables.
parse_registry folds every YAML key except type/files into entry.metadata, so it
no longer hardcodes domain fields (shape/library_id). ~216 -> ~96 lines, 6 classes
-> 2 (both pure-data).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Comment thread.gitignore Outdated
Comment threadsrc/scverse_misc/datasets/_fetcher.py Outdated
Comment threadsrc/scverse_misc/datasets/_fetcher.py Outdated
Co-authored-by: Philipp A. <flying-sheep@web.de>
Comment threadsrc/scverse_misc/datasets/_fetcher.py Outdated

@flying-sheepflying-sheep left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@flying-sheepflying-sheep left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Apart from the above, also needs param docs.

timtreisand others added 2 commits June 18, 2026 17:43
- sphinx_ext: read __scverse_misc_canonical_instance_name__ (the attr the
namespace decorator sets); the old __scverse_misc_namespace_name__ never
existed, so namespace-decorator docstrings were silently never rendered
- datasets: parse_registry drops unknown per-file YAML keys so extras
(e.g. `description`) no longer crash FileEntry(**fd)
- datasets: _load_spatialdata extracts into a per-dataset dir and finds the
store by glob("*.zarr") instead of hardcoding <name>.zarr — decouples from
zip layout, avoids collisions in the shared target
- tests: cover the download closure, both built-in loaders, file(name=...),
extra-key tolerance, and the new glob/error paths (datasets module now 100%)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@flying-sheep
flying-sheep self-requested a review June 18, 2026 16:23

@flying-sheepflying-sheep left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

looks great! one small docs problem:

Comment threadsrc/scverse_misc/datasets/_registry.py Outdated
Comment threadtests/test_datasets.py Outdated
Comment threaddocs/api.md Outdated
flying-sheepand others added 4 commits June 19, 2026 15:53
Co-authored-by: Philipp A. <flying-sheep@web.de>
- parse_registry now warns on (and still drops) unrecognised per-file keys
so typos surface, via a small `_file_entry` helper
- remove unused `calls["processor"]` capture in test_download_drives_pooch
(the FakePup.fetch method itself backs the real pup.fetch call and stays)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@flying-sheep
flying-sheep merged commit 3926c37 into scverse:mainJun 19, 2026
10 checks passed
@timtreis
timtreis deleted the feat/datasets branch June 19, 2026 14:11
timtreis added a commit to timtreis/squidpy that referenced this pull request Jun 23, 2026
Replace squidpy's internal pooch-based registry/downloader with the shared
scverse_misc.datasets system (scverse-misc[datasets]):
- _registry.py: build a scverse_misc DatasetRegistry from datasets.yaml,
folding squidpy-specific shape/library_id into the generic metadata mapping.
Drops squidpy's duplicated FileEntry/DatasetEntry/DatasetRegistry/DatasetType.
- _downloader.py: register squidpy's domain loaders (image -> ImageContainer,
visium_10x -> read.visium, spatialdata -> read_zarr) via register_loader and
override the built-in anndata loader for the shape warning. The pooch
download/verify/extract machinery now lives in scverse-misc.
- _datasets.py: public API unchanged; type dispatch uses plain strings.
- pyproject: drop direct pooch dep (now via scverse-misc[datasets]).
Net ~750 lines deleted. Public API (sq.datasets.*) is unchanged.
Depends on scverse/scverse-misc#40.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
timtreis added a commit to scverse/squidpy that referenced this pull request Aug 11, 2026
…wnloader (#1213)
* refactor(datasets): use scverse-misc dataset registry + downloader
Replace squidpy's internal pooch-based registry/downloader with the shared
scverse_misc.datasets system (scverse-misc[datasets]):
- _registry.py: build a scverse_misc DatasetRegistry from datasets.yaml,
folding squidpy-specific shape/library_id into the generic metadata mapping.
Drops squidpy's duplicated FileEntry/DatasetEntry/DatasetRegistry/DatasetType.
- _downloader.py: register squidpy's domain loaders (image -> ImageContainer,
visium_10x -> read.visium, spatialdata -> read_zarr) via register_loader and
override the built-in anndata loader for the shape warning. The pooch
download/verify/extract machinery now lives in scverse-misc.
- _datasets.py: public API unchanged; type dispatch uses plain strings.
- pyproject: drop direct pooch dep (now via scverse-misc[datasets]).
Net ~750 lines deleted. Public API (sq.datasets.*) is unchanged.
Depends on scverse/scverse-misc#40.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(datasets): use scverse-misc's built-in spatialdata loader
scverse-misc now ships a generic spatialdata loader, so squidpy no longer needs
its own; it registers only its domain loaders (image, visium_10x) plus the
anndata shape-warning override.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(datasets): consistent <datasetdir>/<type>/ cache layout for visium
Drop the redundant 'visium' prefix in visium() so downloads land in
<datasetdir>/visium_10x/<sample>/ like every other type (was doubly nested
under visium/visium_10x). Update the hires-image path assertion accordingly.
Verified: all @internet datasets tests pass (8 passed).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(ci): update prefetch script for the scverse-misc registry
registry.{anndata,image,spatialdata}_datasets were squidpy's old registry
properties, removed in the migration. Use dataset_names(type) instead. Visium
samples now cache to <datasetdir>/visium_10x/<sample>/ via the public API.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(datasets): track scverse-misc slim API (parse_registry + fetch)
scverse-misc dropped its DatasetRegistry/Fetcher/FetchContext classes for a typed
data model + functions. Adapt:
- _registry: parse_registry() -> (base_url, dict[str, DatasetEntry]); get_registry()
returns the dict, get_base_url() the base. shape/library_id/doc_header now live in
entry.metadata.
- _downloader: loaders are (entry, target, download, **kwargs); DatasetDownloader wraps
fetch(); visium uses pooch.Untar instead of a manual tarfile loop.
- _datasets/tests updated to the dict + metadata shape.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* refactor(datasets): collapse downloader wrapper, fix pooch/pyyaml deps
- replace DatasetDownloader class + get_downloader singleton + module
wrapper with a single download() over scverse_misc.datasets.fetch
- restore pooch as a direct dependency (used via pooch.Untar in the
visium loader); drop now-unused direct pyyaml dependency
- drop redundant visium() name validation, keeping its specific message
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* fix(datasets): address review feedback on scverse-misc migration
- visium(): guard against the visium_10x names specifically so a valid-but-
wrong-type name (e.g. "imc") fails with a clear error instead of dying deep
in the anndata loader with an "unexpected keyword argument" TypeError
- get_registry(): return a read-only MappingProxyType so the lru_cached
singleton can't be mutated by callers; rename _parsed() -> _load()
- cap scverse-misc[datasets] to <0.2 (0.1.x API still churning)
- document the <datasetdir>/<type>/ cache-path change (visium/->visium_10x/,
images/->image/) in the release notes
- restore the include_hires_tiff metadata-toggle test; add offline regression
tests for the visium guard; comment the intentional global anndata override
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* chore(datasets): drop notes-dev.md changelog entry (no longer used)
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Selman Özleyen <32667648+selmanozleyen@users.noreply.github.com>
@flying-sheep

Copy link
Copy Markdown
Member

Because I was interested: the pyyaml dependency is pulled in both by dask and zarr, so it costs nothing:

pipdeptree --python=(^hatch env find hatch-test.stable | path join bin/python) -r -p pyyaml,scanpy[dask]PyYAML==6.0.3┣━━ dask==2026.7.1 [requires: PyYAML>=5.4.1]┃ ┗━━ scanpy==1.14.0.dev57+g41165bfc9 [requires: dask>=2024.10, extra: dask]┗━━ donfig==0.8.1.post1 [requires: PyYAML] ┗━━ zarr==3.2.1 [requires: donfig>=0.8] ┣━━ anndata==0.13.2 [requires: zarr>=3.1] ┃ ┗━━ scanpy==1.14.0.dev57+g41165bfc9 [requires: anndata>=0.12.14, extra: dask] ┗━━ scanpy==1.14.0.dev57+g41165bfc9 [requires: zarr>=3.2]

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@timtreis@flying-sheep
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat(datasets): reusable dataset registry + downloader - #40

Merged
flying-sheep merged 26 commits into
scverse:mainfrom
timtreis:feat/datasets
Jun 19, 2026
Merged

feat(datasets): reusable dataset registry + downloader#40
flying-sheep merged 26 commits into
scverse:mainfrom
timtreis:feat/datasets

Conversation

@timtreis

@timtreistimtreis commented Jun 15, 2026

Copy link
Copy Markdown
Member

Motivation

Several scverse packages each reimplement the same thing: a registry of downloadable datasets + a pooch-based downloader with hash verification. squidpy, scanpy and pertpy all roll their own; there is no shared, reusable building block. This PR adds one to scverse-misc so packages can drop their bespoke infra and register only their domain-specific loaders.

fromscverse_misc.datasetsimportFetcher, register_loader@register_loader("spatialdata")defload_sd(ctx):
zip_path=ctx.download(ctx.entry.file(suffix=".zip"))
ctx.extract_archive(zip_path)
importspatialdataassdreturnsd.read_zarr(ctx.target_dir/f"{ctx.entry.name}.zarr")
sdata=Fetcher("datasets.yaml").fetch("cells")

timtreisand others added 2 commits June 15, 2026 12:44
Add a `datasets` subpackage (behind the `datasets` extra) that packages can
share instead of each reimplementing pooch-based dataset downloading:
- DatasetRegistry / DatasetEntry / FileEntry: declarative YAML registry,
supporting both full URLs (e.g. Zenodo) and base_url + s3_key.
- Fetcher: pooch download with SHA-256 verification, URL fallback, caching,
archive extraction (via FetchContext helpers).
- register_loader: pluggable loader registry keyed by the free-form dataset
`type` string, so domain loaders (image, spatialdata, visium, ...) are
registered by the consuming package. Ships a built-in `anndata` loader.
Tests cover registry parsing, URL building, loader dispatch and extraction
(no network). Verified end-to-end against a real S3-hosted SpatialData zip.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@codecov

codecovBot commented Jun 15, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 93.60%. Comparing base (e084b6a) to head (9fa4c8d).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #40 +/- ##
==========================================
+ Coverage 91.36% 93.60% +2.24% 
==========================================
Files 8 11 +3 Lines 440 532 +92 ==========================================
+ Hits 402 498 +96 + Misses 38 34 -4 
Files with missing linesCoverage Δ
src/scverse_misc/datasets/__init__.py100.00% <100.00%> (ø)
src/scverse_misc/datasets/_fetcher.py100.00% <100.00%> (ø)
src/scverse_misc/datasets/_registry.py100.00% <100.00%> (ø)

... and 2 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

timtreisand others added 2 commits June 15, 2026 12:57
- type register_loader via overloads so the decorator preserves loader types
- narrow file() suffix matching for mypy
- positional-only ctx in the built-in anndata loader to match the Loader protocol
- ignore_missing_imports for stubless optional deps (pooch, anndata, yaml) via a
single mypy override
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Multi-file datasets (e.g. a per-sample Visium layout) need to place files in a
subdirectory rather than the shared type cache dir.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- add a built-in spatialdata loader (zip -> .zarr -> read_zarr), behind the new
'spatialdata' extra
- promote anndata to a core dependency so the anndata loader works out of the box
- [datasets] extra is now just the download machinery (pooch, pyyaml, tqdm)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
timtreisand others added 4 commits June 15, 2026 13:37
Matches the Loader protocol; the pre-commit mypy hook type-checks tests too
(local runs over src/ alone missed it).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…r, retries)
- replace the hand-rolled url-fallback loop with pooch.create(...).fetch(),
gaining retry_if_failed for free
- replace manual shutil.unpack_archive with pooch's Unzip/Untar processors
(passed via FetchContext.download(processor=...))
- drop the now-redundant extract_archive helper and module logger
- Fetcher gains a 'retries' arg (pooch retry_if_failed, default 3)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- FileEntry.urls() returned a candidate list but only [0] was ever used and
pooch.create takes one url per key -> collapse to resolve_url() -> str
- remove unused FetchContext.download_all
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Address review feedback (over-engineered): remove DatasetRegistry, Fetcher and
FetchContext. Keep the typed FileEntry/DatasetEntry dataclasses; the registry is
now a plain dict[str, DatasetEntry] from parse_registry(), and downloading is a
fetch() function. Loaders are (entry, target, download, **kwargs) callables.
parse_registry folds every YAML key except type/files into entry.metadata, so it
no longer hardcodes domain fields (shape/library_id). ~216 -> ~96 lines, 6 classes
-> 2 (both pure-data).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Comment thread.gitignore Outdated
Comment threadsrc/scverse_misc/datasets/_fetcher.py Outdated
Comment threadsrc/scverse_misc/datasets/_fetcher.py Outdated
Co-authored-by: Philipp A. <flying-sheep@web.de>
Comment threadsrc/scverse_misc/datasets/_fetcher.py Outdated

@flying-sheepflying-sheep left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@flying-sheepflying-sheep left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Apart from the above, also needs param docs.

timtreisand others added 2 commits June 18, 2026 17:43
- sphinx_ext: read __scverse_misc_canonical_instance_name__ (the attr the
namespace decorator sets); the old __scverse_misc_namespace_name__ never
existed, so namespace-decorator docstrings were silently never rendered
- datasets: parse_registry drops unknown per-file YAML keys so extras
(e.g. `description`) no longer crash FileEntry(**fd)
- datasets: _load_spatialdata extracts into a per-dataset dir and finds the
store by glob("*.zarr") instead of hardcoding <name>.zarr — decouples from
zip layout, avoids collisions in the shared target
- tests: cover the download closure, both built-in loaders, file(name=...),
extra-key tolerance, and the new glob/error paths (datasets module now 100%)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@flying-sheep
flying-sheep self-requested a review June 18, 2026 16:23

@flying-sheepflying-sheep left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

looks great! one small docs problem:

Comment threadsrc/scverse_misc/datasets/_registry.py Outdated
Comment threadtests/test_datasets.py Outdated
Comment threaddocs/api.md Outdated
flying-sheepand others added 4 commits June 19, 2026 15:53
Co-authored-by: Philipp A. <flying-sheep@web.de>
- parse_registry now warns on (and still drops) unrecognised per-file keys
so typos surface, via a small `_file_entry` helper
- remove unused `calls["processor"]` capture in test_download_drives_pooch
(the FakePup.fetch method itself backs the real pup.fetch call and stays)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@flying-sheep
flying-sheep merged commit 3926c37 into scverse:mainJun 19, 2026
10 checks passed
@timtreis
timtreis deleted the feat/datasets branch June 19, 2026 14:11
timtreis added a commit to timtreis/squidpy that referenced this pull request Jun 23, 2026
Replace squidpy's internal pooch-based registry/downloader with the shared
scverse_misc.datasets system (scverse-misc[datasets]):
- _registry.py: build a scverse_misc DatasetRegistry from datasets.yaml,
folding squidpy-specific shape/library_id into the generic metadata mapping.
Drops squidpy's duplicated FileEntry/DatasetEntry/DatasetRegistry/DatasetType.
- _downloader.py: register squidpy's domain loaders (image -> ImageContainer,
visium_10x -> read.visium, spatialdata -> read_zarr) via register_loader and
override the built-in anndata loader for the shape warning. The pooch
download/verify/extract machinery now lives in scverse-misc.
- _datasets.py: public API unchanged; type dispatch uses plain strings.
- pyproject: drop direct pooch dep (now via scverse-misc[datasets]).
Net ~750 lines deleted. Public API (sq.datasets.*) is unchanged.
Depends on scverse/scverse-misc#40.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
timtreis added a commit to scverse/squidpy that referenced this pull request Aug 11, 2026
…wnloader (#1213)
* refactor(datasets): use scverse-misc dataset registry + downloader
Replace squidpy's internal pooch-based registry/downloader with the shared
scverse_misc.datasets system (scverse-misc[datasets]):
- _registry.py: build a scverse_misc DatasetRegistry from datasets.yaml,
folding squidpy-specific shape/library_id into the generic metadata mapping.
Drops squidpy's duplicated FileEntry/DatasetEntry/DatasetRegistry/DatasetType.
- _downloader.py: register squidpy's domain loaders (image -> ImageContainer,
visium_10x -> read.visium, spatialdata -> read_zarr) via register_loader and
override the built-in anndata loader for the shape warning. The pooch
download/verify/extract machinery now lives in scverse-misc.
- _datasets.py: public API unchanged; type dispatch uses plain strings.
- pyproject: drop direct pooch dep (now via scverse-misc[datasets]).
Net ~750 lines deleted. Public API (sq.datasets.*) is unchanged.
Depends on scverse/scverse-misc#40.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(datasets): use scverse-misc's built-in spatialdata loader
scverse-misc now ships a generic spatialdata loader, so squidpy no longer needs
its own; it registers only its domain loaders (image, visium_10x) plus the
anndata shape-warning override.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(datasets): consistent <datasetdir>/<type>/ cache layout for visium
Drop the redundant 'visium' prefix in visium() so downloads land in
<datasetdir>/visium_10x/<sample>/ like every other type (was doubly nested
under visium/visium_10x). Update the hires-image path assertion accordingly.
Verified: all @internet datasets tests pass (8 passed).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(ci): update prefetch script for the scverse-misc registry
registry.{anndata,image,spatialdata}_datasets were squidpy's old registry
properties, removed in the migration. Use dataset_names(type) instead. Visium
samples now cache to <datasetdir>/visium_10x/<sample>/ via the public API.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(datasets): track scverse-misc slim API (parse_registry + fetch)
scverse-misc dropped its DatasetRegistry/Fetcher/FetchContext classes for a typed
data model + functions. Adapt:
- _registry: parse_registry() -> (base_url, dict[str, DatasetEntry]); get_registry()
returns the dict, get_base_url() the base. shape/library_id/doc_header now live in
entry.metadata.
- _downloader: loaders are (entry, target, download, **kwargs); DatasetDownloader wraps
fetch(); visium uses pooch.Untar instead of a manual tarfile loop.
- _datasets/tests updated to the dict + metadata shape.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* refactor(datasets): collapse downloader wrapper, fix pooch/pyyaml deps
- replace DatasetDownloader class + get_downloader singleton + module
wrapper with a single download() over scverse_misc.datasets.fetch
- restore pooch as a direct dependency (used via pooch.Untar in the
visium loader); drop now-unused direct pyyaml dependency
- drop redundant visium() name validation, keeping its specific message
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* fix(datasets): address review feedback on scverse-misc migration
- visium(): guard against the visium_10x names specifically so a valid-but-
wrong-type name (e.g. "imc") fails with a clear error instead of dying deep
in the anndata loader with an "unexpected keyword argument" TypeError
- get_registry(): return a read-only MappingProxyType so the lru_cached
singleton can't be mutated by callers; rename _parsed() -> _load()
- cap scverse-misc[datasets] to <0.2 (0.1.x API still churning)
- document the <datasetdir>/<type>/ cache-path change (visium/->visium_10x/,
images/->image/) in the release notes
- restore the include_hires_tiff metadata-toggle test; add offline regression
tests for the visium guard; comment the intentional global anndata override
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* chore(datasets): drop notes-dev.md changelog entry (no longer used)
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Selman Özleyen <32667648+selmanozleyen@users.noreply.github.com>
@flying-sheep

Copy link
Copy Markdown
Member

Because I was interested: the pyyaml dependency is pulled in both by dask and zarr, so it costs nothing:

pipdeptree --python=(^hatch env find hatch-test.stable | path join bin/python) -r -p pyyaml,scanpy[dask]PyYAML==6.0.3┣━━ dask==2026.7.1 [requires: PyYAML>=5.4.1]┃ ┗━━ scanpy==1.14.0.dev57+g41165bfc9 [requires: dask>=2024.10, extra: dask]┗━━ donfig==0.8.1.post1 [requires: PyYAML] ┗━━ zarr==3.2.1 [requires: donfig>=0.8] ┣━━ anndata==0.13.2 [requires: zarr>=3.1] ┃ ┗━━ scanpy==1.14.0.dev57+g41165bfc9 [requires: anndata>=0.12.14, extra: dask] ┗━━ scanpy==1.14.0.dev57+g41165bfc9 [requires: zarr>=3.2]

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@timtreis@flying-sheep