Skip to content

Repository files navigation

DeepCell Types

DOI

DeepCell Types is a generalized cell-phenotyping model for spatial proteomics. It generalizes across datasets with varying marker panels by matching each image's channels against a marker / cell-type registry that ships with the package (vocab.json), so inference runs on any in-memory image with no extra data download.

License notice. Distributed under a Modified Apache License, Version 2.0 with non-commercial / academic-only carve-outs (see the LICENSE file for the full text). For any other use, including commercial use, contact vanvalenlab@gmail.com.

Installation

Install into a virtual environment (venv, conda/mamba, uv, pixi — your choice):

python -m venv dct-env &&source dct-env/bin/activate
pip install git+https://github.com/vanvalenlab/deepcell-types@master

Download the model

Downloading the checkpoint requires a free access token — register at users.deepcell.org and export it (see docs/site/API-key.md):

export DEEPCELL_ACCESS_TOKEN=<your token>
fromdeepcell_types.utilsimportdownload_model# Downloads the latest checkpoint into ~/.deepcell/models and returns its path.model_path=download_model()

Running inference

Inference needs only the checkpoint and your image as an in-memory array — no TissueNet archive required. predict resolves your channels against the packaged vocab.json automatically:

importtorchfromdeepcell_typesimportpredict# Default GPU if available, else CPU (same result, slower); use "cuda:1" etc. for a specific GPU.device="cuda"iftorch.cuda.is_available() else"cpu"# raw: numpy (C, H, W); mask: 2D label image; channel_names: list[str]; mpp: microns/pixel.# model_name accepts the path from download_model() or a path to a .pt file.labels=predict(raw, mask, channel_names, mpp, model_name=model_path, device=device)

See the tutorial for a complete walk-through.

TissueNet zarr archive (optional)

Only needed for training — inference runs entirely from the packaged vocab.json and never requires it. Registered users can download a public archive from https://users.deepcell.org (see docs/site/API-key.md):

export DEEPCELL_TYPES_ZARR_PATH=/absolute/path/to/tissuenet.zarr

Custom preprocessing (advanced)

When a FOV's predictions look implausible — usually a saturated or high-background channel steering the calls — adapt that FOV's per-channel normalization. Start with the preproc-adapt skill (skills/preproc-adapt/): an agent-driven loop that diagnoses the offending channel/op and iterates the config for you.

It drives predict's optional preprocess hook, which you can also build directly from a bounded set of ops:

fromdeepcell_typesimportpredict, make_preprocessorconfig= [
{"op": "clip_percentile", "p": 99.9},
{"op": "channel_drop", "names": ["NeuN"]}, # drop a confounding marker
{"op": "min_max_normalize"}, # model sees [0, 1]
]
labels=predict(raw, mask, channel_names, mpp, model_name=...,
device=device, preprocess=make_preprocessor(config))

The hook receives the resampled in-vocabulary (C, H, W) array and must return a (C, H, W) array in [0, 1]. With preprocess=None (default) the built-in p99.9 clip + min-max is used.

Training

Install the [train] extra (adds zarr, pandas, scikit-learn, torchmetrics, plotly, …):

pip install "deepcell-types[train] @ git+https://github.com/vanvalenlab/deepcell-types@master"

Entry points under scripts/:

  • train.py — main training loop (stage 1: backbone, weighted sampler on).
  • retrain_head.py — stage 2: freeze the backbone and retrain the residual-MLP head on the natural class distribution (sampler off). This decoupled recipe is the default and produces the best model; the resMLP head is auto-detected at inference.
  • pretrain.py — masked-marker pretraining.
  • predict.py — batched evaluation over a zarr archive.

Training scripts read config from a TissueNet zarr v3 archive; pass --zarr_dir or set DEEPCELL_TYPES_ZARR_PATH. The deepcell_types.training modules (TissueNetConfig, FullImageDataset, FocalLoss, HierarchicalLoss) can be imported directly for custom scripts.

Baselines

All four paper comparison baselines live in deepcell_types.baselines and run via python -m deepcell_types.baselines <name> (no submodules).

Fairness contract. For an apples-to-apples comparison, run every baseline and the main model with the same class-balancing sampler and model-selection split: --class_balance dct (default) and --val_split_file splits/fov_split_valsubset.json (seed 42), then evaluate on splits/fov_split_test_current.json. These flags are opt-in — parity depends on passing them consistently to every method.

  • XGBoost — on mean-marker-intensity features.
    pip install -e ".[baseline-xgboost]"
    python -m deepcell_types.baselines xgboost ... # or: xgboost-tune
  • Nimbus — UNet marker-positivity (Rumberger et al., Nature Methods 2025). Pins nimbus-inference==0.0.5 (requires Python <3.12).
    pip install -e ".[baseline-nimbus]"
    python -m deepcell_types.baselines nimbus ...
  • MAPS — MLP classifier (Nature Communications 2023).
    pip install -e ".[baseline-maps]"
    python -m deepcell_types.baselines maps ...
  • CellSighter — ResNet-50 multiplexed classifier (Amitay et al., Nature Communications 2023); pulls in torchvision.
    pip install -e ".[baseline-cellsighter]"
    python -m deepcell_types.baselines cellsighter ...

Citation

@article{deepcelltypes,
title={Generalized cell phenotyping for spatial proteomics with language-informed vision models},
author={Wang, Xuefei and Dilip, Rohit and Iqbal, Ahamed Raffey and Bussi, Yuval and Brown, Caitlin and Pradhan, Elora and Jain, Yashvardhan and Yu, Kevin and Li, Shenyi and Abt, Martin and Borner, Katy and Keren, Leeat and Yue, Yisong and Barnowski, Ross and Van Valen, David},
journal={bioRxiv},
pages={2026--07},
year={2026},
publisher={Cold Spring Harbor Laboratory},
url={https://www.biorxiv.org/content/10.1101/2024.11.02.621624v4}
}

About

No description, website, or topics provided.

Resources

Stars

10 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - vanvalenlab/deepcell-types · GitHub
Skip to content

Repository files navigation

DeepCell Types

DOI

DeepCell Types is a generalized cell-phenotyping model for spatial proteomics. It generalizes across datasets with varying marker panels by matching each image's channels against a marker / cell-type registry that ships with the package (vocab.json), so inference runs on any in-memory image with no extra data download.

License notice. Distributed under a Modified Apache License, Version 2.0 with non-commercial / academic-only carve-outs (see the LICENSE file for the full text). For any other use, including commercial use, contact vanvalenlab@gmail.com.

Installation

Install into a virtual environment (venv, conda/mamba, uv, pixi — your choice):

python -m venv dct-env &&source dct-env/bin/activate
pip install git+https://github.com/vanvalenlab/deepcell-types@master

Download the model

Downloading the checkpoint requires a free access token — register at users.deepcell.org and export it (see docs/site/API-key.md):

export DEEPCELL_ACCESS_TOKEN=<your token>
fromdeepcell_types.utilsimportdownload_model# Downloads the latest checkpoint into ~/.deepcell/models and returns its path.model_path=download_model()

Running inference

Inference needs only the checkpoint and your image as an in-memory array — no TissueNet archive required. predict resolves your channels against the packaged vocab.json automatically:

importtorchfromdeepcell_typesimportpredict# Default GPU if available, else CPU (same result, slower); use "cuda:1" etc. for a specific GPU.device="cuda"iftorch.cuda.is_available() else"cpu"# raw: numpy (C, H, W); mask: 2D label image; channel_names: list[str]; mpp: microns/pixel.# model_name accepts the path from download_model() or a path to a .pt file.labels=predict(raw, mask, channel_names, mpp, model_name=model_path, device=device)

See the tutorial for a complete walk-through.

TissueNet zarr archive (optional)

Only needed for training — inference runs entirely from the packaged vocab.json and never requires it. Registered users can download a public archive from https://users.deepcell.org (see docs/site/API-key.md):

export DEEPCELL_TYPES_ZARR_PATH=/absolute/path/to/tissuenet.zarr

Custom preprocessing (advanced)

When a FOV's predictions look implausible — usually a saturated or high-background channel steering the calls — adapt that FOV's per-channel normalization. Start with the preproc-adapt skill (skills/preproc-adapt/): an agent-driven loop that diagnoses the offending channel/op and iterates the config for you.

It drives predict's optional preprocess hook, which you can also build directly from a bounded set of ops:

fromdeepcell_typesimportpredict, make_preprocessorconfig= [
{"op": "clip_percentile", "p": 99.9},
{"op": "channel_drop", "names": ["NeuN"]}, # drop a confounding marker
{"op": "min_max_normalize"}, # model sees [0, 1]
]
labels=predict(raw, mask, channel_names, mpp, model_name=...,
device=device, preprocess=make_preprocessor(config))

The hook receives the resampled in-vocabulary (C, H, W) array and must return a (C, H, W) array in [0, 1]. With preprocess=None (default) the built-in p99.9 clip + min-max is used.

Training

Install the [train] extra (adds zarr, pandas, scikit-learn, torchmetrics, plotly, …):

pip install "deepcell-types[train] @ git+https://github.com/vanvalenlab/deepcell-types@master"

Entry points under scripts/:

  • train.py — main training loop (stage 1: backbone, weighted sampler on).
  • retrain_head.py — stage 2: freeze the backbone and retrain the residual-MLP head on the natural class distribution (sampler off). This decoupled recipe is the default and produces the best model; the resMLP head is auto-detected at inference.
  • pretrain.py — masked-marker pretraining.
  • predict.py — batched evaluation over a zarr archive.

Training scripts read config from a TissueNet zarr v3 archive; pass --zarr_dir or set DEEPCELL_TYPES_ZARR_PATH. The deepcell_types.training modules (TissueNetConfig, FullImageDataset, FocalLoss, HierarchicalLoss) can be imported directly for custom scripts.

Baselines

All four paper comparison baselines live in deepcell_types.baselines and run via python -m deepcell_types.baselines <name> (no submodules).

Fairness contract. For an apples-to-apples comparison, run every baseline and the main model with the same class-balancing sampler and model-selection split: --class_balance dct (default) and --val_split_file splits/fov_split_valsubset.json (seed 42), then evaluate on splits/fov_split_test_current.json. These flags are opt-in — parity depends on passing them consistently to every method.

  • XGBoost — on mean-marker-intensity features.
    pip install -e ".[baseline-xgboost]"
    python -m deepcell_types.baselines xgboost ... # or: xgboost-tune
  • Nimbus — UNet marker-positivity (Rumberger et al., Nature Methods 2025). Pins nimbus-inference==0.0.5 (requires Python <3.12).
    pip install -e ".[baseline-nimbus]"
    python -m deepcell_types.baselines nimbus ...
  • MAPS — MLP classifier (Nature Communications 2023).
    pip install -e ".[baseline-maps]"
    python -m deepcell_types.baselines maps ...
  • CellSighter — ResNet-50 multiplexed classifier (Amitay et al., Nature Communications 2023); pulls in torchvision.
    pip install -e ".[baseline-cellsighter]"
    python -m deepcell_types.baselines cellsighter ...

Citation

@article{deepcelltypes,
title={Generalized cell phenotyping for spatial proteomics with language-informed vision models},
author={Wang, Xuefei and Dilip, Rohit and Iqbal, Ahamed Raffey and Bussi, Yuval and Brown, Caitlin and Pradhan, Elora and Jain, Yashvardhan and Yu, Kevin and Li, Shenyi and Abt, Martin and Borner, Katy and Keren, Leeat and Yue, Yisong and Barnowski, Ross and Van Valen, David},
journal={bioRxiv},
pages={2026--07},
year={2026},
publisher={Cold Spring Harbor Laboratory},
url={https://www.biorxiv.org/content/10.1101/2024.11.02.621624v4}
}

About

No description, website, or topics provided.

Resources

Stars

10 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - vanvalenlab/deepcell-types · GitHub
Skip to content

Repository files navigation

DeepCell Types

DOI

DeepCell Types is a generalized cell-phenotyping model for spatial proteomics. It generalizes across datasets with varying marker panels by matching each image's channels against a marker / cell-type registry that ships with the package (vocab.json), so inference runs on any in-memory image with no extra data download.

License notice. Distributed under a Modified Apache License, Version 2.0 with non-commercial / academic-only carve-outs (see the LICENSE file for the full text). For any other use, including commercial use, contact vanvalenlab@gmail.com.

Installation

Install into a virtual environment (venv, conda/mamba, uv, pixi — your choice):

python -m venv dct-env &&source dct-env/bin/activate
pip install git+https://github.com/vanvalenlab/deepcell-types@master

Download the model

Downloading the checkpoint requires a free access token — register at users.deepcell.org and export it (see docs/site/API-key.md):

export DEEPCELL_ACCESS_TOKEN=<your token>
fromdeepcell_types.utilsimportdownload_model# Downloads the latest checkpoint into ~/.deepcell/models and returns its path.model_path=download_model()

Running inference

Inference needs only the checkpoint and your image as an in-memory array — no TissueNet archive required. predict resolves your channels against the packaged vocab.json automatically:

importtorchfromdeepcell_typesimportpredict# Default GPU if available, else CPU (same result, slower); use "cuda:1" etc. for a specific GPU.device="cuda"iftorch.cuda.is_available() else"cpu"# raw: numpy (C, H, W); mask: 2D label image; channel_names: list[str]; mpp: microns/pixel.# model_name accepts the path from download_model() or a path to a .pt file.labels=predict(raw, mask, channel_names, mpp, model_name=model_path, device=device)

See the tutorial for a complete walk-through.

TissueNet zarr archive (optional)

Only needed for training — inference runs entirely from the packaged vocab.json and never requires it. Registered users can download a public archive from https://users.deepcell.org (see docs/site/API-key.md):

export DEEPCELL_TYPES_ZARR_PATH=/absolute/path/to/tissuenet.zarr

Custom preprocessing (advanced)

When a FOV's predictions look implausible — usually a saturated or high-background channel steering the calls — adapt that FOV's per-channel normalization. Start with the preproc-adapt skill (skills/preproc-adapt/): an agent-driven loop that diagnoses the offending channel/op and iterates the config for you.

It drives predict's optional preprocess hook, which you can also build directly from a bounded set of ops:

fromdeepcell_typesimportpredict, make_preprocessorconfig= [
{"op": "clip_percentile", "p": 99.9},
{"op": "channel_drop", "names": ["NeuN"]}, # drop a confounding marker
{"op": "min_max_normalize"}, # model sees [0, 1]
]
labels=predict(raw, mask, channel_names, mpp, model_name=...,
device=device, preprocess=make_preprocessor(config))

The hook receives the resampled in-vocabulary (C, H, W) array and must return a (C, H, W) array in [0, 1]. With preprocess=None (default) the built-in p99.9 clip + min-max is used.

Training

Install the [train] extra (adds zarr, pandas, scikit-learn, torchmetrics, plotly, …):

pip install "deepcell-types[train] @ git+https://github.com/vanvalenlab/deepcell-types@master"

Entry points under scripts/:

  • train.py — main training loop (stage 1: backbone, weighted sampler on).
  • retrain_head.py — stage 2: freeze the backbone and retrain the residual-MLP head on the natural class distribution (sampler off). This decoupled recipe is the default and produces the best model; the resMLP head is auto-detected at inference.
  • pretrain.py — masked-marker pretraining.
  • predict.py — batched evaluation over a zarr archive.

Training scripts read config from a TissueNet zarr v3 archive; pass --zarr_dir or set DEEPCELL_TYPES_ZARR_PATH. The deepcell_types.training modules (TissueNetConfig, FullImageDataset, FocalLoss, HierarchicalLoss) can be imported directly for custom scripts.

Baselines

All four paper comparison baselines live in deepcell_types.baselines and run via python -m deepcell_types.baselines <name> (no submodules).

Fairness contract. For an apples-to-apples comparison, run every baseline and the main model with the same class-balancing sampler and model-selection split: --class_balance dct (default) and --val_split_file splits/fov_split_valsubset.json (seed 42), then evaluate on splits/fov_split_test_current.json. These flags are opt-in — parity depends on passing them consistently to every method.

  • XGBoost — on mean-marker-intensity features.
    pip install -e ".[baseline-xgboost]"
    python -m deepcell_types.baselines xgboost ... # or: xgboost-tune
  • Nimbus — UNet marker-positivity (Rumberger et al., Nature Methods 2025). Pins nimbus-inference==0.0.5 (requires Python <3.12).
    pip install -e ".[baseline-nimbus]"
    python -m deepcell_types.baselines nimbus ...
  • MAPS — MLP classifier (Nature Communications 2023).
    pip install -e ".[baseline-maps]"
    python -m deepcell_types.baselines maps ...
  • CellSighter — ResNet-50 multiplexed classifier (Amitay et al., Nature Communications 2023); pulls in torchvision.
    pip install -e ".[baseline-cellsighter]"
    python -m deepcell_types.baselines cellsighter ...

Citation

@article{deepcelltypes,
title={Generalized cell phenotyping for spatial proteomics with language-informed vision models},
author={Wang, Xuefei and Dilip, Rohit and Iqbal, Ahamed Raffey and Bussi, Yuval and Brown, Caitlin and Pradhan, Elora and Jain, Yashvardhan and Yu, Kevin and Li, Shenyi and Abt, Martin and Borner, Katy and Keren, Leeat and Yue, Yisong and Barnowski, Ross and Van Valen, David},
journal={bioRxiv},
pages={2026--07},
year={2026},
publisher={Cold Spring Harbor Laboratory},
url={https://www.biorxiv.org/content/10.1101/2024.11.02.621624v4}
}

About

No description, website, or topics provided.

Resources

Stars

10 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - vanvalenlab/deepcell-types · GitHub
Skip to content

Repository files navigation

DeepCell Types

DOI

DeepCell Types is a generalized cell-phenotyping model for spatial proteomics. It generalizes across datasets with varying marker panels by matching each image's channels against a marker / cell-type registry that ships with the package (vocab.json), so inference runs on any in-memory image with no extra data download.

License notice. Distributed under a Modified Apache License, Version 2.0 with non-commercial / academic-only carve-outs (see the LICENSE file for the full text). For any other use, including commercial use, contact vanvalenlab@gmail.com.

Installation

Install into a virtual environment (venv, conda/mamba, uv, pixi — your choice):

python -m venv dct-env &&source dct-env/bin/activate
pip install git+https://github.com/vanvalenlab/deepcell-types@master

Download the model

Downloading the checkpoint requires a free access token — register at users.deepcell.org and export it (see docs/site/API-key.md):

export DEEPCELL_ACCESS_TOKEN=<your token>
fromdeepcell_types.utilsimportdownload_model# Downloads the latest checkpoint into ~/.deepcell/models and returns its path.model_path=download_model()

Running inference

Inference needs only the checkpoint and your image as an in-memory array — no TissueNet archive required. predict resolves your channels against the packaged vocab.json automatically:

importtorchfromdeepcell_typesimportpredict# Default GPU if available, else CPU (same result, slower); use "cuda:1" etc. for a specific GPU.device="cuda"iftorch.cuda.is_available() else"cpu"# raw: numpy (C, H, W); mask: 2D label image; channel_names: list[str]; mpp: microns/pixel.# model_name accepts the path from download_model() or a path to a .pt file.labels=predict(raw, mask, channel_names, mpp, model_name=model_path, device=device)

See the tutorial for a complete walk-through.

TissueNet zarr archive (optional)

Only needed for training — inference runs entirely from the packaged vocab.json and never requires it. Registered users can download a public archive from https://users.deepcell.org (see docs/site/API-key.md):

export DEEPCELL_TYPES_ZARR_PATH=/absolute/path/to/tissuenet.zarr

Custom preprocessing (advanced)

When a FOV's predictions look implausible — usually a saturated or high-background channel steering the calls — adapt that FOV's per-channel normalization. Start with the preproc-adapt skill (skills/preproc-adapt/): an agent-driven loop that diagnoses the offending channel/op and iterates the config for you.

It drives predict's optional preprocess hook, which you can also build directly from a bounded set of ops:

fromdeepcell_typesimportpredict, make_preprocessorconfig= [
{"op": "clip_percentile", "p": 99.9},
{"op": "channel_drop", "names": ["NeuN"]}, # drop a confounding marker
{"op": "min_max_normalize"}, # model sees [0, 1]
]
labels=predict(raw, mask, channel_names, mpp, model_name=...,
device=device, preprocess=make_preprocessor(config))

The hook receives the resampled in-vocabulary (C, H, W) array and must return a (C, H, W) array in [0, 1]. With preprocess=None (default) the built-in p99.9 clip + min-max is used.

Training

Install the [train] extra (adds zarr, pandas, scikit-learn, torchmetrics, plotly, …):

pip install "deepcell-types[train] @ git+https://github.com/vanvalenlab/deepcell-types@master"

Entry points under scripts/:

  • train.py — main training loop (stage 1: backbone, weighted sampler on).
  • retrain_head.py — stage 2: freeze the backbone and retrain the residual-MLP head on the natural class distribution (sampler off). This decoupled recipe is the default and produces the best model; the resMLP head is auto-detected at inference.
  • pretrain.py — masked-marker pretraining.
  • predict.py — batched evaluation over a zarr archive.

Training scripts read config from a TissueNet zarr v3 archive; pass --zarr_dir or set DEEPCELL_TYPES_ZARR_PATH. The deepcell_types.training modules (TissueNetConfig, FullImageDataset, FocalLoss, HierarchicalLoss) can be imported directly for custom scripts.

Baselines

All four paper comparison baselines live in deepcell_types.baselines and run via python -m deepcell_types.baselines <name> (no submodules).

Fairness contract. For an apples-to-apples comparison, run every baseline and the main model with the same class-balancing sampler and model-selection split: --class_balance dct (default) and --val_split_file splits/fov_split_valsubset.json (seed 42), then evaluate on splits/fov_split_test_current.json. These flags are opt-in — parity depends on passing them consistently to every method.

  • XGBoost — on mean-marker-intensity features.
    pip install -e ".[baseline-xgboost]"
    python -m deepcell_types.baselines xgboost ... # or: xgboost-tune
  • Nimbus — UNet marker-positivity (Rumberger et al., Nature Methods 2025). Pins nimbus-inference==0.0.5 (requires Python <3.12).
    pip install -e ".[baseline-nimbus]"
    python -m deepcell_types.baselines nimbus ...
  • MAPS — MLP classifier (Nature Communications 2023).
    pip install -e ".[baseline-maps]"
    python -m deepcell_types.baselines maps ...
  • CellSighter — ResNet-50 multiplexed classifier (Amitay et al., Nature Communications 2023); pulls in torchvision.
    pip install -e ".[baseline-cellsighter]"
    python -m deepcell_types.baselines cellsighter ...

Citation

@article{deepcelltypes,
title={Generalized cell phenotyping for spatial proteomics with language-informed vision models},
author={Wang, Xuefei and Dilip, Rohit and Iqbal, Ahamed Raffey and Bussi, Yuval and Brown, Caitlin and Pradhan, Elora and Jain, Yashvardhan and Yu, Kevin and Li, Shenyi and Abt, Martin and Borner, Katy and Keren, Leeat and Yue, Yisong and Barnowski, Ross and Van Valen, David},
journal={bioRxiv},
pages={2026--07},
year={2026},
publisher={Cold Spring Harbor Laboratory},
url={https://www.biorxiv.org/content/10.1101/2024.11.02.621624v4}
}

About

No description, website, or topics provided.

Resources

Stars

10 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - vanvalenlab/deepcell-types · GitHub
Skip to content

Repository files navigation

DeepCell Types

DOI

DeepCell Types is a generalized cell-phenotyping model for spatial proteomics. It generalizes across datasets with varying marker panels by matching each image's channels against a marker / cell-type registry that ships with the package (vocab.json), so inference runs on any in-memory image with no extra data download.

License notice. Distributed under a Modified Apache License, Version 2.0 with non-commercial / academic-only carve-outs (see the LICENSE file for the full text). For any other use, including commercial use, contact vanvalenlab@gmail.com.

Installation

Install into a virtual environment (venv, conda/mamba, uv, pixi — your choice):

python -m venv dct-env &&source dct-env/bin/activate
pip install git+https://github.com/vanvalenlab/deepcell-types@master

Download the model

Downloading the checkpoint requires a free access token — register at users.deepcell.org and export it (see docs/site/API-key.md):

export DEEPCELL_ACCESS_TOKEN=<your token>
fromdeepcell_types.utilsimportdownload_model# Downloads the latest checkpoint into ~/.deepcell/models and returns its path.model_path=download_model()

Running inference

Inference needs only the checkpoint and your image as an in-memory array — no TissueNet archive required. predict resolves your channels against the packaged vocab.json automatically:

importtorchfromdeepcell_typesimportpredict# Default GPU if available, else CPU (same result, slower); use "cuda:1" etc. for a specific GPU.device="cuda"iftorch.cuda.is_available() else"cpu"# raw: numpy (C, H, W); mask: 2D label image; channel_names: list[str]; mpp: microns/pixel.# model_name accepts the path from download_model() or a path to a .pt file.labels=predict(raw, mask, channel_names, mpp, model_name=model_path, device=device)

See the tutorial for a complete walk-through.

TissueNet zarr archive (optional)

Only needed for training — inference runs entirely from the packaged vocab.json and never requires it. Registered users can download a public archive from https://users.deepcell.org (see docs/site/API-key.md):

export DEEPCELL_TYPES_ZARR_PATH=/absolute/path/to/tissuenet.zarr

Custom preprocessing (advanced)

When a FOV's predictions look implausible — usually a saturated or high-background channel steering the calls — adapt that FOV's per-channel normalization. Start with the preproc-adapt skill (skills/preproc-adapt/): an agent-driven loop that diagnoses the offending channel/op and iterates the config for you.

It drives predict's optional preprocess hook, which you can also build directly from a bounded set of ops:

fromdeepcell_typesimportpredict, make_preprocessorconfig= [
{"op": "clip_percentile", "p": 99.9},
{"op": "channel_drop", "names": ["NeuN"]}, # drop a confounding marker
{"op": "min_max_normalize"}, # model sees [0, 1]
]
labels=predict(raw, mask, channel_names, mpp, model_name=...,
device=device, preprocess=make_preprocessor(config))

The hook receives the resampled in-vocabulary (C, H, W) array and must return a (C, H, W) array in [0, 1]. With preprocess=None (default) the built-in p99.9 clip + min-max is used.

Training

Install the [train] extra (adds zarr, pandas, scikit-learn, torchmetrics, plotly, …):

pip install "deepcell-types[train] @ git+https://github.com/vanvalenlab/deepcell-types@master"

Entry points under scripts/:

  • train.py — main training loop (stage 1: backbone, weighted sampler on).
  • retrain_head.py — stage 2: freeze the backbone and retrain the residual-MLP head on the natural class distribution (sampler off). This decoupled recipe is the default and produces the best model; the resMLP head is auto-detected at inference.
  • pretrain.py — masked-marker pretraining.
  • predict.py — batched evaluation over a zarr archive.

Training scripts read config from a TissueNet zarr v3 archive; pass --zarr_dir or set DEEPCELL_TYPES_ZARR_PATH. The deepcell_types.training modules (TissueNetConfig, FullImageDataset, FocalLoss, HierarchicalLoss) can be imported directly for custom scripts.

Baselines

All four paper comparison baselines live in deepcell_types.baselines and run via python -m deepcell_types.baselines <name> (no submodules).

Fairness contract. For an apples-to-apples comparison, run every baseline and the main model with the same class-balancing sampler and model-selection split: --class_balance dct (default) and --val_split_file splits/fov_split_valsubset.json (seed 42), then evaluate on splits/fov_split_test_current.json. These flags are opt-in — parity depends on passing them consistently to every method.

  • XGBoost — on mean-marker-intensity features.
    pip install -e ".[baseline-xgboost]"
    python -m deepcell_types.baselines xgboost ... # or: xgboost-tune
  • Nimbus — UNet marker-positivity (Rumberger et al., Nature Methods 2025). Pins nimbus-inference==0.0.5 (requires Python <3.12).
    pip install -e ".[baseline-nimbus]"
    python -m deepcell_types.baselines nimbus ...
  • MAPS — MLP classifier (Nature Communications 2023).
    pip install -e ".[baseline-maps]"
    python -m deepcell_types.baselines maps ...
  • CellSighter — ResNet-50 multiplexed classifier (Amitay et al., Nature Communications 2023); pulls in torchvision.
    pip install -e ".[baseline-cellsighter]"
    python -m deepcell_types.baselines cellsighter ...

Citation

@article{deepcelltypes,
title={Generalized cell phenotyping for spatial proteomics with language-informed vision models},
author={Wang, Xuefei and Dilip, Rohit and Iqbal, Ahamed Raffey and Bussi, Yuval and Brown, Caitlin and Pradhan, Elora and Jain, Yashvardhan and Yu, Kevin and Li, Shenyi and Abt, Martin and Borner, Katy and Keren, Leeat and Yue, Yisong and Barnowski, Ross and Van Valen, David},
journal={bioRxiv},
pages={2026--07},
year={2026},
publisher={Cold Spring Harbor Laboratory},
url={https://www.biorxiv.org/content/10.1101/2024.11.02.621624v4}
}

About

No description, website, or topics provided.

Resources

Stars

10 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - vanvalenlab/deepcell-types · GitHub
Skip to content

Repository files navigation

DeepCell Types

DOI

DeepCell Types is a generalized cell-phenotyping model for spatial proteomics. It generalizes across datasets with varying marker panels by matching each image's channels against a marker / cell-type registry that ships with the package (vocab.json), so inference runs on any in-memory image with no extra data download.

License notice. Distributed under a Modified Apache License, Version 2.0 with non-commercial / academic-only carve-outs (see the LICENSE file for the full text). For any other use, including commercial use, contact vanvalenlab@gmail.com.

Installation

Install into a virtual environment (venv, conda/mamba, uv, pixi — your choice):

python -m venv dct-env &&source dct-env/bin/activate
pip install git+https://github.com/vanvalenlab/deepcell-types@master

Download the model

Downloading the checkpoint requires a free access token — register at users.deepcell.org and export it (see docs/site/API-key.md):

export DEEPCELL_ACCESS_TOKEN=<your token>
fromdeepcell_types.utilsimportdownload_model# Downloads the latest checkpoint into ~/.deepcell/models and returns its path.model_path=download_model()

Running inference

Inference needs only the checkpoint and your image as an in-memory array — no TissueNet archive required. predict resolves your channels against the packaged vocab.json automatically:

importtorchfromdeepcell_typesimportpredict# Default GPU if available, else CPU (same result, slower); use "cuda:1" etc. for a specific GPU.device="cuda"iftorch.cuda.is_available() else"cpu"# raw: numpy (C, H, W); mask: 2D label image; channel_names: list[str]; mpp: microns/pixel.# model_name accepts the path from download_model() or a path to a .pt file.labels=predict(raw, mask, channel_names, mpp, model_name=model_path, device=device)

See the tutorial for a complete walk-through.

TissueNet zarr archive (optional)

Only needed for training — inference runs entirely from the packaged vocab.json and never requires it. Registered users can download a public archive from https://users.deepcell.org (see docs/site/API-key.md):

export DEEPCELL_TYPES_ZARR_PATH=/absolute/path/to/tissuenet.zarr

Custom preprocessing (advanced)

When a FOV's predictions look implausible — usually a saturated or high-background channel steering the calls — adapt that FOV's per-channel normalization. Start with the preproc-adapt skill (skills/preproc-adapt/): an agent-driven loop that diagnoses the offending channel/op and iterates the config for you.

It drives predict's optional preprocess hook, which you can also build directly from a bounded set of ops:

fromdeepcell_typesimportpredict, make_preprocessorconfig= [
{"op": "clip_percentile", "p": 99.9},
{"op": "channel_drop", "names": ["NeuN"]}, # drop a confounding marker
{"op": "min_max_normalize"}, # model sees [0, 1]
]
labels=predict(raw, mask, channel_names, mpp, model_name=...,
device=device, preprocess=make_preprocessor(config))

The hook receives the resampled in-vocabulary (C, H, W) array and must return a (C, H, W) array in [0, 1]. With preprocess=None (default) the built-in p99.9 clip + min-max is used.

Training

Install the [train] extra (adds zarr, pandas, scikit-learn, torchmetrics, plotly, …):

pip install "deepcell-types[train] @ git+https://github.com/vanvalenlab/deepcell-types@master"

Entry points under scripts/:

  • train.py — main training loop (stage 1: backbone, weighted sampler on).
  • retrain_head.py — stage 2: freeze the backbone and retrain the residual-MLP head on the natural class distribution (sampler off). This decoupled recipe is the default and produces the best model; the resMLP head is auto-detected at inference.
  • pretrain.py — masked-marker pretraining.
  • predict.py — batched evaluation over a zarr archive.

Training scripts read config from a TissueNet zarr v3 archive; pass --zarr_dir or set DEEPCELL_TYPES_ZARR_PATH. The deepcell_types.training modules (TissueNetConfig, FullImageDataset, FocalLoss, HierarchicalLoss) can be imported directly for custom scripts.

Baselines

All four paper comparison baselines live in deepcell_types.baselines and run via python -m deepcell_types.baselines <name> (no submodules).

Fairness contract. For an apples-to-apples comparison, run every baseline and the main model with the same class-balancing sampler and model-selection split: --class_balance dct (default) and --val_split_file splits/fov_split_valsubset.json (seed 42), then evaluate on splits/fov_split_test_current.json. These flags are opt-in — parity depends on passing them consistently to every method.

  • XGBoost — on mean-marker-intensity features.
    pip install -e ".[baseline-xgboost]"
    python -m deepcell_types.baselines xgboost ... # or: xgboost-tune
  • Nimbus — UNet marker-positivity (Rumberger et al., Nature Methods 2025). Pins nimbus-inference==0.0.5 (requires Python <3.12).
    pip install -e ".[baseline-nimbus]"
    python -m deepcell_types.baselines nimbus ...
  • MAPS — MLP classifier (Nature Communications 2023).
    pip install -e ".[baseline-maps]"
    python -m deepcell_types.baselines maps ...
  • CellSighter — ResNet-50 multiplexed classifier (Amitay et al., Nature Communications 2023); pulls in torchvision.
    pip install -e ".[baseline-cellsighter]"
    python -m deepcell_types.baselines cellsighter ...

Citation

@article{deepcelltypes,
title={Generalized cell phenotyping for spatial proteomics with language-informed vision models},
author={Wang, Xuefei and Dilip, Rohit and Iqbal, Ahamed Raffey and Bussi, Yuval and Brown, Caitlin and Pradhan, Elora and Jain, Yashvardhan and Yu, Kevin and Li, Shenyi and Abt, Martin and Borner, Katy and Keren, Leeat and Yue, Yisong and Barnowski, Ross and Van Valen, David},
journal={bioRxiv},
pages={2026--07},
year={2026},
publisher={Cold Spring Harbor Laboratory},
url={https://www.biorxiv.org/content/10.1101/2024.11.02.621624v4}
}

About

No description, website, or topics provided.

Resources

Stars

10 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - vanvalenlab/deepcell-types · GitHub
Skip to content

Repository files navigation

DeepCell Types

DOI

DeepCell Types is a generalized cell-phenotyping model for spatial proteomics. It generalizes across datasets with varying marker panels by matching each image's channels against a marker / cell-type registry that ships with the package (vocab.json), so inference runs on any in-memory image with no extra data download.

License notice. Distributed under a Modified Apache License, Version 2.0 with non-commercial / academic-only carve-outs (see the LICENSE file for the full text). For any other use, including commercial use, contact vanvalenlab@gmail.com.

Installation

Install into a virtual environment (venv, conda/mamba, uv, pixi — your choice):

python -m venv dct-env &&source dct-env/bin/activate
pip install git+https://github.com/vanvalenlab/deepcell-types@master

Download the model

Downloading the checkpoint requires a free access token — register at users.deepcell.org and export it (see docs/site/API-key.md):

export DEEPCELL_ACCESS_TOKEN=<your token>
fromdeepcell_types.utilsimportdownload_model# Downloads the latest checkpoint into ~/.deepcell/models and returns its path.model_path=download_model()

Running inference

Inference needs only the checkpoint and your image as an in-memory array — no TissueNet archive required. predict resolves your channels against the packaged vocab.json automatically:

importtorchfromdeepcell_typesimportpredict# Default GPU if available, else CPU (same result, slower); use "cuda:1" etc. for a specific GPU.device="cuda"iftorch.cuda.is_available() else"cpu"# raw: numpy (C, H, W); mask: 2D label image; channel_names: list[str]; mpp: microns/pixel.# model_name accepts the path from download_model() or a path to a .pt file.labels=predict(raw, mask, channel_names, mpp, model_name=model_path, device=device)

See the tutorial for a complete walk-through.

TissueNet zarr archive (optional)

Only needed for training — inference runs entirely from the packaged vocab.json and never requires it. Registered users can download a public archive from https://users.deepcell.org (see docs/site/API-key.md):

export DEEPCELL_TYPES_ZARR_PATH=/absolute/path/to/tissuenet.zarr

Custom preprocessing (advanced)

When a FOV's predictions look implausible — usually a saturated or high-background channel steering the calls — adapt that FOV's per-channel normalization. Start with the preproc-adapt skill (skills/preproc-adapt/): an agent-driven loop that diagnoses the offending channel/op and iterates the config for you.

It drives predict's optional preprocess hook, which you can also build directly from a bounded set of ops:

fromdeepcell_typesimportpredict, make_preprocessorconfig= [
{"op": "clip_percentile", "p": 99.9},
{"op": "channel_drop", "names": ["NeuN"]}, # drop a confounding marker
{"op": "min_max_normalize"}, # model sees [0, 1]
]
labels=predict(raw, mask, channel_names, mpp, model_name=...,
device=device, preprocess=make_preprocessor(config))

The hook receives the resampled in-vocabulary (C, H, W) array and must return a (C, H, W) array in [0, 1]. With preprocess=None (default) the built-in p99.9 clip + min-max is used.

Training

Install the [train] extra (adds zarr, pandas, scikit-learn, torchmetrics, plotly, …):

pip install "deepcell-types[train] @ git+https://github.com/vanvalenlab/deepcell-types@master"

Entry points under scripts/:

  • train.py — main training loop (stage 1: backbone, weighted sampler on).
  • retrain_head.py — stage 2: freeze the backbone and retrain the residual-MLP head on the natural class distribution (sampler off). This decoupled recipe is the default and produces the best model; the resMLP head is auto-detected at inference.
  • pretrain.py — masked-marker pretraining.
  • predict.py — batched evaluation over a zarr archive.

Training scripts read config from a TissueNet zarr v3 archive; pass --zarr_dir or set DEEPCELL_TYPES_ZARR_PATH. The deepcell_types.training modules (TissueNetConfig, FullImageDataset, FocalLoss, HierarchicalLoss) can be imported directly for custom scripts.

Baselines

All four paper comparison baselines live in deepcell_types.baselines and run via python -m deepcell_types.baselines <name> (no submodules).

Fairness contract. For an apples-to-apples comparison, run every baseline and the main model with the same class-balancing sampler and model-selection split: --class_balance dct (default) and --val_split_file splits/fov_split_valsubset.json (seed 42), then evaluate on splits/fov_split_test_current.json. These flags are opt-in — parity depends on passing them consistently to every method.

  • XGBoost — on mean-marker-intensity features.
    pip install -e ".[baseline-xgboost]"
    python -m deepcell_types.baselines xgboost ... # or: xgboost-tune
  • Nimbus — UNet marker-positivity (Rumberger et al., Nature Methods 2025). Pins nimbus-inference==0.0.5 (requires Python <3.12).
    pip install -e ".[baseline-nimbus]"
    python -m deepcell_types.baselines nimbus ...
  • MAPS — MLP classifier (Nature Communications 2023).
    pip install -e ".[baseline-maps]"
    python -m deepcell_types.baselines maps ...
  • CellSighter — ResNet-50 multiplexed classifier (Amitay et al., Nature Communications 2023); pulls in torchvision.
    pip install -e ".[baseline-cellsighter]"
    python -m deepcell_types.baselines cellsighter ...

Citation

@article{deepcelltypes,
title={Generalized cell phenotyping for spatial proteomics with language-informed vision models},
author={Wang, Xuefei and Dilip, Rohit and Iqbal, Ahamed Raffey and Bussi, Yuval and Brown, Caitlin and Pradhan, Elora and Jain, Yashvardhan and Yu, Kevin and Li, Shenyi and Abt, Martin and Borner, Katy and Keren, Leeat and Yue, Yisong and Barnowski, Ross and Van Valen, David},
journal={bioRxiv},
pages={2026--07},
year={2026},
publisher={Cold Spring Harbor Laboratory},
url={https://www.biorxiv.org/content/10.1101/2024.11.02.621624v4}
}

About

No description, website, or topics provided.

Resources

Stars

10 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - vanvalenlab/deepcell-types · GitHub
Skip to content

Repository files navigation

DeepCell Types

DOI

DeepCell Types is a generalized cell-phenotyping model for spatial proteomics. It generalizes across datasets with varying marker panels by matching each image's channels against a marker / cell-type registry that ships with the package (vocab.json), so inference runs on any in-memory image with no extra data download.

License notice. Distributed under a Modified Apache License, Version 2.0 with non-commercial / academic-only carve-outs (see the LICENSE file for the full text). For any other use, including commercial use, contact vanvalenlab@gmail.com.

Installation

Install into a virtual environment (venv, conda/mamba, uv, pixi — your choice):

python -m venv dct-env &&source dct-env/bin/activate
pip install git+https://github.com/vanvalenlab/deepcell-types@master

Download the model

Downloading the checkpoint requires a free access token — register at users.deepcell.org and export it (see docs/site/API-key.md):

export DEEPCELL_ACCESS_TOKEN=<your token>
fromdeepcell_types.utilsimportdownload_model# Downloads the latest checkpoint into ~/.deepcell/models and returns its path.model_path=download_model()

Running inference

Inference needs only the checkpoint and your image as an in-memory array — no TissueNet archive required. predict resolves your channels against the packaged vocab.json automatically:

importtorchfromdeepcell_typesimportpredict# Default GPU if available, else CPU (same result, slower); use "cuda:1" etc. for a specific GPU.device="cuda"iftorch.cuda.is_available() else"cpu"# raw: numpy (C, H, W); mask: 2D label image; channel_names: list[str]; mpp: microns/pixel.# model_name accepts the path from download_model() or a path to a .pt file.labels=predict(raw, mask, channel_names, mpp, model_name=model_path, device=device)

See the tutorial for a complete walk-through.

TissueNet zarr archive (optional)

Only needed for training — inference runs entirely from the packaged vocab.json and never requires it. Registered users can download a public archive from https://users.deepcell.org (see docs/site/API-key.md):

export DEEPCELL_TYPES_ZARR_PATH=/absolute/path/to/tissuenet.zarr

Custom preprocessing (advanced)

When a FOV's predictions look implausible — usually a saturated or high-background channel steering the calls — adapt that FOV's per-channel normalization. Start with the preproc-adapt skill (skills/preproc-adapt/): an agent-driven loop that diagnoses the offending channel/op and iterates the config for you.

It drives predict's optional preprocess hook, which you can also build directly from a bounded set of ops:

fromdeepcell_typesimportpredict, make_preprocessorconfig= [
{"op": "clip_percentile", "p": 99.9},
{"op": "channel_drop", "names": ["NeuN"]}, # drop a confounding marker
{"op": "min_max_normalize"}, # model sees [0, 1]
]
labels=predict(raw, mask, channel_names, mpp, model_name=...,
device=device, preprocess=make_preprocessor(config))

The hook receives the resampled in-vocabulary (C, H, W) array and must return a (C, H, W) array in [0, 1]. With preprocess=None (default) the built-in p99.9 clip + min-max is used.

Training

Install the [train] extra (adds zarr, pandas, scikit-learn, torchmetrics, plotly, …):

pip install "deepcell-types[train] @ git+https://github.com/vanvalenlab/deepcell-types@master"

Entry points under scripts/:

  • train.py — main training loop (stage 1: backbone, weighted sampler on).
  • retrain_head.py — stage 2: freeze the backbone and retrain the residual-MLP head on the natural class distribution (sampler off). This decoupled recipe is the default and produces the best model; the resMLP head is auto-detected at inference.
  • pretrain.py — masked-marker pretraining.
  • predict.py — batched evaluation over a zarr archive.

Training scripts read config from a TissueNet zarr v3 archive; pass --zarr_dir or set DEEPCELL_TYPES_ZARR_PATH. The deepcell_types.training modules (TissueNetConfig, FullImageDataset, FocalLoss, HierarchicalLoss) can be imported directly for custom scripts.

Baselines

All four paper comparison baselines live in deepcell_types.baselines and run via python -m deepcell_types.baselines <name> (no submodules).

Fairness contract. For an apples-to-apples comparison, run every baseline and the main model with the same class-balancing sampler and model-selection split: --class_balance dct (default) and --val_split_file splits/fov_split_valsubset.json (seed 42), then evaluate on splits/fov_split_test_current.json. These flags are opt-in — parity depends on passing them consistently to every method.

  • XGBoost — on mean-marker-intensity features.
    pip install -e ".[baseline-xgboost]"
    python -m deepcell_types.baselines xgboost ... # or: xgboost-tune
  • Nimbus — UNet marker-positivity (Rumberger et al., Nature Methods 2025). Pins nimbus-inference==0.0.5 (requires Python <3.12).
    pip install -e ".[baseline-nimbus]"
    python -m deepcell_types.baselines nimbus ...
  • MAPS — MLP classifier (Nature Communications 2023).
    pip install -e ".[baseline-maps]"
    python -m deepcell_types.baselines maps ...
  • CellSighter — ResNet-50 multiplexed classifier (Amitay et al., Nature Communications 2023); pulls in torchvision.
    pip install -e ".[baseline-cellsighter]"
    python -m deepcell_types.baselines cellsighter ...

Citation

@article{deepcelltypes,
title={Generalized cell phenotyping for spatial proteomics with language-informed vision models},
author={Wang, Xuefei and Dilip, Rohit and Iqbal, Ahamed Raffey and Bussi, Yuval and Brown, Caitlin and Pradhan, Elora and Jain, Yashvardhan and Yu, Kevin and Li, Shenyi and Abt, Martin and Borner, Katy and Keren, Leeat and Yue, Yisong and Barnowski, Ross and Van Valen, David},
journal={bioRxiv},
pages={2026--07},
year={2026},
publisher={Cold Spring Harbor Laboratory},
url={https://www.biorxiv.org/content/10.1101/2024.11.02.621624v4}
}

About

No description, website, or topics provided.

Resources

Stars

10 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages