Skip to content

Repository files navigation

CellReasoner: A reasoning-enhanced large language model for cell type annotation 🧬🧠


CellReasoner Overview

📌 Table of Contents


🔬 Key Highlights

  • Only a few expert-level reasoning samples are needed to activate reasoning in a 7B LLM.
  • CellReasoner achieves expert-level interpretability and zero-/few-shot generalization.
  • Demonstrated superior performance across various scRNA-seq and scATAC-seq datasets.
  • Compatible with marker-by-marker annotation, ontology mapping, and biological reasoning.

🧠 Less data, more reasoning: CellReasoner achieves accurate, interpretable, and scalable cell annotation with minimal supervision.


🔑 Key Results

ModelScore
Deepseek-V30.50
Deepseek-R10.53
ChatGPT-o30.58
ChatGPT-4o0.63
singleR0.68
CellReasoner-7B0.73
CellReasoner-32B0.74

ModelScore
Deepseek-V30.52
Deepseek-R10.52
ChatGPT-4o0.76
ChatGPT-o30.85
singleR0.83
CellReasoner-7B0.87
CellReasoner-32B0.84

🧠 Model Zoo

Our CellReasoner models are available on Hugging Face 🤗:

ModelBackboneLink
CellReasoner-7BQwen2.5-7B-Instruct🤗
CellReasoner-32BQwQ-32B🤗

🏋️‍♂️ Training

We use the LLaMA-Factory framework for fine-tuning. It offers a flexible and efficient pipeline for supervised fine-tuning, LoRA, and multi-stage training strategies.


📚 Training Data

We adopt a three-stage training strategy combining reasoning scaffold, biological knowledge infusion, and reasoning mode fusion.

Dataset NameTraining StageSamples
CellCoTReasoning Scaffold, Reasoning Mode Fusion380
pancancer38kKnowledge Infusion37,187
pancancer4kInternal test dataset3,800

You can download the datasets from here.

🚀 Usage

🛠️ Step 1: Prepare Conda Environment

Make sure you have a working conda environment with the necessary dependencies installed. We recommend:

conda create -n cellreasoner python=3.11
conda activate cellreasoner
pip install -r requirements.txt

🧪 Step 2: Preprocess Input Data

If your input is in Seurat .rds format, use the R preprocessing script:

Rscript s01.process_rds.R ./demo_data/pbmc_demo.rds ./output/ data/ranked_hvg.list

If your input is in AnnData .h5ad format, use the Python script:

python s01.process_h5ad.py \
--input_file ./demo_data/pbmc_demo.h5ad \
--output_path ./output_h5ad \
--ranked_hvg_list ./data/ranked_hvg.list

Both pipelines will generate the following output files:

output/
├── pbmc_demo.h5
└── pbmc_demo.meta.csv

🧱 Step 3: Build Dataset for CellReasoner

Build the model input file using:

python s02.build_dataset.py \
--h5_path ./output/pbmc_demo.h5 \
--output_path ./output/ \
--meta_file_path ./output/pbmc_demo.meta.csv

If your metadata includes cell type labels (for scoring), specify the column name:

python s02.build_dataset.py \
--h5_path ./output/pbmc_demo.h5 \
--output_path ./output/ \
--meta_file_path ./output/pbmc_demo.meta.csv \
--cell_type_column "seurat_annotations"

This will generate:

output/
└── pbmc_demo_for_CellReasoner.json

🤖 Step 4: Run Inference with CellReasoner

python s03.inference.py \
--model "CellReasoner-7B" \
--output_path "./output" \
--input_json "./output/pbmc_demo_for_CellReasoner.json" \
--batch_size 2

Result:

output/
└── pbmc_demo_CellReasoner_result.csv

📊 Evaluation and Reasoning Visualization

To compute scores, generate plots, or view reasoning outputs, refer to:

s03.inference.ipynb

Citation

@article {Cao2025.05.20.655112,
author = {Cao, Guangshuo and Shen, Yi and Wu, Jianghong and Chao, Haoyu and Chen, Ming and Chen, Dijun},
title = {CellReasoner: A reasoning-enhanced large language model for cell type annotation},
elocation-id = {2025.05.20.655112},
year = {2025},
doi = {10.1101/2025.05.20.655112},
URL = {https://www.biorxiv.org/content/early/2025/05/26/2025.05.20.655112},
eprint = {https://www.biorxiv.org/content/early/2025/05/26/2025.05.20.655112.full.pdf},
journal = {bioRxiv}
}

About

CellReasoner: A reasoning-enhanced large language model for cell type annotation

Resources

Stars

10 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - compbioNJU/CellReasoner: CellReasoner: A reasoning-enhanced large language model for cell type annotation · GitHub
Skip to content

Repository files navigation

CellReasoner: A reasoning-enhanced large language model for cell type annotation 🧬🧠


CellReasoner Overview

📌 Table of Contents


🔬 Key Highlights

  • Only a few expert-level reasoning samples are needed to activate reasoning in a 7B LLM.
  • CellReasoner achieves expert-level interpretability and zero-/few-shot generalization.
  • Demonstrated superior performance across various scRNA-seq and scATAC-seq datasets.
  • Compatible with marker-by-marker annotation, ontology mapping, and biological reasoning.

🧠 Less data, more reasoning: CellReasoner achieves accurate, interpretable, and scalable cell annotation with minimal supervision.


🔑 Key Results

ModelScore
Deepseek-V30.50
Deepseek-R10.53
ChatGPT-o30.58
ChatGPT-4o0.63
singleR0.68
CellReasoner-7B0.73
CellReasoner-32B0.74

ModelScore
Deepseek-V30.52
Deepseek-R10.52
ChatGPT-4o0.76
ChatGPT-o30.85
singleR0.83
CellReasoner-7B0.87
CellReasoner-32B0.84

🧠 Model Zoo

Our CellReasoner models are available on Hugging Face 🤗:

ModelBackboneLink
CellReasoner-7BQwen2.5-7B-Instruct🤗
CellReasoner-32BQwQ-32B🤗

🏋️‍♂️ Training

We use the LLaMA-Factory framework for fine-tuning. It offers a flexible and efficient pipeline for supervised fine-tuning, LoRA, and multi-stage training strategies.


📚 Training Data

We adopt a three-stage training strategy combining reasoning scaffold, biological knowledge infusion, and reasoning mode fusion.

Dataset NameTraining StageSamples
CellCoTReasoning Scaffold, Reasoning Mode Fusion380
pancancer38kKnowledge Infusion37,187
pancancer4kInternal test dataset3,800

You can download the datasets from here.

🚀 Usage

🛠️ Step 1: Prepare Conda Environment

Make sure you have a working conda environment with the necessary dependencies installed. We recommend:

conda create -n cellreasoner python=3.11
conda activate cellreasoner
pip install -r requirements.txt

🧪 Step 2: Preprocess Input Data

If your input is in Seurat .rds format, use the R preprocessing script:

Rscript s01.process_rds.R ./demo_data/pbmc_demo.rds ./output/ data/ranked_hvg.list

If your input is in AnnData .h5ad format, use the Python script:

python s01.process_h5ad.py \
--input_file ./demo_data/pbmc_demo.h5ad \
--output_path ./output_h5ad \
--ranked_hvg_list ./data/ranked_hvg.list

Both pipelines will generate the following output files:

output/
├── pbmc_demo.h5
└── pbmc_demo.meta.csv

🧱 Step 3: Build Dataset for CellReasoner

Build the model input file using:

python s02.build_dataset.py \
--h5_path ./output/pbmc_demo.h5 \
--output_path ./output/ \
--meta_file_path ./output/pbmc_demo.meta.csv

If your metadata includes cell type labels (for scoring), specify the column name:

python s02.build_dataset.py \
--h5_path ./output/pbmc_demo.h5 \
--output_path ./output/ \
--meta_file_path ./output/pbmc_demo.meta.csv \
--cell_type_column "seurat_annotations"

This will generate:

output/
└── pbmc_demo_for_CellReasoner.json

🤖 Step 4: Run Inference with CellReasoner

python s03.inference.py \
--model "CellReasoner-7B" \
--output_path "./output" \
--input_json "./output/pbmc_demo_for_CellReasoner.json" \
--batch_size 2

Result:

output/
└── pbmc_demo_CellReasoner_result.csv

📊 Evaluation and Reasoning Visualization

To compute scores, generate plots, or view reasoning outputs, refer to:

s03.inference.ipynb

Citation

@article {Cao2025.05.20.655112,
author = {Cao, Guangshuo and Shen, Yi and Wu, Jianghong and Chao, Haoyu and Chen, Ming and Chen, Dijun},
title = {CellReasoner: A reasoning-enhanced large language model for cell type annotation},
elocation-id = {2025.05.20.655112},
year = {2025},
doi = {10.1101/2025.05.20.655112},
URL = {https://www.biorxiv.org/content/early/2025/05/26/2025.05.20.655112},
eprint = {https://www.biorxiv.org/content/early/2025/05/26/2025.05.20.655112.full.pdf},
journal = {bioRxiv}
}

About

CellReasoner: A reasoning-enhanced large language model for cell type annotation

Resources

Stars

10 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - compbioNJU/CellReasoner: CellReasoner: A reasoning-enhanced large language model for cell type annotation · GitHub
Skip to content

Repository files navigation

CellReasoner: A reasoning-enhanced large language model for cell type annotation 🧬🧠


CellReasoner Overview

📌 Table of Contents


🔬 Key Highlights

  • Only a few expert-level reasoning samples are needed to activate reasoning in a 7B LLM.
  • CellReasoner achieves expert-level interpretability and zero-/few-shot generalization.
  • Demonstrated superior performance across various scRNA-seq and scATAC-seq datasets.
  • Compatible with marker-by-marker annotation, ontology mapping, and biological reasoning.

🧠 Less data, more reasoning: CellReasoner achieves accurate, interpretable, and scalable cell annotation with minimal supervision.


🔑 Key Results

ModelScore
Deepseek-V30.50
Deepseek-R10.53
ChatGPT-o30.58
ChatGPT-4o0.63
singleR0.68
CellReasoner-7B0.73
CellReasoner-32B0.74

ModelScore
Deepseek-V30.52
Deepseek-R10.52
ChatGPT-4o0.76
ChatGPT-o30.85
singleR0.83
CellReasoner-7B0.87
CellReasoner-32B0.84

🧠 Model Zoo

Our CellReasoner models are available on Hugging Face 🤗:

ModelBackboneLink
CellReasoner-7BQwen2.5-7B-Instruct🤗
CellReasoner-32BQwQ-32B🤗

🏋️‍♂️ Training

We use the LLaMA-Factory framework for fine-tuning. It offers a flexible and efficient pipeline for supervised fine-tuning, LoRA, and multi-stage training strategies.


📚 Training Data

We adopt a three-stage training strategy combining reasoning scaffold, biological knowledge infusion, and reasoning mode fusion.

Dataset NameTraining StageSamples
CellCoTReasoning Scaffold, Reasoning Mode Fusion380
pancancer38kKnowledge Infusion37,187
pancancer4kInternal test dataset3,800

You can download the datasets from here.

🚀 Usage

🛠️ Step 1: Prepare Conda Environment

Make sure you have a working conda environment with the necessary dependencies installed. We recommend:

conda create -n cellreasoner python=3.11
conda activate cellreasoner
pip install -r requirements.txt

🧪 Step 2: Preprocess Input Data

If your input is in Seurat .rds format, use the R preprocessing script:

Rscript s01.process_rds.R ./demo_data/pbmc_demo.rds ./output/ data/ranked_hvg.list

If your input is in AnnData .h5ad format, use the Python script:

python s01.process_h5ad.py \
--input_file ./demo_data/pbmc_demo.h5ad \
--output_path ./output_h5ad \
--ranked_hvg_list ./data/ranked_hvg.list

Both pipelines will generate the following output files:

output/
├── pbmc_demo.h5
└── pbmc_demo.meta.csv

🧱 Step 3: Build Dataset for CellReasoner

Build the model input file using:

python s02.build_dataset.py \
--h5_path ./output/pbmc_demo.h5 \
--output_path ./output/ \
--meta_file_path ./output/pbmc_demo.meta.csv

If your metadata includes cell type labels (for scoring), specify the column name:

python s02.build_dataset.py \
--h5_path ./output/pbmc_demo.h5 \
--output_path ./output/ \
--meta_file_path ./output/pbmc_demo.meta.csv \
--cell_type_column "seurat_annotations"

This will generate:

output/
└── pbmc_demo_for_CellReasoner.json

🤖 Step 4: Run Inference with CellReasoner

python s03.inference.py \
--model "CellReasoner-7B" \
--output_path "./output" \
--input_json "./output/pbmc_demo_for_CellReasoner.json" \
--batch_size 2

Result:

output/
└── pbmc_demo_CellReasoner_result.csv

📊 Evaluation and Reasoning Visualization

To compute scores, generate plots, or view reasoning outputs, refer to:

s03.inference.ipynb

Citation

@article {Cao2025.05.20.655112,
author = {Cao, Guangshuo and Shen, Yi and Wu, Jianghong and Chao, Haoyu and Chen, Ming and Chen, Dijun},
title = {CellReasoner: A reasoning-enhanced large language model for cell type annotation},
elocation-id = {2025.05.20.655112},
year = {2025},
doi = {10.1101/2025.05.20.655112},
URL = {https://www.biorxiv.org/content/early/2025/05/26/2025.05.20.655112},
eprint = {https://www.biorxiv.org/content/early/2025/05/26/2025.05.20.655112.full.pdf},
journal = {bioRxiv}
}

About

CellReasoner: A reasoning-enhanced large language model for cell type annotation

Resources

Stars

10 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - compbioNJU/CellReasoner: CellReasoner: A reasoning-enhanced large language model for cell type annotation · GitHub
Skip to content

Repository files navigation

CellReasoner: A reasoning-enhanced large language model for cell type annotation 🧬🧠


CellReasoner Overview

📌 Table of Contents


🔬 Key Highlights

  • Only a few expert-level reasoning samples are needed to activate reasoning in a 7B LLM.
  • CellReasoner achieves expert-level interpretability and zero-/few-shot generalization.
  • Demonstrated superior performance across various scRNA-seq and scATAC-seq datasets.
  • Compatible with marker-by-marker annotation, ontology mapping, and biological reasoning.

🧠 Less data, more reasoning: CellReasoner achieves accurate, interpretable, and scalable cell annotation with minimal supervision.


🔑 Key Results

ModelScore
Deepseek-V30.50
Deepseek-R10.53
ChatGPT-o30.58
ChatGPT-4o0.63
singleR0.68
CellReasoner-7B0.73
CellReasoner-32B0.74

ModelScore
Deepseek-V30.52
Deepseek-R10.52
ChatGPT-4o0.76
ChatGPT-o30.85
singleR0.83
CellReasoner-7B0.87
CellReasoner-32B0.84

🧠 Model Zoo

Our CellReasoner models are available on Hugging Face 🤗:

ModelBackboneLink
CellReasoner-7BQwen2.5-7B-Instruct🤗
CellReasoner-32BQwQ-32B🤗

🏋️‍♂️ Training

We use the LLaMA-Factory framework for fine-tuning. It offers a flexible and efficient pipeline for supervised fine-tuning, LoRA, and multi-stage training strategies.


📚 Training Data

We adopt a three-stage training strategy combining reasoning scaffold, biological knowledge infusion, and reasoning mode fusion.

Dataset NameTraining StageSamples
CellCoTReasoning Scaffold, Reasoning Mode Fusion380
pancancer38kKnowledge Infusion37,187
pancancer4kInternal test dataset3,800

You can download the datasets from here.

🚀 Usage

🛠️ Step 1: Prepare Conda Environment

Make sure you have a working conda environment with the necessary dependencies installed. We recommend:

conda create -n cellreasoner python=3.11
conda activate cellreasoner
pip install -r requirements.txt

🧪 Step 2: Preprocess Input Data

If your input is in Seurat .rds format, use the R preprocessing script:

Rscript s01.process_rds.R ./demo_data/pbmc_demo.rds ./output/ data/ranked_hvg.list

If your input is in AnnData .h5ad format, use the Python script:

python s01.process_h5ad.py \
--input_file ./demo_data/pbmc_demo.h5ad \
--output_path ./output_h5ad \
--ranked_hvg_list ./data/ranked_hvg.list

Both pipelines will generate the following output files:

output/
├── pbmc_demo.h5
└── pbmc_demo.meta.csv

🧱 Step 3: Build Dataset for CellReasoner

Build the model input file using:

python s02.build_dataset.py \
--h5_path ./output/pbmc_demo.h5 \
--output_path ./output/ \
--meta_file_path ./output/pbmc_demo.meta.csv

If your metadata includes cell type labels (for scoring), specify the column name:

python s02.build_dataset.py \
--h5_path ./output/pbmc_demo.h5 \
--output_path ./output/ \
--meta_file_path ./output/pbmc_demo.meta.csv \
--cell_type_column "seurat_annotations"

This will generate:

output/
└── pbmc_demo_for_CellReasoner.json

🤖 Step 4: Run Inference with CellReasoner

python s03.inference.py \
--model "CellReasoner-7B" \
--output_path "./output" \
--input_json "./output/pbmc_demo_for_CellReasoner.json" \
--batch_size 2

Result:

output/
└── pbmc_demo_CellReasoner_result.csv

📊 Evaluation and Reasoning Visualization

To compute scores, generate plots, or view reasoning outputs, refer to:

s03.inference.ipynb

Citation

@article {Cao2025.05.20.655112,
author = {Cao, Guangshuo and Shen, Yi and Wu, Jianghong and Chao, Haoyu and Chen, Ming and Chen, Dijun},
title = {CellReasoner: A reasoning-enhanced large language model for cell type annotation},
elocation-id = {2025.05.20.655112},
year = {2025},
doi = {10.1101/2025.05.20.655112},
URL = {https://www.biorxiv.org/content/early/2025/05/26/2025.05.20.655112},
eprint = {https://www.biorxiv.org/content/early/2025/05/26/2025.05.20.655112.full.pdf},
journal = {bioRxiv}
}

About

CellReasoner: A reasoning-enhanced large language model for cell type annotation

Resources

Stars

10 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - compbioNJU/CellReasoner: CellReasoner: A reasoning-enhanced large language model for cell type annotation · GitHub
Skip to content

Repository files navigation

CellReasoner: A reasoning-enhanced large language model for cell type annotation 🧬🧠


CellReasoner Overview

📌 Table of Contents


🔬 Key Highlights

  • Only a few expert-level reasoning samples are needed to activate reasoning in a 7B LLM.
  • CellReasoner achieves expert-level interpretability and zero-/few-shot generalization.
  • Demonstrated superior performance across various scRNA-seq and scATAC-seq datasets.
  • Compatible with marker-by-marker annotation, ontology mapping, and biological reasoning.

🧠 Less data, more reasoning: CellReasoner achieves accurate, interpretable, and scalable cell annotation with minimal supervision.


🔑 Key Results

ModelScore
Deepseek-V30.50
Deepseek-R10.53
ChatGPT-o30.58
ChatGPT-4o0.63
singleR0.68
CellReasoner-7B0.73
CellReasoner-32B0.74

ModelScore
Deepseek-V30.52
Deepseek-R10.52
ChatGPT-4o0.76
ChatGPT-o30.85
singleR0.83
CellReasoner-7B0.87
CellReasoner-32B0.84

🧠 Model Zoo

Our CellReasoner models are available on Hugging Face 🤗:

ModelBackboneLink
CellReasoner-7BQwen2.5-7B-Instruct🤗
CellReasoner-32BQwQ-32B🤗

🏋️‍♂️ Training

We use the LLaMA-Factory framework for fine-tuning. It offers a flexible and efficient pipeline for supervised fine-tuning, LoRA, and multi-stage training strategies.


📚 Training Data

We adopt a three-stage training strategy combining reasoning scaffold, biological knowledge infusion, and reasoning mode fusion.

Dataset NameTraining StageSamples
CellCoTReasoning Scaffold, Reasoning Mode Fusion380
pancancer38kKnowledge Infusion37,187
pancancer4kInternal test dataset3,800

You can download the datasets from here.

🚀 Usage

🛠️ Step 1: Prepare Conda Environment

Make sure you have a working conda environment with the necessary dependencies installed. We recommend:

conda create -n cellreasoner python=3.11
conda activate cellreasoner
pip install -r requirements.txt

🧪 Step 2: Preprocess Input Data

If your input is in Seurat .rds format, use the R preprocessing script:

Rscript s01.process_rds.R ./demo_data/pbmc_demo.rds ./output/ data/ranked_hvg.list

If your input is in AnnData .h5ad format, use the Python script:

python s01.process_h5ad.py \
--input_file ./demo_data/pbmc_demo.h5ad \
--output_path ./output_h5ad \
--ranked_hvg_list ./data/ranked_hvg.list

Both pipelines will generate the following output files:

output/
├── pbmc_demo.h5
└── pbmc_demo.meta.csv

🧱 Step 3: Build Dataset for CellReasoner

Build the model input file using:

python s02.build_dataset.py \
--h5_path ./output/pbmc_demo.h5 \
--output_path ./output/ \
--meta_file_path ./output/pbmc_demo.meta.csv

If your metadata includes cell type labels (for scoring), specify the column name:

python s02.build_dataset.py \
--h5_path ./output/pbmc_demo.h5 \
--output_path ./output/ \
--meta_file_path ./output/pbmc_demo.meta.csv \
--cell_type_column "seurat_annotations"

This will generate:

output/
└── pbmc_demo_for_CellReasoner.json

🤖 Step 4: Run Inference with CellReasoner

python s03.inference.py \
--model "CellReasoner-7B" \
--output_path "./output" \
--input_json "./output/pbmc_demo_for_CellReasoner.json" \
--batch_size 2

Result:

output/
└── pbmc_demo_CellReasoner_result.csv

📊 Evaluation and Reasoning Visualization

To compute scores, generate plots, or view reasoning outputs, refer to:

s03.inference.ipynb

Citation

@article {Cao2025.05.20.655112,
author = {Cao, Guangshuo and Shen, Yi and Wu, Jianghong and Chao, Haoyu and Chen, Ming and Chen, Dijun},
title = {CellReasoner: A reasoning-enhanced large language model for cell type annotation},
elocation-id = {2025.05.20.655112},
year = {2025},
doi = {10.1101/2025.05.20.655112},
URL = {https://www.biorxiv.org/content/early/2025/05/26/2025.05.20.655112},
eprint = {https://www.biorxiv.org/content/early/2025/05/26/2025.05.20.655112.full.pdf},
journal = {bioRxiv}
}

About

CellReasoner: A reasoning-enhanced large language model for cell type annotation

Resources

Stars

10 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - compbioNJU/CellReasoner: CellReasoner: A reasoning-enhanced large language model for cell type annotation · GitHub
Skip to content

Repository files navigation

CellReasoner: A reasoning-enhanced large language model for cell type annotation 🧬🧠


CellReasoner Overview

📌 Table of Contents


🔬 Key Highlights

  • Only a few expert-level reasoning samples are needed to activate reasoning in a 7B LLM.
  • CellReasoner achieves expert-level interpretability and zero-/few-shot generalization.
  • Demonstrated superior performance across various scRNA-seq and scATAC-seq datasets.
  • Compatible with marker-by-marker annotation, ontology mapping, and biological reasoning.

🧠 Less data, more reasoning: CellReasoner achieves accurate, interpretable, and scalable cell annotation with minimal supervision.


🔑 Key Results

ModelScore
Deepseek-V30.50
Deepseek-R10.53
ChatGPT-o30.58
ChatGPT-4o0.63
singleR0.68
CellReasoner-7B0.73
CellReasoner-32B0.74

ModelScore
Deepseek-V30.52
Deepseek-R10.52
ChatGPT-4o0.76
ChatGPT-o30.85
singleR0.83
CellReasoner-7B0.87
CellReasoner-32B0.84

🧠 Model Zoo

Our CellReasoner models are available on Hugging Face 🤗:

ModelBackboneLink
CellReasoner-7BQwen2.5-7B-Instruct🤗
CellReasoner-32BQwQ-32B🤗

🏋️‍♂️ Training

We use the LLaMA-Factory framework for fine-tuning. It offers a flexible and efficient pipeline for supervised fine-tuning, LoRA, and multi-stage training strategies.


📚 Training Data

We adopt a three-stage training strategy combining reasoning scaffold, biological knowledge infusion, and reasoning mode fusion.

Dataset NameTraining StageSamples
CellCoTReasoning Scaffold, Reasoning Mode Fusion380
pancancer38kKnowledge Infusion37,187
pancancer4kInternal test dataset3,800

You can download the datasets from here.

🚀 Usage

🛠️ Step 1: Prepare Conda Environment

Make sure you have a working conda environment with the necessary dependencies installed. We recommend:

conda create -n cellreasoner python=3.11
conda activate cellreasoner
pip install -r requirements.txt

🧪 Step 2: Preprocess Input Data

If your input is in Seurat .rds format, use the R preprocessing script:

Rscript s01.process_rds.R ./demo_data/pbmc_demo.rds ./output/ data/ranked_hvg.list

If your input is in AnnData .h5ad format, use the Python script:

python s01.process_h5ad.py \
--input_file ./demo_data/pbmc_demo.h5ad \
--output_path ./output_h5ad \
--ranked_hvg_list ./data/ranked_hvg.list

Both pipelines will generate the following output files:

output/
├── pbmc_demo.h5
└── pbmc_demo.meta.csv

🧱 Step 3: Build Dataset for CellReasoner

Build the model input file using:

python s02.build_dataset.py \
--h5_path ./output/pbmc_demo.h5 \
--output_path ./output/ \
--meta_file_path ./output/pbmc_demo.meta.csv

If your metadata includes cell type labels (for scoring), specify the column name:

python s02.build_dataset.py \
--h5_path ./output/pbmc_demo.h5 \
--output_path ./output/ \
--meta_file_path ./output/pbmc_demo.meta.csv \
--cell_type_column "seurat_annotations"

This will generate:

output/
└── pbmc_demo_for_CellReasoner.json

🤖 Step 4: Run Inference with CellReasoner

python s03.inference.py \
--model "CellReasoner-7B" \
--output_path "./output" \
--input_json "./output/pbmc_demo_for_CellReasoner.json" \
--batch_size 2

Result:

output/
└── pbmc_demo_CellReasoner_result.csv

📊 Evaluation and Reasoning Visualization

To compute scores, generate plots, or view reasoning outputs, refer to:

s03.inference.ipynb

Citation

@article {Cao2025.05.20.655112,
author = {Cao, Guangshuo and Shen, Yi and Wu, Jianghong and Chao, Haoyu and Chen, Ming and Chen, Dijun},
title = {CellReasoner: A reasoning-enhanced large language model for cell type annotation},
elocation-id = {2025.05.20.655112},
year = {2025},
doi = {10.1101/2025.05.20.655112},
URL = {https://www.biorxiv.org/content/early/2025/05/26/2025.05.20.655112},
eprint = {https://www.biorxiv.org/content/early/2025/05/26/2025.05.20.655112.full.pdf},
journal = {bioRxiv}
}

About

CellReasoner: A reasoning-enhanced large language model for cell type annotation

Resources

Stars

10 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - compbioNJU/CellReasoner: CellReasoner: A reasoning-enhanced large language model for cell type annotation · GitHub
Skip to content

Repository files navigation

CellReasoner: A reasoning-enhanced large language model for cell type annotation 🧬🧠


CellReasoner Overview

📌 Table of Contents


🔬 Key Highlights

  • Only a few expert-level reasoning samples are needed to activate reasoning in a 7B LLM.
  • CellReasoner achieves expert-level interpretability and zero-/few-shot generalization.
  • Demonstrated superior performance across various scRNA-seq and scATAC-seq datasets.
  • Compatible with marker-by-marker annotation, ontology mapping, and biological reasoning.

🧠 Less data, more reasoning: CellReasoner achieves accurate, interpretable, and scalable cell annotation with minimal supervision.


🔑 Key Results

ModelScore
Deepseek-V30.50
Deepseek-R10.53
ChatGPT-o30.58
ChatGPT-4o0.63
singleR0.68
CellReasoner-7B0.73
CellReasoner-32B0.74

ModelScore
Deepseek-V30.52
Deepseek-R10.52
ChatGPT-4o0.76
ChatGPT-o30.85
singleR0.83
CellReasoner-7B0.87
CellReasoner-32B0.84

🧠 Model Zoo

Our CellReasoner models are available on Hugging Face 🤗:

ModelBackboneLink
CellReasoner-7BQwen2.5-7B-Instruct🤗
CellReasoner-32BQwQ-32B🤗

🏋️‍♂️ Training

We use the LLaMA-Factory framework for fine-tuning. It offers a flexible and efficient pipeline for supervised fine-tuning, LoRA, and multi-stage training strategies.


📚 Training Data

We adopt a three-stage training strategy combining reasoning scaffold, biological knowledge infusion, and reasoning mode fusion.

Dataset NameTraining StageSamples
CellCoTReasoning Scaffold, Reasoning Mode Fusion380
pancancer38kKnowledge Infusion37,187
pancancer4kInternal test dataset3,800

You can download the datasets from here.

🚀 Usage

🛠️ Step 1: Prepare Conda Environment

Make sure you have a working conda environment with the necessary dependencies installed. We recommend:

conda create -n cellreasoner python=3.11
conda activate cellreasoner
pip install -r requirements.txt

🧪 Step 2: Preprocess Input Data

If your input is in Seurat .rds format, use the R preprocessing script:

Rscript s01.process_rds.R ./demo_data/pbmc_demo.rds ./output/ data/ranked_hvg.list

If your input is in AnnData .h5ad format, use the Python script:

python s01.process_h5ad.py \
--input_file ./demo_data/pbmc_demo.h5ad \
--output_path ./output_h5ad \
--ranked_hvg_list ./data/ranked_hvg.list

Both pipelines will generate the following output files:

output/
├── pbmc_demo.h5
└── pbmc_demo.meta.csv

🧱 Step 3: Build Dataset for CellReasoner

Build the model input file using:

python s02.build_dataset.py \
--h5_path ./output/pbmc_demo.h5 \
--output_path ./output/ \
--meta_file_path ./output/pbmc_demo.meta.csv

If your metadata includes cell type labels (for scoring), specify the column name:

python s02.build_dataset.py \
--h5_path ./output/pbmc_demo.h5 \
--output_path ./output/ \
--meta_file_path ./output/pbmc_demo.meta.csv \
--cell_type_column "seurat_annotations"

This will generate:

output/
└── pbmc_demo_for_CellReasoner.json

🤖 Step 4: Run Inference with CellReasoner

python s03.inference.py \
--model "CellReasoner-7B" \
--output_path "./output" \
--input_json "./output/pbmc_demo_for_CellReasoner.json" \
--batch_size 2

Result:

output/
└── pbmc_demo_CellReasoner_result.csv

📊 Evaluation and Reasoning Visualization

To compute scores, generate plots, or view reasoning outputs, refer to:

s03.inference.ipynb

Citation

@article {Cao2025.05.20.655112,
author = {Cao, Guangshuo and Shen, Yi and Wu, Jianghong and Chao, Haoyu and Chen, Ming and Chen, Dijun},
title = {CellReasoner: A reasoning-enhanced large language model for cell type annotation},
elocation-id = {2025.05.20.655112},
year = {2025},
doi = {10.1101/2025.05.20.655112},
URL = {https://www.biorxiv.org/content/early/2025/05/26/2025.05.20.655112},
eprint = {https://www.biorxiv.org/content/early/2025/05/26/2025.05.20.655112.full.pdf},
journal = {bioRxiv}
}

About

CellReasoner: A reasoning-enhanced large language model for cell type annotation

Resources

Stars

10 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - compbioNJU/CellReasoner: CellReasoner: A reasoning-enhanced large language model for cell type annotation · GitHub
Skip to content

Repository files navigation

CellReasoner: A reasoning-enhanced large language model for cell type annotation 🧬🧠


CellReasoner Overview

📌 Table of Contents


🔬 Key Highlights

  • Only a few expert-level reasoning samples are needed to activate reasoning in a 7B LLM.
  • CellReasoner achieves expert-level interpretability and zero-/few-shot generalization.
  • Demonstrated superior performance across various scRNA-seq and scATAC-seq datasets.
  • Compatible with marker-by-marker annotation, ontology mapping, and biological reasoning.

🧠 Less data, more reasoning: CellReasoner achieves accurate, interpretable, and scalable cell annotation with minimal supervision.


🔑 Key Results

ModelScore
Deepseek-V30.50
Deepseek-R10.53
ChatGPT-o30.58
ChatGPT-4o0.63
singleR0.68
CellReasoner-7B0.73
CellReasoner-32B0.74

ModelScore
Deepseek-V30.52
Deepseek-R10.52
ChatGPT-4o0.76
ChatGPT-o30.85
singleR0.83
CellReasoner-7B0.87
CellReasoner-32B0.84

🧠 Model Zoo

Our CellReasoner models are available on Hugging Face 🤗:

ModelBackboneLink
CellReasoner-7BQwen2.5-7B-Instruct🤗
CellReasoner-32BQwQ-32B🤗

🏋️‍♂️ Training

We use the LLaMA-Factory framework for fine-tuning. It offers a flexible and efficient pipeline for supervised fine-tuning, LoRA, and multi-stage training strategies.


📚 Training Data

We adopt a three-stage training strategy combining reasoning scaffold, biological knowledge infusion, and reasoning mode fusion.

Dataset NameTraining StageSamples
CellCoTReasoning Scaffold, Reasoning Mode Fusion380
pancancer38kKnowledge Infusion37,187
pancancer4kInternal test dataset3,800

You can download the datasets from here.

🚀 Usage

🛠️ Step 1: Prepare Conda Environment

Make sure you have a working conda environment with the necessary dependencies installed. We recommend:

conda create -n cellreasoner python=3.11
conda activate cellreasoner
pip install -r requirements.txt

🧪 Step 2: Preprocess Input Data

If your input is in Seurat .rds format, use the R preprocessing script:

Rscript s01.process_rds.R ./demo_data/pbmc_demo.rds ./output/ data/ranked_hvg.list

If your input is in AnnData .h5ad format, use the Python script:

python s01.process_h5ad.py \
--input_file ./demo_data/pbmc_demo.h5ad \
--output_path ./output_h5ad \
--ranked_hvg_list ./data/ranked_hvg.list

Both pipelines will generate the following output files:

output/
├── pbmc_demo.h5
└── pbmc_demo.meta.csv

🧱 Step 3: Build Dataset for CellReasoner

Build the model input file using:

python s02.build_dataset.py \
--h5_path ./output/pbmc_demo.h5 \
--output_path ./output/ \
--meta_file_path ./output/pbmc_demo.meta.csv

If your metadata includes cell type labels (for scoring), specify the column name:

python s02.build_dataset.py \
--h5_path ./output/pbmc_demo.h5 \
--output_path ./output/ \
--meta_file_path ./output/pbmc_demo.meta.csv \
--cell_type_column "seurat_annotations"

This will generate:

output/
└── pbmc_demo_for_CellReasoner.json

🤖 Step 4: Run Inference with CellReasoner

python s03.inference.py \
--model "CellReasoner-7B" \
--output_path "./output" \
--input_json "./output/pbmc_demo_for_CellReasoner.json" \
--batch_size 2

Result:

output/
└── pbmc_demo_CellReasoner_result.csv

📊 Evaluation and Reasoning Visualization

To compute scores, generate plots, or view reasoning outputs, refer to:

s03.inference.ipynb

Citation

@article {Cao2025.05.20.655112,
author = {Cao, Guangshuo and Shen, Yi and Wu, Jianghong and Chao, Haoyu and Chen, Ming and Chen, Dijun},
title = {CellReasoner: A reasoning-enhanced large language model for cell type annotation},
elocation-id = {2025.05.20.655112},
year = {2025},
doi = {10.1101/2025.05.20.655112},
URL = {https://www.biorxiv.org/content/early/2025/05/26/2025.05.20.655112},
eprint = {https://www.biorxiv.org/content/early/2025/05/26/2025.05.20.655112.full.pdf},
journal = {bioRxiv}
}

About

CellReasoner: A reasoning-enhanced large language model for cell type annotation

Resources

Stars

10 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages