Skip to content

Repository files navigation

RAGulateRAGulate logo

Retrieval-Augmented Generation for Post-hoc Literature-Grounded Assessment of Gene Regulatory Interaction

GitHub issuesPyPI - ProjectDocsDOI

Introduction

RAGulate is a Retrieval-Augmented Generation (RAG) pipeline that integrates domain-specific large language models (LLM) with curated literature to identify, score, and assess inferred regulatory (transcription factor-target gene) interactions in their biological context.

RAGulate pipeline

For further information and example tutorials, please check our documentation:

If you have any questions or concerns, feel free to open an issue.

Requirements

RAGulate is implemented in the LlamaIndex framework. Running RAGulate on CUDA is highly recommended if available.

Before installing and running RAGulate, ensure you have the following libraries installed:

  • PyTorch (version 2.0 or higher)
    Install with the exact command from the PyTorch “Get Started” page for your OS, Python version and (optionally) CUDA toolkit.
  • NumPy (version 1.23 or higher)
  • bm-25 (version 0.2.2 or higher)

You can install these dependencies using pip:

pip install torch numpy

Installation

Option 1:
You can install RAGulate via pip for a lightweight installation:

pip install ragulate-bio

Option 2:
Alternatively, if you want the latest, unreleased version, you can install it directly from the source on GitHub:

pip install git+https://github.com/YDaiLab/RAGulate.git

Import

importragulate_bioasragulate# recommended alias

Note: The PyPI distribution is named ragulate-bio to avoid a name conflict with an unrelated project called ragulate. Always import ragulate_bio in Python (you may alias it to ragulate for convenience)

Option 3 (Coming soon):
For users who prefer Conda or Mamba for environment management, you can install RAGulate along with extra dependencies:

Conda:

conda install -c zandigohar RAGulate

Mamba:

mamba create -n RAGulate -c zandigohar RAGulate

Tutorials

Run the tutorial notebook:

Reproducibility

Large reproducibility resources for RAGulate (including prior regulatory databases, serialized artifacts, and precomputed resources) are hosted separately on Zenodo due to size constraints:

https://doi.org/10.5281/zenodo.18498382

To reproduce the tutorial and experimental results, download the Zenodo dataset and place the files in the expected locations (see config.py), then run the notebooks in docs/ in order:

  1. NCBI-contextualization_Final.ipynb (generates the context-specific pickle)
  2. RAGulate_Modular_Rep.ipynb (full reproducibility run)

FAQ

Q1: Do I need a GPU to run RAGulate?
No, a GPU is not required. However, using a CUDA-enabled GPU is strongly recommended for faster runs, especially with large queries.

Q2: How do I know if I can use a GPU with RAGulate?
There are two quick checks:

  1. System check
    In your terminal, run nvidia-smi. If you see your GPU listed (model, memory, driver version), your machine has an NVIDIA GPU with the driver installed.

  2. Python check
    In a Python shell, run:

    importtorchprint(torch.cuda.is_available()) # True means PyTorch can see your GPUprint(torch.cuda.device_count()) # How many GPUs are usable

Q3: Can I use RAGulate with R-based tools?
RAGulate is written in Python and works directly with Numpy objects.

Q4: What if I also have another package called ragulate installed? RAGulate will warn you if it detects a conflicting installation. We recommend using a clean virtual environment to avoid import clashes.

Q5: How do I cite RAGulate?
See the Citation section below for the latest reference and preprint link.

Q6: How can I reproduce the paper’s results?
See our Reproducibility Guide for step-by-step instructions. Then run RAGulate.

Citation

This repository is under active development. Please cite as:
Zandigohar M. and Dai Y. RAGulate: RAGulate: Retrieval-Augmented Generation for Post-hoc Literature-Grounded Regulatory Assessment. 2026. https://doi.org/10.64898/2026.01.20.700704

Development & Contact

RAGulate was developed and is actively maintained by Mehrdad Zandigohar as part of his PhD research at the University of Illinois Chicago (UIC), in the lab of Dr. Yang Dai.

📬 For private questions, please email: mzandi2@uic.edu

🤝 For collaboration inquiries, please contact PI: Dr. Yang Dai (yangdai@uic.edu)

Contributions, feature suggestions, and feedback are always welcome!

License

The code in RAGulate is licensed under the MIT License, which permits academic and commercial use, modification, and distribution.

Please note that any third-party dependencies bundled with RAGulate may have their own respective licenses.

About

RAGulate is a retrieval-augmented generation (RAG) pipeline that integrates domain-specific language models with curated literature to identify, score, and validate inferred transcription factor-target gene interactions in their biological context.

Resources

Code of conduct

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - YDaiLab/RAGulate: RAGulate is a retrieval-augmented generation (RAG) pipeline that integrates domain-specific language models with curated literature to identify, score, and validate inferred transcription factor-target gene interactions in their biological context. · GitHub
Skip to content

Repository files navigation

RAGulateRAGulate logo

Retrieval-Augmented Generation for Post-hoc Literature-Grounded Assessment of Gene Regulatory Interaction

GitHub issuesPyPI - ProjectDocsDOI

Introduction

RAGulate is a Retrieval-Augmented Generation (RAG) pipeline that integrates domain-specific large language models (LLM) with curated literature to identify, score, and assess inferred regulatory (transcription factor-target gene) interactions in their biological context.

RAGulate pipeline

For further information and example tutorials, please check our documentation:

If you have any questions or concerns, feel free to open an issue.

Requirements

RAGulate is implemented in the LlamaIndex framework. Running RAGulate on CUDA is highly recommended if available.

Before installing and running RAGulate, ensure you have the following libraries installed:

  • PyTorch (version 2.0 or higher)
    Install with the exact command from the PyTorch “Get Started” page for your OS, Python version and (optionally) CUDA toolkit.
  • NumPy (version 1.23 or higher)
  • bm-25 (version 0.2.2 or higher)

You can install these dependencies using pip:

pip install torch numpy

Installation

Option 1:
You can install RAGulate via pip for a lightweight installation:

pip install ragulate-bio

Option 2:
Alternatively, if you want the latest, unreleased version, you can install it directly from the source on GitHub:

pip install git+https://github.com/YDaiLab/RAGulate.git

Import

importragulate_bioasragulate# recommended alias

Note: The PyPI distribution is named ragulate-bio to avoid a name conflict with an unrelated project called ragulate. Always import ragulate_bio in Python (you may alias it to ragulate for convenience)

Option 3 (Coming soon):
For users who prefer Conda or Mamba for environment management, you can install RAGulate along with extra dependencies:

Conda:

conda install -c zandigohar RAGulate

Mamba:

mamba create -n RAGulate -c zandigohar RAGulate

Tutorials

Run the tutorial notebook:

Reproducibility

Large reproducibility resources for RAGulate (including prior regulatory databases, serialized artifacts, and precomputed resources) are hosted separately on Zenodo due to size constraints:

https://doi.org/10.5281/zenodo.18498382

To reproduce the tutorial and experimental results, download the Zenodo dataset and place the files in the expected locations (see config.py), then run the notebooks in docs/ in order:

  1. NCBI-contextualization_Final.ipynb (generates the context-specific pickle)
  2. RAGulate_Modular_Rep.ipynb (full reproducibility run)

FAQ

Q1: Do I need a GPU to run RAGulate?
No, a GPU is not required. However, using a CUDA-enabled GPU is strongly recommended for faster runs, especially with large queries.

Q2: How do I know if I can use a GPU with RAGulate?
There are two quick checks:

  1. System check
    In your terminal, run nvidia-smi. If you see your GPU listed (model, memory, driver version), your machine has an NVIDIA GPU with the driver installed.

  2. Python check
    In a Python shell, run:

    importtorchprint(torch.cuda.is_available()) # True means PyTorch can see your GPUprint(torch.cuda.device_count()) # How many GPUs are usable

Q3: Can I use RAGulate with R-based tools?
RAGulate is written in Python and works directly with Numpy objects.

Q4: What if I also have another package called ragulate installed? RAGulate will warn you if it detects a conflicting installation. We recommend using a clean virtual environment to avoid import clashes.

Q5: How do I cite RAGulate?
See the Citation section below for the latest reference and preprint link.

Q6: How can I reproduce the paper’s results?
See our Reproducibility Guide for step-by-step instructions. Then run RAGulate.

Citation

This repository is under active development. Please cite as:
Zandigohar M. and Dai Y. RAGulate: RAGulate: Retrieval-Augmented Generation for Post-hoc Literature-Grounded Regulatory Assessment. 2026. https://doi.org/10.64898/2026.01.20.700704

Development & Contact

RAGulate was developed and is actively maintained by Mehrdad Zandigohar as part of his PhD research at the University of Illinois Chicago (UIC), in the lab of Dr. Yang Dai.

📬 For private questions, please email: mzandi2@uic.edu

🤝 For collaboration inquiries, please contact PI: Dr. Yang Dai (yangdai@uic.edu)

Contributions, feature suggestions, and feedback are always welcome!

License

The code in RAGulate is licensed under the MIT License, which permits academic and commercial use, modification, and distribution.

Please note that any third-party dependencies bundled with RAGulate may have their own respective licenses.

About

RAGulate is a retrieval-augmented generation (RAG) pipeline that integrates domain-specific language models with curated literature to identify, score, and validate inferred transcription factor-target gene interactions in their biological context.

Resources

Code of conduct

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - YDaiLab/RAGulate: RAGulate is a retrieval-augmented generation (RAG) pipeline that integrates domain-specific language models with curated literature to identify, score, and validate inferred transcription factor-target gene interactions in their biological context. · GitHub
Skip to content

Repository files navigation

RAGulateRAGulate logo

Retrieval-Augmented Generation for Post-hoc Literature-Grounded Assessment of Gene Regulatory Interaction

GitHub issuesPyPI - ProjectDocsDOI

Introduction

RAGulate is a Retrieval-Augmented Generation (RAG) pipeline that integrates domain-specific large language models (LLM) with curated literature to identify, score, and assess inferred regulatory (transcription factor-target gene) interactions in their biological context.

RAGulate pipeline

For further information and example tutorials, please check our documentation:

If you have any questions or concerns, feel free to open an issue.

Requirements

RAGulate is implemented in the LlamaIndex framework. Running RAGulate on CUDA is highly recommended if available.

Before installing and running RAGulate, ensure you have the following libraries installed:

  • PyTorch (version 2.0 or higher)
    Install with the exact command from the PyTorch “Get Started” page for your OS, Python version and (optionally) CUDA toolkit.
  • NumPy (version 1.23 or higher)
  • bm-25 (version 0.2.2 or higher)

You can install these dependencies using pip:

pip install torch numpy

Installation

Option 1:
You can install RAGulate via pip for a lightweight installation:

pip install ragulate-bio

Option 2:
Alternatively, if you want the latest, unreleased version, you can install it directly from the source on GitHub:

pip install git+https://github.com/YDaiLab/RAGulate.git

Import

importragulate_bioasragulate# recommended alias

Note: The PyPI distribution is named ragulate-bio to avoid a name conflict with an unrelated project called ragulate. Always import ragulate_bio in Python (you may alias it to ragulate for convenience)

Option 3 (Coming soon):
For users who prefer Conda or Mamba for environment management, you can install RAGulate along with extra dependencies:

Conda:

conda install -c zandigohar RAGulate

Mamba:

mamba create -n RAGulate -c zandigohar RAGulate

Tutorials

Run the tutorial notebook:

Reproducibility

Large reproducibility resources for RAGulate (including prior regulatory databases, serialized artifacts, and precomputed resources) are hosted separately on Zenodo due to size constraints:

https://doi.org/10.5281/zenodo.18498382

To reproduce the tutorial and experimental results, download the Zenodo dataset and place the files in the expected locations (see config.py), then run the notebooks in docs/ in order:

  1. NCBI-contextualization_Final.ipynb (generates the context-specific pickle)
  2. RAGulate_Modular_Rep.ipynb (full reproducibility run)

FAQ

Q1: Do I need a GPU to run RAGulate?
No, a GPU is not required. However, using a CUDA-enabled GPU is strongly recommended for faster runs, especially with large queries.

Q2: How do I know if I can use a GPU with RAGulate?
There are two quick checks:

  1. System check
    In your terminal, run nvidia-smi. If you see your GPU listed (model, memory, driver version), your machine has an NVIDIA GPU with the driver installed.

  2. Python check
    In a Python shell, run:

    importtorchprint(torch.cuda.is_available()) # True means PyTorch can see your GPUprint(torch.cuda.device_count()) # How many GPUs are usable

Q3: Can I use RAGulate with R-based tools?
RAGulate is written in Python and works directly with Numpy objects.

Q4: What if I also have another package called ragulate installed? RAGulate will warn you if it detects a conflicting installation. We recommend using a clean virtual environment to avoid import clashes.

Q5: How do I cite RAGulate?
See the Citation section below for the latest reference and preprint link.

Q6: How can I reproduce the paper’s results?
See our Reproducibility Guide for step-by-step instructions. Then run RAGulate.

Citation

This repository is under active development. Please cite as:
Zandigohar M. and Dai Y. RAGulate: RAGulate: Retrieval-Augmented Generation for Post-hoc Literature-Grounded Regulatory Assessment. 2026. https://doi.org/10.64898/2026.01.20.700704

Development & Contact

RAGulate was developed and is actively maintained by Mehrdad Zandigohar as part of his PhD research at the University of Illinois Chicago (UIC), in the lab of Dr. Yang Dai.

📬 For private questions, please email: mzandi2@uic.edu

🤝 For collaboration inquiries, please contact PI: Dr. Yang Dai (yangdai@uic.edu)

Contributions, feature suggestions, and feedback are always welcome!

License

The code in RAGulate is licensed under the MIT License, which permits academic and commercial use, modification, and distribution.

Please note that any third-party dependencies bundled with RAGulate may have their own respective licenses.

About

RAGulate is a retrieval-augmented generation (RAG) pipeline that integrates domain-specific language models with curated literature to identify, score, and validate inferred transcription factor-target gene interactions in their biological context.

Resources

Code of conduct

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - YDaiLab/RAGulate: RAGulate is a retrieval-augmented generation (RAG) pipeline that integrates domain-specific language models with curated literature to identify, score, and validate inferred transcription factor-target gene interactions in their biological context. · GitHub
Skip to content

Repository files navigation

RAGulateRAGulate logo

Retrieval-Augmented Generation for Post-hoc Literature-Grounded Assessment of Gene Regulatory Interaction

GitHub issuesPyPI - ProjectDocsDOI

Introduction

RAGulate is a Retrieval-Augmented Generation (RAG) pipeline that integrates domain-specific large language models (LLM) with curated literature to identify, score, and assess inferred regulatory (transcription factor-target gene) interactions in their biological context.

RAGulate pipeline

For further information and example tutorials, please check our documentation:

If you have any questions or concerns, feel free to open an issue.

Requirements

RAGulate is implemented in the LlamaIndex framework. Running RAGulate on CUDA is highly recommended if available.

Before installing and running RAGulate, ensure you have the following libraries installed:

  • PyTorch (version 2.0 or higher)
    Install with the exact command from the PyTorch “Get Started” page for your OS, Python version and (optionally) CUDA toolkit.
  • NumPy (version 1.23 or higher)
  • bm-25 (version 0.2.2 or higher)

You can install these dependencies using pip:

pip install torch numpy

Installation

Option 1:
You can install RAGulate via pip for a lightweight installation:

pip install ragulate-bio

Option 2:
Alternatively, if you want the latest, unreleased version, you can install it directly from the source on GitHub:

pip install git+https://github.com/YDaiLab/RAGulate.git

Import

importragulate_bioasragulate# recommended alias

Note: The PyPI distribution is named ragulate-bio to avoid a name conflict with an unrelated project called ragulate. Always import ragulate_bio in Python (you may alias it to ragulate for convenience)

Option 3 (Coming soon):
For users who prefer Conda or Mamba for environment management, you can install RAGulate along with extra dependencies:

Conda:

conda install -c zandigohar RAGulate

Mamba:

mamba create -n RAGulate -c zandigohar RAGulate

Tutorials

Run the tutorial notebook:

Reproducibility

Large reproducibility resources for RAGulate (including prior regulatory databases, serialized artifacts, and precomputed resources) are hosted separately on Zenodo due to size constraints:

https://doi.org/10.5281/zenodo.18498382

To reproduce the tutorial and experimental results, download the Zenodo dataset and place the files in the expected locations (see config.py), then run the notebooks in docs/ in order:

  1. NCBI-contextualization_Final.ipynb (generates the context-specific pickle)
  2. RAGulate_Modular_Rep.ipynb (full reproducibility run)

FAQ

Q1: Do I need a GPU to run RAGulate?
No, a GPU is not required. However, using a CUDA-enabled GPU is strongly recommended for faster runs, especially with large queries.

Q2: How do I know if I can use a GPU with RAGulate?
There are two quick checks:

  1. System check
    In your terminal, run nvidia-smi. If you see your GPU listed (model, memory, driver version), your machine has an NVIDIA GPU with the driver installed.

  2. Python check
    In a Python shell, run:

    importtorchprint(torch.cuda.is_available()) # True means PyTorch can see your GPUprint(torch.cuda.device_count()) # How many GPUs are usable

Q3: Can I use RAGulate with R-based tools?
RAGulate is written in Python and works directly with Numpy objects.

Q4: What if I also have another package called ragulate installed? RAGulate will warn you if it detects a conflicting installation. We recommend using a clean virtual environment to avoid import clashes.

Q5: How do I cite RAGulate?
See the Citation section below for the latest reference and preprint link.

Q6: How can I reproduce the paper’s results?
See our Reproducibility Guide for step-by-step instructions. Then run RAGulate.

Citation

This repository is under active development. Please cite as:
Zandigohar M. and Dai Y. RAGulate: RAGulate: Retrieval-Augmented Generation for Post-hoc Literature-Grounded Regulatory Assessment. 2026. https://doi.org/10.64898/2026.01.20.700704

Development & Contact

RAGulate was developed and is actively maintained by Mehrdad Zandigohar as part of his PhD research at the University of Illinois Chicago (UIC), in the lab of Dr. Yang Dai.

📬 For private questions, please email: mzandi2@uic.edu

🤝 For collaboration inquiries, please contact PI: Dr. Yang Dai (yangdai@uic.edu)

Contributions, feature suggestions, and feedback are always welcome!

License

The code in RAGulate is licensed under the MIT License, which permits academic and commercial use, modification, and distribution.

Please note that any third-party dependencies bundled with RAGulate may have their own respective licenses.

About

RAGulate is a retrieval-augmented generation (RAG) pipeline that integrates domain-specific language models with curated literature to identify, score, and validate inferred transcription factor-target gene interactions in their biological context.

Resources

Code of conduct

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - YDaiLab/RAGulate: RAGulate is a retrieval-augmented generation (RAG) pipeline that integrates domain-specific language models with curated literature to identify, score, and validate inferred transcription factor-target gene interactions in their biological context. · GitHub
Skip to content

Repository files navigation

RAGulateRAGulate logo

Retrieval-Augmented Generation for Post-hoc Literature-Grounded Assessment of Gene Regulatory Interaction

GitHub issuesPyPI - ProjectDocsDOI

Introduction

RAGulate is a Retrieval-Augmented Generation (RAG) pipeline that integrates domain-specific large language models (LLM) with curated literature to identify, score, and assess inferred regulatory (transcription factor-target gene) interactions in their biological context.

RAGulate pipeline

For further information and example tutorials, please check our documentation:

If you have any questions or concerns, feel free to open an issue.

Requirements

RAGulate is implemented in the LlamaIndex framework. Running RAGulate on CUDA is highly recommended if available.

Before installing and running RAGulate, ensure you have the following libraries installed:

  • PyTorch (version 2.0 or higher)
    Install with the exact command from the PyTorch “Get Started” page for your OS, Python version and (optionally) CUDA toolkit.
  • NumPy (version 1.23 or higher)
  • bm-25 (version 0.2.2 or higher)

You can install these dependencies using pip:

pip install torch numpy

Installation

Option 1:
You can install RAGulate via pip for a lightweight installation:

pip install ragulate-bio

Option 2:
Alternatively, if you want the latest, unreleased version, you can install it directly from the source on GitHub:

pip install git+https://github.com/YDaiLab/RAGulate.git

Import

importragulate_bioasragulate# recommended alias

Note: The PyPI distribution is named ragulate-bio to avoid a name conflict with an unrelated project called ragulate. Always import ragulate_bio in Python (you may alias it to ragulate for convenience)

Option 3 (Coming soon):
For users who prefer Conda or Mamba for environment management, you can install RAGulate along with extra dependencies:

Conda:

conda install -c zandigohar RAGulate

Mamba:

mamba create -n RAGulate -c zandigohar RAGulate

Tutorials

Run the tutorial notebook:

Reproducibility

Large reproducibility resources for RAGulate (including prior regulatory databases, serialized artifacts, and precomputed resources) are hosted separately on Zenodo due to size constraints:

https://doi.org/10.5281/zenodo.18498382

To reproduce the tutorial and experimental results, download the Zenodo dataset and place the files in the expected locations (see config.py), then run the notebooks in docs/ in order:

  1. NCBI-contextualization_Final.ipynb (generates the context-specific pickle)
  2. RAGulate_Modular_Rep.ipynb (full reproducibility run)

FAQ

Q1: Do I need a GPU to run RAGulate?
No, a GPU is not required. However, using a CUDA-enabled GPU is strongly recommended for faster runs, especially with large queries.

Q2: How do I know if I can use a GPU with RAGulate?
There are two quick checks:

  1. System check
    In your terminal, run nvidia-smi. If you see your GPU listed (model, memory, driver version), your machine has an NVIDIA GPU with the driver installed.

  2. Python check
    In a Python shell, run:

    importtorchprint(torch.cuda.is_available()) # True means PyTorch can see your GPUprint(torch.cuda.device_count()) # How many GPUs are usable

Q3: Can I use RAGulate with R-based tools?
RAGulate is written in Python and works directly with Numpy objects.

Q4: What if I also have another package called ragulate installed? RAGulate will warn you if it detects a conflicting installation. We recommend using a clean virtual environment to avoid import clashes.

Q5: How do I cite RAGulate?
See the Citation section below for the latest reference and preprint link.

Q6: How can I reproduce the paper’s results?
See our Reproducibility Guide for step-by-step instructions. Then run RAGulate.

Citation

This repository is under active development. Please cite as:
Zandigohar M. and Dai Y. RAGulate: RAGulate: Retrieval-Augmented Generation for Post-hoc Literature-Grounded Regulatory Assessment. 2026. https://doi.org/10.64898/2026.01.20.700704

Development & Contact

RAGulate was developed and is actively maintained by Mehrdad Zandigohar as part of his PhD research at the University of Illinois Chicago (UIC), in the lab of Dr. Yang Dai.

📬 For private questions, please email: mzandi2@uic.edu

🤝 For collaboration inquiries, please contact PI: Dr. Yang Dai (yangdai@uic.edu)

Contributions, feature suggestions, and feedback are always welcome!

License

The code in RAGulate is licensed under the MIT License, which permits academic and commercial use, modification, and distribution.

Please note that any third-party dependencies bundled with RAGulate may have their own respective licenses.

About

RAGulate is a retrieval-augmented generation (RAG) pipeline that integrates domain-specific language models with curated literature to identify, score, and validate inferred transcription factor-target gene interactions in their biological context.

Resources

Code of conduct

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - YDaiLab/RAGulate: RAGulate is a retrieval-augmented generation (RAG) pipeline that integrates domain-specific language models with curated literature to identify, score, and validate inferred transcription factor-target gene interactions in their biological context. · GitHub
Skip to content

Repository files navigation

RAGulateRAGulate logo

Retrieval-Augmented Generation for Post-hoc Literature-Grounded Assessment of Gene Regulatory Interaction

GitHub issuesPyPI - ProjectDocsDOI

Introduction

RAGulate is a Retrieval-Augmented Generation (RAG) pipeline that integrates domain-specific large language models (LLM) with curated literature to identify, score, and assess inferred regulatory (transcription factor-target gene) interactions in their biological context.

RAGulate pipeline

For further information and example tutorials, please check our documentation:

If you have any questions or concerns, feel free to open an issue.

Requirements

RAGulate is implemented in the LlamaIndex framework. Running RAGulate on CUDA is highly recommended if available.

Before installing and running RAGulate, ensure you have the following libraries installed:

  • PyTorch (version 2.0 or higher)
    Install with the exact command from the PyTorch “Get Started” page for your OS, Python version and (optionally) CUDA toolkit.
  • NumPy (version 1.23 or higher)
  • bm-25 (version 0.2.2 or higher)

You can install these dependencies using pip:

pip install torch numpy

Installation

Option 1:
You can install RAGulate via pip for a lightweight installation:

pip install ragulate-bio

Option 2:
Alternatively, if you want the latest, unreleased version, you can install it directly from the source on GitHub:

pip install git+https://github.com/YDaiLab/RAGulate.git

Import

importragulate_bioasragulate# recommended alias

Note: The PyPI distribution is named ragulate-bio to avoid a name conflict with an unrelated project called ragulate. Always import ragulate_bio in Python (you may alias it to ragulate for convenience)

Option 3 (Coming soon):
For users who prefer Conda or Mamba for environment management, you can install RAGulate along with extra dependencies:

Conda:

conda install -c zandigohar RAGulate

Mamba:

mamba create -n RAGulate -c zandigohar RAGulate

Tutorials

Run the tutorial notebook:

Reproducibility

Large reproducibility resources for RAGulate (including prior regulatory databases, serialized artifacts, and precomputed resources) are hosted separately on Zenodo due to size constraints:

https://doi.org/10.5281/zenodo.18498382

To reproduce the tutorial and experimental results, download the Zenodo dataset and place the files in the expected locations (see config.py), then run the notebooks in docs/ in order:

  1. NCBI-contextualization_Final.ipynb (generates the context-specific pickle)
  2. RAGulate_Modular_Rep.ipynb (full reproducibility run)

FAQ

Q1: Do I need a GPU to run RAGulate?
No, a GPU is not required. However, using a CUDA-enabled GPU is strongly recommended for faster runs, especially with large queries.

Q2: How do I know if I can use a GPU with RAGulate?
There are two quick checks:

  1. System check
    In your terminal, run nvidia-smi. If you see your GPU listed (model, memory, driver version), your machine has an NVIDIA GPU with the driver installed.

  2. Python check
    In a Python shell, run:

    importtorchprint(torch.cuda.is_available()) # True means PyTorch can see your GPUprint(torch.cuda.device_count()) # How many GPUs are usable

Q3: Can I use RAGulate with R-based tools?
RAGulate is written in Python and works directly with Numpy objects.

Q4: What if I also have another package called ragulate installed? RAGulate will warn you if it detects a conflicting installation. We recommend using a clean virtual environment to avoid import clashes.

Q5: How do I cite RAGulate?
See the Citation section below for the latest reference and preprint link.

Q6: How can I reproduce the paper’s results?
See our Reproducibility Guide for step-by-step instructions. Then run RAGulate.

Citation

This repository is under active development. Please cite as:
Zandigohar M. and Dai Y. RAGulate: RAGulate: Retrieval-Augmented Generation for Post-hoc Literature-Grounded Regulatory Assessment. 2026. https://doi.org/10.64898/2026.01.20.700704

Development & Contact

RAGulate was developed and is actively maintained by Mehrdad Zandigohar as part of his PhD research at the University of Illinois Chicago (UIC), in the lab of Dr. Yang Dai.

📬 For private questions, please email: mzandi2@uic.edu

🤝 For collaboration inquiries, please contact PI: Dr. Yang Dai (yangdai@uic.edu)

Contributions, feature suggestions, and feedback are always welcome!

License

The code in RAGulate is licensed under the MIT License, which permits academic and commercial use, modification, and distribution.

Please note that any third-party dependencies bundled with RAGulate may have their own respective licenses.

About

RAGulate is a retrieval-augmented generation (RAG) pipeline that integrates domain-specific language models with curated literature to identify, score, and validate inferred transcription factor-target gene interactions in their biological context.

Resources

Code of conduct

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - YDaiLab/RAGulate: RAGulate is a retrieval-augmented generation (RAG) pipeline that integrates domain-specific language models with curated literature to identify, score, and validate inferred transcription factor-target gene interactions in their biological context. · GitHub
Skip to content

Repository files navigation

RAGulateRAGulate logo

Retrieval-Augmented Generation for Post-hoc Literature-Grounded Assessment of Gene Regulatory Interaction

GitHub issuesPyPI - ProjectDocsDOI

Introduction

RAGulate is a Retrieval-Augmented Generation (RAG) pipeline that integrates domain-specific large language models (LLM) with curated literature to identify, score, and assess inferred regulatory (transcription factor-target gene) interactions in their biological context.

RAGulate pipeline

For further information and example tutorials, please check our documentation:

If you have any questions or concerns, feel free to open an issue.

Requirements

RAGulate is implemented in the LlamaIndex framework. Running RAGulate on CUDA is highly recommended if available.

Before installing and running RAGulate, ensure you have the following libraries installed:

  • PyTorch (version 2.0 or higher)
    Install with the exact command from the PyTorch “Get Started” page for your OS, Python version and (optionally) CUDA toolkit.
  • NumPy (version 1.23 or higher)
  • bm-25 (version 0.2.2 or higher)

You can install these dependencies using pip:

pip install torch numpy

Installation

Option 1:
You can install RAGulate via pip for a lightweight installation:

pip install ragulate-bio

Option 2:
Alternatively, if you want the latest, unreleased version, you can install it directly from the source on GitHub:

pip install git+https://github.com/YDaiLab/RAGulate.git

Import

importragulate_bioasragulate# recommended alias

Note: The PyPI distribution is named ragulate-bio to avoid a name conflict with an unrelated project called ragulate. Always import ragulate_bio in Python (you may alias it to ragulate for convenience)

Option 3 (Coming soon):
For users who prefer Conda or Mamba for environment management, you can install RAGulate along with extra dependencies:

Conda:

conda install -c zandigohar RAGulate

Mamba:

mamba create -n RAGulate -c zandigohar RAGulate

Tutorials

Run the tutorial notebook:

Reproducibility

Large reproducibility resources for RAGulate (including prior regulatory databases, serialized artifacts, and precomputed resources) are hosted separately on Zenodo due to size constraints:

https://doi.org/10.5281/zenodo.18498382

To reproduce the tutorial and experimental results, download the Zenodo dataset and place the files in the expected locations (see config.py), then run the notebooks in docs/ in order:

  1. NCBI-contextualization_Final.ipynb (generates the context-specific pickle)
  2. RAGulate_Modular_Rep.ipynb (full reproducibility run)

FAQ

Q1: Do I need a GPU to run RAGulate?
No, a GPU is not required. However, using a CUDA-enabled GPU is strongly recommended for faster runs, especially with large queries.

Q2: How do I know if I can use a GPU with RAGulate?
There are two quick checks:

  1. System check
    In your terminal, run nvidia-smi. If you see your GPU listed (model, memory, driver version), your machine has an NVIDIA GPU with the driver installed.

  2. Python check
    In a Python shell, run:

    importtorchprint(torch.cuda.is_available()) # True means PyTorch can see your GPUprint(torch.cuda.device_count()) # How many GPUs are usable

Q3: Can I use RAGulate with R-based tools?
RAGulate is written in Python and works directly with Numpy objects.

Q4: What if I also have another package called ragulate installed? RAGulate will warn you if it detects a conflicting installation. We recommend using a clean virtual environment to avoid import clashes.

Q5: How do I cite RAGulate?
See the Citation section below for the latest reference and preprint link.

Q6: How can I reproduce the paper’s results?
See our Reproducibility Guide for step-by-step instructions. Then run RAGulate.

Citation

This repository is under active development. Please cite as:
Zandigohar M. and Dai Y. RAGulate: RAGulate: Retrieval-Augmented Generation for Post-hoc Literature-Grounded Regulatory Assessment. 2026. https://doi.org/10.64898/2026.01.20.700704

Development & Contact

RAGulate was developed and is actively maintained by Mehrdad Zandigohar as part of his PhD research at the University of Illinois Chicago (UIC), in the lab of Dr. Yang Dai.

📬 For private questions, please email: mzandi2@uic.edu

🤝 For collaboration inquiries, please contact PI: Dr. Yang Dai (yangdai@uic.edu)

Contributions, feature suggestions, and feedback are always welcome!

License

The code in RAGulate is licensed under the MIT License, which permits academic and commercial use, modification, and distribution.

Please note that any third-party dependencies bundled with RAGulate may have their own respective licenses.

About

RAGulate is a retrieval-augmented generation (RAG) pipeline that integrates domain-specific language models with curated literature to identify, score, and validate inferred transcription factor-target gene interactions in their biological context.

Resources

Code of conduct

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - YDaiLab/RAGulate: RAGulate is a retrieval-augmented generation (RAG) pipeline that integrates domain-specific language models with curated literature to identify, score, and validate inferred transcription factor-target gene interactions in their biological context. · GitHub
Skip to content

Repository files navigation

RAGulateRAGulate logo

Retrieval-Augmented Generation for Post-hoc Literature-Grounded Assessment of Gene Regulatory Interaction

GitHub issuesPyPI - ProjectDocsDOI

Introduction

RAGulate is a Retrieval-Augmented Generation (RAG) pipeline that integrates domain-specific large language models (LLM) with curated literature to identify, score, and assess inferred regulatory (transcription factor-target gene) interactions in their biological context.

RAGulate pipeline

For further information and example tutorials, please check our documentation:

If you have any questions or concerns, feel free to open an issue.

Requirements

RAGulate is implemented in the LlamaIndex framework. Running RAGulate on CUDA is highly recommended if available.

Before installing and running RAGulate, ensure you have the following libraries installed:

  • PyTorch (version 2.0 or higher)
    Install with the exact command from the PyTorch “Get Started” page for your OS, Python version and (optionally) CUDA toolkit.
  • NumPy (version 1.23 or higher)
  • bm-25 (version 0.2.2 or higher)

You can install these dependencies using pip:

pip install torch numpy

Installation

Option 1:
You can install RAGulate via pip for a lightweight installation:

pip install ragulate-bio

Option 2:
Alternatively, if you want the latest, unreleased version, you can install it directly from the source on GitHub:

pip install git+https://github.com/YDaiLab/RAGulate.git

Import

importragulate_bioasragulate# recommended alias

Note: The PyPI distribution is named ragulate-bio to avoid a name conflict with an unrelated project called ragulate. Always import ragulate_bio in Python (you may alias it to ragulate for convenience)

Option 3 (Coming soon):
For users who prefer Conda or Mamba for environment management, you can install RAGulate along with extra dependencies:

Conda:

conda install -c zandigohar RAGulate

Mamba:

mamba create -n RAGulate -c zandigohar RAGulate

Tutorials

Run the tutorial notebook:

Reproducibility

Large reproducibility resources for RAGulate (including prior regulatory databases, serialized artifacts, and precomputed resources) are hosted separately on Zenodo due to size constraints:

https://doi.org/10.5281/zenodo.18498382

To reproduce the tutorial and experimental results, download the Zenodo dataset and place the files in the expected locations (see config.py), then run the notebooks in docs/ in order:

  1. NCBI-contextualization_Final.ipynb (generates the context-specific pickle)
  2. RAGulate_Modular_Rep.ipynb (full reproducibility run)

FAQ

Q1: Do I need a GPU to run RAGulate?
No, a GPU is not required. However, using a CUDA-enabled GPU is strongly recommended for faster runs, especially with large queries.

Q2: How do I know if I can use a GPU with RAGulate?
There are two quick checks:

  1. System check
    In your terminal, run nvidia-smi. If you see your GPU listed (model, memory, driver version), your machine has an NVIDIA GPU with the driver installed.

  2. Python check
    In a Python shell, run:

    importtorchprint(torch.cuda.is_available()) # True means PyTorch can see your GPUprint(torch.cuda.device_count()) # How many GPUs are usable

Q3: Can I use RAGulate with R-based tools?
RAGulate is written in Python and works directly with Numpy objects.

Q4: What if I also have another package called ragulate installed? RAGulate will warn you if it detects a conflicting installation. We recommend using a clean virtual environment to avoid import clashes.

Q5: How do I cite RAGulate?
See the Citation section below for the latest reference and preprint link.

Q6: How can I reproduce the paper’s results?
See our Reproducibility Guide for step-by-step instructions. Then run RAGulate.

Citation

This repository is under active development. Please cite as:
Zandigohar M. and Dai Y. RAGulate: RAGulate: Retrieval-Augmented Generation for Post-hoc Literature-Grounded Regulatory Assessment. 2026. https://doi.org/10.64898/2026.01.20.700704

Development & Contact

RAGulate was developed and is actively maintained by Mehrdad Zandigohar as part of his PhD research at the University of Illinois Chicago (UIC), in the lab of Dr. Yang Dai.

📬 For private questions, please email: mzandi2@uic.edu

🤝 For collaboration inquiries, please contact PI: Dr. Yang Dai (yangdai@uic.edu)

Contributions, feature suggestions, and feedback are always welcome!

License

The code in RAGulate is licensed under the MIT License, which permits academic and commercial use, modification, and distribution.

Please note that any third-party dependencies bundled with RAGulate may have their own respective licenses.

About

RAGulate is a retrieval-augmented generation (RAG) pipeline that integrates domain-specific language models with curated literature to identify, score, and validate inferred transcription factor-target gene interactions in their biological context.

Resources

Code of conduct

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages