Skip to content

NVIDIA BioNeMo

Welcome to the NVIDIA BioNeMo Github Organization.

The repositories hosted here often accompany our publications and are released open source to the community to accelerate scientific innovation. Unless otherwise noted, the code is provided as-is and is not actively maintained.

Find out more about our research labs at research.nvidia.com/labs/dbr and research.nvidia.com and about our healthcare products at nvidia.com/en-us/industries/healthcare-life-sciences/.

Explore our pinned repositories below to see some of our featured projects!


NVIDIA BioNeMo is an open developer platform for AI-driven life science research.

It provides GPU-accelerated models, tools, and datasets for the entire AI lifecycle, enabling researchers and developers to build, customize, and deploy AI applications that transform physical lab results into the digital insights that drive the next experiment.

The platform is built on five core pillars:

  • Data: Large-scale datasets for training, fine-tuning, and benchmarking models.
  • Models: Open-source models for understanding biological systems, designing novel proteins and small molecules, and optimizing candidates for synthesizability, binding affinity, and molecular properties.
  • Libraries and Tools: Foundational GPU-optimized libraries and kernels for accelerated AI training and inference.
  • Training and Customization: Frameworks and recipes for pretraining, fine-tuning, and adapting models for specialized use cases.
  • Optimized Inference and Deployment: Enterprise-ready NVIDIA inference microservices (NIM) and reference architectures for production use.

Note: Many components of the BioNeMo platform are modular and hosted in their own dedicated GitHub repositories or organizations. This README serves as a central index to guide you to the right tools.

Table of Contents


License

BioNeMo components are generally released under:

Individual components may vary — check each resource for specific license terms.


Data

Unlike natural language models trained on internet-scale data, biology and chemistry lack the critical mass of data required for large, general-purpose foundation models. To address this ecosystem-wide gap, NVIDIA is partnering with leading organizations to create and release open datasets.

DatasetDescription
3D Structures of Protein Complexes
(available through the AlphaFold Database)
Large-scale open database of predicted protein complex structures built with ecosystem partners to accelerate interaction biology and drug discovery. License: CC BY 4.0
Consistency Distilled Synthetic Protein Database455K curated, high-quality protein sequence-structure pairs. Built using ProteinMPNN to generate synthetic sequences for Foldseek AFDB cluster representative structures, then refolded with ESMFold to obtain fully atomistic, self-consistent models. Filtered to pLDDT > 80. License: CC BY 4.0

Models

NVIDIA BioNeMo provides high-quality, fully open-source models — including the full training codebase, pre-trained weights, and research papers — completely free to use. These models are hosted in the NVIDIA-BioNeMo GitHub organization.

These models reflect our active research directions, and we highly encourage community feedback, collaboration, and adaptation to push their capabilities further.

Understand

Use CaseModelDescription
Target Identification / Disease Understanding (RNA)CodonFMCodon-level RNA foundation model trained on 130M protein-coding sequences from 22K+ species. Captures synonymous codon variation for mRNA design, stability modeling, and variant interpretation.
Structure Prediction (RNA)RNAProState-of-the-art RNA 3D structure prediction model. Combines Protenix-based co-folding architectures with RNA foundation models, MSA, and template-based modeling.

Design

Use CaseModelDescription
ProteinsProteina-ComplexaProtein binder design for protein and small molecule targets. Combines a pretrained flow-based generative model (built on La-Proteina) with inference-time optimization for high-quality binder generation.
La-ProteinaAll-atom protein generation using partially latent flow matching. Jointly generates amino acid sequence and full atomistic structure (backbone + side chains) for up to 800 residues. Enables atomistic motif scaffolding for enzyme design.
ProteinaLarge-scale flow-based generative model for protein backbone structures with hierarchical fold-class conditioning and a scalable transformer architecture.
ProtComposerSpatial-layout-conditioned protein structure generation using 3D ellipsoids to control shape and substructure arrangements.
Small MoleculesGenMolFragment-based molecule generation using masked discrete diffusion over SAFE representations. Supports de novo design, scaffold decoration, linker design, motif extension, and lead optimization.
MegalodonTransformer-based 3D molecule generative model using equivariant graph transformer architecture. Generates both 2D topology and 3D structure with physically realistic, low-energy conformations.
AvgFlowEfficient molecular 3D conformer generation using SO(3)-averaged flow-matching and reflow. Architecture-agnostic framework applicable to equivariant and non-equivariant models.

Optimize

Use CaseModelDescription
Property PredictionKERMTPretrained graph neural network for molecular property prediction (ADMET). Multi-task extension of GROVER with accelerated data loading via cuik-molmaker. SOTA on real-world ADMET data.
SynthesizabilityReaSynSynthesis pathway prediction using an encoder-decoder Transformer with Chain-of-Reaction notation. Predicts reaction steps from building blocks to final products, or finds synthesizable analogs for unsynthesizable targets.
Binding EnergyDualBind3D structure-based deep learning model for protein-ligand binding affinity prediction using a dual-loss framework (supervised MSE + unsupervised denoising). Orders of magnitude faster than physics-based FEP methods.

Libraries and Tools

GPU-optimized libraries and tools that integrate into existing workflows. Engineered to be lightweight and specialized for maximum performance without dependency bloat.

TaskToolDescription
Data Processing & AnalysisParabricksGPU-accelerated genomics software suite for rapid secondary analysis of DNA/RNA sequencing data.
nvMolKitGPU-accelerated cheminformatics library for molecular fingerprinting, Tanimoto/cosine similarity, Butina clustering, conformer generation (ETKDGv3), MMFF geometry optimization, and substructure search.
cuik-molmakerMolecular featurization package for converting chemical structures into GNN inputs. Accelerates Chemprop training by 1.6x and inference by 2.4x with 80% memory reduction.
nvQSPGPU-accelerated Quantitative Systems Pharmacology ODE solvers. 77x speedup over CPU for virtual patient simulations with bit-exact FP64 reproducibility.
Training & InferencecuEquivarianceCUDA-X library with optimized kernels for efficient training of geometry-aware equivariant neural networks (AlphaFold-like and molecular structure models).
BioNeMo-SCDLScalable, memory-efficient data loader for training large single-cell models. Part of BioNeMo Framework.
BioNeMo-MoCoFramework for constructing generative models (diffusion, flow-matching) using continuous and discrete interpolants. Part of BioNeMo Framework.
BioNeMo-NoodlesEfficient genomic data handling with memory-mapped access to FASTA files. Part of BioNeMo Framework.

Training and Customization

BioNeMo provides frameworks and recipes for pretraining, fine-tuning, and adapting biomolecular AI models at scale on GPU infrastructure.

ToolDescription
BioNeMo FrameworkReference training implementations and ready-to-run examples showing how to achieve lower-precision training, maximum scaling & throughput for models like Llama3, ESM2, Evo2, CodonFM, and Geneformer using FSDP and TransformerEngine.
Context Parallelism (boltz-cp)Long-sequence parallelism for protein structure prediction models. Distributes activation tensors across GPUs to overcome single-GPU memory limits for large biomolecules.

Documentation: docs.nvidia.com/bionemo-framework


Optimized Inference and Deployment

BioNeMo NIM microservices are enterprise-ready inference microservices with built-in API endpoints. Each NIM includes algorithmic, system, and runtime optimizations into a prebuilt container — go from zero to inference in minutes.

NIMDescription
OpenFold33D structure prediction for molecular complexes (proteins, DNA, RNA, ligands)
OpenFold2Protein structure prediction from sequence, MSAs, and templates
Boltz-2Biomolecular complex structure prediction
Evo2-40BGenomic foundation model with long-context sequence understanding
MSA SearchMultiple sequence alignment generation from query sequences
ProteinMPNNAmino acid sequence design for protein backbones
RFDiffusionGenerative model for protein backbone and binder design
GenMolFragment-based small molecule generation
DiffDockMolecular blind docking for predicting protein-ligand binding poses
MolMIMMolecular generation optimized for user-defined drug properties

Browse all available NIM microservices: build.nvidia.com/explore/biology

NIM microservices can be deployed self-hosted via Docker or Kubernetes, or on cloud platforms including AWS, Google Cloud, Microsoft Azure, and NVIDIA DGX Cloud.


Workflow Examples and Community Contributions

Application-level examples showing how BioNeMo platform components work together:

Note: If you have an example you'd like to contribute, we'd love to include it. Please get started by opening a GitHub issue and we'll reach out to you.

Pinned Loading

  1. bionemo-recipesbionemo-recipesPublic

    BioNeMo Recipes: For building and adapting AI models in drug discovery at scale

    Python 848 177

  2. Proteina-ComplexaProteina-ComplexaPublic

    Generative model for protein binder design for protein and small molecule targets. Combines a pretrained flow-based generative model (built on La-Proteina) with inference-time optimization.

    Python 421 73

  3. nvMolKitnvMolKitPublic

    A high-performance, GPU-accelerated library for key computational chemistry tasks, such as molecular similarity, conformer generation, and geometry relaxation.

    Cuda 273 27

  4. genmolgenmolPublic

    GenMol is a generative AI model for creating novel molecules. It utilizes masked discrete diffusion and fragment-based generation to create valid molecules, which are encoded in the SAFE molecular …

    Python 199 35

  5. RNAProRNAProPublic

    RNAPro is a state-of-the-art RNA 3D folding model developed in collaboration with the hosts and winners of the Stanford RNA 3D Folding Kaggle competition.

    Python 95 20

  6. KERMTKERMTPublic

    KERMT is a pretrained graph neural network model for molecular property prediction.

    Python 96 17

Repositories

Showing 10 of 23 repositories

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…

Most used topics

Loading…

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
NVIDIA BioNeMo · GitHub
Skip to content

NVIDIA BioNeMo

Welcome to the NVIDIA BioNeMo Github Organization.

The repositories hosted here often accompany our publications and are released open source to the community to accelerate scientific innovation. Unless otherwise noted, the code is provided as-is and is not actively maintained.

Find out more about our research labs at research.nvidia.com/labs/dbr and research.nvidia.com and about our healthcare products at nvidia.com/en-us/industries/healthcare-life-sciences/.

Explore our pinned repositories below to see some of our featured projects!


NVIDIA BioNeMo is an open developer platform for AI-driven life science research.

It provides GPU-accelerated models, tools, and datasets for the entire AI lifecycle, enabling researchers and developers to build, customize, and deploy AI applications that transform physical lab results into the digital insights that drive the next experiment.

The platform is built on five core pillars:

  • Data: Large-scale datasets for training, fine-tuning, and benchmarking models.
  • Models: Open-source models for understanding biological systems, designing novel proteins and small molecules, and optimizing candidates for synthesizability, binding affinity, and molecular properties.
  • Libraries and Tools: Foundational GPU-optimized libraries and kernels for accelerated AI training and inference.
  • Training and Customization: Frameworks and recipes for pretraining, fine-tuning, and adapting models for specialized use cases.
  • Optimized Inference and Deployment: Enterprise-ready NVIDIA inference microservices (NIM) and reference architectures for production use.

Note: Many components of the BioNeMo platform are modular and hosted in their own dedicated GitHub repositories or organizations. This README serves as a central index to guide you to the right tools.

Table of Contents


License

BioNeMo components are generally released under:

Individual components may vary — check each resource for specific license terms.


Data

Unlike natural language models trained on internet-scale data, biology and chemistry lack the critical mass of data required for large, general-purpose foundation models. To address this ecosystem-wide gap, NVIDIA is partnering with leading organizations to create and release open datasets.

DatasetDescription
3D Structures of Protein Complexes
(available through the AlphaFold Database)
Large-scale open database of predicted protein complex structures built with ecosystem partners to accelerate interaction biology and drug discovery. License: CC BY 4.0
Consistency Distilled Synthetic Protein Database455K curated, high-quality protein sequence-structure pairs. Built using ProteinMPNN to generate synthetic sequences for Foldseek AFDB cluster representative structures, then refolded with ESMFold to obtain fully atomistic, self-consistent models. Filtered to pLDDT > 80. License: CC BY 4.0

Models

NVIDIA BioNeMo provides high-quality, fully open-source models — including the full training codebase, pre-trained weights, and research papers — completely free to use. These models are hosted in the NVIDIA-BioNeMo GitHub organization.

These models reflect our active research directions, and we highly encourage community feedback, collaboration, and adaptation to push their capabilities further.

Understand

Use CaseModelDescription
Target Identification / Disease Understanding (RNA)CodonFMCodon-level RNA foundation model trained on 130M protein-coding sequences from 22K+ species. Captures synonymous codon variation for mRNA design, stability modeling, and variant interpretation.
Structure Prediction (RNA)RNAProState-of-the-art RNA 3D structure prediction model. Combines Protenix-based co-folding architectures with RNA foundation models, MSA, and template-based modeling.

Design

Use CaseModelDescription
ProteinsProteina-ComplexaProtein binder design for protein and small molecule targets. Combines a pretrained flow-based generative model (built on La-Proteina) with inference-time optimization for high-quality binder generation.
La-ProteinaAll-atom protein generation using partially latent flow matching. Jointly generates amino acid sequence and full atomistic structure (backbone + side chains) for up to 800 residues. Enables atomistic motif scaffolding for enzyme design.
ProteinaLarge-scale flow-based generative model for protein backbone structures with hierarchical fold-class conditioning and a scalable transformer architecture.
ProtComposerSpatial-layout-conditioned protein structure generation using 3D ellipsoids to control shape and substructure arrangements.
Small MoleculesGenMolFragment-based molecule generation using masked discrete diffusion over SAFE representations. Supports de novo design, scaffold decoration, linker design, motif extension, and lead optimization.
MegalodonTransformer-based 3D molecule generative model using equivariant graph transformer architecture. Generates both 2D topology and 3D structure with physically realistic, low-energy conformations.
AvgFlowEfficient molecular 3D conformer generation using SO(3)-averaged flow-matching and reflow. Architecture-agnostic framework applicable to equivariant and non-equivariant models.

Optimize

Use CaseModelDescription
Property PredictionKERMTPretrained graph neural network for molecular property prediction (ADMET). Multi-task extension of GROVER with accelerated data loading via cuik-molmaker. SOTA on real-world ADMET data.
SynthesizabilityReaSynSynthesis pathway prediction using an encoder-decoder Transformer with Chain-of-Reaction notation. Predicts reaction steps from building blocks to final products, or finds synthesizable analogs for unsynthesizable targets.
Binding EnergyDualBind3D structure-based deep learning model for protein-ligand binding affinity prediction using a dual-loss framework (supervised MSE + unsupervised denoising). Orders of magnitude faster than physics-based FEP methods.

Libraries and Tools

GPU-optimized libraries and tools that integrate into existing workflows. Engineered to be lightweight and specialized for maximum performance without dependency bloat.

TaskToolDescription
Data Processing & AnalysisParabricksGPU-accelerated genomics software suite for rapid secondary analysis of DNA/RNA sequencing data.
nvMolKitGPU-accelerated cheminformatics library for molecular fingerprinting, Tanimoto/cosine similarity, Butina clustering, conformer generation (ETKDGv3), MMFF geometry optimization, and substructure search.
cuik-molmakerMolecular featurization package for converting chemical structures into GNN inputs. Accelerates Chemprop training by 1.6x and inference by 2.4x with 80% memory reduction.
nvQSPGPU-accelerated Quantitative Systems Pharmacology ODE solvers. 77x speedup over CPU for virtual patient simulations with bit-exact FP64 reproducibility.
Training & InferencecuEquivarianceCUDA-X library with optimized kernels for efficient training of geometry-aware equivariant neural networks (AlphaFold-like and molecular structure models).
BioNeMo-SCDLScalable, memory-efficient data loader for training large single-cell models. Part of BioNeMo Framework.
BioNeMo-MoCoFramework for constructing generative models (diffusion, flow-matching) using continuous and discrete interpolants. Part of BioNeMo Framework.
BioNeMo-NoodlesEfficient genomic data handling with memory-mapped access to FASTA files. Part of BioNeMo Framework.

Training and Customization

BioNeMo provides frameworks and recipes for pretraining, fine-tuning, and adapting biomolecular AI models at scale on GPU infrastructure.

ToolDescription
BioNeMo FrameworkReference training implementations and ready-to-run examples showing how to achieve lower-precision training, maximum scaling & throughput for models like Llama3, ESM2, Evo2, CodonFM, and Geneformer using FSDP and TransformerEngine.
Context Parallelism (boltz-cp)Long-sequence parallelism for protein structure prediction models. Distributes activation tensors across GPUs to overcome single-GPU memory limits for large biomolecules.

Documentation: docs.nvidia.com/bionemo-framework


Optimized Inference and Deployment

BioNeMo NIM microservices are enterprise-ready inference microservices with built-in API endpoints. Each NIM includes algorithmic, system, and runtime optimizations into a prebuilt container — go from zero to inference in minutes.

NIMDescription
OpenFold33D structure prediction for molecular complexes (proteins, DNA, RNA, ligands)
OpenFold2Protein structure prediction from sequence, MSAs, and templates
Boltz-2Biomolecular complex structure prediction
Evo2-40BGenomic foundation model with long-context sequence understanding
MSA SearchMultiple sequence alignment generation from query sequences
ProteinMPNNAmino acid sequence design for protein backbones
RFDiffusionGenerative model for protein backbone and binder design
GenMolFragment-based small molecule generation
DiffDockMolecular blind docking for predicting protein-ligand binding poses
MolMIMMolecular generation optimized for user-defined drug properties

Browse all available NIM microservices: build.nvidia.com/explore/biology

NIM microservices can be deployed self-hosted via Docker or Kubernetes, or on cloud platforms including AWS, Google Cloud, Microsoft Azure, and NVIDIA DGX Cloud.


Workflow Examples and Community Contributions

Application-level examples showing how BioNeMo platform components work together:

Note: If you have an example you'd like to contribute, we'd love to include it. Please get started by opening a GitHub issue and we'll reach out to you.

Pinned Loading

  1. bionemo-recipesbionemo-recipesPublic

    BioNeMo Recipes: For building and adapting AI models in drug discovery at scale

    Python 848 177

  2. Proteina-ComplexaProteina-ComplexaPublic

    Generative model for protein binder design for protein and small molecule targets. Combines a pretrained flow-based generative model (built on La-Proteina) with inference-time optimization.

    Python 421 73

  3. nvMolKitnvMolKitPublic

    A high-performance, GPU-accelerated library for key computational chemistry tasks, such as molecular similarity, conformer generation, and geometry relaxation.

    Cuda 273 27

  4. genmolgenmolPublic

    GenMol is a generative AI model for creating novel molecules. It utilizes masked discrete diffusion and fragment-based generation to create valid molecules, which are encoded in the SAFE molecular …

    Python 199 35

  5. RNAProRNAProPublic

    RNAPro is a state-of-the-art RNA 3D folding model developed in collaboration with the hosts and winners of the Stanford RNA 3D Folding Kaggle competition.

    Python 95 20

  6. KERMTKERMTPublic

    KERMT is a pretrained graph neural network model for molecular property prediction.

    Python 96 17

Repositories

Showing 10 of 23 repositories

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…

Most used topics

Loading…

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' NVIDIA BioNeMo · GitHub
Skip to content

NVIDIA BioNeMo

Welcome to the NVIDIA BioNeMo Github Organization.

The repositories hosted here often accompany our publications and are released open source to the community to accelerate scientific innovation. Unless otherwise noted, the code is provided as-is and is not actively maintained.

Find out more about our research labs at research.nvidia.com/labs/dbr and research.nvidia.com and about our healthcare products at nvidia.com/en-us/industries/healthcare-life-sciences/.

Explore our pinned repositories below to see some of our featured projects!


NVIDIA BioNeMo is an open developer platform for AI-driven life science research.

It provides GPU-accelerated models, tools, and datasets for the entire AI lifecycle, enabling researchers and developers to build, customize, and deploy AI applications that transform physical lab results into the digital insights that drive the next experiment.

The platform is built on five core pillars:

  • Data: Large-scale datasets for training, fine-tuning, and benchmarking models.
  • Models: Open-source models for understanding biological systems, designing novel proteins and small molecules, and optimizing candidates for synthesizability, binding affinity, and molecular properties.
  • Libraries and Tools: Foundational GPU-optimized libraries and kernels for accelerated AI training and inference.
  • Training and Customization: Frameworks and recipes for pretraining, fine-tuning, and adapting models for specialized use cases.
  • Optimized Inference and Deployment: Enterprise-ready NVIDIA inference microservices (NIM) and reference architectures for production use.

Note: Many components of the BioNeMo platform are modular and hosted in their own dedicated GitHub repositories or organizations. This README serves as a central index to guide you to the right tools.

Table of Contents


License

BioNeMo components are generally released under:

Individual components may vary — check each resource for specific license terms.


Data

Unlike natural language models trained on internet-scale data, biology and chemistry lack the critical mass of data required for large, general-purpose foundation models. To address this ecosystem-wide gap, NVIDIA is partnering with leading organizations to create and release open datasets.

DatasetDescription
3D Structures of Protein Complexes
(available through the AlphaFold Database)
Large-scale open database of predicted protein complex structures built with ecosystem partners to accelerate interaction biology and drug discovery. License: CC BY 4.0
Consistency Distilled Synthetic Protein Database455K curated, high-quality protein sequence-structure pairs. Built using ProteinMPNN to generate synthetic sequences for Foldseek AFDB cluster representative structures, then refolded with ESMFold to obtain fully atomistic, self-consistent models. Filtered to pLDDT > 80. License: CC BY 4.0

Models

NVIDIA BioNeMo provides high-quality, fully open-source models — including the full training codebase, pre-trained weights, and research papers — completely free to use. These models are hosted in the NVIDIA-BioNeMo GitHub organization.

These models reflect our active research directions, and we highly encourage community feedback, collaboration, and adaptation to push their capabilities further.

Understand

Use CaseModelDescription
Target Identification / Disease Understanding (RNA)CodonFMCodon-level RNA foundation model trained on 130M protein-coding sequences from 22K+ species. Captures synonymous codon variation for mRNA design, stability modeling, and variant interpretation.
Structure Prediction (RNA)RNAProState-of-the-art RNA 3D structure prediction model. Combines Protenix-based co-folding architectures with RNA foundation models, MSA, and template-based modeling.

Design

Use CaseModelDescription
ProteinsProteina-ComplexaProtein binder design for protein and small molecule targets. Combines a pretrained flow-based generative model (built on La-Proteina) with inference-time optimization for high-quality binder generation.
La-ProteinaAll-atom protein generation using partially latent flow matching. Jointly generates amino acid sequence and full atomistic structure (backbone + side chains) for up to 800 residues. Enables atomistic motif scaffolding for enzyme design.
ProteinaLarge-scale flow-based generative model for protein backbone structures with hierarchical fold-class conditioning and a scalable transformer architecture.
ProtComposerSpatial-layout-conditioned protein structure generation using 3D ellipsoids to control shape and substructure arrangements.
Small MoleculesGenMolFragment-based molecule generation using masked discrete diffusion over SAFE representations. Supports de novo design, scaffold decoration, linker design, motif extension, and lead optimization.
MegalodonTransformer-based 3D molecule generative model using equivariant graph transformer architecture. Generates both 2D topology and 3D structure with physically realistic, low-energy conformations.
AvgFlowEfficient molecular 3D conformer generation using SO(3)-averaged flow-matching and reflow. Architecture-agnostic framework applicable to equivariant and non-equivariant models.

Optimize

Use CaseModelDescription
Property PredictionKERMTPretrained graph neural network for molecular property prediction (ADMET). Multi-task extension of GROVER with accelerated data loading via cuik-molmaker. SOTA on real-world ADMET data.
SynthesizabilityReaSynSynthesis pathway prediction using an encoder-decoder Transformer with Chain-of-Reaction notation. Predicts reaction steps from building blocks to final products, or finds synthesizable analogs for unsynthesizable targets.
Binding EnergyDualBind3D structure-based deep learning model for protein-ligand binding affinity prediction using a dual-loss framework (supervised MSE + unsupervised denoising). Orders of magnitude faster than physics-based FEP methods.

Libraries and Tools

GPU-optimized libraries and tools that integrate into existing workflows. Engineered to be lightweight and specialized for maximum performance without dependency bloat.

TaskToolDescription
Data Processing & AnalysisParabricksGPU-accelerated genomics software suite for rapid secondary analysis of DNA/RNA sequencing data.
nvMolKitGPU-accelerated cheminformatics library for molecular fingerprinting, Tanimoto/cosine similarity, Butina clustering, conformer generation (ETKDGv3), MMFF geometry optimization, and substructure search.
cuik-molmakerMolecular featurization package for converting chemical structures into GNN inputs. Accelerates Chemprop training by 1.6x and inference by 2.4x with 80% memory reduction.
nvQSPGPU-accelerated Quantitative Systems Pharmacology ODE solvers. 77x speedup over CPU for virtual patient simulations with bit-exact FP64 reproducibility.
Training & InferencecuEquivarianceCUDA-X library with optimized kernels for efficient training of geometry-aware equivariant neural networks (AlphaFold-like and molecular structure models).
BioNeMo-SCDLScalable, memory-efficient data loader for training large single-cell models. Part of BioNeMo Framework.
BioNeMo-MoCoFramework for constructing generative models (diffusion, flow-matching) using continuous and discrete interpolants. Part of BioNeMo Framework.
BioNeMo-NoodlesEfficient genomic data handling with memory-mapped access to FASTA files. Part of BioNeMo Framework.

Training and Customization

BioNeMo provides frameworks and recipes for pretraining, fine-tuning, and adapting biomolecular AI models at scale on GPU infrastructure.

ToolDescription
BioNeMo FrameworkReference training implementations and ready-to-run examples showing how to achieve lower-precision training, maximum scaling & throughput for models like Llama3, ESM2, Evo2, CodonFM, and Geneformer using FSDP and TransformerEngine.
Context Parallelism (boltz-cp)Long-sequence parallelism for protein structure prediction models. Distributes activation tensors across GPUs to overcome single-GPU memory limits for large biomolecules.

Documentation: docs.nvidia.com/bionemo-framework


Optimized Inference and Deployment

BioNeMo NIM microservices are enterprise-ready inference microservices with built-in API endpoints. Each NIM includes algorithmic, system, and runtime optimizations into a prebuilt container — go from zero to inference in minutes.

NIMDescription
OpenFold33D structure prediction for molecular complexes (proteins, DNA, RNA, ligands)
OpenFold2Protein structure prediction from sequence, MSAs, and templates
Boltz-2Biomolecular complex structure prediction
Evo2-40BGenomic foundation model with long-context sequence understanding
MSA SearchMultiple sequence alignment generation from query sequences
ProteinMPNNAmino acid sequence design for protein backbones
RFDiffusionGenerative model for protein backbone and binder design
GenMolFragment-based small molecule generation
DiffDockMolecular blind docking for predicting protein-ligand binding poses
MolMIMMolecular generation optimized for user-defined drug properties

Browse all available NIM microservices: build.nvidia.com/explore/biology

NIM microservices can be deployed self-hosted via Docker or Kubernetes, or on cloud platforms including AWS, Google Cloud, Microsoft Azure, and NVIDIA DGX Cloud.


Workflow Examples and Community Contributions

Application-level examples showing how BioNeMo platform components work together:

Note: If you have an example you'd like to contribute, we'd love to include it. Please get started by opening a GitHub issue and we'll reach out to you.

Pinned Loading

  1. bionemo-recipesbionemo-recipesPublic

    BioNeMo Recipes: For building and adapting AI models in drug discovery at scale

    Python 848 177

  2. Proteina-ComplexaProteina-ComplexaPublic

    Generative model for protein binder design for protein and small molecule targets. Combines a pretrained flow-based generative model (built on La-Proteina) with inference-time optimization.

    Python 421 73

  3. nvMolKitnvMolKitPublic

    A high-performance, GPU-accelerated library for key computational chemistry tasks, such as molecular similarity, conformer generation, and geometry relaxation.

    Cuda 273 27

  4. genmolgenmolPublic

    GenMol is a generative AI model for creating novel molecules. It utilizes masked discrete diffusion and fragment-based generation to create valid molecules, which are encoded in the SAFE molecular …

    Python 199 35

  5. RNAProRNAProPublic

    RNAPro is a state-of-the-art RNA 3D folding model developed in collaboration with the hosts and winners of the Stanford RNA 3D Folding Kaggle competition.

    Python 95 20

  6. KERMTKERMTPublic

    KERMT is a pretrained graph neural network model for molecular property prediction.

    Python 96 17

Repositories

Showing 10 of 23 repositories

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…

Most used topics

Loading…

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' NVIDIA BioNeMo · GitHub
Skip to content

NVIDIA BioNeMo

Welcome to the NVIDIA BioNeMo Github Organization.

The repositories hosted here often accompany our publications and are released open source to the community to accelerate scientific innovation. Unless otherwise noted, the code is provided as-is and is not actively maintained.

Find out more about our research labs at research.nvidia.com/labs/dbr and research.nvidia.com and about our healthcare products at nvidia.com/en-us/industries/healthcare-life-sciences/.

Explore our pinned repositories below to see some of our featured projects!


NVIDIA BioNeMo is an open developer platform for AI-driven life science research.

It provides GPU-accelerated models, tools, and datasets for the entire AI lifecycle, enabling researchers and developers to build, customize, and deploy AI applications that transform physical lab results into the digital insights that drive the next experiment.

The platform is built on five core pillars:

  • Data: Large-scale datasets for training, fine-tuning, and benchmarking models.
  • Models: Open-source models for understanding biological systems, designing novel proteins and small molecules, and optimizing candidates for synthesizability, binding affinity, and molecular properties.
  • Libraries and Tools: Foundational GPU-optimized libraries and kernels for accelerated AI training and inference.
  • Training and Customization: Frameworks and recipes for pretraining, fine-tuning, and adapting models for specialized use cases.
  • Optimized Inference and Deployment: Enterprise-ready NVIDIA inference microservices (NIM) and reference architectures for production use.

Note: Many components of the BioNeMo platform are modular and hosted in their own dedicated GitHub repositories or organizations. This README serves as a central index to guide you to the right tools.

Table of Contents


License

BioNeMo components are generally released under:

Individual components may vary — check each resource for specific license terms.


Data

Unlike natural language models trained on internet-scale data, biology and chemistry lack the critical mass of data required for large, general-purpose foundation models. To address this ecosystem-wide gap, NVIDIA is partnering with leading organizations to create and release open datasets.

DatasetDescription
3D Structures of Protein Complexes
(available through the AlphaFold Database)
Large-scale open database of predicted protein complex structures built with ecosystem partners to accelerate interaction biology and drug discovery. License: CC BY 4.0
Consistency Distilled Synthetic Protein Database455K curated, high-quality protein sequence-structure pairs. Built using ProteinMPNN to generate synthetic sequences for Foldseek AFDB cluster representative structures, then refolded with ESMFold to obtain fully atomistic, self-consistent models. Filtered to pLDDT > 80. License: CC BY 4.0

Models

NVIDIA BioNeMo provides high-quality, fully open-source models — including the full training codebase, pre-trained weights, and research papers — completely free to use. These models are hosted in the NVIDIA-BioNeMo GitHub organization.

These models reflect our active research directions, and we highly encourage community feedback, collaboration, and adaptation to push their capabilities further.

Understand

Use CaseModelDescription
Target Identification / Disease Understanding (RNA)CodonFMCodon-level RNA foundation model trained on 130M protein-coding sequences from 22K+ species. Captures synonymous codon variation for mRNA design, stability modeling, and variant interpretation.
Structure Prediction (RNA)RNAProState-of-the-art RNA 3D structure prediction model. Combines Protenix-based co-folding architectures with RNA foundation models, MSA, and template-based modeling.

Design

Use CaseModelDescription
ProteinsProteina-ComplexaProtein binder design for protein and small molecule targets. Combines a pretrained flow-based generative model (built on La-Proteina) with inference-time optimization for high-quality binder generation.
La-ProteinaAll-atom protein generation using partially latent flow matching. Jointly generates amino acid sequence and full atomistic structure (backbone + side chains) for up to 800 residues. Enables atomistic motif scaffolding for enzyme design.
ProteinaLarge-scale flow-based generative model for protein backbone structures with hierarchical fold-class conditioning and a scalable transformer architecture.
ProtComposerSpatial-layout-conditioned protein structure generation using 3D ellipsoids to control shape and substructure arrangements.
Small MoleculesGenMolFragment-based molecule generation using masked discrete diffusion over SAFE representations. Supports de novo design, scaffold decoration, linker design, motif extension, and lead optimization.
MegalodonTransformer-based 3D molecule generative model using equivariant graph transformer architecture. Generates both 2D topology and 3D structure with physically realistic, low-energy conformations.
AvgFlowEfficient molecular 3D conformer generation using SO(3)-averaged flow-matching and reflow. Architecture-agnostic framework applicable to equivariant and non-equivariant models.

Optimize

Use CaseModelDescription
Property PredictionKERMTPretrained graph neural network for molecular property prediction (ADMET). Multi-task extension of GROVER with accelerated data loading via cuik-molmaker. SOTA on real-world ADMET data.
SynthesizabilityReaSynSynthesis pathway prediction using an encoder-decoder Transformer with Chain-of-Reaction notation. Predicts reaction steps from building blocks to final products, or finds synthesizable analogs for unsynthesizable targets.
Binding EnergyDualBind3D structure-based deep learning model for protein-ligand binding affinity prediction using a dual-loss framework (supervised MSE + unsupervised denoising). Orders of magnitude faster than physics-based FEP methods.

Libraries and Tools

GPU-optimized libraries and tools that integrate into existing workflows. Engineered to be lightweight and specialized for maximum performance without dependency bloat.

TaskToolDescription
Data Processing & AnalysisParabricksGPU-accelerated genomics software suite for rapid secondary analysis of DNA/RNA sequencing data.
nvMolKitGPU-accelerated cheminformatics library for molecular fingerprinting, Tanimoto/cosine similarity, Butina clustering, conformer generation (ETKDGv3), MMFF geometry optimization, and substructure search.
cuik-molmakerMolecular featurization package for converting chemical structures into GNN inputs. Accelerates Chemprop training by 1.6x and inference by 2.4x with 80% memory reduction.
nvQSPGPU-accelerated Quantitative Systems Pharmacology ODE solvers. 77x speedup over CPU for virtual patient simulations with bit-exact FP64 reproducibility.
Training & InferencecuEquivarianceCUDA-X library with optimized kernels for efficient training of geometry-aware equivariant neural networks (AlphaFold-like and molecular structure models).
BioNeMo-SCDLScalable, memory-efficient data loader for training large single-cell models. Part of BioNeMo Framework.
BioNeMo-MoCoFramework for constructing generative models (diffusion, flow-matching) using continuous and discrete interpolants. Part of BioNeMo Framework.
BioNeMo-NoodlesEfficient genomic data handling with memory-mapped access to FASTA files. Part of BioNeMo Framework.

Training and Customization

BioNeMo provides frameworks and recipes for pretraining, fine-tuning, and adapting biomolecular AI models at scale on GPU infrastructure.

ToolDescription
BioNeMo FrameworkReference training implementations and ready-to-run examples showing how to achieve lower-precision training, maximum scaling & throughput for models like Llama3, ESM2, Evo2, CodonFM, and Geneformer using FSDP and TransformerEngine.
Context Parallelism (boltz-cp)Long-sequence parallelism for protein structure prediction models. Distributes activation tensors across GPUs to overcome single-GPU memory limits for large biomolecules.

Documentation: docs.nvidia.com/bionemo-framework


Optimized Inference and Deployment

BioNeMo NIM microservices are enterprise-ready inference microservices with built-in API endpoints. Each NIM includes algorithmic, system, and runtime optimizations into a prebuilt container — go from zero to inference in minutes.

NIMDescription
OpenFold33D structure prediction for molecular complexes (proteins, DNA, RNA, ligands)
OpenFold2Protein structure prediction from sequence, MSAs, and templates
Boltz-2Biomolecular complex structure prediction
Evo2-40BGenomic foundation model with long-context sequence understanding
MSA SearchMultiple sequence alignment generation from query sequences
ProteinMPNNAmino acid sequence design for protein backbones
RFDiffusionGenerative model for protein backbone and binder design
GenMolFragment-based small molecule generation
DiffDockMolecular blind docking for predicting protein-ligand binding poses
MolMIMMolecular generation optimized for user-defined drug properties

Browse all available NIM microservices: build.nvidia.com/explore/biology

NIM microservices can be deployed self-hosted via Docker or Kubernetes, or on cloud platforms including AWS, Google Cloud, Microsoft Azure, and NVIDIA DGX Cloud.


Workflow Examples and Community Contributions

Application-level examples showing how BioNeMo platform components work together:

Note: If you have an example you'd like to contribute, we'd love to include it. Please get started by opening a GitHub issue and we'll reach out to you.

Pinned Loading

  1. bionemo-recipesbionemo-recipesPublic

    BioNeMo Recipes: For building and adapting AI models in drug discovery at scale

    Python 848 177

  2. Proteina-ComplexaProteina-ComplexaPublic

    Generative model for protein binder design for protein and small molecule targets. Combines a pretrained flow-based generative model (built on La-Proteina) with inference-time optimization.

    Python 421 73

  3. nvMolKitnvMolKitPublic

    A high-performance, GPU-accelerated library for key computational chemistry tasks, such as molecular similarity, conformer generation, and geometry relaxation.

    Cuda 273 27

  4. genmolgenmolPublic

    GenMol is a generative AI model for creating novel molecules. It utilizes masked discrete diffusion and fragment-based generation to create valid molecules, which are encoded in the SAFE molecular …

    Python 199 35

  5. RNAProRNAProPublic

    RNAPro is a state-of-the-art RNA 3D folding model developed in collaboration with the hosts and winners of the Stanford RNA 3D Folding Kaggle competition.

    Python 95 20

  6. KERMTKERMTPublic

    KERMT is a pretrained graph neural network model for molecular property prediction.

    Python 96 17

Repositories

Showing 10 of 23 repositories

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…

Most used topics

Loading…

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' NVIDIA BioNeMo · GitHub
Skip to content

NVIDIA BioNeMo

Welcome to the NVIDIA BioNeMo Github Organization.

The repositories hosted here often accompany our publications and are released open source to the community to accelerate scientific innovation. Unless otherwise noted, the code is provided as-is and is not actively maintained.

Find out more about our research labs at research.nvidia.com/labs/dbr and research.nvidia.com and about our healthcare products at nvidia.com/en-us/industries/healthcare-life-sciences/.

Explore our pinned repositories below to see some of our featured projects!


NVIDIA BioNeMo is an open developer platform for AI-driven life science research.

It provides GPU-accelerated models, tools, and datasets for the entire AI lifecycle, enabling researchers and developers to build, customize, and deploy AI applications that transform physical lab results into the digital insights that drive the next experiment.

The platform is built on five core pillars:

  • Data: Large-scale datasets for training, fine-tuning, and benchmarking models.
  • Models: Open-source models for understanding biological systems, designing novel proteins and small molecules, and optimizing candidates for synthesizability, binding affinity, and molecular properties.
  • Libraries and Tools: Foundational GPU-optimized libraries and kernels for accelerated AI training and inference.
  • Training and Customization: Frameworks and recipes for pretraining, fine-tuning, and adapting models for specialized use cases.
  • Optimized Inference and Deployment: Enterprise-ready NVIDIA inference microservices (NIM) and reference architectures for production use.

Note: Many components of the BioNeMo platform are modular and hosted in their own dedicated GitHub repositories or organizations. This README serves as a central index to guide you to the right tools.

Table of Contents


License

BioNeMo components are generally released under:

Individual components may vary — check each resource for specific license terms.


Data

Unlike natural language models trained on internet-scale data, biology and chemistry lack the critical mass of data required for large, general-purpose foundation models. To address this ecosystem-wide gap, NVIDIA is partnering with leading organizations to create and release open datasets.

DatasetDescription
3D Structures of Protein Complexes
(available through the AlphaFold Database)
Large-scale open database of predicted protein complex structures built with ecosystem partners to accelerate interaction biology and drug discovery. License: CC BY 4.0
Consistency Distilled Synthetic Protein Database455K curated, high-quality protein sequence-structure pairs. Built using ProteinMPNN to generate synthetic sequences for Foldseek AFDB cluster representative structures, then refolded with ESMFold to obtain fully atomistic, self-consistent models. Filtered to pLDDT > 80. License: CC BY 4.0

Models

NVIDIA BioNeMo provides high-quality, fully open-source models — including the full training codebase, pre-trained weights, and research papers — completely free to use. These models are hosted in the NVIDIA-BioNeMo GitHub organization.

These models reflect our active research directions, and we highly encourage community feedback, collaboration, and adaptation to push their capabilities further.

Understand

Use CaseModelDescription
Target Identification / Disease Understanding (RNA)CodonFMCodon-level RNA foundation model trained on 130M protein-coding sequences from 22K+ species. Captures synonymous codon variation for mRNA design, stability modeling, and variant interpretation.
Structure Prediction (RNA)RNAProState-of-the-art RNA 3D structure prediction model. Combines Protenix-based co-folding architectures with RNA foundation models, MSA, and template-based modeling.

Design

Use CaseModelDescription
ProteinsProteina-ComplexaProtein binder design for protein and small molecule targets. Combines a pretrained flow-based generative model (built on La-Proteina) with inference-time optimization for high-quality binder generation.
La-ProteinaAll-atom protein generation using partially latent flow matching. Jointly generates amino acid sequence and full atomistic structure (backbone + side chains) for up to 800 residues. Enables atomistic motif scaffolding for enzyme design.
ProteinaLarge-scale flow-based generative model for protein backbone structures with hierarchical fold-class conditioning and a scalable transformer architecture.
ProtComposerSpatial-layout-conditioned protein structure generation using 3D ellipsoids to control shape and substructure arrangements.
Small MoleculesGenMolFragment-based molecule generation using masked discrete diffusion over SAFE representations. Supports de novo design, scaffold decoration, linker design, motif extension, and lead optimization.
MegalodonTransformer-based 3D molecule generative model using equivariant graph transformer architecture. Generates both 2D topology and 3D structure with physically realistic, low-energy conformations.
AvgFlowEfficient molecular 3D conformer generation using SO(3)-averaged flow-matching and reflow. Architecture-agnostic framework applicable to equivariant and non-equivariant models.

Optimize

Use CaseModelDescription
Property PredictionKERMTPretrained graph neural network for molecular property prediction (ADMET). Multi-task extension of GROVER with accelerated data loading via cuik-molmaker. SOTA on real-world ADMET data.
SynthesizabilityReaSynSynthesis pathway prediction using an encoder-decoder Transformer with Chain-of-Reaction notation. Predicts reaction steps from building blocks to final products, or finds synthesizable analogs for unsynthesizable targets.
Binding EnergyDualBind3D structure-based deep learning model for protein-ligand binding affinity prediction using a dual-loss framework (supervised MSE + unsupervised denoising). Orders of magnitude faster than physics-based FEP methods.

Libraries and Tools

GPU-optimized libraries and tools that integrate into existing workflows. Engineered to be lightweight and specialized for maximum performance without dependency bloat.

TaskToolDescription
Data Processing & AnalysisParabricksGPU-accelerated genomics software suite for rapid secondary analysis of DNA/RNA sequencing data.
nvMolKitGPU-accelerated cheminformatics library for molecular fingerprinting, Tanimoto/cosine similarity, Butina clustering, conformer generation (ETKDGv3), MMFF geometry optimization, and substructure search.
cuik-molmakerMolecular featurization package for converting chemical structures into GNN inputs. Accelerates Chemprop training by 1.6x and inference by 2.4x with 80% memory reduction.
nvQSPGPU-accelerated Quantitative Systems Pharmacology ODE solvers. 77x speedup over CPU for virtual patient simulations with bit-exact FP64 reproducibility.
Training & InferencecuEquivarianceCUDA-X library with optimized kernels for efficient training of geometry-aware equivariant neural networks (AlphaFold-like and molecular structure models).
BioNeMo-SCDLScalable, memory-efficient data loader for training large single-cell models. Part of BioNeMo Framework.
BioNeMo-MoCoFramework for constructing generative models (diffusion, flow-matching) using continuous and discrete interpolants. Part of BioNeMo Framework.
BioNeMo-NoodlesEfficient genomic data handling with memory-mapped access to FASTA files. Part of BioNeMo Framework.

Training and Customization

BioNeMo provides frameworks and recipes for pretraining, fine-tuning, and adapting biomolecular AI models at scale on GPU infrastructure.

ToolDescription
BioNeMo FrameworkReference training implementations and ready-to-run examples showing how to achieve lower-precision training, maximum scaling & throughput for models like Llama3, ESM2, Evo2, CodonFM, and Geneformer using FSDP and TransformerEngine.
Context Parallelism (boltz-cp)Long-sequence parallelism for protein structure prediction models. Distributes activation tensors across GPUs to overcome single-GPU memory limits for large biomolecules.

Documentation: docs.nvidia.com/bionemo-framework


Optimized Inference and Deployment

BioNeMo NIM microservices are enterprise-ready inference microservices with built-in API endpoints. Each NIM includes algorithmic, system, and runtime optimizations into a prebuilt container — go from zero to inference in minutes.

NIMDescription
OpenFold33D structure prediction for molecular complexes (proteins, DNA, RNA, ligands)
OpenFold2Protein structure prediction from sequence, MSAs, and templates
Boltz-2Biomolecular complex structure prediction
Evo2-40BGenomic foundation model with long-context sequence understanding
MSA SearchMultiple sequence alignment generation from query sequences
ProteinMPNNAmino acid sequence design for protein backbones
RFDiffusionGenerative model for protein backbone and binder design
GenMolFragment-based small molecule generation
DiffDockMolecular blind docking for predicting protein-ligand binding poses
MolMIMMolecular generation optimized for user-defined drug properties

Browse all available NIM microservices: build.nvidia.com/explore/biology

NIM microservices can be deployed self-hosted via Docker or Kubernetes, or on cloud platforms including AWS, Google Cloud, Microsoft Azure, and NVIDIA DGX Cloud.


Workflow Examples and Community Contributions

Application-level examples showing how BioNeMo platform components work together:

Note: If you have an example you'd like to contribute, we'd love to include it. Please get started by opening a GitHub issue and we'll reach out to you.

Pinned Loading

  1. bionemo-recipesbionemo-recipesPublic

    BioNeMo Recipes: For building and adapting AI models in drug discovery at scale

    Python 848 177

  2. Proteina-ComplexaProteina-ComplexaPublic

    Generative model for protein binder design for protein and small molecule targets. Combines a pretrained flow-based generative model (built on La-Proteina) with inference-time optimization.

    Python 421 73

  3. nvMolKitnvMolKitPublic

    A high-performance, GPU-accelerated library for key computational chemistry tasks, such as molecular similarity, conformer generation, and geometry relaxation.

    Cuda 273 27

  4. genmolgenmolPublic

    GenMol is a generative AI model for creating novel molecules. It utilizes masked discrete diffusion and fragment-based generation to create valid molecules, which are encoded in the SAFE molecular …

    Python 199 35

  5. RNAProRNAProPublic

    RNAPro is a state-of-the-art RNA 3D folding model developed in collaboration with the hosts and winners of the Stanford RNA 3D Folding Kaggle competition.

    Python 95 20

  6. KERMTKERMTPublic

    KERMT is a pretrained graph neural network model for molecular property prediction.

    Python 96 17

Repositories

Showing 10 of 23 repositories

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…

Most used topics

Loading…

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' NVIDIA BioNeMo · GitHub
Skip to content

NVIDIA BioNeMo

Welcome to the NVIDIA BioNeMo Github Organization.

The repositories hosted here often accompany our publications and are released open source to the community to accelerate scientific innovation. Unless otherwise noted, the code is provided as-is and is not actively maintained.

Find out more about our research labs at research.nvidia.com/labs/dbr and research.nvidia.com and about our healthcare products at nvidia.com/en-us/industries/healthcare-life-sciences/.

Explore our pinned repositories below to see some of our featured projects!


NVIDIA BioNeMo is an open developer platform for AI-driven life science research.

It provides GPU-accelerated models, tools, and datasets for the entire AI lifecycle, enabling researchers and developers to build, customize, and deploy AI applications that transform physical lab results into the digital insights that drive the next experiment.

The platform is built on five core pillars:

  • Data: Large-scale datasets for training, fine-tuning, and benchmarking models.
  • Models: Open-source models for understanding biological systems, designing novel proteins and small molecules, and optimizing candidates for synthesizability, binding affinity, and molecular properties.
  • Libraries and Tools: Foundational GPU-optimized libraries and kernels for accelerated AI training and inference.
  • Training and Customization: Frameworks and recipes for pretraining, fine-tuning, and adapting models for specialized use cases.
  • Optimized Inference and Deployment: Enterprise-ready NVIDIA inference microservices (NIM) and reference architectures for production use.

Note: Many components of the BioNeMo platform are modular and hosted in their own dedicated GitHub repositories or organizations. This README serves as a central index to guide you to the right tools.

Table of Contents


License

BioNeMo components are generally released under:

Individual components may vary — check each resource for specific license terms.


Data

Unlike natural language models trained on internet-scale data, biology and chemistry lack the critical mass of data required for large, general-purpose foundation models. To address this ecosystem-wide gap, NVIDIA is partnering with leading organizations to create and release open datasets.

DatasetDescription
3D Structures of Protein Complexes
(available through the AlphaFold Database)
Large-scale open database of predicted protein complex structures built with ecosystem partners to accelerate interaction biology and drug discovery. License: CC BY 4.0
Consistency Distilled Synthetic Protein Database455K curated, high-quality protein sequence-structure pairs. Built using ProteinMPNN to generate synthetic sequences for Foldseek AFDB cluster representative structures, then refolded with ESMFold to obtain fully atomistic, self-consistent models. Filtered to pLDDT > 80. License: CC BY 4.0

Models

NVIDIA BioNeMo provides high-quality, fully open-source models — including the full training codebase, pre-trained weights, and research papers — completely free to use. These models are hosted in the NVIDIA-BioNeMo GitHub organization.

These models reflect our active research directions, and we highly encourage community feedback, collaboration, and adaptation to push their capabilities further.

Understand

Use CaseModelDescription
Target Identification / Disease Understanding (RNA)CodonFMCodon-level RNA foundation model trained on 130M protein-coding sequences from 22K+ species. Captures synonymous codon variation for mRNA design, stability modeling, and variant interpretation.
Structure Prediction (RNA)RNAProState-of-the-art RNA 3D structure prediction model. Combines Protenix-based co-folding architectures with RNA foundation models, MSA, and template-based modeling.

Design

Use CaseModelDescription
ProteinsProteina-ComplexaProtein binder design for protein and small molecule targets. Combines a pretrained flow-based generative model (built on La-Proteina) with inference-time optimization for high-quality binder generation.
La-ProteinaAll-atom protein generation using partially latent flow matching. Jointly generates amino acid sequence and full atomistic structure (backbone + side chains) for up to 800 residues. Enables atomistic motif scaffolding for enzyme design.
ProteinaLarge-scale flow-based generative model for protein backbone structures with hierarchical fold-class conditioning and a scalable transformer architecture.
ProtComposerSpatial-layout-conditioned protein structure generation using 3D ellipsoids to control shape and substructure arrangements.
Small MoleculesGenMolFragment-based molecule generation using masked discrete diffusion over SAFE representations. Supports de novo design, scaffold decoration, linker design, motif extension, and lead optimization.
MegalodonTransformer-based 3D molecule generative model using equivariant graph transformer architecture. Generates both 2D topology and 3D structure with physically realistic, low-energy conformations.
AvgFlowEfficient molecular 3D conformer generation using SO(3)-averaged flow-matching and reflow. Architecture-agnostic framework applicable to equivariant and non-equivariant models.

Optimize

Use CaseModelDescription
Property PredictionKERMTPretrained graph neural network for molecular property prediction (ADMET). Multi-task extension of GROVER with accelerated data loading via cuik-molmaker. SOTA on real-world ADMET data.
SynthesizabilityReaSynSynthesis pathway prediction using an encoder-decoder Transformer with Chain-of-Reaction notation. Predicts reaction steps from building blocks to final products, or finds synthesizable analogs for unsynthesizable targets.
Binding EnergyDualBind3D structure-based deep learning model for protein-ligand binding affinity prediction using a dual-loss framework (supervised MSE + unsupervised denoising). Orders of magnitude faster than physics-based FEP methods.

Libraries and Tools

GPU-optimized libraries and tools that integrate into existing workflows. Engineered to be lightweight and specialized for maximum performance without dependency bloat.

TaskToolDescription
Data Processing & AnalysisParabricksGPU-accelerated genomics software suite for rapid secondary analysis of DNA/RNA sequencing data.
nvMolKitGPU-accelerated cheminformatics library for molecular fingerprinting, Tanimoto/cosine similarity, Butina clustering, conformer generation (ETKDGv3), MMFF geometry optimization, and substructure search.
cuik-molmakerMolecular featurization package for converting chemical structures into GNN inputs. Accelerates Chemprop training by 1.6x and inference by 2.4x with 80% memory reduction.
nvQSPGPU-accelerated Quantitative Systems Pharmacology ODE solvers. 77x speedup over CPU for virtual patient simulations with bit-exact FP64 reproducibility.
Training & InferencecuEquivarianceCUDA-X library with optimized kernels for efficient training of geometry-aware equivariant neural networks (AlphaFold-like and molecular structure models).
BioNeMo-SCDLScalable, memory-efficient data loader for training large single-cell models. Part of BioNeMo Framework.
BioNeMo-MoCoFramework for constructing generative models (diffusion, flow-matching) using continuous and discrete interpolants. Part of BioNeMo Framework.
BioNeMo-NoodlesEfficient genomic data handling with memory-mapped access to FASTA files. Part of BioNeMo Framework.

Training and Customization

BioNeMo provides frameworks and recipes for pretraining, fine-tuning, and adapting biomolecular AI models at scale on GPU infrastructure.

ToolDescription
BioNeMo FrameworkReference training implementations and ready-to-run examples showing how to achieve lower-precision training, maximum scaling & throughput for models like Llama3, ESM2, Evo2, CodonFM, and Geneformer using FSDP and TransformerEngine.
Context Parallelism (boltz-cp)Long-sequence parallelism for protein structure prediction models. Distributes activation tensors across GPUs to overcome single-GPU memory limits for large biomolecules.

Documentation: docs.nvidia.com/bionemo-framework


Optimized Inference and Deployment

BioNeMo NIM microservices are enterprise-ready inference microservices with built-in API endpoints. Each NIM includes algorithmic, system, and runtime optimizations into a prebuilt container — go from zero to inference in minutes.

NIMDescription
OpenFold33D structure prediction for molecular complexes (proteins, DNA, RNA, ligands)
OpenFold2Protein structure prediction from sequence, MSAs, and templates
Boltz-2Biomolecular complex structure prediction
Evo2-40BGenomic foundation model with long-context sequence understanding
MSA SearchMultiple sequence alignment generation from query sequences
ProteinMPNNAmino acid sequence design for protein backbones
RFDiffusionGenerative model for protein backbone and binder design
GenMolFragment-based small molecule generation
DiffDockMolecular blind docking for predicting protein-ligand binding poses
MolMIMMolecular generation optimized for user-defined drug properties

Browse all available NIM microservices: build.nvidia.com/explore/biology

NIM microservices can be deployed self-hosted via Docker or Kubernetes, or on cloud platforms including AWS, Google Cloud, Microsoft Azure, and NVIDIA DGX Cloud.


Workflow Examples and Community Contributions

Application-level examples showing how BioNeMo platform components work together:

Note: If you have an example you'd like to contribute, we'd love to include it. Please get started by opening a GitHub issue and we'll reach out to you.

Pinned Loading

  1. bionemo-recipesbionemo-recipesPublic

    BioNeMo Recipes: For building and adapting AI models in drug discovery at scale

    Python 848 177

  2. Proteina-ComplexaProteina-ComplexaPublic

    Generative model for protein binder design for protein and small molecule targets. Combines a pretrained flow-based generative model (built on La-Proteina) with inference-time optimization.

    Python 421 73

  3. nvMolKitnvMolKitPublic

    A high-performance, GPU-accelerated library for key computational chemistry tasks, such as molecular similarity, conformer generation, and geometry relaxation.

    Cuda 273 27

  4. genmolgenmolPublic

    GenMol is a generative AI model for creating novel molecules. It utilizes masked discrete diffusion and fragment-based generation to create valid molecules, which are encoded in the SAFE molecular …

    Python 199 35

  5. RNAProRNAProPublic

    RNAPro is a state-of-the-art RNA 3D folding model developed in collaboration with the hosts and winners of the Stanford RNA 3D Folding Kaggle competition.

    Python 95 20

  6. KERMTKERMTPublic

    KERMT is a pretrained graph neural network model for molecular property prediction.

    Python 96 17

Repositories

Showing 10 of 23 repositories

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…

Most used topics

Loading…

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' NVIDIA BioNeMo · GitHub
Skip to content

NVIDIA BioNeMo

Welcome to the NVIDIA BioNeMo Github Organization.

The repositories hosted here often accompany our publications and are released open source to the community to accelerate scientific innovation. Unless otherwise noted, the code is provided as-is and is not actively maintained.

Find out more about our research labs at research.nvidia.com/labs/dbr and research.nvidia.com and about our healthcare products at nvidia.com/en-us/industries/healthcare-life-sciences/.

Explore our pinned repositories below to see some of our featured projects!


NVIDIA BioNeMo is an open developer platform for AI-driven life science research.

It provides GPU-accelerated models, tools, and datasets for the entire AI lifecycle, enabling researchers and developers to build, customize, and deploy AI applications that transform physical lab results into the digital insights that drive the next experiment.

The platform is built on five core pillars:

  • Data: Large-scale datasets for training, fine-tuning, and benchmarking models.
  • Models: Open-source models for understanding biological systems, designing novel proteins and small molecules, and optimizing candidates for synthesizability, binding affinity, and molecular properties.
  • Libraries and Tools: Foundational GPU-optimized libraries and kernels for accelerated AI training and inference.
  • Training and Customization: Frameworks and recipes for pretraining, fine-tuning, and adapting models for specialized use cases.
  • Optimized Inference and Deployment: Enterprise-ready NVIDIA inference microservices (NIM) and reference architectures for production use.

Note: Many components of the BioNeMo platform are modular and hosted in their own dedicated GitHub repositories or organizations. This README serves as a central index to guide you to the right tools.

Table of Contents


License

BioNeMo components are generally released under:

Individual components may vary — check each resource for specific license terms.


Data

Unlike natural language models trained on internet-scale data, biology and chemistry lack the critical mass of data required for large, general-purpose foundation models. To address this ecosystem-wide gap, NVIDIA is partnering with leading organizations to create and release open datasets.

DatasetDescription
3D Structures of Protein Complexes
(available through the AlphaFold Database)
Large-scale open database of predicted protein complex structures built with ecosystem partners to accelerate interaction biology and drug discovery. License: CC BY 4.0
Consistency Distilled Synthetic Protein Database455K curated, high-quality protein sequence-structure pairs. Built using ProteinMPNN to generate synthetic sequences for Foldseek AFDB cluster representative structures, then refolded with ESMFold to obtain fully atomistic, self-consistent models. Filtered to pLDDT > 80. License: CC BY 4.0

Models

NVIDIA BioNeMo provides high-quality, fully open-source models — including the full training codebase, pre-trained weights, and research papers — completely free to use. These models are hosted in the NVIDIA-BioNeMo GitHub organization.

These models reflect our active research directions, and we highly encourage community feedback, collaboration, and adaptation to push their capabilities further.

Understand

Use CaseModelDescription
Target Identification / Disease Understanding (RNA)CodonFMCodon-level RNA foundation model trained on 130M protein-coding sequences from 22K+ species. Captures synonymous codon variation for mRNA design, stability modeling, and variant interpretation.
Structure Prediction (RNA)RNAProState-of-the-art RNA 3D structure prediction model. Combines Protenix-based co-folding architectures with RNA foundation models, MSA, and template-based modeling.

Design

Use CaseModelDescription
ProteinsProteina-ComplexaProtein binder design for protein and small molecule targets. Combines a pretrained flow-based generative model (built on La-Proteina) with inference-time optimization for high-quality binder generation.
La-ProteinaAll-atom protein generation using partially latent flow matching. Jointly generates amino acid sequence and full atomistic structure (backbone + side chains) for up to 800 residues. Enables atomistic motif scaffolding for enzyme design.
ProteinaLarge-scale flow-based generative model for protein backbone structures with hierarchical fold-class conditioning and a scalable transformer architecture.
ProtComposerSpatial-layout-conditioned protein structure generation using 3D ellipsoids to control shape and substructure arrangements.
Small MoleculesGenMolFragment-based molecule generation using masked discrete diffusion over SAFE representations. Supports de novo design, scaffold decoration, linker design, motif extension, and lead optimization.
MegalodonTransformer-based 3D molecule generative model using equivariant graph transformer architecture. Generates both 2D topology and 3D structure with physically realistic, low-energy conformations.
AvgFlowEfficient molecular 3D conformer generation using SO(3)-averaged flow-matching and reflow. Architecture-agnostic framework applicable to equivariant and non-equivariant models.

Optimize

Use CaseModelDescription
Property PredictionKERMTPretrained graph neural network for molecular property prediction (ADMET). Multi-task extension of GROVER with accelerated data loading via cuik-molmaker. SOTA on real-world ADMET data.
SynthesizabilityReaSynSynthesis pathway prediction using an encoder-decoder Transformer with Chain-of-Reaction notation. Predicts reaction steps from building blocks to final products, or finds synthesizable analogs for unsynthesizable targets.
Binding EnergyDualBind3D structure-based deep learning model for protein-ligand binding affinity prediction using a dual-loss framework (supervised MSE + unsupervised denoising). Orders of magnitude faster than physics-based FEP methods.

Libraries and Tools

GPU-optimized libraries and tools that integrate into existing workflows. Engineered to be lightweight and specialized for maximum performance without dependency bloat.

TaskToolDescription
Data Processing & AnalysisParabricksGPU-accelerated genomics software suite for rapid secondary analysis of DNA/RNA sequencing data.
nvMolKitGPU-accelerated cheminformatics library for molecular fingerprinting, Tanimoto/cosine similarity, Butina clustering, conformer generation (ETKDGv3), MMFF geometry optimization, and substructure search.
cuik-molmakerMolecular featurization package for converting chemical structures into GNN inputs. Accelerates Chemprop training by 1.6x and inference by 2.4x with 80% memory reduction.
nvQSPGPU-accelerated Quantitative Systems Pharmacology ODE solvers. 77x speedup over CPU for virtual patient simulations with bit-exact FP64 reproducibility.
Training & InferencecuEquivarianceCUDA-X library with optimized kernels for efficient training of geometry-aware equivariant neural networks (AlphaFold-like and molecular structure models).
BioNeMo-SCDLScalable, memory-efficient data loader for training large single-cell models. Part of BioNeMo Framework.
BioNeMo-MoCoFramework for constructing generative models (diffusion, flow-matching) using continuous and discrete interpolants. Part of BioNeMo Framework.
BioNeMo-NoodlesEfficient genomic data handling with memory-mapped access to FASTA files. Part of BioNeMo Framework.

Training and Customization

BioNeMo provides frameworks and recipes for pretraining, fine-tuning, and adapting biomolecular AI models at scale on GPU infrastructure.

ToolDescription
BioNeMo FrameworkReference training implementations and ready-to-run examples showing how to achieve lower-precision training, maximum scaling & throughput for models like Llama3, ESM2, Evo2, CodonFM, and Geneformer using FSDP and TransformerEngine.
Context Parallelism (boltz-cp)Long-sequence parallelism for protein structure prediction models. Distributes activation tensors across GPUs to overcome single-GPU memory limits for large biomolecules.

Documentation: docs.nvidia.com/bionemo-framework


Optimized Inference and Deployment

BioNeMo NIM microservices are enterprise-ready inference microservices with built-in API endpoints. Each NIM includes algorithmic, system, and runtime optimizations into a prebuilt container — go from zero to inference in minutes.

NIMDescription
OpenFold33D structure prediction for molecular complexes (proteins, DNA, RNA, ligands)
OpenFold2Protein structure prediction from sequence, MSAs, and templates
Boltz-2Biomolecular complex structure prediction
Evo2-40BGenomic foundation model with long-context sequence understanding
MSA SearchMultiple sequence alignment generation from query sequences
ProteinMPNNAmino acid sequence design for protein backbones
RFDiffusionGenerative model for protein backbone and binder design
GenMolFragment-based small molecule generation
DiffDockMolecular blind docking for predicting protein-ligand binding poses
MolMIMMolecular generation optimized for user-defined drug properties

Browse all available NIM microservices: build.nvidia.com/explore/biology

NIM microservices can be deployed self-hosted via Docker or Kubernetes, or on cloud platforms including AWS, Google Cloud, Microsoft Azure, and NVIDIA DGX Cloud.


Workflow Examples and Community Contributions

Application-level examples showing how BioNeMo platform components work together:

Note: If you have an example you'd like to contribute, we'd love to include it. Please get started by opening a GitHub issue and we'll reach out to you.

Pinned Loading

  1. bionemo-recipesbionemo-recipesPublic

    BioNeMo Recipes: For building and adapting AI models in drug discovery at scale

    Python 848 177

  2. Proteina-ComplexaProteina-ComplexaPublic

    Generative model for protein binder design for protein and small molecule targets. Combines a pretrained flow-based generative model (built on La-Proteina) with inference-time optimization.

    Python 421 73

  3. nvMolKitnvMolKitPublic

    A high-performance, GPU-accelerated library for key computational chemistry tasks, such as molecular similarity, conformer generation, and geometry relaxation.

    Cuda 273 27

  4. genmolgenmolPublic

    GenMol is a generative AI model for creating novel molecules. It utilizes masked discrete diffusion and fragment-based generation to create valid molecules, which are encoded in the SAFE molecular …

    Python 199 35

  5. RNAProRNAProPublic

    RNAPro is a state-of-the-art RNA 3D folding model developed in collaboration with the hosts and winners of the Stanford RNA 3D Folding Kaggle competition.

    Python 95 20

  6. KERMTKERMTPublic

    KERMT is a pretrained graph neural network model for molecular property prediction.

    Python 96 17

Repositories

Showing 10 of 23 repositories

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…

Most used topics

Loading…

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); NVIDIA BioNeMo · GitHub
Skip to content

NVIDIA BioNeMo

Welcome to the NVIDIA BioNeMo Github Organization.

The repositories hosted here often accompany our publications and are released open source to the community to accelerate scientific innovation. Unless otherwise noted, the code is provided as-is and is not actively maintained.

Find out more about our research labs at research.nvidia.com/labs/dbr and research.nvidia.com and about our healthcare products at nvidia.com/en-us/industries/healthcare-life-sciences/.

Explore our pinned repositories below to see some of our featured projects!


NVIDIA BioNeMo is an open developer platform for AI-driven life science research.

It provides GPU-accelerated models, tools, and datasets for the entire AI lifecycle, enabling researchers and developers to build, customize, and deploy AI applications that transform physical lab results into the digital insights that drive the next experiment.

The platform is built on five core pillars:

  • Data: Large-scale datasets for training, fine-tuning, and benchmarking models.
  • Models: Open-source models for understanding biological systems, designing novel proteins and small molecules, and optimizing candidates for synthesizability, binding affinity, and molecular properties.
  • Libraries and Tools: Foundational GPU-optimized libraries and kernels for accelerated AI training and inference.
  • Training and Customization: Frameworks and recipes for pretraining, fine-tuning, and adapting models for specialized use cases.
  • Optimized Inference and Deployment: Enterprise-ready NVIDIA inference microservices (NIM) and reference architectures for production use.

Note: Many components of the BioNeMo platform are modular and hosted in their own dedicated GitHub repositories or organizations. This README serves as a central index to guide you to the right tools.

Table of Contents


License

BioNeMo components are generally released under:

Individual components may vary — check each resource for specific license terms.


Data

Unlike natural language models trained on internet-scale data, biology and chemistry lack the critical mass of data required for large, general-purpose foundation models. To address this ecosystem-wide gap, NVIDIA is partnering with leading organizations to create and release open datasets.

DatasetDescription
3D Structures of Protein Complexes
(available through the AlphaFold Database)
Large-scale open database of predicted protein complex structures built with ecosystem partners to accelerate interaction biology and drug discovery. License: CC BY 4.0
Consistency Distilled Synthetic Protein Database455K curated, high-quality protein sequence-structure pairs. Built using ProteinMPNN to generate synthetic sequences for Foldseek AFDB cluster representative structures, then refolded with ESMFold to obtain fully atomistic, self-consistent models. Filtered to pLDDT > 80. License: CC BY 4.0

Models

NVIDIA BioNeMo provides high-quality, fully open-source models — including the full training codebase, pre-trained weights, and research papers — completely free to use. These models are hosted in the NVIDIA-BioNeMo GitHub organization.

These models reflect our active research directions, and we highly encourage community feedback, collaboration, and adaptation to push their capabilities further.

Understand

Use CaseModelDescription
Target Identification / Disease Understanding (RNA)CodonFMCodon-level RNA foundation model trained on 130M protein-coding sequences from 22K+ species. Captures synonymous codon variation for mRNA design, stability modeling, and variant interpretation.
Structure Prediction (RNA)RNAProState-of-the-art RNA 3D structure prediction model. Combines Protenix-based co-folding architectures with RNA foundation models, MSA, and template-based modeling.

Design

Use CaseModelDescription
ProteinsProteina-ComplexaProtein binder design for protein and small molecule targets. Combines a pretrained flow-based generative model (built on La-Proteina) with inference-time optimization for high-quality binder generation.
La-ProteinaAll-atom protein generation using partially latent flow matching. Jointly generates amino acid sequence and full atomistic structure (backbone + side chains) for up to 800 residues. Enables atomistic motif scaffolding for enzyme design.
ProteinaLarge-scale flow-based generative model for protein backbone structures with hierarchical fold-class conditioning and a scalable transformer architecture.
ProtComposerSpatial-layout-conditioned protein structure generation using 3D ellipsoids to control shape and substructure arrangements.
Small MoleculesGenMolFragment-based molecule generation using masked discrete diffusion over SAFE representations. Supports de novo design, scaffold decoration, linker design, motif extension, and lead optimization.
MegalodonTransformer-based 3D molecule generative model using equivariant graph transformer architecture. Generates both 2D topology and 3D structure with physically realistic, low-energy conformations.
AvgFlowEfficient molecular 3D conformer generation using SO(3)-averaged flow-matching and reflow. Architecture-agnostic framework applicable to equivariant and non-equivariant models.

Optimize

Use CaseModelDescription
Property PredictionKERMTPretrained graph neural network for molecular property prediction (ADMET). Multi-task extension of GROVER with accelerated data loading via cuik-molmaker. SOTA on real-world ADMET data.
SynthesizabilityReaSynSynthesis pathway prediction using an encoder-decoder Transformer with Chain-of-Reaction notation. Predicts reaction steps from building blocks to final products, or finds synthesizable analogs for unsynthesizable targets.
Binding EnergyDualBind3D structure-based deep learning model for protein-ligand binding affinity prediction using a dual-loss framework (supervised MSE + unsupervised denoising). Orders of magnitude faster than physics-based FEP methods.

Libraries and Tools

GPU-optimized libraries and tools that integrate into existing workflows. Engineered to be lightweight and specialized for maximum performance without dependency bloat.

TaskToolDescription
Data Processing & AnalysisParabricksGPU-accelerated genomics software suite for rapid secondary analysis of DNA/RNA sequencing data.
nvMolKitGPU-accelerated cheminformatics library for molecular fingerprinting, Tanimoto/cosine similarity, Butina clustering, conformer generation (ETKDGv3), MMFF geometry optimization, and substructure search.
cuik-molmakerMolecular featurization package for converting chemical structures into GNN inputs. Accelerates Chemprop training by 1.6x and inference by 2.4x with 80% memory reduction.
nvQSPGPU-accelerated Quantitative Systems Pharmacology ODE solvers. 77x speedup over CPU for virtual patient simulations with bit-exact FP64 reproducibility.
Training & InferencecuEquivarianceCUDA-X library with optimized kernels for efficient training of geometry-aware equivariant neural networks (AlphaFold-like and molecular structure models).
BioNeMo-SCDLScalable, memory-efficient data loader for training large single-cell models. Part of BioNeMo Framework.
BioNeMo-MoCoFramework for constructing generative models (diffusion, flow-matching) using continuous and discrete interpolants. Part of BioNeMo Framework.
BioNeMo-NoodlesEfficient genomic data handling with memory-mapped access to FASTA files. Part of BioNeMo Framework.

Training and Customization

BioNeMo provides frameworks and recipes for pretraining, fine-tuning, and adapting biomolecular AI models at scale on GPU infrastructure.

ToolDescription
BioNeMo FrameworkReference training implementations and ready-to-run examples showing how to achieve lower-precision training, maximum scaling & throughput for models like Llama3, ESM2, Evo2, CodonFM, and Geneformer using FSDP and TransformerEngine.
Context Parallelism (boltz-cp)Long-sequence parallelism for protein structure prediction models. Distributes activation tensors across GPUs to overcome single-GPU memory limits for large biomolecules.

Documentation: docs.nvidia.com/bionemo-framework


Optimized Inference and Deployment

BioNeMo NIM microservices are enterprise-ready inference microservices with built-in API endpoints. Each NIM includes algorithmic, system, and runtime optimizations into a prebuilt container — go from zero to inference in minutes.

NIMDescription
OpenFold33D structure prediction for molecular complexes (proteins, DNA, RNA, ligands)
OpenFold2Protein structure prediction from sequence, MSAs, and templates
Boltz-2Biomolecular complex structure prediction
Evo2-40BGenomic foundation model with long-context sequence understanding
MSA SearchMultiple sequence alignment generation from query sequences
ProteinMPNNAmino acid sequence design for protein backbones
RFDiffusionGenerative model for protein backbone and binder design
GenMolFragment-based small molecule generation
DiffDockMolecular blind docking for predicting protein-ligand binding poses
MolMIMMolecular generation optimized for user-defined drug properties

Browse all available NIM microservices: build.nvidia.com/explore/biology

NIM microservices can be deployed self-hosted via Docker or Kubernetes, or on cloud platforms including AWS, Google Cloud, Microsoft Azure, and NVIDIA DGX Cloud.


Workflow Examples and Community Contributions

Application-level examples showing how BioNeMo platform components work together:

Note: If you have an example you'd like to contribute, we'd love to include it. Please get started by opening a GitHub issue and we'll reach out to you.

Pinned Loading

  1. bionemo-recipesbionemo-recipesPublic

    BioNeMo Recipes: For building and adapting AI models in drug discovery at scale

    Python 848 177

  2. Proteina-ComplexaProteina-ComplexaPublic

    Generative model for protein binder design for protein and small molecule targets. Combines a pretrained flow-based generative model (built on La-Proteina) with inference-time optimization.

    Python 421 73

  3. nvMolKitnvMolKitPublic

    A high-performance, GPU-accelerated library for key computational chemistry tasks, such as molecular similarity, conformer generation, and geometry relaxation.

    Cuda 273 27

  4. genmolgenmolPublic

    GenMol is a generative AI model for creating novel molecules. It utilizes masked discrete diffusion and fragment-based generation to create valid molecules, which are encoded in the SAFE molecular …

    Python 199 35

  5. RNAProRNAProPublic

    RNAPro is a state-of-the-art RNA 3D folding model developed in collaboration with the hosts and winners of the Stanford RNA 3D Folding Kaggle competition.

    Python 95 20

  6. KERMTKERMTPublic

    KERMT is a pretrained graph neural network model for molecular property prediction.

    Python 96 17

Repositories

Showing 10 of 23 repositories

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…

Most used topics

Loading…