Skip to content

Repository files navigation

Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow

This is the official code repository for the paper titled Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow (ICML 2025).

Contribution

  • We propose SO(3)-Averaged Flow: A novel flow-matching objective that analytically computes the probability flow from noise to all rotations of the data. When the "correctness" of samples is rotational invariant (such as conformer generation), SO(3)-Averaged Flow improves training efficiency by eliminating the need for rotational data augmentation and further improves model performance.
  • We propose to use reflow+distillation to reduce the number of sampling steps of flow-based model for conformer generation and maintain high generation quality.
  • We provide a JAX implementation of the diffusion transformer with pairwise biased attention architecture. It is powerful and scalable for generative modeling of molecules.

Installation

Clone this repository:

git clone https://github.com/NVIDIA-Digital-Bio/avgflow.git
cd avgflow

Run the following command to create a conda environment and install the dependencies:

conda env create -f env.yml
conda activate avgflow

Pretrained Checkpoints

We provide 4 model weights through the NVIDIA NGC, including:

  1. 52M DiT trained with AvgFlow objective (Link)
  2. 52M DiT finetuned with reflow for few-step generation (Link)
  3. 52M DiT finetuned with reflow+distillation for 1-step generation (Link)
  4. 64M DiT trained with AvgFlow objective. (Link)

If you have NGC CLI tool installed, you can run the following command to download the checkpoints:

bash scripts/download_ckpts.sh

Otherwise, you can create a checkpoints directory by:

mkdir -p checkpoints

and download the checkpoints from the NGC pages above.

Sampling

The model can be used to generate conformers given: 1. a single SMILES string, or 2. a CSV file containing a batch of SMILES strings and the number of conformers to be generated for each molecule.

For conformer generation of a single molecule, run:

python avgflow/generate_from_smiles.py \
--config PATH/TO/CONFIG.yaml \
--smiles SMILES_STRING \
--num_confs N_CONF \
--output_dir PATH/TO/OUTPUT/DIRECTORY \

Example can be found in example/sampling/gen_smiles.sh, which generates 40 conformers for molecule C#CCNC(=O)C1=C[C@@H](c2ccc(Br)cc2)C[C@@H](OCc2ccc(CO)cc2)O1.

For conformer generation of a batch of molecules in a csv file, run:

python avgflow/generate_from_csv.py \
--config PATH/TO/CONFIG.yaml \
--smiles_csv PATH/TO/SMILES.csv \
--output_dir PATH/TO/OUTPUT/DIRECTORY \

Example can be found in example/sampling/gen_csv.sh, which generate various number of conformers for molecules in example/data/toy_gen_csv.csv. Please follow the format of example/data/toy_gen_csv.csv to construct your own csv for sampling.

The config yaml files define the model architecture to be initialized and checkpoint to be loaded. We provide 4 config files for the 4 checkpoints we released:

  1. config/generation_config/avgflow_52m_gen.yaml for the 52M DiT trained with AvgFlow objective.
  2. config/generation_config/avgflow_64m_gen.yaml for the 64M DiT trained with AvgFlow objective.
  3. config/generation_config/avgflow_52m_reflow_gen.yaml for the 52M DiT finetuned with reflow for few-step generation.
  4. config/generation_config/avgflow_52m_distill_gen.yaml for the 52M DiT finetuned with reflow+distillation for 1-step generation.

Please choose the config based on your checkpoint choice, and note that the distilled checkpoint only works with 1-step generation.

Training

Preparation of training data

Each molecule with ground truth conformers in the dataset has to be preprocessed before training. We recommend to have a dictionary for each molecule that contains at least 2 keys:

  1. features: Features computed from the 2D molecular graph using data_preprocessing.preprocess.mol2features
  2. conformers: np.array with dimension [C, N, 3], where C is the number of conformers and N is the number of atoms in the molecule.

For reflow/distill finetuning which requires ($X'_0$, $X'_1$) pairs, we recommend the preprocessed dictionary to contain 2 other keys instead of conformer:

  1. x0s: np.array with dimension [C, N, 3], where C is the number of ($X'_0$, $X'_1$) pairs and N is the number of atoms in the molecule. Gaussian noise at $t=0$.
  2. x1s: np.array with dimension [C, N, 3], where C is the number of ($X'_0$, $X'_1$) pairs and N is the number of atoms in the molecule. Model generated conformer at $t=1$ from each corresponding x0s.

Please refer to example/data/generate_toy_dataset.ipynb for example of creating the toy training and finetuning datasets.

Launch training

Follow the following steps to launch training:

  1. Prepare dataset as illustrated above.
  2. Prepare training config. See example in config/train_config/avgflow_52m_train_toy.yaml.
  3. (Optional) Change how the preprocessed dataset is loaded in avgflow/train.py line 45-58. You may parallelize the data loading for large training dataset.
  4. Launch training (see example in example/train/train_toy.sh) with:
python avgflow/train.py --config PATH/TO/CONFIG.yaml 

Follow the same procedure and use avgflow/reflow_finetune.py for finetuning.

License

Copyright @ 2025, NVIDIA Corporation. All rights reserved.
The source code is made available under Apache-2.0.
The model weights are made available under the NVIDIA Open Model License.

Citation

If you find this repository and our paper useful, please cite our work through:

@article{cao2025efficient,
title = {Efficient Molecular Conformer Generation with SO (3)-Averaged Flow Matching and Reflow},
author = {Cao, Zhonglin and Geiger, Mario and Costa, Allan Dos Santos and Reidenbach, Danny and Kreis, Karsten and Geffner, Tomas and Pellegrini, Franco and Zhou, Guoqing and Kucukbenli, Emine},
journal = {arXiv preprint arXiv:2507.09785},
year = {2025}
}

Disclaimer

This project will download and install additional third-party open source software projects. Review the license terms of these open source projects before use.

About

Official code repository for the paper titled "Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow" (ICML 2025)

Resources

Stars

16 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - NVIDIA-BioNeMo/avgflow: Official code repository for the paper titled "Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow" (ICML 2025) · GitHub
Skip to content

Repository files navigation

Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow

This is the official code repository for the paper titled Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow (ICML 2025).

Contribution

  • We propose SO(3)-Averaged Flow: A novel flow-matching objective that analytically computes the probability flow from noise to all rotations of the data. When the "correctness" of samples is rotational invariant (such as conformer generation), SO(3)-Averaged Flow improves training efficiency by eliminating the need for rotational data augmentation and further improves model performance.
  • We propose to use reflow+distillation to reduce the number of sampling steps of flow-based model for conformer generation and maintain high generation quality.
  • We provide a JAX implementation of the diffusion transformer with pairwise biased attention architecture. It is powerful and scalable for generative modeling of molecules.

Installation

Clone this repository:

git clone https://github.com/NVIDIA-Digital-Bio/avgflow.git
cd avgflow

Run the following command to create a conda environment and install the dependencies:

conda env create -f env.yml
conda activate avgflow

Pretrained Checkpoints

We provide 4 model weights through the NVIDIA NGC, including:

  1. 52M DiT trained with AvgFlow objective (Link)
  2. 52M DiT finetuned with reflow for few-step generation (Link)
  3. 52M DiT finetuned with reflow+distillation for 1-step generation (Link)
  4. 64M DiT trained with AvgFlow objective. (Link)

If you have NGC CLI tool installed, you can run the following command to download the checkpoints:

bash scripts/download_ckpts.sh

Otherwise, you can create a checkpoints directory by:

mkdir -p checkpoints

and download the checkpoints from the NGC pages above.

Sampling

The model can be used to generate conformers given: 1. a single SMILES string, or 2. a CSV file containing a batch of SMILES strings and the number of conformers to be generated for each molecule.

For conformer generation of a single molecule, run:

python avgflow/generate_from_smiles.py \
--config PATH/TO/CONFIG.yaml \
--smiles SMILES_STRING \
--num_confs N_CONF \
--output_dir PATH/TO/OUTPUT/DIRECTORY \

Example can be found in example/sampling/gen_smiles.sh, which generates 40 conformers for molecule C#CCNC(=O)C1=C[C@@H](c2ccc(Br)cc2)C[C@@H](OCc2ccc(CO)cc2)O1.

For conformer generation of a batch of molecules in a csv file, run:

python avgflow/generate_from_csv.py \
--config PATH/TO/CONFIG.yaml \
--smiles_csv PATH/TO/SMILES.csv \
--output_dir PATH/TO/OUTPUT/DIRECTORY \

Example can be found in example/sampling/gen_csv.sh, which generate various number of conformers for molecules in example/data/toy_gen_csv.csv. Please follow the format of example/data/toy_gen_csv.csv to construct your own csv for sampling.

The config yaml files define the model architecture to be initialized and checkpoint to be loaded. We provide 4 config files for the 4 checkpoints we released:

  1. config/generation_config/avgflow_52m_gen.yaml for the 52M DiT trained with AvgFlow objective.
  2. config/generation_config/avgflow_64m_gen.yaml for the 64M DiT trained with AvgFlow objective.
  3. config/generation_config/avgflow_52m_reflow_gen.yaml for the 52M DiT finetuned with reflow for few-step generation.
  4. config/generation_config/avgflow_52m_distill_gen.yaml for the 52M DiT finetuned with reflow+distillation for 1-step generation.

Please choose the config based on your checkpoint choice, and note that the distilled checkpoint only works with 1-step generation.

Training

Preparation of training data

Each molecule with ground truth conformers in the dataset has to be preprocessed before training. We recommend to have a dictionary for each molecule that contains at least 2 keys:

  1. features: Features computed from the 2D molecular graph using data_preprocessing.preprocess.mol2features
  2. conformers: np.array with dimension [C, N, 3], where C is the number of conformers and N is the number of atoms in the molecule.

For reflow/distill finetuning which requires ($X'_0$, $X'_1$) pairs, we recommend the preprocessed dictionary to contain 2 other keys instead of conformer:

  1. x0s: np.array with dimension [C, N, 3], where C is the number of ($X'_0$, $X'_1$) pairs and N is the number of atoms in the molecule. Gaussian noise at $t=0$.
  2. x1s: np.array with dimension [C, N, 3], where C is the number of ($X'_0$, $X'_1$) pairs and N is the number of atoms in the molecule. Model generated conformer at $t=1$ from each corresponding x0s.

Please refer to example/data/generate_toy_dataset.ipynb for example of creating the toy training and finetuning datasets.

Launch training

Follow the following steps to launch training:

  1. Prepare dataset as illustrated above.
  2. Prepare training config. See example in config/train_config/avgflow_52m_train_toy.yaml.
  3. (Optional) Change how the preprocessed dataset is loaded in avgflow/train.py line 45-58. You may parallelize the data loading for large training dataset.
  4. Launch training (see example in example/train/train_toy.sh) with:
python avgflow/train.py --config PATH/TO/CONFIG.yaml 

Follow the same procedure and use avgflow/reflow_finetune.py for finetuning.

License

Copyright @ 2025, NVIDIA Corporation. All rights reserved.
The source code is made available under Apache-2.0.
The model weights are made available under the NVIDIA Open Model License.

Citation

If you find this repository and our paper useful, please cite our work through:

@article{cao2025efficient,
title = {Efficient Molecular Conformer Generation with SO (3)-Averaged Flow Matching and Reflow},
author = {Cao, Zhonglin and Geiger, Mario and Costa, Allan Dos Santos and Reidenbach, Danny and Kreis, Karsten and Geffner, Tomas and Pellegrini, Franco and Zhou, Guoqing and Kucukbenli, Emine},
journal = {arXiv preprint arXiv:2507.09785},
year = {2025}
}

Disclaimer

This project will download and install additional third-party open source software projects. Review the license terms of these open source projects before use.

About

Official code repository for the paper titled "Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow" (ICML 2025)

Resources

Stars

16 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - NVIDIA-BioNeMo/avgflow: Official code repository for the paper titled "Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow" (ICML 2025) · GitHub
Skip to content

Repository files navigation

Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow

This is the official code repository for the paper titled Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow (ICML 2025).

Contribution

  • We propose SO(3)-Averaged Flow: A novel flow-matching objective that analytically computes the probability flow from noise to all rotations of the data. When the "correctness" of samples is rotational invariant (such as conformer generation), SO(3)-Averaged Flow improves training efficiency by eliminating the need for rotational data augmentation and further improves model performance.
  • We propose to use reflow+distillation to reduce the number of sampling steps of flow-based model for conformer generation and maintain high generation quality.
  • We provide a JAX implementation of the diffusion transformer with pairwise biased attention architecture. It is powerful and scalable for generative modeling of molecules.

Installation

Clone this repository:

git clone https://github.com/NVIDIA-Digital-Bio/avgflow.git
cd avgflow

Run the following command to create a conda environment and install the dependencies:

conda env create -f env.yml
conda activate avgflow

Pretrained Checkpoints

We provide 4 model weights through the NVIDIA NGC, including:

  1. 52M DiT trained with AvgFlow objective (Link)
  2. 52M DiT finetuned with reflow for few-step generation (Link)
  3. 52M DiT finetuned with reflow+distillation for 1-step generation (Link)
  4. 64M DiT trained with AvgFlow objective. (Link)

If you have NGC CLI tool installed, you can run the following command to download the checkpoints:

bash scripts/download_ckpts.sh

Otherwise, you can create a checkpoints directory by:

mkdir -p checkpoints

and download the checkpoints from the NGC pages above.

Sampling

The model can be used to generate conformers given: 1. a single SMILES string, or 2. a CSV file containing a batch of SMILES strings and the number of conformers to be generated for each molecule.

For conformer generation of a single molecule, run:

python avgflow/generate_from_smiles.py \
--config PATH/TO/CONFIG.yaml \
--smiles SMILES_STRING \
--num_confs N_CONF \
--output_dir PATH/TO/OUTPUT/DIRECTORY \

Example can be found in example/sampling/gen_smiles.sh, which generates 40 conformers for molecule C#CCNC(=O)C1=C[C@@H](c2ccc(Br)cc2)C[C@@H](OCc2ccc(CO)cc2)O1.

For conformer generation of a batch of molecules in a csv file, run:

python avgflow/generate_from_csv.py \
--config PATH/TO/CONFIG.yaml \
--smiles_csv PATH/TO/SMILES.csv \
--output_dir PATH/TO/OUTPUT/DIRECTORY \

Example can be found in example/sampling/gen_csv.sh, which generate various number of conformers for molecules in example/data/toy_gen_csv.csv. Please follow the format of example/data/toy_gen_csv.csv to construct your own csv for sampling.

The config yaml files define the model architecture to be initialized and checkpoint to be loaded. We provide 4 config files for the 4 checkpoints we released:

  1. config/generation_config/avgflow_52m_gen.yaml for the 52M DiT trained with AvgFlow objective.
  2. config/generation_config/avgflow_64m_gen.yaml for the 64M DiT trained with AvgFlow objective.
  3. config/generation_config/avgflow_52m_reflow_gen.yaml for the 52M DiT finetuned with reflow for few-step generation.
  4. config/generation_config/avgflow_52m_distill_gen.yaml for the 52M DiT finetuned with reflow+distillation for 1-step generation.

Please choose the config based on your checkpoint choice, and note that the distilled checkpoint only works with 1-step generation.

Training

Preparation of training data

Each molecule with ground truth conformers in the dataset has to be preprocessed before training. We recommend to have a dictionary for each molecule that contains at least 2 keys:

  1. features: Features computed from the 2D molecular graph using data_preprocessing.preprocess.mol2features
  2. conformers: np.array with dimension [C, N, 3], where C is the number of conformers and N is the number of atoms in the molecule.

For reflow/distill finetuning which requires ($X'_0$, $X'_1$) pairs, we recommend the preprocessed dictionary to contain 2 other keys instead of conformer:

  1. x0s: np.array with dimension [C, N, 3], where C is the number of ($X'_0$, $X'_1$) pairs and N is the number of atoms in the molecule. Gaussian noise at $t=0$.
  2. x1s: np.array with dimension [C, N, 3], where C is the number of ($X'_0$, $X'_1$) pairs and N is the number of atoms in the molecule. Model generated conformer at $t=1$ from each corresponding x0s.

Please refer to example/data/generate_toy_dataset.ipynb for example of creating the toy training and finetuning datasets.

Launch training

Follow the following steps to launch training:

  1. Prepare dataset as illustrated above.
  2. Prepare training config. See example in config/train_config/avgflow_52m_train_toy.yaml.
  3. (Optional) Change how the preprocessed dataset is loaded in avgflow/train.py line 45-58. You may parallelize the data loading for large training dataset.
  4. Launch training (see example in example/train/train_toy.sh) with:
python avgflow/train.py --config PATH/TO/CONFIG.yaml 

Follow the same procedure and use avgflow/reflow_finetune.py for finetuning.

License

Copyright @ 2025, NVIDIA Corporation. All rights reserved.
The source code is made available under Apache-2.0.
The model weights are made available under the NVIDIA Open Model License.

Citation

If you find this repository and our paper useful, please cite our work through:

@article{cao2025efficient,
title = {Efficient Molecular Conformer Generation with SO (3)-Averaged Flow Matching and Reflow},
author = {Cao, Zhonglin and Geiger, Mario and Costa, Allan Dos Santos and Reidenbach, Danny and Kreis, Karsten and Geffner, Tomas and Pellegrini, Franco and Zhou, Guoqing and Kucukbenli, Emine},
journal = {arXiv preprint arXiv:2507.09785},
year = {2025}
}

Disclaimer

This project will download and install additional third-party open source software projects. Review the license terms of these open source projects before use.

About

Official code repository for the paper titled "Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow" (ICML 2025)

Resources

Stars

16 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - NVIDIA-BioNeMo/avgflow: Official code repository for the paper titled "Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow" (ICML 2025) · GitHub
Skip to content

Repository files navigation

Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow

This is the official code repository for the paper titled Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow (ICML 2025).

Contribution

  • We propose SO(3)-Averaged Flow: A novel flow-matching objective that analytically computes the probability flow from noise to all rotations of the data. When the "correctness" of samples is rotational invariant (such as conformer generation), SO(3)-Averaged Flow improves training efficiency by eliminating the need for rotational data augmentation and further improves model performance.
  • We propose to use reflow+distillation to reduce the number of sampling steps of flow-based model for conformer generation and maintain high generation quality.
  • We provide a JAX implementation of the diffusion transformer with pairwise biased attention architecture. It is powerful and scalable for generative modeling of molecules.

Installation

Clone this repository:

git clone https://github.com/NVIDIA-Digital-Bio/avgflow.git
cd avgflow

Run the following command to create a conda environment and install the dependencies:

conda env create -f env.yml
conda activate avgflow

Pretrained Checkpoints

We provide 4 model weights through the NVIDIA NGC, including:

  1. 52M DiT trained with AvgFlow objective (Link)
  2. 52M DiT finetuned with reflow for few-step generation (Link)
  3. 52M DiT finetuned with reflow+distillation for 1-step generation (Link)
  4. 64M DiT trained with AvgFlow objective. (Link)

If you have NGC CLI tool installed, you can run the following command to download the checkpoints:

bash scripts/download_ckpts.sh

Otherwise, you can create a checkpoints directory by:

mkdir -p checkpoints

and download the checkpoints from the NGC pages above.

Sampling

The model can be used to generate conformers given: 1. a single SMILES string, or 2. a CSV file containing a batch of SMILES strings and the number of conformers to be generated for each molecule.

For conformer generation of a single molecule, run:

python avgflow/generate_from_smiles.py \
--config PATH/TO/CONFIG.yaml \
--smiles SMILES_STRING \
--num_confs N_CONF \
--output_dir PATH/TO/OUTPUT/DIRECTORY \

Example can be found in example/sampling/gen_smiles.sh, which generates 40 conformers for molecule C#CCNC(=O)C1=C[C@@H](c2ccc(Br)cc2)C[C@@H](OCc2ccc(CO)cc2)O1.

For conformer generation of a batch of molecules in a csv file, run:

python avgflow/generate_from_csv.py \
--config PATH/TO/CONFIG.yaml \
--smiles_csv PATH/TO/SMILES.csv \
--output_dir PATH/TO/OUTPUT/DIRECTORY \

Example can be found in example/sampling/gen_csv.sh, which generate various number of conformers for molecules in example/data/toy_gen_csv.csv. Please follow the format of example/data/toy_gen_csv.csv to construct your own csv for sampling.

The config yaml files define the model architecture to be initialized and checkpoint to be loaded. We provide 4 config files for the 4 checkpoints we released:

  1. config/generation_config/avgflow_52m_gen.yaml for the 52M DiT trained with AvgFlow objective.
  2. config/generation_config/avgflow_64m_gen.yaml for the 64M DiT trained with AvgFlow objective.
  3. config/generation_config/avgflow_52m_reflow_gen.yaml for the 52M DiT finetuned with reflow for few-step generation.
  4. config/generation_config/avgflow_52m_distill_gen.yaml for the 52M DiT finetuned with reflow+distillation for 1-step generation.

Please choose the config based on your checkpoint choice, and note that the distilled checkpoint only works with 1-step generation.

Training

Preparation of training data

Each molecule with ground truth conformers in the dataset has to be preprocessed before training. We recommend to have a dictionary for each molecule that contains at least 2 keys:

  1. features: Features computed from the 2D molecular graph using data_preprocessing.preprocess.mol2features
  2. conformers: np.array with dimension [C, N, 3], where C is the number of conformers and N is the number of atoms in the molecule.

For reflow/distill finetuning which requires ($X'_0$, $X'_1$) pairs, we recommend the preprocessed dictionary to contain 2 other keys instead of conformer:

  1. x0s: np.array with dimension [C, N, 3], where C is the number of ($X'_0$, $X'_1$) pairs and N is the number of atoms in the molecule. Gaussian noise at $t=0$.
  2. x1s: np.array with dimension [C, N, 3], where C is the number of ($X'_0$, $X'_1$) pairs and N is the number of atoms in the molecule. Model generated conformer at $t=1$ from each corresponding x0s.

Please refer to example/data/generate_toy_dataset.ipynb for example of creating the toy training and finetuning datasets.

Launch training

Follow the following steps to launch training:

  1. Prepare dataset as illustrated above.
  2. Prepare training config. See example in config/train_config/avgflow_52m_train_toy.yaml.
  3. (Optional) Change how the preprocessed dataset is loaded in avgflow/train.py line 45-58. You may parallelize the data loading for large training dataset.
  4. Launch training (see example in example/train/train_toy.sh) with:
python avgflow/train.py --config PATH/TO/CONFIG.yaml 

Follow the same procedure and use avgflow/reflow_finetune.py for finetuning.

License

Copyright @ 2025, NVIDIA Corporation. All rights reserved.
The source code is made available under Apache-2.0.
The model weights are made available under the NVIDIA Open Model License.

Citation

If you find this repository and our paper useful, please cite our work through:

@article{cao2025efficient,
title = {Efficient Molecular Conformer Generation with SO (3)-Averaged Flow Matching and Reflow},
author = {Cao, Zhonglin and Geiger, Mario and Costa, Allan Dos Santos and Reidenbach, Danny and Kreis, Karsten and Geffner, Tomas and Pellegrini, Franco and Zhou, Guoqing and Kucukbenli, Emine},
journal = {arXiv preprint arXiv:2507.09785},
year = {2025}
}

Disclaimer

This project will download and install additional third-party open source software projects. Review the license terms of these open source projects before use.

About

Official code repository for the paper titled "Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow" (ICML 2025)

Resources

Stars

16 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - NVIDIA-BioNeMo/avgflow: Official code repository for the paper titled "Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow" (ICML 2025) · GitHub
Skip to content

Repository files navigation

Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow

This is the official code repository for the paper titled Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow (ICML 2025).

Contribution

  • We propose SO(3)-Averaged Flow: A novel flow-matching objective that analytically computes the probability flow from noise to all rotations of the data. When the "correctness" of samples is rotational invariant (such as conformer generation), SO(3)-Averaged Flow improves training efficiency by eliminating the need for rotational data augmentation and further improves model performance.
  • We propose to use reflow+distillation to reduce the number of sampling steps of flow-based model for conformer generation and maintain high generation quality.
  • We provide a JAX implementation of the diffusion transformer with pairwise biased attention architecture. It is powerful and scalable for generative modeling of molecules.

Installation

Clone this repository:

git clone https://github.com/NVIDIA-Digital-Bio/avgflow.git
cd avgflow

Run the following command to create a conda environment and install the dependencies:

conda env create -f env.yml
conda activate avgflow

Pretrained Checkpoints

We provide 4 model weights through the NVIDIA NGC, including:

  1. 52M DiT trained with AvgFlow objective (Link)
  2. 52M DiT finetuned with reflow for few-step generation (Link)
  3. 52M DiT finetuned with reflow+distillation for 1-step generation (Link)
  4. 64M DiT trained with AvgFlow objective. (Link)

If you have NGC CLI tool installed, you can run the following command to download the checkpoints:

bash scripts/download_ckpts.sh

Otherwise, you can create a checkpoints directory by:

mkdir -p checkpoints

and download the checkpoints from the NGC pages above.

Sampling

The model can be used to generate conformers given: 1. a single SMILES string, or 2. a CSV file containing a batch of SMILES strings and the number of conformers to be generated for each molecule.

For conformer generation of a single molecule, run:

python avgflow/generate_from_smiles.py \
--config PATH/TO/CONFIG.yaml \
--smiles SMILES_STRING \
--num_confs N_CONF \
--output_dir PATH/TO/OUTPUT/DIRECTORY \

Example can be found in example/sampling/gen_smiles.sh, which generates 40 conformers for molecule C#CCNC(=O)C1=C[C@@H](c2ccc(Br)cc2)C[C@@H](OCc2ccc(CO)cc2)O1.

For conformer generation of a batch of molecules in a csv file, run:

python avgflow/generate_from_csv.py \
--config PATH/TO/CONFIG.yaml \
--smiles_csv PATH/TO/SMILES.csv \
--output_dir PATH/TO/OUTPUT/DIRECTORY \

Example can be found in example/sampling/gen_csv.sh, which generate various number of conformers for molecules in example/data/toy_gen_csv.csv. Please follow the format of example/data/toy_gen_csv.csv to construct your own csv for sampling.

The config yaml files define the model architecture to be initialized and checkpoint to be loaded. We provide 4 config files for the 4 checkpoints we released:

  1. config/generation_config/avgflow_52m_gen.yaml for the 52M DiT trained with AvgFlow objective.
  2. config/generation_config/avgflow_64m_gen.yaml for the 64M DiT trained with AvgFlow objective.
  3. config/generation_config/avgflow_52m_reflow_gen.yaml for the 52M DiT finetuned with reflow for few-step generation.
  4. config/generation_config/avgflow_52m_distill_gen.yaml for the 52M DiT finetuned with reflow+distillation for 1-step generation.

Please choose the config based on your checkpoint choice, and note that the distilled checkpoint only works with 1-step generation.

Training

Preparation of training data

Each molecule with ground truth conformers in the dataset has to be preprocessed before training. We recommend to have a dictionary for each molecule that contains at least 2 keys:

  1. features: Features computed from the 2D molecular graph using data_preprocessing.preprocess.mol2features
  2. conformers: np.array with dimension [C, N, 3], where C is the number of conformers and N is the number of atoms in the molecule.

For reflow/distill finetuning which requires ($X'_0$, $X'_1$) pairs, we recommend the preprocessed dictionary to contain 2 other keys instead of conformer:

  1. x0s: np.array with dimension [C, N, 3], where C is the number of ($X'_0$, $X'_1$) pairs and N is the number of atoms in the molecule. Gaussian noise at $t=0$.
  2. x1s: np.array with dimension [C, N, 3], where C is the number of ($X'_0$, $X'_1$) pairs and N is the number of atoms in the molecule. Model generated conformer at $t=1$ from each corresponding x0s.

Please refer to example/data/generate_toy_dataset.ipynb for example of creating the toy training and finetuning datasets.

Launch training

Follow the following steps to launch training:

  1. Prepare dataset as illustrated above.
  2. Prepare training config. See example in config/train_config/avgflow_52m_train_toy.yaml.
  3. (Optional) Change how the preprocessed dataset is loaded in avgflow/train.py line 45-58. You may parallelize the data loading for large training dataset.
  4. Launch training (see example in example/train/train_toy.sh) with:
python avgflow/train.py --config PATH/TO/CONFIG.yaml 

Follow the same procedure and use avgflow/reflow_finetune.py for finetuning.

License

Copyright @ 2025, NVIDIA Corporation. All rights reserved.
The source code is made available under Apache-2.0.
The model weights are made available under the NVIDIA Open Model License.

Citation

If you find this repository and our paper useful, please cite our work through:

@article{cao2025efficient,
title = {Efficient Molecular Conformer Generation with SO (3)-Averaged Flow Matching and Reflow},
author = {Cao, Zhonglin and Geiger, Mario and Costa, Allan Dos Santos and Reidenbach, Danny and Kreis, Karsten and Geffner, Tomas and Pellegrini, Franco and Zhou, Guoqing and Kucukbenli, Emine},
journal = {arXiv preprint arXiv:2507.09785},
year = {2025}
}

Disclaimer

This project will download and install additional third-party open source software projects. Review the license terms of these open source projects before use.

About

Official code repository for the paper titled "Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow" (ICML 2025)

Resources

Stars

16 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - NVIDIA-BioNeMo/avgflow: Official code repository for the paper titled "Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow" (ICML 2025) · GitHub
Skip to content

Repository files navigation

Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow

This is the official code repository for the paper titled Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow (ICML 2025).

Contribution

  • We propose SO(3)-Averaged Flow: A novel flow-matching objective that analytically computes the probability flow from noise to all rotations of the data. When the "correctness" of samples is rotational invariant (such as conformer generation), SO(3)-Averaged Flow improves training efficiency by eliminating the need for rotational data augmentation and further improves model performance.
  • We propose to use reflow+distillation to reduce the number of sampling steps of flow-based model for conformer generation and maintain high generation quality.
  • We provide a JAX implementation of the diffusion transformer with pairwise biased attention architecture. It is powerful and scalable for generative modeling of molecules.

Installation

Clone this repository:

git clone https://github.com/NVIDIA-Digital-Bio/avgflow.git
cd avgflow

Run the following command to create a conda environment and install the dependencies:

conda env create -f env.yml
conda activate avgflow

Pretrained Checkpoints

We provide 4 model weights through the NVIDIA NGC, including:

  1. 52M DiT trained with AvgFlow objective (Link)
  2. 52M DiT finetuned with reflow for few-step generation (Link)
  3. 52M DiT finetuned with reflow+distillation for 1-step generation (Link)
  4. 64M DiT trained with AvgFlow objective. (Link)

If you have NGC CLI tool installed, you can run the following command to download the checkpoints:

bash scripts/download_ckpts.sh

Otherwise, you can create a checkpoints directory by:

mkdir -p checkpoints

and download the checkpoints from the NGC pages above.

Sampling

The model can be used to generate conformers given: 1. a single SMILES string, or 2. a CSV file containing a batch of SMILES strings and the number of conformers to be generated for each molecule.

For conformer generation of a single molecule, run:

python avgflow/generate_from_smiles.py \
--config PATH/TO/CONFIG.yaml \
--smiles SMILES_STRING \
--num_confs N_CONF \
--output_dir PATH/TO/OUTPUT/DIRECTORY \

Example can be found in example/sampling/gen_smiles.sh, which generates 40 conformers for molecule C#CCNC(=O)C1=C[C@@H](c2ccc(Br)cc2)C[C@@H](OCc2ccc(CO)cc2)O1.

For conformer generation of a batch of molecules in a csv file, run:

python avgflow/generate_from_csv.py \
--config PATH/TO/CONFIG.yaml \
--smiles_csv PATH/TO/SMILES.csv \
--output_dir PATH/TO/OUTPUT/DIRECTORY \

Example can be found in example/sampling/gen_csv.sh, which generate various number of conformers for molecules in example/data/toy_gen_csv.csv. Please follow the format of example/data/toy_gen_csv.csv to construct your own csv for sampling.

The config yaml files define the model architecture to be initialized and checkpoint to be loaded. We provide 4 config files for the 4 checkpoints we released:

  1. config/generation_config/avgflow_52m_gen.yaml for the 52M DiT trained with AvgFlow objective.
  2. config/generation_config/avgflow_64m_gen.yaml for the 64M DiT trained with AvgFlow objective.
  3. config/generation_config/avgflow_52m_reflow_gen.yaml for the 52M DiT finetuned with reflow for few-step generation.
  4. config/generation_config/avgflow_52m_distill_gen.yaml for the 52M DiT finetuned with reflow+distillation for 1-step generation.

Please choose the config based on your checkpoint choice, and note that the distilled checkpoint only works with 1-step generation.

Training

Preparation of training data

Each molecule with ground truth conformers in the dataset has to be preprocessed before training. We recommend to have a dictionary for each molecule that contains at least 2 keys:

  1. features: Features computed from the 2D molecular graph using data_preprocessing.preprocess.mol2features
  2. conformers: np.array with dimension [C, N, 3], where C is the number of conformers and N is the number of atoms in the molecule.

For reflow/distill finetuning which requires ($X'_0$, $X'_1$) pairs, we recommend the preprocessed dictionary to contain 2 other keys instead of conformer:

  1. x0s: np.array with dimension [C, N, 3], where C is the number of ($X'_0$, $X'_1$) pairs and N is the number of atoms in the molecule. Gaussian noise at $t=0$.
  2. x1s: np.array with dimension [C, N, 3], where C is the number of ($X'_0$, $X'_1$) pairs and N is the number of atoms in the molecule. Model generated conformer at $t=1$ from each corresponding x0s.

Please refer to example/data/generate_toy_dataset.ipynb for example of creating the toy training and finetuning datasets.

Launch training

Follow the following steps to launch training:

  1. Prepare dataset as illustrated above.
  2. Prepare training config. See example in config/train_config/avgflow_52m_train_toy.yaml.
  3. (Optional) Change how the preprocessed dataset is loaded in avgflow/train.py line 45-58. You may parallelize the data loading for large training dataset.
  4. Launch training (see example in example/train/train_toy.sh) with:
python avgflow/train.py --config PATH/TO/CONFIG.yaml 

Follow the same procedure and use avgflow/reflow_finetune.py for finetuning.

License

Copyright @ 2025, NVIDIA Corporation. All rights reserved.
The source code is made available under Apache-2.0.
The model weights are made available under the NVIDIA Open Model License.

Citation

If you find this repository and our paper useful, please cite our work through:

@article{cao2025efficient,
title = {Efficient Molecular Conformer Generation with SO (3)-Averaged Flow Matching and Reflow},
author = {Cao, Zhonglin and Geiger, Mario and Costa, Allan Dos Santos and Reidenbach, Danny and Kreis, Karsten and Geffner, Tomas and Pellegrini, Franco and Zhou, Guoqing and Kucukbenli, Emine},
journal = {arXiv preprint arXiv:2507.09785},
year = {2025}
}

Disclaimer

This project will download and install additional third-party open source software projects. Review the license terms of these open source projects before use.

About

Official code repository for the paper titled "Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow" (ICML 2025)

Resources

Stars

16 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - NVIDIA-BioNeMo/avgflow: Official code repository for the paper titled "Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow" (ICML 2025) · GitHub
Skip to content

Repository files navigation

Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow

This is the official code repository for the paper titled Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow (ICML 2025).

Contribution

  • We propose SO(3)-Averaged Flow: A novel flow-matching objective that analytically computes the probability flow from noise to all rotations of the data. When the "correctness" of samples is rotational invariant (such as conformer generation), SO(3)-Averaged Flow improves training efficiency by eliminating the need for rotational data augmentation and further improves model performance.
  • We propose to use reflow+distillation to reduce the number of sampling steps of flow-based model for conformer generation and maintain high generation quality.
  • We provide a JAX implementation of the diffusion transformer with pairwise biased attention architecture. It is powerful and scalable for generative modeling of molecules.

Installation

Clone this repository:

git clone https://github.com/NVIDIA-Digital-Bio/avgflow.git
cd avgflow

Run the following command to create a conda environment and install the dependencies:

conda env create -f env.yml
conda activate avgflow

Pretrained Checkpoints

We provide 4 model weights through the NVIDIA NGC, including:

  1. 52M DiT trained with AvgFlow objective (Link)
  2. 52M DiT finetuned with reflow for few-step generation (Link)
  3. 52M DiT finetuned with reflow+distillation for 1-step generation (Link)
  4. 64M DiT trained with AvgFlow objective. (Link)

If you have NGC CLI tool installed, you can run the following command to download the checkpoints:

bash scripts/download_ckpts.sh

Otherwise, you can create a checkpoints directory by:

mkdir -p checkpoints

and download the checkpoints from the NGC pages above.

Sampling

The model can be used to generate conformers given: 1. a single SMILES string, or 2. a CSV file containing a batch of SMILES strings and the number of conformers to be generated for each molecule.

For conformer generation of a single molecule, run:

python avgflow/generate_from_smiles.py \
--config PATH/TO/CONFIG.yaml \
--smiles SMILES_STRING \
--num_confs N_CONF \
--output_dir PATH/TO/OUTPUT/DIRECTORY \

Example can be found in example/sampling/gen_smiles.sh, which generates 40 conformers for molecule C#CCNC(=O)C1=C[C@@H](c2ccc(Br)cc2)C[C@@H](OCc2ccc(CO)cc2)O1.

For conformer generation of a batch of molecules in a csv file, run:

python avgflow/generate_from_csv.py \
--config PATH/TO/CONFIG.yaml \
--smiles_csv PATH/TO/SMILES.csv \
--output_dir PATH/TO/OUTPUT/DIRECTORY \

Example can be found in example/sampling/gen_csv.sh, which generate various number of conformers for molecules in example/data/toy_gen_csv.csv. Please follow the format of example/data/toy_gen_csv.csv to construct your own csv for sampling.

The config yaml files define the model architecture to be initialized and checkpoint to be loaded. We provide 4 config files for the 4 checkpoints we released:

  1. config/generation_config/avgflow_52m_gen.yaml for the 52M DiT trained with AvgFlow objective.
  2. config/generation_config/avgflow_64m_gen.yaml for the 64M DiT trained with AvgFlow objective.
  3. config/generation_config/avgflow_52m_reflow_gen.yaml for the 52M DiT finetuned with reflow for few-step generation.
  4. config/generation_config/avgflow_52m_distill_gen.yaml for the 52M DiT finetuned with reflow+distillation for 1-step generation.

Please choose the config based on your checkpoint choice, and note that the distilled checkpoint only works with 1-step generation.

Training

Preparation of training data

Each molecule with ground truth conformers in the dataset has to be preprocessed before training. We recommend to have a dictionary for each molecule that contains at least 2 keys:

  1. features: Features computed from the 2D molecular graph using data_preprocessing.preprocess.mol2features
  2. conformers: np.array with dimension [C, N, 3], where C is the number of conformers and N is the number of atoms in the molecule.

For reflow/distill finetuning which requires ($X'_0$, $X'_1$) pairs, we recommend the preprocessed dictionary to contain 2 other keys instead of conformer:

  1. x0s: np.array with dimension [C, N, 3], where C is the number of ($X'_0$, $X'_1$) pairs and N is the number of atoms in the molecule. Gaussian noise at $t=0$.
  2. x1s: np.array with dimension [C, N, 3], where C is the number of ($X'_0$, $X'_1$) pairs and N is the number of atoms in the molecule. Model generated conformer at $t=1$ from each corresponding x0s.

Please refer to example/data/generate_toy_dataset.ipynb for example of creating the toy training and finetuning datasets.

Launch training

Follow the following steps to launch training:

  1. Prepare dataset as illustrated above.
  2. Prepare training config. See example in config/train_config/avgflow_52m_train_toy.yaml.
  3. (Optional) Change how the preprocessed dataset is loaded in avgflow/train.py line 45-58. You may parallelize the data loading for large training dataset.
  4. Launch training (see example in example/train/train_toy.sh) with:
python avgflow/train.py --config PATH/TO/CONFIG.yaml 

Follow the same procedure and use avgflow/reflow_finetune.py for finetuning.

License

Copyright @ 2025, NVIDIA Corporation. All rights reserved.
The source code is made available under Apache-2.0.
The model weights are made available under the NVIDIA Open Model License.

Citation

If you find this repository and our paper useful, please cite our work through:

@article{cao2025efficient,
title = {Efficient Molecular Conformer Generation with SO (3)-Averaged Flow Matching and Reflow},
author = {Cao, Zhonglin and Geiger, Mario and Costa, Allan Dos Santos and Reidenbach, Danny and Kreis, Karsten and Geffner, Tomas and Pellegrini, Franco and Zhou, Guoqing and Kucukbenli, Emine},
journal = {arXiv preprint arXiv:2507.09785},
year = {2025}
}

Disclaimer

This project will download and install additional third-party open source software projects. Review the license terms of these open source projects before use.

About

Official code repository for the paper titled "Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow" (ICML 2025)

Resources

Stars

16 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - NVIDIA-BioNeMo/avgflow: Official code repository for the paper titled "Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow" (ICML 2025) · GitHub
Skip to content

Repository files navigation

Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow

This is the official code repository for the paper titled Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow (ICML 2025).

Contribution

  • We propose SO(3)-Averaged Flow: A novel flow-matching objective that analytically computes the probability flow from noise to all rotations of the data. When the "correctness" of samples is rotational invariant (such as conformer generation), SO(3)-Averaged Flow improves training efficiency by eliminating the need for rotational data augmentation and further improves model performance.
  • We propose to use reflow+distillation to reduce the number of sampling steps of flow-based model for conformer generation and maintain high generation quality.
  • We provide a JAX implementation of the diffusion transformer with pairwise biased attention architecture. It is powerful and scalable for generative modeling of molecules.

Installation

Clone this repository:

git clone https://github.com/NVIDIA-Digital-Bio/avgflow.git
cd avgflow

Run the following command to create a conda environment and install the dependencies:

conda env create -f env.yml
conda activate avgflow

Pretrained Checkpoints

We provide 4 model weights through the NVIDIA NGC, including:

  1. 52M DiT trained with AvgFlow objective (Link)
  2. 52M DiT finetuned with reflow for few-step generation (Link)
  3. 52M DiT finetuned with reflow+distillation for 1-step generation (Link)
  4. 64M DiT trained with AvgFlow objective. (Link)

If you have NGC CLI tool installed, you can run the following command to download the checkpoints:

bash scripts/download_ckpts.sh

Otherwise, you can create a checkpoints directory by:

mkdir -p checkpoints

and download the checkpoints from the NGC pages above.

Sampling

The model can be used to generate conformers given: 1. a single SMILES string, or 2. a CSV file containing a batch of SMILES strings and the number of conformers to be generated for each molecule.

For conformer generation of a single molecule, run:

python avgflow/generate_from_smiles.py \
--config PATH/TO/CONFIG.yaml \
--smiles SMILES_STRING \
--num_confs N_CONF \
--output_dir PATH/TO/OUTPUT/DIRECTORY \

Example can be found in example/sampling/gen_smiles.sh, which generates 40 conformers for molecule C#CCNC(=O)C1=C[C@@H](c2ccc(Br)cc2)C[C@@H](OCc2ccc(CO)cc2)O1.

For conformer generation of a batch of molecules in a csv file, run:

python avgflow/generate_from_csv.py \
--config PATH/TO/CONFIG.yaml \
--smiles_csv PATH/TO/SMILES.csv \
--output_dir PATH/TO/OUTPUT/DIRECTORY \

Example can be found in example/sampling/gen_csv.sh, which generate various number of conformers for molecules in example/data/toy_gen_csv.csv. Please follow the format of example/data/toy_gen_csv.csv to construct your own csv for sampling.

The config yaml files define the model architecture to be initialized and checkpoint to be loaded. We provide 4 config files for the 4 checkpoints we released:

  1. config/generation_config/avgflow_52m_gen.yaml for the 52M DiT trained with AvgFlow objective.
  2. config/generation_config/avgflow_64m_gen.yaml for the 64M DiT trained with AvgFlow objective.
  3. config/generation_config/avgflow_52m_reflow_gen.yaml for the 52M DiT finetuned with reflow for few-step generation.
  4. config/generation_config/avgflow_52m_distill_gen.yaml for the 52M DiT finetuned with reflow+distillation for 1-step generation.

Please choose the config based on your checkpoint choice, and note that the distilled checkpoint only works with 1-step generation.

Training

Preparation of training data

Each molecule with ground truth conformers in the dataset has to be preprocessed before training. We recommend to have a dictionary for each molecule that contains at least 2 keys:

  1. features: Features computed from the 2D molecular graph using data_preprocessing.preprocess.mol2features
  2. conformers: np.array with dimension [C, N, 3], where C is the number of conformers and N is the number of atoms in the molecule.

For reflow/distill finetuning which requires ($X'_0$, $X'_1$) pairs, we recommend the preprocessed dictionary to contain 2 other keys instead of conformer:

  1. x0s: np.array with dimension [C, N, 3], where C is the number of ($X'_0$, $X'_1$) pairs and N is the number of atoms in the molecule. Gaussian noise at $t=0$.
  2. x1s: np.array with dimension [C, N, 3], where C is the number of ($X'_0$, $X'_1$) pairs and N is the number of atoms in the molecule. Model generated conformer at $t=1$ from each corresponding x0s.

Please refer to example/data/generate_toy_dataset.ipynb for example of creating the toy training and finetuning datasets.

Launch training

Follow the following steps to launch training:

  1. Prepare dataset as illustrated above.
  2. Prepare training config. See example in config/train_config/avgflow_52m_train_toy.yaml.
  3. (Optional) Change how the preprocessed dataset is loaded in avgflow/train.py line 45-58. You may parallelize the data loading for large training dataset.
  4. Launch training (see example in example/train/train_toy.sh) with:
python avgflow/train.py --config PATH/TO/CONFIG.yaml 

Follow the same procedure and use avgflow/reflow_finetune.py for finetuning.

License

Copyright @ 2025, NVIDIA Corporation. All rights reserved.
The source code is made available under Apache-2.0.
The model weights are made available under the NVIDIA Open Model License.

Citation

If you find this repository and our paper useful, please cite our work through:

@article{cao2025efficient,
title = {Efficient Molecular Conformer Generation with SO (3)-Averaged Flow Matching and Reflow},
author = {Cao, Zhonglin and Geiger, Mario and Costa, Allan Dos Santos and Reidenbach, Danny and Kreis, Karsten and Geffner, Tomas and Pellegrini, Franco and Zhou, Guoqing and Kucukbenli, Emine},
journal = {arXiv preprint arXiv:2507.09785},
year = {2025}
}

Disclaimer

This project will download and install additional third-party open source software projects. Review the license terms of these open source projects before use.

About

Official code repository for the paper titled "Efficient Molecular Conformer Generation with SO(3) Averaged Flow-Matching and Reflow" (ICML 2025)

Resources

Stars

16 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages