Repository files navigation

MIST: Molecular Insight SMILES Transformer

GitHub LicensearXiv:2409.15370Model on HF

MIST is a family of molecular foundation models for molecular property prediction. The models were pre-trained on Smirk 😏 tokenized SMILES strings from the Enamine REAL Space dataset using the Masked Language Modeling (MLM) objective, then fine-tuned for downstream prediction tasks.

Installation

The following provides installation instructions for the top-level package (electrolyte_fm), optional add-ons for our various additional analysis and downstream applications (See ./opt may require additional configuration.

  1. Install uv and julia (only needed for /opt tasks)
  2. Instantiate the environment: uv sync
  3. Use submit/submit.py to submit a training job or checkout one of our applications in ./opt

You may need to install rust if pre-built wheels for smirk are not available on PyPI. Feel free to open an issue to request additional pre-built wheels.

Installation took under a minute on a MacBook Pro 2023. However, installation times depend heavily on internet connection and local uv cache. Pulling the NVIDIA packages can easily add 5-10 minutes without a uv cache. Installing the Julia environments (for example in /opt) can takes ~5 mins but may take longer on limited hardware (e.g. GitHub runners take ~18 mins).

Local

  1. Install rust and uv
  2. Run uv sync

Not all configurations will work locally. For example, pre-training configuration files typically use DeepSpeed, which requires NVIDIA GPUs. Some adaptation may be required to run outside of an NVIDIA GPU cluster.

Polaris

  1. Install rust and uv

  2. Load conda

module purge
module use /soft/modulefiles/
module --ignore_cache load conda/2024-04-29
conda activate base
  1. Install the environment
uv sync

Artemis

Same as above except:

  1. Skip loading conda (just use uv)
  2. Ensure a module for CUDA@12.2 exists, may need to install with spack (make sure buildable: True)

Apptainer

  1. Install or load from a module Apptainer
  2. Build the image bash container/build.sh, once build relocate the image mv /tmp/mist.sif ./mist.sif
  3. Run training within the image apptainer run --nv mist.sif python train.py ...

See submit/dgx.j2 or submit/delta.j2 for a more complete example of using the container

System Requirements

Hardware

Generally, running the code here requires access to GPUs and ideally a dedicated NVIDIA GPU cluster. However, much of the (non-training) code can be run on a single NVIDIA A40 GPU. Notably, for the MIST-28M model, CPU or MPS inference is viable.

Software Dependencies

All software dependencies can be found in the pyproject.toml (python) or Project.toml (julia) files. Specific versions are detailed in the lock files.

MIST has been tested on the following primary dependencies:

DependencyVersions
Ubuntu24.04
Python3.12, 3.13
NVIDIA CUDA12.8.0.038
NVIDIA cuBLAS12.8.3.14
NVIDIA cuDNN9.7.0.66
NVIDIA NCCL2.25.1
NVIDIA GPUA40, A100, H100
PyTorch2.6

Model Weights

Model weights can be retrieved from Zenodo (fine-tuned only) or Hugging Face (pre-trained and fine-tuned).

Pre-training Dataset

MIST was pre-trained on Enamine REAL Space as provided by Enamine. We are currently working to secure permission to publish a subset of that dataset; however, MIST can be trained on a collection of text files of newline-delimited SMILES. We have uploaded an example dataset to Zenodo.

Fine-Tuning Datasets

All fine-tuning datasets are documented in Supplementary Section C.1. Datasets that have not already been publicly released elsewhere can be found on Zenodo.

Demonstration Code

Examples of using or training the MIST models can be found:

Submitting Jobs

We use a python script (submit/submit.py) to template training jobs for submission on HPC systems across multiple sites. Templates may need to be modified for your particular HPC cluster, but should provide a starting point.

source ./activate # Activate Environment
./submit/submit.py ./submit/polaris.j2 --data ./submit/pretrain.yaml | qsub

See submit/submit.py --help for more info

Note: ./activate is used to activate the python virtual environment and set various environment variables.

Pre-training and fine-tuning parameters are documented in our paper's methods section. The configuration for the MIST-1.8B is provided for reference. Configurations for all other models, including our fine-tuning configs, are documented in training logs.

Development

Pre-commit

We use pre-commit to preform various linting checks on the code. To enable:

  1. Install poetry (See above)
  2. Run pre-commit: uv run pre-commit
  3. Run before committing: uv run pre-commit install --allow-missing-config

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

MIST: Molecular Insight SMILES Transformer

GitHub LicensearXiv:2409.15370Model on HF

MIST is a family of molecular foundation models for molecular property prediction. The models were pre-trained on Smirk 😏 tokenized SMILES strings from the Enamine REAL Space dataset using the Masked Language Modeling (MLM) objective, then fine-tuned for downstream prediction tasks.

Installation

The following provides installation instructions for the top-level package (electrolyte_fm), optional add-ons for our various additional analysis and downstream applications (See ./opt may require additional configuration.

  1. Install uv and julia (only needed for /opt tasks)
  2. Instantiate the environment: uv sync
  3. Use submit/submit.py to submit a training job or checkout one of our applications in ./opt

You may need to install rust if pre-built wheels for smirk are not available on PyPI. Feel free to open an issue to request additional pre-built wheels.

Installation took under a minute on a MacBook Pro 2023. However, installation times depend heavily on internet connection and local uv cache. Pulling the NVIDIA packages can easily add 5-10 minutes without a uv cache. Installing the Julia environments (for example in /opt) can takes ~5 mins but may take longer on limited hardware (e.g. GitHub runners take ~18 mins).

Local

  1. Install rust and uv
  2. Run uv sync

Not all configurations will work locally. For example, pre-training configuration files typically use DeepSpeed, which requires NVIDIA GPUs. Some adaptation may be required to run outside of an NVIDIA GPU cluster.

Polaris

  1. Install rust and uv

  2. Load conda

module purge
module use /soft/modulefiles/
module --ignore_cache load conda/2024-04-29
conda activate base
  1. Install the environment
uv sync

Artemis

Same as above except:

  1. Skip loading conda (just use uv)
  2. Ensure a module for CUDA@12.2 exists, may need to install with spack (make sure buildable: True)

Apptainer

  1. Install or load from a module Apptainer
  2. Build the image bash container/build.sh, once build relocate the image mv /tmp/mist.sif ./mist.sif
  3. Run training within the image apptainer run --nv mist.sif python train.py ...

See submit/dgx.j2 or submit/delta.j2 for a more complete example of using the container

System Requirements

Hardware

Generally, running the code here requires access to GPUs and ideally a dedicated NVIDIA GPU cluster. However, much of the (non-training) code can be run on a single NVIDIA A40 GPU. Notably, for the MIST-28M model, CPU or MPS inference is viable.

Software Dependencies

All software dependencies can be found in the pyproject.toml (python) or Project.toml (julia) files. Specific versions are detailed in the lock files.

MIST has been tested on the following primary dependencies:

DependencyVersions
Ubuntu24.04
Python3.12, 3.13
NVIDIA CUDA12.8.0.038
NVIDIA cuBLAS12.8.3.14
NVIDIA cuDNN9.7.0.66
NVIDIA NCCL2.25.1
NVIDIA GPUA40, A100, H100
PyTorch2.6

Model Weights

Model weights can be retrieved from Zenodo (fine-tuned only) or Hugging Face (pre-trained and fine-tuned).

Pre-training Dataset

MIST was pre-trained on Enamine REAL Space as provided by Enamine. We are currently working to secure permission to publish a subset of that dataset; however, MIST can be trained on a collection of text files of newline-delimited SMILES. We have uploaded an example dataset to Zenodo.

Fine-Tuning Datasets

All fine-tuning datasets are documented in Supplementary Section C.1. Datasets that have not already been publicly released elsewhere can be found on Zenodo.

Demonstration Code

Examples of using or training the MIST models can be found:

Submitting Jobs

We use a python script (submit/submit.py) to template training jobs for submission on HPC systems across multiple sites. Templates may need to be modified for your particular HPC cluster, but should provide a starting point.

source ./activate # Activate Environment
./submit/submit.py ./submit/polaris.j2 --data ./submit/pretrain.yaml | qsub

See submit/submit.py --help for more info

Note: ./activate is used to activate the python virtual environment and set various environment variables.

Pre-training and fine-tuning parameters are documented in our paper's methods section. The configuration for the MIST-1.8B is provided for reference. Configurations for all other models, including our fine-tuning configs, are documented in training logs.

Development

Pre-commit

We use pre-commit to preform various linting checks on the code. To enable:

  1. Install poetry (See above)
  2. Run pre-commit: uv run pre-commit
  3. Run before committing: uv run pre-commit install --allow-missing-config

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

MIST: Molecular Insight SMILES Transformer

GitHub LicensearXiv:2409.15370Model on HF

MIST is a family of molecular foundation models for molecular property prediction. The models were pre-trained on Smirk 😏 tokenized SMILES strings from the Enamine REAL Space dataset using the Masked Language Modeling (MLM) objective, then fine-tuned for downstream prediction tasks.

Installation

The following provides installation instructions for the top-level package (electrolyte_fm), optional add-ons for our various additional analysis and downstream applications (See ./opt may require additional configuration.

  1. Install uv and julia (only needed for /opt tasks)
  2. Instantiate the environment: uv sync
  3. Use submit/submit.py to submit a training job or checkout one of our applications in ./opt

You may need to install rust if pre-built wheels for smirk are not available on PyPI. Feel free to open an issue to request additional pre-built wheels.

Installation took under a minute on a MacBook Pro 2023. However, installation times depend heavily on internet connection and local uv cache. Pulling the NVIDIA packages can easily add 5-10 minutes without a uv cache. Installing the Julia environments (for example in /opt) can takes ~5 mins but may take longer on limited hardware (e.g. GitHub runners take ~18 mins).

Local

  1. Install rust and uv
  2. Run uv sync

Not all configurations will work locally. For example, pre-training configuration files typically use DeepSpeed, which requires NVIDIA GPUs. Some adaptation may be required to run outside of an NVIDIA GPU cluster.

Polaris

  1. Install rust and uv

  2. Load conda

module purge
module use /soft/modulefiles/
module --ignore_cache load conda/2024-04-29
conda activate base
  1. Install the environment
uv sync

Artemis

Same as above except:

  1. Skip loading conda (just use uv)
  2. Ensure a module for CUDA@12.2 exists, may need to install with spack (make sure buildable: True)

Apptainer

  1. Install or load from a module Apptainer
  2. Build the image bash container/build.sh, once build relocate the image mv /tmp/mist.sif ./mist.sif
  3. Run training within the image apptainer run --nv mist.sif python train.py ...

See submit/dgx.j2 or submit/delta.j2 for a more complete example of using the container

System Requirements

Hardware

Generally, running the code here requires access to GPUs and ideally a dedicated NVIDIA GPU cluster. However, much of the (non-training) code can be run on a single NVIDIA A40 GPU. Notably, for the MIST-28M model, CPU or MPS inference is viable.

Software Dependencies

All software dependencies can be found in the pyproject.toml (python) or Project.toml (julia) files. Specific versions are detailed in the lock files.

MIST has been tested on the following primary dependencies:

DependencyVersions
Ubuntu24.04
Python3.12, 3.13
NVIDIA CUDA12.8.0.038
NVIDIA cuBLAS12.8.3.14
NVIDIA cuDNN9.7.0.66
NVIDIA NCCL2.25.1
NVIDIA GPUA40, A100, H100
PyTorch2.6

Model Weights

Model weights can be retrieved from Zenodo (fine-tuned only) or Hugging Face (pre-trained and fine-tuned).

Pre-training Dataset

MIST was pre-trained on Enamine REAL Space as provided by Enamine. We are currently working to secure permission to publish a subset of that dataset; however, MIST can be trained on a collection of text files of newline-delimited SMILES. We have uploaded an example dataset to Zenodo.

Fine-Tuning Datasets

All fine-tuning datasets are documented in Supplementary Section C.1. Datasets that have not already been publicly released elsewhere can be found on Zenodo.

Demonstration Code

Examples of using or training the MIST models can be found:

Submitting Jobs

We use a python script (submit/submit.py) to template training jobs for submission on HPC systems across multiple sites. Templates may need to be modified for your particular HPC cluster, but should provide a starting point.

source ./activate # Activate Environment
./submit/submit.py ./submit/polaris.j2 --data ./submit/pretrain.yaml | qsub

See submit/submit.py --help for more info

Note: ./activate is used to activate the python virtual environment and set various environment variables.

Pre-training and fine-tuning parameters are documented in our paper's methods section. The configuration for the MIST-1.8B is provided for reference. Configurations for all other models, including our fine-tuning configs, are documented in training logs.

Development

Pre-commit

We use pre-commit to preform various linting checks on the code. To enable:

  1. Install poetry (See above)
  2. Run pre-commit: uv run pre-commit
  3. Run before committing: uv run pre-commit install --allow-missing-config

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

MIST: Molecular Insight SMILES Transformer

GitHub LicensearXiv:2409.15370Model on HF

MIST is a family of molecular foundation models for molecular property prediction. The models were pre-trained on Smirk 😏 tokenized SMILES strings from the Enamine REAL Space dataset using the Masked Language Modeling (MLM) objective, then fine-tuned for downstream prediction tasks.

Installation

The following provides installation instructions for the top-level package (electrolyte_fm), optional add-ons for our various additional analysis and downstream applications (See ./opt may require additional configuration.

  1. Install uv and julia (only needed for /opt tasks)
  2. Instantiate the environment: uv sync
  3. Use submit/submit.py to submit a training job or checkout one of our applications in ./opt

You may need to install rust if pre-built wheels for smirk are not available on PyPI. Feel free to open an issue to request additional pre-built wheels.

Installation took under a minute on a MacBook Pro 2023. However, installation times depend heavily on internet connection and local uv cache. Pulling the NVIDIA packages can easily add 5-10 minutes without a uv cache. Installing the Julia environments (for example in /opt) can takes ~5 mins but may take longer on limited hardware (e.g. GitHub runners take ~18 mins).

Local

  1. Install rust and uv
  2. Run uv sync

Not all configurations will work locally. For example, pre-training configuration files typically use DeepSpeed, which requires NVIDIA GPUs. Some adaptation may be required to run outside of an NVIDIA GPU cluster.

Polaris

  1. Install rust and uv

  2. Load conda

module purge
module use /soft/modulefiles/
module --ignore_cache load conda/2024-04-29
conda activate base
  1. Install the environment
uv sync

Artemis

Same as above except:

  1. Skip loading conda (just use uv)
  2. Ensure a module for CUDA@12.2 exists, may need to install with spack (make sure buildable: True)

Apptainer

  1. Install or load from a module Apptainer
  2. Build the image bash container/build.sh, once build relocate the image mv /tmp/mist.sif ./mist.sif
  3. Run training within the image apptainer run --nv mist.sif python train.py ...

See submit/dgx.j2 or submit/delta.j2 for a more complete example of using the container

System Requirements

Hardware

Generally, running the code here requires access to GPUs and ideally a dedicated NVIDIA GPU cluster. However, much of the (non-training) code can be run on a single NVIDIA A40 GPU. Notably, for the MIST-28M model, CPU or MPS inference is viable.

Software Dependencies

All software dependencies can be found in the pyproject.toml (python) or Project.toml (julia) files. Specific versions are detailed in the lock files.

MIST has been tested on the following primary dependencies:

DependencyVersions
Ubuntu24.04
Python3.12, 3.13
NVIDIA CUDA12.8.0.038
NVIDIA cuBLAS12.8.3.14
NVIDIA cuDNN9.7.0.66
NVIDIA NCCL2.25.1
NVIDIA GPUA40, A100, H100
PyTorch2.6

Model Weights

Model weights can be retrieved from Zenodo (fine-tuned only) or Hugging Face (pre-trained and fine-tuned).

Pre-training Dataset

MIST was pre-trained on Enamine REAL Space as provided by Enamine. We are currently working to secure permission to publish a subset of that dataset; however, MIST can be trained on a collection of text files of newline-delimited SMILES. We have uploaded an example dataset to Zenodo.

Fine-Tuning Datasets

All fine-tuning datasets are documented in Supplementary Section C.1. Datasets that have not already been publicly released elsewhere can be found on Zenodo.

Demonstration Code

Examples of using or training the MIST models can be found:

Submitting Jobs

We use a python script (submit/submit.py) to template training jobs for submission on HPC systems across multiple sites. Templates may need to be modified for your particular HPC cluster, but should provide a starting point.

source ./activate # Activate Environment
./submit/submit.py ./submit/polaris.j2 --data ./submit/pretrain.yaml | qsub

See submit/submit.py --help for more info

Note: ./activate is used to activate the python virtual environment and set various environment variables.

Pre-training and fine-tuning parameters are documented in our paper's methods section. The configuration for the MIST-1.8B is provided for reference. Configurations for all other models, including our fine-tuning configs, are documented in training logs.

Development

Pre-commit

We use pre-commit to preform various linting checks on the code. To enable:

  1. Install poetry (See above)
  2. Run pre-commit: uv run pre-commit
  3. Run before committing: uv run pre-commit install --allow-missing-config

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

MIST: Molecular Insight SMILES Transformer

GitHub LicensearXiv:2409.15370Model on HF

MIST is a family of molecular foundation models for molecular property prediction. The models were pre-trained on Smirk 😏 tokenized SMILES strings from the Enamine REAL Space dataset using the Masked Language Modeling (MLM) objective, then fine-tuned for downstream prediction tasks.

Installation

The following provides installation instructions for the top-level package (electrolyte_fm), optional add-ons for our various additional analysis and downstream applications (See ./opt may require additional configuration.

  1. Install uv and julia (only needed for /opt tasks)
  2. Instantiate the environment: uv sync
  3. Use submit/submit.py to submit a training job or checkout one of our applications in ./opt

You may need to install rust if pre-built wheels for smirk are not available on PyPI. Feel free to open an issue to request additional pre-built wheels.

Installation took under a minute on a MacBook Pro 2023. However, installation times depend heavily on internet connection and local uv cache. Pulling the NVIDIA packages can easily add 5-10 minutes without a uv cache. Installing the Julia environments (for example in /opt) can takes ~5 mins but may take longer on limited hardware (e.g. GitHub runners take ~18 mins).

Local

  1. Install rust and uv
  2. Run uv sync

Not all configurations will work locally. For example, pre-training configuration files typically use DeepSpeed, which requires NVIDIA GPUs. Some adaptation may be required to run outside of an NVIDIA GPU cluster.

Polaris

  1. Install rust and uv

  2. Load conda

module purge
module use /soft/modulefiles/
module --ignore_cache load conda/2024-04-29
conda activate base
  1. Install the environment
uv sync

Artemis

Same as above except:

  1. Skip loading conda (just use uv)
  2. Ensure a module for CUDA@12.2 exists, may need to install with spack (make sure buildable: True)

Apptainer

  1. Install or load from a module Apptainer
  2. Build the image bash container/build.sh, once build relocate the image mv /tmp/mist.sif ./mist.sif
  3. Run training within the image apptainer run --nv mist.sif python train.py ...

See submit/dgx.j2 or submit/delta.j2 for a more complete example of using the container

System Requirements

Hardware

Generally, running the code here requires access to GPUs and ideally a dedicated NVIDIA GPU cluster. However, much of the (non-training) code can be run on a single NVIDIA A40 GPU. Notably, for the MIST-28M model, CPU or MPS inference is viable.

Software Dependencies

All software dependencies can be found in the pyproject.toml (python) or Project.toml (julia) files. Specific versions are detailed in the lock files.

MIST has been tested on the following primary dependencies:

DependencyVersions
Ubuntu24.04
Python3.12, 3.13
NVIDIA CUDA12.8.0.038
NVIDIA cuBLAS12.8.3.14
NVIDIA cuDNN9.7.0.66
NVIDIA NCCL2.25.1
NVIDIA GPUA40, A100, H100
PyTorch2.6

Model Weights

Model weights can be retrieved from Zenodo (fine-tuned only) or Hugging Face (pre-trained and fine-tuned).

Pre-training Dataset

MIST was pre-trained on Enamine REAL Space as provided by Enamine. We are currently working to secure permission to publish a subset of that dataset; however, MIST can be trained on a collection of text files of newline-delimited SMILES. We have uploaded an example dataset to Zenodo.

Fine-Tuning Datasets

All fine-tuning datasets are documented in Supplementary Section C.1. Datasets that have not already been publicly released elsewhere can be found on Zenodo.

Demonstration Code

Examples of using or training the MIST models can be found:

Submitting Jobs

We use a python script (submit/submit.py) to template training jobs for submission on HPC systems across multiple sites. Templates may need to be modified for your particular HPC cluster, but should provide a starting point.

source ./activate # Activate Environment
./submit/submit.py ./submit/polaris.j2 --data ./submit/pretrain.yaml | qsub

See submit/submit.py --help for more info

Note: ./activate is used to activate the python virtual environment and set various environment variables.

Pre-training and fine-tuning parameters are documented in our paper's methods section. The configuration for the MIST-1.8B is provided for reference. Configurations for all other models, including our fine-tuning configs, are documented in training logs.

Development

Pre-commit

We use pre-commit to preform various linting checks on the code. To enable:

  1. Install poetry (See above)
  2. Run pre-commit: uv run pre-commit
  3. Run before committing: uv run pre-commit install --allow-missing-config

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

MIST: Molecular Insight SMILES Transformer

GitHub LicensearXiv:2409.15370Model on HF

MIST is a family of molecular foundation models for molecular property prediction. The models were pre-trained on Smirk 😏 tokenized SMILES strings from the Enamine REAL Space dataset using the Masked Language Modeling (MLM) objective, then fine-tuned for downstream prediction tasks.

Installation

The following provides installation instructions for the top-level package (electrolyte_fm), optional add-ons for our various additional analysis and downstream applications (See ./opt may require additional configuration.

  1. Install uv and julia (only needed for /opt tasks)
  2. Instantiate the environment: uv sync
  3. Use submit/submit.py to submit a training job or checkout one of our applications in ./opt

You may need to install rust if pre-built wheels for smirk are not available on PyPI. Feel free to open an issue to request additional pre-built wheels.

Installation took under a minute on a MacBook Pro 2023. However, installation times depend heavily on internet connection and local uv cache. Pulling the NVIDIA packages can easily add 5-10 minutes without a uv cache. Installing the Julia environments (for example in /opt) can takes ~5 mins but may take longer on limited hardware (e.g. GitHub runners take ~18 mins).

Local

  1. Install rust and uv
  2. Run uv sync

Not all configurations will work locally. For example, pre-training configuration files typically use DeepSpeed, which requires NVIDIA GPUs. Some adaptation may be required to run outside of an NVIDIA GPU cluster.

Polaris

  1. Install rust and uv

  2. Load conda

module purge
module use /soft/modulefiles/
module --ignore_cache load conda/2024-04-29
conda activate base
  1. Install the environment
uv sync

Artemis

Same as above except:

  1. Skip loading conda (just use uv)
  2. Ensure a module for CUDA@12.2 exists, may need to install with spack (make sure buildable: True)

Apptainer

  1. Install or load from a module Apptainer
  2. Build the image bash container/build.sh, once build relocate the image mv /tmp/mist.sif ./mist.sif
  3. Run training within the image apptainer run --nv mist.sif python train.py ...

See submit/dgx.j2 or submit/delta.j2 for a more complete example of using the container

System Requirements

Hardware

Generally, running the code here requires access to GPUs and ideally a dedicated NVIDIA GPU cluster. However, much of the (non-training) code can be run on a single NVIDIA A40 GPU. Notably, for the MIST-28M model, CPU or MPS inference is viable.

Software Dependencies

All software dependencies can be found in the pyproject.toml (python) or Project.toml (julia) files. Specific versions are detailed in the lock files.

MIST has been tested on the following primary dependencies:

DependencyVersions
Ubuntu24.04
Python3.12, 3.13
NVIDIA CUDA12.8.0.038
NVIDIA cuBLAS12.8.3.14
NVIDIA cuDNN9.7.0.66
NVIDIA NCCL2.25.1
NVIDIA GPUA40, A100, H100
PyTorch2.6

Model Weights

Model weights can be retrieved from Zenodo (fine-tuned only) or Hugging Face (pre-trained and fine-tuned).

Pre-training Dataset

MIST was pre-trained on Enamine REAL Space as provided by Enamine. We are currently working to secure permission to publish a subset of that dataset; however, MIST can be trained on a collection of text files of newline-delimited SMILES. We have uploaded an example dataset to Zenodo.

Fine-Tuning Datasets

All fine-tuning datasets are documented in Supplementary Section C.1. Datasets that have not already been publicly released elsewhere can be found on Zenodo.

Demonstration Code

Examples of using or training the MIST models can be found:

Submitting Jobs

We use a python script (submit/submit.py) to template training jobs for submission on HPC systems across multiple sites. Templates may need to be modified for your particular HPC cluster, but should provide a starting point.

source ./activate # Activate Environment
./submit/submit.py ./submit/polaris.j2 --data ./submit/pretrain.yaml | qsub

See submit/submit.py --help for more info

Note: ./activate is used to activate the python virtual environment and set various environment variables.

Pre-training and fine-tuning parameters are documented in our paper's methods section. The configuration for the MIST-1.8B is provided for reference. Configurations for all other models, including our fine-tuning configs, are documented in training logs.

Development

Pre-commit

We use pre-commit to preform various linting checks on the code. To enable:

  1. Install poetry (See above)
  2. Run pre-commit: uv run pre-commit
  3. Run before committing: uv run pre-commit install --allow-missing-config

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

MIST: Molecular Insight SMILES Transformer

GitHub LicensearXiv:2409.15370Model on HF

MIST is a family of molecular foundation models for molecular property prediction. The models were pre-trained on Smirk 😏 tokenized SMILES strings from the Enamine REAL Space dataset using the Masked Language Modeling (MLM) objective, then fine-tuned for downstream prediction tasks.

Installation

The following provides installation instructions for the top-level package (electrolyte_fm), optional add-ons for our various additional analysis and downstream applications (See ./opt may require additional configuration.

  1. Install uv and julia (only needed for /opt tasks)
  2. Instantiate the environment: uv sync
  3. Use submit/submit.py to submit a training job or checkout one of our applications in ./opt

You may need to install rust if pre-built wheels for smirk are not available on PyPI. Feel free to open an issue to request additional pre-built wheels.

Installation took under a minute on a MacBook Pro 2023. However, installation times depend heavily on internet connection and local uv cache. Pulling the NVIDIA packages can easily add 5-10 minutes without a uv cache. Installing the Julia environments (for example in /opt) can takes ~5 mins but may take longer on limited hardware (e.g. GitHub runners take ~18 mins).

Local

  1. Install rust and uv
  2. Run uv sync

Not all configurations will work locally. For example, pre-training configuration files typically use DeepSpeed, which requires NVIDIA GPUs. Some adaptation may be required to run outside of an NVIDIA GPU cluster.

Polaris

  1. Install rust and uv

  2. Load conda

module purge
module use /soft/modulefiles/
module --ignore_cache load conda/2024-04-29
conda activate base
  1. Install the environment
uv sync

Artemis

Same as above except:

  1. Skip loading conda (just use uv)
  2. Ensure a module for CUDA@12.2 exists, may need to install with spack (make sure buildable: True)

Apptainer

  1. Install or load from a module Apptainer
  2. Build the image bash container/build.sh, once build relocate the image mv /tmp/mist.sif ./mist.sif
  3. Run training within the image apptainer run --nv mist.sif python train.py ...

See submit/dgx.j2 or submit/delta.j2 for a more complete example of using the container

System Requirements

Hardware

Generally, running the code here requires access to GPUs and ideally a dedicated NVIDIA GPU cluster. However, much of the (non-training) code can be run on a single NVIDIA A40 GPU. Notably, for the MIST-28M model, CPU or MPS inference is viable.

Software Dependencies

All software dependencies can be found in the pyproject.toml (python) or Project.toml (julia) files. Specific versions are detailed in the lock files.

MIST has been tested on the following primary dependencies:

DependencyVersions
Ubuntu24.04
Python3.12, 3.13
NVIDIA CUDA12.8.0.038
NVIDIA cuBLAS12.8.3.14
NVIDIA cuDNN9.7.0.66
NVIDIA NCCL2.25.1
NVIDIA GPUA40, A100, H100
PyTorch2.6

Model Weights

Model weights can be retrieved from Zenodo (fine-tuned only) or Hugging Face (pre-trained and fine-tuned).

Pre-training Dataset

MIST was pre-trained on Enamine REAL Space as provided by Enamine. We are currently working to secure permission to publish a subset of that dataset; however, MIST can be trained on a collection of text files of newline-delimited SMILES. We have uploaded an example dataset to Zenodo.

Fine-Tuning Datasets

All fine-tuning datasets are documented in Supplementary Section C.1. Datasets that have not already been publicly released elsewhere can be found on Zenodo.

Demonstration Code

Examples of using or training the MIST models can be found:

Submitting Jobs

We use a python script (submit/submit.py) to template training jobs for submission on HPC systems across multiple sites. Templates may need to be modified for your particular HPC cluster, but should provide a starting point.

source ./activate # Activate Environment
./submit/submit.py ./submit/polaris.j2 --data ./submit/pretrain.yaml | qsub

See submit/submit.py --help for more info

Note: ./activate is used to activate the python virtual environment and set various environment variables.

Pre-training and fine-tuning parameters are documented in our paper's methods section. The configuration for the MIST-1.8B is provided for reference. Configurations for all other models, including our fine-tuning configs, are documented in training logs.

Development

Pre-commit

We use pre-commit to preform various linting checks on the code. To enable:

  1. Install poetry (See above)
  2. Run pre-commit: uv run pre-commit
  3. Run before committing: uv run pre-commit install --allow-missing-config

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

MIST: Molecular Insight SMILES Transformer

GitHub LicensearXiv:2409.15370Model on HF

MIST is a family of molecular foundation models for molecular property prediction. The models were pre-trained on Smirk 😏 tokenized SMILES strings from the Enamine REAL Space dataset using the Masked Language Modeling (MLM) objective, then fine-tuned for downstream prediction tasks.

Installation

The following provides installation instructions for the top-level package (electrolyte_fm), optional add-ons for our various additional analysis and downstream applications (See ./opt may require additional configuration.

  1. Install uv and julia (only needed for /opt tasks)
  2. Instantiate the environment: uv sync
  3. Use submit/submit.py to submit a training job or checkout one of our applications in ./opt

You may need to install rust if pre-built wheels for smirk are not available on PyPI. Feel free to open an issue to request additional pre-built wheels.

Installation took under a minute on a MacBook Pro 2023. However, installation times depend heavily on internet connection and local uv cache. Pulling the NVIDIA packages can easily add 5-10 minutes without a uv cache. Installing the Julia environments (for example in /opt) can takes ~5 mins but may take longer on limited hardware (e.g. GitHub runners take ~18 mins).

Local

  1. Install rust and uv
  2. Run uv sync

Not all configurations will work locally. For example, pre-training configuration files typically use DeepSpeed, which requires NVIDIA GPUs. Some adaptation may be required to run outside of an NVIDIA GPU cluster.

Polaris

  1. Install rust and uv

  2. Load conda

module purge
module use /soft/modulefiles/
module --ignore_cache load conda/2024-04-29
conda activate base
  1. Install the environment
uv sync

Artemis

Same as above except:

  1. Skip loading conda (just use uv)
  2. Ensure a module for CUDA@12.2 exists, may need to install with spack (make sure buildable: True)

Apptainer

  1. Install or load from a module Apptainer
  2. Build the image bash container/build.sh, once build relocate the image mv /tmp/mist.sif ./mist.sif
  3. Run training within the image apptainer run --nv mist.sif python train.py ...

See submit/dgx.j2 or submit/delta.j2 for a more complete example of using the container

System Requirements

Hardware

Generally, running the code here requires access to GPUs and ideally a dedicated NVIDIA GPU cluster. However, much of the (non-training) code can be run on a single NVIDIA A40 GPU. Notably, for the MIST-28M model, CPU or MPS inference is viable.

Software Dependencies

All software dependencies can be found in the pyproject.toml (python) or Project.toml (julia) files. Specific versions are detailed in the lock files.

MIST has been tested on the following primary dependencies:

DependencyVersions
Ubuntu24.04
Python3.12, 3.13
NVIDIA CUDA12.8.0.038
NVIDIA cuBLAS12.8.3.14
NVIDIA cuDNN9.7.0.66
NVIDIA NCCL2.25.1
NVIDIA GPUA40, A100, H100
PyTorch2.6

Model Weights

Model weights can be retrieved from Zenodo (fine-tuned only) or Hugging Face (pre-trained and fine-tuned).

Pre-training Dataset

MIST was pre-trained on Enamine REAL Space as provided by Enamine. We are currently working to secure permission to publish a subset of that dataset; however, MIST can be trained on a collection of text files of newline-delimited SMILES. We have uploaded an example dataset to Zenodo.

Fine-Tuning Datasets

All fine-tuning datasets are documented in Supplementary Section C.1. Datasets that have not already been publicly released elsewhere can be found on Zenodo.

Demonstration Code

Examples of using or training the MIST models can be found:

Submitting Jobs

We use a python script (submit/submit.py) to template training jobs for submission on HPC systems across multiple sites. Templates may need to be modified for your particular HPC cluster, but should provide a starting point.

source ./activate # Activate Environment
./submit/submit.py ./submit/polaris.j2 --data ./submit/pretrain.yaml | qsub

See submit/submit.py --help for more info

Note: ./activate is used to activate the python virtual environment and set various environment variables.

Pre-training and fine-tuning parameters are documented in our paper's methods section. The configuration for the MIST-1.8B is provided for reference. Configurations for all other models, including our fine-tuning configs, are documented in training logs.

Development

Pre-commit

We use pre-commit to preform various linting checks on the code. To enable:

  1. Install poetry (See above)
  2. Run pre-commit: uv run pre-commit
  3. Run before committing: uv run pre-commit install --allow-missing-config

Releases

Packages

Used by

Contributors

Languages