Repository files navigation

SynNet

This repo contains the code and analysis scripts for our amortized approach to synthetic tree generation using neural networks. Our model can serve as both a synthesis planning tool and as a tool for synthesizable molecular design.

The method is described in detail in the publication "Amortized tree generation for bottom-up synthesis planning and synthesizable molecular design" available on the arXiv and summarized below.

Summary

We model synthetic pathways as tree structures called synthetic trees. A synthetic tree has a single root node and one or more child nodes. Every node is chemical molecule:

  • The root node is the final product molecule
  • The leaf nodes consist of purchasable building blocks.
  • All other inner nodes are constrained to be a product of allowed chemical reactions.

At a high level, each synthetic tree is constructed one reaction step at a time in a bottom-up manner, that is starting from purchasable building blocks.

Overview

The model consists of four modules, each containing a multi-layer perceptron (MLP):

  1. An Action Type selection function that classifies action types among the four possible actions (“Add”, “Expand”, “Merge”, and “End”) in building the synthetic tree. Each action increases the depth of the synthetic tree by one.

  2. A First Reactant selection function that selects the first reactant. A MLP predicts a molecular embedding and a first reactant is identified from the pool of building blocks through a k-nearest neighbors (k-NN) search.

  3. A Reaction selection function whose output is a probability distribution over available reaction templates. Inapplicable reactions are masked based on reactant 1. A suitable template is then sampled using a greedy search.

  4. A Second Reactant selection function that identifies the second reactant if the sampled template is bi-molecular. The model predicts an embedding for the second reactant, and a candidate is then sampled via a k-NN search from the masked set of building blocks.

the model

These four modules predict the probability distributions of actions to be taken within a single reaction step, and determine the nodes to be added to the synthetic tree under construction. All of these networks are conditioned on the target molecule embedding.

Synthesis planning

This task is to infer the synthetic pathway to a given target molecule. We formulate this problem as generating a synthetic tree such that the product molecule it produces (i.e., the molecule at the root node) matches the desired target molecule.

For this task, we can take a molecular embedding for the desired product, and use it as input to our model to produce a synthetic tree. If the desired product is successfully recovered, then the final root molecule will match the desired molecule used to create the input embedding. If the desired product is not successully recovered, it is possible the final root molecule may still be similar to the desired molecule used to create the input embedding, and thus our tool can also be used for synthesizable analog recommendation.

the generation process

Synthesizable molecular design

This task is to optimize a molecular structure with respect to an oracle function (e.g. bioactivity), while ensuring the synthetic accessibility of the molecules. We formulate this problem as optimizing the structure of a synthetic tree with respect to the desired properties of the product molecule it produces.

To do this, we optimize the molecular embedding of the molecule using a genetic algorithm and the desired oracle function. The optimized molecule embedding can then be used as input to our model to produce a synthetic tree, where the final root molecule corresponds to the optimized molecule.

Setup instructions

Environment

Conda is used to create the environment for running SynNet.

# Install environment from file
conda env create -f environment.yml

Before running any SynNet code, activate the environment and install this package in development mode:

source activate synnet
pip install -e .

The model implementations can be found in src/syn_net/models/.

The pre-processing and analysis scripts are in scripts/.

Train the model from scratch

Before training any models, you will first need to some data preprocessing. Please see INSTRUCTIONS.md for a complete guide.

Data

SynNet relies on two datasources:

  1. reaction templates and
  2. building blocks.

The data used for the publication are 1) the Hartenfeller-Button reaction templates, which are available under data/assets/reaction-templates/hb.txt and 2) Enamine building blocks. The building blocks are not freely available.

To obtain the data, go to https://enamine.net/building-blocks/building-blocks-catalog. We used the "Building Blocks, US Stock" data. You need to first register and then request access to download the dataset. The people from enamine.net manually approve you, so please be nice and patient.

Reproducing results

Before running anything, set up the environment as decribed above.

Using pre-trained models

We have made available a set of pre-trained models at the following link. The pretrained models correspond to the Action, Reactant 1, Reaction, and Reactant 2 networks, trained on the Hartenfeller-Button dataset and Enamine building blocks using radius 2, length 4096 Morgan fingerprints for the molecular node embeddings, and length 256 fingerprints for the k-NN search. For further details, please see the publication.

To download the pre-trained model to ./checkpoints:

# Download
wget -O hb_fp_2_4096_256.tar.gz https://figshare.com/ndownloader/files/31067692
# Extract
tar -vxf hb_fp_2_4096_256.tar.gz
# Rename files to match new scripts (...)
mv hb_fp_2_4096_256/ checkpoints/
formodelin"act""rt1""rxn""rt2"do
mkdir checkpoints/$model
mv "checkpoints/$model.ckpt""checkpoints/$model/ckpts.dummy-val_loss=0.00.ckpt"done
rm -f hb_fp_2_4096_256.tar.gz

The following scripts are run from the command line. Use python some_script.py --help or check the source code to see the instructions of each argument.

Prerequisites

In addition to the necessary data, we will need to pre-compute an embedding of the building blocks. To do so, please follow steps 0-2 from the INSTRUCTIONS.md. Then, replace the environment variables in the commands below.

Synthesis Planning

To perform synthesis planning described in the main text:

python scripts/20-predict-targets.py \
--building-blocks-file $BUILDING_BLOCKS_FILE \
--rxns-collection-file $RXN_COLLECTION_FILE \
--embeddings-knn-file $EMBEDDINGS_KNN_FILE \
--data "data/assets/molecules/sample-targets.txt" \
--ckpt-dir "checkpoints/" \
--output-dir "results/demo-inference/"

This script will feed a list of ten molecules to SynNet.

Synthesizable Molecular Design

To perform synthesizable molecular design, run:

python scripts/optimize_ga.py \
--ckpt-dir "checkpoints/" \
--building-blocks-file $BUILDING_BLOCKS_FILE \
--rxns-collection-file $RXN_COLLECTION_FILE \
--embeddings-knn-file $EMBEDDINGS_KNN_FILE \
--input-file path/to/zinc.csv \
--radius 2 --nbits 4096 \
--num_population 128 --num_offspring 512 --num_gen 200 --objective gsk \
--ncpu 32

This script uses a genetic algorithm to optimize molecular embeddings and returns the predicted synthetic trees for the optimized molecular embedding.

Note: input-file contains the seed molecules in CSV format for an initial run, and as a pre-saved numpy array of the population for restarting the run. If omitted, a random fingerprint will be chosen.

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

SynNet

This repo contains the code and analysis scripts for our amortized approach to synthetic tree generation using neural networks. Our model can serve as both a synthesis planning tool and as a tool for synthesizable molecular design.

The method is described in detail in the publication "Amortized tree generation for bottom-up synthesis planning and synthesizable molecular design" available on the arXiv and summarized below.

Summary

We model synthetic pathways as tree structures called synthetic trees. A synthetic tree has a single root node and one or more child nodes. Every node is chemical molecule:

  • The root node is the final product molecule
  • The leaf nodes consist of purchasable building blocks.
  • All other inner nodes are constrained to be a product of allowed chemical reactions.

At a high level, each synthetic tree is constructed one reaction step at a time in a bottom-up manner, that is starting from purchasable building blocks.

Overview

The model consists of four modules, each containing a multi-layer perceptron (MLP):

  1. An Action Type selection function that classifies action types among the four possible actions (“Add”, “Expand”, “Merge”, and “End”) in building the synthetic tree. Each action increases the depth of the synthetic tree by one.

  2. A First Reactant selection function that selects the first reactant. A MLP predicts a molecular embedding and a first reactant is identified from the pool of building blocks through a k-nearest neighbors (k-NN) search.

  3. A Reaction selection function whose output is a probability distribution over available reaction templates. Inapplicable reactions are masked based on reactant 1. A suitable template is then sampled using a greedy search.

  4. A Second Reactant selection function that identifies the second reactant if the sampled template is bi-molecular. The model predicts an embedding for the second reactant, and a candidate is then sampled via a k-NN search from the masked set of building blocks.

the model

These four modules predict the probability distributions of actions to be taken within a single reaction step, and determine the nodes to be added to the synthetic tree under construction. All of these networks are conditioned on the target molecule embedding.

Synthesis planning

This task is to infer the synthetic pathway to a given target molecule. We formulate this problem as generating a synthetic tree such that the product molecule it produces (i.e., the molecule at the root node) matches the desired target molecule.

For this task, we can take a molecular embedding for the desired product, and use it as input to our model to produce a synthetic tree. If the desired product is successfully recovered, then the final root molecule will match the desired molecule used to create the input embedding. If the desired product is not successully recovered, it is possible the final root molecule may still be similar to the desired molecule used to create the input embedding, and thus our tool can also be used for synthesizable analog recommendation.

the generation process

Synthesizable molecular design

This task is to optimize a molecular structure with respect to an oracle function (e.g. bioactivity), while ensuring the synthetic accessibility of the molecules. We formulate this problem as optimizing the structure of a synthetic tree with respect to the desired properties of the product molecule it produces.

To do this, we optimize the molecular embedding of the molecule using a genetic algorithm and the desired oracle function. The optimized molecule embedding can then be used as input to our model to produce a synthetic tree, where the final root molecule corresponds to the optimized molecule.

Setup instructions

Environment

Conda is used to create the environment for running SynNet.

# Install environment from file
conda env create -f environment.yml

Before running any SynNet code, activate the environment and install this package in development mode:

source activate synnet
pip install -e .

The model implementations can be found in src/syn_net/models/.

The pre-processing and analysis scripts are in scripts/.

Train the model from scratch

Before training any models, you will first need to some data preprocessing. Please see INSTRUCTIONS.md for a complete guide.

Data

SynNet relies on two datasources:

  1. reaction templates and
  2. building blocks.

The data used for the publication are 1) the Hartenfeller-Button reaction templates, which are available under data/assets/reaction-templates/hb.txt and 2) Enamine building blocks. The building blocks are not freely available.

To obtain the data, go to https://enamine.net/building-blocks/building-blocks-catalog. We used the "Building Blocks, US Stock" data. You need to first register and then request access to download the dataset. The people from enamine.net manually approve you, so please be nice and patient.

Reproducing results

Before running anything, set up the environment as decribed above.

Using pre-trained models

We have made available a set of pre-trained models at the following link. The pretrained models correspond to the Action, Reactant 1, Reaction, and Reactant 2 networks, trained on the Hartenfeller-Button dataset and Enamine building blocks using radius 2, length 4096 Morgan fingerprints for the molecular node embeddings, and length 256 fingerprints for the k-NN search. For further details, please see the publication.

To download the pre-trained model to ./checkpoints:

# Download
wget -O hb_fp_2_4096_256.tar.gz https://figshare.com/ndownloader/files/31067692
# Extract
tar -vxf hb_fp_2_4096_256.tar.gz
# Rename files to match new scripts (...)
mv hb_fp_2_4096_256/ checkpoints/
formodelin"act""rt1""rxn""rt2"do
mkdir checkpoints/$model
mv "checkpoints/$model.ckpt""checkpoints/$model/ckpts.dummy-val_loss=0.00.ckpt"done
rm -f hb_fp_2_4096_256.tar.gz

The following scripts are run from the command line. Use python some_script.py --help or check the source code to see the instructions of each argument.

Prerequisites

In addition to the necessary data, we will need to pre-compute an embedding of the building blocks. To do so, please follow steps 0-2 from the INSTRUCTIONS.md. Then, replace the environment variables in the commands below.

Synthesis Planning

To perform synthesis planning described in the main text:

python scripts/20-predict-targets.py \
--building-blocks-file $BUILDING_BLOCKS_FILE \
--rxns-collection-file $RXN_COLLECTION_FILE \
--embeddings-knn-file $EMBEDDINGS_KNN_FILE \
--data "data/assets/molecules/sample-targets.txt" \
--ckpt-dir "checkpoints/" \
--output-dir "results/demo-inference/"

This script will feed a list of ten molecules to SynNet.

Synthesizable Molecular Design

To perform synthesizable molecular design, run:

python scripts/optimize_ga.py \
--ckpt-dir "checkpoints/" \
--building-blocks-file $BUILDING_BLOCKS_FILE \
--rxns-collection-file $RXN_COLLECTION_FILE \
--embeddings-knn-file $EMBEDDINGS_KNN_FILE \
--input-file path/to/zinc.csv \
--radius 2 --nbits 4096 \
--num_population 128 --num_offspring 512 --num_gen 200 --objective gsk \
--ncpu 32

This script uses a genetic algorithm to optimize molecular embeddings and returns the predicted synthetic trees for the optimized molecular embedding.

Note: input-file contains the seed molecules in CSV format for an initial run, and as a pre-saved numpy array of the population for restarting the run. If omitted, a random fingerprint will be chosen.

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SynNet

This repo contains the code and analysis scripts for our amortized approach to synthetic tree generation using neural networks. Our model can serve as both a synthesis planning tool and as a tool for synthesizable molecular design.

The method is described in detail in the publication "Amortized tree generation for bottom-up synthesis planning and synthesizable molecular design" available on the arXiv and summarized below.

Summary

We model synthetic pathways as tree structures called synthetic trees. A synthetic tree has a single root node and one or more child nodes. Every node is chemical molecule:

  • The root node is the final product molecule
  • The leaf nodes consist of purchasable building blocks.
  • All other inner nodes are constrained to be a product of allowed chemical reactions.

At a high level, each synthetic tree is constructed one reaction step at a time in a bottom-up manner, that is starting from purchasable building blocks.

Overview

The model consists of four modules, each containing a multi-layer perceptron (MLP):

  1. An Action Type selection function that classifies action types among the four possible actions (“Add”, “Expand”, “Merge”, and “End”) in building the synthetic tree. Each action increases the depth of the synthetic tree by one.

  2. A First Reactant selection function that selects the first reactant. A MLP predicts a molecular embedding and a first reactant is identified from the pool of building blocks through a k-nearest neighbors (k-NN) search.

  3. A Reaction selection function whose output is a probability distribution over available reaction templates. Inapplicable reactions are masked based on reactant 1. A suitable template is then sampled using a greedy search.

  4. A Second Reactant selection function that identifies the second reactant if the sampled template is bi-molecular. The model predicts an embedding for the second reactant, and a candidate is then sampled via a k-NN search from the masked set of building blocks.

the model

These four modules predict the probability distributions of actions to be taken within a single reaction step, and determine the nodes to be added to the synthetic tree under construction. All of these networks are conditioned on the target molecule embedding.

Synthesis planning

This task is to infer the synthetic pathway to a given target molecule. We formulate this problem as generating a synthetic tree such that the product molecule it produces (i.e., the molecule at the root node) matches the desired target molecule.

For this task, we can take a molecular embedding for the desired product, and use it as input to our model to produce a synthetic tree. If the desired product is successfully recovered, then the final root molecule will match the desired molecule used to create the input embedding. If the desired product is not successully recovered, it is possible the final root molecule may still be similar to the desired molecule used to create the input embedding, and thus our tool can also be used for synthesizable analog recommendation.

the generation process

Synthesizable molecular design

This task is to optimize a molecular structure with respect to an oracle function (e.g. bioactivity), while ensuring the synthetic accessibility of the molecules. We formulate this problem as optimizing the structure of a synthetic tree with respect to the desired properties of the product molecule it produces.

To do this, we optimize the molecular embedding of the molecule using a genetic algorithm and the desired oracle function. The optimized molecule embedding can then be used as input to our model to produce a synthetic tree, where the final root molecule corresponds to the optimized molecule.

Setup instructions

Environment

Conda is used to create the environment for running SynNet.

# Install environment from file
conda env create -f environment.yml

Before running any SynNet code, activate the environment and install this package in development mode:

source activate synnet
pip install -e .

The model implementations can be found in src/syn_net/models/.

The pre-processing and analysis scripts are in scripts/.

Train the model from scratch

Before training any models, you will first need to some data preprocessing. Please see INSTRUCTIONS.md for a complete guide.

Data

SynNet relies on two datasources:

  1. reaction templates and
  2. building blocks.

The data used for the publication are 1) the Hartenfeller-Button reaction templates, which are available under data/assets/reaction-templates/hb.txt and 2) Enamine building blocks. The building blocks are not freely available.

To obtain the data, go to https://enamine.net/building-blocks/building-blocks-catalog. We used the "Building Blocks, US Stock" data. You need to first register and then request access to download the dataset. The people from enamine.net manually approve you, so please be nice and patient.

Reproducing results

Before running anything, set up the environment as decribed above.

Using pre-trained models

We have made available a set of pre-trained models at the following link. The pretrained models correspond to the Action, Reactant 1, Reaction, and Reactant 2 networks, trained on the Hartenfeller-Button dataset and Enamine building blocks using radius 2, length 4096 Morgan fingerprints for the molecular node embeddings, and length 256 fingerprints for the k-NN search. For further details, please see the publication.

To download the pre-trained model to ./checkpoints:

# Download
wget -O hb_fp_2_4096_256.tar.gz https://figshare.com/ndownloader/files/31067692
# Extract
tar -vxf hb_fp_2_4096_256.tar.gz
# Rename files to match new scripts (...)
mv hb_fp_2_4096_256/ checkpoints/
formodelin"act""rt1""rxn""rt2"do
mkdir checkpoints/$model
mv "checkpoints/$model.ckpt""checkpoints/$model/ckpts.dummy-val_loss=0.00.ckpt"done
rm -f hb_fp_2_4096_256.tar.gz

The following scripts are run from the command line. Use python some_script.py --help or check the source code to see the instructions of each argument.

Prerequisites

In addition to the necessary data, we will need to pre-compute an embedding of the building blocks. To do so, please follow steps 0-2 from the INSTRUCTIONS.md. Then, replace the environment variables in the commands below.

Synthesis Planning

To perform synthesis planning described in the main text:

python scripts/20-predict-targets.py \
--building-blocks-file $BUILDING_BLOCKS_FILE \
--rxns-collection-file $RXN_COLLECTION_FILE \
--embeddings-knn-file $EMBEDDINGS_KNN_FILE \
--data "data/assets/molecules/sample-targets.txt" \
--ckpt-dir "checkpoints/" \
--output-dir "results/demo-inference/"

This script will feed a list of ten molecules to SynNet.

Synthesizable Molecular Design

To perform synthesizable molecular design, run:

python scripts/optimize_ga.py \
--ckpt-dir "checkpoints/" \
--building-blocks-file $BUILDING_BLOCKS_FILE \
--rxns-collection-file $RXN_COLLECTION_FILE \
--embeddings-knn-file $EMBEDDINGS_KNN_FILE \
--input-file path/to/zinc.csv \
--radius 2 --nbits 4096 \
--num_population 128 --num_offspring 512 --num_gen 200 --objective gsk \
--ncpu 32

This script uses a genetic algorithm to optimize molecular embeddings and returns the predicted synthetic trees for the optimized molecular embedding.

Note: input-file contains the seed molecules in CSV format for an initial run, and as a pre-saved numpy array of the population for restarting the run. If omitted, a random fingerprint will be chosen.

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SynNet

This repo contains the code and analysis scripts for our amortized approach to synthetic tree generation using neural networks. Our model can serve as both a synthesis planning tool and as a tool for synthesizable molecular design.

The method is described in detail in the publication "Amortized tree generation for bottom-up synthesis planning and synthesizable molecular design" available on the arXiv and summarized below.

Summary

We model synthetic pathways as tree structures called synthetic trees. A synthetic tree has a single root node and one or more child nodes. Every node is chemical molecule:

  • The root node is the final product molecule
  • The leaf nodes consist of purchasable building blocks.
  • All other inner nodes are constrained to be a product of allowed chemical reactions.

At a high level, each synthetic tree is constructed one reaction step at a time in a bottom-up manner, that is starting from purchasable building blocks.

Overview

The model consists of four modules, each containing a multi-layer perceptron (MLP):

  1. An Action Type selection function that classifies action types among the four possible actions (“Add”, “Expand”, “Merge”, and “End”) in building the synthetic tree. Each action increases the depth of the synthetic tree by one.

  2. A First Reactant selection function that selects the first reactant. A MLP predicts a molecular embedding and a first reactant is identified from the pool of building blocks through a k-nearest neighbors (k-NN) search.

  3. A Reaction selection function whose output is a probability distribution over available reaction templates. Inapplicable reactions are masked based on reactant 1. A suitable template is then sampled using a greedy search.

  4. A Second Reactant selection function that identifies the second reactant if the sampled template is bi-molecular. The model predicts an embedding for the second reactant, and a candidate is then sampled via a k-NN search from the masked set of building blocks.

the model

These four modules predict the probability distributions of actions to be taken within a single reaction step, and determine the nodes to be added to the synthetic tree under construction. All of these networks are conditioned on the target molecule embedding.

Synthesis planning

This task is to infer the synthetic pathway to a given target molecule. We formulate this problem as generating a synthetic tree such that the product molecule it produces (i.e., the molecule at the root node) matches the desired target molecule.

For this task, we can take a molecular embedding for the desired product, and use it as input to our model to produce a synthetic tree. If the desired product is successfully recovered, then the final root molecule will match the desired molecule used to create the input embedding. If the desired product is not successully recovered, it is possible the final root molecule may still be similar to the desired molecule used to create the input embedding, and thus our tool can also be used for synthesizable analog recommendation.

the generation process

Synthesizable molecular design

This task is to optimize a molecular structure with respect to an oracle function (e.g. bioactivity), while ensuring the synthetic accessibility of the molecules. We formulate this problem as optimizing the structure of a synthetic tree with respect to the desired properties of the product molecule it produces.

To do this, we optimize the molecular embedding of the molecule using a genetic algorithm and the desired oracle function. The optimized molecule embedding can then be used as input to our model to produce a synthetic tree, where the final root molecule corresponds to the optimized molecule.

Setup instructions

Environment

Conda is used to create the environment for running SynNet.

# Install environment from file
conda env create -f environment.yml

Before running any SynNet code, activate the environment and install this package in development mode:

source activate synnet
pip install -e .

The model implementations can be found in src/syn_net/models/.

The pre-processing and analysis scripts are in scripts/.

Train the model from scratch

Before training any models, you will first need to some data preprocessing. Please see INSTRUCTIONS.md for a complete guide.

Data

SynNet relies on two datasources:

  1. reaction templates and
  2. building blocks.

The data used for the publication are 1) the Hartenfeller-Button reaction templates, which are available under data/assets/reaction-templates/hb.txt and 2) Enamine building blocks. The building blocks are not freely available.

To obtain the data, go to https://enamine.net/building-blocks/building-blocks-catalog. We used the "Building Blocks, US Stock" data. You need to first register and then request access to download the dataset. The people from enamine.net manually approve you, so please be nice and patient.

Reproducing results

Before running anything, set up the environment as decribed above.

Using pre-trained models

We have made available a set of pre-trained models at the following link. The pretrained models correspond to the Action, Reactant 1, Reaction, and Reactant 2 networks, trained on the Hartenfeller-Button dataset and Enamine building blocks using radius 2, length 4096 Morgan fingerprints for the molecular node embeddings, and length 256 fingerprints for the k-NN search. For further details, please see the publication.

To download the pre-trained model to ./checkpoints:

# Download
wget -O hb_fp_2_4096_256.tar.gz https://figshare.com/ndownloader/files/31067692
# Extract
tar -vxf hb_fp_2_4096_256.tar.gz
# Rename files to match new scripts (...)
mv hb_fp_2_4096_256/ checkpoints/
formodelin"act""rt1""rxn""rt2"do
mkdir checkpoints/$model
mv "checkpoints/$model.ckpt""checkpoints/$model/ckpts.dummy-val_loss=0.00.ckpt"done
rm -f hb_fp_2_4096_256.tar.gz

The following scripts are run from the command line. Use python some_script.py --help or check the source code to see the instructions of each argument.

Prerequisites

In addition to the necessary data, we will need to pre-compute an embedding of the building blocks. To do so, please follow steps 0-2 from the INSTRUCTIONS.md. Then, replace the environment variables in the commands below.

Synthesis Planning

To perform synthesis planning described in the main text:

python scripts/20-predict-targets.py \
--building-blocks-file $BUILDING_BLOCKS_FILE \
--rxns-collection-file $RXN_COLLECTION_FILE \
--embeddings-knn-file $EMBEDDINGS_KNN_FILE \
--data "data/assets/molecules/sample-targets.txt" \
--ckpt-dir "checkpoints/" \
--output-dir "results/demo-inference/"

This script will feed a list of ten molecules to SynNet.

Synthesizable Molecular Design

To perform synthesizable molecular design, run:

python scripts/optimize_ga.py \
--ckpt-dir "checkpoints/" \
--building-blocks-file $BUILDING_BLOCKS_FILE \
--rxns-collection-file $RXN_COLLECTION_FILE \
--embeddings-knn-file $EMBEDDINGS_KNN_FILE \
--input-file path/to/zinc.csv \
--radius 2 --nbits 4096 \
--num_population 128 --num_offspring 512 --num_gen 200 --objective gsk \
--ncpu 32

This script uses a genetic algorithm to optimize molecular embeddings and returns the predicted synthetic trees for the optimized molecular embedding.

Note: input-file contains the seed molecules in CSV format for an initial run, and as a pre-saved numpy array of the population for restarting the run. If omitted, a random fingerprint will be chosen.

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

SynNet

This repo contains the code and analysis scripts for our amortized approach to synthetic tree generation using neural networks. Our model can serve as both a synthesis planning tool and as a tool for synthesizable molecular design.

The method is described in detail in the publication "Amortized tree generation for bottom-up synthesis planning and synthesizable molecular design" available on the arXiv and summarized below.

Summary

We model synthetic pathways as tree structures called synthetic trees. A synthetic tree has a single root node and one or more child nodes. Every node is chemical molecule:

  • The root node is the final product molecule
  • The leaf nodes consist of purchasable building blocks.
  • All other inner nodes are constrained to be a product of allowed chemical reactions.

At a high level, each synthetic tree is constructed one reaction step at a time in a bottom-up manner, that is starting from purchasable building blocks.

Overview

The model consists of four modules, each containing a multi-layer perceptron (MLP):

  1. An Action Type selection function that classifies action types among the four possible actions (“Add”, “Expand”, “Merge”, and “End”) in building the synthetic tree. Each action increases the depth of the synthetic tree by one.

  2. A First Reactant selection function that selects the first reactant. A MLP predicts a molecular embedding and a first reactant is identified from the pool of building blocks through a k-nearest neighbors (k-NN) search.

  3. A Reaction selection function whose output is a probability distribution over available reaction templates. Inapplicable reactions are masked based on reactant 1. A suitable template is then sampled using a greedy search.

  4. A Second Reactant selection function that identifies the second reactant if the sampled template is bi-molecular. The model predicts an embedding for the second reactant, and a candidate is then sampled via a k-NN search from the masked set of building blocks.

the model

These four modules predict the probability distributions of actions to be taken within a single reaction step, and determine the nodes to be added to the synthetic tree under construction. All of these networks are conditioned on the target molecule embedding.

Synthesis planning

This task is to infer the synthetic pathway to a given target molecule. We formulate this problem as generating a synthetic tree such that the product molecule it produces (i.e., the molecule at the root node) matches the desired target molecule.

For this task, we can take a molecular embedding for the desired product, and use it as input to our model to produce a synthetic tree. If the desired product is successfully recovered, then the final root molecule will match the desired molecule used to create the input embedding. If the desired product is not successully recovered, it is possible the final root molecule may still be similar to the desired molecule used to create the input embedding, and thus our tool can also be used for synthesizable analog recommendation.

the generation process

Synthesizable molecular design

This task is to optimize a molecular structure with respect to an oracle function (e.g. bioactivity), while ensuring the synthetic accessibility of the molecules. We formulate this problem as optimizing the structure of a synthetic tree with respect to the desired properties of the product molecule it produces.

To do this, we optimize the molecular embedding of the molecule using a genetic algorithm and the desired oracle function. The optimized molecule embedding can then be used as input to our model to produce a synthetic tree, where the final root molecule corresponds to the optimized molecule.

Setup instructions

Environment

Conda is used to create the environment for running SynNet.

# Install environment from file
conda env create -f environment.yml

Before running any SynNet code, activate the environment and install this package in development mode:

source activate synnet
pip install -e .

The model implementations can be found in src/syn_net/models/.

The pre-processing and analysis scripts are in scripts/.

Train the model from scratch

Before training any models, you will first need to some data preprocessing. Please see INSTRUCTIONS.md for a complete guide.

Data

SynNet relies on two datasources:

  1. reaction templates and
  2. building blocks.

The data used for the publication are 1) the Hartenfeller-Button reaction templates, which are available under data/assets/reaction-templates/hb.txt and 2) Enamine building blocks. The building blocks are not freely available.

To obtain the data, go to https://enamine.net/building-blocks/building-blocks-catalog. We used the "Building Blocks, US Stock" data. You need to first register and then request access to download the dataset. The people from enamine.net manually approve you, so please be nice and patient.

Reproducing results

Before running anything, set up the environment as decribed above.

Using pre-trained models

We have made available a set of pre-trained models at the following link. The pretrained models correspond to the Action, Reactant 1, Reaction, and Reactant 2 networks, trained on the Hartenfeller-Button dataset and Enamine building blocks using radius 2, length 4096 Morgan fingerprints for the molecular node embeddings, and length 256 fingerprints for the k-NN search. For further details, please see the publication.

To download the pre-trained model to ./checkpoints:

# Download
wget -O hb_fp_2_4096_256.tar.gz https://figshare.com/ndownloader/files/31067692
# Extract
tar -vxf hb_fp_2_4096_256.tar.gz
# Rename files to match new scripts (...)
mv hb_fp_2_4096_256/ checkpoints/
formodelin"act""rt1""rxn""rt2"do
mkdir checkpoints/$model
mv "checkpoints/$model.ckpt""checkpoints/$model/ckpts.dummy-val_loss=0.00.ckpt"done
rm -f hb_fp_2_4096_256.tar.gz

The following scripts are run from the command line. Use python some_script.py --help or check the source code to see the instructions of each argument.

Prerequisites

In addition to the necessary data, we will need to pre-compute an embedding of the building blocks. To do so, please follow steps 0-2 from the INSTRUCTIONS.md. Then, replace the environment variables in the commands below.

Synthesis Planning

To perform synthesis planning described in the main text:

python scripts/20-predict-targets.py \
--building-blocks-file $BUILDING_BLOCKS_FILE \
--rxns-collection-file $RXN_COLLECTION_FILE \
--embeddings-knn-file $EMBEDDINGS_KNN_FILE \
--data "data/assets/molecules/sample-targets.txt" \
--ckpt-dir "checkpoints/" \
--output-dir "results/demo-inference/"

This script will feed a list of ten molecules to SynNet.

Synthesizable Molecular Design

To perform synthesizable molecular design, run:

python scripts/optimize_ga.py \
--ckpt-dir "checkpoints/" \
--building-blocks-file $BUILDING_BLOCKS_FILE \
--rxns-collection-file $RXN_COLLECTION_FILE \
--embeddings-knn-file $EMBEDDINGS_KNN_FILE \
--input-file path/to/zinc.csv \
--radius 2 --nbits 4096 \
--num_population 128 --num_offspring 512 --num_gen 200 --objective gsk \
--ncpu 32

This script uses a genetic algorithm to optimize molecular embeddings and returns the predicted synthetic trees for the optimized molecular embedding.

Note: input-file contains the seed molecules in CSV format for an initial run, and as a pre-saved numpy array of the population for restarting the run. If omitted, a random fingerprint will be chosen.

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SynNet

This repo contains the code and analysis scripts for our amortized approach to synthetic tree generation using neural networks. Our model can serve as both a synthesis planning tool and as a tool for synthesizable molecular design.

The method is described in detail in the publication "Amortized tree generation for bottom-up synthesis planning and synthesizable molecular design" available on the arXiv and summarized below.

Summary

We model synthetic pathways as tree structures called synthetic trees. A synthetic tree has a single root node and one or more child nodes. Every node is chemical molecule:

  • The root node is the final product molecule
  • The leaf nodes consist of purchasable building blocks.
  • All other inner nodes are constrained to be a product of allowed chemical reactions.

At a high level, each synthetic tree is constructed one reaction step at a time in a bottom-up manner, that is starting from purchasable building blocks.

Overview

The model consists of four modules, each containing a multi-layer perceptron (MLP):

  1. An Action Type selection function that classifies action types among the four possible actions (“Add”, “Expand”, “Merge”, and “End”) in building the synthetic tree. Each action increases the depth of the synthetic tree by one.

  2. A First Reactant selection function that selects the first reactant. A MLP predicts a molecular embedding and a first reactant is identified from the pool of building blocks through a k-nearest neighbors (k-NN) search.

  3. A Reaction selection function whose output is a probability distribution over available reaction templates. Inapplicable reactions are masked based on reactant 1. A suitable template is then sampled using a greedy search.

  4. A Second Reactant selection function that identifies the second reactant if the sampled template is bi-molecular. The model predicts an embedding for the second reactant, and a candidate is then sampled via a k-NN search from the masked set of building blocks.

the model

These four modules predict the probability distributions of actions to be taken within a single reaction step, and determine the nodes to be added to the synthetic tree under construction. All of these networks are conditioned on the target molecule embedding.

Synthesis planning

This task is to infer the synthetic pathway to a given target molecule. We formulate this problem as generating a synthetic tree such that the product molecule it produces (i.e., the molecule at the root node) matches the desired target molecule.

For this task, we can take a molecular embedding for the desired product, and use it as input to our model to produce a synthetic tree. If the desired product is successfully recovered, then the final root molecule will match the desired molecule used to create the input embedding. If the desired product is not successully recovered, it is possible the final root molecule may still be similar to the desired molecule used to create the input embedding, and thus our tool can also be used for synthesizable analog recommendation.

the generation process

Synthesizable molecular design

This task is to optimize a molecular structure with respect to an oracle function (e.g. bioactivity), while ensuring the synthetic accessibility of the molecules. We formulate this problem as optimizing the structure of a synthetic tree with respect to the desired properties of the product molecule it produces.

To do this, we optimize the molecular embedding of the molecule using a genetic algorithm and the desired oracle function. The optimized molecule embedding can then be used as input to our model to produce a synthetic tree, where the final root molecule corresponds to the optimized molecule.

Setup instructions

Environment

Conda is used to create the environment for running SynNet.

# Install environment from file
conda env create -f environment.yml

Before running any SynNet code, activate the environment and install this package in development mode:

source activate synnet
pip install -e .

The model implementations can be found in src/syn_net/models/.

The pre-processing and analysis scripts are in scripts/.

Train the model from scratch

Before training any models, you will first need to some data preprocessing. Please see INSTRUCTIONS.md for a complete guide.

Data

SynNet relies on two datasources:

  1. reaction templates and
  2. building blocks.

The data used for the publication are 1) the Hartenfeller-Button reaction templates, which are available under data/assets/reaction-templates/hb.txt and 2) Enamine building blocks. The building blocks are not freely available.

To obtain the data, go to https://enamine.net/building-blocks/building-blocks-catalog. We used the "Building Blocks, US Stock" data. You need to first register and then request access to download the dataset. The people from enamine.net manually approve you, so please be nice and patient.

Reproducing results

Before running anything, set up the environment as decribed above.

Using pre-trained models

We have made available a set of pre-trained models at the following link. The pretrained models correspond to the Action, Reactant 1, Reaction, and Reactant 2 networks, trained on the Hartenfeller-Button dataset and Enamine building blocks using radius 2, length 4096 Morgan fingerprints for the molecular node embeddings, and length 256 fingerprints for the k-NN search. For further details, please see the publication.

To download the pre-trained model to ./checkpoints:

# Download
wget -O hb_fp_2_4096_256.tar.gz https://figshare.com/ndownloader/files/31067692
# Extract
tar -vxf hb_fp_2_4096_256.tar.gz
# Rename files to match new scripts (...)
mv hb_fp_2_4096_256/ checkpoints/
formodelin"act""rt1""rxn""rt2"do
mkdir checkpoints/$model
mv "checkpoints/$model.ckpt""checkpoints/$model/ckpts.dummy-val_loss=0.00.ckpt"done
rm -f hb_fp_2_4096_256.tar.gz

The following scripts are run from the command line. Use python some_script.py --help or check the source code to see the instructions of each argument.

Prerequisites

In addition to the necessary data, we will need to pre-compute an embedding of the building blocks. To do so, please follow steps 0-2 from the INSTRUCTIONS.md. Then, replace the environment variables in the commands below.

Synthesis Planning

To perform synthesis planning described in the main text:

python scripts/20-predict-targets.py \
--building-blocks-file $BUILDING_BLOCKS_FILE \
--rxns-collection-file $RXN_COLLECTION_FILE \
--embeddings-knn-file $EMBEDDINGS_KNN_FILE \
--data "data/assets/molecules/sample-targets.txt" \
--ckpt-dir "checkpoints/" \
--output-dir "results/demo-inference/"

This script will feed a list of ten molecules to SynNet.

Synthesizable Molecular Design

To perform synthesizable molecular design, run:

python scripts/optimize_ga.py \
--ckpt-dir "checkpoints/" \
--building-blocks-file $BUILDING_BLOCKS_FILE \
--rxns-collection-file $RXN_COLLECTION_FILE \
--embeddings-knn-file $EMBEDDINGS_KNN_FILE \
--input-file path/to/zinc.csv \
--radius 2 --nbits 4096 \
--num_population 128 --num_offspring 512 --num_gen 200 --objective gsk \
--ncpu 32

This script uses a genetic algorithm to optimize molecular embeddings and returns the predicted synthetic trees for the optimized molecular embedding.

Note: input-file contains the seed molecules in CSV format for an initial run, and as a pre-saved numpy array of the population for restarting the run. If omitted, a random fingerprint will be chosen.

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SynNet

This repo contains the code and analysis scripts for our amortized approach to synthetic tree generation using neural networks. Our model can serve as both a synthesis planning tool and as a tool for synthesizable molecular design.

The method is described in detail in the publication "Amortized tree generation for bottom-up synthesis planning and synthesizable molecular design" available on the arXiv and summarized below.

Summary

We model synthetic pathways as tree structures called synthetic trees. A synthetic tree has a single root node and one or more child nodes. Every node is chemical molecule:

  • The root node is the final product molecule
  • The leaf nodes consist of purchasable building blocks.
  • All other inner nodes are constrained to be a product of allowed chemical reactions.

At a high level, each synthetic tree is constructed one reaction step at a time in a bottom-up manner, that is starting from purchasable building blocks.

Overview

The model consists of four modules, each containing a multi-layer perceptron (MLP):

  1. An Action Type selection function that classifies action types among the four possible actions (“Add”, “Expand”, “Merge”, and “End”) in building the synthetic tree. Each action increases the depth of the synthetic tree by one.

  2. A First Reactant selection function that selects the first reactant. A MLP predicts a molecular embedding and a first reactant is identified from the pool of building blocks through a k-nearest neighbors (k-NN) search.

  3. A Reaction selection function whose output is a probability distribution over available reaction templates. Inapplicable reactions are masked based on reactant 1. A suitable template is then sampled using a greedy search.

  4. A Second Reactant selection function that identifies the second reactant if the sampled template is bi-molecular. The model predicts an embedding for the second reactant, and a candidate is then sampled via a k-NN search from the masked set of building blocks.

the model

These four modules predict the probability distributions of actions to be taken within a single reaction step, and determine the nodes to be added to the synthetic tree under construction. All of these networks are conditioned on the target molecule embedding.

Synthesis planning

This task is to infer the synthetic pathway to a given target molecule. We formulate this problem as generating a synthetic tree such that the product molecule it produces (i.e., the molecule at the root node) matches the desired target molecule.

For this task, we can take a molecular embedding for the desired product, and use it as input to our model to produce a synthetic tree. If the desired product is successfully recovered, then the final root molecule will match the desired molecule used to create the input embedding. If the desired product is not successully recovered, it is possible the final root molecule may still be similar to the desired molecule used to create the input embedding, and thus our tool can also be used for synthesizable analog recommendation.

the generation process

Synthesizable molecular design

This task is to optimize a molecular structure with respect to an oracle function (e.g. bioactivity), while ensuring the synthetic accessibility of the molecules. We formulate this problem as optimizing the structure of a synthetic tree with respect to the desired properties of the product molecule it produces.

To do this, we optimize the molecular embedding of the molecule using a genetic algorithm and the desired oracle function. The optimized molecule embedding can then be used as input to our model to produce a synthetic tree, where the final root molecule corresponds to the optimized molecule.

Setup instructions

Environment

Conda is used to create the environment for running SynNet.

# Install environment from file
conda env create -f environment.yml

Before running any SynNet code, activate the environment and install this package in development mode:

source activate synnet
pip install -e .

The model implementations can be found in src/syn_net/models/.

The pre-processing and analysis scripts are in scripts/.

Train the model from scratch

Before training any models, you will first need to some data preprocessing. Please see INSTRUCTIONS.md for a complete guide.

Data

SynNet relies on two datasources:

  1. reaction templates and
  2. building blocks.

The data used for the publication are 1) the Hartenfeller-Button reaction templates, which are available under data/assets/reaction-templates/hb.txt and 2) Enamine building blocks. The building blocks are not freely available.

To obtain the data, go to https://enamine.net/building-blocks/building-blocks-catalog. We used the "Building Blocks, US Stock" data. You need to first register and then request access to download the dataset. The people from enamine.net manually approve you, so please be nice and patient.

Reproducing results

Before running anything, set up the environment as decribed above.

Using pre-trained models

We have made available a set of pre-trained models at the following link. The pretrained models correspond to the Action, Reactant 1, Reaction, and Reactant 2 networks, trained on the Hartenfeller-Button dataset and Enamine building blocks using radius 2, length 4096 Morgan fingerprints for the molecular node embeddings, and length 256 fingerprints for the k-NN search. For further details, please see the publication.

To download the pre-trained model to ./checkpoints:

# Download
wget -O hb_fp_2_4096_256.tar.gz https://figshare.com/ndownloader/files/31067692
# Extract
tar -vxf hb_fp_2_4096_256.tar.gz
# Rename files to match new scripts (...)
mv hb_fp_2_4096_256/ checkpoints/
formodelin"act""rt1""rxn""rt2"do
mkdir checkpoints/$model
mv "checkpoints/$model.ckpt""checkpoints/$model/ckpts.dummy-val_loss=0.00.ckpt"done
rm -f hb_fp_2_4096_256.tar.gz

The following scripts are run from the command line. Use python some_script.py --help or check the source code to see the instructions of each argument.

Prerequisites

In addition to the necessary data, we will need to pre-compute an embedding of the building blocks. To do so, please follow steps 0-2 from the INSTRUCTIONS.md. Then, replace the environment variables in the commands below.

Synthesis Planning

To perform synthesis planning described in the main text:

python scripts/20-predict-targets.py \
--building-blocks-file $BUILDING_BLOCKS_FILE \
--rxns-collection-file $RXN_COLLECTION_FILE \
--embeddings-knn-file $EMBEDDINGS_KNN_FILE \
--data "data/assets/molecules/sample-targets.txt" \
--ckpt-dir "checkpoints/" \
--output-dir "results/demo-inference/"

This script will feed a list of ten molecules to SynNet.

Synthesizable Molecular Design

To perform synthesizable molecular design, run:

python scripts/optimize_ga.py \
--ckpt-dir "checkpoints/" \
--building-blocks-file $BUILDING_BLOCKS_FILE \
--rxns-collection-file $RXN_COLLECTION_FILE \
--embeddings-knn-file $EMBEDDINGS_KNN_FILE \
--input-file path/to/zinc.csv \
--radius 2 --nbits 4096 \
--num_population 128 --num_offspring 512 --num_gen 200 --objective gsk \
--ncpu 32

This script uses a genetic algorithm to optimize molecular embeddings and returns the predicted synthetic trees for the optimized molecular embedding.

Note: input-file contains the seed molecules in CSV format for an initial run, and as a pre-saved numpy array of the population for restarting the run. If omitted, a random fingerprint will be chosen.

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

SynNet

This repo contains the code and analysis scripts for our amortized approach to synthetic tree generation using neural networks. Our model can serve as both a synthesis planning tool and as a tool for synthesizable molecular design.

The method is described in detail in the publication "Amortized tree generation for bottom-up synthesis planning and synthesizable molecular design" available on the arXiv and summarized below.

Summary

We model synthetic pathways as tree structures called synthetic trees. A synthetic tree has a single root node and one or more child nodes. Every node is chemical molecule:

  • The root node is the final product molecule
  • The leaf nodes consist of purchasable building blocks.
  • All other inner nodes are constrained to be a product of allowed chemical reactions.

At a high level, each synthetic tree is constructed one reaction step at a time in a bottom-up manner, that is starting from purchasable building blocks.

Overview

The model consists of four modules, each containing a multi-layer perceptron (MLP):

  1. An Action Type selection function that classifies action types among the four possible actions (“Add”, “Expand”, “Merge”, and “End”) in building the synthetic tree. Each action increases the depth of the synthetic tree by one.

  2. A First Reactant selection function that selects the first reactant. A MLP predicts a molecular embedding and a first reactant is identified from the pool of building blocks through a k-nearest neighbors (k-NN) search.

  3. A Reaction selection function whose output is a probability distribution over available reaction templates. Inapplicable reactions are masked based on reactant 1. A suitable template is then sampled using a greedy search.

  4. A Second Reactant selection function that identifies the second reactant if the sampled template is bi-molecular. The model predicts an embedding for the second reactant, and a candidate is then sampled via a k-NN search from the masked set of building blocks.

the model

These four modules predict the probability distributions of actions to be taken within a single reaction step, and determine the nodes to be added to the synthetic tree under construction. All of these networks are conditioned on the target molecule embedding.

Synthesis planning

This task is to infer the synthetic pathway to a given target molecule. We formulate this problem as generating a synthetic tree such that the product molecule it produces (i.e., the molecule at the root node) matches the desired target molecule.

For this task, we can take a molecular embedding for the desired product, and use it as input to our model to produce a synthetic tree. If the desired product is successfully recovered, then the final root molecule will match the desired molecule used to create the input embedding. If the desired product is not successully recovered, it is possible the final root molecule may still be similar to the desired molecule used to create the input embedding, and thus our tool can also be used for synthesizable analog recommendation.

the generation process

Synthesizable molecular design

This task is to optimize a molecular structure with respect to an oracle function (e.g. bioactivity), while ensuring the synthetic accessibility of the molecules. We formulate this problem as optimizing the structure of a synthetic tree with respect to the desired properties of the product molecule it produces.

To do this, we optimize the molecular embedding of the molecule using a genetic algorithm and the desired oracle function. The optimized molecule embedding can then be used as input to our model to produce a synthetic tree, where the final root molecule corresponds to the optimized molecule.

Setup instructions

Environment

Conda is used to create the environment for running SynNet.

# Install environment from file
conda env create -f environment.yml

Before running any SynNet code, activate the environment and install this package in development mode:

source activate synnet
pip install -e .

The model implementations can be found in src/syn_net/models/.

The pre-processing and analysis scripts are in scripts/.

Train the model from scratch

Before training any models, you will first need to some data preprocessing. Please see INSTRUCTIONS.md for a complete guide.

Data

SynNet relies on two datasources:

  1. reaction templates and
  2. building blocks.

The data used for the publication are 1) the Hartenfeller-Button reaction templates, which are available under data/assets/reaction-templates/hb.txt and 2) Enamine building blocks. The building blocks are not freely available.

To obtain the data, go to https://enamine.net/building-blocks/building-blocks-catalog. We used the "Building Blocks, US Stock" data. You need to first register and then request access to download the dataset. The people from enamine.net manually approve you, so please be nice and patient.

Reproducing results

Before running anything, set up the environment as decribed above.

Using pre-trained models

We have made available a set of pre-trained models at the following link. The pretrained models correspond to the Action, Reactant 1, Reaction, and Reactant 2 networks, trained on the Hartenfeller-Button dataset and Enamine building blocks using radius 2, length 4096 Morgan fingerprints for the molecular node embeddings, and length 256 fingerprints for the k-NN search. For further details, please see the publication.

To download the pre-trained model to ./checkpoints:

# Download
wget -O hb_fp_2_4096_256.tar.gz https://figshare.com/ndownloader/files/31067692
# Extract
tar -vxf hb_fp_2_4096_256.tar.gz
# Rename files to match new scripts (...)
mv hb_fp_2_4096_256/ checkpoints/
formodelin"act""rt1""rxn""rt2"do
mkdir checkpoints/$model
mv "checkpoints/$model.ckpt""checkpoints/$model/ckpts.dummy-val_loss=0.00.ckpt"done
rm -f hb_fp_2_4096_256.tar.gz

The following scripts are run from the command line. Use python some_script.py --help or check the source code to see the instructions of each argument.

Prerequisites

In addition to the necessary data, we will need to pre-compute an embedding of the building blocks. To do so, please follow steps 0-2 from the INSTRUCTIONS.md. Then, replace the environment variables in the commands below.

Synthesis Planning

To perform synthesis planning described in the main text:

python scripts/20-predict-targets.py \
--building-blocks-file $BUILDING_BLOCKS_FILE \
--rxns-collection-file $RXN_COLLECTION_FILE \
--embeddings-knn-file $EMBEDDINGS_KNN_FILE \
--data "data/assets/molecules/sample-targets.txt" \
--ckpt-dir "checkpoints/" \
--output-dir "results/demo-inference/"

This script will feed a list of ten molecules to SynNet.

Synthesizable Molecular Design

To perform synthesizable molecular design, run:

python scripts/optimize_ga.py \
--ckpt-dir "checkpoints/" \
--building-blocks-file $BUILDING_BLOCKS_FILE \
--rxns-collection-file $RXN_COLLECTION_FILE \
--embeddings-knn-file $EMBEDDINGS_KNN_FILE \
--input-file path/to/zinc.csv \
--radius 2 --nbits 4096 \
--num_population 128 --num_offspring 512 --num_gen 200 --objective gsk \
--ncpu 32

This script uses a genetic algorithm to optimize molecular embeddings and returns the predicted synthetic trees for the optimized molecular embedding.

Note: input-file contains the seed molecules in CSV format for an initial run, and as a pre-saved numpy array of the population for restarting the run. If omitted, a random fingerprint will be chosen.

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages