Latest commit

History

772 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Test Image 1

The ATOM Modeling PipeLine (AMPL; https://github.com/ATOMconsortium/AMPL) is an open-source, modular, extensible software pipeline for building and sharing models to advance in silico drug discovery. To see the list of AMPL parameters, please check this link, https://github.com/ATOMconsortium/AMPL/blob/master/atomsci/ddm/docs/PARAMETERS.md

This page contains a collection of AMPL-COLAB tutorial notebooks.

+ Please note that if you have trouble opening up any of the following notebooks, please go to, https://nbviewer.jupyter.org/, and paste the notebook link to view the contents.

0. Basic Google COLAB Introduction (Works best with Google Chrome)

  • Tutorial-00: Basic COLAB tutorial. For all the COLAB tutorials, click on the tutorial link, and then click on "Open in Colab" baner. You can open and run the notebook from the browser. If you want to save your edits to the notebook, you need to save a copy in your Google Drive. Usually, Google COLAB saves the notebook files under the "My Drive > Colab Notebooks" folder

1. Data Collection and creating Machine-Learning ready datasets:

The data that we gather for modeling is small-molecule/drug binding data. The following links will introduce some of the concepts and outcome measures related to this topic:

For the tutorials, we will use the small-molecule binding data obtained from either one of the following resources, ChEMBL (https://www.ebi.ac.uk/chembl/), Drug Target Commons (DTC; https://drugtargetcommons.fimm.fi/) and Excape-DB (https://solr.ideaconsult.net/search/excape/).

Click here to learn about single-target focussed data.

Explore HTR3A binding data from ExCAPE-DB

Explore HTR3A binding data from Drug Target Commons database

  • Tutorial-03: (Time: ~ 4 minutes) This COLAB notebook will use AMPL for Data cleaning, EDA and clustering of HTR3A protein data from Drug Target Commons (DTC)
  • Tutorial-04: (Time: ~ 10 minutes) This COLAB notebook will use AMPL for Data curation of HTR3A protein data from Drug Target Commons (DTC)

Curating, merging and visualizing two datasets

  • Tutorial-05: (Time: ~ 4 minutes) This COLAB notebook will use AMPL to upload datasets (small-molecule activity data from ChEMBL), clean, merge and do some basic Exploratory Data Analysis.
  • Tutorial-06: (Time: ~ 4 minutes) This COLAB notebook with use AMPL to merge HTR3A binding data from two different data sources, DTC and ExCAPE-DB.

Exploratory Data Analysis (EDA) Notebooks

  • Tutorial-07: (Time: ~ 4 minutes). The notebook uses HTR3A as the protein target. The notebook accomplishes the following tasks:
    • Uses AMPL software
    • Reads in data from three database sources: ChEMBL, Excape-DB and DTC
    • Cleans, standardizes and analyzes the data
    • Merges and harmonizes to create a dataset
  • Tutorial-08: Exploratory Data Analysis-Regression
  • Tutorial-09: Exploratory Data Analysis-Regression

2. Model training and tuning:

Random Forest modeling to predict solubility

  • Tutorial-10: (Time: ~ 2 minutes): Simple supervised learning example. AMPL will read the public data (117 chemical compounds), curate, fit a Random Forest model to predict solubility and test the model. For additional information on the dataset, please check this publication,https://pubmed.ncbi.nlm.nih.gov/15154768/Delaney

Graph Convolution modeling to predict SCN5A binding affinities

  • Tutorial-11: (Mode: AMPL_GPU; Time: ~ 18 minutes): This COLAB notebook will use AMPL for predicting binding affinities -pIC50 values- of ligands that could bind to human Sodium channel protein type 5 subunit alpha protein (Gene: SCN5A) using Graph Convolutional Network Model. ChEMBL database is the data source of binding affinities (pIC50) Test Image 1

3. Hyper-parameter Optimization (HPO), Uncertainty Quantification (UQ), and using metrics for analyzing model performance.

This notebook also explores AMPL functions for saving and loading prebuild AMPL models for analysis.

  • Tutorial-12 Hyper-parameter Optimization (HPO) and Uncertainty Quantification (UQ).
  • Tutorial-13 Notebook includes HPO Grid Search on three different modeling methods (Random Forest, NN and XGBoost).

4. Creating high-quality models

  • Tutorial-12 Notebook provides the framework for visualizing the results of HPO results and use them to identify best models.

5. Model Inference:

AMPL Workshops

Workshop date, June 05, 2021: Protein Target-focussed Binding Data Curation, Exploratory Data Analysis and Featurization using AMPL. Please note that Google Chrome browser works best with the COLAB Jupyter notebooks

Supporting links

Similar chemoinformatics, drug-discovery software tools:

Chemoinformatics databases

Acknowledgements

Most of the tutorial code chunks came from multple Jupyter notebooks generously shared by the ATOM team.

  • Amanda Paulson
  • Ben Madej
  • Da Shi
  • Hiran Ranganathan
  • Jessica Mauvais
  • Jonathan Allen
  • Kevin Mcloughlin
  • Sarangan Ravichandran
  • Stewart He
  • Ya Ju Fan
  • Contributions from the following student programs:

About

AMPL software tutorials

Resources

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

772 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Test Image 1

The ATOM Modeling PipeLine (AMPL; https://github.com/ATOMconsortium/AMPL) is an open-source, modular, extensible software pipeline for building and sharing models to advance in silico drug discovery. To see the list of AMPL parameters, please check this link, https://github.com/ATOMconsortium/AMPL/blob/master/atomsci/ddm/docs/PARAMETERS.md

This page contains a collection of AMPL-COLAB tutorial notebooks.

+ Please note that if you have trouble opening up any of the following notebooks, please go to, https://nbviewer.jupyter.org/, and paste the notebook link to view the contents.

0. Basic Google COLAB Introduction (Works best with Google Chrome)

  • Tutorial-00: Basic COLAB tutorial. For all the COLAB tutorials, click on the tutorial link, and then click on "Open in Colab" baner. You can open and run the notebook from the browser. If you want to save your edits to the notebook, you need to save a copy in your Google Drive. Usually, Google COLAB saves the notebook files under the "My Drive > Colab Notebooks" folder

1. Data Collection and creating Machine-Learning ready datasets:

The data that we gather for modeling is small-molecule/drug binding data. The following links will introduce some of the concepts and outcome measures related to this topic:

For the tutorials, we will use the small-molecule binding data obtained from either one of the following resources, ChEMBL (https://www.ebi.ac.uk/chembl/), Drug Target Commons (DTC; https://drugtargetcommons.fimm.fi/) and Excape-DB (https://solr.ideaconsult.net/search/excape/).

Click here to learn about single-target focussed data.

Explore HTR3A binding data from ExCAPE-DB

Explore HTR3A binding data from Drug Target Commons database

  • Tutorial-03: (Time: ~ 4 minutes) This COLAB notebook will use AMPL for Data cleaning, EDA and clustering of HTR3A protein data from Drug Target Commons (DTC)
  • Tutorial-04: (Time: ~ 10 minutes) This COLAB notebook will use AMPL for Data curation of HTR3A protein data from Drug Target Commons (DTC)

Curating, merging and visualizing two datasets

  • Tutorial-05: (Time: ~ 4 minutes) This COLAB notebook will use AMPL to upload datasets (small-molecule activity data from ChEMBL), clean, merge and do some basic Exploratory Data Analysis.
  • Tutorial-06: (Time: ~ 4 minutes) This COLAB notebook with use AMPL to merge HTR3A binding data from two different data sources, DTC and ExCAPE-DB.

Exploratory Data Analysis (EDA) Notebooks

  • Tutorial-07: (Time: ~ 4 minutes). The notebook uses HTR3A as the protein target. The notebook accomplishes the following tasks:
    • Uses AMPL software
    • Reads in data from three database sources: ChEMBL, Excape-DB and DTC
    • Cleans, standardizes and analyzes the data
    • Merges and harmonizes to create a dataset
  • Tutorial-08: Exploratory Data Analysis-Regression
  • Tutorial-09: Exploratory Data Analysis-Regression

2. Model training and tuning:

Random Forest modeling to predict solubility

  • Tutorial-10: (Time: ~ 2 minutes): Simple supervised learning example. AMPL will read the public data (117 chemical compounds), curate, fit a Random Forest model to predict solubility and test the model. For additional information on the dataset, please check this publication,https://pubmed.ncbi.nlm.nih.gov/15154768/Delaney

Graph Convolution modeling to predict SCN5A binding affinities

  • Tutorial-11: (Mode: AMPL_GPU; Time: ~ 18 minutes): This COLAB notebook will use AMPL for predicting binding affinities -pIC50 values- of ligands that could bind to human Sodium channel protein type 5 subunit alpha protein (Gene: SCN5A) using Graph Convolutional Network Model. ChEMBL database is the data source of binding affinities (pIC50) Test Image 1

3. Hyper-parameter Optimization (HPO), Uncertainty Quantification (UQ), and using metrics for analyzing model performance.

This notebook also explores AMPL functions for saving and loading prebuild AMPL models for analysis.

  • Tutorial-12 Hyper-parameter Optimization (HPO) and Uncertainty Quantification (UQ).
  • Tutorial-13 Notebook includes HPO Grid Search on three different modeling methods (Random Forest, NN and XGBoost).

4. Creating high-quality models

  • Tutorial-12 Notebook provides the framework for visualizing the results of HPO results and use them to identify best models.

5. Model Inference:

AMPL Workshops

Workshop date, June 05, 2021: Protein Target-focussed Binding Data Curation, Exploratory Data Analysis and Featurization using AMPL. Please note that Google Chrome browser works best with the COLAB Jupyter notebooks

Supporting links

Similar chemoinformatics, drug-discovery software tools:

Chemoinformatics databases

Acknowledgements

Most of the tutorial code chunks came from multple Jupyter notebooks generously shared by the ATOM team.

  • Amanda Paulson
  • Ben Madej
  • Da Shi
  • Hiran Ranganathan
  • Jessica Mauvais
  • Jonathan Allen
  • Kevin Mcloughlin
  • Sarangan Ravichandran
  • Stewart He
  • Ya Ju Fan
  • Contributions from the following student programs:

About

AMPL software tutorials

Resources

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

772 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Test Image 1

The ATOM Modeling PipeLine (AMPL; https://github.com/ATOMconsortium/AMPL) is an open-source, modular, extensible software pipeline for building and sharing models to advance in silico drug discovery. To see the list of AMPL parameters, please check this link, https://github.com/ATOMconsortium/AMPL/blob/master/atomsci/ddm/docs/PARAMETERS.md

This page contains a collection of AMPL-COLAB tutorial notebooks.

+ Please note that if you have trouble opening up any of the following notebooks, please go to, https://nbviewer.jupyter.org/, and paste the notebook link to view the contents.

0. Basic Google COLAB Introduction (Works best with Google Chrome)

  • Tutorial-00: Basic COLAB tutorial. For all the COLAB tutorials, click on the tutorial link, and then click on "Open in Colab" baner. You can open and run the notebook from the browser. If you want to save your edits to the notebook, you need to save a copy in your Google Drive. Usually, Google COLAB saves the notebook files under the "My Drive > Colab Notebooks" folder

1. Data Collection and creating Machine-Learning ready datasets:

The data that we gather for modeling is small-molecule/drug binding data. The following links will introduce some of the concepts and outcome measures related to this topic:

For the tutorials, we will use the small-molecule binding data obtained from either one of the following resources, ChEMBL (https://www.ebi.ac.uk/chembl/), Drug Target Commons (DTC; https://drugtargetcommons.fimm.fi/) and Excape-DB (https://solr.ideaconsult.net/search/excape/).

Click here to learn about single-target focussed data.

Explore HTR3A binding data from ExCAPE-DB

Explore HTR3A binding data from Drug Target Commons database

  • Tutorial-03: (Time: ~ 4 minutes) This COLAB notebook will use AMPL for Data cleaning, EDA and clustering of HTR3A protein data from Drug Target Commons (DTC)
  • Tutorial-04: (Time: ~ 10 minutes) This COLAB notebook will use AMPL for Data curation of HTR3A protein data from Drug Target Commons (DTC)

Curating, merging and visualizing two datasets

  • Tutorial-05: (Time: ~ 4 minutes) This COLAB notebook will use AMPL to upload datasets (small-molecule activity data from ChEMBL), clean, merge and do some basic Exploratory Data Analysis.
  • Tutorial-06: (Time: ~ 4 minutes) This COLAB notebook with use AMPL to merge HTR3A binding data from two different data sources, DTC and ExCAPE-DB.

Exploratory Data Analysis (EDA) Notebooks

  • Tutorial-07: (Time: ~ 4 minutes). The notebook uses HTR3A as the protein target. The notebook accomplishes the following tasks:
    • Uses AMPL software
    • Reads in data from three database sources: ChEMBL, Excape-DB and DTC
    • Cleans, standardizes and analyzes the data
    • Merges and harmonizes to create a dataset
  • Tutorial-08: Exploratory Data Analysis-Regression
  • Tutorial-09: Exploratory Data Analysis-Regression

2. Model training and tuning:

Random Forest modeling to predict solubility

  • Tutorial-10: (Time: ~ 2 minutes): Simple supervised learning example. AMPL will read the public data (117 chemical compounds), curate, fit a Random Forest model to predict solubility and test the model. For additional information on the dataset, please check this publication,https://pubmed.ncbi.nlm.nih.gov/15154768/Delaney

Graph Convolution modeling to predict SCN5A binding affinities

  • Tutorial-11: (Mode: AMPL_GPU; Time: ~ 18 minutes): This COLAB notebook will use AMPL for predicting binding affinities -pIC50 values- of ligands that could bind to human Sodium channel protein type 5 subunit alpha protein (Gene: SCN5A) using Graph Convolutional Network Model. ChEMBL database is the data source of binding affinities (pIC50) Test Image 1

3. Hyper-parameter Optimization (HPO), Uncertainty Quantification (UQ), and using metrics for analyzing model performance.

This notebook also explores AMPL functions for saving and loading prebuild AMPL models for analysis.

  • Tutorial-12 Hyper-parameter Optimization (HPO) and Uncertainty Quantification (UQ).
  • Tutorial-13 Notebook includes HPO Grid Search on three different modeling methods (Random Forest, NN and XGBoost).

4. Creating high-quality models

  • Tutorial-12 Notebook provides the framework for visualizing the results of HPO results and use them to identify best models.

5. Model Inference:

AMPL Workshops

Workshop date, June 05, 2021: Protein Target-focussed Binding Data Curation, Exploratory Data Analysis and Featurization using AMPL. Please note that Google Chrome browser works best with the COLAB Jupyter notebooks

Supporting links

Similar chemoinformatics, drug-discovery software tools:

Chemoinformatics databases

Acknowledgements

Most of the tutorial code chunks came from multple Jupyter notebooks generously shared by the ATOM team.

  • Amanda Paulson
  • Ben Madej
  • Da Shi
  • Hiran Ranganathan
  • Jessica Mauvais
  • Jonathan Allen
  • Kevin Mcloughlin
  • Sarangan Ravichandran
  • Stewart He
  • Ya Ju Fan
  • Contributions from the following student programs:

About

AMPL software tutorials

Resources

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

772 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Test Image 1

The ATOM Modeling PipeLine (AMPL; https://github.com/ATOMconsortium/AMPL) is an open-source, modular, extensible software pipeline for building and sharing models to advance in silico drug discovery. To see the list of AMPL parameters, please check this link, https://github.com/ATOMconsortium/AMPL/blob/master/atomsci/ddm/docs/PARAMETERS.md

This page contains a collection of AMPL-COLAB tutorial notebooks.

+ Please note that if you have trouble opening up any of the following notebooks, please go to, https://nbviewer.jupyter.org/, and paste the notebook link to view the contents.

0. Basic Google COLAB Introduction (Works best with Google Chrome)

  • Tutorial-00: Basic COLAB tutorial. For all the COLAB tutorials, click on the tutorial link, and then click on "Open in Colab" baner. You can open and run the notebook from the browser. If you want to save your edits to the notebook, you need to save a copy in your Google Drive. Usually, Google COLAB saves the notebook files under the "My Drive > Colab Notebooks" folder

1. Data Collection and creating Machine-Learning ready datasets:

The data that we gather for modeling is small-molecule/drug binding data. The following links will introduce some of the concepts and outcome measures related to this topic:

For the tutorials, we will use the small-molecule binding data obtained from either one of the following resources, ChEMBL (https://www.ebi.ac.uk/chembl/), Drug Target Commons (DTC; https://drugtargetcommons.fimm.fi/) and Excape-DB (https://solr.ideaconsult.net/search/excape/).

Click here to learn about single-target focussed data.

Explore HTR3A binding data from ExCAPE-DB

Explore HTR3A binding data from Drug Target Commons database

  • Tutorial-03: (Time: ~ 4 minutes) This COLAB notebook will use AMPL for Data cleaning, EDA and clustering of HTR3A protein data from Drug Target Commons (DTC)
  • Tutorial-04: (Time: ~ 10 minutes) This COLAB notebook will use AMPL for Data curation of HTR3A protein data from Drug Target Commons (DTC)

Curating, merging and visualizing two datasets

  • Tutorial-05: (Time: ~ 4 minutes) This COLAB notebook will use AMPL to upload datasets (small-molecule activity data from ChEMBL), clean, merge and do some basic Exploratory Data Analysis.
  • Tutorial-06: (Time: ~ 4 minutes) This COLAB notebook with use AMPL to merge HTR3A binding data from two different data sources, DTC and ExCAPE-DB.

Exploratory Data Analysis (EDA) Notebooks

  • Tutorial-07: (Time: ~ 4 minutes). The notebook uses HTR3A as the protein target. The notebook accomplishes the following tasks:
    • Uses AMPL software
    • Reads in data from three database sources: ChEMBL, Excape-DB and DTC
    • Cleans, standardizes and analyzes the data
    • Merges and harmonizes to create a dataset
  • Tutorial-08: Exploratory Data Analysis-Regression
  • Tutorial-09: Exploratory Data Analysis-Regression

2. Model training and tuning:

Random Forest modeling to predict solubility

  • Tutorial-10: (Time: ~ 2 minutes): Simple supervised learning example. AMPL will read the public data (117 chemical compounds), curate, fit a Random Forest model to predict solubility and test the model. For additional information on the dataset, please check this publication,https://pubmed.ncbi.nlm.nih.gov/15154768/Delaney

Graph Convolution modeling to predict SCN5A binding affinities

  • Tutorial-11: (Mode: AMPL_GPU; Time: ~ 18 minutes): This COLAB notebook will use AMPL for predicting binding affinities -pIC50 values- of ligands that could bind to human Sodium channel protein type 5 subunit alpha protein (Gene: SCN5A) using Graph Convolutional Network Model. ChEMBL database is the data source of binding affinities (pIC50) Test Image 1

3. Hyper-parameter Optimization (HPO), Uncertainty Quantification (UQ), and using metrics for analyzing model performance.

This notebook also explores AMPL functions for saving and loading prebuild AMPL models for analysis.

  • Tutorial-12 Hyper-parameter Optimization (HPO) and Uncertainty Quantification (UQ).
  • Tutorial-13 Notebook includes HPO Grid Search on three different modeling methods (Random Forest, NN and XGBoost).

4. Creating high-quality models

  • Tutorial-12 Notebook provides the framework for visualizing the results of HPO results and use them to identify best models.

5. Model Inference:

AMPL Workshops

Workshop date, June 05, 2021: Protein Target-focussed Binding Data Curation, Exploratory Data Analysis and Featurization using AMPL. Please note that Google Chrome browser works best with the COLAB Jupyter notebooks

Supporting links

Similar chemoinformatics, drug-discovery software tools:

Chemoinformatics databases

Acknowledgements

Most of the tutorial code chunks came from multple Jupyter notebooks generously shared by the ATOM team.

  • Amanda Paulson
  • Ben Madej
  • Da Shi
  • Hiran Ranganathan
  • Jessica Mauvais
  • Jonathan Allen
  • Kevin Mcloughlin
  • Sarangan Ravichandran
  • Stewart He
  • Ya Ju Fan
  • Contributions from the following student programs:

About

AMPL software tutorials

Resources

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

772 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Test Image 1

The ATOM Modeling PipeLine (AMPL; https://github.com/ATOMconsortium/AMPL) is an open-source, modular, extensible software pipeline for building and sharing models to advance in silico drug discovery. To see the list of AMPL parameters, please check this link, https://github.com/ATOMconsortium/AMPL/blob/master/atomsci/ddm/docs/PARAMETERS.md

This page contains a collection of AMPL-COLAB tutorial notebooks.

+ Please note that if you have trouble opening up any of the following notebooks, please go to, https://nbviewer.jupyter.org/, and paste the notebook link to view the contents.

0. Basic Google COLAB Introduction (Works best with Google Chrome)

  • Tutorial-00: Basic COLAB tutorial. For all the COLAB tutorials, click on the tutorial link, and then click on "Open in Colab" baner. You can open and run the notebook from the browser. If you want to save your edits to the notebook, you need to save a copy in your Google Drive. Usually, Google COLAB saves the notebook files under the "My Drive > Colab Notebooks" folder

1. Data Collection and creating Machine-Learning ready datasets:

The data that we gather for modeling is small-molecule/drug binding data. The following links will introduce some of the concepts and outcome measures related to this topic:

For the tutorials, we will use the small-molecule binding data obtained from either one of the following resources, ChEMBL (https://www.ebi.ac.uk/chembl/), Drug Target Commons (DTC; https://drugtargetcommons.fimm.fi/) and Excape-DB (https://solr.ideaconsult.net/search/excape/).

Click here to learn about single-target focussed data.

Explore HTR3A binding data from ExCAPE-DB

Explore HTR3A binding data from Drug Target Commons database

  • Tutorial-03: (Time: ~ 4 minutes) This COLAB notebook will use AMPL for Data cleaning, EDA and clustering of HTR3A protein data from Drug Target Commons (DTC)
  • Tutorial-04: (Time: ~ 10 minutes) This COLAB notebook will use AMPL for Data curation of HTR3A protein data from Drug Target Commons (DTC)

Curating, merging and visualizing two datasets

  • Tutorial-05: (Time: ~ 4 minutes) This COLAB notebook will use AMPL to upload datasets (small-molecule activity data from ChEMBL), clean, merge and do some basic Exploratory Data Analysis.
  • Tutorial-06: (Time: ~ 4 minutes) This COLAB notebook with use AMPL to merge HTR3A binding data from two different data sources, DTC and ExCAPE-DB.

Exploratory Data Analysis (EDA) Notebooks

  • Tutorial-07: (Time: ~ 4 minutes). The notebook uses HTR3A as the protein target. The notebook accomplishes the following tasks:
    • Uses AMPL software
    • Reads in data from three database sources: ChEMBL, Excape-DB and DTC
    • Cleans, standardizes and analyzes the data
    • Merges and harmonizes to create a dataset
  • Tutorial-08: Exploratory Data Analysis-Regression
  • Tutorial-09: Exploratory Data Analysis-Regression

2. Model training and tuning:

Random Forest modeling to predict solubility

  • Tutorial-10: (Time: ~ 2 minutes): Simple supervised learning example. AMPL will read the public data (117 chemical compounds), curate, fit a Random Forest model to predict solubility and test the model. For additional information on the dataset, please check this publication,https://pubmed.ncbi.nlm.nih.gov/15154768/Delaney

Graph Convolution modeling to predict SCN5A binding affinities

  • Tutorial-11: (Mode: AMPL_GPU; Time: ~ 18 minutes): This COLAB notebook will use AMPL for predicting binding affinities -pIC50 values- of ligands that could bind to human Sodium channel protein type 5 subunit alpha protein (Gene: SCN5A) using Graph Convolutional Network Model. ChEMBL database is the data source of binding affinities (pIC50) Test Image 1

3. Hyper-parameter Optimization (HPO), Uncertainty Quantification (UQ), and using metrics for analyzing model performance.

This notebook also explores AMPL functions for saving and loading prebuild AMPL models for analysis.

  • Tutorial-12 Hyper-parameter Optimization (HPO) and Uncertainty Quantification (UQ).
  • Tutorial-13 Notebook includes HPO Grid Search on three different modeling methods (Random Forest, NN and XGBoost).

4. Creating high-quality models

  • Tutorial-12 Notebook provides the framework for visualizing the results of HPO results and use them to identify best models.

5. Model Inference:

AMPL Workshops

Workshop date, June 05, 2021: Protein Target-focussed Binding Data Curation, Exploratory Data Analysis and Featurization using AMPL. Please note that Google Chrome browser works best with the COLAB Jupyter notebooks

Supporting links

Similar chemoinformatics, drug-discovery software tools:

Chemoinformatics databases

Acknowledgements

Most of the tutorial code chunks came from multple Jupyter notebooks generously shared by the ATOM team.

  • Amanda Paulson
  • Ben Madej
  • Da Shi
  • Hiran Ranganathan
  • Jessica Mauvais
  • Jonathan Allen
  • Kevin Mcloughlin
  • Sarangan Ravichandran
  • Stewart He
  • Ya Ju Fan
  • Contributions from the following student programs:

About

AMPL software tutorials

Resources

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

772 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Test Image 1

The ATOM Modeling PipeLine (AMPL; https://github.com/ATOMconsortium/AMPL) is an open-source, modular, extensible software pipeline for building and sharing models to advance in silico drug discovery. To see the list of AMPL parameters, please check this link, https://github.com/ATOMconsortium/AMPL/blob/master/atomsci/ddm/docs/PARAMETERS.md

This page contains a collection of AMPL-COLAB tutorial notebooks.

+ Please note that if you have trouble opening up any of the following notebooks, please go to, https://nbviewer.jupyter.org/, and paste the notebook link to view the contents.

0. Basic Google COLAB Introduction (Works best with Google Chrome)

  • Tutorial-00: Basic COLAB tutorial. For all the COLAB tutorials, click on the tutorial link, and then click on "Open in Colab" baner. You can open and run the notebook from the browser. If you want to save your edits to the notebook, you need to save a copy in your Google Drive. Usually, Google COLAB saves the notebook files under the "My Drive > Colab Notebooks" folder

1. Data Collection and creating Machine-Learning ready datasets:

The data that we gather for modeling is small-molecule/drug binding data. The following links will introduce some of the concepts and outcome measures related to this topic:

For the tutorials, we will use the small-molecule binding data obtained from either one of the following resources, ChEMBL (https://www.ebi.ac.uk/chembl/), Drug Target Commons (DTC; https://drugtargetcommons.fimm.fi/) and Excape-DB (https://solr.ideaconsult.net/search/excape/).

Click here to learn about single-target focussed data.

Explore HTR3A binding data from ExCAPE-DB

Explore HTR3A binding data from Drug Target Commons database

  • Tutorial-03: (Time: ~ 4 minutes) This COLAB notebook will use AMPL for Data cleaning, EDA and clustering of HTR3A protein data from Drug Target Commons (DTC)
  • Tutorial-04: (Time: ~ 10 minutes) This COLAB notebook will use AMPL for Data curation of HTR3A protein data from Drug Target Commons (DTC)

Curating, merging and visualizing two datasets

  • Tutorial-05: (Time: ~ 4 minutes) This COLAB notebook will use AMPL to upload datasets (small-molecule activity data from ChEMBL), clean, merge and do some basic Exploratory Data Analysis.
  • Tutorial-06: (Time: ~ 4 minutes) This COLAB notebook with use AMPL to merge HTR3A binding data from two different data sources, DTC and ExCAPE-DB.

Exploratory Data Analysis (EDA) Notebooks

  • Tutorial-07: (Time: ~ 4 minutes). The notebook uses HTR3A as the protein target. The notebook accomplishes the following tasks:
    • Uses AMPL software
    • Reads in data from three database sources: ChEMBL, Excape-DB and DTC
    • Cleans, standardizes and analyzes the data
    • Merges and harmonizes to create a dataset
  • Tutorial-08: Exploratory Data Analysis-Regression
  • Tutorial-09: Exploratory Data Analysis-Regression

2. Model training and tuning:

Random Forest modeling to predict solubility

  • Tutorial-10: (Time: ~ 2 minutes): Simple supervised learning example. AMPL will read the public data (117 chemical compounds), curate, fit a Random Forest model to predict solubility and test the model. For additional information on the dataset, please check this publication,https://pubmed.ncbi.nlm.nih.gov/15154768/Delaney

Graph Convolution modeling to predict SCN5A binding affinities

  • Tutorial-11: (Mode: AMPL_GPU; Time: ~ 18 minutes): This COLAB notebook will use AMPL for predicting binding affinities -pIC50 values- of ligands that could bind to human Sodium channel protein type 5 subunit alpha protein (Gene: SCN5A) using Graph Convolutional Network Model. ChEMBL database is the data source of binding affinities (pIC50) Test Image 1

3. Hyper-parameter Optimization (HPO), Uncertainty Quantification (UQ), and using metrics for analyzing model performance.

This notebook also explores AMPL functions for saving and loading prebuild AMPL models for analysis.

  • Tutorial-12 Hyper-parameter Optimization (HPO) and Uncertainty Quantification (UQ).
  • Tutorial-13 Notebook includes HPO Grid Search on three different modeling methods (Random Forest, NN and XGBoost).

4. Creating high-quality models

  • Tutorial-12 Notebook provides the framework for visualizing the results of HPO results and use them to identify best models.

5. Model Inference:

AMPL Workshops

Workshop date, June 05, 2021: Protein Target-focussed Binding Data Curation, Exploratory Data Analysis and Featurization using AMPL. Please note that Google Chrome browser works best with the COLAB Jupyter notebooks

Supporting links

Similar chemoinformatics, drug-discovery software tools:

Chemoinformatics databases

Acknowledgements

Most of the tutorial code chunks came from multple Jupyter notebooks generously shared by the ATOM team.

  • Amanda Paulson
  • Ben Madej
  • Da Shi
  • Hiran Ranganathan
  • Jessica Mauvais
  • Jonathan Allen
  • Kevin Mcloughlin
  • Sarangan Ravichandran
  • Stewart He
  • Ya Ju Fan
  • Contributions from the following student programs:

About

AMPL software tutorials

Resources

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

772 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Test Image 1

The ATOM Modeling PipeLine (AMPL; https://github.com/ATOMconsortium/AMPL) is an open-source, modular, extensible software pipeline for building and sharing models to advance in silico drug discovery. To see the list of AMPL parameters, please check this link, https://github.com/ATOMconsortium/AMPL/blob/master/atomsci/ddm/docs/PARAMETERS.md

This page contains a collection of AMPL-COLAB tutorial notebooks.

+ Please note that if you have trouble opening up any of the following notebooks, please go to, https://nbviewer.jupyter.org/, and paste the notebook link to view the contents.

0. Basic Google COLAB Introduction (Works best with Google Chrome)

  • Tutorial-00: Basic COLAB tutorial. For all the COLAB tutorials, click on the tutorial link, and then click on "Open in Colab" baner. You can open and run the notebook from the browser. If you want to save your edits to the notebook, you need to save a copy in your Google Drive. Usually, Google COLAB saves the notebook files under the "My Drive > Colab Notebooks" folder

1. Data Collection and creating Machine-Learning ready datasets:

The data that we gather for modeling is small-molecule/drug binding data. The following links will introduce some of the concepts and outcome measures related to this topic:

For the tutorials, we will use the small-molecule binding data obtained from either one of the following resources, ChEMBL (https://www.ebi.ac.uk/chembl/), Drug Target Commons (DTC; https://drugtargetcommons.fimm.fi/) and Excape-DB (https://solr.ideaconsult.net/search/excape/).

Click here to learn about single-target focussed data.

Explore HTR3A binding data from ExCAPE-DB

Explore HTR3A binding data from Drug Target Commons database

  • Tutorial-03: (Time: ~ 4 minutes) This COLAB notebook will use AMPL for Data cleaning, EDA and clustering of HTR3A protein data from Drug Target Commons (DTC)
  • Tutorial-04: (Time: ~ 10 minutes) This COLAB notebook will use AMPL for Data curation of HTR3A protein data from Drug Target Commons (DTC)

Curating, merging and visualizing two datasets

  • Tutorial-05: (Time: ~ 4 minutes) This COLAB notebook will use AMPL to upload datasets (small-molecule activity data from ChEMBL), clean, merge and do some basic Exploratory Data Analysis.
  • Tutorial-06: (Time: ~ 4 minutes) This COLAB notebook with use AMPL to merge HTR3A binding data from two different data sources, DTC and ExCAPE-DB.

Exploratory Data Analysis (EDA) Notebooks

  • Tutorial-07: (Time: ~ 4 minutes). The notebook uses HTR3A as the protein target. The notebook accomplishes the following tasks:
    • Uses AMPL software
    • Reads in data from three database sources: ChEMBL, Excape-DB and DTC
    • Cleans, standardizes and analyzes the data
    • Merges and harmonizes to create a dataset
  • Tutorial-08: Exploratory Data Analysis-Regression
  • Tutorial-09: Exploratory Data Analysis-Regression

2. Model training and tuning:

Random Forest modeling to predict solubility

  • Tutorial-10: (Time: ~ 2 minutes): Simple supervised learning example. AMPL will read the public data (117 chemical compounds), curate, fit a Random Forest model to predict solubility and test the model. For additional information on the dataset, please check this publication,https://pubmed.ncbi.nlm.nih.gov/15154768/Delaney

Graph Convolution modeling to predict SCN5A binding affinities

  • Tutorial-11: (Mode: AMPL_GPU; Time: ~ 18 minutes): This COLAB notebook will use AMPL for predicting binding affinities -pIC50 values- of ligands that could bind to human Sodium channel protein type 5 subunit alpha protein (Gene: SCN5A) using Graph Convolutional Network Model. ChEMBL database is the data source of binding affinities (pIC50) Test Image 1

3. Hyper-parameter Optimization (HPO), Uncertainty Quantification (UQ), and using metrics for analyzing model performance.

This notebook also explores AMPL functions for saving and loading prebuild AMPL models for analysis.

  • Tutorial-12 Hyper-parameter Optimization (HPO) and Uncertainty Quantification (UQ).
  • Tutorial-13 Notebook includes HPO Grid Search on three different modeling methods (Random Forest, NN and XGBoost).

4. Creating high-quality models

  • Tutorial-12 Notebook provides the framework for visualizing the results of HPO results and use them to identify best models.

5. Model Inference:

AMPL Workshops

Workshop date, June 05, 2021: Protein Target-focussed Binding Data Curation, Exploratory Data Analysis and Featurization using AMPL. Please note that Google Chrome browser works best with the COLAB Jupyter notebooks

Supporting links

Similar chemoinformatics, drug-discovery software tools:

Chemoinformatics databases

Acknowledgements

Most of the tutorial code chunks came from multple Jupyter notebooks generously shared by the ATOM team.

  • Amanda Paulson
  • Ben Madej
  • Da Shi
  • Hiran Ranganathan
  • Jessica Mauvais
  • Jonathan Allen
  • Kevin Mcloughlin
  • Sarangan Ravichandran
  • Stewart He
  • Ya Ju Fan
  • Contributions from the following student programs:

About

AMPL software tutorials

Resources

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

772 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Test Image 1

The ATOM Modeling PipeLine (AMPL; https://github.com/ATOMconsortium/AMPL) is an open-source, modular, extensible software pipeline for building and sharing models to advance in silico drug discovery. To see the list of AMPL parameters, please check this link, https://github.com/ATOMconsortium/AMPL/blob/master/atomsci/ddm/docs/PARAMETERS.md

This page contains a collection of AMPL-COLAB tutorial notebooks.

+ Please note that if you have trouble opening up any of the following notebooks, please go to, https://nbviewer.jupyter.org/, and paste the notebook link to view the contents.

0. Basic Google COLAB Introduction (Works best with Google Chrome)

  • Tutorial-00: Basic COLAB tutorial. For all the COLAB tutorials, click on the tutorial link, and then click on "Open in Colab" baner. You can open and run the notebook from the browser. If you want to save your edits to the notebook, you need to save a copy in your Google Drive. Usually, Google COLAB saves the notebook files under the "My Drive > Colab Notebooks" folder

1. Data Collection and creating Machine-Learning ready datasets:

The data that we gather for modeling is small-molecule/drug binding data. The following links will introduce some of the concepts and outcome measures related to this topic:

For the tutorials, we will use the small-molecule binding data obtained from either one of the following resources, ChEMBL (https://www.ebi.ac.uk/chembl/), Drug Target Commons (DTC; https://drugtargetcommons.fimm.fi/) and Excape-DB (https://solr.ideaconsult.net/search/excape/).

Click here to learn about single-target focussed data.

Explore HTR3A binding data from ExCAPE-DB

Explore HTR3A binding data from Drug Target Commons database

  • Tutorial-03: (Time: ~ 4 minutes) This COLAB notebook will use AMPL for Data cleaning, EDA and clustering of HTR3A protein data from Drug Target Commons (DTC)
  • Tutorial-04: (Time: ~ 10 minutes) This COLAB notebook will use AMPL for Data curation of HTR3A protein data from Drug Target Commons (DTC)

Curating, merging and visualizing two datasets

  • Tutorial-05: (Time: ~ 4 minutes) This COLAB notebook will use AMPL to upload datasets (small-molecule activity data from ChEMBL), clean, merge and do some basic Exploratory Data Analysis.
  • Tutorial-06: (Time: ~ 4 minutes) This COLAB notebook with use AMPL to merge HTR3A binding data from two different data sources, DTC and ExCAPE-DB.

Exploratory Data Analysis (EDA) Notebooks

  • Tutorial-07: (Time: ~ 4 minutes). The notebook uses HTR3A as the protein target. The notebook accomplishes the following tasks:
    • Uses AMPL software
    • Reads in data from three database sources: ChEMBL, Excape-DB and DTC
    • Cleans, standardizes and analyzes the data
    • Merges and harmonizes to create a dataset
  • Tutorial-08: Exploratory Data Analysis-Regression
  • Tutorial-09: Exploratory Data Analysis-Regression

2. Model training and tuning:

Random Forest modeling to predict solubility

  • Tutorial-10: (Time: ~ 2 minutes): Simple supervised learning example. AMPL will read the public data (117 chemical compounds), curate, fit a Random Forest model to predict solubility and test the model. For additional information on the dataset, please check this publication,https://pubmed.ncbi.nlm.nih.gov/15154768/Delaney

Graph Convolution modeling to predict SCN5A binding affinities

  • Tutorial-11: (Mode: AMPL_GPU; Time: ~ 18 minutes): This COLAB notebook will use AMPL for predicting binding affinities -pIC50 values- of ligands that could bind to human Sodium channel protein type 5 subunit alpha protein (Gene: SCN5A) using Graph Convolutional Network Model. ChEMBL database is the data source of binding affinities (pIC50) Test Image 1

3. Hyper-parameter Optimization (HPO), Uncertainty Quantification (UQ), and using metrics for analyzing model performance.

This notebook also explores AMPL functions for saving and loading prebuild AMPL models for analysis.

  • Tutorial-12 Hyper-parameter Optimization (HPO) and Uncertainty Quantification (UQ).
  • Tutorial-13 Notebook includes HPO Grid Search on three different modeling methods (Random Forest, NN and XGBoost).

4. Creating high-quality models

  • Tutorial-12 Notebook provides the framework for visualizing the results of HPO results and use them to identify best models.

5. Model Inference:

AMPL Workshops

Workshop date, June 05, 2021: Protein Target-focussed Binding Data Curation, Exploratory Data Analysis and Featurization using AMPL. Please note that Google Chrome browser works best with the COLAB Jupyter notebooks

Supporting links

Similar chemoinformatics, drug-discovery software tools:

Chemoinformatics databases

Acknowledgements

Most of the tutorial code chunks came from multple Jupyter notebooks generously shared by the ATOM team.

  • Amanda Paulson
  • Ben Madej
  • Da Shi
  • Hiran Ranganathan
  • Jessica Mauvais
  • Jonathan Allen
  • Kevin Mcloughlin
  • Sarangan Ravichandran
  • Stewart He
  • Ya Ju Fan
  • Contributions from the following student programs:

About

AMPL software tutorials

Resources

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages