Skip to content

Repository files navigation

Adversarial attack of sequence-free enhancer prediction identifies patterns of chromatin architecture

Jamil Gafur1, Olivia W Lang2, & William KM Lai2,3

1 Department of Computer Science, University of Iowa, Iowa City, Iowa 52242, USA

2 Department of Molecular Biology and Genetics, Cornell University, Ithaca, NY 14850, USA

3 Department of Computational Biology, Cornell University, Ithaca, NY 14850, USA

Abstract

The wide range of cellular complexity created by multicellular organisms is due in large part to the intricate and synergistic interplay of regulatory complexes throughout the eukaryotic genome. These regulatory elements ‘enhance’ specific gene programs and have been shown to operate in diverse networks that are distinct across cell states of the same organism. Attempts to characterize and predict enhancers have typically focused on leveraging information-dense DNA sequence in parallel with epigenomic assays. We examined the viability of enhancer prediction using only a minimal set of epigenomic datasets without direct DNA information. We demonstrate that chromatin datasets are sufficient to identify enhancers genome-wide with high accuracy. By training networks leveraging data from multiple cell types simultaneously, we generated a cell-type invariant enhancer prediction platform that utilized only the patterns of protein binding for inference. We also showed the utility of swarm-based adversarial attacks (APSO) to deconvolute trained genomic neural networks for the first time. Critically, unlike saliency mapping or other game-theory based approaches, APSO is completely network-architecture independent and can be applied to any prediction engine to derive the features that drive inference.

Network Training and Analysis

This repo has been designed to completely re-generate the analysis and figures contained within the manuscript (DOI: XXXXXX). Random seeds have been set as neccessary to maximize reproducibility. Scripts should be executed in numerical order per sub-folder.

1. Preprocessing

  • Navigate to the Preprocessing folder and follow the instructions

2. Building the environment

conda create --prefix ~/work/ChromEnh python=3.9
conda activate ~/work/ChromEnh
conda install pytorch torchvision torchaudio pytorch-cuda=11.8 -c pytorch -c nvidia

Validate pytorch is able to see GPU:

python
import torch
torch.cuda.is_available()
pip install h5py tqdm matplotlib scikit-learn seaborn

The APSO code should be cloned from here:

git clone https://github.com/EpiGenomicsCode/Adversarial_Observation

And moved into the 'util' folder.

3. Download data

Shell scripts should be executed sequentially in:

  • 00_preprocessing

4. Train networks

There are currently three different studies that can be performed:

  1. Chromosome dropout traininig
  2. Cell line independent training
  3. Large model training

These networks can be executed asynchronously in:

  • 01_network_training

5. Explainable AI analysis

Shell scripts should be executed sequentially in:

  • 03_xai

6. Large network analysis

Shell scripts should be executed sequentially in:

  • 04_largenetwork

FAQ

  1. How do I add my own model?
  • Navigate to the util/Model folder and add your model as a new python file
  • Edit the loadModel function in the util.py file

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - EpiGenomicsCode/ChromEnhancer · GitHub
Skip to content

Repository files navigation

Adversarial attack of sequence-free enhancer prediction identifies patterns of chromatin architecture

Jamil Gafur1, Olivia W Lang2, & William KM Lai2,3

1 Department of Computer Science, University of Iowa, Iowa City, Iowa 52242, USA

2 Department of Molecular Biology and Genetics, Cornell University, Ithaca, NY 14850, USA

3 Department of Computational Biology, Cornell University, Ithaca, NY 14850, USA

Abstract

The wide range of cellular complexity created by multicellular organisms is due in large part to the intricate and synergistic interplay of regulatory complexes throughout the eukaryotic genome. These regulatory elements ‘enhance’ specific gene programs and have been shown to operate in diverse networks that are distinct across cell states of the same organism. Attempts to characterize and predict enhancers have typically focused on leveraging information-dense DNA sequence in parallel with epigenomic assays. We examined the viability of enhancer prediction using only a minimal set of epigenomic datasets without direct DNA information. We demonstrate that chromatin datasets are sufficient to identify enhancers genome-wide with high accuracy. By training networks leveraging data from multiple cell types simultaneously, we generated a cell-type invariant enhancer prediction platform that utilized only the patterns of protein binding for inference. We also showed the utility of swarm-based adversarial attacks (APSO) to deconvolute trained genomic neural networks for the first time. Critically, unlike saliency mapping or other game-theory based approaches, APSO is completely network-architecture independent and can be applied to any prediction engine to derive the features that drive inference.

Network Training and Analysis

This repo has been designed to completely re-generate the analysis and figures contained within the manuscript (DOI: XXXXXX). Random seeds have been set as neccessary to maximize reproducibility. Scripts should be executed in numerical order per sub-folder.

1. Preprocessing

  • Navigate to the Preprocessing folder and follow the instructions

2. Building the environment

conda create --prefix ~/work/ChromEnh python=3.9
conda activate ~/work/ChromEnh
conda install pytorch torchvision torchaudio pytorch-cuda=11.8 -c pytorch -c nvidia

Validate pytorch is able to see GPU:

python
import torch
torch.cuda.is_available()
pip install h5py tqdm matplotlib scikit-learn seaborn

The APSO code should be cloned from here:

git clone https://github.com/EpiGenomicsCode/Adversarial_Observation

And moved into the 'util' folder.

3. Download data

Shell scripts should be executed sequentially in:

  • 00_preprocessing

4. Train networks

There are currently three different studies that can be performed:

  1. Chromosome dropout traininig
  2. Cell line independent training
  3. Large model training

These networks can be executed asynchronously in:

  • 01_network_training

5. Explainable AI analysis

Shell scripts should be executed sequentially in:

  • 03_xai

6. Large network analysis

Shell scripts should be executed sequentially in:

  • 04_largenetwork

FAQ

  1. How do I add my own model?
  • Navigate to the util/Model folder and add your model as a new python file
  • Edit the loadModel function in the util.py file

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - EpiGenomicsCode/ChromEnhancer · GitHub
Skip to content

Repository files navigation

Adversarial attack of sequence-free enhancer prediction identifies patterns of chromatin architecture

Jamil Gafur1, Olivia W Lang2, & William KM Lai2,3

1 Department of Computer Science, University of Iowa, Iowa City, Iowa 52242, USA

2 Department of Molecular Biology and Genetics, Cornell University, Ithaca, NY 14850, USA

3 Department of Computational Biology, Cornell University, Ithaca, NY 14850, USA

Abstract

The wide range of cellular complexity created by multicellular organisms is due in large part to the intricate and synergistic interplay of regulatory complexes throughout the eukaryotic genome. These regulatory elements ‘enhance’ specific gene programs and have been shown to operate in diverse networks that are distinct across cell states of the same organism. Attempts to characterize and predict enhancers have typically focused on leveraging information-dense DNA sequence in parallel with epigenomic assays. We examined the viability of enhancer prediction using only a minimal set of epigenomic datasets without direct DNA information. We demonstrate that chromatin datasets are sufficient to identify enhancers genome-wide with high accuracy. By training networks leveraging data from multiple cell types simultaneously, we generated a cell-type invariant enhancer prediction platform that utilized only the patterns of protein binding for inference. We also showed the utility of swarm-based adversarial attacks (APSO) to deconvolute trained genomic neural networks for the first time. Critically, unlike saliency mapping or other game-theory based approaches, APSO is completely network-architecture independent and can be applied to any prediction engine to derive the features that drive inference.

Network Training and Analysis

This repo has been designed to completely re-generate the analysis and figures contained within the manuscript (DOI: XXXXXX). Random seeds have been set as neccessary to maximize reproducibility. Scripts should be executed in numerical order per sub-folder.

1. Preprocessing

  • Navigate to the Preprocessing folder and follow the instructions

2. Building the environment

conda create --prefix ~/work/ChromEnh python=3.9
conda activate ~/work/ChromEnh
conda install pytorch torchvision torchaudio pytorch-cuda=11.8 -c pytorch -c nvidia

Validate pytorch is able to see GPU:

python
import torch
torch.cuda.is_available()
pip install h5py tqdm matplotlib scikit-learn seaborn

The APSO code should be cloned from here:

git clone https://github.com/EpiGenomicsCode/Adversarial_Observation

And moved into the 'util' folder.

3. Download data

Shell scripts should be executed sequentially in:

  • 00_preprocessing

4. Train networks

There are currently three different studies that can be performed:

  1. Chromosome dropout traininig
  2. Cell line independent training
  3. Large model training

These networks can be executed asynchronously in:

  • 01_network_training

5. Explainable AI analysis

Shell scripts should be executed sequentially in:

  • 03_xai

6. Large network analysis

Shell scripts should be executed sequentially in:

  • 04_largenetwork

FAQ

  1. How do I add my own model?
  • Navigate to the util/Model folder and add your model as a new python file
  • Edit the loadModel function in the util.py file

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - EpiGenomicsCode/ChromEnhancer · GitHub
Skip to content

Repository files navigation

Adversarial attack of sequence-free enhancer prediction identifies patterns of chromatin architecture

Jamil Gafur1, Olivia W Lang2, & William KM Lai2,3

1 Department of Computer Science, University of Iowa, Iowa City, Iowa 52242, USA

2 Department of Molecular Biology and Genetics, Cornell University, Ithaca, NY 14850, USA

3 Department of Computational Biology, Cornell University, Ithaca, NY 14850, USA

Abstract

The wide range of cellular complexity created by multicellular organisms is due in large part to the intricate and synergistic interplay of regulatory complexes throughout the eukaryotic genome. These regulatory elements ‘enhance’ specific gene programs and have been shown to operate in diverse networks that are distinct across cell states of the same organism. Attempts to characterize and predict enhancers have typically focused on leveraging information-dense DNA sequence in parallel with epigenomic assays. We examined the viability of enhancer prediction using only a minimal set of epigenomic datasets without direct DNA information. We demonstrate that chromatin datasets are sufficient to identify enhancers genome-wide with high accuracy. By training networks leveraging data from multiple cell types simultaneously, we generated a cell-type invariant enhancer prediction platform that utilized only the patterns of protein binding for inference. We also showed the utility of swarm-based adversarial attacks (APSO) to deconvolute trained genomic neural networks for the first time. Critically, unlike saliency mapping or other game-theory based approaches, APSO is completely network-architecture independent and can be applied to any prediction engine to derive the features that drive inference.

Network Training and Analysis

This repo has been designed to completely re-generate the analysis and figures contained within the manuscript (DOI: XXXXXX). Random seeds have been set as neccessary to maximize reproducibility. Scripts should be executed in numerical order per sub-folder.

1. Preprocessing

  • Navigate to the Preprocessing folder and follow the instructions

2. Building the environment

conda create --prefix ~/work/ChromEnh python=3.9
conda activate ~/work/ChromEnh
conda install pytorch torchvision torchaudio pytorch-cuda=11.8 -c pytorch -c nvidia

Validate pytorch is able to see GPU:

python
import torch
torch.cuda.is_available()
pip install h5py tqdm matplotlib scikit-learn seaborn

The APSO code should be cloned from here:

git clone https://github.com/EpiGenomicsCode/Adversarial_Observation

And moved into the 'util' folder.

3. Download data

Shell scripts should be executed sequentially in:

  • 00_preprocessing

4. Train networks

There are currently three different studies that can be performed:

  1. Chromosome dropout traininig
  2. Cell line independent training
  3. Large model training

These networks can be executed asynchronously in:

  • 01_network_training

5. Explainable AI analysis

Shell scripts should be executed sequentially in:

  • 03_xai

6. Large network analysis

Shell scripts should be executed sequentially in:

  • 04_largenetwork

FAQ

  1. How do I add my own model?
  • Navigate to the util/Model folder and add your model as a new python file
  • Edit the loadModel function in the util.py file

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - EpiGenomicsCode/ChromEnhancer · GitHub
Skip to content

Repository files navigation

Adversarial attack of sequence-free enhancer prediction identifies patterns of chromatin architecture

Jamil Gafur1, Olivia W Lang2, & William KM Lai2,3

1 Department of Computer Science, University of Iowa, Iowa City, Iowa 52242, USA

2 Department of Molecular Biology and Genetics, Cornell University, Ithaca, NY 14850, USA

3 Department of Computational Biology, Cornell University, Ithaca, NY 14850, USA

Abstract

The wide range of cellular complexity created by multicellular organisms is due in large part to the intricate and synergistic interplay of regulatory complexes throughout the eukaryotic genome. These regulatory elements ‘enhance’ specific gene programs and have been shown to operate in diverse networks that are distinct across cell states of the same organism. Attempts to characterize and predict enhancers have typically focused on leveraging information-dense DNA sequence in parallel with epigenomic assays. We examined the viability of enhancer prediction using only a minimal set of epigenomic datasets without direct DNA information. We demonstrate that chromatin datasets are sufficient to identify enhancers genome-wide with high accuracy. By training networks leveraging data from multiple cell types simultaneously, we generated a cell-type invariant enhancer prediction platform that utilized only the patterns of protein binding for inference. We also showed the utility of swarm-based adversarial attacks (APSO) to deconvolute trained genomic neural networks for the first time. Critically, unlike saliency mapping or other game-theory based approaches, APSO is completely network-architecture independent and can be applied to any prediction engine to derive the features that drive inference.

Network Training and Analysis

This repo has been designed to completely re-generate the analysis and figures contained within the manuscript (DOI: XXXXXX). Random seeds have been set as neccessary to maximize reproducibility. Scripts should be executed in numerical order per sub-folder.

1. Preprocessing

  • Navigate to the Preprocessing folder and follow the instructions

2. Building the environment

conda create --prefix ~/work/ChromEnh python=3.9
conda activate ~/work/ChromEnh
conda install pytorch torchvision torchaudio pytorch-cuda=11.8 -c pytorch -c nvidia

Validate pytorch is able to see GPU:

python
import torch
torch.cuda.is_available()
pip install h5py tqdm matplotlib scikit-learn seaborn

The APSO code should be cloned from here:

git clone https://github.com/EpiGenomicsCode/Adversarial_Observation

And moved into the 'util' folder.

3. Download data

Shell scripts should be executed sequentially in:

  • 00_preprocessing

4. Train networks

There are currently three different studies that can be performed:

  1. Chromosome dropout traininig
  2. Cell line independent training
  3. Large model training

These networks can be executed asynchronously in:

  • 01_network_training

5. Explainable AI analysis

Shell scripts should be executed sequentially in:

  • 03_xai

6. Large network analysis

Shell scripts should be executed sequentially in:

  • 04_largenetwork

FAQ

  1. How do I add my own model?
  • Navigate to the util/Model folder and add your model as a new python file
  • Edit the loadModel function in the util.py file

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - EpiGenomicsCode/ChromEnhancer · GitHub
Skip to content

Repository files navigation

Adversarial attack of sequence-free enhancer prediction identifies patterns of chromatin architecture

Jamil Gafur1, Olivia W Lang2, & William KM Lai2,3

1 Department of Computer Science, University of Iowa, Iowa City, Iowa 52242, USA

2 Department of Molecular Biology and Genetics, Cornell University, Ithaca, NY 14850, USA

3 Department of Computational Biology, Cornell University, Ithaca, NY 14850, USA

Abstract

The wide range of cellular complexity created by multicellular organisms is due in large part to the intricate and synergistic interplay of regulatory complexes throughout the eukaryotic genome. These regulatory elements ‘enhance’ specific gene programs and have been shown to operate in diverse networks that are distinct across cell states of the same organism. Attempts to characterize and predict enhancers have typically focused on leveraging information-dense DNA sequence in parallel with epigenomic assays. We examined the viability of enhancer prediction using only a minimal set of epigenomic datasets without direct DNA information. We demonstrate that chromatin datasets are sufficient to identify enhancers genome-wide with high accuracy. By training networks leveraging data from multiple cell types simultaneously, we generated a cell-type invariant enhancer prediction platform that utilized only the patterns of protein binding for inference. We also showed the utility of swarm-based adversarial attacks (APSO) to deconvolute trained genomic neural networks for the first time. Critically, unlike saliency mapping or other game-theory based approaches, APSO is completely network-architecture independent and can be applied to any prediction engine to derive the features that drive inference.

Network Training and Analysis

This repo has been designed to completely re-generate the analysis and figures contained within the manuscript (DOI: XXXXXX). Random seeds have been set as neccessary to maximize reproducibility. Scripts should be executed in numerical order per sub-folder.

1. Preprocessing

  • Navigate to the Preprocessing folder and follow the instructions

2. Building the environment

conda create --prefix ~/work/ChromEnh python=3.9
conda activate ~/work/ChromEnh
conda install pytorch torchvision torchaudio pytorch-cuda=11.8 -c pytorch -c nvidia

Validate pytorch is able to see GPU:

python
import torch
torch.cuda.is_available()
pip install h5py tqdm matplotlib scikit-learn seaborn

The APSO code should be cloned from here:

git clone https://github.com/EpiGenomicsCode/Adversarial_Observation

And moved into the 'util' folder.

3. Download data

Shell scripts should be executed sequentially in:

  • 00_preprocessing

4. Train networks

There are currently three different studies that can be performed:

  1. Chromosome dropout traininig
  2. Cell line independent training
  3. Large model training

These networks can be executed asynchronously in:

  • 01_network_training

5. Explainable AI analysis

Shell scripts should be executed sequentially in:

  • 03_xai

6. Large network analysis

Shell scripts should be executed sequentially in:

  • 04_largenetwork

FAQ

  1. How do I add my own model?
  • Navigate to the util/Model folder and add your model as a new python file
  • Edit the loadModel function in the util.py file

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); })(); GitHub - EpiGenomicsCode/ChromEnhancer · GitHub
Skip to content

Repository files navigation

Adversarial attack of sequence-free enhancer prediction identifies patterns of chromatin architecture

Jamil Gafur1, Olivia W Lang2, & William KM Lai2,3

1 Department of Computer Science, University of Iowa, Iowa City, Iowa 52242, USA

2 Department of Molecular Biology and Genetics, Cornell University, Ithaca, NY 14850, USA

3 Department of Computational Biology, Cornell University, Ithaca, NY 14850, USA

Abstract

The wide range of cellular complexity created by multicellular organisms is due in large part to the intricate and synergistic interplay of regulatory complexes throughout the eukaryotic genome. These regulatory elements ‘enhance’ specific gene programs and have been shown to operate in diverse networks that are distinct across cell states of the same organism. Attempts to characterize and predict enhancers have typically focused on leveraging information-dense DNA sequence in parallel with epigenomic assays. We examined the viability of enhancer prediction using only a minimal set of epigenomic datasets without direct DNA information. We demonstrate that chromatin datasets are sufficient to identify enhancers genome-wide with high accuracy. By training networks leveraging data from multiple cell types simultaneously, we generated a cell-type invariant enhancer prediction platform that utilized only the patterns of protein binding for inference. We also showed the utility of swarm-based adversarial attacks (APSO) to deconvolute trained genomic neural networks for the first time. Critically, unlike saliency mapping or other game-theory based approaches, APSO is completely network-architecture independent and can be applied to any prediction engine to derive the features that drive inference.

Network Training and Analysis

This repo has been designed to completely re-generate the analysis and figures contained within the manuscript (DOI: XXXXXX). Random seeds have been set as neccessary to maximize reproducibility. Scripts should be executed in numerical order per sub-folder.

1. Preprocessing

  • Navigate to the Preprocessing folder and follow the instructions

2. Building the environment

conda create --prefix ~/work/ChromEnh python=3.9
conda activate ~/work/ChromEnh
conda install pytorch torchvision torchaudio pytorch-cuda=11.8 -c pytorch -c nvidia

Validate pytorch is able to see GPU:

python
import torch
torch.cuda.is_available()
pip install h5py tqdm matplotlib scikit-learn seaborn

The APSO code should be cloned from here:

git clone https://github.com/EpiGenomicsCode/Adversarial_Observation

And moved into the 'util' folder.

3. Download data

Shell scripts should be executed sequentially in:

  • 00_preprocessing

4. Train networks

There are currently three different studies that can be performed:

  1. Chromosome dropout traininig
  2. Cell line independent training
  3. Large model training

These networks can be executed asynchronously in:

  • 01_network_training

5. Explainable AI analysis

Shell scripts should be executed sequentially in:

  • 03_xai

6. Large network analysis

Shell scripts should be executed sequentially in:

  • 04_largenetwork

FAQ

  1. How do I add my own model?
  • Navigate to the util/Model folder and add your model as a new python file
  • Edit the loadModel function in the util.py file

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages