Repository files navigation

Gene Expression Prediction Using 1D Convolution Network

This is the official repository for the paper Gene Expression Prediction Using a Deep 1D Convolution Neural Network. This repository contains the basic code to rreplicate our set of experiments.

There were two sets of experiments which we ran:

  • Each cell type is taken individually and a model is trained and tested upon it
  • The train and validation data is mixed and only one model is trained on this entire new data set and then testing is performed individually on each cell type

Data

The data for these experiments can be downloaded from here. The data contains gene expression values and the five kinds of histone modifications for 56 cell types. Each cell type contains 19802 gene samples. For each gene, the region containing 10,000 base pairs surrounding the Transcription Start Site (TSS) was divided into 100 bins of 100 base pairs each and for each bin, five different modification signals were recorded.

data

Histone modifications for one cell type

Model

The model created is a 1D convolutional neural network followed by a dense network. We use a 1D convolution network to map the spatial sequence of the histone modification signals. For more details please have a look at the paper.

model

The proposed network architecture

Usage

There are five python files in this repository

  • Convnet.py - This defines the model architecture along with the training and testing loops
  • data.py - This script loads and transforms the data in a format suitable for ingestion
  • utils.py - This describes some of the utility functions
  • roc_auc_callback.py - This defines a area under the ROC curve callback to keep track of how the model performance changes across various epochs
  • main.py - This is the main file which defines various hyperparameters and other configurations. It also runs a complete training and evaluation step for the data

Only the main.py is sufficient to train and test our model. There are various settings available in the main.py which can be accessed as

python main.py -h
usage: main.py [-h] [--epochs EPOCHS] [--batch_size BATCH_SIZE] [--lr LR] [--lr_decay LR_DECAY] [--data_dir DATA_DIR] [--debug] [--save_dir SAVE_DIR] [--mix MIX]
1D Convolutional Network on Histone Modification Signals.
optional arguments:
-h, --help show this help message and exit
--epochs EPOCHS
--batch_size BATCH_SIZE
--lr LR Initial learning rate
--lr_decay LR_DECAY The value multiplied by lr at each epoch. Set a larger value for larger epochs
--data_dir DATA_DIR The directory of the data files
--debug Save weights by TensorBoard
--save_dir SAVE_DIR
--mix MIX Train on the data by mixing all cell types

To train and test with default values use

python main.py

Or if you need to provide different values of some parameters, you can use

python main.py --epochs 10 --lr 0.001

Once the training and testing is finsished the details about the training loss, validation loss and other evaluation metrics like auc are stored in the logs folder or in folders with names suggesting the parameters of the run. Those saved results could be used to plot training curves and evaluate the model.

Results

This model was the first generalized deep learning model for Gene Expression Prediction which even outperfomed the state of the art. This method was computationally very inexpensive as the model trained is pretty small and the number of training steps to be performed to achieve better than the state of the art performance is also small. Earlier methods used to train different models for different cell types but this approach only requires training the model once and then predictions can be made on various cell types.

result

AUC score vs Cell Types (comparison of our method with different baselines)

References

  • V. Chaubey, M. S. Nair and G. N. Pillai, "Gene Expression Prediction Using a Deep 1D Convolution Neural Network," 2019 IEEE Symposium Series on Computational Intelligence (SSCI), Xiamen, China, 2019, pp. 1383-1389, doi: 10.1109/SSCI44817.2019.9002669.
  • Ritambhara Singh, Jack Lanchantin, Arshdeep Sekhon, & Yanjun Qi. (2019). Dataset for DeepChrome and AttentiveChrome [Data set]. Zenodo. http://doi.org/10.5281/zenodo.2652278

About

Official repository for the paper "Gene Expression Prediction Using 1D Convolution Neural Network"

Topics

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Gene Expression Prediction Using 1D Convolution Network

This is the official repository for the paper Gene Expression Prediction Using a Deep 1D Convolution Neural Network. This repository contains the basic code to rreplicate our set of experiments.

There were two sets of experiments which we ran:

  • Each cell type is taken individually and a model is trained and tested upon it
  • The train and validation data is mixed and only one model is trained on this entire new data set and then testing is performed individually on each cell type

Data

The data for these experiments can be downloaded from here. The data contains gene expression values and the five kinds of histone modifications for 56 cell types. Each cell type contains 19802 gene samples. For each gene, the region containing 10,000 base pairs surrounding the Transcription Start Site (TSS) was divided into 100 bins of 100 base pairs each and for each bin, five different modification signals were recorded.

data

Histone modifications for one cell type

Model

The model created is a 1D convolutional neural network followed by a dense network. We use a 1D convolution network to map the spatial sequence of the histone modification signals. For more details please have a look at the paper.

model

The proposed network architecture

Usage

There are five python files in this repository

  • Convnet.py - This defines the model architecture along with the training and testing loops
  • data.py - This script loads and transforms the data in a format suitable for ingestion
  • utils.py - This describes some of the utility functions
  • roc_auc_callback.py - This defines a area under the ROC curve callback to keep track of how the model performance changes across various epochs
  • main.py - This is the main file which defines various hyperparameters and other configurations. It also runs a complete training and evaluation step for the data

Only the main.py is sufficient to train and test our model. There are various settings available in the main.py which can be accessed as

python main.py -h
usage: main.py [-h] [--epochs EPOCHS] [--batch_size BATCH_SIZE] [--lr LR] [--lr_decay LR_DECAY] [--data_dir DATA_DIR] [--debug] [--save_dir SAVE_DIR] [--mix MIX]
1D Convolutional Network on Histone Modification Signals.
optional arguments:
-h, --help show this help message and exit
--epochs EPOCHS
--batch_size BATCH_SIZE
--lr LR Initial learning rate
--lr_decay LR_DECAY The value multiplied by lr at each epoch. Set a larger value for larger epochs
--data_dir DATA_DIR The directory of the data files
--debug Save weights by TensorBoard
--save_dir SAVE_DIR
--mix MIX Train on the data by mixing all cell types

To train and test with default values use

python main.py

Or if you need to provide different values of some parameters, you can use

python main.py --epochs 10 --lr 0.001

Once the training and testing is finsished the details about the training loss, validation loss and other evaluation metrics like auc are stored in the logs folder or in folders with names suggesting the parameters of the run. Those saved results could be used to plot training curves and evaluate the model.

Results

This model was the first generalized deep learning model for Gene Expression Prediction which even outperfomed the state of the art. This method was computationally very inexpensive as the model trained is pretty small and the number of training steps to be performed to achieve better than the state of the art performance is also small. Earlier methods used to train different models for different cell types but this approach only requires training the model once and then predictions can be made on various cell types.

result

AUC score vs Cell Types (comparison of our method with different baselines)

References

  • V. Chaubey, M. S. Nair and G. N. Pillai, "Gene Expression Prediction Using a Deep 1D Convolution Neural Network," 2019 IEEE Symposium Series on Computational Intelligence (SSCI), Xiamen, China, 2019, pp. 1383-1389, doi: 10.1109/SSCI44817.2019.9002669.
  • Ritambhara Singh, Jack Lanchantin, Arshdeep Sekhon, & Yanjun Qi. (2019). Dataset for DeepChrome and AttentiveChrome [Data set]. Zenodo. http://doi.org/10.5281/zenodo.2652278

About

Official repository for the paper "Gene Expression Prediction Using 1D Convolution Neural Network"

Topics

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Gene Expression Prediction Using 1D Convolution Network

This is the official repository for the paper Gene Expression Prediction Using a Deep 1D Convolution Neural Network. This repository contains the basic code to rreplicate our set of experiments.

There were two sets of experiments which we ran:

  • Each cell type is taken individually and a model is trained and tested upon it
  • The train and validation data is mixed and only one model is trained on this entire new data set and then testing is performed individually on each cell type

Data

The data for these experiments can be downloaded from here. The data contains gene expression values and the five kinds of histone modifications for 56 cell types. Each cell type contains 19802 gene samples. For each gene, the region containing 10,000 base pairs surrounding the Transcription Start Site (TSS) was divided into 100 bins of 100 base pairs each and for each bin, five different modification signals were recorded.

data

Histone modifications for one cell type

Model

The model created is a 1D convolutional neural network followed by a dense network. We use a 1D convolution network to map the spatial sequence of the histone modification signals. For more details please have a look at the paper.

model

The proposed network architecture

Usage

There are five python files in this repository

  • Convnet.py - This defines the model architecture along with the training and testing loops
  • data.py - This script loads and transforms the data in a format suitable for ingestion
  • utils.py - This describes some of the utility functions
  • roc_auc_callback.py - This defines a area under the ROC curve callback to keep track of how the model performance changes across various epochs
  • main.py - This is the main file which defines various hyperparameters and other configurations. It also runs a complete training and evaluation step for the data

Only the main.py is sufficient to train and test our model. There are various settings available in the main.py which can be accessed as

python main.py -h
usage: main.py [-h] [--epochs EPOCHS] [--batch_size BATCH_SIZE] [--lr LR] [--lr_decay LR_DECAY] [--data_dir DATA_DIR] [--debug] [--save_dir SAVE_DIR] [--mix MIX]
1D Convolutional Network on Histone Modification Signals.
optional arguments:
-h, --help show this help message and exit
--epochs EPOCHS
--batch_size BATCH_SIZE
--lr LR Initial learning rate
--lr_decay LR_DECAY The value multiplied by lr at each epoch. Set a larger value for larger epochs
--data_dir DATA_DIR The directory of the data files
--debug Save weights by TensorBoard
--save_dir SAVE_DIR
--mix MIX Train on the data by mixing all cell types

To train and test with default values use

python main.py

Or if you need to provide different values of some parameters, you can use

python main.py --epochs 10 --lr 0.001

Once the training and testing is finsished the details about the training loss, validation loss and other evaluation metrics like auc are stored in the logs folder or in folders with names suggesting the parameters of the run. Those saved results could be used to plot training curves and evaluate the model.

Results

This model was the first generalized deep learning model for Gene Expression Prediction which even outperfomed the state of the art. This method was computationally very inexpensive as the model trained is pretty small and the number of training steps to be performed to achieve better than the state of the art performance is also small. Earlier methods used to train different models for different cell types but this approach only requires training the model once and then predictions can be made on various cell types.

result

AUC score vs Cell Types (comparison of our method with different baselines)

References

  • V. Chaubey, M. S. Nair and G. N. Pillai, "Gene Expression Prediction Using a Deep 1D Convolution Neural Network," 2019 IEEE Symposium Series on Computational Intelligence (SSCI), Xiamen, China, 2019, pp. 1383-1389, doi: 10.1109/SSCI44817.2019.9002669.
  • Ritambhara Singh, Jack Lanchantin, Arshdeep Sekhon, & Yanjun Qi. (2019). Dataset for DeepChrome and AttentiveChrome [Data set]. Zenodo. http://doi.org/10.5281/zenodo.2652278

About

Official repository for the paper "Gene Expression Prediction Using 1D Convolution Neural Network"

Topics

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Gene Expression Prediction Using 1D Convolution Network

This is the official repository for the paper Gene Expression Prediction Using a Deep 1D Convolution Neural Network. This repository contains the basic code to rreplicate our set of experiments.

There were two sets of experiments which we ran:

  • Each cell type is taken individually and a model is trained and tested upon it
  • The train and validation data is mixed and only one model is trained on this entire new data set and then testing is performed individually on each cell type

Data

The data for these experiments can be downloaded from here. The data contains gene expression values and the five kinds of histone modifications for 56 cell types. Each cell type contains 19802 gene samples. For each gene, the region containing 10,000 base pairs surrounding the Transcription Start Site (TSS) was divided into 100 bins of 100 base pairs each and for each bin, five different modification signals were recorded.

data

Histone modifications for one cell type

Model

The model created is a 1D convolutional neural network followed by a dense network. We use a 1D convolution network to map the spatial sequence of the histone modification signals. For more details please have a look at the paper.

model

The proposed network architecture

Usage

There are five python files in this repository

  • Convnet.py - This defines the model architecture along with the training and testing loops
  • data.py - This script loads and transforms the data in a format suitable for ingestion
  • utils.py - This describes some of the utility functions
  • roc_auc_callback.py - This defines a area under the ROC curve callback to keep track of how the model performance changes across various epochs
  • main.py - This is the main file which defines various hyperparameters and other configurations. It also runs a complete training and evaluation step for the data

Only the main.py is sufficient to train and test our model. There are various settings available in the main.py which can be accessed as

python main.py -h
usage: main.py [-h] [--epochs EPOCHS] [--batch_size BATCH_SIZE] [--lr LR] [--lr_decay LR_DECAY] [--data_dir DATA_DIR] [--debug] [--save_dir SAVE_DIR] [--mix MIX]
1D Convolutional Network on Histone Modification Signals.
optional arguments:
-h, --help show this help message and exit
--epochs EPOCHS
--batch_size BATCH_SIZE
--lr LR Initial learning rate
--lr_decay LR_DECAY The value multiplied by lr at each epoch. Set a larger value for larger epochs
--data_dir DATA_DIR The directory of the data files
--debug Save weights by TensorBoard
--save_dir SAVE_DIR
--mix MIX Train on the data by mixing all cell types

To train and test with default values use

python main.py

Or if you need to provide different values of some parameters, you can use

python main.py --epochs 10 --lr 0.001

Once the training and testing is finsished the details about the training loss, validation loss and other evaluation metrics like auc are stored in the logs folder or in folders with names suggesting the parameters of the run. Those saved results could be used to plot training curves and evaluate the model.

Results

This model was the first generalized deep learning model for Gene Expression Prediction which even outperfomed the state of the art. This method was computationally very inexpensive as the model trained is pretty small and the number of training steps to be performed to achieve better than the state of the art performance is also small. Earlier methods used to train different models for different cell types but this approach only requires training the model once and then predictions can be made on various cell types.

result

AUC score vs Cell Types (comparison of our method with different baselines)

References

  • V. Chaubey, M. S. Nair and G. N. Pillai, "Gene Expression Prediction Using a Deep 1D Convolution Neural Network," 2019 IEEE Symposium Series on Computational Intelligence (SSCI), Xiamen, China, 2019, pp. 1383-1389, doi: 10.1109/SSCI44817.2019.9002669.
  • Ritambhara Singh, Jack Lanchantin, Arshdeep Sekhon, & Yanjun Qi. (2019). Dataset for DeepChrome and AttentiveChrome [Data set]. Zenodo. http://doi.org/10.5281/zenodo.2652278

About

Official repository for the paper "Gene Expression Prediction Using 1D Convolution Neural Network"

Topics

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Gene Expression Prediction Using 1D Convolution Network

This is the official repository for the paper Gene Expression Prediction Using a Deep 1D Convolution Neural Network. This repository contains the basic code to rreplicate our set of experiments.

There were two sets of experiments which we ran:

  • Each cell type is taken individually and a model is trained and tested upon it
  • The train and validation data is mixed and only one model is trained on this entire new data set and then testing is performed individually on each cell type

Data

The data for these experiments can be downloaded from here. The data contains gene expression values and the five kinds of histone modifications for 56 cell types. Each cell type contains 19802 gene samples. For each gene, the region containing 10,000 base pairs surrounding the Transcription Start Site (TSS) was divided into 100 bins of 100 base pairs each and for each bin, five different modification signals were recorded.

data

Histone modifications for one cell type

Model

The model created is a 1D convolutional neural network followed by a dense network. We use a 1D convolution network to map the spatial sequence of the histone modification signals. For more details please have a look at the paper.

model

The proposed network architecture

Usage

There are five python files in this repository

  • Convnet.py - This defines the model architecture along with the training and testing loops
  • data.py - This script loads and transforms the data in a format suitable for ingestion
  • utils.py - This describes some of the utility functions
  • roc_auc_callback.py - This defines a area under the ROC curve callback to keep track of how the model performance changes across various epochs
  • main.py - This is the main file which defines various hyperparameters and other configurations. It also runs a complete training and evaluation step for the data

Only the main.py is sufficient to train and test our model. There are various settings available in the main.py which can be accessed as

python main.py -h
usage: main.py [-h] [--epochs EPOCHS] [--batch_size BATCH_SIZE] [--lr LR] [--lr_decay LR_DECAY] [--data_dir DATA_DIR] [--debug] [--save_dir SAVE_DIR] [--mix MIX]
1D Convolutional Network on Histone Modification Signals.
optional arguments:
-h, --help show this help message and exit
--epochs EPOCHS
--batch_size BATCH_SIZE
--lr LR Initial learning rate
--lr_decay LR_DECAY The value multiplied by lr at each epoch. Set a larger value for larger epochs
--data_dir DATA_DIR The directory of the data files
--debug Save weights by TensorBoard
--save_dir SAVE_DIR
--mix MIX Train on the data by mixing all cell types

To train and test with default values use

python main.py

Or if you need to provide different values of some parameters, you can use

python main.py --epochs 10 --lr 0.001

Once the training and testing is finsished the details about the training loss, validation loss and other evaluation metrics like auc are stored in the logs folder or in folders with names suggesting the parameters of the run. Those saved results could be used to plot training curves and evaluate the model.

Results

This model was the first generalized deep learning model for Gene Expression Prediction which even outperfomed the state of the art. This method was computationally very inexpensive as the model trained is pretty small and the number of training steps to be performed to achieve better than the state of the art performance is also small. Earlier methods used to train different models for different cell types but this approach only requires training the model once and then predictions can be made on various cell types.

result

AUC score vs Cell Types (comparison of our method with different baselines)

References

  • V. Chaubey, M. S. Nair and G. N. Pillai, "Gene Expression Prediction Using a Deep 1D Convolution Neural Network," 2019 IEEE Symposium Series on Computational Intelligence (SSCI), Xiamen, China, 2019, pp. 1383-1389, doi: 10.1109/SSCI44817.2019.9002669.
  • Ritambhara Singh, Jack Lanchantin, Arshdeep Sekhon, & Yanjun Qi. (2019). Dataset for DeepChrome and AttentiveChrome [Data set]. Zenodo. http://doi.org/10.5281/zenodo.2652278

About

Official repository for the paper "Gene Expression Prediction Using 1D Convolution Neural Network"

Topics

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Gene Expression Prediction Using 1D Convolution Network

This is the official repository for the paper Gene Expression Prediction Using a Deep 1D Convolution Neural Network. This repository contains the basic code to rreplicate our set of experiments.

There were two sets of experiments which we ran:

  • Each cell type is taken individually and a model is trained and tested upon it
  • The train and validation data is mixed and only one model is trained on this entire new data set and then testing is performed individually on each cell type

Data

The data for these experiments can be downloaded from here. The data contains gene expression values and the five kinds of histone modifications for 56 cell types. Each cell type contains 19802 gene samples. For each gene, the region containing 10,000 base pairs surrounding the Transcription Start Site (TSS) was divided into 100 bins of 100 base pairs each and for each bin, five different modification signals were recorded.

data

Histone modifications for one cell type

Model

The model created is a 1D convolutional neural network followed by a dense network. We use a 1D convolution network to map the spatial sequence of the histone modification signals. For more details please have a look at the paper.

model

The proposed network architecture

Usage

There are five python files in this repository

  • Convnet.py - This defines the model architecture along with the training and testing loops
  • data.py - This script loads and transforms the data in a format suitable for ingestion
  • utils.py - This describes some of the utility functions
  • roc_auc_callback.py - This defines a area under the ROC curve callback to keep track of how the model performance changes across various epochs
  • main.py - This is the main file which defines various hyperparameters and other configurations. It also runs a complete training and evaluation step for the data

Only the main.py is sufficient to train and test our model. There are various settings available in the main.py which can be accessed as

python main.py -h
usage: main.py [-h] [--epochs EPOCHS] [--batch_size BATCH_SIZE] [--lr LR] [--lr_decay LR_DECAY] [--data_dir DATA_DIR] [--debug] [--save_dir SAVE_DIR] [--mix MIX]
1D Convolutional Network on Histone Modification Signals.
optional arguments:
-h, --help show this help message and exit
--epochs EPOCHS
--batch_size BATCH_SIZE
--lr LR Initial learning rate
--lr_decay LR_DECAY The value multiplied by lr at each epoch. Set a larger value for larger epochs
--data_dir DATA_DIR The directory of the data files
--debug Save weights by TensorBoard
--save_dir SAVE_DIR
--mix MIX Train on the data by mixing all cell types

To train and test with default values use

python main.py

Or if you need to provide different values of some parameters, you can use

python main.py --epochs 10 --lr 0.001

Once the training and testing is finsished the details about the training loss, validation loss and other evaluation metrics like auc are stored in the logs folder or in folders with names suggesting the parameters of the run. Those saved results could be used to plot training curves and evaluate the model.

Results

This model was the first generalized deep learning model for Gene Expression Prediction which even outperfomed the state of the art. This method was computationally very inexpensive as the model trained is pretty small and the number of training steps to be performed to achieve better than the state of the art performance is also small. Earlier methods used to train different models for different cell types but this approach only requires training the model once and then predictions can be made on various cell types.

result

AUC score vs Cell Types (comparison of our method with different baselines)

References

  • V. Chaubey, M. S. Nair and G. N. Pillai, "Gene Expression Prediction Using a Deep 1D Convolution Neural Network," 2019 IEEE Symposium Series on Computational Intelligence (SSCI), Xiamen, China, 2019, pp. 1383-1389, doi: 10.1109/SSCI44817.2019.9002669.
  • Ritambhara Singh, Jack Lanchantin, Arshdeep Sekhon, & Yanjun Qi. (2019). Dataset for DeepChrome and AttentiveChrome [Data set]. Zenodo. http://doi.org/10.5281/zenodo.2652278

About

Official repository for the paper "Gene Expression Prediction Using 1D Convolution Neural Network"

Topics

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Gene Expression Prediction Using 1D Convolution Network

This is the official repository for the paper Gene Expression Prediction Using a Deep 1D Convolution Neural Network. This repository contains the basic code to rreplicate our set of experiments.

There were two sets of experiments which we ran:

  • Each cell type is taken individually and a model is trained and tested upon it
  • The train and validation data is mixed and only one model is trained on this entire new data set and then testing is performed individually on each cell type

Data

The data for these experiments can be downloaded from here. The data contains gene expression values and the five kinds of histone modifications for 56 cell types. Each cell type contains 19802 gene samples. For each gene, the region containing 10,000 base pairs surrounding the Transcription Start Site (TSS) was divided into 100 bins of 100 base pairs each and for each bin, five different modification signals were recorded.

data

Histone modifications for one cell type

Model

The model created is a 1D convolutional neural network followed by a dense network. We use a 1D convolution network to map the spatial sequence of the histone modification signals. For more details please have a look at the paper.

model

The proposed network architecture

Usage

There are five python files in this repository

  • Convnet.py - This defines the model architecture along with the training and testing loops
  • data.py - This script loads and transforms the data in a format suitable for ingestion
  • utils.py - This describes some of the utility functions
  • roc_auc_callback.py - This defines a area under the ROC curve callback to keep track of how the model performance changes across various epochs
  • main.py - This is the main file which defines various hyperparameters and other configurations. It also runs a complete training and evaluation step for the data

Only the main.py is sufficient to train and test our model. There are various settings available in the main.py which can be accessed as

python main.py -h
usage: main.py [-h] [--epochs EPOCHS] [--batch_size BATCH_SIZE] [--lr LR] [--lr_decay LR_DECAY] [--data_dir DATA_DIR] [--debug] [--save_dir SAVE_DIR] [--mix MIX]
1D Convolutional Network on Histone Modification Signals.
optional arguments:
-h, --help show this help message and exit
--epochs EPOCHS
--batch_size BATCH_SIZE
--lr LR Initial learning rate
--lr_decay LR_DECAY The value multiplied by lr at each epoch. Set a larger value for larger epochs
--data_dir DATA_DIR The directory of the data files
--debug Save weights by TensorBoard
--save_dir SAVE_DIR
--mix MIX Train on the data by mixing all cell types

To train and test with default values use

python main.py

Or if you need to provide different values of some parameters, you can use

python main.py --epochs 10 --lr 0.001

Once the training and testing is finsished the details about the training loss, validation loss and other evaluation metrics like auc are stored in the logs folder or in folders with names suggesting the parameters of the run. Those saved results could be used to plot training curves and evaluate the model.

Results

This model was the first generalized deep learning model for Gene Expression Prediction which even outperfomed the state of the art. This method was computationally very inexpensive as the model trained is pretty small and the number of training steps to be performed to achieve better than the state of the art performance is also small. Earlier methods used to train different models for different cell types but this approach only requires training the model once and then predictions can be made on various cell types.

result

AUC score vs Cell Types (comparison of our method with different baselines)

References

  • V. Chaubey, M. S. Nair and G. N. Pillai, "Gene Expression Prediction Using a Deep 1D Convolution Neural Network," 2019 IEEE Symposium Series on Computational Intelligence (SSCI), Xiamen, China, 2019, pp. 1383-1389, doi: 10.1109/SSCI44817.2019.9002669.
  • Ritambhara Singh, Jack Lanchantin, Arshdeep Sekhon, & Yanjun Qi. (2019). Dataset for DeepChrome and AttentiveChrome [Data set]. Zenodo. http://doi.org/10.5281/zenodo.2652278

About

Official repository for the paper "Gene Expression Prediction Using 1D Convolution Neural Network"

Topics

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Gene Expression Prediction Using 1D Convolution Network

This is the official repository for the paper Gene Expression Prediction Using a Deep 1D Convolution Neural Network. This repository contains the basic code to rreplicate our set of experiments.

There were two sets of experiments which we ran:

  • Each cell type is taken individually and a model is trained and tested upon it
  • The train and validation data is mixed and only one model is trained on this entire new data set and then testing is performed individually on each cell type

Data

The data for these experiments can be downloaded from here. The data contains gene expression values and the five kinds of histone modifications for 56 cell types. Each cell type contains 19802 gene samples. For each gene, the region containing 10,000 base pairs surrounding the Transcription Start Site (TSS) was divided into 100 bins of 100 base pairs each and for each bin, five different modification signals were recorded.

data

Histone modifications for one cell type

Model

The model created is a 1D convolutional neural network followed by a dense network. We use a 1D convolution network to map the spatial sequence of the histone modification signals. For more details please have a look at the paper.

model

The proposed network architecture

Usage

There are five python files in this repository

  • Convnet.py - This defines the model architecture along with the training and testing loops
  • data.py - This script loads and transforms the data in a format suitable for ingestion
  • utils.py - This describes some of the utility functions
  • roc_auc_callback.py - This defines a area under the ROC curve callback to keep track of how the model performance changes across various epochs
  • main.py - This is the main file which defines various hyperparameters and other configurations. It also runs a complete training and evaluation step for the data

Only the main.py is sufficient to train and test our model. There are various settings available in the main.py which can be accessed as

python main.py -h
usage: main.py [-h] [--epochs EPOCHS] [--batch_size BATCH_SIZE] [--lr LR] [--lr_decay LR_DECAY] [--data_dir DATA_DIR] [--debug] [--save_dir SAVE_DIR] [--mix MIX]
1D Convolutional Network on Histone Modification Signals.
optional arguments:
-h, --help show this help message and exit
--epochs EPOCHS
--batch_size BATCH_SIZE
--lr LR Initial learning rate
--lr_decay LR_DECAY The value multiplied by lr at each epoch. Set a larger value for larger epochs
--data_dir DATA_DIR The directory of the data files
--debug Save weights by TensorBoard
--save_dir SAVE_DIR
--mix MIX Train on the data by mixing all cell types

To train and test with default values use

python main.py

Or if you need to provide different values of some parameters, you can use

python main.py --epochs 10 --lr 0.001

Once the training and testing is finsished the details about the training loss, validation loss and other evaluation metrics like auc are stored in the logs folder or in folders with names suggesting the parameters of the run. Those saved results could be used to plot training curves and evaluate the model.

Results

This model was the first generalized deep learning model for Gene Expression Prediction which even outperfomed the state of the art. This method was computationally very inexpensive as the model trained is pretty small and the number of training steps to be performed to achieve better than the state of the art performance is also small. Earlier methods used to train different models for different cell types but this approach only requires training the model once and then predictions can be made on various cell types.

result

AUC score vs Cell Types (comparison of our method with different baselines)

References

  • V. Chaubey, M. S. Nair and G. N. Pillai, "Gene Expression Prediction Using a Deep 1D Convolution Neural Network," 2019 IEEE Symposium Series on Computational Intelligence (SSCI), Xiamen, China, 2019, pp. 1383-1389, doi: 10.1109/SSCI44817.2019.9002669.
  • Ritambhara Singh, Jack Lanchantin, Arshdeep Sekhon, & Yanjun Qi. (2019). Dataset for DeepChrome and AttentiveChrome [Data set]. Zenodo. http://doi.org/10.5281/zenodo.2652278

About

Official repository for the paper "Gene Expression Prediction Using 1D Convolution Neural Network"

Topics

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages