Skip to content

Latest commit

History

27 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

alt text

Meta-Signer

Meta-Signer is a machine learning aggregated approach for feature evaluation of metagenomic datasets. Random forest, support vector machines, logistic regression, and multi-layer neural networks. Features are then aggregated across models and partitions into a single ranked list of the top k features.

Execution:

We provide a python environment which can be imported using the Conda python package manager.

Deep learning models are built using Tensorflow. Meta-Signer was designed using Tensorflow v1.14.0.

To fully utilize GPUs for faster training of the deep learning models, users will need to be sure that both CUDA and cuDNN are properly installed.

Other dependencies should be downloaded upon importing the provided environment.

Clone Repository

git clone https://github.com/YDaiLab/Meta-Signer.git
cd Meta-Signer

Import Conda Environment

conda env create -f meta-signer.yml
source activate meta-signer

Meta-Signer's Required Input

To use Meta-Signer on a dataset, first create a directory in the data folder. This directory requires two files:

FileDescription
abundance.tsvA tab separated file where each row is a feature and each column is a sample. The first column should be the feature ID. There should be no header of sample IDs
labels.txtA text file where each row is the sample class value. Rows should be in the same order as columns found in abundance.tsv

Examples can be found in the PRISM and PRISM_3 datasets provided.

Set configuration settings

Meta-Signer offers a flexible framework which can be customized in the configuration file. The configuration file offers the following parameters:

Evaluation
NumberTestSplitsNumber of partitions for cross-validation
NumberRunsNumber of indepenendant iterations of cross-validation to run
NormalizationNormalization method applied to data (Standard or MinMax)
DataSetDirectory in data directory to load data from
FilterThreshCountRemove features who are present in fewer than the specified fraction of samples
FilterThreshMeanRemove features with a mean value less than the specified value
MaxKThe maximum number of features to generate in the rank aggregation
AggregateMethodThe method used for rank aggregation (GA or CE)
RF
TrainUse Random Forest for feature ranking and aggregation
NumberTreesNumber of decision trees per forest
ValidationModelsNumber of partitions for internal cross-validation for tuning
SVM
TrainUse SVM for feature ranking and aggregation
MaxIterationsMaximum number of iterations to train
GridCVNumber of partitions for internal cross-validation for tuning
Logistic Regression
TrainUse logistic regression for feature ranking and aggregation
MaxIterationsMaximum number of iterations to train
GridCVNumber of partitions for internal cross-validation for tuning
MLPNN
TrainUse MLPNN for feature ranking and aggregation
LearningRateLearning rate for neural network models
BatchSizeSize of each batch during neural network training
PatienceNumber of epochs to stop training after no improvement

Run the Meta-Signer pipeline:

Once the configuration is set to desired values, generate the aggregated feature list using:

cd src
python generate_feature_ranking.py

Upon completion, Meta-Signer will generate a directory in the results folder with the same name as set to the DataSet flag in the configuration file. This directory will contain important files of interest including:

FileDescription
training_performance.htmlA portable HTML file showing cross-validated evaluation of ML methods
feature_evaluation/ensemble_rank_table.csvranked lists of features for each method and each cross-validated run
feature_evaluation/aggregated_rank_table.csvAggregated ranked list of features
prediction_evaluation/results.tsvResults table for cross-validated evaluation of ML methods

Once the features have been aggregated into a single ranked list, the user can decide on how many features to use for the final training of ML models. Meta-Signer can generate these final trained ML models using a user specified number of features using:

cd src
python generate_models.py <DataSet><k>

Where DataSet is the directory in the results folder to use and k is the final number of features to use during training. Additionally, the models can be trained on an external datset using:

cd src
python generate_models.py <DataSet><k> -e <ExternalDataSet>

Where ExternalDataSet is a directory in the data folder with abundance.tsv and labels.txt files.

Upon completion, Meta-Signer will create a directory within the dataset's results directory that will contain:

FileDescription
feature_ranking.htmlA portable HTML file the ranked features up to the specified value of k
rf_model.pklThe trained random forest model in pickle format
logistic_regression_model.pklThe trained logistic regression model in pickle format
svm_model.pklThe trained SVM model in pickle format
mlpnn.h5The trained neural network model in H5 format
training_results.tsvThe performance of trained models on the training set
external_results.tsvThe performance of trained models on the external test set

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - YDaiLab/Meta-Signer · GitHub
Skip to content

Latest commit

History

27 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

alt text

Meta-Signer

Meta-Signer is a machine learning aggregated approach for feature evaluation of metagenomic datasets. Random forest, support vector machines, logistic regression, and multi-layer neural networks. Features are then aggregated across models and partitions into a single ranked list of the top k features.

Execution:

We provide a python environment which can be imported using the Conda python package manager.

Deep learning models are built using Tensorflow. Meta-Signer was designed using Tensorflow v1.14.0.

To fully utilize GPUs for faster training of the deep learning models, users will need to be sure that both CUDA and cuDNN are properly installed.

Other dependencies should be downloaded upon importing the provided environment.

Clone Repository

git clone https://github.com/YDaiLab/Meta-Signer.git
cd Meta-Signer

Import Conda Environment

conda env create -f meta-signer.yml
source activate meta-signer

Meta-Signer's Required Input

To use Meta-Signer on a dataset, first create a directory in the data folder. This directory requires two files:

FileDescription
abundance.tsvA tab separated file where each row is a feature and each column is a sample. The first column should be the feature ID. There should be no header of sample IDs
labels.txtA text file where each row is the sample class value. Rows should be in the same order as columns found in abundance.tsv

Examples can be found in the PRISM and PRISM_3 datasets provided.

Set configuration settings

Meta-Signer offers a flexible framework which can be customized in the configuration file. The configuration file offers the following parameters:

Evaluation
NumberTestSplitsNumber of partitions for cross-validation
NumberRunsNumber of indepenendant iterations of cross-validation to run
NormalizationNormalization method applied to data (Standard or MinMax)
DataSetDirectory in data directory to load data from
FilterThreshCountRemove features who are present in fewer than the specified fraction of samples
FilterThreshMeanRemove features with a mean value less than the specified value
MaxKThe maximum number of features to generate in the rank aggregation
AggregateMethodThe method used for rank aggregation (GA or CE)
RF
TrainUse Random Forest for feature ranking and aggregation
NumberTreesNumber of decision trees per forest
ValidationModelsNumber of partitions for internal cross-validation for tuning
SVM
TrainUse SVM for feature ranking and aggregation
MaxIterationsMaximum number of iterations to train
GridCVNumber of partitions for internal cross-validation for tuning
Logistic Regression
TrainUse logistic regression for feature ranking and aggregation
MaxIterationsMaximum number of iterations to train
GridCVNumber of partitions for internal cross-validation for tuning
MLPNN
TrainUse MLPNN for feature ranking and aggregation
LearningRateLearning rate for neural network models
BatchSizeSize of each batch during neural network training
PatienceNumber of epochs to stop training after no improvement

Run the Meta-Signer pipeline:

Once the configuration is set to desired values, generate the aggregated feature list using:

cd src
python generate_feature_ranking.py

Upon completion, Meta-Signer will generate a directory in the results folder with the same name as set to the DataSet flag in the configuration file. This directory will contain important files of interest including:

FileDescription
training_performance.htmlA portable HTML file showing cross-validated evaluation of ML methods
feature_evaluation/ensemble_rank_table.csvranked lists of features for each method and each cross-validated run
feature_evaluation/aggregated_rank_table.csvAggregated ranked list of features
prediction_evaluation/results.tsvResults table for cross-validated evaluation of ML methods

Once the features have been aggregated into a single ranked list, the user can decide on how many features to use for the final training of ML models. Meta-Signer can generate these final trained ML models using a user specified number of features using:

cd src
python generate_models.py <DataSet><k>

Where DataSet is the directory in the results folder to use and k is the final number of features to use during training. Additionally, the models can be trained on an external datset using:

cd src
python generate_models.py <DataSet><k> -e <ExternalDataSet>

Where ExternalDataSet is a directory in the data folder with abundance.tsv and labels.txt files.

Upon completion, Meta-Signer will create a directory within the dataset's results directory that will contain:

FileDescription
feature_ranking.htmlA portable HTML file the ranked features up to the specified value of k
rf_model.pklThe trained random forest model in pickle format
logistic_regression_model.pklThe trained logistic regression model in pickle format
svm_model.pklThe trained SVM model in pickle format
mlpnn.h5The trained neural network model in H5 format
training_results.tsvThe performance of trained models on the training set
external_results.tsvThe performance of trained models on the external test set

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - YDaiLab/Meta-Signer · GitHub
Skip to content

Latest commit

History

27 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

alt text

Meta-Signer

Meta-Signer is a machine learning aggregated approach for feature evaluation of metagenomic datasets. Random forest, support vector machines, logistic regression, and multi-layer neural networks. Features are then aggregated across models and partitions into a single ranked list of the top k features.

Execution:

We provide a python environment which can be imported using the Conda python package manager.

Deep learning models are built using Tensorflow. Meta-Signer was designed using Tensorflow v1.14.0.

To fully utilize GPUs for faster training of the deep learning models, users will need to be sure that both CUDA and cuDNN are properly installed.

Other dependencies should be downloaded upon importing the provided environment.

Clone Repository

git clone https://github.com/YDaiLab/Meta-Signer.git
cd Meta-Signer

Import Conda Environment

conda env create -f meta-signer.yml
source activate meta-signer

Meta-Signer's Required Input

To use Meta-Signer on a dataset, first create a directory in the data folder. This directory requires two files:

FileDescription
abundance.tsvA tab separated file where each row is a feature and each column is a sample. The first column should be the feature ID. There should be no header of sample IDs
labels.txtA text file where each row is the sample class value. Rows should be in the same order as columns found in abundance.tsv

Examples can be found in the PRISM and PRISM_3 datasets provided.

Set configuration settings

Meta-Signer offers a flexible framework which can be customized in the configuration file. The configuration file offers the following parameters:

Evaluation
NumberTestSplitsNumber of partitions for cross-validation
NumberRunsNumber of indepenendant iterations of cross-validation to run
NormalizationNormalization method applied to data (Standard or MinMax)
DataSetDirectory in data directory to load data from
FilterThreshCountRemove features who are present in fewer than the specified fraction of samples
FilterThreshMeanRemove features with a mean value less than the specified value
MaxKThe maximum number of features to generate in the rank aggregation
AggregateMethodThe method used for rank aggregation (GA or CE)
RF
TrainUse Random Forest for feature ranking and aggregation
NumberTreesNumber of decision trees per forest
ValidationModelsNumber of partitions for internal cross-validation for tuning
SVM
TrainUse SVM for feature ranking and aggregation
MaxIterationsMaximum number of iterations to train
GridCVNumber of partitions for internal cross-validation for tuning
Logistic Regression
TrainUse logistic regression for feature ranking and aggregation
MaxIterationsMaximum number of iterations to train
GridCVNumber of partitions for internal cross-validation for tuning
MLPNN
TrainUse MLPNN for feature ranking and aggregation
LearningRateLearning rate for neural network models
BatchSizeSize of each batch during neural network training
PatienceNumber of epochs to stop training after no improvement

Run the Meta-Signer pipeline:

Once the configuration is set to desired values, generate the aggregated feature list using:

cd src
python generate_feature_ranking.py

Upon completion, Meta-Signer will generate a directory in the results folder with the same name as set to the DataSet flag in the configuration file. This directory will contain important files of interest including:

FileDescription
training_performance.htmlA portable HTML file showing cross-validated evaluation of ML methods
feature_evaluation/ensemble_rank_table.csvranked lists of features for each method and each cross-validated run
feature_evaluation/aggregated_rank_table.csvAggregated ranked list of features
prediction_evaluation/results.tsvResults table for cross-validated evaluation of ML methods

Once the features have been aggregated into a single ranked list, the user can decide on how many features to use for the final training of ML models. Meta-Signer can generate these final trained ML models using a user specified number of features using:

cd src
python generate_models.py <DataSet><k>

Where DataSet is the directory in the results folder to use and k is the final number of features to use during training. Additionally, the models can be trained on an external datset using:

cd src
python generate_models.py <DataSet><k> -e <ExternalDataSet>

Where ExternalDataSet is a directory in the data folder with abundance.tsv and labels.txt files.

Upon completion, Meta-Signer will create a directory within the dataset's results directory that will contain:

FileDescription
feature_ranking.htmlA portable HTML file the ranked features up to the specified value of k
rf_model.pklThe trained random forest model in pickle format
logistic_regression_model.pklThe trained logistic regression model in pickle format
svm_model.pklThe trained SVM model in pickle format
mlpnn.h5The trained neural network model in H5 format
training_results.tsvThe performance of trained models on the training set
external_results.tsvThe performance of trained models on the external test set

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - YDaiLab/Meta-Signer · GitHub
Skip to content

Latest commit

History

27 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

alt text

Meta-Signer

Meta-Signer is a machine learning aggregated approach for feature evaluation of metagenomic datasets. Random forest, support vector machines, logistic regression, and multi-layer neural networks. Features are then aggregated across models and partitions into a single ranked list of the top k features.

Execution:

We provide a python environment which can be imported using the Conda python package manager.

Deep learning models are built using Tensorflow. Meta-Signer was designed using Tensorflow v1.14.0.

To fully utilize GPUs for faster training of the deep learning models, users will need to be sure that both CUDA and cuDNN are properly installed.

Other dependencies should be downloaded upon importing the provided environment.

Clone Repository

git clone https://github.com/YDaiLab/Meta-Signer.git
cd Meta-Signer

Import Conda Environment

conda env create -f meta-signer.yml
source activate meta-signer

Meta-Signer's Required Input

To use Meta-Signer on a dataset, first create a directory in the data folder. This directory requires two files:

FileDescription
abundance.tsvA tab separated file where each row is a feature and each column is a sample. The first column should be the feature ID. There should be no header of sample IDs
labels.txtA text file where each row is the sample class value. Rows should be in the same order as columns found in abundance.tsv

Examples can be found in the PRISM and PRISM_3 datasets provided.

Set configuration settings

Meta-Signer offers a flexible framework which can be customized in the configuration file. The configuration file offers the following parameters:

Evaluation
NumberTestSplitsNumber of partitions for cross-validation
NumberRunsNumber of indepenendant iterations of cross-validation to run
NormalizationNormalization method applied to data (Standard or MinMax)
DataSetDirectory in data directory to load data from
FilterThreshCountRemove features who are present in fewer than the specified fraction of samples
FilterThreshMeanRemove features with a mean value less than the specified value
MaxKThe maximum number of features to generate in the rank aggregation
AggregateMethodThe method used for rank aggregation (GA or CE)
RF
TrainUse Random Forest for feature ranking and aggregation
NumberTreesNumber of decision trees per forest
ValidationModelsNumber of partitions for internal cross-validation for tuning
SVM
TrainUse SVM for feature ranking and aggregation
MaxIterationsMaximum number of iterations to train
GridCVNumber of partitions for internal cross-validation for tuning
Logistic Regression
TrainUse logistic regression for feature ranking and aggregation
MaxIterationsMaximum number of iterations to train
GridCVNumber of partitions for internal cross-validation for tuning
MLPNN
TrainUse MLPNN for feature ranking and aggregation
LearningRateLearning rate for neural network models
BatchSizeSize of each batch during neural network training
PatienceNumber of epochs to stop training after no improvement

Run the Meta-Signer pipeline:

Once the configuration is set to desired values, generate the aggregated feature list using:

cd src
python generate_feature_ranking.py

Upon completion, Meta-Signer will generate a directory in the results folder with the same name as set to the DataSet flag in the configuration file. This directory will contain important files of interest including:

FileDescription
training_performance.htmlA portable HTML file showing cross-validated evaluation of ML methods
feature_evaluation/ensemble_rank_table.csvranked lists of features for each method and each cross-validated run
feature_evaluation/aggregated_rank_table.csvAggregated ranked list of features
prediction_evaluation/results.tsvResults table for cross-validated evaluation of ML methods

Once the features have been aggregated into a single ranked list, the user can decide on how many features to use for the final training of ML models. Meta-Signer can generate these final trained ML models using a user specified number of features using:

cd src
python generate_models.py <DataSet><k>

Where DataSet is the directory in the results folder to use and k is the final number of features to use during training. Additionally, the models can be trained on an external datset using:

cd src
python generate_models.py <DataSet><k> -e <ExternalDataSet>

Where ExternalDataSet is a directory in the data folder with abundance.tsv and labels.txt files.

Upon completion, Meta-Signer will create a directory within the dataset's results directory that will contain:

FileDescription
feature_ranking.htmlA portable HTML file the ranked features up to the specified value of k
rf_model.pklThe trained random forest model in pickle format
logistic_regression_model.pklThe trained logistic regression model in pickle format
svm_model.pklThe trained SVM model in pickle format
mlpnn.h5The trained neural network model in H5 format
training_results.tsvThe performance of trained models on the training set
external_results.tsvThe performance of trained models on the external test set

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - YDaiLab/Meta-Signer · GitHub
Skip to content

Latest commit

History

27 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

alt text

Meta-Signer

Meta-Signer is a machine learning aggregated approach for feature evaluation of metagenomic datasets. Random forest, support vector machines, logistic regression, and multi-layer neural networks. Features are then aggregated across models and partitions into a single ranked list of the top k features.

Execution:

We provide a python environment which can be imported using the Conda python package manager.

Deep learning models are built using Tensorflow. Meta-Signer was designed using Tensorflow v1.14.0.

To fully utilize GPUs for faster training of the deep learning models, users will need to be sure that both CUDA and cuDNN are properly installed.

Other dependencies should be downloaded upon importing the provided environment.

Clone Repository

git clone https://github.com/YDaiLab/Meta-Signer.git
cd Meta-Signer

Import Conda Environment

conda env create -f meta-signer.yml
source activate meta-signer

Meta-Signer's Required Input

To use Meta-Signer on a dataset, first create a directory in the data folder. This directory requires two files:

FileDescription
abundance.tsvA tab separated file where each row is a feature and each column is a sample. The first column should be the feature ID. There should be no header of sample IDs
labels.txtA text file where each row is the sample class value. Rows should be in the same order as columns found in abundance.tsv

Examples can be found in the PRISM and PRISM_3 datasets provided.

Set configuration settings

Meta-Signer offers a flexible framework which can be customized in the configuration file. The configuration file offers the following parameters:

Evaluation
NumberTestSplitsNumber of partitions for cross-validation
NumberRunsNumber of indepenendant iterations of cross-validation to run
NormalizationNormalization method applied to data (Standard or MinMax)
DataSetDirectory in data directory to load data from
FilterThreshCountRemove features who are present in fewer than the specified fraction of samples
FilterThreshMeanRemove features with a mean value less than the specified value
MaxKThe maximum number of features to generate in the rank aggregation
AggregateMethodThe method used for rank aggregation (GA or CE)
RF
TrainUse Random Forest for feature ranking and aggregation
NumberTreesNumber of decision trees per forest
ValidationModelsNumber of partitions for internal cross-validation for tuning
SVM
TrainUse SVM for feature ranking and aggregation
MaxIterationsMaximum number of iterations to train
GridCVNumber of partitions for internal cross-validation for tuning
Logistic Regression
TrainUse logistic regression for feature ranking and aggregation
MaxIterationsMaximum number of iterations to train
GridCVNumber of partitions for internal cross-validation for tuning
MLPNN
TrainUse MLPNN for feature ranking and aggregation
LearningRateLearning rate for neural network models
BatchSizeSize of each batch during neural network training
PatienceNumber of epochs to stop training after no improvement

Run the Meta-Signer pipeline:

Once the configuration is set to desired values, generate the aggregated feature list using:

cd src
python generate_feature_ranking.py

Upon completion, Meta-Signer will generate a directory in the results folder with the same name as set to the DataSet flag in the configuration file. This directory will contain important files of interest including:

FileDescription
training_performance.htmlA portable HTML file showing cross-validated evaluation of ML methods
feature_evaluation/ensemble_rank_table.csvranked lists of features for each method and each cross-validated run
feature_evaluation/aggregated_rank_table.csvAggregated ranked list of features
prediction_evaluation/results.tsvResults table for cross-validated evaluation of ML methods

Once the features have been aggregated into a single ranked list, the user can decide on how many features to use for the final training of ML models. Meta-Signer can generate these final trained ML models using a user specified number of features using:

cd src
python generate_models.py <DataSet><k>

Where DataSet is the directory in the results folder to use and k is the final number of features to use during training. Additionally, the models can be trained on an external datset using:

cd src
python generate_models.py <DataSet><k> -e <ExternalDataSet>

Where ExternalDataSet is a directory in the data folder with abundance.tsv and labels.txt files.

Upon completion, Meta-Signer will create a directory within the dataset's results directory that will contain:

FileDescription
feature_ranking.htmlA portable HTML file the ranked features up to the specified value of k
rf_model.pklThe trained random forest model in pickle format
logistic_regression_model.pklThe trained logistic regression model in pickle format
svm_model.pklThe trained SVM model in pickle format
mlpnn.h5The trained neural network model in H5 format
training_results.tsvThe performance of trained models on the training set
external_results.tsvThe performance of trained models on the external test set

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - YDaiLab/Meta-Signer · GitHub
Skip to content

Latest commit

History

27 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

alt text

Meta-Signer

Meta-Signer is a machine learning aggregated approach for feature evaluation of metagenomic datasets. Random forest, support vector machines, logistic regression, and multi-layer neural networks. Features are then aggregated across models and partitions into a single ranked list of the top k features.

Execution:

We provide a python environment which can be imported using the Conda python package manager.

Deep learning models are built using Tensorflow. Meta-Signer was designed using Tensorflow v1.14.0.

To fully utilize GPUs for faster training of the deep learning models, users will need to be sure that both CUDA and cuDNN are properly installed.

Other dependencies should be downloaded upon importing the provided environment.

Clone Repository

git clone https://github.com/YDaiLab/Meta-Signer.git
cd Meta-Signer

Import Conda Environment

conda env create -f meta-signer.yml
source activate meta-signer

Meta-Signer's Required Input

To use Meta-Signer on a dataset, first create a directory in the data folder. This directory requires two files:

FileDescription
abundance.tsvA tab separated file where each row is a feature and each column is a sample. The first column should be the feature ID. There should be no header of sample IDs
labels.txtA text file where each row is the sample class value. Rows should be in the same order as columns found in abundance.tsv

Examples can be found in the PRISM and PRISM_3 datasets provided.

Set configuration settings

Meta-Signer offers a flexible framework which can be customized in the configuration file. The configuration file offers the following parameters:

Evaluation
NumberTestSplitsNumber of partitions for cross-validation
NumberRunsNumber of indepenendant iterations of cross-validation to run
NormalizationNormalization method applied to data (Standard or MinMax)
DataSetDirectory in data directory to load data from
FilterThreshCountRemove features who are present in fewer than the specified fraction of samples
FilterThreshMeanRemove features with a mean value less than the specified value
MaxKThe maximum number of features to generate in the rank aggregation
AggregateMethodThe method used for rank aggregation (GA or CE)
RF
TrainUse Random Forest for feature ranking and aggregation
NumberTreesNumber of decision trees per forest
ValidationModelsNumber of partitions for internal cross-validation for tuning
SVM
TrainUse SVM for feature ranking and aggregation
MaxIterationsMaximum number of iterations to train
GridCVNumber of partitions for internal cross-validation for tuning
Logistic Regression
TrainUse logistic regression for feature ranking and aggregation
MaxIterationsMaximum number of iterations to train
GridCVNumber of partitions for internal cross-validation for tuning
MLPNN
TrainUse MLPNN for feature ranking and aggregation
LearningRateLearning rate for neural network models
BatchSizeSize of each batch during neural network training
PatienceNumber of epochs to stop training after no improvement

Run the Meta-Signer pipeline:

Once the configuration is set to desired values, generate the aggregated feature list using:

cd src
python generate_feature_ranking.py

Upon completion, Meta-Signer will generate a directory in the results folder with the same name as set to the DataSet flag in the configuration file. This directory will contain important files of interest including:

FileDescription
training_performance.htmlA portable HTML file showing cross-validated evaluation of ML methods
feature_evaluation/ensemble_rank_table.csvranked lists of features for each method and each cross-validated run
feature_evaluation/aggregated_rank_table.csvAggregated ranked list of features
prediction_evaluation/results.tsvResults table for cross-validated evaluation of ML methods

Once the features have been aggregated into a single ranked list, the user can decide on how many features to use for the final training of ML models. Meta-Signer can generate these final trained ML models using a user specified number of features using:

cd src
python generate_models.py <DataSet><k>

Where DataSet is the directory in the results folder to use and k is the final number of features to use during training. Additionally, the models can be trained on an external datset using:

cd src
python generate_models.py <DataSet><k> -e <ExternalDataSet>

Where ExternalDataSet is a directory in the data folder with abundance.tsv and labels.txt files.

Upon completion, Meta-Signer will create a directory within the dataset's results directory that will contain:

FileDescription
feature_ranking.htmlA portable HTML file the ranked features up to the specified value of k
rf_model.pklThe trained random forest model in pickle format
logistic_regression_model.pklThe trained logistic regression model in pickle format
svm_model.pklThe trained SVM model in pickle format
mlpnn.h5The trained neural network model in H5 format
training_results.tsvThe performance of trained models on the training set
external_results.tsvThe performance of trained models on the external test set

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - YDaiLab/Meta-Signer · GitHub
Skip to content

Latest commit

History

27 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

alt text

Meta-Signer

Meta-Signer is a machine learning aggregated approach for feature evaluation of metagenomic datasets. Random forest, support vector machines, logistic regression, and multi-layer neural networks. Features are then aggregated across models and partitions into a single ranked list of the top k features.

Execution:

We provide a python environment which can be imported using the Conda python package manager.

Deep learning models are built using Tensorflow. Meta-Signer was designed using Tensorflow v1.14.0.

To fully utilize GPUs for faster training of the deep learning models, users will need to be sure that both CUDA and cuDNN are properly installed.

Other dependencies should be downloaded upon importing the provided environment.

Clone Repository

git clone https://github.com/YDaiLab/Meta-Signer.git
cd Meta-Signer

Import Conda Environment

conda env create -f meta-signer.yml
source activate meta-signer

Meta-Signer's Required Input

To use Meta-Signer on a dataset, first create a directory in the data folder. This directory requires two files:

FileDescription
abundance.tsvA tab separated file where each row is a feature and each column is a sample. The first column should be the feature ID. There should be no header of sample IDs
labels.txtA text file where each row is the sample class value. Rows should be in the same order as columns found in abundance.tsv

Examples can be found in the PRISM and PRISM_3 datasets provided.

Set configuration settings

Meta-Signer offers a flexible framework which can be customized in the configuration file. The configuration file offers the following parameters:

Evaluation
NumberTestSplitsNumber of partitions for cross-validation
NumberRunsNumber of indepenendant iterations of cross-validation to run
NormalizationNormalization method applied to data (Standard or MinMax)
DataSetDirectory in data directory to load data from
FilterThreshCountRemove features who are present in fewer than the specified fraction of samples
FilterThreshMeanRemove features with a mean value less than the specified value
MaxKThe maximum number of features to generate in the rank aggregation
AggregateMethodThe method used for rank aggregation (GA or CE)
RF
TrainUse Random Forest for feature ranking and aggregation
NumberTreesNumber of decision trees per forest
ValidationModelsNumber of partitions for internal cross-validation for tuning
SVM
TrainUse SVM for feature ranking and aggregation
MaxIterationsMaximum number of iterations to train
GridCVNumber of partitions for internal cross-validation for tuning
Logistic Regression
TrainUse logistic regression for feature ranking and aggregation
MaxIterationsMaximum number of iterations to train
GridCVNumber of partitions for internal cross-validation for tuning
MLPNN
TrainUse MLPNN for feature ranking and aggregation
LearningRateLearning rate for neural network models
BatchSizeSize of each batch during neural network training
PatienceNumber of epochs to stop training after no improvement

Run the Meta-Signer pipeline:

Once the configuration is set to desired values, generate the aggregated feature list using:

cd src
python generate_feature_ranking.py

Upon completion, Meta-Signer will generate a directory in the results folder with the same name as set to the DataSet flag in the configuration file. This directory will contain important files of interest including:

FileDescription
training_performance.htmlA portable HTML file showing cross-validated evaluation of ML methods
feature_evaluation/ensemble_rank_table.csvranked lists of features for each method and each cross-validated run
feature_evaluation/aggregated_rank_table.csvAggregated ranked list of features
prediction_evaluation/results.tsvResults table for cross-validated evaluation of ML methods

Once the features have been aggregated into a single ranked list, the user can decide on how many features to use for the final training of ML models. Meta-Signer can generate these final trained ML models using a user specified number of features using:

cd src
python generate_models.py <DataSet><k>

Where DataSet is the directory in the results folder to use and k is the final number of features to use during training. Additionally, the models can be trained on an external datset using:

cd src
python generate_models.py <DataSet><k> -e <ExternalDataSet>

Where ExternalDataSet is a directory in the data folder with abundance.tsv and labels.txt files.

Upon completion, Meta-Signer will create a directory within the dataset's results directory that will contain:

FileDescription
feature_ranking.htmlA portable HTML file the ranked features up to the specified value of k
rf_model.pklThe trained random forest model in pickle format
logistic_regression_model.pklThe trained logistic regression model in pickle format
svm_model.pklThe trained SVM model in pickle format
mlpnn.h5The trained neural network model in H5 format
training_results.tsvThe performance of trained models on the training set
external_results.tsvThe performance of trained models on the external test set

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - YDaiLab/Meta-Signer · GitHub
Skip to content

Latest commit

History

27 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

alt text

Meta-Signer

Meta-Signer is a machine learning aggregated approach for feature evaluation of metagenomic datasets. Random forest, support vector machines, logistic regression, and multi-layer neural networks. Features are then aggregated across models and partitions into a single ranked list of the top k features.

Execution:

We provide a python environment which can be imported using the Conda python package manager.

Deep learning models are built using Tensorflow. Meta-Signer was designed using Tensorflow v1.14.0.

To fully utilize GPUs for faster training of the deep learning models, users will need to be sure that both CUDA and cuDNN are properly installed.

Other dependencies should be downloaded upon importing the provided environment.

Clone Repository

git clone https://github.com/YDaiLab/Meta-Signer.git
cd Meta-Signer

Import Conda Environment

conda env create -f meta-signer.yml
source activate meta-signer

Meta-Signer's Required Input

To use Meta-Signer on a dataset, first create a directory in the data folder. This directory requires two files:

FileDescription
abundance.tsvA tab separated file where each row is a feature and each column is a sample. The first column should be the feature ID. There should be no header of sample IDs
labels.txtA text file where each row is the sample class value. Rows should be in the same order as columns found in abundance.tsv

Examples can be found in the PRISM and PRISM_3 datasets provided.

Set configuration settings

Meta-Signer offers a flexible framework which can be customized in the configuration file. The configuration file offers the following parameters:

Evaluation
NumberTestSplitsNumber of partitions for cross-validation
NumberRunsNumber of indepenendant iterations of cross-validation to run
NormalizationNormalization method applied to data (Standard or MinMax)
DataSetDirectory in data directory to load data from
FilterThreshCountRemove features who are present in fewer than the specified fraction of samples
FilterThreshMeanRemove features with a mean value less than the specified value
MaxKThe maximum number of features to generate in the rank aggregation
AggregateMethodThe method used for rank aggregation (GA or CE)
RF
TrainUse Random Forest for feature ranking and aggregation
NumberTreesNumber of decision trees per forest
ValidationModelsNumber of partitions for internal cross-validation for tuning
SVM
TrainUse SVM for feature ranking and aggregation
MaxIterationsMaximum number of iterations to train
GridCVNumber of partitions for internal cross-validation for tuning
Logistic Regression
TrainUse logistic regression for feature ranking and aggregation
MaxIterationsMaximum number of iterations to train
GridCVNumber of partitions for internal cross-validation for tuning
MLPNN
TrainUse MLPNN for feature ranking and aggregation
LearningRateLearning rate for neural network models
BatchSizeSize of each batch during neural network training
PatienceNumber of epochs to stop training after no improvement

Run the Meta-Signer pipeline:

Once the configuration is set to desired values, generate the aggregated feature list using:

cd src
python generate_feature_ranking.py

Upon completion, Meta-Signer will generate a directory in the results folder with the same name as set to the DataSet flag in the configuration file. This directory will contain important files of interest including:

FileDescription
training_performance.htmlA portable HTML file showing cross-validated evaluation of ML methods
feature_evaluation/ensemble_rank_table.csvranked lists of features for each method and each cross-validated run
feature_evaluation/aggregated_rank_table.csvAggregated ranked list of features
prediction_evaluation/results.tsvResults table for cross-validated evaluation of ML methods

Once the features have been aggregated into a single ranked list, the user can decide on how many features to use for the final training of ML models. Meta-Signer can generate these final trained ML models using a user specified number of features using:

cd src
python generate_models.py <DataSet><k>

Where DataSet is the directory in the results folder to use and k is the final number of features to use during training. Additionally, the models can be trained on an external datset using:

cd src
python generate_models.py <DataSet><k> -e <ExternalDataSet>

Where ExternalDataSet is a directory in the data folder with abundance.tsv and labels.txt files.

Upon completion, Meta-Signer will create a directory within the dataset's results directory that will contain:

FileDescription
feature_ranking.htmlA portable HTML file the ranked features up to the specified value of k
rf_model.pklThe trained random forest model in pickle format
logistic_regression_model.pklThe trained logistic regression model in pickle format
svm_model.pklThe trained SVM model in pickle format
mlpnn.h5The trained neural network model in H5 format
training_results.tsvThe performance of trained models on the training set
external_results.tsvThe performance of trained models on the external test set

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages