Skip to content

Latest commit

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

AMPHORA
AMPHORA is an Automated Phylogenomic Inference Pipeline for bacterial sequences. From a given a set of protein sequences, it automatically identifies 31 phylogenetic marker genes. It then generates high-quality multiple sequence alignments for these genes and make tree-based phylotype assignments.
CITATION
==============================================================
Please cite AMPHORA as: Martin Wu and Jonathan A Eisen. A simple, fast, and accurate method of phylogenomic inference Genome Biology 2008, 9:R151
SYSTEM REQUIREMENT
============================================================== Linux OS (kernel version 2.6 or later)
The following software is required by the AMPHORA package. They need to be downloaded and installed separately from AMPHORA. 1. Perl 5.8.8 or later (www.perl.org) 2. Bioperl core package 1.5.2 or later (www.bioperl.org) 3. HMMER (hmmer.janelia.org) 4. WU BLAST (blast.wustl.edu)
The following software is included in the AMPHORA distribution. Their source codes have been slightly modified to suit the needs of AMPHORA. 1. seqboot (evolution.genetics.washington.edu/phylip) 2. quicktree (www.sanger.ac.uk/Software/analysis/quicktree) 3. raxml (icwww.epfl.ch/~stamatak/index-Dateien/Page443.htm)
LICENSE
==============================================================
AMPHORA Copyright 2008 by Martin Wu
AMPHORA is free software: you may redistribute it and/or modify its under the terms of the GNU General Public License as published by the Free Software Foundation; either version 2 of the License, or any later version.
AMPHORA is distributed in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU General Public License for more details (http://www.gnu.org/licenses/).
INSTALLATION
============================================================== 1. Unpack the package tar -xzvf AMPHORA.tar.gz
2. Install AMPHORA to a user specified directory. Make sure you have write permission to the directory.
cd AMPHORA perl INSTALL.pl -AMPHORA_home user-specified-directory -Bioperl_home path-to-bioperl
PACKAGE CONTENTS
============================================================== After successful installation, there should be several folders in the home directory of AMPHORA
1. Marker It contains curated seed multiple sequence alignments of the phylogenetic markers in GDE format (with embedded masks) and associated Hidden Markov Models. Currently there are 31 protein marker genes. They are dnaG, frr, infC, nusA, pgk, pyrG, rplA, rplB, rplC, rplD, rplE, rplF, rplK, rplL, rplM, rplN, rplP, rplS, rplT, rpmA, rpoB, rpsB, rpsC, rpsE, rpsI, rpsJ, rpsK, rpsM, rpsS, smpB, tsf.
2. Reference It contains protein sequences of the marker genes from all complete bacterial genomes. It also contains a bacterial genome tree that was made from the concatenated protein sequences of all the marker genes. Reference trees for each marker gene were derived from the genome tree by replacing the species names with their corresponding gene names and keeping the topology intact.
3. Taxonomy The NCBI taxonomy database (ftp://ftp.ncbi.nih.gov/pub/taxonomy/) with minor modifications. The changes are listed in the file change.note
4. Scripts Perl scripts for identifying markers, generating trimmed multiple sequence alignments, and assigning phylotypes based on phylogenetic inferences.
5. bin Helper programs
USING AMPHORA
==============================================================
1. Phylotying bacterial sequences
Usage: AMPHORA_home/Scripts/Phylotyping.pl <options> protein-sequence-file output-file
Options:
-Replicates: number of bootstrap replicates
-BootstrapCutoff: normalized to 100 replicates (1-100)% default 70
Output:
foo.pep (identified maker sequences in fasta format, i.e. rpoB.pep)
foo.aln (aligned and trimmed marker sequences, i.e., rpoB.aln)
Trees/foo.tree (generated phylogenetic trees)
2. Identify marker sequences
Usage: perl AMPHORA_home/Scripts/MarkerScanner.pl protein-sequence-file
Output: foo.pep (identified maker sequences in fasta format, i.e., rpoB.pep)
3. Align and trim the marker sequences
Usage: perl AMPHORA_home/Scripts/MarkerAlignTrim.pl
Options: -Partial: query sequences contain partial genes
-Trim: trim the alignment using masks embedded within the marker database
-Strict: use a conservative mask
-Directory: the file directory where marker sequences are located. Default: current directory
-Help: print the help message
Output:
foo.aln (aligned and trimmed marker sequences, i.e., rpoB.aln)

About

Automated Phylogenomic Inference Pipeline

Resources

Stars

5 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - koadman/AMPHORA: Automated Phylogenomic Inference Pipeline · GitHub
Skip to content

Latest commit

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

AMPHORA
AMPHORA is an Automated Phylogenomic Inference Pipeline for bacterial sequences. From a given a set of protein sequences, it automatically identifies 31 phylogenetic marker genes. It then generates high-quality multiple sequence alignments for these genes and make tree-based phylotype assignments.
CITATION
==============================================================
Please cite AMPHORA as: Martin Wu and Jonathan A Eisen. A simple, fast, and accurate method of phylogenomic inference Genome Biology 2008, 9:R151
SYSTEM REQUIREMENT
============================================================== Linux OS (kernel version 2.6 or later)
The following software is required by the AMPHORA package. They need to be downloaded and installed separately from AMPHORA. 1. Perl 5.8.8 or later (www.perl.org) 2. Bioperl core package 1.5.2 or later (www.bioperl.org) 3. HMMER (hmmer.janelia.org) 4. WU BLAST (blast.wustl.edu)
The following software is included in the AMPHORA distribution. Their source codes have been slightly modified to suit the needs of AMPHORA. 1. seqboot (evolution.genetics.washington.edu/phylip) 2. quicktree (www.sanger.ac.uk/Software/analysis/quicktree) 3. raxml (icwww.epfl.ch/~stamatak/index-Dateien/Page443.htm)
LICENSE
==============================================================
AMPHORA Copyright 2008 by Martin Wu
AMPHORA is free software: you may redistribute it and/or modify its under the terms of the GNU General Public License as published by the Free Software Foundation; either version 2 of the License, or any later version.
AMPHORA is distributed in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU General Public License for more details (http://www.gnu.org/licenses/).
INSTALLATION
============================================================== 1. Unpack the package tar -xzvf AMPHORA.tar.gz
2. Install AMPHORA to a user specified directory. Make sure you have write permission to the directory.
cd AMPHORA perl INSTALL.pl -AMPHORA_home user-specified-directory -Bioperl_home path-to-bioperl
PACKAGE CONTENTS
============================================================== After successful installation, there should be several folders in the home directory of AMPHORA
1. Marker It contains curated seed multiple sequence alignments of the phylogenetic markers in GDE format (with embedded masks) and associated Hidden Markov Models. Currently there are 31 protein marker genes. They are dnaG, frr, infC, nusA, pgk, pyrG, rplA, rplB, rplC, rplD, rplE, rplF, rplK, rplL, rplM, rplN, rplP, rplS, rplT, rpmA, rpoB, rpsB, rpsC, rpsE, rpsI, rpsJ, rpsK, rpsM, rpsS, smpB, tsf.
2. Reference It contains protein sequences of the marker genes from all complete bacterial genomes. It also contains a bacterial genome tree that was made from the concatenated protein sequences of all the marker genes. Reference trees for each marker gene were derived from the genome tree by replacing the species names with their corresponding gene names and keeping the topology intact.
3. Taxonomy The NCBI taxonomy database (ftp://ftp.ncbi.nih.gov/pub/taxonomy/) with minor modifications. The changes are listed in the file change.note
4. Scripts Perl scripts for identifying markers, generating trimmed multiple sequence alignments, and assigning phylotypes based on phylogenetic inferences.
5. bin Helper programs
USING AMPHORA
==============================================================
1. Phylotying bacterial sequences
Usage: AMPHORA_home/Scripts/Phylotyping.pl <options> protein-sequence-file output-file
Options:
-Replicates: number of bootstrap replicates
-BootstrapCutoff: normalized to 100 replicates (1-100)% default 70
Output:
foo.pep (identified maker sequences in fasta format, i.e. rpoB.pep)
foo.aln (aligned and trimmed marker sequences, i.e., rpoB.aln)
Trees/foo.tree (generated phylogenetic trees)
2. Identify marker sequences
Usage: perl AMPHORA_home/Scripts/MarkerScanner.pl protein-sequence-file
Output: foo.pep (identified maker sequences in fasta format, i.e., rpoB.pep)
3. Align and trim the marker sequences
Usage: perl AMPHORA_home/Scripts/MarkerAlignTrim.pl
Options: -Partial: query sequences contain partial genes
-Trim: trim the alignment using masks embedded within the marker database
-Strict: use a conservative mask
-Directory: the file directory where marker sequences are located. Default: current directory
-Help: print the help message
Output:
foo.aln (aligned and trimmed marker sequences, i.e., rpoB.aln)

About

Automated Phylogenomic Inference Pipeline

Resources

Stars

5 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - koadman/AMPHORA: Automated Phylogenomic Inference Pipeline · GitHub
Skip to content

Latest commit

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

AMPHORA
AMPHORA is an Automated Phylogenomic Inference Pipeline for bacterial sequences. From a given a set of protein sequences, it automatically identifies 31 phylogenetic marker genes. It then generates high-quality multiple sequence alignments for these genes and make tree-based phylotype assignments.
CITATION
==============================================================
Please cite AMPHORA as: Martin Wu and Jonathan A Eisen. A simple, fast, and accurate method of phylogenomic inference Genome Biology 2008, 9:R151
SYSTEM REQUIREMENT
============================================================== Linux OS (kernel version 2.6 or later)
The following software is required by the AMPHORA package. They need to be downloaded and installed separately from AMPHORA. 1. Perl 5.8.8 or later (www.perl.org) 2. Bioperl core package 1.5.2 or later (www.bioperl.org) 3. HMMER (hmmer.janelia.org) 4. WU BLAST (blast.wustl.edu)
The following software is included in the AMPHORA distribution. Their source codes have been slightly modified to suit the needs of AMPHORA. 1. seqboot (evolution.genetics.washington.edu/phylip) 2. quicktree (www.sanger.ac.uk/Software/analysis/quicktree) 3. raxml (icwww.epfl.ch/~stamatak/index-Dateien/Page443.htm)
LICENSE
==============================================================
AMPHORA Copyright 2008 by Martin Wu
AMPHORA is free software: you may redistribute it and/or modify its under the terms of the GNU General Public License as published by the Free Software Foundation; either version 2 of the License, or any later version.
AMPHORA is distributed in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU General Public License for more details (http://www.gnu.org/licenses/).
INSTALLATION
============================================================== 1. Unpack the package tar -xzvf AMPHORA.tar.gz
2. Install AMPHORA to a user specified directory. Make sure you have write permission to the directory.
cd AMPHORA perl INSTALL.pl -AMPHORA_home user-specified-directory -Bioperl_home path-to-bioperl
PACKAGE CONTENTS
============================================================== After successful installation, there should be several folders in the home directory of AMPHORA
1. Marker It contains curated seed multiple sequence alignments of the phylogenetic markers in GDE format (with embedded masks) and associated Hidden Markov Models. Currently there are 31 protein marker genes. They are dnaG, frr, infC, nusA, pgk, pyrG, rplA, rplB, rplC, rplD, rplE, rplF, rplK, rplL, rplM, rplN, rplP, rplS, rplT, rpmA, rpoB, rpsB, rpsC, rpsE, rpsI, rpsJ, rpsK, rpsM, rpsS, smpB, tsf.
2. Reference It contains protein sequences of the marker genes from all complete bacterial genomes. It also contains a bacterial genome tree that was made from the concatenated protein sequences of all the marker genes. Reference trees for each marker gene were derived from the genome tree by replacing the species names with their corresponding gene names and keeping the topology intact.
3. Taxonomy The NCBI taxonomy database (ftp://ftp.ncbi.nih.gov/pub/taxonomy/) with minor modifications. The changes are listed in the file change.note
4. Scripts Perl scripts for identifying markers, generating trimmed multiple sequence alignments, and assigning phylotypes based on phylogenetic inferences.
5. bin Helper programs
USING AMPHORA
==============================================================
1. Phylotying bacterial sequences
Usage: AMPHORA_home/Scripts/Phylotyping.pl <options> protein-sequence-file output-file
Options:
-Replicates: number of bootstrap replicates
-BootstrapCutoff: normalized to 100 replicates (1-100)% default 70
Output:
foo.pep (identified maker sequences in fasta format, i.e. rpoB.pep)
foo.aln (aligned and trimmed marker sequences, i.e., rpoB.aln)
Trees/foo.tree (generated phylogenetic trees)
2. Identify marker sequences
Usage: perl AMPHORA_home/Scripts/MarkerScanner.pl protein-sequence-file
Output: foo.pep (identified maker sequences in fasta format, i.e., rpoB.pep)
3. Align and trim the marker sequences
Usage: perl AMPHORA_home/Scripts/MarkerAlignTrim.pl
Options: -Partial: query sequences contain partial genes
-Trim: trim the alignment using masks embedded within the marker database
-Strict: use a conservative mask
-Directory: the file directory where marker sequences are located. Default: current directory
-Help: print the help message
Output:
foo.aln (aligned and trimmed marker sequences, i.e., rpoB.aln)

About

Automated Phylogenomic Inference Pipeline

Resources

Stars

5 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - koadman/AMPHORA: Automated Phylogenomic Inference Pipeline · GitHub
Skip to content

Latest commit

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

AMPHORA
AMPHORA is an Automated Phylogenomic Inference Pipeline for bacterial sequences. From a given a set of protein sequences, it automatically identifies 31 phylogenetic marker genes. It then generates high-quality multiple sequence alignments for these genes and make tree-based phylotype assignments.
CITATION
==============================================================
Please cite AMPHORA as: Martin Wu and Jonathan A Eisen. A simple, fast, and accurate method of phylogenomic inference Genome Biology 2008, 9:R151
SYSTEM REQUIREMENT
============================================================== Linux OS (kernel version 2.6 or later)
The following software is required by the AMPHORA package. They need to be downloaded and installed separately from AMPHORA. 1. Perl 5.8.8 or later (www.perl.org) 2. Bioperl core package 1.5.2 or later (www.bioperl.org) 3. HMMER (hmmer.janelia.org) 4. WU BLAST (blast.wustl.edu)
The following software is included in the AMPHORA distribution. Their source codes have been slightly modified to suit the needs of AMPHORA. 1. seqboot (evolution.genetics.washington.edu/phylip) 2. quicktree (www.sanger.ac.uk/Software/analysis/quicktree) 3. raxml (icwww.epfl.ch/~stamatak/index-Dateien/Page443.htm)
LICENSE
==============================================================
AMPHORA Copyright 2008 by Martin Wu
AMPHORA is free software: you may redistribute it and/or modify its under the terms of the GNU General Public License as published by the Free Software Foundation; either version 2 of the License, or any later version.
AMPHORA is distributed in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU General Public License for more details (http://www.gnu.org/licenses/).
INSTALLATION
============================================================== 1. Unpack the package tar -xzvf AMPHORA.tar.gz
2. Install AMPHORA to a user specified directory. Make sure you have write permission to the directory.
cd AMPHORA perl INSTALL.pl -AMPHORA_home user-specified-directory -Bioperl_home path-to-bioperl
PACKAGE CONTENTS
============================================================== After successful installation, there should be several folders in the home directory of AMPHORA
1. Marker It contains curated seed multiple sequence alignments of the phylogenetic markers in GDE format (with embedded masks) and associated Hidden Markov Models. Currently there are 31 protein marker genes. They are dnaG, frr, infC, nusA, pgk, pyrG, rplA, rplB, rplC, rplD, rplE, rplF, rplK, rplL, rplM, rplN, rplP, rplS, rplT, rpmA, rpoB, rpsB, rpsC, rpsE, rpsI, rpsJ, rpsK, rpsM, rpsS, smpB, tsf.
2. Reference It contains protein sequences of the marker genes from all complete bacterial genomes. It also contains a bacterial genome tree that was made from the concatenated protein sequences of all the marker genes. Reference trees for each marker gene were derived from the genome tree by replacing the species names with their corresponding gene names and keeping the topology intact.
3. Taxonomy The NCBI taxonomy database (ftp://ftp.ncbi.nih.gov/pub/taxonomy/) with minor modifications. The changes are listed in the file change.note
4. Scripts Perl scripts for identifying markers, generating trimmed multiple sequence alignments, and assigning phylotypes based on phylogenetic inferences.
5. bin Helper programs
USING AMPHORA
==============================================================
1. Phylotying bacterial sequences
Usage: AMPHORA_home/Scripts/Phylotyping.pl <options> protein-sequence-file output-file
Options:
-Replicates: number of bootstrap replicates
-BootstrapCutoff: normalized to 100 replicates (1-100)% default 70
Output:
foo.pep (identified maker sequences in fasta format, i.e. rpoB.pep)
foo.aln (aligned and trimmed marker sequences, i.e., rpoB.aln)
Trees/foo.tree (generated phylogenetic trees)
2. Identify marker sequences
Usage: perl AMPHORA_home/Scripts/MarkerScanner.pl protein-sequence-file
Output: foo.pep (identified maker sequences in fasta format, i.e., rpoB.pep)
3. Align and trim the marker sequences
Usage: perl AMPHORA_home/Scripts/MarkerAlignTrim.pl
Options: -Partial: query sequences contain partial genes
-Trim: trim the alignment using masks embedded within the marker database
-Strict: use a conservative mask
-Directory: the file directory where marker sequences are located. Default: current directory
-Help: print the help message
Output:
foo.aln (aligned and trimmed marker sequences, i.e., rpoB.aln)

About

Automated Phylogenomic Inference Pipeline

Resources

Stars

5 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - koadman/AMPHORA: Automated Phylogenomic Inference Pipeline · GitHub
Skip to content

Latest commit

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

AMPHORA
AMPHORA is an Automated Phylogenomic Inference Pipeline for bacterial sequences. From a given a set of protein sequences, it automatically identifies 31 phylogenetic marker genes. It then generates high-quality multiple sequence alignments for these genes and make tree-based phylotype assignments.
CITATION
==============================================================
Please cite AMPHORA as: Martin Wu and Jonathan A Eisen. A simple, fast, and accurate method of phylogenomic inference Genome Biology 2008, 9:R151
SYSTEM REQUIREMENT
============================================================== Linux OS (kernel version 2.6 or later)
The following software is required by the AMPHORA package. They need to be downloaded and installed separately from AMPHORA. 1. Perl 5.8.8 or later (www.perl.org) 2. Bioperl core package 1.5.2 or later (www.bioperl.org) 3. HMMER (hmmer.janelia.org) 4. WU BLAST (blast.wustl.edu)
The following software is included in the AMPHORA distribution. Their source codes have been slightly modified to suit the needs of AMPHORA. 1. seqboot (evolution.genetics.washington.edu/phylip) 2. quicktree (www.sanger.ac.uk/Software/analysis/quicktree) 3. raxml (icwww.epfl.ch/~stamatak/index-Dateien/Page443.htm)
LICENSE
==============================================================
AMPHORA Copyright 2008 by Martin Wu
AMPHORA is free software: you may redistribute it and/or modify its under the terms of the GNU General Public License as published by the Free Software Foundation; either version 2 of the License, or any later version.
AMPHORA is distributed in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU General Public License for more details (http://www.gnu.org/licenses/).
INSTALLATION
============================================================== 1. Unpack the package tar -xzvf AMPHORA.tar.gz
2. Install AMPHORA to a user specified directory. Make sure you have write permission to the directory.
cd AMPHORA perl INSTALL.pl -AMPHORA_home user-specified-directory -Bioperl_home path-to-bioperl
PACKAGE CONTENTS
============================================================== After successful installation, there should be several folders in the home directory of AMPHORA
1. Marker It contains curated seed multiple sequence alignments of the phylogenetic markers in GDE format (with embedded masks) and associated Hidden Markov Models. Currently there are 31 protein marker genes. They are dnaG, frr, infC, nusA, pgk, pyrG, rplA, rplB, rplC, rplD, rplE, rplF, rplK, rplL, rplM, rplN, rplP, rplS, rplT, rpmA, rpoB, rpsB, rpsC, rpsE, rpsI, rpsJ, rpsK, rpsM, rpsS, smpB, tsf.
2. Reference It contains protein sequences of the marker genes from all complete bacterial genomes. It also contains a bacterial genome tree that was made from the concatenated protein sequences of all the marker genes. Reference trees for each marker gene were derived from the genome tree by replacing the species names with their corresponding gene names and keeping the topology intact.
3. Taxonomy The NCBI taxonomy database (ftp://ftp.ncbi.nih.gov/pub/taxonomy/) with minor modifications. The changes are listed in the file change.note
4. Scripts Perl scripts for identifying markers, generating trimmed multiple sequence alignments, and assigning phylotypes based on phylogenetic inferences.
5. bin Helper programs
USING AMPHORA
==============================================================
1. Phylotying bacterial sequences
Usage: AMPHORA_home/Scripts/Phylotyping.pl <options> protein-sequence-file output-file
Options:
-Replicates: number of bootstrap replicates
-BootstrapCutoff: normalized to 100 replicates (1-100)% default 70
Output:
foo.pep (identified maker sequences in fasta format, i.e. rpoB.pep)
foo.aln (aligned and trimmed marker sequences, i.e., rpoB.aln)
Trees/foo.tree (generated phylogenetic trees)
2. Identify marker sequences
Usage: perl AMPHORA_home/Scripts/MarkerScanner.pl protein-sequence-file
Output: foo.pep (identified maker sequences in fasta format, i.e., rpoB.pep)
3. Align and trim the marker sequences
Usage: perl AMPHORA_home/Scripts/MarkerAlignTrim.pl
Options: -Partial: query sequences contain partial genes
-Trim: trim the alignment using masks embedded within the marker database
-Strict: use a conservative mask
-Directory: the file directory where marker sequences are located. Default: current directory
-Help: print the help message
Output:
foo.aln (aligned and trimmed marker sequences, i.e., rpoB.aln)

About

Automated Phylogenomic Inference Pipeline

Resources

Stars

5 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - koadman/AMPHORA: Automated Phylogenomic Inference Pipeline · GitHub
Skip to content

Latest commit

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

AMPHORA
AMPHORA is an Automated Phylogenomic Inference Pipeline for bacterial sequences. From a given a set of protein sequences, it automatically identifies 31 phylogenetic marker genes. It then generates high-quality multiple sequence alignments for these genes and make tree-based phylotype assignments.
CITATION
==============================================================
Please cite AMPHORA as: Martin Wu and Jonathan A Eisen. A simple, fast, and accurate method of phylogenomic inference Genome Biology 2008, 9:R151
SYSTEM REQUIREMENT
============================================================== Linux OS (kernel version 2.6 or later)
The following software is required by the AMPHORA package. They need to be downloaded and installed separately from AMPHORA. 1. Perl 5.8.8 or later (www.perl.org) 2. Bioperl core package 1.5.2 or later (www.bioperl.org) 3. HMMER (hmmer.janelia.org) 4. WU BLAST (blast.wustl.edu)
The following software is included in the AMPHORA distribution. Their source codes have been slightly modified to suit the needs of AMPHORA. 1. seqboot (evolution.genetics.washington.edu/phylip) 2. quicktree (www.sanger.ac.uk/Software/analysis/quicktree) 3. raxml (icwww.epfl.ch/~stamatak/index-Dateien/Page443.htm)
LICENSE
==============================================================
AMPHORA Copyright 2008 by Martin Wu
AMPHORA is free software: you may redistribute it and/or modify its under the terms of the GNU General Public License as published by the Free Software Foundation; either version 2 of the License, or any later version.
AMPHORA is distributed in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU General Public License for more details (http://www.gnu.org/licenses/).
INSTALLATION
============================================================== 1. Unpack the package tar -xzvf AMPHORA.tar.gz
2. Install AMPHORA to a user specified directory. Make sure you have write permission to the directory.
cd AMPHORA perl INSTALL.pl -AMPHORA_home user-specified-directory -Bioperl_home path-to-bioperl
PACKAGE CONTENTS
============================================================== After successful installation, there should be several folders in the home directory of AMPHORA
1. Marker It contains curated seed multiple sequence alignments of the phylogenetic markers in GDE format (with embedded masks) and associated Hidden Markov Models. Currently there are 31 protein marker genes. They are dnaG, frr, infC, nusA, pgk, pyrG, rplA, rplB, rplC, rplD, rplE, rplF, rplK, rplL, rplM, rplN, rplP, rplS, rplT, rpmA, rpoB, rpsB, rpsC, rpsE, rpsI, rpsJ, rpsK, rpsM, rpsS, smpB, tsf.
2. Reference It contains protein sequences of the marker genes from all complete bacterial genomes. It also contains a bacterial genome tree that was made from the concatenated protein sequences of all the marker genes. Reference trees for each marker gene were derived from the genome tree by replacing the species names with their corresponding gene names and keeping the topology intact.
3. Taxonomy The NCBI taxonomy database (ftp://ftp.ncbi.nih.gov/pub/taxonomy/) with minor modifications. The changes are listed in the file change.note
4. Scripts Perl scripts for identifying markers, generating trimmed multiple sequence alignments, and assigning phylotypes based on phylogenetic inferences.
5. bin Helper programs
USING AMPHORA
==============================================================
1. Phylotying bacterial sequences
Usage: AMPHORA_home/Scripts/Phylotyping.pl <options> protein-sequence-file output-file
Options:
-Replicates: number of bootstrap replicates
-BootstrapCutoff: normalized to 100 replicates (1-100)% default 70
Output:
foo.pep (identified maker sequences in fasta format, i.e. rpoB.pep)
foo.aln (aligned and trimmed marker sequences, i.e., rpoB.aln)
Trees/foo.tree (generated phylogenetic trees)
2. Identify marker sequences
Usage: perl AMPHORA_home/Scripts/MarkerScanner.pl protein-sequence-file
Output: foo.pep (identified maker sequences in fasta format, i.e., rpoB.pep)
3. Align and trim the marker sequences
Usage: perl AMPHORA_home/Scripts/MarkerAlignTrim.pl
Options: -Partial: query sequences contain partial genes
-Trim: trim the alignment using masks embedded within the marker database
-Strict: use a conservative mask
-Directory: the file directory where marker sequences are located. Default: current directory
-Help: print the help message
Output:
foo.aln (aligned and trimmed marker sequences, i.e., rpoB.aln)

About

Automated Phylogenomic Inference Pipeline

Resources

Stars

5 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); })(); GitHub - koadman/AMPHORA: Automated Phylogenomic Inference Pipeline · GitHub
Skip to content

Latest commit

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

AMPHORA
AMPHORA is an Automated Phylogenomic Inference Pipeline for bacterial sequences. From a given a set of protein sequences, it automatically identifies 31 phylogenetic marker genes. It then generates high-quality multiple sequence alignments for these genes and make tree-based phylotype assignments.
CITATION
==============================================================
Please cite AMPHORA as: Martin Wu and Jonathan A Eisen. A simple, fast, and accurate method of phylogenomic inference Genome Biology 2008, 9:R151
SYSTEM REQUIREMENT
============================================================== Linux OS (kernel version 2.6 or later)
The following software is required by the AMPHORA package. They need to be downloaded and installed separately from AMPHORA. 1. Perl 5.8.8 or later (www.perl.org) 2. Bioperl core package 1.5.2 or later (www.bioperl.org) 3. HMMER (hmmer.janelia.org) 4. WU BLAST (blast.wustl.edu)
The following software is included in the AMPHORA distribution. Their source codes have been slightly modified to suit the needs of AMPHORA. 1. seqboot (evolution.genetics.washington.edu/phylip) 2. quicktree (www.sanger.ac.uk/Software/analysis/quicktree) 3. raxml (icwww.epfl.ch/~stamatak/index-Dateien/Page443.htm)
LICENSE
==============================================================
AMPHORA Copyright 2008 by Martin Wu
AMPHORA is free software: you may redistribute it and/or modify its under the terms of the GNU General Public License as published by the Free Software Foundation; either version 2 of the License, or any later version.
AMPHORA is distributed in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU General Public License for more details (http://www.gnu.org/licenses/).
INSTALLATION
============================================================== 1. Unpack the package tar -xzvf AMPHORA.tar.gz
2. Install AMPHORA to a user specified directory. Make sure you have write permission to the directory.
cd AMPHORA perl INSTALL.pl -AMPHORA_home user-specified-directory -Bioperl_home path-to-bioperl
PACKAGE CONTENTS
============================================================== After successful installation, there should be several folders in the home directory of AMPHORA
1. Marker It contains curated seed multiple sequence alignments of the phylogenetic markers in GDE format (with embedded masks) and associated Hidden Markov Models. Currently there are 31 protein marker genes. They are dnaG, frr, infC, nusA, pgk, pyrG, rplA, rplB, rplC, rplD, rplE, rplF, rplK, rplL, rplM, rplN, rplP, rplS, rplT, rpmA, rpoB, rpsB, rpsC, rpsE, rpsI, rpsJ, rpsK, rpsM, rpsS, smpB, tsf.
2. Reference It contains protein sequences of the marker genes from all complete bacterial genomes. It also contains a bacterial genome tree that was made from the concatenated protein sequences of all the marker genes. Reference trees for each marker gene were derived from the genome tree by replacing the species names with their corresponding gene names and keeping the topology intact.
3. Taxonomy The NCBI taxonomy database (ftp://ftp.ncbi.nih.gov/pub/taxonomy/) with minor modifications. The changes are listed in the file change.note
4. Scripts Perl scripts for identifying markers, generating trimmed multiple sequence alignments, and assigning phylotypes based on phylogenetic inferences.
5. bin Helper programs
USING AMPHORA
==============================================================
1. Phylotying bacterial sequences
Usage: AMPHORA_home/Scripts/Phylotyping.pl <options> protein-sequence-file output-file
Options:
-Replicates: number of bootstrap replicates
-BootstrapCutoff: normalized to 100 replicates (1-100)% default 70
Output:
foo.pep (identified maker sequences in fasta format, i.e. rpoB.pep)
foo.aln (aligned and trimmed marker sequences, i.e., rpoB.aln)
Trees/foo.tree (generated phylogenetic trees)
2. Identify marker sequences
Usage: perl AMPHORA_home/Scripts/MarkerScanner.pl protein-sequence-file
Output: foo.pep (identified maker sequences in fasta format, i.e., rpoB.pep)
3. Align and trim the marker sequences
Usage: perl AMPHORA_home/Scripts/MarkerAlignTrim.pl
Options: -Partial: query sequences contain partial genes
-Trim: trim the alignment using masks embedded within the marker database
-Strict: use a conservative mask
-Directory: the file directory where marker sequences are located. Default: current directory
-Help: print the help message
Output:
foo.aln (aligned and trimmed marker sequences, i.e., rpoB.aln)

About

Automated Phylogenomic Inference Pipeline

Resources

Stars

5 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages