Latest commit

History

59 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

fatools: A python utility package for working with large-scale fasta sequences

A total of 30 utilities/options for common operarations with fasta sequences are currently organized under 6 main subcommands. The utilities range from searching specific sequence entries from a large set of fasta sequences based on ID or a text string in the defline or a sequence motif to spliting a large sequence file into small chuncks, reporting summary stats of a genome assembly, and filtering sequences by length or gap size or redundant entries based on ID or sequences. Furthermore, by allowing input from and output to stdout, multiple processes can be done sequencially in one line of commpands via use of pipe (|). Make sure you have python 2 installed to run fatools.

List of subcommands

Typing 'fatools ' displays the list of subcomands; typing 'fatools ' displays the detailed utilities/options for the subcommand.

convert

-r print sequence in revevrse compliment.

-N convert all non-ACGT letters to N.

-R remove all non-ACGT letters.

-U to upper case

-u to lower case


extract

-F N extract the first N fasta entries, if N is larger than the total number of entries, then print to the last entries.

-S N extract from the Nth entry to the last entry.

-L N extract last N sequence entries. If N is larger than the total number of entries, then print all entries.

Use -S N -F M for entries from N to M; use -F N and -L N to extract both the first and last N entries.

-f N extract first N bp, prints the entire sequence if N is larger than the total length.

-s N extract sequence up to to N bp.

-l N extract last N bp, prints the entire sequence if N is larger than the total length.

Use -s N -f M for sequence from N to M bp; use -f N and -l N to extract both the first and last N bp as one sequence separated by a space.

Note: The -f, -s, and -l options were designed for working with a single long sequences, even though they will work for multiple sequences by applying the same operation to all sequences.


filter

-g N skip sequences with N or more Ns.

-r 1/2 1: skip redundant entry based ID; 2: keep redundant entries by adding a serial number to the identical IDs to make each ID unique.

-R 1/N skip redundant entries based on sequence. 1: use the entire sequence; N: use only the first and last N bases.

-l N skip sequences shorter than N bp.

-L N skip sequences longer than N bp.

use -l N -L M for sequences with length from N to M bp (inclusive).

In all options, '-e' can be added to print the skipped entries in STDERR, which can be captured using 2>[skipped.fa].


report

-f print fasta entries as in the input.

-F print fasta entries with all sequence in one line.

-n print sequences without the defline.

-d print deflines in short form (part before the first space).

-D print deflines in the original form.

-c print the total number of fasta entries in the input.

-l print short defline +[\t] length.

-L print original defline +[\t] length.

-s print sequence summary statistics including N50.

-S print sequence summary statistics plus detailed gap info.
Use -h with -s and -S to disable the header above the outputs
Use -H to print parameters in human friendly form.


search

-s string: search for entries containing "string" in the sequence.

-d string: search for entries containing "string" in the defline: Default is for exact match; use "/string" to search for entries with "string" as part of the ID.

-F file: search for sequences based on a list of IDs in the file (one ID/line).
Can use -D to specify delimiter in the defline. Default is space or '|' or end of line;
use -i to specify the field number, default is 1.

-1 print only the 1st match for -d and -s.

-v use with -s, -d or -F to negate the search.


split

-G N split each of the sequences in the input file as non-gap fragments.
"N" is the number of consecutive Ns base, default is 1;
Use -G N with -t to print just the gap positions.

-n N split the input sequences into chunks, each containing N fasta entries (the last chunk may be less).

-N N split the input sequences into N chunks, each containing equal number of entries (last one may be smaller).

-M N split the input sequences into chunks at ~N MB (million bp) in size (last chunk may be smaller).

-o file: prefix for output files (serial numbers added to prefix; required).


Making fatools executable and available from any directory

Linux

  1. Open the terminal and navigate to where you have downloaded fatools
  2. Find where python2 is installed in your system which python. Usually you would get something like /usr/bin/python
  3. Copy this output to the beggining of fatools as #!/usr/bin/python.
  4. Run the following command to make the script executable chmod +x fatools
  5. Add fatools to your bin directory or any other directories included in your $PATH
  6. You should be able to run fatools from anywhere!

Windows

  1. Type 'control panel' in the Windows search bar
  2. Go to System and Security > System > Advanced System settings > Environment variables
  3. Under system variables, select 'Path' and click 'Edit' and then 'New'
  4. Add the path of where you have fatools located and click OK
  5. You should be able to run fatools from anywhere!

Examples

Navigate to the exampleFiles directory in this repository. In there, a fasta file (exampleFasta.fa) and a file containing a list of IDs (IDlist.txt) from exampleFasta.fa.

To extract the fasta sequences from exampleFasta.fasta based on the list of IDs:
fatools search -F IDlist.txt exampleFasta.fa

NP_001245510.1 notch, isoform B [Drosophila melanogaster]
NNMQSQRSRRRSRAPNTWICFWINKMHAVASLPASLPLLLLTLAFANLPNTVRGTDTALVAASCTSVGCQNG
GTCVTQLNGKTYCACDSHYVGDYCEHRNPCNSMRCQNGGTCQVTFRNGRPGISCKCPLGFDESLCEIAVP
NACDHVTCLNGGTCQLKTLEEYTCACANGYTGERCETKNLCASSPCRNGATCTALAGSSSFTCSCPPGFT... DY343456.1 Macropus rufus BRCA1 (BRCA1) gene, partial cds
CAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAGTCTGGATGAAAGTAAGGAAATATGTAGTGCTGGA
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGTAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCA FY343456.1 Macropus rufus BRCA1 (BRCA1) gene, partial cds
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGCAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAG
TCTGGATGAAAGTAAGGAAATATGTAGTGCTGGAAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCAT 

To report summary statistics:
fatools report -s exampleFasta.fa

Total 2,868 bps from qualified 5 sequences (5 total); length average: 573 (210-1262) bp; N50: 698 bp

To get fasta sequences with a specific maximum length filter

python fatools filter -L250 exampleFasta.fa

DY343456.1
CAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAGTCTGGATGAAAGTAAGGAAATATGTAGTGCTGGA
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGTAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCA
FY343456.1
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGCAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAG
TCTGGATGAAAGTAAGGAAATATGTAGTGCTGGAAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCAT 

You can combine multiple utilities using pipe "|". Let's say you want to see the short defline and the length of the first 3 fasta sequences in a fasta file.

fatools extract -F3 sequenceTesting2.txt | fatools report -l -

AY211956.1Macropus(BRCA1)gene,partialcds 698 NP_001245510.1 1262 DY343456.1 210

Or extract the sequences from a large sequence set for a list of IDs and then search sequences with a specific sequence by using

fatools search -F IDlist.txt exampleFasta.fa |fatools -S AAATAAA -

About

a python tool package for working with fasta sequences

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

59 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

fatools: A python utility package for working with large-scale fasta sequences

A total of 30 utilities/options for common operarations with fasta sequences are currently organized under 6 main subcommands. The utilities range from searching specific sequence entries from a large set of fasta sequences based on ID or a text string in the defline or a sequence motif to spliting a large sequence file into small chuncks, reporting summary stats of a genome assembly, and filtering sequences by length or gap size or redundant entries based on ID or sequences. Furthermore, by allowing input from and output to stdout, multiple processes can be done sequencially in one line of commpands via use of pipe (|). Make sure you have python 2 installed to run fatools.

List of subcommands

Typing 'fatools ' displays the list of subcomands; typing 'fatools ' displays the detailed utilities/options for the subcommand.

convert

-r print sequence in revevrse compliment.

-N convert all non-ACGT letters to N.

-R remove all non-ACGT letters.

-U to upper case

-u to lower case


extract

-F N extract the first N fasta entries, if N is larger than the total number of entries, then print to the last entries.

-S N extract from the Nth entry to the last entry.

-L N extract last N sequence entries. If N is larger than the total number of entries, then print all entries.

Use -S N -F M for entries from N to M; use -F N and -L N to extract both the first and last N entries.

-f N extract first N bp, prints the entire sequence if N is larger than the total length.

-s N extract sequence up to to N bp.

-l N extract last N bp, prints the entire sequence if N is larger than the total length.

Use -s N -f M for sequence from N to M bp; use -f N and -l N to extract both the first and last N bp as one sequence separated by a space.

Note: The -f, -s, and -l options were designed for working with a single long sequences, even though they will work for multiple sequences by applying the same operation to all sequences.


filter

-g N skip sequences with N or more Ns.

-r 1/2 1: skip redundant entry based ID; 2: keep redundant entries by adding a serial number to the identical IDs to make each ID unique.

-R 1/N skip redundant entries based on sequence. 1: use the entire sequence; N: use only the first and last N bases.

-l N skip sequences shorter than N bp.

-L N skip sequences longer than N bp.

use -l N -L M for sequences with length from N to M bp (inclusive).

In all options, '-e' can be added to print the skipped entries in STDERR, which can be captured using 2>[skipped.fa].


report

-f print fasta entries as in the input.

-F print fasta entries with all sequence in one line.

-n print sequences without the defline.

-d print deflines in short form (part before the first space).

-D print deflines in the original form.

-c print the total number of fasta entries in the input.

-l print short defline +[\t] length.

-L print original defline +[\t] length.

-s print sequence summary statistics including N50.

-S print sequence summary statistics plus detailed gap info.
Use -h with -s and -S to disable the header above the outputs
Use -H to print parameters in human friendly form.


search

-s string: search for entries containing "string" in the sequence.

-d string: search for entries containing "string" in the defline: Default is for exact match; use "/string" to search for entries with "string" as part of the ID.

-F file: search for sequences based on a list of IDs in the file (one ID/line).
Can use -D to specify delimiter in the defline. Default is space or '|' or end of line;
use -i to specify the field number, default is 1.

-1 print only the 1st match for -d and -s.

-v use with -s, -d or -F to negate the search.


split

-G N split each of the sequences in the input file as non-gap fragments.
"N" is the number of consecutive Ns base, default is 1;
Use -G N with -t to print just the gap positions.

-n N split the input sequences into chunks, each containing N fasta entries (the last chunk may be less).

-N N split the input sequences into N chunks, each containing equal number of entries (last one may be smaller).

-M N split the input sequences into chunks at ~N MB (million bp) in size (last chunk may be smaller).

-o file: prefix for output files (serial numbers added to prefix; required).


Making fatools executable and available from any directory

Linux

  1. Open the terminal and navigate to where you have downloaded fatools
  2. Find where python2 is installed in your system which python. Usually you would get something like /usr/bin/python
  3. Copy this output to the beggining of fatools as #!/usr/bin/python.
  4. Run the following command to make the script executable chmod +x fatools
  5. Add fatools to your bin directory or any other directories included in your $PATH
  6. You should be able to run fatools from anywhere!

Windows

  1. Type 'control panel' in the Windows search bar
  2. Go to System and Security > System > Advanced System settings > Environment variables
  3. Under system variables, select 'Path' and click 'Edit' and then 'New'
  4. Add the path of where you have fatools located and click OK
  5. You should be able to run fatools from anywhere!

Examples

Navigate to the exampleFiles directory in this repository. In there, a fasta file (exampleFasta.fa) and a file containing a list of IDs (IDlist.txt) from exampleFasta.fa.

To extract the fasta sequences from exampleFasta.fasta based on the list of IDs:
fatools search -F IDlist.txt exampleFasta.fa

NP_001245510.1 notch, isoform B [Drosophila melanogaster]
NNMQSQRSRRRSRAPNTWICFWINKMHAVASLPASLPLLLLTLAFANLPNTVRGTDTALVAASCTSVGCQNG
GTCVTQLNGKTYCACDSHYVGDYCEHRNPCNSMRCQNGGTCQVTFRNGRPGISCKCPLGFDESLCEIAVP
NACDHVTCLNGGTCQLKTLEEYTCACANGYTGERCETKNLCASSPCRNGATCTALAGSSSFTCSCPPGFT... DY343456.1 Macropus rufus BRCA1 (BRCA1) gene, partial cds
CAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAGTCTGGATGAAAGTAAGGAAATATGTAGTGCTGGA
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGTAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCA FY343456.1 Macropus rufus BRCA1 (BRCA1) gene, partial cds
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGCAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAG
TCTGGATGAAAGTAAGGAAATATGTAGTGCTGGAAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCAT 

To report summary statistics:
fatools report -s exampleFasta.fa

Total 2,868 bps from qualified 5 sequences (5 total); length average: 573 (210-1262) bp; N50: 698 bp

To get fasta sequences with a specific maximum length filter

python fatools filter -L250 exampleFasta.fa

DY343456.1
CAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAGTCTGGATGAAAGTAAGGAAATATGTAGTGCTGGA
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGTAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCA
FY343456.1
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGCAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAG
TCTGGATGAAAGTAAGGAAATATGTAGTGCTGGAAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCAT 

You can combine multiple utilities using pipe "|". Let's say you want to see the short defline and the length of the first 3 fasta sequences in a fasta file.

fatools extract -F3 sequenceTesting2.txt | fatools report -l -

AY211956.1Macropus(BRCA1)gene,partialcds 698 NP_001245510.1 1262 DY343456.1 210

Or extract the sequences from a large sequence set for a list of IDs and then search sequences with a specific sequence by using

fatools search -F IDlist.txt exampleFasta.fa |fatools -S AAATAAA -

About

a python tool package for working with fasta sequences

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

59 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

fatools: A python utility package for working with large-scale fasta sequences

A total of 30 utilities/options for common operarations with fasta sequences are currently organized under 6 main subcommands. The utilities range from searching specific sequence entries from a large set of fasta sequences based on ID or a text string in the defline or a sequence motif to spliting a large sequence file into small chuncks, reporting summary stats of a genome assembly, and filtering sequences by length or gap size or redundant entries based on ID or sequences. Furthermore, by allowing input from and output to stdout, multiple processes can be done sequencially in one line of commpands via use of pipe (|). Make sure you have python 2 installed to run fatools.

List of subcommands

Typing 'fatools ' displays the list of subcomands; typing 'fatools ' displays the detailed utilities/options for the subcommand.

convert

-r print sequence in revevrse compliment.

-N convert all non-ACGT letters to N.

-R remove all non-ACGT letters.

-U to upper case

-u to lower case


extract

-F N extract the first N fasta entries, if N is larger than the total number of entries, then print to the last entries.

-S N extract from the Nth entry to the last entry.

-L N extract last N sequence entries. If N is larger than the total number of entries, then print all entries.

Use -S N -F M for entries from N to M; use -F N and -L N to extract both the first and last N entries.

-f N extract first N bp, prints the entire sequence if N is larger than the total length.

-s N extract sequence up to to N bp.

-l N extract last N bp, prints the entire sequence if N is larger than the total length.

Use -s N -f M for sequence from N to M bp; use -f N and -l N to extract both the first and last N bp as one sequence separated by a space.

Note: The -f, -s, and -l options were designed for working with a single long sequences, even though they will work for multiple sequences by applying the same operation to all sequences.


filter

-g N skip sequences with N or more Ns.

-r 1/2 1: skip redundant entry based ID; 2: keep redundant entries by adding a serial number to the identical IDs to make each ID unique.

-R 1/N skip redundant entries based on sequence. 1: use the entire sequence; N: use only the first and last N bases.

-l N skip sequences shorter than N bp.

-L N skip sequences longer than N bp.

use -l N -L M for sequences with length from N to M bp (inclusive).

In all options, '-e' can be added to print the skipped entries in STDERR, which can be captured using 2>[skipped.fa].


report

-f print fasta entries as in the input.

-F print fasta entries with all sequence in one line.

-n print sequences without the defline.

-d print deflines in short form (part before the first space).

-D print deflines in the original form.

-c print the total number of fasta entries in the input.

-l print short defline +[\t] length.

-L print original defline +[\t] length.

-s print sequence summary statistics including N50.

-S print sequence summary statistics plus detailed gap info.
Use -h with -s and -S to disable the header above the outputs
Use -H to print parameters in human friendly form.


search

-s string: search for entries containing "string" in the sequence.

-d string: search for entries containing "string" in the defline: Default is for exact match; use "/string" to search for entries with "string" as part of the ID.

-F file: search for sequences based on a list of IDs in the file (one ID/line).
Can use -D to specify delimiter in the defline. Default is space or '|' or end of line;
use -i to specify the field number, default is 1.

-1 print only the 1st match for -d and -s.

-v use with -s, -d or -F to negate the search.


split

-G N split each of the sequences in the input file as non-gap fragments.
"N" is the number of consecutive Ns base, default is 1;
Use -G N with -t to print just the gap positions.

-n N split the input sequences into chunks, each containing N fasta entries (the last chunk may be less).

-N N split the input sequences into N chunks, each containing equal number of entries (last one may be smaller).

-M N split the input sequences into chunks at ~N MB (million bp) in size (last chunk may be smaller).

-o file: prefix for output files (serial numbers added to prefix; required).


Making fatools executable and available from any directory

Linux

  1. Open the terminal and navigate to where you have downloaded fatools
  2. Find where python2 is installed in your system which python. Usually you would get something like /usr/bin/python
  3. Copy this output to the beggining of fatools as #!/usr/bin/python.
  4. Run the following command to make the script executable chmod +x fatools
  5. Add fatools to your bin directory or any other directories included in your $PATH
  6. You should be able to run fatools from anywhere!

Windows

  1. Type 'control panel' in the Windows search bar
  2. Go to System and Security > System > Advanced System settings > Environment variables
  3. Under system variables, select 'Path' and click 'Edit' and then 'New'
  4. Add the path of where you have fatools located and click OK
  5. You should be able to run fatools from anywhere!

Examples

Navigate to the exampleFiles directory in this repository. In there, a fasta file (exampleFasta.fa) and a file containing a list of IDs (IDlist.txt) from exampleFasta.fa.

To extract the fasta sequences from exampleFasta.fasta based on the list of IDs:
fatools search -F IDlist.txt exampleFasta.fa

NP_001245510.1 notch, isoform B [Drosophila melanogaster]
NNMQSQRSRRRSRAPNTWICFWINKMHAVASLPASLPLLLLTLAFANLPNTVRGTDTALVAASCTSVGCQNG
GTCVTQLNGKTYCACDSHYVGDYCEHRNPCNSMRCQNGGTCQVTFRNGRPGISCKCPLGFDESLCEIAVP
NACDHVTCLNGGTCQLKTLEEYTCACANGYTGERCETKNLCASSPCRNGATCTALAGSSSFTCSCPPGFT... DY343456.1 Macropus rufus BRCA1 (BRCA1) gene, partial cds
CAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAGTCTGGATGAAAGTAAGGAAATATGTAGTGCTGGA
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGTAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCA FY343456.1 Macropus rufus BRCA1 (BRCA1) gene, partial cds
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGCAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAG
TCTGGATGAAAGTAAGGAAATATGTAGTGCTGGAAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCAT 

To report summary statistics:
fatools report -s exampleFasta.fa

Total 2,868 bps from qualified 5 sequences (5 total); length average: 573 (210-1262) bp; N50: 698 bp

To get fasta sequences with a specific maximum length filter

python fatools filter -L250 exampleFasta.fa

DY343456.1
CAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAGTCTGGATGAAAGTAAGGAAATATGTAGTGCTGGA
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGTAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCA
FY343456.1
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGCAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAG
TCTGGATGAAAGTAAGGAAATATGTAGTGCTGGAAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCAT 

You can combine multiple utilities using pipe "|". Let's say you want to see the short defline and the length of the first 3 fasta sequences in a fasta file.

fatools extract -F3 sequenceTesting2.txt | fatools report -l -

AY211956.1Macropus(BRCA1)gene,partialcds 698 NP_001245510.1 1262 DY343456.1 210

Or extract the sequences from a large sequence set for a list of IDs and then search sequences with a specific sequence by using

fatools search -F IDlist.txt exampleFasta.fa |fatools -S AAATAAA -

About

a python tool package for working with fasta sequences

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

59 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

fatools: A python utility package for working with large-scale fasta sequences

A total of 30 utilities/options for common operarations with fasta sequences are currently organized under 6 main subcommands. The utilities range from searching specific sequence entries from a large set of fasta sequences based on ID or a text string in the defline or a sequence motif to spliting a large sequence file into small chuncks, reporting summary stats of a genome assembly, and filtering sequences by length or gap size or redundant entries based on ID or sequences. Furthermore, by allowing input from and output to stdout, multiple processes can be done sequencially in one line of commpands via use of pipe (|). Make sure you have python 2 installed to run fatools.

List of subcommands

Typing 'fatools ' displays the list of subcomands; typing 'fatools ' displays the detailed utilities/options for the subcommand.

convert

-r print sequence in revevrse compliment.

-N convert all non-ACGT letters to N.

-R remove all non-ACGT letters.

-U to upper case

-u to lower case


extract

-F N extract the first N fasta entries, if N is larger than the total number of entries, then print to the last entries.

-S N extract from the Nth entry to the last entry.

-L N extract last N sequence entries. If N is larger than the total number of entries, then print all entries.

Use -S N -F M for entries from N to M; use -F N and -L N to extract both the first and last N entries.

-f N extract first N bp, prints the entire sequence if N is larger than the total length.

-s N extract sequence up to to N bp.

-l N extract last N bp, prints the entire sequence if N is larger than the total length.

Use -s N -f M for sequence from N to M bp; use -f N and -l N to extract both the first and last N bp as one sequence separated by a space.

Note: The -f, -s, and -l options were designed for working with a single long sequences, even though they will work for multiple sequences by applying the same operation to all sequences.


filter

-g N skip sequences with N or more Ns.

-r 1/2 1: skip redundant entry based ID; 2: keep redundant entries by adding a serial number to the identical IDs to make each ID unique.

-R 1/N skip redundant entries based on sequence. 1: use the entire sequence; N: use only the first and last N bases.

-l N skip sequences shorter than N bp.

-L N skip sequences longer than N bp.

use -l N -L M for sequences with length from N to M bp (inclusive).

In all options, '-e' can be added to print the skipped entries in STDERR, which can be captured using 2>[skipped.fa].


report

-f print fasta entries as in the input.

-F print fasta entries with all sequence in one line.

-n print sequences without the defline.

-d print deflines in short form (part before the first space).

-D print deflines in the original form.

-c print the total number of fasta entries in the input.

-l print short defline +[\t] length.

-L print original defline +[\t] length.

-s print sequence summary statistics including N50.

-S print sequence summary statistics plus detailed gap info.
Use -h with -s and -S to disable the header above the outputs
Use -H to print parameters in human friendly form.


search

-s string: search for entries containing "string" in the sequence.

-d string: search for entries containing "string" in the defline: Default is for exact match; use "/string" to search for entries with "string" as part of the ID.

-F file: search for sequences based on a list of IDs in the file (one ID/line).
Can use -D to specify delimiter in the defline. Default is space or '|' or end of line;
use -i to specify the field number, default is 1.

-1 print only the 1st match for -d and -s.

-v use with -s, -d or -F to negate the search.


split

-G N split each of the sequences in the input file as non-gap fragments.
"N" is the number of consecutive Ns base, default is 1;
Use -G N with -t to print just the gap positions.

-n N split the input sequences into chunks, each containing N fasta entries (the last chunk may be less).

-N N split the input sequences into N chunks, each containing equal number of entries (last one may be smaller).

-M N split the input sequences into chunks at ~N MB (million bp) in size (last chunk may be smaller).

-o file: prefix for output files (serial numbers added to prefix; required).


Making fatools executable and available from any directory

Linux

  1. Open the terminal and navigate to where you have downloaded fatools
  2. Find where python2 is installed in your system which python. Usually you would get something like /usr/bin/python
  3. Copy this output to the beggining of fatools as #!/usr/bin/python.
  4. Run the following command to make the script executable chmod +x fatools
  5. Add fatools to your bin directory or any other directories included in your $PATH
  6. You should be able to run fatools from anywhere!

Windows

  1. Type 'control panel' in the Windows search bar
  2. Go to System and Security > System > Advanced System settings > Environment variables
  3. Under system variables, select 'Path' and click 'Edit' and then 'New'
  4. Add the path of where you have fatools located and click OK
  5. You should be able to run fatools from anywhere!

Examples

Navigate to the exampleFiles directory in this repository. In there, a fasta file (exampleFasta.fa) and a file containing a list of IDs (IDlist.txt) from exampleFasta.fa.

To extract the fasta sequences from exampleFasta.fasta based on the list of IDs:
fatools search -F IDlist.txt exampleFasta.fa

NP_001245510.1 notch, isoform B [Drosophila melanogaster]
NNMQSQRSRRRSRAPNTWICFWINKMHAVASLPASLPLLLLTLAFANLPNTVRGTDTALVAASCTSVGCQNG
GTCVTQLNGKTYCACDSHYVGDYCEHRNPCNSMRCQNGGTCQVTFRNGRPGISCKCPLGFDESLCEIAVP
NACDHVTCLNGGTCQLKTLEEYTCACANGYTGERCETKNLCASSPCRNGATCTALAGSSSFTCSCPPGFT... DY343456.1 Macropus rufus BRCA1 (BRCA1) gene, partial cds
CAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAGTCTGGATGAAAGTAAGGAAATATGTAGTGCTGGA
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGTAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCA FY343456.1 Macropus rufus BRCA1 (BRCA1) gene, partial cds
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGCAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAG
TCTGGATGAAAGTAAGGAAATATGTAGTGCTGGAAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCAT 

To report summary statistics:
fatools report -s exampleFasta.fa

Total 2,868 bps from qualified 5 sequences (5 total); length average: 573 (210-1262) bp; N50: 698 bp

To get fasta sequences with a specific maximum length filter

python fatools filter -L250 exampleFasta.fa

DY343456.1
CAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAGTCTGGATGAAAGTAAGGAAATATGTAGTGCTGGA
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGTAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCA
FY343456.1
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGCAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAG
TCTGGATGAAAGTAAGGAAATATGTAGTGCTGGAAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCAT 

You can combine multiple utilities using pipe "|". Let's say you want to see the short defline and the length of the first 3 fasta sequences in a fasta file.

fatools extract -F3 sequenceTesting2.txt | fatools report -l -

AY211956.1Macropus(BRCA1)gene,partialcds 698 NP_001245510.1 1262 DY343456.1 210

Or extract the sequences from a large sequence set for a list of IDs and then search sequences with a specific sequence by using

fatools search -F IDlist.txt exampleFasta.fa |fatools -S AAATAAA -

About

a python tool package for working with fasta sequences

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

59 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

fatools: A python utility package for working with large-scale fasta sequences

A total of 30 utilities/options for common operarations with fasta sequences are currently organized under 6 main subcommands. The utilities range from searching specific sequence entries from a large set of fasta sequences based on ID or a text string in the defline or a sequence motif to spliting a large sequence file into small chuncks, reporting summary stats of a genome assembly, and filtering sequences by length or gap size or redundant entries based on ID or sequences. Furthermore, by allowing input from and output to stdout, multiple processes can be done sequencially in one line of commpands via use of pipe (|). Make sure you have python 2 installed to run fatools.

List of subcommands

Typing 'fatools ' displays the list of subcomands; typing 'fatools ' displays the detailed utilities/options for the subcommand.

convert

-r print sequence in revevrse compliment.

-N convert all non-ACGT letters to N.

-R remove all non-ACGT letters.

-U to upper case

-u to lower case


extract

-F N extract the first N fasta entries, if N is larger than the total number of entries, then print to the last entries.

-S N extract from the Nth entry to the last entry.

-L N extract last N sequence entries. If N is larger than the total number of entries, then print all entries.

Use -S N -F M for entries from N to M; use -F N and -L N to extract both the first and last N entries.

-f N extract first N bp, prints the entire sequence if N is larger than the total length.

-s N extract sequence up to to N bp.

-l N extract last N bp, prints the entire sequence if N is larger than the total length.

Use -s N -f M for sequence from N to M bp; use -f N and -l N to extract both the first and last N bp as one sequence separated by a space.

Note: The -f, -s, and -l options were designed for working with a single long sequences, even though they will work for multiple sequences by applying the same operation to all sequences.


filter

-g N skip sequences with N or more Ns.

-r 1/2 1: skip redundant entry based ID; 2: keep redundant entries by adding a serial number to the identical IDs to make each ID unique.

-R 1/N skip redundant entries based on sequence. 1: use the entire sequence; N: use only the first and last N bases.

-l N skip sequences shorter than N bp.

-L N skip sequences longer than N bp.

use -l N -L M for sequences with length from N to M bp (inclusive).

In all options, '-e' can be added to print the skipped entries in STDERR, which can be captured using 2>[skipped.fa].


report

-f print fasta entries as in the input.

-F print fasta entries with all sequence in one line.

-n print sequences without the defline.

-d print deflines in short form (part before the first space).

-D print deflines in the original form.

-c print the total number of fasta entries in the input.

-l print short defline +[\t] length.

-L print original defline +[\t] length.

-s print sequence summary statistics including N50.

-S print sequence summary statistics plus detailed gap info.
Use -h with -s and -S to disable the header above the outputs
Use -H to print parameters in human friendly form.


search

-s string: search for entries containing "string" in the sequence.

-d string: search for entries containing "string" in the defline: Default is for exact match; use "/string" to search for entries with "string" as part of the ID.

-F file: search for sequences based on a list of IDs in the file (one ID/line).
Can use -D to specify delimiter in the defline. Default is space or '|' or end of line;
use -i to specify the field number, default is 1.

-1 print only the 1st match for -d and -s.

-v use with -s, -d or -F to negate the search.


split

-G N split each of the sequences in the input file as non-gap fragments.
"N" is the number of consecutive Ns base, default is 1;
Use -G N with -t to print just the gap positions.

-n N split the input sequences into chunks, each containing N fasta entries (the last chunk may be less).

-N N split the input sequences into N chunks, each containing equal number of entries (last one may be smaller).

-M N split the input sequences into chunks at ~N MB (million bp) in size (last chunk may be smaller).

-o file: prefix for output files (serial numbers added to prefix; required).


Making fatools executable and available from any directory

Linux

  1. Open the terminal and navigate to where you have downloaded fatools
  2. Find where python2 is installed in your system which python. Usually you would get something like /usr/bin/python
  3. Copy this output to the beggining of fatools as #!/usr/bin/python.
  4. Run the following command to make the script executable chmod +x fatools
  5. Add fatools to your bin directory or any other directories included in your $PATH
  6. You should be able to run fatools from anywhere!

Windows

  1. Type 'control panel' in the Windows search bar
  2. Go to System and Security > System > Advanced System settings > Environment variables
  3. Under system variables, select 'Path' and click 'Edit' and then 'New'
  4. Add the path of where you have fatools located and click OK
  5. You should be able to run fatools from anywhere!

Examples

Navigate to the exampleFiles directory in this repository. In there, a fasta file (exampleFasta.fa) and a file containing a list of IDs (IDlist.txt) from exampleFasta.fa.

To extract the fasta sequences from exampleFasta.fasta based on the list of IDs:
fatools search -F IDlist.txt exampleFasta.fa

NP_001245510.1 notch, isoform B [Drosophila melanogaster]
NNMQSQRSRRRSRAPNTWICFWINKMHAVASLPASLPLLLLTLAFANLPNTVRGTDTALVAASCTSVGCQNG
GTCVTQLNGKTYCACDSHYVGDYCEHRNPCNSMRCQNGGTCQVTFRNGRPGISCKCPLGFDESLCEIAVP
NACDHVTCLNGGTCQLKTLEEYTCACANGYTGERCETKNLCASSPCRNGATCTALAGSSSFTCSCPPGFT... DY343456.1 Macropus rufus BRCA1 (BRCA1) gene, partial cds
CAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAGTCTGGATGAAAGTAAGGAAATATGTAGTGCTGGA
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGTAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCA FY343456.1 Macropus rufus BRCA1 (BRCA1) gene, partial cds
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGCAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAG
TCTGGATGAAAGTAAGGAAATATGTAGTGCTGGAAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCAT 

To report summary statistics:
fatools report -s exampleFasta.fa

Total 2,868 bps from qualified 5 sequences (5 total); length average: 573 (210-1262) bp; N50: 698 bp

To get fasta sequences with a specific maximum length filter

python fatools filter -L250 exampleFasta.fa

DY343456.1
CAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAGTCTGGATGAAAGTAAGGAAATATGTAGTGCTGGA
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGTAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCA
FY343456.1
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGCAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAG
TCTGGATGAAAGTAAGGAAATATGTAGTGCTGGAAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCAT 

You can combine multiple utilities using pipe "|". Let's say you want to see the short defline and the length of the first 3 fasta sequences in a fasta file.

fatools extract -F3 sequenceTesting2.txt | fatools report -l -

AY211956.1Macropus(BRCA1)gene,partialcds 698 NP_001245510.1 1262 DY343456.1 210

Or extract the sequences from a large sequence set for a list of IDs and then search sequences with a specific sequence by using

fatools search -F IDlist.txt exampleFasta.fa |fatools -S AAATAAA -

About

a python tool package for working with fasta sequences

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

59 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

fatools: A python utility package for working with large-scale fasta sequences

A total of 30 utilities/options for common operarations with fasta sequences are currently organized under 6 main subcommands. The utilities range from searching specific sequence entries from a large set of fasta sequences based on ID or a text string in the defline or a sequence motif to spliting a large sequence file into small chuncks, reporting summary stats of a genome assembly, and filtering sequences by length or gap size or redundant entries based on ID or sequences. Furthermore, by allowing input from and output to stdout, multiple processes can be done sequencially in one line of commpands via use of pipe (|). Make sure you have python 2 installed to run fatools.

List of subcommands

Typing 'fatools ' displays the list of subcomands; typing 'fatools ' displays the detailed utilities/options for the subcommand.

convert

-r print sequence in revevrse compliment.

-N convert all non-ACGT letters to N.

-R remove all non-ACGT letters.

-U to upper case

-u to lower case


extract

-F N extract the first N fasta entries, if N is larger than the total number of entries, then print to the last entries.

-S N extract from the Nth entry to the last entry.

-L N extract last N sequence entries. If N is larger than the total number of entries, then print all entries.

Use -S N -F M for entries from N to M; use -F N and -L N to extract both the first and last N entries.

-f N extract first N bp, prints the entire sequence if N is larger than the total length.

-s N extract sequence up to to N bp.

-l N extract last N bp, prints the entire sequence if N is larger than the total length.

Use -s N -f M for sequence from N to M bp; use -f N and -l N to extract both the first and last N bp as one sequence separated by a space.

Note: The -f, -s, and -l options were designed for working with a single long sequences, even though they will work for multiple sequences by applying the same operation to all sequences.


filter

-g N skip sequences with N or more Ns.

-r 1/2 1: skip redundant entry based ID; 2: keep redundant entries by adding a serial number to the identical IDs to make each ID unique.

-R 1/N skip redundant entries based on sequence. 1: use the entire sequence; N: use only the first and last N bases.

-l N skip sequences shorter than N bp.

-L N skip sequences longer than N bp.

use -l N -L M for sequences with length from N to M bp (inclusive).

In all options, '-e' can be added to print the skipped entries in STDERR, which can be captured using 2>[skipped.fa].


report

-f print fasta entries as in the input.

-F print fasta entries with all sequence in one line.

-n print sequences without the defline.

-d print deflines in short form (part before the first space).

-D print deflines in the original form.

-c print the total number of fasta entries in the input.

-l print short defline +[\t] length.

-L print original defline +[\t] length.

-s print sequence summary statistics including N50.

-S print sequence summary statistics plus detailed gap info.
Use -h with -s and -S to disable the header above the outputs
Use -H to print parameters in human friendly form.


search

-s string: search for entries containing "string" in the sequence.

-d string: search for entries containing "string" in the defline: Default is for exact match; use "/string" to search for entries with "string" as part of the ID.

-F file: search for sequences based on a list of IDs in the file (one ID/line).
Can use -D to specify delimiter in the defline. Default is space or '|' or end of line;
use -i to specify the field number, default is 1.

-1 print only the 1st match for -d and -s.

-v use with -s, -d or -F to negate the search.


split

-G N split each of the sequences in the input file as non-gap fragments.
"N" is the number of consecutive Ns base, default is 1;
Use -G N with -t to print just the gap positions.

-n N split the input sequences into chunks, each containing N fasta entries (the last chunk may be less).

-N N split the input sequences into N chunks, each containing equal number of entries (last one may be smaller).

-M N split the input sequences into chunks at ~N MB (million bp) in size (last chunk may be smaller).

-o file: prefix for output files (serial numbers added to prefix; required).


Making fatools executable and available from any directory

Linux

  1. Open the terminal and navigate to where you have downloaded fatools
  2. Find where python2 is installed in your system which python. Usually you would get something like /usr/bin/python
  3. Copy this output to the beggining of fatools as #!/usr/bin/python.
  4. Run the following command to make the script executable chmod +x fatools
  5. Add fatools to your bin directory or any other directories included in your $PATH
  6. You should be able to run fatools from anywhere!

Windows

  1. Type 'control panel' in the Windows search bar
  2. Go to System and Security > System > Advanced System settings > Environment variables
  3. Under system variables, select 'Path' and click 'Edit' and then 'New'
  4. Add the path of where you have fatools located and click OK
  5. You should be able to run fatools from anywhere!

Examples

Navigate to the exampleFiles directory in this repository. In there, a fasta file (exampleFasta.fa) and a file containing a list of IDs (IDlist.txt) from exampleFasta.fa.

To extract the fasta sequences from exampleFasta.fasta based on the list of IDs:
fatools search -F IDlist.txt exampleFasta.fa

NP_001245510.1 notch, isoform B [Drosophila melanogaster]
NNMQSQRSRRRSRAPNTWICFWINKMHAVASLPASLPLLLLTLAFANLPNTVRGTDTALVAASCTSVGCQNG
GTCVTQLNGKTYCACDSHYVGDYCEHRNPCNSMRCQNGGTCQVTFRNGRPGISCKCPLGFDESLCEIAVP
NACDHVTCLNGGTCQLKTLEEYTCACANGYTGERCETKNLCASSPCRNGATCTALAGSSSFTCSCPPGFT... DY343456.1 Macropus rufus BRCA1 (BRCA1) gene, partial cds
CAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAGTCTGGATGAAAGTAAGGAAATATGTAGTGCTGGA
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGTAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCA FY343456.1 Macropus rufus BRCA1 (BRCA1) gene, partial cds
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGCAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAG
TCTGGATGAAAGTAAGGAAATATGTAGTGCTGGAAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCAT 

To report summary statistics:
fatools report -s exampleFasta.fa

Total 2,868 bps from qualified 5 sequences (5 total); length average: 573 (210-1262) bp; N50: 698 bp

To get fasta sequences with a specific maximum length filter

python fatools filter -L250 exampleFasta.fa

DY343456.1
CAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAGTCTGGATGAAAGTAAGGAAATATGTAGTGCTGGA
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGTAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCA
FY343456.1
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGCAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAG
TCTGGATGAAAGTAAGGAAATATGTAGTGCTGGAAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCAT 

You can combine multiple utilities using pipe "|". Let's say you want to see the short defline and the length of the first 3 fasta sequences in a fasta file.

fatools extract -F3 sequenceTesting2.txt | fatools report -l -

AY211956.1Macropus(BRCA1)gene,partialcds 698 NP_001245510.1 1262 DY343456.1 210

Or extract the sequences from a large sequence set for a list of IDs and then search sequences with a specific sequence by using

fatools search -F IDlist.txt exampleFasta.fa |fatools -S AAATAAA -

About

a python tool package for working with fasta sequences

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

59 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

fatools: A python utility package for working with large-scale fasta sequences

A total of 30 utilities/options for common operarations with fasta sequences are currently organized under 6 main subcommands. The utilities range from searching specific sequence entries from a large set of fasta sequences based on ID or a text string in the defline or a sequence motif to spliting a large sequence file into small chuncks, reporting summary stats of a genome assembly, and filtering sequences by length or gap size or redundant entries based on ID or sequences. Furthermore, by allowing input from and output to stdout, multiple processes can be done sequencially in one line of commpands via use of pipe (|). Make sure you have python 2 installed to run fatools.

List of subcommands

Typing 'fatools ' displays the list of subcomands; typing 'fatools ' displays the detailed utilities/options for the subcommand.

convert

-r print sequence in revevrse compliment.

-N convert all non-ACGT letters to N.

-R remove all non-ACGT letters.

-U to upper case

-u to lower case


extract

-F N extract the first N fasta entries, if N is larger than the total number of entries, then print to the last entries.

-S N extract from the Nth entry to the last entry.

-L N extract last N sequence entries. If N is larger than the total number of entries, then print all entries.

Use -S N -F M for entries from N to M; use -F N and -L N to extract both the first and last N entries.

-f N extract first N bp, prints the entire sequence if N is larger than the total length.

-s N extract sequence up to to N bp.

-l N extract last N bp, prints the entire sequence if N is larger than the total length.

Use -s N -f M for sequence from N to M bp; use -f N and -l N to extract both the first and last N bp as one sequence separated by a space.

Note: The -f, -s, and -l options were designed for working with a single long sequences, even though they will work for multiple sequences by applying the same operation to all sequences.


filter

-g N skip sequences with N or more Ns.

-r 1/2 1: skip redundant entry based ID; 2: keep redundant entries by adding a serial number to the identical IDs to make each ID unique.

-R 1/N skip redundant entries based on sequence. 1: use the entire sequence; N: use only the first and last N bases.

-l N skip sequences shorter than N bp.

-L N skip sequences longer than N bp.

use -l N -L M for sequences with length from N to M bp (inclusive).

In all options, '-e' can be added to print the skipped entries in STDERR, which can be captured using 2>[skipped.fa].


report

-f print fasta entries as in the input.

-F print fasta entries with all sequence in one line.

-n print sequences without the defline.

-d print deflines in short form (part before the first space).

-D print deflines in the original form.

-c print the total number of fasta entries in the input.

-l print short defline +[\t] length.

-L print original defline +[\t] length.

-s print sequence summary statistics including N50.

-S print sequence summary statistics plus detailed gap info.
Use -h with -s and -S to disable the header above the outputs
Use -H to print parameters in human friendly form.


search

-s string: search for entries containing "string" in the sequence.

-d string: search for entries containing "string" in the defline: Default is for exact match; use "/string" to search for entries with "string" as part of the ID.

-F file: search for sequences based on a list of IDs in the file (one ID/line).
Can use -D to specify delimiter in the defline. Default is space or '|' or end of line;
use -i to specify the field number, default is 1.

-1 print only the 1st match for -d and -s.

-v use with -s, -d or -F to negate the search.


split

-G N split each of the sequences in the input file as non-gap fragments.
"N" is the number of consecutive Ns base, default is 1;
Use -G N with -t to print just the gap positions.

-n N split the input sequences into chunks, each containing N fasta entries (the last chunk may be less).

-N N split the input sequences into N chunks, each containing equal number of entries (last one may be smaller).

-M N split the input sequences into chunks at ~N MB (million bp) in size (last chunk may be smaller).

-o file: prefix for output files (serial numbers added to prefix; required).


Making fatools executable and available from any directory

Linux

  1. Open the terminal and navigate to where you have downloaded fatools
  2. Find where python2 is installed in your system which python. Usually you would get something like /usr/bin/python
  3. Copy this output to the beggining of fatools as #!/usr/bin/python.
  4. Run the following command to make the script executable chmod +x fatools
  5. Add fatools to your bin directory or any other directories included in your $PATH
  6. You should be able to run fatools from anywhere!

Windows

  1. Type 'control panel' in the Windows search bar
  2. Go to System and Security > System > Advanced System settings > Environment variables
  3. Under system variables, select 'Path' and click 'Edit' and then 'New'
  4. Add the path of where you have fatools located and click OK
  5. You should be able to run fatools from anywhere!

Examples

Navigate to the exampleFiles directory in this repository. In there, a fasta file (exampleFasta.fa) and a file containing a list of IDs (IDlist.txt) from exampleFasta.fa.

To extract the fasta sequences from exampleFasta.fasta based on the list of IDs:
fatools search -F IDlist.txt exampleFasta.fa

NP_001245510.1 notch, isoform B [Drosophila melanogaster]
NNMQSQRSRRRSRAPNTWICFWINKMHAVASLPASLPLLLLTLAFANLPNTVRGTDTALVAASCTSVGCQNG
GTCVTQLNGKTYCACDSHYVGDYCEHRNPCNSMRCQNGGTCQVTFRNGRPGISCKCPLGFDESLCEIAVP
NACDHVTCLNGGTCQLKTLEEYTCACANGYTGERCETKNLCASSPCRNGATCTALAGSSSFTCSCPPGFT... DY343456.1 Macropus rufus BRCA1 (BRCA1) gene, partial cds
CAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAGTCTGGATGAAAGTAAGGAAATATGTAGTGCTGGA
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGTAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCA FY343456.1 Macropus rufus BRCA1 (BRCA1) gene, partial cds
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGCAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAG
TCTGGATGAAAGTAAGGAAATATGTAGTGCTGGAAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCAT 

To report summary statistics:
fatools report -s exampleFasta.fa

Total 2,868 bps from qualified 5 sequences (5 total); length average: 573 (210-1262) bp; N50: 698 bp

To get fasta sequences with a specific maximum length filter

python fatools filter -L250 exampleFasta.fa

DY343456.1
CAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAGTCTGGATGAAAGTAAGGAAATATGTAGTGCTGGA
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGTAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCA
FY343456.1
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGCAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAG
TCTGGATGAAAGTAAGGAAATATGTAGTGCTGGAAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCAT 

You can combine multiple utilities using pipe "|". Let's say you want to see the short defline and the length of the first 3 fasta sequences in a fasta file.

fatools extract -F3 sequenceTesting2.txt | fatools report -l -

AY211956.1Macropus(BRCA1)gene,partialcds 698 NP_001245510.1 1262 DY343456.1 210

Or extract the sequences from a large sequence set for a list of IDs and then search sequences with a specific sequence by using

fatools search -F IDlist.txt exampleFasta.fa |fatools -S AAATAAA -

About

a python tool package for working with fasta sequences

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

59 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

fatools: A python utility package for working with large-scale fasta sequences

A total of 30 utilities/options for common operarations with fasta sequences are currently organized under 6 main subcommands. The utilities range from searching specific sequence entries from a large set of fasta sequences based on ID or a text string in the defline or a sequence motif to spliting a large sequence file into small chuncks, reporting summary stats of a genome assembly, and filtering sequences by length or gap size or redundant entries based on ID or sequences. Furthermore, by allowing input from and output to stdout, multiple processes can be done sequencially in one line of commpands via use of pipe (|). Make sure you have python 2 installed to run fatools.

List of subcommands

Typing 'fatools ' displays the list of subcomands; typing 'fatools ' displays the detailed utilities/options for the subcommand.

convert

-r print sequence in revevrse compliment.

-N convert all non-ACGT letters to N.

-R remove all non-ACGT letters.

-U to upper case

-u to lower case


extract

-F N extract the first N fasta entries, if N is larger than the total number of entries, then print to the last entries.

-S N extract from the Nth entry to the last entry.

-L N extract last N sequence entries. If N is larger than the total number of entries, then print all entries.

Use -S N -F M for entries from N to M; use -F N and -L N to extract both the first and last N entries.

-f N extract first N bp, prints the entire sequence if N is larger than the total length.

-s N extract sequence up to to N bp.

-l N extract last N bp, prints the entire sequence if N is larger than the total length.

Use -s N -f M for sequence from N to M bp; use -f N and -l N to extract both the first and last N bp as one sequence separated by a space.

Note: The -f, -s, and -l options were designed for working with a single long sequences, even though they will work for multiple sequences by applying the same operation to all sequences.


filter

-g N skip sequences with N or more Ns.

-r 1/2 1: skip redundant entry based ID; 2: keep redundant entries by adding a serial number to the identical IDs to make each ID unique.

-R 1/N skip redundant entries based on sequence. 1: use the entire sequence; N: use only the first and last N bases.

-l N skip sequences shorter than N bp.

-L N skip sequences longer than N bp.

use -l N -L M for sequences with length from N to M bp (inclusive).

In all options, '-e' can be added to print the skipped entries in STDERR, which can be captured using 2>[skipped.fa].


report

-f print fasta entries as in the input.

-F print fasta entries with all sequence in one line.

-n print sequences without the defline.

-d print deflines in short form (part before the first space).

-D print deflines in the original form.

-c print the total number of fasta entries in the input.

-l print short defline +[\t] length.

-L print original defline +[\t] length.

-s print sequence summary statistics including N50.

-S print sequence summary statistics plus detailed gap info.
Use -h with -s and -S to disable the header above the outputs
Use -H to print parameters in human friendly form.


search

-s string: search for entries containing "string" in the sequence.

-d string: search for entries containing "string" in the defline: Default is for exact match; use "/string" to search for entries with "string" as part of the ID.

-F file: search for sequences based on a list of IDs in the file (one ID/line).
Can use -D to specify delimiter in the defline. Default is space or '|' or end of line;
use -i to specify the field number, default is 1.

-1 print only the 1st match for -d and -s.

-v use with -s, -d or -F to negate the search.


split

-G N split each of the sequences in the input file as non-gap fragments.
"N" is the number of consecutive Ns base, default is 1;
Use -G N with -t to print just the gap positions.

-n N split the input sequences into chunks, each containing N fasta entries (the last chunk may be less).

-N N split the input sequences into N chunks, each containing equal number of entries (last one may be smaller).

-M N split the input sequences into chunks at ~N MB (million bp) in size (last chunk may be smaller).

-o file: prefix for output files (serial numbers added to prefix; required).


Making fatools executable and available from any directory

Linux

  1. Open the terminal and navigate to where you have downloaded fatools
  2. Find where python2 is installed in your system which python. Usually you would get something like /usr/bin/python
  3. Copy this output to the beggining of fatools as #!/usr/bin/python.
  4. Run the following command to make the script executable chmod +x fatools
  5. Add fatools to your bin directory or any other directories included in your $PATH
  6. You should be able to run fatools from anywhere!

Windows

  1. Type 'control panel' in the Windows search bar
  2. Go to System and Security > System > Advanced System settings > Environment variables
  3. Under system variables, select 'Path' and click 'Edit' and then 'New'
  4. Add the path of where you have fatools located and click OK
  5. You should be able to run fatools from anywhere!

Examples

Navigate to the exampleFiles directory in this repository. In there, a fasta file (exampleFasta.fa) and a file containing a list of IDs (IDlist.txt) from exampleFasta.fa.

To extract the fasta sequences from exampleFasta.fasta based on the list of IDs:
fatools search -F IDlist.txt exampleFasta.fa

NP_001245510.1 notch, isoform B [Drosophila melanogaster]
NNMQSQRSRRRSRAPNTWICFWINKMHAVASLPASLPLLLLTLAFANLPNTVRGTDTALVAASCTSVGCQNG
GTCVTQLNGKTYCACDSHYVGDYCEHRNPCNSMRCQNGGTCQVTFRNGRPGISCKCPLGFDESLCEIAVP
NACDHVTCLNGGTCQLKTLEEYTCACANGYTGERCETKNLCASSPCRNGATCTALAGSSSFTCSCPPGFT... DY343456.1 Macropus rufus BRCA1 (BRCA1) gene, partial cds
CAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAGTCTGGATGAAAGTAAGGAAATATGTAGTGCTGGA
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGTAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCA FY343456.1 Macropus rufus BRCA1 (BRCA1) gene, partial cds
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGCAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAG
TCTGGATGAAAGTAAGGAAATATGTAGTGCTGGAAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCAT 

To report summary statistics:
fatools report -s exampleFasta.fa

Total 2,868 bps from qualified 5 sequences (5 total); length average: 573 (210-1262) bp; N50: 698 bp

To get fasta sequences with a specific maximum length filter

python fatools filter -L250 exampleFasta.fa

DY343456.1
CAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAGTCTGGATGAAAGTAAGGAAATATGTAGTGCTGGA
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGTAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCA
FY343456.1
TGTGGCACAGATGCTCGTGCCACCTCATTACTTCCTGAAACCACCAGCTTATCGCCCAACACAGACCGAA
TGAATGTAGAAAAGGCTGAACTCTGTAATAAAAGCAAACAGCCTGGCTTAGCAAAAAACCAACAGAGCAG
TCTGGATGAAAGTAAGGAAATATGTAGTGCTGGAAAGACCCTGGGTGCCCATGAGCTGAATGCCCATCAT 

You can combine multiple utilities using pipe "|". Let's say you want to see the short defline and the length of the first 3 fasta sequences in a fasta file.

fatools extract -F3 sequenceTesting2.txt | fatools report -l -

AY211956.1Macropus(BRCA1)gene,partialcds 698 NP_001245510.1 1262 DY343456.1 210

Or extract the sequences from a large sequence set for a list of IDs and then search sequences with a specific sequence by using

fatools search -F IDlist.txt exampleFasta.fa |fatools -S AAATAAA -

About

a python tool package for working with fasta sequences

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages