Skip to content

Repository files navigation

FastaParser

python versionsbuild statuscode style: blackcoveragedocs statuslicense

pypipypi downloads

A Python FASTA file Parser and Writer.

The FASTA file format is a standard text-based format for representing nucleotide and aminoacid sequences (usual file extensions include: .fasta, .fna, .ffn, .faa and .frn). FastaParser is able to parse such files and extract the biological sequences within into Python objects. It can also handle and manipulate such sequences as well as write sequences to new or existing FASTA files.

Installation

With pip:

$ pip install fastaparser

Usage

Read FASTA files

Generate python objects from FASTA files:

>>>importfastaparser>>>withopen("fasta_file.fasta") asfasta_file:
parser=fastaparser.Reader(fasta_file)
forseqinparser:
# seq is a FastaSequence objectprint('ID:', seq.id)
print('Description:', seq.description)
print('Sequence:', seq.sequence_as_string())
print()

output:

ID: sp|P04439|HLAA_HUMAN
Description: HLA class I histocompatibility antigen, A alpha chain OS=Homo sapi...
Sequence: MAVMAPRTLLLLLSGALALTQTWAGSHSMRYFFTSVSRPGRGEPRFIAVGYVDDTQFVRFDSDAASQRM...
ID: sp|P15822|ZEP1_HUMAN
Description: Zinc finger protein 40 OS=Homo sapiens OX=9606 GN=HIVEP1 PE=1 SV=3...
Sequence: MPRTKQIHPRNLRDKIEEAQKELNGAEVSKKEILQAGVKGTSESLKGVKRKKIVAENHLKKIPKSPLRN...

or just parse FASTA headers and sequences, which is much faster but less feature rich:

>>>importfastaparser>>>withopen("fasta_file.fasta") asfasta_file:
parser=fastaparser.Reader(fasta_file, parse_method='quick')
forseqinparser:
# seq is a namedtuple('Fasta', ['header', 'sequence'])print('Header:', seq.header)
print('Sequence:', seq.sequence)
print()

output:

Header: >sp|P04439|HLAA_HUMAN HLA class I histocompatibility antigen, A alpha c...
Sequence: MAVMAPRTLLLLLSGALALTQTWAGSHSMRYFFTSVSRPGRGEPRFIAVGYVDDTQFVRFDSDAASQRM...
Header: >sp|P15822|ZEP1_HUMAN Zinc finger protein 40 OS=Homo sapiens OX=9606 GN...
Sequence: MPRTKQIHPRNLRDKIEEAQKELNGAEVSKKEILQAGVKGTSESLKGVKRKKIVAENHLKKIPKSPLRN...

Write FASTA files

Create FASTA files from FastaSequence objects:

>>>importfastaparser>>>withopen("fasta_file.fasta", 'w') asfasta_file:
writer=fastaparser.Writer(fasta_file)
fasta_sequence=fastaparser.FastaSequence(
sequence='ACTGCTGCTAGCTAGC',
id='id123',
description='test sequence'
)
writer.writefasta(fasta_sequence)

or single header and sequence strings:

>>>importfastaparser>>>withopen("fasta_file.fasta", 'w') asfasta_file:
writer=fastaparser.Writer(fasta_file)
writer.writefasta(('id123 test sequence', 'ACTGCTGCTAGCTAGC'))

Documentation

Documentation for FastaParser is available here: https://fastaparser.readthedocs.io/en/latest

Releases

Sponsor this project

Used by

Contributors

Languages