Skip to content

Repository files navigation

pyTMHMM

pyTMHMM is a Python 3.5+/Cython implementation of the transmembrane helix predictor using a hidden Markov model (TMHMM) originally described in:

E.L. Sonnhammer, G. von Heijne, and A. Krogh. A hidden Markov model for predicting transmembrane helices in protein sequences. In J. Glasgow, T. Littlejohn, F. Major, R. Lathrop, D. Sankoff, and C. Sensen, editors, Proceedings of the Sixth International Conference on Intelligent Systems for Molecular Biology, pages 175-182, Menlo Park, CA, 1998. AAAI Press. PMID 9783223

History

Dan Søndergaard is the original author of this package and his repository is now archived. Dan wrote this code for a few reasons:

  • the source code is not available as part of the publication
  • the downloadable binaries are Linux-only
  • the downloadable binaries may not be redistributed, so it's not possible to put them in a Docker image or a VM for other people to use
  • there is a need to predict transmembrane helices in a scripted, automated way

This Python implementation includes a parser for the undocumented file format used to describe the model and a fast Cython implementation of the Viterbi algorithm used to perform the annotation. The tool will output files similar to the files produced by the original TMHMM implementation.

Incompatibilities

  • The original TMHMM implementation handles ambigious characters and gaps in an undocumented way. However, pyTMHMM does not attempt to handle such characters at all and will fail. A possible fix is to replace those characters with something also based on expert/domain knowledge.

Installation

This package supports Python 3.5 or greater. Install with:

> pip install pyTMHMM

Command Line Usage

> pyTMHMM -h
usage: pyTMHMM [-h] -f SEQUENCE_FILE [-m MODEL_FILE] [-p]
required arguments:
-f SEQUENCE_FILE, --file SEQUENCE_FILE
path to file in fasta format with sequences
optional arguments:
-h, --help show this help message and exit
-m MODEL_FILE, --model MODEL_FILE
path to the model to use (default: TMHMM2.0.model)
-p, --plot plot posterior probabilies

The -p/--plot option requires matplotlib.

The input sequence file should have one or more sequences in Fasta format, for example:

> head PAR3_HUMAN.fasta
>sp|O00254|PAR3_HUMAN Proteinase-activated receptor 3 OS=Homo sapiens OX=9606 GN=F2RL2 PE=1 SV=1
MKALIFAAAGLLLLLPTFCQSGMENDTNNLAKPTLPIKTFRGAPPNSFEEFPFSALEGWT
GATITVKIKCPEESASHLHVKNATMGYLTSSLSTKLIPAIYLLVFVVGVPANAVTLWMLF
FRTRSICTTVFYTNLAIADFLFCVTLPFKIAYHLNGNNWVFGEVLCRATTVIFYGNMYCS
ILLLACISINRYLAIVHPFTYRGLPKHTYALVTCGLVWATVFLYMLPFFILKQEYYLVQP
DITTCHDVHNTCESSSPFQLYYFISLAFFGFLIPFVLIIYCYAAIIRTLNAYDHRWLWYV
KASLLILVIFTICFAPSNIILIIHHANYYYNNTDGLYFIYLIALCLGSLNSCLDPFLYFL
MSKTRNHSTAYLTK

Example command:

> pyTMHMM -f PAR3_HUMAN.fasta

This produces three files for each sequence in the Fasta file, named by id.

Summary file

The coordinates of the predicted domains:

> cat sp|O00254|PAR3_HUMAN.summary 0 97 outside
98 120 transmembrane helix
121 128 inside
129 151 transmembrane helix
152 165 outside
166 188 transmembrane helix
189 207 inside
208 230 transmembrane helix
231 259 outside
260 282 transmembrane helix
283 302 inside
303 322 transmembrane helix
323 336 outside
337 359 transmembrane helix
360 373 inside

Annotation file

An annotated sequence in Fasta-like format:

> cat sp|O00254|PAR3_HUMAN.annotation >sp|O00254|PAR3_HUMAN Proteinase-activated receptor 3 OS=Homo sapiens OX=9606 GN=F2RL2 PE=1 SV=1
OOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOO
OOOOOOOOOOOOOOOOOOOMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiMMMMMMMMMMMMMMMMMMMMMMMoooooo
ooooooooMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiiiiiiiiiiiiMMMMMMMMMMMMMMMMMMMMMMMoooooo
oooooooooooooooooooooooMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiiiiiiiiiiiiiMMMMMMMMMMMMM
MMMMMMMooooooooooooooMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiiiiiii

Posterior probabilities file

A file containing the posterior probabilities for each label:

> head sp|O00254|PAR3_HUMAN.plot inside membrane outside
0.6417636608794935 0.0 0.3582363391205064
0.693933311909457 0.006819179965744769 0.2992475081247982
0.3041488405999551 0.36045181385397806 0.3353993455460668
0.15867304975718463 0.5320740444690139 0.3092529057738015
0.011878169861623369 0.8126781067794638 0.1754437233589128
0.009103844612501565 0.7722962064006578 0.21859994898684057
0.0008287471596339259 0.6966223976666195 0.3025488551737467
0.0007860447761827514 0.7122010989508554 0.2870128562729619
0.0006349307902653272 0.712364526792757 0.28700054241697776

Optional plot file

If the -p flag is set and matplotlib is installed a plot in PDF format is made:

"TM domains in PAR3_HUMAN"

doc/sp|O00254|PAR3_HUMAN.pdf

API Usage

You can also use pyTMHMM as a library:

import pyTMHMM
annotation, posterior = pyTMHMM.predict(sequence_string)

This returns the annotation as a string and the posterior probabilities for each label as a numpy array with shape (len(sequence), 3) where column 0, 1 and 2 corresponds to being inside, transmembrane and outside, respectively.

If you don't need the posterior probabilities set compute_posterior=False, this will save computation:

annotation = pyTMHMM.predict(
sequence_string, compute_posterior=False
)

About

Python 3.5+ implementation of TMHMM (TransMembrane helix Hidden Markov Model)

Resources

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - bosborne/pyTMHMM: Python 3.5+ implementation of TMHMM (TransMembrane helix Hidden Markov Model) · GitHub
Skip to content

Repository files navigation

pyTMHMM

pyTMHMM is a Python 3.5+/Cython implementation of the transmembrane helix predictor using a hidden Markov model (TMHMM) originally described in:

E.L. Sonnhammer, G. von Heijne, and A. Krogh. A hidden Markov model for predicting transmembrane helices in protein sequences. In J. Glasgow, T. Littlejohn, F. Major, R. Lathrop, D. Sankoff, and C. Sensen, editors, Proceedings of the Sixth International Conference on Intelligent Systems for Molecular Biology, pages 175-182, Menlo Park, CA, 1998. AAAI Press. PMID 9783223

History

Dan Søndergaard is the original author of this package and his repository is now archived. Dan wrote this code for a few reasons:

  • the source code is not available as part of the publication
  • the downloadable binaries are Linux-only
  • the downloadable binaries may not be redistributed, so it's not possible to put them in a Docker image or a VM for other people to use
  • there is a need to predict transmembrane helices in a scripted, automated way

This Python implementation includes a parser for the undocumented file format used to describe the model and a fast Cython implementation of the Viterbi algorithm used to perform the annotation. The tool will output files similar to the files produced by the original TMHMM implementation.

Incompatibilities

  • The original TMHMM implementation handles ambigious characters and gaps in an undocumented way. However, pyTMHMM does not attempt to handle such characters at all and will fail. A possible fix is to replace those characters with something also based on expert/domain knowledge.

Installation

This package supports Python 3.5 or greater. Install with:

> pip install pyTMHMM

Command Line Usage

> pyTMHMM -h
usage: pyTMHMM [-h] -f SEQUENCE_FILE [-m MODEL_FILE] [-p]
required arguments:
-f SEQUENCE_FILE, --file SEQUENCE_FILE
path to file in fasta format with sequences
optional arguments:
-h, --help show this help message and exit
-m MODEL_FILE, --model MODEL_FILE
path to the model to use (default: TMHMM2.0.model)
-p, --plot plot posterior probabilies

The -p/--plot option requires matplotlib.

The input sequence file should have one or more sequences in Fasta format, for example:

> head PAR3_HUMAN.fasta
>sp|O00254|PAR3_HUMAN Proteinase-activated receptor 3 OS=Homo sapiens OX=9606 GN=F2RL2 PE=1 SV=1
MKALIFAAAGLLLLLPTFCQSGMENDTNNLAKPTLPIKTFRGAPPNSFEEFPFSALEGWT
GATITVKIKCPEESASHLHVKNATMGYLTSSLSTKLIPAIYLLVFVVGVPANAVTLWMLF
FRTRSICTTVFYTNLAIADFLFCVTLPFKIAYHLNGNNWVFGEVLCRATTVIFYGNMYCS
ILLLACISINRYLAIVHPFTYRGLPKHTYALVTCGLVWATVFLYMLPFFILKQEYYLVQP
DITTCHDVHNTCESSSPFQLYYFISLAFFGFLIPFVLIIYCYAAIIRTLNAYDHRWLWYV
KASLLILVIFTICFAPSNIILIIHHANYYYNNTDGLYFIYLIALCLGSLNSCLDPFLYFL
MSKTRNHSTAYLTK

Example command:

> pyTMHMM -f PAR3_HUMAN.fasta

This produces three files for each sequence in the Fasta file, named by id.

Summary file

The coordinates of the predicted domains:

> cat sp|O00254|PAR3_HUMAN.summary 0 97 outside
98 120 transmembrane helix
121 128 inside
129 151 transmembrane helix
152 165 outside
166 188 transmembrane helix
189 207 inside
208 230 transmembrane helix
231 259 outside
260 282 transmembrane helix
283 302 inside
303 322 transmembrane helix
323 336 outside
337 359 transmembrane helix
360 373 inside

Annotation file

An annotated sequence in Fasta-like format:

> cat sp|O00254|PAR3_HUMAN.annotation >sp|O00254|PAR3_HUMAN Proteinase-activated receptor 3 OS=Homo sapiens OX=9606 GN=F2RL2 PE=1 SV=1
OOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOO
OOOOOOOOOOOOOOOOOOOMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiMMMMMMMMMMMMMMMMMMMMMMMoooooo
ooooooooMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiiiiiiiiiiiiMMMMMMMMMMMMMMMMMMMMMMMoooooo
oooooooooooooooooooooooMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiiiiiiiiiiiiiMMMMMMMMMMMMM
MMMMMMMooooooooooooooMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiiiiiii

Posterior probabilities file

A file containing the posterior probabilities for each label:

> head sp|O00254|PAR3_HUMAN.plot inside membrane outside
0.6417636608794935 0.0 0.3582363391205064
0.693933311909457 0.006819179965744769 0.2992475081247982
0.3041488405999551 0.36045181385397806 0.3353993455460668
0.15867304975718463 0.5320740444690139 0.3092529057738015
0.011878169861623369 0.8126781067794638 0.1754437233589128
0.009103844612501565 0.7722962064006578 0.21859994898684057
0.0008287471596339259 0.6966223976666195 0.3025488551737467
0.0007860447761827514 0.7122010989508554 0.2870128562729619
0.0006349307902653272 0.712364526792757 0.28700054241697776

Optional plot file

If the -p flag is set and matplotlib is installed a plot in PDF format is made:

"TM domains in PAR3_HUMAN"

doc/sp|O00254|PAR3_HUMAN.pdf

API Usage

You can also use pyTMHMM as a library:

import pyTMHMM
annotation, posterior = pyTMHMM.predict(sequence_string)

This returns the annotation as a string and the posterior probabilities for each label as a numpy array with shape (len(sequence), 3) where column 0, 1 and 2 corresponds to being inside, transmembrane and outside, respectively.

If you don't need the posterior probabilities set compute_posterior=False, this will save computation:

annotation = pyTMHMM.predict(
sequence_string, compute_posterior=False
)

About

Python 3.5+ implementation of TMHMM (TransMembrane helix Hidden Markov Model)

Resources

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - bosborne/pyTMHMM: Python 3.5+ implementation of TMHMM (TransMembrane helix Hidden Markov Model) · GitHub
Skip to content

Repository files navigation

pyTMHMM

pyTMHMM is a Python 3.5+/Cython implementation of the transmembrane helix predictor using a hidden Markov model (TMHMM) originally described in:

E.L. Sonnhammer, G. von Heijne, and A. Krogh. A hidden Markov model for predicting transmembrane helices in protein sequences. In J. Glasgow, T. Littlejohn, F. Major, R. Lathrop, D. Sankoff, and C. Sensen, editors, Proceedings of the Sixth International Conference on Intelligent Systems for Molecular Biology, pages 175-182, Menlo Park, CA, 1998. AAAI Press. PMID 9783223

History

Dan Søndergaard is the original author of this package and his repository is now archived. Dan wrote this code for a few reasons:

  • the source code is not available as part of the publication
  • the downloadable binaries are Linux-only
  • the downloadable binaries may not be redistributed, so it's not possible to put them in a Docker image or a VM for other people to use
  • there is a need to predict transmembrane helices in a scripted, automated way

This Python implementation includes a parser for the undocumented file format used to describe the model and a fast Cython implementation of the Viterbi algorithm used to perform the annotation. The tool will output files similar to the files produced by the original TMHMM implementation.

Incompatibilities

  • The original TMHMM implementation handles ambigious characters and gaps in an undocumented way. However, pyTMHMM does not attempt to handle such characters at all and will fail. A possible fix is to replace those characters with something also based on expert/domain knowledge.

Installation

This package supports Python 3.5 or greater. Install with:

> pip install pyTMHMM

Command Line Usage

> pyTMHMM -h
usage: pyTMHMM [-h] -f SEQUENCE_FILE [-m MODEL_FILE] [-p]
required arguments:
-f SEQUENCE_FILE, --file SEQUENCE_FILE
path to file in fasta format with sequences
optional arguments:
-h, --help show this help message and exit
-m MODEL_FILE, --model MODEL_FILE
path to the model to use (default: TMHMM2.0.model)
-p, --plot plot posterior probabilies

The -p/--plot option requires matplotlib.

The input sequence file should have one or more sequences in Fasta format, for example:

> head PAR3_HUMAN.fasta
>sp|O00254|PAR3_HUMAN Proteinase-activated receptor 3 OS=Homo sapiens OX=9606 GN=F2RL2 PE=1 SV=1
MKALIFAAAGLLLLLPTFCQSGMENDTNNLAKPTLPIKTFRGAPPNSFEEFPFSALEGWT
GATITVKIKCPEESASHLHVKNATMGYLTSSLSTKLIPAIYLLVFVVGVPANAVTLWMLF
FRTRSICTTVFYTNLAIADFLFCVTLPFKIAYHLNGNNWVFGEVLCRATTVIFYGNMYCS
ILLLACISINRYLAIVHPFTYRGLPKHTYALVTCGLVWATVFLYMLPFFILKQEYYLVQP
DITTCHDVHNTCESSSPFQLYYFISLAFFGFLIPFVLIIYCYAAIIRTLNAYDHRWLWYV
KASLLILVIFTICFAPSNIILIIHHANYYYNNTDGLYFIYLIALCLGSLNSCLDPFLYFL
MSKTRNHSTAYLTK

Example command:

> pyTMHMM -f PAR3_HUMAN.fasta

This produces three files for each sequence in the Fasta file, named by id.

Summary file

The coordinates of the predicted domains:

> cat sp|O00254|PAR3_HUMAN.summary 0 97 outside
98 120 transmembrane helix
121 128 inside
129 151 transmembrane helix
152 165 outside
166 188 transmembrane helix
189 207 inside
208 230 transmembrane helix
231 259 outside
260 282 transmembrane helix
283 302 inside
303 322 transmembrane helix
323 336 outside
337 359 transmembrane helix
360 373 inside

Annotation file

An annotated sequence in Fasta-like format:

> cat sp|O00254|PAR3_HUMAN.annotation >sp|O00254|PAR3_HUMAN Proteinase-activated receptor 3 OS=Homo sapiens OX=9606 GN=F2RL2 PE=1 SV=1
OOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOO
OOOOOOOOOOOOOOOOOOOMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiMMMMMMMMMMMMMMMMMMMMMMMoooooo
ooooooooMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiiiiiiiiiiiiMMMMMMMMMMMMMMMMMMMMMMMoooooo
oooooooooooooooooooooooMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiiiiiiiiiiiiiMMMMMMMMMMMMM
MMMMMMMooooooooooooooMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiiiiiii

Posterior probabilities file

A file containing the posterior probabilities for each label:

> head sp|O00254|PAR3_HUMAN.plot inside membrane outside
0.6417636608794935 0.0 0.3582363391205064
0.693933311909457 0.006819179965744769 0.2992475081247982
0.3041488405999551 0.36045181385397806 0.3353993455460668
0.15867304975718463 0.5320740444690139 0.3092529057738015
0.011878169861623369 0.8126781067794638 0.1754437233589128
0.009103844612501565 0.7722962064006578 0.21859994898684057
0.0008287471596339259 0.6966223976666195 0.3025488551737467
0.0007860447761827514 0.7122010989508554 0.2870128562729619
0.0006349307902653272 0.712364526792757 0.28700054241697776

Optional plot file

If the -p flag is set and matplotlib is installed a plot in PDF format is made:

"TM domains in PAR3_HUMAN"

doc/sp|O00254|PAR3_HUMAN.pdf

API Usage

You can also use pyTMHMM as a library:

import pyTMHMM
annotation, posterior = pyTMHMM.predict(sequence_string)

This returns the annotation as a string and the posterior probabilities for each label as a numpy array with shape (len(sequence), 3) where column 0, 1 and 2 corresponds to being inside, transmembrane and outside, respectively.

If you don't need the posterior probabilities set compute_posterior=False, this will save computation:

annotation = pyTMHMM.predict(
sequence_string, compute_posterior=False
)

About

Python 3.5+ implementation of TMHMM (TransMembrane helix Hidden Markov Model)

Resources

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - bosborne/pyTMHMM: Python 3.5+ implementation of TMHMM (TransMembrane helix Hidden Markov Model) · GitHub
Skip to content

Repository files navigation

pyTMHMM

pyTMHMM is a Python 3.5+/Cython implementation of the transmembrane helix predictor using a hidden Markov model (TMHMM) originally described in:

E.L. Sonnhammer, G. von Heijne, and A. Krogh. A hidden Markov model for predicting transmembrane helices in protein sequences. In J. Glasgow, T. Littlejohn, F. Major, R. Lathrop, D. Sankoff, and C. Sensen, editors, Proceedings of the Sixth International Conference on Intelligent Systems for Molecular Biology, pages 175-182, Menlo Park, CA, 1998. AAAI Press. PMID 9783223

History

Dan Søndergaard is the original author of this package and his repository is now archived. Dan wrote this code for a few reasons:

  • the source code is not available as part of the publication
  • the downloadable binaries are Linux-only
  • the downloadable binaries may not be redistributed, so it's not possible to put them in a Docker image or a VM for other people to use
  • there is a need to predict transmembrane helices in a scripted, automated way

This Python implementation includes a parser for the undocumented file format used to describe the model and a fast Cython implementation of the Viterbi algorithm used to perform the annotation. The tool will output files similar to the files produced by the original TMHMM implementation.

Incompatibilities

  • The original TMHMM implementation handles ambigious characters and gaps in an undocumented way. However, pyTMHMM does not attempt to handle such characters at all and will fail. A possible fix is to replace those characters with something also based on expert/domain knowledge.

Installation

This package supports Python 3.5 or greater. Install with:

> pip install pyTMHMM

Command Line Usage

> pyTMHMM -h
usage: pyTMHMM [-h] -f SEQUENCE_FILE [-m MODEL_FILE] [-p]
required arguments:
-f SEQUENCE_FILE, --file SEQUENCE_FILE
path to file in fasta format with sequences
optional arguments:
-h, --help show this help message and exit
-m MODEL_FILE, --model MODEL_FILE
path to the model to use (default: TMHMM2.0.model)
-p, --plot plot posterior probabilies

The -p/--plot option requires matplotlib.

The input sequence file should have one or more sequences in Fasta format, for example:

> head PAR3_HUMAN.fasta
>sp|O00254|PAR3_HUMAN Proteinase-activated receptor 3 OS=Homo sapiens OX=9606 GN=F2RL2 PE=1 SV=1
MKALIFAAAGLLLLLPTFCQSGMENDTNNLAKPTLPIKTFRGAPPNSFEEFPFSALEGWT
GATITVKIKCPEESASHLHVKNATMGYLTSSLSTKLIPAIYLLVFVVGVPANAVTLWMLF
FRTRSICTTVFYTNLAIADFLFCVTLPFKIAYHLNGNNWVFGEVLCRATTVIFYGNMYCS
ILLLACISINRYLAIVHPFTYRGLPKHTYALVTCGLVWATVFLYMLPFFILKQEYYLVQP
DITTCHDVHNTCESSSPFQLYYFISLAFFGFLIPFVLIIYCYAAIIRTLNAYDHRWLWYV
KASLLILVIFTICFAPSNIILIIHHANYYYNNTDGLYFIYLIALCLGSLNSCLDPFLYFL
MSKTRNHSTAYLTK

Example command:

> pyTMHMM -f PAR3_HUMAN.fasta

This produces three files for each sequence in the Fasta file, named by id.

Summary file

The coordinates of the predicted domains:

> cat sp|O00254|PAR3_HUMAN.summary 0 97 outside
98 120 transmembrane helix
121 128 inside
129 151 transmembrane helix
152 165 outside
166 188 transmembrane helix
189 207 inside
208 230 transmembrane helix
231 259 outside
260 282 transmembrane helix
283 302 inside
303 322 transmembrane helix
323 336 outside
337 359 transmembrane helix
360 373 inside

Annotation file

An annotated sequence in Fasta-like format:

> cat sp|O00254|PAR3_HUMAN.annotation >sp|O00254|PAR3_HUMAN Proteinase-activated receptor 3 OS=Homo sapiens OX=9606 GN=F2RL2 PE=1 SV=1
OOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOO
OOOOOOOOOOOOOOOOOOOMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiMMMMMMMMMMMMMMMMMMMMMMMoooooo
ooooooooMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiiiiiiiiiiiiMMMMMMMMMMMMMMMMMMMMMMMoooooo
oooooooooooooooooooooooMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiiiiiiiiiiiiiMMMMMMMMMMMMM
MMMMMMMooooooooooooooMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiiiiiii

Posterior probabilities file

A file containing the posterior probabilities for each label:

> head sp|O00254|PAR3_HUMAN.plot inside membrane outside
0.6417636608794935 0.0 0.3582363391205064
0.693933311909457 0.006819179965744769 0.2992475081247982
0.3041488405999551 0.36045181385397806 0.3353993455460668
0.15867304975718463 0.5320740444690139 0.3092529057738015
0.011878169861623369 0.8126781067794638 0.1754437233589128
0.009103844612501565 0.7722962064006578 0.21859994898684057
0.0008287471596339259 0.6966223976666195 0.3025488551737467
0.0007860447761827514 0.7122010989508554 0.2870128562729619
0.0006349307902653272 0.712364526792757 0.28700054241697776

Optional plot file

If the -p flag is set and matplotlib is installed a plot in PDF format is made:

"TM domains in PAR3_HUMAN"

doc/sp|O00254|PAR3_HUMAN.pdf

API Usage

You can also use pyTMHMM as a library:

import pyTMHMM
annotation, posterior = pyTMHMM.predict(sequence_string)

This returns the annotation as a string and the posterior probabilities for each label as a numpy array with shape (len(sequence), 3) where column 0, 1 and 2 corresponds to being inside, transmembrane and outside, respectively.

If you don't need the posterior probabilities set compute_posterior=False, this will save computation:

annotation = pyTMHMM.predict(
sequence_string, compute_posterior=False
)

About

Python 3.5+ implementation of TMHMM (TransMembrane helix Hidden Markov Model)

Resources

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - bosborne/pyTMHMM: Python 3.5+ implementation of TMHMM (TransMembrane helix Hidden Markov Model) · GitHub
Skip to content

Repository files navigation

pyTMHMM

pyTMHMM is a Python 3.5+/Cython implementation of the transmembrane helix predictor using a hidden Markov model (TMHMM) originally described in:

E.L. Sonnhammer, G. von Heijne, and A. Krogh. A hidden Markov model for predicting transmembrane helices in protein sequences. In J. Glasgow, T. Littlejohn, F. Major, R. Lathrop, D. Sankoff, and C. Sensen, editors, Proceedings of the Sixth International Conference on Intelligent Systems for Molecular Biology, pages 175-182, Menlo Park, CA, 1998. AAAI Press. PMID 9783223

History

Dan Søndergaard is the original author of this package and his repository is now archived. Dan wrote this code for a few reasons:

  • the source code is not available as part of the publication
  • the downloadable binaries are Linux-only
  • the downloadable binaries may not be redistributed, so it's not possible to put them in a Docker image or a VM for other people to use
  • there is a need to predict transmembrane helices in a scripted, automated way

This Python implementation includes a parser for the undocumented file format used to describe the model and a fast Cython implementation of the Viterbi algorithm used to perform the annotation. The tool will output files similar to the files produced by the original TMHMM implementation.

Incompatibilities

  • The original TMHMM implementation handles ambigious characters and gaps in an undocumented way. However, pyTMHMM does not attempt to handle such characters at all and will fail. A possible fix is to replace those characters with something also based on expert/domain knowledge.

Installation

This package supports Python 3.5 or greater. Install with:

> pip install pyTMHMM

Command Line Usage

> pyTMHMM -h
usage: pyTMHMM [-h] -f SEQUENCE_FILE [-m MODEL_FILE] [-p]
required arguments:
-f SEQUENCE_FILE, --file SEQUENCE_FILE
path to file in fasta format with sequences
optional arguments:
-h, --help show this help message and exit
-m MODEL_FILE, --model MODEL_FILE
path to the model to use (default: TMHMM2.0.model)
-p, --plot plot posterior probabilies

The -p/--plot option requires matplotlib.

The input sequence file should have one or more sequences in Fasta format, for example:

> head PAR3_HUMAN.fasta
>sp|O00254|PAR3_HUMAN Proteinase-activated receptor 3 OS=Homo sapiens OX=9606 GN=F2RL2 PE=1 SV=1
MKALIFAAAGLLLLLPTFCQSGMENDTNNLAKPTLPIKTFRGAPPNSFEEFPFSALEGWT
GATITVKIKCPEESASHLHVKNATMGYLTSSLSTKLIPAIYLLVFVVGVPANAVTLWMLF
FRTRSICTTVFYTNLAIADFLFCVTLPFKIAYHLNGNNWVFGEVLCRATTVIFYGNMYCS
ILLLACISINRYLAIVHPFTYRGLPKHTYALVTCGLVWATVFLYMLPFFILKQEYYLVQP
DITTCHDVHNTCESSSPFQLYYFISLAFFGFLIPFVLIIYCYAAIIRTLNAYDHRWLWYV
KASLLILVIFTICFAPSNIILIIHHANYYYNNTDGLYFIYLIALCLGSLNSCLDPFLYFL
MSKTRNHSTAYLTK

Example command:

> pyTMHMM -f PAR3_HUMAN.fasta

This produces three files for each sequence in the Fasta file, named by id.

Summary file

The coordinates of the predicted domains:

> cat sp|O00254|PAR3_HUMAN.summary 0 97 outside
98 120 transmembrane helix
121 128 inside
129 151 transmembrane helix
152 165 outside
166 188 transmembrane helix
189 207 inside
208 230 transmembrane helix
231 259 outside
260 282 transmembrane helix
283 302 inside
303 322 transmembrane helix
323 336 outside
337 359 transmembrane helix
360 373 inside

Annotation file

An annotated sequence in Fasta-like format:

> cat sp|O00254|PAR3_HUMAN.annotation >sp|O00254|PAR3_HUMAN Proteinase-activated receptor 3 OS=Homo sapiens OX=9606 GN=F2RL2 PE=1 SV=1
OOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOO
OOOOOOOOOOOOOOOOOOOMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiMMMMMMMMMMMMMMMMMMMMMMMoooooo
ooooooooMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiiiiiiiiiiiiMMMMMMMMMMMMMMMMMMMMMMMoooooo
oooooooooooooooooooooooMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiiiiiiiiiiiiiMMMMMMMMMMMMM
MMMMMMMooooooooooooooMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiiiiiii

Posterior probabilities file

A file containing the posterior probabilities for each label:

> head sp|O00254|PAR3_HUMAN.plot inside membrane outside
0.6417636608794935 0.0 0.3582363391205064
0.693933311909457 0.006819179965744769 0.2992475081247982
0.3041488405999551 0.36045181385397806 0.3353993455460668
0.15867304975718463 0.5320740444690139 0.3092529057738015
0.011878169861623369 0.8126781067794638 0.1754437233589128
0.009103844612501565 0.7722962064006578 0.21859994898684057
0.0008287471596339259 0.6966223976666195 0.3025488551737467
0.0007860447761827514 0.7122010989508554 0.2870128562729619
0.0006349307902653272 0.712364526792757 0.28700054241697776

Optional plot file

If the -p flag is set and matplotlib is installed a plot in PDF format is made:

"TM domains in PAR3_HUMAN"

doc/sp|O00254|PAR3_HUMAN.pdf

API Usage

You can also use pyTMHMM as a library:

import pyTMHMM
annotation, posterior = pyTMHMM.predict(sequence_string)

This returns the annotation as a string and the posterior probabilities for each label as a numpy array with shape (len(sequence), 3) where column 0, 1 and 2 corresponds to being inside, transmembrane and outside, respectively.

If you don't need the posterior probabilities set compute_posterior=False, this will save computation:

annotation = pyTMHMM.predict(
sequence_string, compute_posterior=False
)

About

Python 3.5+ implementation of TMHMM (TransMembrane helix Hidden Markov Model)

Resources

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - bosborne/pyTMHMM: Python 3.5+ implementation of TMHMM (TransMembrane helix Hidden Markov Model) · GitHub
Skip to content

Repository files navigation

pyTMHMM

pyTMHMM is a Python 3.5+/Cython implementation of the transmembrane helix predictor using a hidden Markov model (TMHMM) originally described in:

E.L. Sonnhammer, G. von Heijne, and A. Krogh. A hidden Markov model for predicting transmembrane helices in protein sequences. In J. Glasgow, T. Littlejohn, F. Major, R. Lathrop, D. Sankoff, and C. Sensen, editors, Proceedings of the Sixth International Conference on Intelligent Systems for Molecular Biology, pages 175-182, Menlo Park, CA, 1998. AAAI Press. PMID 9783223

History

Dan Søndergaard is the original author of this package and his repository is now archived. Dan wrote this code for a few reasons:

  • the source code is not available as part of the publication
  • the downloadable binaries are Linux-only
  • the downloadable binaries may not be redistributed, so it's not possible to put them in a Docker image or a VM for other people to use
  • there is a need to predict transmembrane helices in a scripted, automated way

This Python implementation includes a parser for the undocumented file format used to describe the model and a fast Cython implementation of the Viterbi algorithm used to perform the annotation. The tool will output files similar to the files produced by the original TMHMM implementation.

Incompatibilities

  • The original TMHMM implementation handles ambigious characters and gaps in an undocumented way. However, pyTMHMM does not attempt to handle such characters at all and will fail. A possible fix is to replace those characters with something also based on expert/domain knowledge.

Installation

This package supports Python 3.5 or greater. Install with:

> pip install pyTMHMM

Command Line Usage

> pyTMHMM -h
usage: pyTMHMM [-h] -f SEQUENCE_FILE [-m MODEL_FILE] [-p]
required arguments:
-f SEQUENCE_FILE, --file SEQUENCE_FILE
path to file in fasta format with sequences
optional arguments:
-h, --help show this help message and exit
-m MODEL_FILE, --model MODEL_FILE
path to the model to use (default: TMHMM2.0.model)
-p, --plot plot posterior probabilies

The -p/--plot option requires matplotlib.

The input sequence file should have one or more sequences in Fasta format, for example:

> head PAR3_HUMAN.fasta
>sp|O00254|PAR3_HUMAN Proteinase-activated receptor 3 OS=Homo sapiens OX=9606 GN=F2RL2 PE=1 SV=1
MKALIFAAAGLLLLLPTFCQSGMENDTNNLAKPTLPIKTFRGAPPNSFEEFPFSALEGWT
GATITVKIKCPEESASHLHVKNATMGYLTSSLSTKLIPAIYLLVFVVGVPANAVTLWMLF
FRTRSICTTVFYTNLAIADFLFCVTLPFKIAYHLNGNNWVFGEVLCRATTVIFYGNMYCS
ILLLACISINRYLAIVHPFTYRGLPKHTYALVTCGLVWATVFLYMLPFFILKQEYYLVQP
DITTCHDVHNTCESSSPFQLYYFISLAFFGFLIPFVLIIYCYAAIIRTLNAYDHRWLWYV
KASLLILVIFTICFAPSNIILIIHHANYYYNNTDGLYFIYLIALCLGSLNSCLDPFLYFL
MSKTRNHSTAYLTK

Example command:

> pyTMHMM -f PAR3_HUMAN.fasta

This produces three files for each sequence in the Fasta file, named by id.

Summary file

The coordinates of the predicted domains:

> cat sp|O00254|PAR3_HUMAN.summary 0 97 outside
98 120 transmembrane helix
121 128 inside
129 151 transmembrane helix
152 165 outside
166 188 transmembrane helix
189 207 inside
208 230 transmembrane helix
231 259 outside
260 282 transmembrane helix
283 302 inside
303 322 transmembrane helix
323 336 outside
337 359 transmembrane helix
360 373 inside

Annotation file

An annotated sequence in Fasta-like format:

> cat sp|O00254|PAR3_HUMAN.annotation >sp|O00254|PAR3_HUMAN Proteinase-activated receptor 3 OS=Homo sapiens OX=9606 GN=F2RL2 PE=1 SV=1
OOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOO
OOOOOOOOOOOOOOOOOOOMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiMMMMMMMMMMMMMMMMMMMMMMMoooooo
ooooooooMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiiiiiiiiiiiiMMMMMMMMMMMMMMMMMMMMMMMoooooo
oooooooooooooooooooooooMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiiiiiiiiiiiiiMMMMMMMMMMMMM
MMMMMMMooooooooooooooMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiiiiiii

Posterior probabilities file

A file containing the posterior probabilities for each label:

> head sp|O00254|PAR3_HUMAN.plot inside membrane outside
0.6417636608794935 0.0 0.3582363391205064
0.693933311909457 0.006819179965744769 0.2992475081247982
0.3041488405999551 0.36045181385397806 0.3353993455460668
0.15867304975718463 0.5320740444690139 0.3092529057738015
0.011878169861623369 0.8126781067794638 0.1754437233589128
0.009103844612501565 0.7722962064006578 0.21859994898684057
0.0008287471596339259 0.6966223976666195 0.3025488551737467
0.0007860447761827514 0.7122010989508554 0.2870128562729619
0.0006349307902653272 0.712364526792757 0.28700054241697776

Optional plot file

If the -p flag is set and matplotlib is installed a plot in PDF format is made:

"TM domains in PAR3_HUMAN"

doc/sp|O00254|PAR3_HUMAN.pdf

API Usage

You can also use pyTMHMM as a library:

import pyTMHMM
annotation, posterior = pyTMHMM.predict(sequence_string)

This returns the annotation as a string and the posterior probabilities for each label as a numpy array with shape (len(sequence), 3) where column 0, 1 and 2 corresponds to being inside, transmembrane and outside, respectively.

If you don't need the posterior probabilities set compute_posterior=False, this will save computation:

annotation = pyTMHMM.predict(
sequence_string, compute_posterior=False
)

About

Python 3.5+ implementation of TMHMM (TransMembrane helix Hidden Markov Model)

Resources

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - bosborne/pyTMHMM: Python 3.5+ implementation of TMHMM (TransMembrane helix Hidden Markov Model) · GitHub
Skip to content

Repository files navigation

pyTMHMM

pyTMHMM is a Python 3.5+/Cython implementation of the transmembrane helix predictor using a hidden Markov model (TMHMM) originally described in:

E.L. Sonnhammer, G. von Heijne, and A. Krogh. A hidden Markov model for predicting transmembrane helices in protein sequences. In J. Glasgow, T. Littlejohn, F. Major, R. Lathrop, D. Sankoff, and C. Sensen, editors, Proceedings of the Sixth International Conference on Intelligent Systems for Molecular Biology, pages 175-182, Menlo Park, CA, 1998. AAAI Press. PMID 9783223

History

Dan Søndergaard is the original author of this package and his repository is now archived. Dan wrote this code for a few reasons:

  • the source code is not available as part of the publication
  • the downloadable binaries are Linux-only
  • the downloadable binaries may not be redistributed, so it's not possible to put them in a Docker image or a VM for other people to use
  • there is a need to predict transmembrane helices in a scripted, automated way

This Python implementation includes a parser for the undocumented file format used to describe the model and a fast Cython implementation of the Viterbi algorithm used to perform the annotation. The tool will output files similar to the files produced by the original TMHMM implementation.

Incompatibilities

  • The original TMHMM implementation handles ambigious characters and gaps in an undocumented way. However, pyTMHMM does not attempt to handle such characters at all and will fail. A possible fix is to replace those characters with something also based on expert/domain knowledge.

Installation

This package supports Python 3.5 or greater. Install with:

> pip install pyTMHMM

Command Line Usage

> pyTMHMM -h
usage: pyTMHMM [-h] -f SEQUENCE_FILE [-m MODEL_FILE] [-p]
required arguments:
-f SEQUENCE_FILE, --file SEQUENCE_FILE
path to file in fasta format with sequences
optional arguments:
-h, --help show this help message and exit
-m MODEL_FILE, --model MODEL_FILE
path to the model to use (default: TMHMM2.0.model)
-p, --plot plot posterior probabilies

The -p/--plot option requires matplotlib.

The input sequence file should have one or more sequences in Fasta format, for example:

> head PAR3_HUMAN.fasta
>sp|O00254|PAR3_HUMAN Proteinase-activated receptor 3 OS=Homo sapiens OX=9606 GN=F2RL2 PE=1 SV=1
MKALIFAAAGLLLLLPTFCQSGMENDTNNLAKPTLPIKTFRGAPPNSFEEFPFSALEGWT
GATITVKIKCPEESASHLHVKNATMGYLTSSLSTKLIPAIYLLVFVVGVPANAVTLWMLF
FRTRSICTTVFYTNLAIADFLFCVTLPFKIAYHLNGNNWVFGEVLCRATTVIFYGNMYCS
ILLLACISINRYLAIVHPFTYRGLPKHTYALVTCGLVWATVFLYMLPFFILKQEYYLVQP
DITTCHDVHNTCESSSPFQLYYFISLAFFGFLIPFVLIIYCYAAIIRTLNAYDHRWLWYV
KASLLILVIFTICFAPSNIILIIHHANYYYNNTDGLYFIYLIALCLGSLNSCLDPFLYFL
MSKTRNHSTAYLTK

Example command:

> pyTMHMM -f PAR3_HUMAN.fasta

This produces three files for each sequence in the Fasta file, named by id.

Summary file

The coordinates of the predicted domains:

> cat sp|O00254|PAR3_HUMAN.summary 0 97 outside
98 120 transmembrane helix
121 128 inside
129 151 transmembrane helix
152 165 outside
166 188 transmembrane helix
189 207 inside
208 230 transmembrane helix
231 259 outside
260 282 transmembrane helix
283 302 inside
303 322 transmembrane helix
323 336 outside
337 359 transmembrane helix
360 373 inside

Annotation file

An annotated sequence in Fasta-like format:

> cat sp|O00254|PAR3_HUMAN.annotation >sp|O00254|PAR3_HUMAN Proteinase-activated receptor 3 OS=Homo sapiens OX=9606 GN=F2RL2 PE=1 SV=1
OOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOO
OOOOOOOOOOOOOOOOOOOMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiMMMMMMMMMMMMMMMMMMMMMMMoooooo
ooooooooMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiiiiiiiiiiiiMMMMMMMMMMMMMMMMMMMMMMMoooooo
oooooooooooooooooooooooMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiiiiiiiiiiiiiMMMMMMMMMMMMM
MMMMMMMooooooooooooooMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiiiiiii

Posterior probabilities file

A file containing the posterior probabilities for each label:

> head sp|O00254|PAR3_HUMAN.plot inside membrane outside
0.6417636608794935 0.0 0.3582363391205064
0.693933311909457 0.006819179965744769 0.2992475081247982
0.3041488405999551 0.36045181385397806 0.3353993455460668
0.15867304975718463 0.5320740444690139 0.3092529057738015
0.011878169861623369 0.8126781067794638 0.1754437233589128
0.009103844612501565 0.7722962064006578 0.21859994898684057
0.0008287471596339259 0.6966223976666195 0.3025488551737467
0.0007860447761827514 0.7122010989508554 0.2870128562729619
0.0006349307902653272 0.712364526792757 0.28700054241697776

Optional plot file

If the -p flag is set and matplotlib is installed a plot in PDF format is made:

"TM domains in PAR3_HUMAN"

doc/sp|O00254|PAR3_HUMAN.pdf

API Usage

You can also use pyTMHMM as a library:

import pyTMHMM
annotation, posterior = pyTMHMM.predict(sequence_string)

This returns the annotation as a string and the posterior probabilities for each label as a numpy array with shape (len(sequence), 3) where column 0, 1 and 2 corresponds to being inside, transmembrane and outside, respectively.

If you don't need the posterior probabilities set compute_posterior=False, this will save computation:

annotation = pyTMHMM.predict(
sequence_string, compute_posterior=False
)

About

Python 3.5+ implementation of TMHMM (TransMembrane helix Hidden Markov Model)

Resources

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - bosborne/pyTMHMM: Python 3.5+ implementation of TMHMM (TransMembrane helix Hidden Markov Model) · GitHub
Skip to content

Repository files navigation

pyTMHMM

pyTMHMM is a Python 3.5+/Cython implementation of the transmembrane helix predictor using a hidden Markov model (TMHMM) originally described in:

E.L. Sonnhammer, G. von Heijne, and A. Krogh. A hidden Markov model for predicting transmembrane helices in protein sequences. In J. Glasgow, T. Littlejohn, F. Major, R. Lathrop, D. Sankoff, and C. Sensen, editors, Proceedings of the Sixth International Conference on Intelligent Systems for Molecular Biology, pages 175-182, Menlo Park, CA, 1998. AAAI Press. PMID 9783223

History

Dan Søndergaard is the original author of this package and his repository is now archived. Dan wrote this code for a few reasons:

  • the source code is not available as part of the publication
  • the downloadable binaries are Linux-only
  • the downloadable binaries may not be redistributed, so it's not possible to put them in a Docker image or a VM for other people to use
  • there is a need to predict transmembrane helices in a scripted, automated way

This Python implementation includes a parser for the undocumented file format used to describe the model and a fast Cython implementation of the Viterbi algorithm used to perform the annotation. The tool will output files similar to the files produced by the original TMHMM implementation.

Incompatibilities

  • The original TMHMM implementation handles ambigious characters and gaps in an undocumented way. However, pyTMHMM does not attempt to handle such characters at all and will fail. A possible fix is to replace those characters with something also based on expert/domain knowledge.

Installation

This package supports Python 3.5 or greater. Install with:

> pip install pyTMHMM

Command Line Usage

> pyTMHMM -h
usage: pyTMHMM [-h] -f SEQUENCE_FILE [-m MODEL_FILE] [-p]
required arguments:
-f SEQUENCE_FILE, --file SEQUENCE_FILE
path to file in fasta format with sequences
optional arguments:
-h, --help show this help message and exit
-m MODEL_FILE, --model MODEL_FILE
path to the model to use (default: TMHMM2.0.model)
-p, --plot plot posterior probabilies

The -p/--plot option requires matplotlib.

The input sequence file should have one or more sequences in Fasta format, for example:

> head PAR3_HUMAN.fasta
>sp|O00254|PAR3_HUMAN Proteinase-activated receptor 3 OS=Homo sapiens OX=9606 GN=F2RL2 PE=1 SV=1
MKALIFAAAGLLLLLPTFCQSGMENDTNNLAKPTLPIKTFRGAPPNSFEEFPFSALEGWT
GATITVKIKCPEESASHLHVKNATMGYLTSSLSTKLIPAIYLLVFVVGVPANAVTLWMLF
FRTRSICTTVFYTNLAIADFLFCVTLPFKIAYHLNGNNWVFGEVLCRATTVIFYGNMYCS
ILLLACISINRYLAIVHPFTYRGLPKHTYALVTCGLVWATVFLYMLPFFILKQEYYLVQP
DITTCHDVHNTCESSSPFQLYYFISLAFFGFLIPFVLIIYCYAAIIRTLNAYDHRWLWYV
KASLLILVIFTICFAPSNIILIIHHANYYYNNTDGLYFIYLIALCLGSLNSCLDPFLYFL
MSKTRNHSTAYLTK

Example command:

> pyTMHMM -f PAR3_HUMAN.fasta

This produces three files for each sequence in the Fasta file, named by id.

Summary file

The coordinates of the predicted domains:

> cat sp|O00254|PAR3_HUMAN.summary 0 97 outside
98 120 transmembrane helix
121 128 inside
129 151 transmembrane helix
152 165 outside
166 188 transmembrane helix
189 207 inside
208 230 transmembrane helix
231 259 outside
260 282 transmembrane helix
283 302 inside
303 322 transmembrane helix
323 336 outside
337 359 transmembrane helix
360 373 inside

Annotation file

An annotated sequence in Fasta-like format:

> cat sp|O00254|PAR3_HUMAN.annotation >sp|O00254|PAR3_HUMAN Proteinase-activated receptor 3 OS=Homo sapiens OX=9606 GN=F2RL2 PE=1 SV=1
OOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOO
OOOOOOOOOOOOOOOOOOOMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiMMMMMMMMMMMMMMMMMMMMMMMoooooo
ooooooooMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiiiiiiiiiiiiMMMMMMMMMMMMMMMMMMMMMMMoooooo
oooooooooooooooooooooooMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiiiiiiiiiiiiiMMMMMMMMMMMMM
MMMMMMMooooooooooooooMMMMMMMMMMMMMMMMMMMMMMMiiiiiiiiiiiiii

Posterior probabilities file

A file containing the posterior probabilities for each label:

> head sp|O00254|PAR3_HUMAN.plot inside membrane outside
0.6417636608794935 0.0 0.3582363391205064
0.693933311909457 0.006819179965744769 0.2992475081247982
0.3041488405999551 0.36045181385397806 0.3353993455460668
0.15867304975718463 0.5320740444690139 0.3092529057738015
0.011878169861623369 0.8126781067794638 0.1754437233589128
0.009103844612501565 0.7722962064006578 0.21859994898684057
0.0008287471596339259 0.6966223976666195 0.3025488551737467
0.0007860447761827514 0.7122010989508554 0.2870128562729619
0.0006349307902653272 0.712364526792757 0.28700054241697776

Optional plot file

If the -p flag is set and matplotlib is installed a plot in PDF format is made:

"TM domains in PAR3_HUMAN"

doc/sp|O00254|PAR3_HUMAN.pdf

API Usage

You can also use pyTMHMM as a library:

import pyTMHMM
annotation, posterior = pyTMHMM.predict(sequence_string)

This returns the annotation as a string and the posterior probabilities for each label as a numpy array with shape (len(sequence), 3) where column 0, 1 and 2 corresponds to being inside, transmembrane and outside, respectively.

If you don't need the posterior probabilities set compute_posterior=False, this will save computation:

annotation = pyTMHMM.predict(
sequence_string, compute_posterior=False
)

About

Python 3.5+ implementation of TMHMM (TransMembrane helix Hidden Markov Model)

Resources

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages