Repository files navigation

Profluent-E1

This repository contains the code for the Profluent E1 family of models - our best in class single sequence and retrieval augmented protein representation models. They are designed to be drop-in replacement for ESM family of models. See Section on licenses for the license details.

Changelog

  • 2026-03-09: Modify MSA Sampling code to fix bug. Files modified: src/E1/msa_sampling.py.

Available Models

Model NameModel ParametersHuggingFace Link
E1-150m150MProfluent-Bio/E1-150m
E1-300m300MProfluent-Bio/E1-300m
E1-600m600MProfluent-Bio/E1-600m

Installation

To use the code in this repository, install the dependencies using the following command:

git clone https://github.com/Profluent-AI/E1.git
cd E1 && pip3 install -e .

If you are using GPUs with cuda capability 8.0 or higher (Ampere architecture or higher), we also recommend installing the flash-attn package:

pip3 install flash-attn --no-build-isolation

Compute Requirements

While the models can be run on both CPU and GPU, we recommend using a GPU for faster inference. In addition, the model was trained with BF16 precision, so we recommend using a GPU with BF16 support (CUDA Capability 8.0 or higher; for example, NVIDIA L40, A100, H100, etc.) since that also allows the inference performance to improve using flash attention.

Interactive Usage

The model weights are hosted on Hugging Face and will be downloaded automatically in the following code. The model can be used in both single sequence and retrieval augmented mode. You can use ? as mask token. To use the model in retrieval augmented mode, you can prepend your query with homolog sequences separated by commas.

importtorchfromE1.batch_preparerimportE1BatchPreparerfromE1.modelingimportE1ForMaskedLMmodel=E1ForMaskedLM.from_pretrained("Profluent-Bio/E1-300m").to("cuda:0")
model.eval()
sequences= ["AAAAA?C", "MFCATEEKL,MCCASDF,MFCC?SEF"]
batch_preparer=E1BatchPreparer()
batch=batch_preparer.get_batch_kwargs(sequences, device="cuda:0")
withtorch.autocast("cuda", dtype=torch.bfloat16, enabled=True):
outputs=model(
input_ids=batch["input_ids"],
within_seq_position_ids=batch["within_seq_position_ids"],
global_position_ids=batch["global_position_ids"],
sequence_ids=batch["sequence_ids"],
past_key_values=None,
use_cache=False,
output_attentions=False,
output_hidden_states=False,
)
logits: torch.Tensor=outputs.logits# (B, L, V)embeddings: torch.Tensor=outputs.embeddings# (B, L, E)print(logits)
print(embeddings)
# Boolean Selectors of shape (B, L) to get relevant tokens from logits/embeddings# last_sequence_selector: True for tokens that are part of the last sequence (including boundary tokens) in case of multi-sequence input.last_sequence_selector=batch["sequence_ids"] ==batch["sequence_ids"].max(dim=1)[0][:, None]
# residue_selector: True for tokens that are part of the input sequence i.e not boundary tokens like 1, 2, <bos>, <eos>, <pad>, etc.residue_selector=~(batch_preparer.get_boundary_token_mask(batch["input_ids"]))
# last_sequence_residue_selector: True for residues that are part of the last sequence (excluding boundary tokens)last_sequence_residue_selector=last_sequence_selector&residue_selector# Will yield embeddings for ["AAAAA?C", "MFCC?SEF"] while throwing away embeddings for the boundary tokens# and for homologous sequences in the second instance. Can do similar for logits.last_sequence_embeddings= [embeddings[i, last_sequence_residue_selector[i]] foriinrange(embeddings.shape[0])]

See the cookbook/basic.ipynbOpen In Colab file for a more complete example. We also provide a notebook for embedding analysisOpen In Colab to demonstrate how to use the model to get sequence embeddings using E1Predictor class.

You can also use the cookbook/zero_shot_fitness_prediction.ipynbOpen In Colab file to predict the fitness of zero-shot substitution mutants against a wild type parent.

For performing in-silico site saturation mutagenesis using masked-marginal scoring method, see the cookbook/site_saturation_mutagenesis.ipynbOpen In Colab file. We score all possible single mutants of a given parent sequence in both single sequence and retrieval augmented mode.

Comparison to other models

Average Spearman on Substitution Assays in Protein Gym BenchmarkUnsupervised Contact Map Prediction on CAMEO
Protein Gym ResultsCAMEO Results

Licenses

Your use of the Profluent-E1 model code is governed by the Apache License 2.0, while your use of the Profluent-E1 model weights and the full release of the Profluent-E1 model is governed by a similarly permissive license with additional attribution requirements - see the NOTICE file for details. You can use, share, and modify Profluent-E1 for free, but you must follow our ATTRIBUTION guidelines to give credit, include the license when you share and follow some other basic rules. Profluent is not responsible for what you build, and may terminate your rights to use Profluent-E1 if you breach the license.

  1. Code in src/E1/model/flash_attention_utils.py is adapted from flash-attention project under BSD-3-Clause license.

About

Profluent-E1 family of Protein Encoder Models

Resources

Stars

115 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all \u003cpre\u003e\u003ccode\u003e blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks"); } } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); } })(); (function(){ try { var __m = "github.com"; var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Profluent-E1

This repository contains the code for the Profluent E1 family of models - our best in class single sequence and retrieval augmented protein representation models. They are designed to be drop-in replacement for ESM family of models. See Section on licenses for the license details.

Changelog

  • 2026-03-09: Modify MSA Sampling code to fix bug. Files modified: src/E1/msa_sampling.py.

Available Models

Model NameModel ParametersHuggingFace Link
E1-150m150MProfluent-Bio/E1-150m
E1-300m300MProfluent-Bio/E1-300m
E1-600m600MProfluent-Bio/E1-600m

Installation

To use the code in this repository, install the dependencies using the following command:

git clone https://github.com/Profluent-AI/E1.git
cd E1 && pip3 install -e .

If you are using GPUs with cuda capability 8.0 or higher (Ampere architecture or higher), we also recommend installing the flash-attn package:

pip3 install flash-attn --no-build-isolation

Compute Requirements

While the models can be run on both CPU and GPU, we recommend using a GPU for faster inference. In addition, the model was trained with BF16 precision, so we recommend using a GPU with BF16 support (CUDA Capability 8.0 or higher; for example, NVIDIA L40, A100, H100, etc.) since that also allows the inference performance to improve using flash attention.

Interactive Usage

The model weights are hosted on Hugging Face and will be downloaded automatically in the following code. The model can be used in both single sequence and retrieval augmented mode. You can use ? as mask token. To use the model in retrieval augmented mode, you can prepend your query with homolog sequences separated by commas.

importtorchfromE1.batch_preparerimportE1BatchPreparerfromE1.modelingimportE1ForMaskedLMmodel=E1ForMaskedLM.from_pretrained("Profluent-Bio/E1-300m").to("cuda:0")
model.eval()
sequences= ["AAAAA?C", "MFCATEEKL,MCCASDF,MFCC?SEF"]
batch_preparer=E1BatchPreparer()
batch=batch_preparer.get_batch_kwargs(sequences, device="cuda:0")
withtorch.autocast("cuda", dtype=torch.bfloat16, enabled=True):
outputs=model(
input_ids=batch["input_ids"],
within_seq_position_ids=batch["within_seq_position_ids"],
global_position_ids=batch["global_position_ids"],
sequence_ids=batch["sequence_ids"],
past_key_values=None,
use_cache=False,
output_attentions=False,
output_hidden_states=False,
)
logits: torch.Tensor=outputs.logits# (B, L, V)embeddings: torch.Tensor=outputs.embeddings# (B, L, E)print(logits)
print(embeddings)
# Boolean Selectors of shape (B, L) to get relevant tokens from logits/embeddings# last_sequence_selector: True for tokens that are part of the last sequence (including boundary tokens) in case of multi-sequence input.last_sequence_selector=batch["sequence_ids"] ==batch["sequence_ids"].max(dim=1)[0][:, None]
# residue_selector: True for tokens that are part of the input sequence i.e not boundary tokens like 1, 2, <bos>, <eos>, <pad>, etc.residue_selector=~(batch_preparer.get_boundary_token_mask(batch["input_ids"]))
# last_sequence_residue_selector: True for residues that are part of the last sequence (excluding boundary tokens)last_sequence_residue_selector=last_sequence_selector&residue_selector# Will yield embeddings for ["AAAAA?C", "MFCC?SEF"] while throwing away embeddings for the boundary tokens# and for homologous sequences in the second instance. Can do similar for logits.last_sequence_embeddings= [embeddings[i, last_sequence_residue_selector[i]] foriinrange(embeddings.shape[0])]

See the cookbook/basic.ipynbOpen In Colab file for a more complete example. We also provide a notebook for embedding analysisOpen In Colab to demonstrate how to use the model to get sequence embeddings using E1Predictor class.

You can also use the cookbook/zero_shot_fitness_prediction.ipynbOpen In Colab file to predict the fitness of zero-shot substitution mutants against a wild type parent.

For performing in-silico site saturation mutagenesis using masked-marginal scoring method, see the cookbook/site_saturation_mutagenesis.ipynbOpen In Colab file. We score all possible single mutants of a given parent sequence in both single sequence and retrieval augmented mode.

Comparison to other models

Average Spearman on Substitution Assays in Protein Gym BenchmarkUnsupervised Contact Map Prediction on CAMEO
Protein Gym ResultsCAMEO Results

Licenses

Your use of the Profluent-E1 model code is governed by the Apache License 2.0, while your use of the Profluent-E1 model weights and the full release of the Profluent-E1 model is governed by a similarly permissive license with additional attribution requirements - see the NOTICE file for details. You can use, share, and modify Profluent-E1 for free, but you must follow our ATTRIBUTION guidelines to give credit, include the license when you share and follow some other basic rules. Profluent is not responsible for what you build, and may terminate your rights to use Profluent-E1 if you breach the license.

  1. Code in src/E1/model/flash_attention_utils.py is adapted from flash-attention project under BSD-3-Clause license.

About

Profluent-E1 family of Protein Encoder Models

Resources

Stars

115 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Profluent-E1

This repository contains the code for the Profluent E1 family of models - our best in class single sequence and retrieval augmented protein representation models. They are designed to be drop-in replacement for ESM family of models. See Section on licenses for the license details.

Changelog

  • 2026-03-09: Modify MSA Sampling code to fix bug. Files modified: src/E1/msa_sampling.py.

Available Models

Model NameModel ParametersHuggingFace Link
E1-150m150MProfluent-Bio/E1-150m
E1-300m300MProfluent-Bio/E1-300m
E1-600m600MProfluent-Bio/E1-600m

Installation

To use the code in this repository, install the dependencies using the following command:

git clone https://github.com/Profluent-AI/E1.git
cd E1 && pip3 install -e .

If you are using GPUs with cuda capability 8.0 or higher (Ampere architecture or higher), we also recommend installing the flash-attn package:

pip3 install flash-attn --no-build-isolation

Compute Requirements

While the models can be run on both CPU and GPU, we recommend using a GPU for faster inference. In addition, the model was trained with BF16 precision, so we recommend using a GPU with BF16 support (CUDA Capability 8.0 or higher; for example, NVIDIA L40, A100, H100, etc.) since that also allows the inference performance to improve using flash attention.

Interactive Usage

The model weights are hosted on Hugging Face and will be downloaded automatically in the following code. The model can be used in both single sequence and retrieval augmented mode. You can use ? as mask token. To use the model in retrieval augmented mode, you can prepend your query with homolog sequences separated by commas.

importtorchfromE1.batch_preparerimportE1BatchPreparerfromE1.modelingimportE1ForMaskedLMmodel=E1ForMaskedLM.from_pretrained("Profluent-Bio/E1-300m").to("cuda:0")
model.eval()
sequences= ["AAAAA?C", "MFCATEEKL,MCCASDF,MFCC?SEF"]
batch_preparer=E1BatchPreparer()
batch=batch_preparer.get_batch_kwargs(sequences, device="cuda:0")
withtorch.autocast("cuda", dtype=torch.bfloat16, enabled=True):
outputs=model(
input_ids=batch["input_ids"],
within_seq_position_ids=batch["within_seq_position_ids"],
global_position_ids=batch["global_position_ids"],
sequence_ids=batch["sequence_ids"],
past_key_values=None,
use_cache=False,
output_attentions=False,
output_hidden_states=False,
)
logits: torch.Tensor=outputs.logits# (B, L, V)embeddings: torch.Tensor=outputs.embeddings# (B, L, E)print(logits)
print(embeddings)
# Boolean Selectors of shape (B, L) to get relevant tokens from logits/embeddings# last_sequence_selector: True for tokens that are part of the last sequence (including boundary tokens) in case of multi-sequence input.last_sequence_selector=batch["sequence_ids"] ==batch["sequence_ids"].max(dim=1)[0][:, None]
# residue_selector: True for tokens that are part of the input sequence i.e not boundary tokens like 1, 2, <bos>, <eos>, <pad>, etc.residue_selector=~(batch_preparer.get_boundary_token_mask(batch["input_ids"]))
# last_sequence_residue_selector: True for residues that are part of the last sequence (excluding boundary tokens)last_sequence_residue_selector=last_sequence_selector&residue_selector# Will yield embeddings for ["AAAAA?C", "MFCC?SEF"] while throwing away embeddings for the boundary tokens# and for homologous sequences in the second instance. Can do similar for logits.last_sequence_embeddings= [embeddings[i, last_sequence_residue_selector[i]] foriinrange(embeddings.shape[0])]

See the cookbook/basic.ipynbOpen In Colab file for a more complete example. We also provide a notebook for embedding analysisOpen In Colab to demonstrate how to use the model to get sequence embeddings using E1Predictor class.

You can also use the cookbook/zero_shot_fitness_prediction.ipynbOpen In Colab file to predict the fitness of zero-shot substitution mutants against a wild type parent.

For performing in-silico site saturation mutagenesis using masked-marginal scoring method, see the cookbook/site_saturation_mutagenesis.ipynbOpen In Colab file. We score all possible single mutants of a given parent sequence in both single sequence and retrieval augmented mode.

Comparison to other models

Average Spearman on Substitution Assays in Protein Gym BenchmarkUnsupervised Contact Map Prediction on CAMEO
Protein Gym ResultsCAMEO Results

Licenses

Your use of the Profluent-E1 model code is governed by the Apache License 2.0, while your use of the Profluent-E1 model weights and the full release of the Profluent-E1 model is governed by a similarly permissive license with additional attribution requirements - see the NOTICE file for details. You can use, share, and modify Profluent-E1 for free, but you must follow our ATTRIBUTION guidelines to give credit, include the license when you share and follow some other basic rules. Profluent is not responsible for what you build, and may terminate your rights to use Profluent-E1 if you breach the license.

  1. Code in src/E1/model/flash_attention_utils.py is adapted from flash-attention project under BSD-3-Clause license.

About

Profluent-E1 family of Protein Encoder Models

Resources

Stars

115 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length \u003e 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Profluent-E1

This repository contains the code for the Profluent E1 family of models - our best in class single sequence and retrieval augmented protein representation models. They are designed to be drop-in replacement for ESM family of models. See Section on licenses for the license details.

Changelog

  • 2026-03-09: Modify MSA Sampling code to fix bug. Files modified: src/E1/msa_sampling.py.

Available Models

Model NameModel ParametersHuggingFace Link
E1-150m150MProfluent-Bio/E1-150m
E1-300m300MProfluent-Bio/E1-300m
E1-600m600MProfluent-Bio/E1-600m

Installation

To use the code in this repository, install the dependencies using the following command:

git clone https://github.com/Profluent-AI/E1.git
cd E1 && pip3 install -e .

If you are using GPUs with cuda capability 8.0 or higher (Ampere architecture or higher), we also recommend installing the flash-attn package:

pip3 install flash-attn --no-build-isolation

Compute Requirements

While the models can be run on both CPU and GPU, we recommend using a GPU for faster inference. In addition, the model was trained with BF16 precision, so we recommend using a GPU with BF16 support (CUDA Capability 8.0 or higher; for example, NVIDIA L40, A100, H100, etc.) since that also allows the inference performance to improve using flash attention.

Interactive Usage

The model weights are hosted on Hugging Face and will be downloaded automatically in the following code. The model can be used in both single sequence and retrieval augmented mode. You can use ? as mask token. To use the model in retrieval augmented mode, you can prepend your query with homolog sequences separated by commas.

importtorchfromE1.batch_preparerimportE1BatchPreparerfromE1.modelingimportE1ForMaskedLMmodel=E1ForMaskedLM.from_pretrained("Profluent-Bio/E1-300m").to("cuda:0")
model.eval()
sequences= ["AAAAA?C", "MFCATEEKL,MCCASDF,MFCC?SEF"]
batch_preparer=E1BatchPreparer()
batch=batch_preparer.get_batch_kwargs(sequences, device="cuda:0")
withtorch.autocast("cuda", dtype=torch.bfloat16, enabled=True):
outputs=model(
input_ids=batch["input_ids"],
within_seq_position_ids=batch["within_seq_position_ids"],
global_position_ids=batch["global_position_ids"],
sequence_ids=batch["sequence_ids"],
past_key_values=None,
use_cache=False,
output_attentions=False,
output_hidden_states=False,
)
logits: torch.Tensor=outputs.logits# (B, L, V)embeddings: torch.Tensor=outputs.embeddings# (B, L, E)print(logits)
print(embeddings)
# Boolean Selectors of shape (B, L) to get relevant tokens from logits/embeddings# last_sequence_selector: True for tokens that are part of the last sequence (including boundary tokens) in case of multi-sequence input.last_sequence_selector=batch["sequence_ids"] ==batch["sequence_ids"].max(dim=1)[0][:, None]
# residue_selector: True for tokens that are part of the input sequence i.e not boundary tokens like 1, 2, <bos>, <eos>, <pad>, etc.residue_selector=~(batch_preparer.get_boundary_token_mask(batch["input_ids"]))
# last_sequence_residue_selector: True for residues that are part of the last sequence (excluding boundary tokens)last_sequence_residue_selector=last_sequence_selector&residue_selector# Will yield embeddings for ["AAAAA?C", "MFCC?SEF"] while throwing away embeddings for the boundary tokens# and for homologous sequences in the second instance. Can do similar for logits.last_sequence_embeddings= [embeddings[i, last_sequence_residue_selector[i]] foriinrange(embeddings.shape[0])]

See the cookbook/basic.ipynbOpen In Colab file for a more complete example. We also provide a notebook for embedding analysisOpen In Colab to demonstrate how to use the model to get sequence embeddings using E1Predictor class.

You can also use the cookbook/zero_shot_fitness_prediction.ipynbOpen In Colab file to predict the fitness of zero-shot substitution mutants against a wild type parent.

For performing in-silico site saturation mutagenesis using masked-marginal scoring method, see the cookbook/site_saturation_mutagenesis.ipynbOpen In Colab file. We score all possible single mutants of a given parent sequence in both single sequence and retrieval augmented mode.

Comparison to other models

Average Spearman on Substitution Assays in Protein Gym BenchmarkUnsupervised Contact Map Prediction on CAMEO
Protein Gym ResultsCAMEO Results

Licenses

Your use of the Profluent-E1 model code is governed by the Apache License 2.0, while your use of the Profluent-E1 model weights and the full release of the Profluent-E1 model is governed by a similarly permissive license with additional attribution requirements - see the NOTICE file for details. You can use, share, and modify Profluent-E1 for free, but you must follow our ATTRIBUTION guidelines to give credit, include the license when you share and follow some other basic rules. Profluent is not responsible for what you build, and may terminate your rights to use Profluent-E1 if you breach the license.

  1. Code in src/E1/model/flash_attention_utils.py is adapted from flash-attention project under BSD-3-Clause license.

About

Profluent-E1 family of Protein Encoder Models

Resources

Stars

115 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Profluent-E1

This repository contains the code for the Profluent E1 family of models - our best in class single sequence and retrieval augmented protein representation models. They are designed to be drop-in replacement for ESM family of models. See Section on licenses for the license details.

Changelog

  • 2026-03-09: Modify MSA Sampling code to fix bug. Files modified: src/E1/msa_sampling.py.

Available Models

Model NameModel ParametersHuggingFace Link
E1-150m150MProfluent-Bio/E1-150m
E1-300m300MProfluent-Bio/E1-300m
E1-600m600MProfluent-Bio/E1-600m

Installation

To use the code in this repository, install the dependencies using the following command:

git clone https://github.com/Profluent-AI/E1.git
cd E1 && pip3 install -e .

If you are using GPUs with cuda capability 8.0 or higher (Ampere architecture or higher), we also recommend installing the flash-attn package:

pip3 install flash-attn --no-build-isolation

Compute Requirements

While the models can be run on both CPU and GPU, we recommend using a GPU for faster inference. In addition, the model was trained with BF16 precision, so we recommend using a GPU with BF16 support (CUDA Capability 8.0 or higher; for example, NVIDIA L40, A100, H100, etc.) since that also allows the inference performance to improve using flash attention.

Interactive Usage

The model weights are hosted on Hugging Face and will be downloaded automatically in the following code. The model can be used in both single sequence and retrieval augmented mode. You can use ? as mask token. To use the model in retrieval augmented mode, you can prepend your query with homolog sequences separated by commas.

importtorchfromE1.batch_preparerimportE1BatchPreparerfromE1.modelingimportE1ForMaskedLMmodel=E1ForMaskedLM.from_pretrained("Profluent-Bio/E1-300m").to("cuda:0")
model.eval()
sequences= ["AAAAA?C", "MFCATEEKL,MCCASDF,MFCC?SEF"]
batch_preparer=E1BatchPreparer()
batch=batch_preparer.get_batch_kwargs(sequences, device="cuda:0")
withtorch.autocast("cuda", dtype=torch.bfloat16, enabled=True):
outputs=model(
input_ids=batch["input_ids"],
within_seq_position_ids=batch["within_seq_position_ids"],
global_position_ids=batch["global_position_ids"],
sequence_ids=batch["sequence_ids"],
past_key_values=None,
use_cache=False,
output_attentions=False,
output_hidden_states=False,
)
logits: torch.Tensor=outputs.logits# (B, L, V)embeddings: torch.Tensor=outputs.embeddings# (B, L, E)print(logits)
print(embeddings)
# Boolean Selectors of shape (B, L) to get relevant tokens from logits/embeddings# last_sequence_selector: True for tokens that are part of the last sequence (including boundary tokens) in case of multi-sequence input.last_sequence_selector=batch["sequence_ids"] ==batch["sequence_ids"].max(dim=1)[0][:, None]
# residue_selector: True for tokens that are part of the input sequence i.e not boundary tokens like 1, 2, <bos>, <eos>, <pad>, etc.residue_selector=~(batch_preparer.get_boundary_token_mask(batch["input_ids"]))
# last_sequence_residue_selector: True for residues that are part of the last sequence (excluding boundary tokens)last_sequence_residue_selector=last_sequence_selector&residue_selector# Will yield embeddings for ["AAAAA?C", "MFCC?SEF"] while throwing away embeddings for the boundary tokens# and for homologous sequences in the second instance. Can do similar for logits.last_sequence_embeddings= [embeddings[i, last_sequence_residue_selector[i]] foriinrange(embeddings.shape[0])]

See the cookbook/basic.ipynbOpen In Colab file for a more complete example. We also provide a notebook for embedding analysisOpen In Colab to demonstrate how to use the model to get sequence embeddings using E1Predictor class.

You can also use the cookbook/zero_shot_fitness_prediction.ipynbOpen In Colab file to predict the fitness of zero-shot substitution mutants against a wild type parent.

For performing in-silico site saturation mutagenesis using masked-marginal scoring method, see the cookbook/site_saturation_mutagenesis.ipynbOpen In Colab file. We score all possible single mutants of a given parent sequence in both single sequence and retrieval augmented mode.

Comparison to other models

Average Spearman on Substitution Assays in Protein Gym BenchmarkUnsupervised Contact Map Prediction on CAMEO
Protein Gym ResultsCAMEO Results

Licenses

Your use of the Profluent-E1 model code is governed by the Apache License 2.0, while your use of the Profluent-E1 model weights and the full release of the Profluent-E1 model is governed by a similarly permissive license with additional attribution requirements - see the NOTICE file for details. You can use, share, and modify Profluent-E1 for free, but you must follow our ATTRIBUTION guidelines to give credit, include the license when you share and follow some other basic rules. Profluent is not responsible for what you build, and may terminate your rights to use Profluent-E1 if you breach the license.

  1. Code in src/E1/model/flash_attention_utils.py is adapted from flash-attention project under BSD-3-Clause license.

About

Profluent-E1 family of Protein Encoder Models

Resources

Stars

115 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Profluent-E1

This repository contains the code for the Profluent E1 family of models - our best in class single sequence and retrieval augmented protein representation models. They are designed to be drop-in replacement for ESM family of models. See Section on licenses for the license details.

Changelog

  • 2026-03-09: Modify MSA Sampling code to fix bug. Files modified: src/E1/msa_sampling.py.

Available Models

Model NameModel ParametersHuggingFace Link
E1-150m150MProfluent-Bio/E1-150m
E1-300m300MProfluent-Bio/E1-300m
E1-600m600MProfluent-Bio/E1-600m

Installation

To use the code in this repository, install the dependencies using the following command:

git clone https://github.com/Profluent-AI/E1.git
cd E1 && pip3 install -e .

If you are using GPUs with cuda capability 8.0 or higher (Ampere architecture or higher), we also recommend installing the flash-attn package:

pip3 install flash-attn --no-build-isolation

Compute Requirements

While the models can be run on both CPU and GPU, we recommend using a GPU for faster inference. In addition, the model was trained with BF16 precision, so we recommend using a GPU with BF16 support (CUDA Capability 8.0 or higher; for example, NVIDIA L40, A100, H100, etc.) since that also allows the inference performance to improve using flash attention.

Interactive Usage

The model weights are hosted on Hugging Face and will be downloaded automatically in the following code. The model can be used in both single sequence and retrieval augmented mode. You can use ? as mask token. To use the model in retrieval augmented mode, you can prepend your query with homolog sequences separated by commas.

importtorchfromE1.batch_preparerimportE1BatchPreparerfromE1.modelingimportE1ForMaskedLMmodel=E1ForMaskedLM.from_pretrained("Profluent-Bio/E1-300m").to("cuda:0")
model.eval()
sequences= ["AAAAA?C", "MFCATEEKL,MCCASDF,MFCC?SEF"]
batch_preparer=E1BatchPreparer()
batch=batch_preparer.get_batch_kwargs(sequences, device="cuda:0")
withtorch.autocast("cuda", dtype=torch.bfloat16, enabled=True):
outputs=model(
input_ids=batch["input_ids"],
within_seq_position_ids=batch["within_seq_position_ids"],
global_position_ids=batch["global_position_ids"],
sequence_ids=batch["sequence_ids"],
past_key_values=None,
use_cache=False,
output_attentions=False,
output_hidden_states=False,
)
logits: torch.Tensor=outputs.logits# (B, L, V)embeddings: torch.Tensor=outputs.embeddings# (B, L, E)print(logits)
print(embeddings)
# Boolean Selectors of shape (B, L) to get relevant tokens from logits/embeddings# last_sequence_selector: True for tokens that are part of the last sequence (including boundary tokens) in case of multi-sequence input.last_sequence_selector=batch["sequence_ids"] ==batch["sequence_ids"].max(dim=1)[0][:, None]
# residue_selector: True for tokens that are part of the input sequence i.e not boundary tokens like 1, 2, <bos>, <eos>, <pad>, etc.residue_selector=~(batch_preparer.get_boundary_token_mask(batch["input_ids"]))
# last_sequence_residue_selector: True for residues that are part of the last sequence (excluding boundary tokens)last_sequence_residue_selector=last_sequence_selector&residue_selector# Will yield embeddings for ["AAAAA?C", "MFCC?SEF"] while throwing away embeddings for the boundary tokens# and for homologous sequences in the second instance. Can do similar for logits.last_sequence_embeddings= [embeddings[i, last_sequence_residue_selector[i]] foriinrange(embeddings.shape[0])]

See the cookbook/basic.ipynbOpen In Colab file for a more complete example. We also provide a notebook for embedding analysisOpen In Colab to demonstrate how to use the model to get sequence embeddings using E1Predictor class.

You can also use the cookbook/zero_shot_fitness_prediction.ipynbOpen In Colab file to predict the fitness of zero-shot substitution mutants against a wild type parent.

For performing in-silico site saturation mutagenesis using masked-marginal scoring method, see the cookbook/site_saturation_mutagenesis.ipynbOpen In Colab file. We score all possible single mutants of a given parent sequence in both single sequence and retrieval augmented mode.

Comparison to other models

Average Spearman on Substitution Assays in Protein Gym BenchmarkUnsupervised Contact Map Prediction on CAMEO
Protein Gym ResultsCAMEO Results

Licenses

Your use of the Profluent-E1 model code is governed by the Apache License 2.0, while your use of the Profluent-E1 model weights and the full release of the Profluent-E1 model is governed by a similarly permissive license with additional attribution requirements - see the NOTICE file for details. You can use, share, and modify Profluent-E1 for free, but you must follow our ATTRIBUTION guidelines to give credit, include the license when you share and follow some other basic rules. Profluent is not responsible for what you build, and may terminate your rights to use Profluent-E1 if you breach the license.

  1. Code in src/E1/model/flash_attention_utils.py is adapted from flash-attention project under BSD-3-Clause license.

About

Profluent-E1 family of Protein Encoder Models

Resources

Stars

115 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Profluent-E1

This repository contains the code for the Profluent E1 family of models - our best in class single sequence and retrieval augmented protein representation models. They are designed to be drop-in replacement for ESM family of models. See Section on licenses for the license details.

Changelog

  • 2026-03-09: Modify MSA Sampling code to fix bug. Files modified: src/E1/msa_sampling.py.

Available Models

Model NameModel ParametersHuggingFace Link
E1-150m150MProfluent-Bio/E1-150m
E1-300m300MProfluent-Bio/E1-300m
E1-600m600MProfluent-Bio/E1-600m

Installation

To use the code in this repository, install the dependencies using the following command:

git clone https://github.com/Profluent-AI/E1.git
cd E1 && pip3 install -e .

If you are using GPUs with cuda capability 8.0 or higher (Ampere architecture or higher), we also recommend installing the flash-attn package:

pip3 install flash-attn --no-build-isolation

Compute Requirements

While the models can be run on both CPU and GPU, we recommend using a GPU for faster inference. In addition, the model was trained with BF16 precision, so we recommend using a GPU with BF16 support (CUDA Capability 8.0 or higher; for example, NVIDIA L40, A100, H100, etc.) since that also allows the inference performance to improve using flash attention.

Interactive Usage

The model weights are hosted on Hugging Face and will be downloaded automatically in the following code. The model can be used in both single sequence and retrieval augmented mode. You can use ? as mask token. To use the model in retrieval augmented mode, you can prepend your query with homolog sequences separated by commas.

importtorchfromE1.batch_preparerimportE1BatchPreparerfromE1.modelingimportE1ForMaskedLMmodel=E1ForMaskedLM.from_pretrained("Profluent-Bio/E1-300m").to("cuda:0")
model.eval()
sequences= ["AAAAA?C", "MFCATEEKL,MCCASDF,MFCC?SEF"]
batch_preparer=E1BatchPreparer()
batch=batch_preparer.get_batch_kwargs(sequences, device="cuda:0")
withtorch.autocast("cuda", dtype=torch.bfloat16, enabled=True):
outputs=model(
input_ids=batch["input_ids"],
within_seq_position_ids=batch["within_seq_position_ids"],
global_position_ids=batch["global_position_ids"],
sequence_ids=batch["sequence_ids"],
past_key_values=None,
use_cache=False,
output_attentions=False,
output_hidden_states=False,
)
logits: torch.Tensor=outputs.logits# (B, L, V)embeddings: torch.Tensor=outputs.embeddings# (B, L, E)print(logits)
print(embeddings)
# Boolean Selectors of shape (B, L) to get relevant tokens from logits/embeddings# last_sequence_selector: True for tokens that are part of the last sequence (including boundary tokens) in case of multi-sequence input.last_sequence_selector=batch["sequence_ids"] ==batch["sequence_ids"].max(dim=1)[0][:, None]
# residue_selector: True for tokens that are part of the input sequence i.e not boundary tokens like 1, 2, <bos>, <eos>, <pad>, etc.residue_selector=~(batch_preparer.get_boundary_token_mask(batch["input_ids"]))
# last_sequence_residue_selector: True for residues that are part of the last sequence (excluding boundary tokens)last_sequence_residue_selector=last_sequence_selector&residue_selector# Will yield embeddings for ["AAAAA?C", "MFCC?SEF"] while throwing away embeddings for the boundary tokens# and for homologous sequences in the second instance. Can do similar for logits.last_sequence_embeddings= [embeddings[i, last_sequence_residue_selector[i]] foriinrange(embeddings.shape[0])]

See the cookbook/basic.ipynbOpen In Colab file for a more complete example. We also provide a notebook for embedding analysisOpen In Colab to demonstrate how to use the model to get sequence embeddings using E1Predictor class.

You can also use the cookbook/zero_shot_fitness_prediction.ipynbOpen In Colab file to predict the fitness of zero-shot substitution mutants against a wild type parent.

For performing in-silico site saturation mutagenesis using masked-marginal scoring method, see the cookbook/site_saturation_mutagenesis.ipynbOpen In Colab file. We score all possible single mutants of a given parent sequence in both single sequence and retrieval augmented mode.

Comparison to other models

Average Spearman on Substitution Assays in Protein Gym BenchmarkUnsupervised Contact Map Prediction on CAMEO
Protein Gym ResultsCAMEO Results

Licenses

Your use of the Profluent-E1 model code is governed by the Apache License 2.0, while your use of the Profluent-E1 model weights and the full release of the Profluent-E1 model is governed by a similarly permissive license with additional attribution requirements - see the NOTICE file for details. You can use, share, and modify Profluent-E1 for free, but you must follow our ATTRIBUTION guidelines to give credit, include the license when you share and follow some other basic rules. Profluent is not responsible for what you build, and may terminate your rights to use Profluent-E1 if you breach the license.

  1. Code in src/E1/model/flash_attention_utils.py is adapted from flash-attention project under BSD-3-Clause license.

About

Profluent-E1 family of Protein Encoder Models

Resources

Stars

115 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Profluent-E1

This repository contains the code for the Profluent E1 family of models - our best in class single sequence and retrieval augmented protein representation models. They are designed to be drop-in replacement for ESM family of models. See Section on licenses for the license details.

Changelog

  • 2026-03-09: Modify MSA Sampling code to fix bug. Files modified: src/E1/msa_sampling.py.

Available Models

Model NameModel ParametersHuggingFace Link
E1-150m150MProfluent-Bio/E1-150m
E1-300m300MProfluent-Bio/E1-300m
E1-600m600MProfluent-Bio/E1-600m

Installation

To use the code in this repository, install the dependencies using the following command:

git clone https://github.com/Profluent-AI/E1.git
cd E1 && pip3 install -e .

If you are using GPUs with cuda capability 8.0 or higher (Ampere architecture or higher), we also recommend installing the flash-attn package:

pip3 install flash-attn --no-build-isolation

Compute Requirements

While the models can be run on both CPU and GPU, we recommend using a GPU for faster inference. In addition, the model was trained with BF16 precision, so we recommend using a GPU with BF16 support (CUDA Capability 8.0 or higher; for example, NVIDIA L40, A100, H100, etc.) since that also allows the inference performance to improve using flash attention.

Interactive Usage

The model weights are hosted on Hugging Face and will be downloaded automatically in the following code. The model can be used in both single sequence and retrieval augmented mode. You can use ? as mask token. To use the model in retrieval augmented mode, you can prepend your query with homolog sequences separated by commas.

importtorchfromE1.batch_preparerimportE1BatchPreparerfromE1.modelingimportE1ForMaskedLMmodel=E1ForMaskedLM.from_pretrained("Profluent-Bio/E1-300m").to("cuda:0")
model.eval()
sequences= ["AAAAA?C", "MFCATEEKL,MCCASDF,MFCC?SEF"]
batch_preparer=E1BatchPreparer()
batch=batch_preparer.get_batch_kwargs(sequences, device="cuda:0")
withtorch.autocast("cuda", dtype=torch.bfloat16, enabled=True):
outputs=model(
input_ids=batch["input_ids"],
within_seq_position_ids=batch["within_seq_position_ids"],
global_position_ids=batch["global_position_ids"],
sequence_ids=batch["sequence_ids"],
past_key_values=None,
use_cache=False,
output_attentions=False,
output_hidden_states=False,
)
logits: torch.Tensor=outputs.logits# (B, L, V)embeddings: torch.Tensor=outputs.embeddings# (B, L, E)print(logits)
print(embeddings)
# Boolean Selectors of shape (B, L) to get relevant tokens from logits/embeddings# last_sequence_selector: True for tokens that are part of the last sequence (including boundary tokens) in case of multi-sequence input.last_sequence_selector=batch["sequence_ids"] ==batch["sequence_ids"].max(dim=1)[0][:, None]
# residue_selector: True for tokens that are part of the input sequence i.e not boundary tokens like 1, 2, <bos>, <eos>, <pad>, etc.residue_selector=~(batch_preparer.get_boundary_token_mask(batch["input_ids"]))
# last_sequence_residue_selector: True for residues that are part of the last sequence (excluding boundary tokens)last_sequence_residue_selector=last_sequence_selector&residue_selector# Will yield embeddings for ["AAAAA?C", "MFCC?SEF"] while throwing away embeddings for the boundary tokens# and for homologous sequences in the second instance. Can do similar for logits.last_sequence_embeddings= [embeddings[i, last_sequence_residue_selector[i]] foriinrange(embeddings.shape[0])]

See the cookbook/basic.ipynbOpen In Colab file for a more complete example. We also provide a notebook for embedding analysisOpen In Colab to demonstrate how to use the model to get sequence embeddings using E1Predictor class.

You can also use the cookbook/zero_shot_fitness_prediction.ipynbOpen In Colab file to predict the fitness of zero-shot substitution mutants against a wild type parent.

For performing in-silico site saturation mutagenesis using masked-marginal scoring method, see the cookbook/site_saturation_mutagenesis.ipynbOpen In Colab file. We score all possible single mutants of a given parent sequence in both single sequence and retrieval augmented mode.

Comparison to other models

Average Spearman on Substitution Assays in Protein Gym BenchmarkUnsupervised Contact Map Prediction on CAMEO
Protein Gym ResultsCAMEO Results

Licenses

Your use of the Profluent-E1 model code is governed by the Apache License 2.0, while your use of the Profluent-E1 model weights and the full release of the Profluent-E1 model is governed by a similarly permissive license with additional attribution requirements - see the NOTICE file for details. You can use, share, and modify Profluent-E1 for free, but you must follow our ATTRIBUTION guidelines to give credit, include the license when you share and follow some other basic rules. Profluent is not responsible for what you build, and may terminate your rights to use Profluent-E1 if you breach the license.

  1. Code in src/E1/model/flash_attention_utils.py is adapted from flash-attention project under BSD-3-Clause license.

About

Profluent-E1 family of Protein Encoder Models

Resources

Stars

115 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages