Repository files navigation

Protein2PAM: Protein Language Models for CRISPR-Cas PAM Prediction

PythonbioRxivLicense: Models (CC BY–NC 4.0)License: Code (Polyform Noncommercial 1.0.0)

This repository contains code for Protein2PAM, a tool that predicts CRISPR-Cas PAMs from protein sequences using protein language models. The models and their applications are described in Nayfach, S., Bhatnagar, A., Novichkov, A., et al. (2025). For a browser-based interface, try the Protein2PAM Webserver.

The figure below shows the overall Protein2PAM workflow.

flowchart LR
A[CRISPR-Cas protein] --> B[Protein2PAM model] --> C[Predicted PAM motif]
style A fill:#0f0f0f,stroke:#00ff99,stroke-width:2px,color:#00ff99
style B fill:#0f0f0f,stroke:#00ff99,stroke-width:2px,color:#00ff99
style C fill:#0f0f0f,stroke:#00ff99,stroke-width:2px,color:#00ff99
Loading

This repo has two Protein2PAM implementations:

  • A full implementation in bare PyTorch (under the protein2pam.models submodule) which includes uncertainty estimation and visualization utilities, but requires users to manually download certain modeling resources.
  • A Huggingface implementation (under the protein2pam.huggingface submodule) that includes only the PAM prediction model without uncertainty estimation. This implementation is intended for more low-level usage, but it is integrated with the Huggingface Hub for easier downloading of model weights. You can see the model collection here.

Contents

Prerequisites

Before installing Protein2PAM, ensure you have:

  • Python 3.8+ Available from python.org.
  • NVIDIA GPU with CUDA support:
    Refer to the official NVIDIA installation guide or see our step-by-step instructions here. Use nvidia-smi to ensure GPU(s) are available on your system.
  • pip If pip is not already installed, you can install it using:
    python3 -m ensurepip --upgrade
  • Mamba or Conda:
    You can install mamba using:
    wget https://github.com/conda-forge/miniforge/releases/latest/download/Miniforge3-Linux-x86_64.sh
    bash Miniforge3-Linux-x86_64.sh

Quickstart

Download repo

git clone https://github.com/Profluent-AI/protein2pam
cd protein2pam

Create environment and install

mamba env create -f environment.yml
source activate protein2pam
pip install -e .

Download and unpack database (only necessary if using the protein2pam.models implementation, not the protein2pam.huggingface implementation)

wget https://storage.googleapis.com/protein2pam-x83y9z7q4k/protein2pam_db_v1.0.tar
tar -xvf protein2pam_db_v1.0.tar

Usage

Python usage

Below are examples of how to run Protein2PAM locally through the Python API.

Import protein2pam and list available models:

importprotein2pamprotein2pam.model_info()

For the example, we'll use the main cas9 model, the PID sequence of SpCas9, and the database located at ./protein2pam_db_v1.0

proteins= ["VQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD"]
oracle=protein2pam.PAMOracle(
model_name="cas9", data_dir="../protein2pam_db_v1.0"
)
predictions=oracle.evaluate(proteins)
forpredictioninpredictions:
print(prediction)

We can plot PAM logos using:

forpaminpredictions:
pam.plot_logo(
file="example.png",
side="downstream",
title="example"
)

Alternatively, if you would like to directly use a Huggingface implementation of the PAM prediction model, you can call

importtorchimporttorch.nn.functionalasFfromprotein2pam.huggingfaceimportEsmForSequenceClassification, get_tokenizer# Tokenize proteins so they can be consumed by the modelproteins= ["VQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD"]
tokenizer=get_tokenizer()
encodings=tokenizer.encode_batch(proteins)
device=torch.device("cuda") iftorch.cuda.is_available() elsetorch.device("cpu")
input_batch=dict(
input_ids=torch.tensor([encoding.idsforencodinginencodings], device=device),
attention_mask=torch.tensor([encoding.attention_maskforencodinginencodings], device=device),
)
# Initialize model. You can use any of the supported model names (below) instead of cas9.model=EsmForSequenceClassification.from_pretrained("Profluent-Bio/protein2pam-cas9", device_map=device)
# Get a batch of PAM probability matrices (as a torch tensor) from the model# dimension 0 is batch size, dimension 1 is sequence position, and dimension 2 is nucleotide ID (ordered ACGT)output=model(**input_batch)
pam_probability_matrix=F.softmax(output.logits, dim=-1)

Models

Models used by the Protein2PAM webserver

Model NameInput Protein/DomainCRISPR TypeSamplesLiterature Datasets
cas8Cas8 or Cas10dType I28,410
cas9Cas9 PI-domainType II15,8431-9
cas12Cas12 proteinType V1,72010-21

These models are exposed on the Protein2PAM Webserver. See Data for details on training samples.

Additional models

In addition to the models deployed on the webserver, we also trained several variants for benchmarking and ablation studies, summarized below:

Model NameInput Protein/DomainCRISPR TypeSamplesLiterature DatasetsNotes
cas9_fullCas9 proteinType II15,8431-9Full-length Cas9 trained with literature PAMs
cas9_full_nolitCas9 proteinType II15,731Full-length Cas9, no literature PAMs
cas9_pid_nolitCas9 PI-domainType II15,731No literature PAMs
cas9_pid_nmeCas9 PI-domainType II15,8431-9Nme orthologs upweighted
cas12_no_litCas12 proteinType V1,675Trained without literature PAMs

Data

Sequences and PAM profiles used for training can be downloaded from Hugging Face and Google Cloud:

wget https://storage.googleapis.com/protein2pam-x83y9z7q4k/protein2pam_train_seqs.tsv

The training set consists of protein:PAM pairs curated from evolutionary data and supplemented with experimental PAMs reported in the literature (see REFERENCES.md)

License

This repository contains both software and trained models under different licenses.

The full license texts are available in LICENSES.md.

For commercial licensing inquiries, please contact partnerships@profluent.bio.

Cite this work

If you use Protein2PAM in your research, please cite the following preprint

About

Predicting CRISPR–Cas PAM specificity with protein language models

Resources

Stars

15 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Protein2PAM: Protein Language Models for CRISPR-Cas PAM Prediction

PythonbioRxivLicense: Models (CC BY–NC 4.0)License: Code (Polyform Noncommercial 1.0.0)

This repository contains code for Protein2PAM, a tool that predicts CRISPR-Cas PAMs from protein sequences using protein language models. The models and their applications are described in Nayfach, S., Bhatnagar, A., Novichkov, A., et al. (2025). For a browser-based interface, try the Protein2PAM Webserver.

The figure below shows the overall Protein2PAM workflow.

flowchart LR
A[CRISPR-Cas protein] --> B[Protein2PAM model] --> C[Predicted PAM motif]
style A fill:#0f0f0f,stroke:#00ff99,stroke-width:2px,color:#00ff99
style B fill:#0f0f0f,stroke:#00ff99,stroke-width:2px,color:#00ff99
style C fill:#0f0f0f,stroke:#00ff99,stroke-width:2px,color:#00ff99
Loading

This repo has two Protein2PAM implementations:

  • A full implementation in bare PyTorch (under the protein2pam.models submodule) which includes uncertainty estimation and visualization utilities, but requires users to manually download certain modeling resources.
  • A Huggingface implementation (under the protein2pam.huggingface submodule) that includes only the PAM prediction model without uncertainty estimation. This implementation is intended for more low-level usage, but it is integrated with the Huggingface Hub for easier downloading of model weights. You can see the model collection here.

Contents

Prerequisites

Before installing Protein2PAM, ensure you have:

  • Python 3.8+ Available from python.org.
  • NVIDIA GPU with CUDA support:
    Refer to the official NVIDIA installation guide or see our step-by-step instructions here. Use nvidia-smi to ensure GPU(s) are available on your system.
  • pip If pip is not already installed, you can install it using:
    python3 -m ensurepip --upgrade
  • Mamba or Conda:
    You can install mamba using:
    wget https://github.com/conda-forge/miniforge/releases/latest/download/Miniforge3-Linux-x86_64.sh
    bash Miniforge3-Linux-x86_64.sh

Quickstart

Download repo

git clone https://github.com/Profluent-AI/protein2pam
cd protein2pam

Create environment and install

mamba env create -f environment.yml
source activate protein2pam
pip install -e .

Download and unpack database (only necessary if using the protein2pam.models implementation, not the protein2pam.huggingface implementation)

wget https://storage.googleapis.com/protein2pam-x83y9z7q4k/protein2pam_db_v1.0.tar
tar -xvf protein2pam_db_v1.0.tar

Usage

Python usage

Below are examples of how to run Protein2PAM locally through the Python API.

Import protein2pam and list available models:

importprotein2pamprotein2pam.model_info()

For the example, we'll use the main cas9 model, the PID sequence of SpCas9, and the database located at ./protein2pam_db_v1.0

proteins= ["VQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD"]
oracle=protein2pam.PAMOracle(
model_name="cas9", data_dir="../protein2pam_db_v1.0"
)
predictions=oracle.evaluate(proteins)
forpredictioninpredictions:
print(prediction)

We can plot PAM logos using:

forpaminpredictions:
pam.plot_logo(
file="example.png",
side="downstream",
title="example"
)

Alternatively, if you would like to directly use a Huggingface implementation of the PAM prediction model, you can call

importtorchimporttorch.nn.functionalasFfromprotein2pam.huggingfaceimportEsmForSequenceClassification, get_tokenizer# Tokenize proteins so they can be consumed by the modelproteins= ["VQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD"]
tokenizer=get_tokenizer()
encodings=tokenizer.encode_batch(proteins)
device=torch.device("cuda") iftorch.cuda.is_available() elsetorch.device("cpu")
input_batch=dict(
input_ids=torch.tensor([encoding.idsforencodinginencodings], device=device),
attention_mask=torch.tensor([encoding.attention_maskforencodinginencodings], device=device),
)
# Initialize model. You can use any of the supported model names (below) instead of cas9.model=EsmForSequenceClassification.from_pretrained("Profluent-Bio/protein2pam-cas9", device_map=device)
# Get a batch of PAM probability matrices (as a torch tensor) from the model# dimension 0 is batch size, dimension 1 is sequence position, and dimension 2 is nucleotide ID (ordered ACGT)output=model(**input_batch)
pam_probability_matrix=F.softmax(output.logits, dim=-1)

Models

Models used by the Protein2PAM webserver

Model NameInput Protein/DomainCRISPR TypeSamplesLiterature Datasets
cas8Cas8 or Cas10dType I28,410
cas9Cas9 PI-domainType II15,8431-9
cas12Cas12 proteinType V1,72010-21

These models are exposed on the Protein2PAM Webserver. See Data for details on training samples.

Additional models

In addition to the models deployed on the webserver, we also trained several variants for benchmarking and ablation studies, summarized below:

Model NameInput Protein/DomainCRISPR TypeSamplesLiterature DatasetsNotes
cas9_fullCas9 proteinType II15,8431-9Full-length Cas9 trained with literature PAMs
cas9_full_nolitCas9 proteinType II15,731Full-length Cas9, no literature PAMs
cas9_pid_nolitCas9 PI-domainType II15,731No literature PAMs
cas9_pid_nmeCas9 PI-domainType II15,8431-9Nme orthologs upweighted
cas12_no_litCas12 proteinType V1,675Trained without literature PAMs

Data

Sequences and PAM profiles used for training can be downloaded from Hugging Face and Google Cloud:

wget https://storage.googleapis.com/protein2pam-x83y9z7q4k/protein2pam_train_seqs.tsv

The training set consists of protein:PAM pairs curated from evolutionary data and supplemented with experimental PAMs reported in the literature (see REFERENCES.md)

License

This repository contains both software and trained models under different licenses.

The full license texts are available in LICENSES.md.

For commercial licensing inquiries, please contact partnerships@profluent.bio.

Cite this work

If you use Protein2PAM in your research, please cite the following preprint

About

Predicting CRISPR–Cas PAM specificity with protein language models

Resources

Stars

15 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Protein2PAM: Protein Language Models for CRISPR-Cas PAM Prediction

PythonbioRxivLicense: Models (CC BY–NC 4.0)License: Code (Polyform Noncommercial 1.0.0)

This repository contains code for Protein2PAM, a tool that predicts CRISPR-Cas PAMs from protein sequences using protein language models. The models and their applications are described in Nayfach, S., Bhatnagar, A., Novichkov, A., et al. (2025). For a browser-based interface, try the Protein2PAM Webserver.

The figure below shows the overall Protein2PAM workflow.

flowchart LR
A[CRISPR-Cas protein] --> B[Protein2PAM model] --> C[Predicted PAM motif]
style A fill:#0f0f0f,stroke:#00ff99,stroke-width:2px,color:#00ff99
style B fill:#0f0f0f,stroke:#00ff99,stroke-width:2px,color:#00ff99
style C fill:#0f0f0f,stroke:#00ff99,stroke-width:2px,color:#00ff99
Loading

This repo has two Protein2PAM implementations:

  • A full implementation in bare PyTorch (under the protein2pam.models submodule) which includes uncertainty estimation and visualization utilities, but requires users to manually download certain modeling resources.
  • A Huggingface implementation (under the protein2pam.huggingface submodule) that includes only the PAM prediction model without uncertainty estimation. This implementation is intended for more low-level usage, but it is integrated with the Huggingface Hub for easier downloading of model weights. You can see the model collection here.

Contents

Prerequisites

Before installing Protein2PAM, ensure you have:

  • Python 3.8+ Available from python.org.
  • NVIDIA GPU with CUDA support:
    Refer to the official NVIDIA installation guide or see our step-by-step instructions here. Use nvidia-smi to ensure GPU(s) are available on your system.
  • pip If pip is not already installed, you can install it using:
    python3 -m ensurepip --upgrade
  • Mamba or Conda:
    You can install mamba using:
    wget https://github.com/conda-forge/miniforge/releases/latest/download/Miniforge3-Linux-x86_64.sh
    bash Miniforge3-Linux-x86_64.sh

Quickstart

Download repo

git clone https://github.com/Profluent-AI/protein2pam
cd protein2pam

Create environment and install

mamba env create -f environment.yml
source activate protein2pam
pip install -e .

Download and unpack database (only necessary if using the protein2pam.models implementation, not the protein2pam.huggingface implementation)

wget https://storage.googleapis.com/protein2pam-x83y9z7q4k/protein2pam_db_v1.0.tar
tar -xvf protein2pam_db_v1.0.tar

Usage

Python usage

Below are examples of how to run Protein2PAM locally through the Python API.

Import protein2pam and list available models:

importprotein2pamprotein2pam.model_info()

For the example, we'll use the main cas9 model, the PID sequence of SpCas9, and the database located at ./protein2pam_db_v1.0

proteins= ["VQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD"]
oracle=protein2pam.PAMOracle(
model_name="cas9", data_dir="../protein2pam_db_v1.0"
)
predictions=oracle.evaluate(proteins)
forpredictioninpredictions:
print(prediction)

We can plot PAM logos using:

forpaminpredictions:
pam.plot_logo(
file="example.png",
side="downstream",
title="example"
)

Alternatively, if you would like to directly use a Huggingface implementation of the PAM prediction model, you can call

importtorchimporttorch.nn.functionalasFfromprotein2pam.huggingfaceimportEsmForSequenceClassification, get_tokenizer# Tokenize proteins so they can be consumed by the modelproteins= ["VQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD"]
tokenizer=get_tokenizer()
encodings=tokenizer.encode_batch(proteins)
device=torch.device("cuda") iftorch.cuda.is_available() elsetorch.device("cpu")
input_batch=dict(
input_ids=torch.tensor([encoding.idsforencodinginencodings], device=device),
attention_mask=torch.tensor([encoding.attention_maskforencodinginencodings], device=device),
)
# Initialize model. You can use any of the supported model names (below) instead of cas9.model=EsmForSequenceClassification.from_pretrained("Profluent-Bio/protein2pam-cas9", device_map=device)
# Get a batch of PAM probability matrices (as a torch tensor) from the model# dimension 0 is batch size, dimension 1 is sequence position, and dimension 2 is nucleotide ID (ordered ACGT)output=model(**input_batch)
pam_probability_matrix=F.softmax(output.logits, dim=-1)

Models

Models used by the Protein2PAM webserver

Model NameInput Protein/DomainCRISPR TypeSamplesLiterature Datasets
cas8Cas8 or Cas10dType I28,410
cas9Cas9 PI-domainType II15,8431-9
cas12Cas12 proteinType V1,72010-21

These models are exposed on the Protein2PAM Webserver. See Data for details on training samples.

Additional models

In addition to the models deployed on the webserver, we also trained several variants for benchmarking and ablation studies, summarized below:

Model NameInput Protein/DomainCRISPR TypeSamplesLiterature DatasetsNotes
cas9_fullCas9 proteinType II15,8431-9Full-length Cas9 trained with literature PAMs
cas9_full_nolitCas9 proteinType II15,731Full-length Cas9, no literature PAMs
cas9_pid_nolitCas9 PI-domainType II15,731No literature PAMs
cas9_pid_nmeCas9 PI-domainType II15,8431-9Nme orthologs upweighted
cas12_no_litCas12 proteinType V1,675Trained without literature PAMs

Data

Sequences and PAM profiles used for training can be downloaded from Hugging Face and Google Cloud:

wget https://storage.googleapis.com/protein2pam-x83y9z7q4k/protein2pam_train_seqs.tsv

The training set consists of protein:PAM pairs curated from evolutionary data and supplemented with experimental PAMs reported in the literature (see REFERENCES.md)

License

This repository contains both software and trained models under different licenses.

The full license texts are available in LICENSES.md.

For commercial licensing inquiries, please contact partnerships@profluent.bio.

Cite this work

If you use Protein2PAM in your research, please cite the following preprint

About

Predicting CRISPR–Cas PAM specificity with protein language models

Resources

Stars

15 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Protein2PAM: Protein Language Models for CRISPR-Cas PAM Prediction

PythonbioRxivLicense: Models (CC BY–NC 4.0)License: Code (Polyform Noncommercial 1.0.0)

This repository contains code for Protein2PAM, a tool that predicts CRISPR-Cas PAMs from protein sequences using protein language models. The models and their applications are described in Nayfach, S., Bhatnagar, A., Novichkov, A., et al. (2025). For a browser-based interface, try the Protein2PAM Webserver.

The figure below shows the overall Protein2PAM workflow.

flowchart LR
A[CRISPR-Cas protein] --> B[Protein2PAM model] --> C[Predicted PAM motif]
style A fill:#0f0f0f,stroke:#00ff99,stroke-width:2px,color:#00ff99
style B fill:#0f0f0f,stroke:#00ff99,stroke-width:2px,color:#00ff99
style C fill:#0f0f0f,stroke:#00ff99,stroke-width:2px,color:#00ff99
Loading

This repo has two Protein2PAM implementations:

  • A full implementation in bare PyTorch (under the protein2pam.models submodule) which includes uncertainty estimation and visualization utilities, but requires users to manually download certain modeling resources.
  • A Huggingface implementation (under the protein2pam.huggingface submodule) that includes only the PAM prediction model without uncertainty estimation. This implementation is intended for more low-level usage, but it is integrated with the Huggingface Hub for easier downloading of model weights. You can see the model collection here.

Contents

Prerequisites

Before installing Protein2PAM, ensure you have:

  • Python 3.8+ Available from python.org.
  • NVIDIA GPU with CUDA support:
    Refer to the official NVIDIA installation guide or see our step-by-step instructions here. Use nvidia-smi to ensure GPU(s) are available on your system.
  • pip If pip is not already installed, you can install it using:
    python3 -m ensurepip --upgrade
  • Mamba or Conda:
    You can install mamba using:
    wget https://github.com/conda-forge/miniforge/releases/latest/download/Miniforge3-Linux-x86_64.sh
    bash Miniforge3-Linux-x86_64.sh

Quickstart

Download repo

git clone https://github.com/Profluent-AI/protein2pam
cd protein2pam

Create environment and install

mamba env create -f environment.yml
source activate protein2pam
pip install -e .

Download and unpack database (only necessary if using the protein2pam.models implementation, not the protein2pam.huggingface implementation)

wget https://storage.googleapis.com/protein2pam-x83y9z7q4k/protein2pam_db_v1.0.tar
tar -xvf protein2pam_db_v1.0.tar

Usage

Python usage

Below are examples of how to run Protein2PAM locally through the Python API.

Import protein2pam and list available models:

importprotein2pamprotein2pam.model_info()

For the example, we'll use the main cas9 model, the PID sequence of SpCas9, and the database located at ./protein2pam_db_v1.0

proteins= ["VQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD"]
oracle=protein2pam.PAMOracle(
model_name="cas9", data_dir="../protein2pam_db_v1.0"
)
predictions=oracle.evaluate(proteins)
forpredictioninpredictions:
print(prediction)

We can plot PAM logos using:

forpaminpredictions:
pam.plot_logo(
file="example.png",
side="downstream",
title="example"
)

Alternatively, if you would like to directly use a Huggingface implementation of the PAM prediction model, you can call

importtorchimporttorch.nn.functionalasFfromprotein2pam.huggingfaceimportEsmForSequenceClassification, get_tokenizer# Tokenize proteins so they can be consumed by the modelproteins= ["VQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD"]
tokenizer=get_tokenizer()
encodings=tokenizer.encode_batch(proteins)
device=torch.device("cuda") iftorch.cuda.is_available() elsetorch.device("cpu")
input_batch=dict(
input_ids=torch.tensor([encoding.idsforencodinginencodings], device=device),
attention_mask=torch.tensor([encoding.attention_maskforencodinginencodings], device=device),
)
# Initialize model. You can use any of the supported model names (below) instead of cas9.model=EsmForSequenceClassification.from_pretrained("Profluent-Bio/protein2pam-cas9", device_map=device)
# Get a batch of PAM probability matrices (as a torch tensor) from the model# dimension 0 is batch size, dimension 1 is sequence position, and dimension 2 is nucleotide ID (ordered ACGT)output=model(**input_batch)
pam_probability_matrix=F.softmax(output.logits, dim=-1)

Models

Models used by the Protein2PAM webserver

Model NameInput Protein/DomainCRISPR TypeSamplesLiterature Datasets
cas8Cas8 or Cas10dType I28,410
cas9Cas9 PI-domainType II15,8431-9
cas12Cas12 proteinType V1,72010-21

These models are exposed on the Protein2PAM Webserver. See Data for details on training samples.

Additional models

In addition to the models deployed on the webserver, we also trained several variants for benchmarking and ablation studies, summarized below:

Model NameInput Protein/DomainCRISPR TypeSamplesLiterature DatasetsNotes
cas9_fullCas9 proteinType II15,8431-9Full-length Cas9 trained with literature PAMs
cas9_full_nolitCas9 proteinType II15,731Full-length Cas9, no literature PAMs
cas9_pid_nolitCas9 PI-domainType II15,731No literature PAMs
cas9_pid_nmeCas9 PI-domainType II15,8431-9Nme orthologs upweighted
cas12_no_litCas12 proteinType V1,675Trained without literature PAMs

Data

Sequences and PAM profiles used for training can be downloaded from Hugging Face and Google Cloud:

wget https://storage.googleapis.com/protein2pam-x83y9z7q4k/protein2pam_train_seqs.tsv

The training set consists of protein:PAM pairs curated from evolutionary data and supplemented with experimental PAMs reported in the literature (see REFERENCES.md)

License

This repository contains both software and trained models under different licenses.

The full license texts are available in LICENSES.md.

For commercial licensing inquiries, please contact partnerships@profluent.bio.

Cite this work

If you use Protein2PAM in your research, please cite the following preprint

About

Predicting CRISPR–Cas PAM specificity with protein language models

Resources

Stars

15 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Protein2PAM: Protein Language Models for CRISPR-Cas PAM Prediction

PythonbioRxivLicense: Models (CC BY–NC 4.0)License: Code (Polyform Noncommercial 1.0.0)

This repository contains code for Protein2PAM, a tool that predicts CRISPR-Cas PAMs from protein sequences using protein language models. The models and their applications are described in Nayfach, S., Bhatnagar, A., Novichkov, A., et al. (2025). For a browser-based interface, try the Protein2PAM Webserver.

The figure below shows the overall Protein2PAM workflow.

flowchart LR
A[CRISPR-Cas protein] --> B[Protein2PAM model] --> C[Predicted PAM motif]
style A fill:#0f0f0f,stroke:#00ff99,stroke-width:2px,color:#00ff99
style B fill:#0f0f0f,stroke:#00ff99,stroke-width:2px,color:#00ff99
style C fill:#0f0f0f,stroke:#00ff99,stroke-width:2px,color:#00ff99
Loading

This repo has two Protein2PAM implementations:

  • A full implementation in bare PyTorch (under the protein2pam.models submodule) which includes uncertainty estimation and visualization utilities, but requires users to manually download certain modeling resources.
  • A Huggingface implementation (under the protein2pam.huggingface submodule) that includes only the PAM prediction model without uncertainty estimation. This implementation is intended for more low-level usage, but it is integrated with the Huggingface Hub for easier downloading of model weights. You can see the model collection here.

Contents

Prerequisites

Before installing Protein2PAM, ensure you have:

  • Python 3.8+ Available from python.org.
  • NVIDIA GPU with CUDA support:
    Refer to the official NVIDIA installation guide or see our step-by-step instructions here. Use nvidia-smi to ensure GPU(s) are available on your system.
  • pip If pip is not already installed, you can install it using:
    python3 -m ensurepip --upgrade
  • Mamba or Conda:
    You can install mamba using:
    wget https://github.com/conda-forge/miniforge/releases/latest/download/Miniforge3-Linux-x86_64.sh
    bash Miniforge3-Linux-x86_64.sh

Quickstart

Download repo

git clone https://github.com/Profluent-AI/protein2pam
cd protein2pam

Create environment and install

mamba env create -f environment.yml
source activate protein2pam
pip install -e .

Download and unpack database (only necessary if using the protein2pam.models implementation, not the protein2pam.huggingface implementation)

wget https://storage.googleapis.com/protein2pam-x83y9z7q4k/protein2pam_db_v1.0.tar
tar -xvf protein2pam_db_v1.0.tar

Usage

Python usage

Below are examples of how to run Protein2PAM locally through the Python API.

Import protein2pam and list available models:

importprotein2pamprotein2pam.model_info()

For the example, we'll use the main cas9 model, the PID sequence of SpCas9, and the database located at ./protein2pam_db_v1.0

proteins= ["VQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD"]
oracle=protein2pam.PAMOracle(
model_name="cas9", data_dir="../protein2pam_db_v1.0"
)
predictions=oracle.evaluate(proteins)
forpredictioninpredictions:
print(prediction)

We can plot PAM logos using:

forpaminpredictions:
pam.plot_logo(
file="example.png",
side="downstream",
title="example"
)

Alternatively, if you would like to directly use a Huggingface implementation of the PAM prediction model, you can call

importtorchimporttorch.nn.functionalasFfromprotein2pam.huggingfaceimportEsmForSequenceClassification, get_tokenizer# Tokenize proteins so they can be consumed by the modelproteins= ["VQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD"]
tokenizer=get_tokenizer()
encodings=tokenizer.encode_batch(proteins)
device=torch.device("cuda") iftorch.cuda.is_available() elsetorch.device("cpu")
input_batch=dict(
input_ids=torch.tensor([encoding.idsforencodinginencodings], device=device),
attention_mask=torch.tensor([encoding.attention_maskforencodinginencodings], device=device),
)
# Initialize model. You can use any of the supported model names (below) instead of cas9.model=EsmForSequenceClassification.from_pretrained("Profluent-Bio/protein2pam-cas9", device_map=device)
# Get a batch of PAM probability matrices (as a torch tensor) from the model# dimension 0 is batch size, dimension 1 is sequence position, and dimension 2 is nucleotide ID (ordered ACGT)output=model(**input_batch)
pam_probability_matrix=F.softmax(output.logits, dim=-1)

Models

Models used by the Protein2PAM webserver

Model NameInput Protein/DomainCRISPR TypeSamplesLiterature Datasets
cas8Cas8 or Cas10dType I28,410
cas9Cas9 PI-domainType II15,8431-9
cas12Cas12 proteinType V1,72010-21

These models are exposed on the Protein2PAM Webserver. See Data for details on training samples.

Additional models

In addition to the models deployed on the webserver, we also trained several variants for benchmarking and ablation studies, summarized below:

Model NameInput Protein/DomainCRISPR TypeSamplesLiterature DatasetsNotes
cas9_fullCas9 proteinType II15,8431-9Full-length Cas9 trained with literature PAMs
cas9_full_nolitCas9 proteinType II15,731Full-length Cas9, no literature PAMs
cas9_pid_nolitCas9 PI-domainType II15,731No literature PAMs
cas9_pid_nmeCas9 PI-domainType II15,8431-9Nme orthologs upweighted
cas12_no_litCas12 proteinType V1,675Trained without literature PAMs

Data

Sequences and PAM profiles used for training can be downloaded from Hugging Face and Google Cloud:

wget https://storage.googleapis.com/protein2pam-x83y9z7q4k/protein2pam_train_seqs.tsv

The training set consists of protein:PAM pairs curated from evolutionary data and supplemented with experimental PAMs reported in the literature (see REFERENCES.md)

License

This repository contains both software and trained models under different licenses.

The full license texts are available in LICENSES.md.

For commercial licensing inquiries, please contact partnerships@profluent.bio.

Cite this work

If you use Protein2PAM in your research, please cite the following preprint

About

Predicting CRISPR–Cas PAM specificity with protein language models

Resources

Stars

15 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Protein2PAM: Protein Language Models for CRISPR-Cas PAM Prediction

PythonbioRxivLicense: Models (CC BY–NC 4.0)License: Code (Polyform Noncommercial 1.0.0)

This repository contains code for Protein2PAM, a tool that predicts CRISPR-Cas PAMs from protein sequences using protein language models. The models and their applications are described in Nayfach, S., Bhatnagar, A., Novichkov, A., et al. (2025). For a browser-based interface, try the Protein2PAM Webserver.

The figure below shows the overall Protein2PAM workflow.

flowchart LR
A[CRISPR-Cas protein] --> B[Protein2PAM model] --> C[Predicted PAM motif]
style A fill:#0f0f0f,stroke:#00ff99,stroke-width:2px,color:#00ff99
style B fill:#0f0f0f,stroke:#00ff99,stroke-width:2px,color:#00ff99
style C fill:#0f0f0f,stroke:#00ff99,stroke-width:2px,color:#00ff99
Loading

This repo has two Protein2PAM implementations:

  • A full implementation in bare PyTorch (under the protein2pam.models submodule) which includes uncertainty estimation and visualization utilities, but requires users to manually download certain modeling resources.
  • A Huggingface implementation (under the protein2pam.huggingface submodule) that includes only the PAM prediction model without uncertainty estimation. This implementation is intended for more low-level usage, but it is integrated with the Huggingface Hub for easier downloading of model weights. You can see the model collection here.

Contents

Prerequisites

Before installing Protein2PAM, ensure you have:

  • Python 3.8+ Available from python.org.
  • NVIDIA GPU with CUDA support:
    Refer to the official NVIDIA installation guide or see our step-by-step instructions here. Use nvidia-smi to ensure GPU(s) are available on your system.
  • pip If pip is not already installed, you can install it using:
    python3 -m ensurepip --upgrade
  • Mamba or Conda:
    You can install mamba using:
    wget https://github.com/conda-forge/miniforge/releases/latest/download/Miniforge3-Linux-x86_64.sh
    bash Miniforge3-Linux-x86_64.sh

Quickstart

Download repo

git clone https://github.com/Profluent-AI/protein2pam
cd protein2pam

Create environment and install

mamba env create -f environment.yml
source activate protein2pam
pip install -e .

Download and unpack database (only necessary if using the protein2pam.models implementation, not the protein2pam.huggingface implementation)

wget https://storage.googleapis.com/protein2pam-x83y9z7q4k/protein2pam_db_v1.0.tar
tar -xvf protein2pam_db_v1.0.tar

Usage

Python usage

Below are examples of how to run Protein2PAM locally through the Python API.

Import protein2pam and list available models:

importprotein2pamprotein2pam.model_info()

For the example, we'll use the main cas9 model, the PID sequence of SpCas9, and the database located at ./protein2pam_db_v1.0

proteins= ["VQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD"]
oracle=protein2pam.PAMOracle(
model_name="cas9", data_dir="../protein2pam_db_v1.0"
)
predictions=oracle.evaluate(proteins)
forpredictioninpredictions:
print(prediction)

We can plot PAM logos using:

forpaminpredictions:
pam.plot_logo(
file="example.png",
side="downstream",
title="example"
)

Alternatively, if you would like to directly use a Huggingface implementation of the PAM prediction model, you can call

importtorchimporttorch.nn.functionalasFfromprotein2pam.huggingfaceimportEsmForSequenceClassification, get_tokenizer# Tokenize proteins so they can be consumed by the modelproteins= ["VQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD"]
tokenizer=get_tokenizer()
encodings=tokenizer.encode_batch(proteins)
device=torch.device("cuda") iftorch.cuda.is_available() elsetorch.device("cpu")
input_batch=dict(
input_ids=torch.tensor([encoding.idsforencodinginencodings], device=device),
attention_mask=torch.tensor([encoding.attention_maskforencodinginencodings], device=device),
)
# Initialize model. You can use any of the supported model names (below) instead of cas9.model=EsmForSequenceClassification.from_pretrained("Profluent-Bio/protein2pam-cas9", device_map=device)
# Get a batch of PAM probability matrices (as a torch tensor) from the model# dimension 0 is batch size, dimension 1 is sequence position, and dimension 2 is nucleotide ID (ordered ACGT)output=model(**input_batch)
pam_probability_matrix=F.softmax(output.logits, dim=-1)

Models

Models used by the Protein2PAM webserver

Model NameInput Protein/DomainCRISPR TypeSamplesLiterature Datasets
cas8Cas8 or Cas10dType I28,410
cas9Cas9 PI-domainType II15,8431-9
cas12Cas12 proteinType V1,72010-21

These models are exposed on the Protein2PAM Webserver. See Data for details on training samples.

Additional models

In addition to the models deployed on the webserver, we also trained several variants for benchmarking and ablation studies, summarized below:

Model NameInput Protein/DomainCRISPR TypeSamplesLiterature DatasetsNotes
cas9_fullCas9 proteinType II15,8431-9Full-length Cas9 trained with literature PAMs
cas9_full_nolitCas9 proteinType II15,731Full-length Cas9, no literature PAMs
cas9_pid_nolitCas9 PI-domainType II15,731No literature PAMs
cas9_pid_nmeCas9 PI-domainType II15,8431-9Nme orthologs upweighted
cas12_no_litCas12 proteinType V1,675Trained without literature PAMs

Data

Sequences and PAM profiles used for training can be downloaded from Hugging Face and Google Cloud:

wget https://storage.googleapis.com/protein2pam-x83y9z7q4k/protein2pam_train_seqs.tsv

The training set consists of protein:PAM pairs curated from evolutionary data and supplemented with experimental PAMs reported in the literature (see REFERENCES.md)

License

This repository contains both software and trained models under different licenses.

The full license texts are available in LICENSES.md.

For commercial licensing inquiries, please contact partnerships@profluent.bio.

Cite this work

If you use Protein2PAM in your research, please cite the following preprint

About

Predicting CRISPR–Cas PAM specificity with protein language models

Resources

Stars

15 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Protein2PAM: Protein Language Models for CRISPR-Cas PAM Prediction

PythonbioRxivLicense: Models (CC BY–NC 4.0)License: Code (Polyform Noncommercial 1.0.0)

This repository contains code for Protein2PAM, a tool that predicts CRISPR-Cas PAMs from protein sequences using protein language models. The models and their applications are described in Nayfach, S., Bhatnagar, A., Novichkov, A., et al. (2025). For a browser-based interface, try the Protein2PAM Webserver.

The figure below shows the overall Protein2PAM workflow.

flowchart LR
A[CRISPR-Cas protein] --> B[Protein2PAM model] --> C[Predicted PAM motif]
style A fill:#0f0f0f,stroke:#00ff99,stroke-width:2px,color:#00ff99
style B fill:#0f0f0f,stroke:#00ff99,stroke-width:2px,color:#00ff99
style C fill:#0f0f0f,stroke:#00ff99,stroke-width:2px,color:#00ff99
Loading

This repo has two Protein2PAM implementations:

  • A full implementation in bare PyTorch (under the protein2pam.models submodule) which includes uncertainty estimation and visualization utilities, but requires users to manually download certain modeling resources.
  • A Huggingface implementation (under the protein2pam.huggingface submodule) that includes only the PAM prediction model without uncertainty estimation. This implementation is intended for more low-level usage, but it is integrated with the Huggingface Hub for easier downloading of model weights. You can see the model collection here.

Contents

Prerequisites

Before installing Protein2PAM, ensure you have:

  • Python 3.8+ Available from python.org.
  • NVIDIA GPU with CUDA support:
    Refer to the official NVIDIA installation guide or see our step-by-step instructions here. Use nvidia-smi to ensure GPU(s) are available on your system.
  • pip If pip is not already installed, you can install it using:
    python3 -m ensurepip --upgrade
  • Mamba or Conda:
    You can install mamba using:
    wget https://github.com/conda-forge/miniforge/releases/latest/download/Miniforge3-Linux-x86_64.sh
    bash Miniforge3-Linux-x86_64.sh

Quickstart

Download repo

git clone https://github.com/Profluent-AI/protein2pam
cd protein2pam

Create environment and install

mamba env create -f environment.yml
source activate protein2pam
pip install -e .

Download and unpack database (only necessary if using the protein2pam.models implementation, not the protein2pam.huggingface implementation)

wget https://storage.googleapis.com/protein2pam-x83y9z7q4k/protein2pam_db_v1.0.tar
tar -xvf protein2pam_db_v1.0.tar

Usage

Python usage

Below are examples of how to run Protein2PAM locally through the Python API.

Import protein2pam and list available models:

importprotein2pamprotein2pam.model_info()

For the example, we'll use the main cas9 model, the PID sequence of SpCas9, and the database located at ./protein2pam_db_v1.0

proteins= ["VQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD"]
oracle=protein2pam.PAMOracle(
model_name="cas9", data_dir="../protein2pam_db_v1.0"
)
predictions=oracle.evaluate(proteins)
forpredictioninpredictions:
print(prediction)

We can plot PAM logos using:

forpaminpredictions:
pam.plot_logo(
file="example.png",
side="downstream",
title="example"
)

Alternatively, if you would like to directly use a Huggingface implementation of the PAM prediction model, you can call

importtorchimporttorch.nn.functionalasFfromprotein2pam.huggingfaceimportEsmForSequenceClassification, get_tokenizer# Tokenize proteins so they can be consumed by the modelproteins= ["VQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD"]
tokenizer=get_tokenizer()
encodings=tokenizer.encode_batch(proteins)
device=torch.device("cuda") iftorch.cuda.is_available() elsetorch.device("cpu")
input_batch=dict(
input_ids=torch.tensor([encoding.idsforencodinginencodings], device=device),
attention_mask=torch.tensor([encoding.attention_maskforencodinginencodings], device=device),
)
# Initialize model. You can use any of the supported model names (below) instead of cas9.model=EsmForSequenceClassification.from_pretrained("Profluent-Bio/protein2pam-cas9", device_map=device)
# Get a batch of PAM probability matrices (as a torch tensor) from the model# dimension 0 is batch size, dimension 1 is sequence position, and dimension 2 is nucleotide ID (ordered ACGT)output=model(**input_batch)
pam_probability_matrix=F.softmax(output.logits, dim=-1)

Models

Models used by the Protein2PAM webserver

Model NameInput Protein/DomainCRISPR TypeSamplesLiterature Datasets
cas8Cas8 or Cas10dType I28,410
cas9Cas9 PI-domainType II15,8431-9
cas12Cas12 proteinType V1,72010-21

These models are exposed on the Protein2PAM Webserver. See Data for details on training samples.

Additional models

In addition to the models deployed on the webserver, we also trained several variants for benchmarking and ablation studies, summarized below:

Model NameInput Protein/DomainCRISPR TypeSamplesLiterature DatasetsNotes
cas9_fullCas9 proteinType II15,8431-9Full-length Cas9 trained with literature PAMs
cas9_full_nolitCas9 proteinType II15,731Full-length Cas9, no literature PAMs
cas9_pid_nolitCas9 PI-domainType II15,731No literature PAMs
cas9_pid_nmeCas9 PI-domainType II15,8431-9Nme orthologs upweighted
cas12_no_litCas12 proteinType V1,675Trained without literature PAMs

Data

Sequences and PAM profiles used for training can be downloaded from Hugging Face and Google Cloud:

wget https://storage.googleapis.com/protein2pam-x83y9z7q4k/protein2pam_train_seqs.tsv

The training set consists of protein:PAM pairs curated from evolutionary data and supplemented with experimental PAMs reported in the literature (see REFERENCES.md)

License

This repository contains both software and trained models under different licenses.

The full license texts are available in LICENSES.md.

For commercial licensing inquiries, please contact partnerships@profluent.bio.

Cite this work

If you use Protein2PAM in your research, please cite the following preprint

About

Predicting CRISPR–Cas PAM specificity with protein language models

Resources

Stars

15 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Protein2PAM: Protein Language Models for CRISPR-Cas PAM Prediction

PythonbioRxivLicense: Models (CC BY–NC 4.0)License: Code (Polyform Noncommercial 1.0.0)

This repository contains code for Protein2PAM, a tool that predicts CRISPR-Cas PAMs from protein sequences using protein language models. The models and their applications are described in Nayfach, S., Bhatnagar, A., Novichkov, A., et al. (2025). For a browser-based interface, try the Protein2PAM Webserver.

The figure below shows the overall Protein2PAM workflow.

flowchart LR
A[CRISPR-Cas protein] --> B[Protein2PAM model] --> C[Predicted PAM motif]
style A fill:#0f0f0f,stroke:#00ff99,stroke-width:2px,color:#00ff99
style B fill:#0f0f0f,stroke:#00ff99,stroke-width:2px,color:#00ff99
style C fill:#0f0f0f,stroke:#00ff99,stroke-width:2px,color:#00ff99
Loading

This repo has two Protein2PAM implementations:

  • A full implementation in bare PyTorch (under the protein2pam.models submodule) which includes uncertainty estimation and visualization utilities, but requires users to manually download certain modeling resources.
  • A Huggingface implementation (under the protein2pam.huggingface submodule) that includes only the PAM prediction model without uncertainty estimation. This implementation is intended for more low-level usage, but it is integrated with the Huggingface Hub for easier downloading of model weights. You can see the model collection here.

Contents

Prerequisites

Before installing Protein2PAM, ensure you have:

  • Python 3.8+ Available from python.org.
  • NVIDIA GPU with CUDA support:
    Refer to the official NVIDIA installation guide or see our step-by-step instructions here. Use nvidia-smi to ensure GPU(s) are available on your system.
  • pip If pip is not already installed, you can install it using:
    python3 -m ensurepip --upgrade
  • Mamba or Conda:
    You can install mamba using:
    wget https://github.com/conda-forge/miniforge/releases/latest/download/Miniforge3-Linux-x86_64.sh
    bash Miniforge3-Linux-x86_64.sh

Quickstart

Download repo

git clone https://github.com/Profluent-AI/protein2pam
cd protein2pam

Create environment and install

mamba env create -f environment.yml
source activate protein2pam
pip install -e .

Download and unpack database (only necessary if using the protein2pam.models implementation, not the protein2pam.huggingface implementation)

wget https://storage.googleapis.com/protein2pam-x83y9z7q4k/protein2pam_db_v1.0.tar
tar -xvf protein2pam_db_v1.0.tar

Usage

Python usage

Below are examples of how to run Protein2PAM locally through the Python API.

Import protein2pam and list available models:

importprotein2pamprotein2pam.model_info()

For the example, we'll use the main cas9 model, the PID sequence of SpCas9, and the database located at ./protein2pam_db_v1.0

proteins= ["VQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD"]
oracle=protein2pam.PAMOracle(
model_name="cas9", data_dir="../protein2pam_db_v1.0"
)
predictions=oracle.evaluate(proteins)
forpredictioninpredictions:
print(prediction)

We can plot PAM logos using:

forpaminpredictions:
pam.plot_logo(
file="example.png",
side="downstream",
title="example"
)

Alternatively, if you would like to directly use a Huggingface implementation of the PAM prediction model, you can call

importtorchimporttorch.nn.functionalasFfromprotein2pam.huggingfaceimportEsmForSequenceClassification, get_tokenizer# Tokenize proteins so they can be consumed by the modelproteins= ["VQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD"]
tokenizer=get_tokenizer()
encodings=tokenizer.encode_batch(proteins)
device=torch.device("cuda") iftorch.cuda.is_available() elsetorch.device("cpu")
input_batch=dict(
input_ids=torch.tensor([encoding.idsforencodinginencodings], device=device),
attention_mask=torch.tensor([encoding.attention_maskforencodinginencodings], device=device),
)
# Initialize model. You can use any of the supported model names (below) instead of cas9.model=EsmForSequenceClassification.from_pretrained("Profluent-Bio/protein2pam-cas9", device_map=device)
# Get a batch of PAM probability matrices (as a torch tensor) from the model# dimension 0 is batch size, dimension 1 is sequence position, and dimension 2 is nucleotide ID (ordered ACGT)output=model(**input_batch)
pam_probability_matrix=F.softmax(output.logits, dim=-1)

Models

Models used by the Protein2PAM webserver

Model NameInput Protein/DomainCRISPR TypeSamplesLiterature Datasets
cas8Cas8 or Cas10dType I28,410
cas9Cas9 PI-domainType II15,8431-9
cas12Cas12 proteinType V1,72010-21

These models are exposed on the Protein2PAM Webserver. See Data for details on training samples.

Additional models

In addition to the models deployed on the webserver, we also trained several variants for benchmarking and ablation studies, summarized below:

Model NameInput Protein/DomainCRISPR TypeSamplesLiterature DatasetsNotes
cas9_fullCas9 proteinType II15,8431-9Full-length Cas9 trained with literature PAMs
cas9_full_nolitCas9 proteinType II15,731Full-length Cas9, no literature PAMs
cas9_pid_nolitCas9 PI-domainType II15,731No literature PAMs
cas9_pid_nmeCas9 PI-domainType II15,8431-9Nme orthologs upweighted
cas12_no_litCas12 proteinType V1,675Trained without literature PAMs

Data

Sequences and PAM profiles used for training can be downloaded from Hugging Face and Google Cloud:

wget https://storage.googleapis.com/protein2pam-x83y9z7q4k/protein2pam_train_seqs.tsv

The training set consists of protein:PAM pairs curated from evolutionary data and supplemented with experimental PAMs reported in the literature (see REFERENCES.md)

License

This repository contains both software and trained models under different licenses.

The full license texts are available in LICENSES.md.

For commercial licensing inquiries, please contact partnerships@profluent.bio.

Cite this work

If you use Protein2PAM in your research, please cite the following preprint

About

Predicting CRISPR–Cas PAM specificity with protein language models

Resources

Stars

15 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages