Repository files navigation

ProGen3

This repository contains code to access the ProGen3 family of models as well as cli interface to run these models for scoring and generation. The model weights are released under a non-commercial usage license. Please see the LICENSE section below.

Models Available

More info on ProGen3 Models can be found here. In the subsequent sections, you can provide a model name from following table where needed.

Model NameModel ParametersHuggingFace Link
Profluent-Bio/progen3-112m112MProfluent-Bio/progen3-112m
Profluent-Bio/progen3-219m219MProfluent-Bio/progen3-219m
Profluent-Bio/progen3-339m339MProfluent-Bio/progen3-339m
Profluent-Bio/progen3-762m762MProfluent-Bio/progen3-762m
Profluent-Bio/progen3-1b1BProfluent-Bio/progen3-1b
Profluent-Bio/progen3-3b3BProfluent-Bio/progen3-3b

Installation

Local Usage

ProGen3 family of models require atleast one GPU device (we have tested it only on a A100/H100 with atleast 40GB of VRAM; GPUs from prior generations may not work due to lack of support for bf16 precision and flash attention kernels) be available.

  1. Clone this repo and run bash setup.sh

Docker Usage

We also provide a docker image ghcr.io/profluent-ai/progen3:v0.1.0 that comes with certain libraries pre-installed to reduce installation time. This image still needs to be run on a machine with GPU device.

Within the container, follow the steps in the section Local Usage.

Interactive Usage

importtorchfromprogen3.modelingimportProGen3ForCausalLMfromprogen3.batch_preparerimportProGen3BatchPreparerfromprogen3.scorerimportProGen3Scorermodel=ProGen3ForCausalLM.from_pretrained("Profluent-Bio/progen3-3b", torch_dtype=torch.bfloat16)
model=model.eval().to("cuda:0")
batch_preparer=ProGen3BatchPreparer()
# Direct Usagesequence="MALWMRLLPLLALLALWGPDPAAAFVNQHLCGSHLVEALYLVCGERGFFYTPKTRREAEDLQVGQVELGGGPGAGSLQPLALEGSLQKRGIVEQCCTSICSLYQLENYCN"inputs=batch_preparer.get_batch_kwargs([sequence], device="cuda:0", reverse=False)
outputs=model(**inputs, return_dict=True)
print(outputs.logits)
# Usage with scorer (returns averaged log likelihood of forward and reverse direction)# Would suggest using Scoring CLI below if scoring very large number of sequencesscorer=ProGen3Scorer(model=model)
scores=scorer.score_batch(sequences=[sequence])
print(scores["log_likelihood"][0])

Scoring CLI

Prepare the fasta file (described in the next section) containing sequences to be scored. Then, run the following command (modifying the arguments as necessary):

torchrun --nproc-per-node=gpu -m progen3.tools.score \
--fasta-path sequences.fasta \
--output-path scores.csv \
--model-name Profluent-Bio/progen3-3b \
--fsdp

See an example in directory examples/protein_gym_zero_shot on how we use it to evaluate spearman score on protein gym assays.

bash examples/protein_gym_zero_shot/download_and_prepare.sh
# Run for individual assay
bash examples/protein_gym_zero_shot/run.sh A0A140D2T1_ZIKV_Sourisseau_2019 Profluent-Bio/progen3-3b
# Run for all assays and print aggregated spearman
bash examples/protein_gym_zero_shot/run.sh all Profluent-Bio/progen3-3b

Input File Formats

  • fasta-path : A path to fasta file

Each fasta file should be of form:

>seq_id_1
AASNNMETYR
>seq_id_2
MKKLPSDDEFGHKKL

where each sequence has a unique id and maximum length of 8190 residues. Once the scoring is complete, the output csv file will have following columns:

  • sequence_id : The sequence id corresponding to each sequence in the input fasta file
  • log_likelihood : The log likelihood of the sequence.
  • perplexity : The perplexity of the sequence.

Generation CLI

Prepare the prompt file (described in the next section). Then run the following command (modifying the arguments as necessary):

mkdir -p generation_outputs/
torchrun --nproc-per-node=gpu -m progen3.tools.generate \
--prompt-file examples/generations/prompts.csv \
--output-dir generation_outputs/ \
--model-name Profluent-Bio/progen3-3b \
--n-per-prompt 5000 \
--fsdp \
--temperature 0.85 \
--top-p 0.95

You will get two files per prompt in the generation_outputs directory after completion. First is {id}.gen.fasta which contains the raw tokens generated by the model for each generated sequence (not containing the prompt). The second is {id}.seq.fasta where the generations are combined with prompt properly to output a valid N-to-C terminal protein sequence. Note, the seq.fasta files only contain a subset of gen.fasta file generations since some of the generations in gen.fasta file may be invalid.

See an example in directory examples/generations of a prompt file for generations for Deaminase given forward and reverse prefixes and for infilling a section of PETase (again both in forward and reverse direction).

Prompt File Format

The prompt file should be a .csv file. Required columns: id, sequence, min_new_tokens, max_new_tokens. (min/max new tokens are added to the sequence length). id is the unique identifier for each prompt sequence.

sequence format: <1/2><prompt>

  • <1/2><prompt> : The prompt to generate from. Assuming you want to generate for original sequence AASNNMETYR, the prompt should be:
    • For CLM in N-to-C (forward) direction: residues from N direction (e.g. 1AAS)
    • For CLM in C-to-N (reverse) direction: residues from C direction (e.g. 2RYTE)
    • For unconditional generation, you can leave the prompt part empty and only provide the direction indicator. e.g. 1 for N-to-C (forward) direction and 2 for C-to-N (reverse) direction.

Info on GLM prompt preparation

Assume you have a original protein sequence of length N. Let's say you want to fill in the span from original residue from <start pos> to <end pos> (start inclusive, end exclusive, 0-indexed) but shorten/lengthen it to (on average) M residues. For the N-to-C (forward) direction, the prompt should be:

1<original protein sequence in N-to-C direction>[GLM]<start pos>-<end pos>-<M>

For the C-to-N (reverse) direction, the prompt should be:

2<original protein sequence in C-to-N direction>[GLM]<N - end pos>-<N - start pos>-<M>

where <N - end pos> and <N - start pos> is the start and end of the span from the opposite direction.

Note: you would want to set the max_new_tokens to the maximum length of the infill you want (this would be a hard constraint compared to <M> in the prompt format which is a soft length constraint).

Example: Consider you want to fill in the NNMET part of AASNNMETYR with, on average, 4 residues instead of 5. So the start pos is 3 and the end pos is 8.

If you want to fill in the NNMET part in N-to-C direction, the prompt should be:

1AASNNMETYR[GLM]3-8-4

If you want to fill in the NNMET part in C-to-N direction, the prompt should be:

2RYTEMNNSAA[GLM]2-7-4

LICENSE

The code in this repository in released under Apache 2.0 License.

The model weights are released under CC BY-NC-SA 4.0 License.

About

Public Release of ProGen3 family of Models

Resources

Stars

114 stars

Watchers

6 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all \u003cpre\u003e\u003ccode\u003e blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks"); } } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); } })(); (function(){ try { var __m = "github.com"; var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

ProGen3

This repository contains code to access the ProGen3 family of models as well as cli interface to run these models for scoring and generation. The model weights are released under a non-commercial usage license. Please see the LICENSE section below.

Models Available

More info on ProGen3 Models can be found here. In the subsequent sections, you can provide a model name from following table where needed.

Model NameModel ParametersHuggingFace Link
Profluent-Bio/progen3-112m112MProfluent-Bio/progen3-112m
Profluent-Bio/progen3-219m219MProfluent-Bio/progen3-219m
Profluent-Bio/progen3-339m339MProfluent-Bio/progen3-339m
Profluent-Bio/progen3-762m762MProfluent-Bio/progen3-762m
Profluent-Bio/progen3-1b1BProfluent-Bio/progen3-1b
Profluent-Bio/progen3-3b3BProfluent-Bio/progen3-3b

Installation

Local Usage

ProGen3 family of models require atleast one GPU device (we have tested it only on a A100/H100 with atleast 40GB of VRAM; GPUs from prior generations may not work due to lack of support for bf16 precision and flash attention kernels) be available.

  1. Clone this repo and run bash setup.sh

Docker Usage

We also provide a docker image ghcr.io/profluent-ai/progen3:v0.1.0 that comes with certain libraries pre-installed to reduce installation time. This image still needs to be run on a machine with GPU device.

Within the container, follow the steps in the section Local Usage.

Interactive Usage

importtorchfromprogen3.modelingimportProGen3ForCausalLMfromprogen3.batch_preparerimportProGen3BatchPreparerfromprogen3.scorerimportProGen3Scorermodel=ProGen3ForCausalLM.from_pretrained("Profluent-Bio/progen3-3b", torch_dtype=torch.bfloat16)
model=model.eval().to("cuda:0")
batch_preparer=ProGen3BatchPreparer()
# Direct Usagesequence="MALWMRLLPLLALLALWGPDPAAAFVNQHLCGSHLVEALYLVCGERGFFYTPKTRREAEDLQVGQVELGGGPGAGSLQPLALEGSLQKRGIVEQCCTSICSLYQLENYCN"inputs=batch_preparer.get_batch_kwargs([sequence], device="cuda:0", reverse=False)
outputs=model(**inputs, return_dict=True)
print(outputs.logits)
# Usage with scorer (returns averaged log likelihood of forward and reverse direction)# Would suggest using Scoring CLI below if scoring very large number of sequencesscorer=ProGen3Scorer(model=model)
scores=scorer.score_batch(sequences=[sequence])
print(scores["log_likelihood"][0])

Scoring CLI

Prepare the fasta file (described in the next section) containing sequences to be scored. Then, run the following command (modifying the arguments as necessary):

torchrun --nproc-per-node=gpu -m progen3.tools.score \
--fasta-path sequences.fasta \
--output-path scores.csv \
--model-name Profluent-Bio/progen3-3b \
--fsdp

See an example in directory examples/protein_gym_zero_shot on how we use it to evaluate spearman score on protein gym assays.

bash examples/protein_gym_zero_shot/download_and_prepare.sh
# Run for individual assay
bash examples/protein_gym_zero_shot/run.sh A0A140D2T1_ZIKV_Sourisseau_2019 Profluent-Bio/progen3-3b
# Run for all assays and print aggregated spearman
bash examples/protein_gym_zero_shot/run.sh all Profluent-Bio/progen3-3b

Input File Formats

  • fasta-path : A path to fasta file

Each fasta file should be of form:

>seq_id_1
AASNNMETYR
>seq_id_2
MKKLPSDDEFGHKKL

where each sequence has a unique id and maximum length of 8190 residues. Once the scoring is complete, the output csv file will have following columns:

  • sequence_id : The sequence id corresponding to each sequence in the input fasta file
  • log_likelihood : The log likelihood of the sequence.
  • perplexity : The perplexity of the sequence.

Generation CLI

Prepare the prompt file (described in the next section). Then run the following command (modifying the arguments as necessary):

mkdir -p generation_outputs/
torchrun --nproc-per-node=gpu -m progen3.tools.generate \
--prompt-file examples/generations/prompts.csv \
--output-dir generation_outputs/ \
--model-name Profluent-Bio/progen3-3b \
--n-per-prompt 5000 \
--fsdp \
--temperature 0.85 \
--top-p 0.95

You will get two files per prompt in the generation_outputs directory after completion. First is {id}.gen.fasta which contains the raw tokens generated by the model for each generated sequence (not containing the prompt). The second is {id}.seq.fasta where the generations are combined with prompt properly to output a valid N-to-C terminal protein sequence. Note, the seq.fasta files only contain a subset of gen.fasta file generations since some of the generations in gen.fasta file may be invalid.

See an example in directory examples/generations of a prompt file for generations for Deaminase given forward and reverse prefixes and for infilling a section of PETase (again both in forward and reverse direction).

Prompt File Format

The prompt file should be a .csv file. Required columns: id, sequence, min_new_tokens, max_new_tokens. (min/max new tokens are added to the sequence length). id is the unique identifier for each prompt sequence.

sequence format: <1/2><prompt>

  • <1/2><prompt> : The prompt to generate from. Assuming you want to generate for original sequence AASNNMETYR, the prompt should be:
    • For CLM in N-to-C (forward) direction: residues from N direction (e.g. 1AAS)
    • For CLM in C-to-N (reverse) direction: residues from C direction (e.g. 2RYTE)
    • For unconditional generation, you can leave the prompt part empty and only provide the direction indicator. e.g. 1 for N-to-C (forward) direction and 2 for C-to-N (reverse) direction.

Info on GLM prompt preparation

Assume you have a original protein sequence of length N. Let's say you want to fill in the span from original residue from <start pos> to <end pos> (start inclusive, end exclusive, 0-indexed) but shorten/lengthen it to (on average) M residues. For the N-to-C (forward) direction, the prompt should be:

1<original protein sequence in N-to-C direction>[GLM]<start pos>-<end pos>-<M>

For the C-to-N (reverse) direction, the prompt should be:

2<original protein sequence in C-to-N direction>[GLM]<N - end pos>-<N - start pos>-<M>

where <N - end pos> and <N - start pos> is the start and end of the span from the opposite direction.

Note: you would want to set the max_new_tokens to the maximum length of the infill you want (this would be a hard constraint compared to <M> in the prompt format which is a soft length constraint).

Example: Consider you want to fill in the NNMET part of AASNNMETYR with, on average, 4 residues instead of 5. So the start pos is 3 and the end pos is 8.

If you want to fill in the NNMET part in N-to-C direction, the prompt should be:

1AASNNMETYR[GLM]3-8-4

If you want to fill in the NNMET part in C-to-N direction, the prompt should be:

2RYTEMNNSAA[GLM]2-7-4

LICENSE

The code in this repository in released under Apache 2.0 License.

The model weights are released under CC BY-NC-SA 4.0 License.

About

Public Release of ProGen3 family of Models

Resources

Stars

114 stars

Watchers

6 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

ProGen3

This repository contains code to access the ProGen3 family of models as well as cli interface to run these models for scoring and generation. The model weights are released under a non-commercial usage license. Please see the LICENSE section below.

Models Available

More info on ProGen3 Models can be found here. In the subsequent sections, you can provide a model name from following table where needed.

Model NameModel ParametersHuggingFace Link
Profluent-Bio/progen3-112m112MProfluent-Bio/progen3-112m
Profluent-Bio/progen3-219m219MProfluent-Bio/progen3-219m
Profluent-Bio/progen3-339m339MProfluent-Bio/progen3-339m
Profluent-Bio/progen3-762m762MProfluent-Bio/progen3-762m
Profluent-Bio/progen3-1b1BProfluent-Bio/progen3-1b
Profluent-Bio/progen3-3b3BProfluent-Bio/progen3-3b

Installation

Local Usage

ProGen3 family of models require atleast one GPU device (we have tested it only on a A100/H100 with atleast 40GB of VRAM; GPUs from prior generations may not work due to lack of support for bf16 precision and flash attention kernels) be available.

  1. Clone this repo and run bash setup.sh

Docker Usage

We also provide a docker image ghcr.io/profluent-ai/progen3:v0.1.0 that comes with certain libraries pre-installed to reduce installation time. This image still needs to be run on a machine with GPU device.

Within the container, follow the steps in the section Local Usage.

Interactive Usage

importtorchfromprogen3.modelingimportProGen3ForCausalLMfromprogen3.batch_preparerimportProGen3BatchPreparerfromprogen3.scorerimportProGen3Scorermodel=ProGen3ForCausalLM.from_pretrained("Profluent-Bio/progen3-3b", torch_dtype=torch.bfloat16)
model=model.eval().to("cuda:0")
batch_preparer=ProGen3BatchPreparer()
# Direct Usagesequence="MALWMRLLPLLALLALWGPDPAAAFVNQHLCGSHLVEALYLVCGERGFFYTPKTRREAEDLQVGQVELGGGPGAGSLQPLALEGSLQKRGIVEQCCTSICSLYQLENYCN"inputs=batch_preparer.get_batch_kwargs([sequence], device="cuda:0", reverse=False)
outputs=model(**inputs, return_dict=True)
print(outputs.logits)
# Usage with scorer (returns averaged log likelihood of forward and reverse direction)# Would suggest using Scoring CLI below if scoring very large number of sequencesscorer=ProGen3Scorer(model=model)
scores=scorer.score_batch(sequences=[sequence])
print(scores["log_likelihood"][0])

Scoring CLI

Prepare the fasta file (described in the next section) containing sequences to be scored. Then, run the following command (modifying the arguments as necessary):

torchrun --nproc-per-node=gpu -m progen3.tools.score \
--fasta-path sequences.fasta \
--output-path scores.csv \
--model-name Profluent-Bio/progen3-3b \
--fsdp

See an example in directory examples/protein_gym_zero_shot on how we use it to evaluate spearman score on protein gym assays.

bash examples/protein_gym_zero_shot/download_and_prepare.sh
# Run for individual assay
bash examples/protein_gym_zero_shot/run.sh A0A140D2T1_ZIKV_Sourisseau_2019 Profluent-Bio/progen3-3b
# Run for all assays and print aggregated spearman
bash examples/protein_gym_zero_shot/run.sh all Profluent-Bio/progen3-3b

Input File Formats

  • fasta-path : A path to fasta file

Each fasta file should be of form:

>seq_id_1
AASNNMETYR
>seq_id_2
MKKLPSDDEFGHKKL

where each sequence has a unique id and maximum length of 8190 residues. Once the scoring is complete, the output csv file will have following columns:

  • sequence_id : The sequence id corresponding to each sequence in the input fasta file
  • log_likelihood : The log likelihood of the sequence.
  • perplexity : The perplexity of the sequence.

Generation CLI

Prepare the prompt file (described in the next section). Then run the following command (modifying the arguments as necessary):

mkdir -p generation_outputs/
torchrun --nproc-per-node=gpu -m progen3.tools.generate \
--prompt-file examples/generations/prompts.csv \
--output-dir generation_outputs/ \
--model-name Profluent-Bio/progen3-3b \
--n-per-prompt 5000 \
--fsdp \
--temperature 0.85 \
--top-p 0.95

You will get two files per prompt in the generation_outputs directory after completion. First is {id}.gen.fasta which contains the raw tokens generated by the model for each generated sequence (not containing the prompt). The second is {id}.seq.fasta where the generations are combined with prompt properly to output a valid N-to-C terminal protein sequence. Note, the seq.fasta files only contain a subset of gen.fasta file generations since some of the generations in gen.fasta file may be invalid.

See an example in directory examples/generations of a prompt file for generations for Deaminase given forward and reverse prefixes and for infilling a section of PETase (again both in forward and reverse direction).

Prompt File Format

The prompt file should be a .csv file. Required columns: id, sequence, min_new_tokens, max_new_tokens. (min/max new tokens are added to the sequence length). id is the unique identifier for each prompt sequence.

sequence format: <1/2><prompt>

  • <1/2><prompt> : The prompt to generate from. Assuming you want to generate for original sequence AASNNMETYR, the prompt should be:
    • For CLM in N-to-C (forward) direction: residues from N direction (e.g. 1AAS)
    • For CLM in C-to-N (reverse) direction: residues from C direction (e.g. 2RYTE)
    • For unconditional generation, you can leave the prompt part empty and only provide the direction indicator. e.g. 1 for N-to-C (forward) direction and 2 for C-to-N (reverse) direction.

Info on GLM prompt preparation

Assume you have a original protein sequence of length N. Let's say you want to fill in the span from original residue from <start pos> to <end pos> (start inclusive, end exclusive, 0-indexed) but shorten/lengthen it to (on average) M residues. For the N-to-C (forward) direction, the prompt should be:

1<original protein sequence in N-to-C direction>[GLM]<start pos>-<end pos>-<M>

For the C-to-N (reverse) direction, the prompt should be:

2<original protein sequence in C-to-N direction>[GLM]<N - end pos>-<N - start pos>-<M>

where <N - end pos> and <N - start pos> is the start and end of the span from the opposite direction.

Note: you would want to set the max_new_tokens to the maximum length of the infill you want (this would be a hard constraint compared to <M> in the prompt format which is a soft length constraint).

Example: Consider you want to fill in the NNMET part of AASNNMETYR with, on average, 4 residues instead of 5. So the start pos is 3 and the end pos is 8.

If you want to fill in the NNMET part in N-to-C direction, the prompt should be:

1AASNNMETYR[GLM]3-8-4

If you want to fill in the NNMET part in C-to-N direction, the prompt should be:

2RYTEMNNSAA[GLM]2-7-4

LICENSE

The code in this repository in released under Apache 2.0 License.

The model weights are released under CC BY-NC-SA 4.0 License.

About

Public Release of ProGen3 family of Models

Resources

Stars

114 stars

Watchers

6 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length \u003e 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

ProGen3

This repository contains code to access the ProGen3 family of models as well as cli interface to run these models for scoring and generation. The model weights are released under a non-commercial usage license. Please see the LICENSE section below.

Models Available

More info on ProGen3 Models can be found here. In the subsequent sections, you can provide a model name from following table where needed.

Model NameModel ParametersHuggingFace Link
Profluent-Bio/progen3-112m112MProfluent-Bio/progen3-112m
Profluent-Bio/progen3-219m219MProfluent-Bio/progen3-219m
Profluent-Bio/progen3-339m339MProfluent-Bio/progen3-339m
Profluent-Bio/progen3-762m762MProfluent-Bio/progen3-762m
Profluent-Bio/progen3-1b1BProfluent-Bio/progen3-1b
Profluent-Bio/progen3-3b3BProfluent-Bio/progen3-3b

Installation

Local Usage

ProGen3 family of models require atleast one GPU device (we have tested it only on a A100/H100 with atleast 40GB of VRAM; GPUs from prior generations may not work due to lack of support for bf16 precision and flash attention kernels) be available.

  1. Clone this repo and run bash setup.sh

Docker Usage

We also provide a docker image ghcr.io/profluent-ai/progen3:v0.1.0 that comes with certain libraries pre-installed to reduce installation time. This image still needs to be run on a machine with GPU device.

Within the container, follow the steps in the section Local Usage.

Interactive Usage

importtorchfromprogen3.modelingimportProGen3ForCausalLMfromprogen3.batch_preparerimportProGen3BatchPreparerfromprogen3.scorerimportProGen3Scorermodel=ProGen3ForCausalLM.from_pretrained("Profluent-Bio/progen3-3b", torch_dtype=torch.bfloat16)
model=model.eval().to("cuda:0")
batch_preparer=ProGen3BatchPreparer()
# Direct Usagesequence="MALWMRLLPLLALLALWGPDPAAAFVNQHLCGSHLVEALYLVCGERGFFYTPKTRREAEDLQVGQVELGGGPGAGSLQPLALEGSLQKRGIVEQCCTSICSLYQLENYCN"inputs=batch_preparer.get_batch_kwargs([sequence], device="cuda:0", reverse=False)
outputs=model(**inputs, return_dict=True)
print(outputs.logits)
# Usage with scorer (returns averaged log likelihood of forward and reverse direction)# Would suggest using Scoring CLI below if scoring very large number of sequencesscorer=ProGen3Scorer(model=model)
scores=scorer.score_batch(sequences=[sequence])
print(scores["log_likelihood"][0])

Scoring CLI

Prepare the fasta file (described in the next section) containing sequences to be scored. Then, run the following command (modifying the arguments as necessary):

torchrun --nproc-per-node=gpu -m progen3.tools.score \
--fasta-path sequences.fasta \
--output-path scores.csv \
--model-name Profluent-Bio/progen3-3b \
--fsdp

See an example in directory examples/protein_gym_zero_shot on how we use it to evaluate spearman score on protein gym assays.

bash examples/protein_gym_zero_shot/download_and_prepare.sh
# Run for individual assay
bash examples/protein_gym_zero_shot/run.sh A0A140D2T1_ZIKV_Sourisseau_2019 Profluent-Bio/progen3-3b
# Run for all assays and print aggregated spearman
bash examples/protein_gym_zero_shot/run.sh all Profluent-Bio/progen3-3b

Input File Formats

  • fasta-path : A path to fasta file

Each fasta file should be of form:

>seq_id_1
AASNNMETYR
>seq_id_2
MKKLPSDDEFGHKKL

where each sequence has a unique id and maximum length of 8190 residues. Once the scoring is complete, the output csv file will have following columns:

  • sequence_id : The sequence id corresponding to each sequence in the input fasta file
  • log_likelihood : The log likelihood of the sequence.
  • perplexity : The perplexity of the sequence.

Generation CLI

Prepare the prompt file (described in the next section). Then run the following command (modifying the arguments as necessary):

mkdir -p generation_outputs/
torchrun --nproc-per-node=gpu -m progen3.tools.generate \
--prompt-file examples/generations/prompts.csv \
--output-dir generation_outputs/ \
--model-name Profluent-Bio/progen3-3b \
--n-per-prompt 5000 \
--fsdp \
--temperature 0.85 \
--top-p 0.95

You will get two files per prompt in the generation_outputs directory after completion. First is {id}.gen.fasta which contains the raw tokens generated by the model for each generated sequence (not containing the prompt). The second is {id}.seq.fasta where the generations are combined with prompt properly to output a valid N-to-C terminal protein sequence. Note, the seq.fasta files only contain a subset of gen.fasta file generations since some of the generations in gen.fasta file may be invalid.

See an example in directory examples/generations of a prompt file for generations for Deaminase given forward and reverse prefixes and for infilling a section of PETase (again both in forward and reverse direction).

Prompt File Format

The prompt file should be a .csv file. Required columns: id, sequence, min_new_tokens, max_new_tokens. (min/max new tokens are added to the sequence length). id is the unique identifier for each prompt sequence.

sequence format: <1/2><prompt>

  • <1/2><prompt> : The prompt to generate from. Assuming you want to generate for original sequence AASNNMETYR, the prompt should be:
    • For CLM in N-to-C (forward) direction: residues from N direction (e.g. 1AAS)
    • For CLM in C-to-N (reverse) direction: residues from C direction (e.g. 2RYTE)
    • For unconditional generation, you can leave the prompt part empty and only provide the direction indicator. e.g. 1 for N-to-C (forward) direction and 2 for C-to-N (reverse) direction.

Info on GLM prompt preparation

Assume you have a original protein sequence of length N. Let's say you want to fill in the span from original residue from <start pos> to <end pos> (start inclusive, end exclusive, 0-indexed) but shorten/lengthen it to (on average) M residues. For the N-to-C (forward) direction, the prompt should be:

1<original protein sequence in N-to-C direction>[GLM]<start pos>-<end pos>-<M>

For the C-to-N (reverse) direction, the prompt should be:

2<original protein sequence in C-to-N direction>[GLM]<N - end pos>-<N - start pos>-<M>

where <N - end pos> and <N - start pos> is the start and end of the span from the opposite direction.

Note: you would want to set the max_new_tokens to the maximum length of the infill you want (this would be a hard constraint compared to <M> in the prompt format which is a soft length constraint).

Example: Consider you want to fill in the NNMET part of AASNNMETYR with, on average, 4 residues instead of 5. So the start pos is 3 and the end pos is 8.

If you want to fill in the NNMET part in N-to-C direction, the prompt should be:

1AASNNMETYR[GLM]3-8-4

If you want to fill in the NNMET part in C-to-N direction, the prompt should be:

2RYTEMNNSAA[GLM]2-7-4

LICENSE

The code in this repository in released under Apache 2.0 License.

The model weights are released under CC BY-NC-SA 4.0 License.

About

Public Release of ProGen3 family of Models

Resources

Stars

114 stars

Watchers

6 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

ProGen3

This repository contains code to access the ProGen3 family of models as well as cli interface to run these models for scoring and generation. The model weights are released under a non-commercial usage license. Please see the LICENSE section below.

Models Available

More info on ProGen3 Models can be found here. In the subsequent sections, you can provide a model name from following table where needed.

Model NameModel ParametersHuggingFace Link
Profluent-Bio/progen3-112m112MProfluent-Bio/progen3-112m
Profluent-Bio/progen3-219m219MProfluent-Bio/progen3-219m
Profluent-Bio/progen3-339m339MProfluent-Bio/progen3-339m
Profluent-Bio/progen3-762m762MProfluent-Bio/progen3-762m
Profluent-Bio/progen3-1b1BProfluent-Bio/progen3-1b
Profluent-Bio/progen3-3b3BProfluent-Bio/progen3-3b

Installation

Local Usage

ProGen3 family of models require atleast one GPU device (we have tested it only on a A100/H100 with atleast 40GB of VRAM; GPUs from prior generations may not work due to lack of support for bf16 precision and flash attention kernels) be available.

  1. Clone this repo and run bash setup.sh

Docker Usage

We also provide a docker image ghcr.io/profluent-ai/progen3:v0.1.0 that comes with certain libraries pre-installed to reduce installation time. This image still needs to be run on a machine with GPU device.

Within the container, follow the steps in the section Local Usage.

Interactive Usage

importtorchfromprogen3.modelingimportProGen3ForCausalLMfromprogen3.batch_preparerimportProGen3BatchPreparerfromprogen3.scorerimportProGen3Scorermodel=ProGen3ForCausalLM.from_pretrained("Profluent-Bio/progen3-3b", torch_dtype=torch.bfloat16)
model=model.eval().to("cuda:0")
batch_preparer=ProGen3BatchPreparer()
# Direct Usagesequence="MALWMRLLPLLALLALWGPDPAAAFVNQHLCGSHLVEALYLVCGERGFFYTPKTRREAEDLQVGQVELGGGPGAGSLQPLALEGSLQKRGIVEQCCTSICSLYQLENYCN"inputs=batch_preparer.get_batch_kwargs([sequence], device="cuda:0", reverse=False)
outputs=model(**inputs, return_dict=True)
print(outputs.logits)
# Usage with scorer (returns averaged log likelihood of forward and reverse direction)# Would suggest using Scoring CLI below if scoring very large number of sequencesscorer=ProGen3Scorer(model=model)
scores=scorer.score_batch(sequences=[sequence])
print(scores["log_likelihood"][0])

Scoring CLI

Prepare the fasta file (described in the next section) containing sequences to be scored. Then, run the following command (modifying the arguments as necessary):

torchrun --nproc-per-node=gpu -m progen3.tools.score \
--fasta-path sequences.fasta \
--output-path scores.csv \
--model-name Profluent-Bio/progen3-3b \
--fsdp

See an example in directory examples/protein_gym_zero_shot on how we use it to evaluate spearman score on protein gym assays.

bash examples/protein_gym_zero_shot/download_and_prepare.sh
# Run for individual assay
bash examples/protein_gym_zero_shot/run.sh A0A140D2T1_ZIKV_Sourisseau_2019 Profluent-Bio/progen3-3b
# Run for all assays and print aggregated spearman
bash examples/protein_gym_zero_shot/run.sh all Profluent-Bio/progen3-3b

Input File Formats

  • fasta-path : A path to fasta file

Each fasta file should be of form:

>seq_id_1
AASNNMETYR
>seq_id_2
MKKLPSDDEFGHKKL

where each sequence has a unique id and maximum length of 8190 residues. Once the scoring is complete, the output csv file will have following columns:

  • sequence_id : The sequence id corresponding to each sequence in the input fasta file
  • log_likelihood : The log likelihood of the sequence.
  • perplexity : The perplexity of the sequence.

Generation CLI

Prepare the prompt file (described in the next section). Then run the following command (modifying the arguments as necessary):

mkdir -p generation_outputs/
torchrun --nproc-per-node=gpu -m progen3.tools.generate \
--prompt-file examples/generations/prompts.csv \
--output-dir generation_outputs/ \
--model-name Profluent-Bio/progen3-3b \
--n-per-prompt 5000 \
--fsdp \
--temperature 0.85 \
--top-p 0.95

You will get two files per prompt in the generation_outputs directory after completion. First is {id}.gen.fasta which contains the raw tokens generated by the model for each generated sequence (not containing the prompt). The second is {id}.seq.fasta where the generations are combined with prompt properly to output a valid N-to-C terminal protein sequence. Note, the seq.fasta files only contain a subset of gen.fasta file generations since some of the generations in gen.fasta file may be invalid.

See an example in directory examples/generations of a prompt file for generations for Deaminase given forward and reverse prefixes and for infilling a section of PETase (again both in forward and reverse direction).

Prompt File Format

The prompt file should be a .csv file. Required columns: id, sequence, min_new_tokens, max_new_tokens. (min/max new tokens are added to the sequence length). id is the unique identifier for each prompt sequence.

sequence format: <1/2><prompt>

  • <1/2><prompt> : The prompt to generate from. Assuming you want to generate for original sequence AASNNMETYR, the prompt should be:
    • For CLM in N-to-C (forward) direction: residues from N direction (e.g. 1AAS)
    • For CLM in C-to-N (reverse) direction: residues from C direction (e.g. 2RYTE)
    • For unconditional generation, you can leave the prompt part empty and only provide the direction indicator. e.g. 1 for N-to-C (forward) direction and 2 for C-to-N (reverse) direction.

Info on GLM prompt preparation

Assume you have a original protein sequence of length N. Let's say you want to fill in the span from original residue from <start pos> to <end pos> (start inclusive, end exclusive, 0-indexed) but shorten/lengthen it to (on average) M residues. For the N-to-C (forward) direction, the prompt should be:

1<original protein sequence in N-to-C direction>[GLM]<start pos>-<end pos>-<M>

For the C-to-N (reverse) direction, the prompt should be:

2<original protein sequence in C-to-N direction>[GLM]<N - end pos>-<N - start pos>-<M>

where <N - end pos> and <N - start pos> is the start and end of the span from the opposite direction.

Note: you would want to set the max_new_tokens to the maximum length of the infill you want (this would be a hard constraint compared to <M> in the prompt format which is a soft length constraint).

Example: Consider you want to fill in the NNMET part of AASNNMETYR with, on average, 4 residues instead of 5. So the start pos is 3 and the end pos is 8.

If you want to fill in the NNMET part in N-to-C direction, the prompt should be:

1AASNNMETYR[GLM]3-8-4

If you want to fill in the NNMET part in C-to-N direction, the prompt should be:

2RYTEMNNSAA[GLM]2-7-4

LICENSE

The code in this repository in released under Apache 2.0 License.

The model weights are released under CC BY-NC-SA 4.0 License.

About

Public Release of ProGen3 family of Models

Resources

Stars

114 stars

Watchers

6 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

ProGen3

This repository contains code to access the ProGen3 family of models as well as cli interface to run these models for scoring and generation. The model weights are released under a non-commercial usage license. Please see the LICENSE section below.

Models Available

More info on ProGen3 Models can be found here. In the subsequent sections, you can provide a model name from following table where needed.

Model NameModel ParametersHuggingFace Link
Profluent-Bio/progen3-112m112MProfluent-Bio/progen3-112m
Profluent-Bio/progen3-219m219MProfluent-Bio/progen3-219m
Profluent-Bio/progen3-339m339MProfluent-Bio/progen3-339m
Profluent-Bio/progen3-762m762MProfluent-Bio/progen3-762m
Profluent-Bio/progen3-1b1BProfluent-Bio/progen3-1b
Profluent-Bio/progen3-3b3BProfluent-Bio/progen3-3b

Installation

Local Usage

ProGen3 family of models require atleast one GPU device (we have tested it only on a A100/H100 with atleast 40GB of VRAM; GPUs from prior generations may not work due to lack of support for bf16 precision and flash attention kernels) be available.

  1. Clone this repo and run bash setup.sh

Docker Usage

We also provide a docker image ghcr.io/profluent-ai/progen3:v0.1.0 that comes with certain libraries pre-installed to reduce installation time. This image still needs to be run on a machine with GPU device.

Within the container, follow the steps in the section Local Usage.

Interactive Usage

importtorchfromprogen3.modelingimportProGen3ForCausalLMfromprogen3.batch_preparerimportProGen3BatchPreparerfromprogen3.scorerimportProGen3Scorermodel=ProGen3ForCausalLM.from_pretrained("Profluent-Bio/progen3-3b", torch_dtype=torch.bfloat16)
model=model.eval().to("cuda:0")
batch_preparer=ProGen3BatchPreparer()
# Direct Usagesequence="MALWMRLLPLLALLALWGPDPAAAFVNQHLCGSHLVEALYLVCGERGFFYTPKTRREAEDLQVGQVELGGGPGAGSLQPLALEGSLQKRGIVEQCCTSICSLYQLENYCN"inputs=batch_preparer.get_batch_kwargs([sequence], device="cuda:0", reverse=False)
outputs=model(**inputs, return_dict=True)
print(outputs.logits)
# Usage with scorer (returns averaged log likelihood of forward and reverse direction)# Would suggest using Scoring CLI below if scoring very large number of sequencesscorer=ProGen3Scorer(model=model)
scores=scorer.score_batch(sequences=[sequence])
print(scores["log_likelihood"][0])

Scoring CLI

Prepare the fasta file (described in the next section) containing sequences to be scored. Then, run the following command (modifying the arguments as necessary):

torchrun --nproc-per-node=gpu -m progen3.tools.score \
--fasta-path sequences.fasta \
--output-path scores.csv \
--model-name Profluent-Bio/progen3-3b \
--fsdp

See an example in directory examples/protein_gym_zero_shot on how we use it to evaluate spearman score on protein gym assays.

bash examples/protein_gym_zero_shot/download_and_prepare.sh
# Run for individual assay
bash examples/protein_gym_zero_shot/run.sh A0A140D2T1_ZIKV_Sourisseau_2019 Profluent-Bio/progen3-3b
# Run for all assays and print aggregated spearman
bash examples/protein_gym_zero_shot/run.sh all Profluent-Bio/progen3-3b

Input File Formats

  • fasta-path : A path to fasta file

Each fasta file should be of form:

>seq_id_1
AASNNMETYR
>seq_id_2
MKKLPSDDEFGHKKL

where each sequence has a unique id and maximum length of 8190 residues. Once the scoring is complete, the output csv file will have following columns:

  • sequence_id : The sequence id corresponding to each sequence in the input fasta file
  • log_likelihood : The log likelihood of the sequence.
  • perplexity : The perplexity of the sequence.

Generation CLI

Prepare the prompt file (described in the next section). Then run the following command (modifying the arguments as necessary):

mkdir -p generation_outputs/
torchrun --nproc-per-node=gpu -m progen3.tools.generate \
--prompt-file examples/generations/prompts.csv \
--output-dir generation_outputs/ \
--model-name Profluent-Bio/progen3-3b \
--n-per-prompt 5000 \
--fsdp \
--temperature 0.85 \
--top-p 0.95

You will get two files per prompt in the generation_outputs directory after completion. First is {id}.gen.fasta which contains the raw tokens generated by the model for each generated sequence (not containing the prompt). The second is {id}.seq.fasta where the generations are combined with prompt properly to output a valid N-to-C terminal protein sequence. Note, the seq.fasta files only contain a subset of gen.fasta file generations since some of the generations in gen.fasta file may be invalid.

See an example in directory examples/generations of a prompt file for generations for Deaminase given forward and reverse prefixes and for infilling a section of PETase (again both in forward and reverse direction).

Prompt File Format

The prompt file should be a .csv file. Required columns: id, sequence, min_new_tokens, max_new_tokens. (min/max new tokens are added to the sequence length). id is the unique identifier for each prompt sequence.

sequence format: <1/2><prompt>

  • <1/2><prompt> : The prompt to generate from. Assuming you want to generate for original sequence AASNNMETYR, the prompt should be:
    • For CLM in N-to-C (forward) direction: residues from N direction (e.g. 1AAS)
    • For CLM in C-to-N (reverse) direction: residues from C direction (e.g. 2RYTE)
    • For unconditional generation, you can leave the prompt part empty and only provide the direction indicator. e.g. 1 for N-to-C (forward) direction and 2 for C-to-N (reverse) direction.

Info on GLM prompt preparation

Assume you have a original protein sequence of length N. Let's say you want to fill in the span from original residue from <start pos> to <end pos> (start inclusive, end exclusive, 0-indexed) but shorten/lengthen it to (on average) M residues. For the N-to-C (forward) direction, the prompt should be:

1<original protein sequence in N-to-C direction>[GLM]<start pos>-<end pos>-<M>

For the C-to-N (reverse) direction, the prompt should be:

2<original protein sequence in C-to-N direction>[GLM]<N - end pos>-<N - start pos>-<M>

where <N - end pos> and <N - start pos> is the start and end of the span from the opposite direction.

Note: you would want to set the max_new_tokens to the maximum length of the infill you want (this would be a hard constraint compared to <M> in the prompt format which is a soft length constraint).

Example: Consider you want to fill in the NNMET part of AASNNMETYR with, on average, 4 residues instead of 5. So the start pos is 3 and the end pos is 8.

If you want to fill in the NNMET part in N-to-C direction, the prompt should be:

1AASNNMETYR[GLM]3-8-4

If you want to fill in the NNMET part in C-to-N direction, the prompt should be:

2RYTEMNNSAA[GLM]2-7-4

LICENSE

The code in this repository in released under Apache 2.0 License.

The model weights are released under CC BY-NC-SA 4.0 License.

About

Public Release of ProGen3 family of Models

Resources

Stars

114 stars

Watchers

6 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

ProGen3

This repository contains code to access the ProGen3 family of models as well as cli interface to run these models for scoring and generation. The model weights are released under a non-commercial usage license. Please see the LICENSE section below.

Models Available

More info on ProGen3 Models can be found here. In the subsequent sections, you can provide a model name from following table where needed.

Model NameModel ParametersHuggingFace Link
Profluent-Bio/progen3-112m112MProfluent-Bio/progen3-112m
Profluent-Bio/progen3-219m219MProfluent-Bio/progen3-219m
Profluent-Bio/progen3-339m339MProfluent-Bio/progen3-339m
Profluent-Bio/progen3-762m762MProfluent-Bio/progen3-762m
Profluent-Bio/progen3-1b1BProfluent-Bio/progen3-1b
Profluent-Bio/progen3-3b3BProfluent-Bio/progen3-3b

Installation

Local Usage

ProGen3 family of models require atleast one GPU device (we have tested it only on a A100/H100 with atleast 40GB of VRAM; GPUs from prior generations may not work due to lack of support for bf16 precision and flash attention kernels) be available.

  1. Clone this repo and run bash setup.sh

Docker Usage

We also provide a docker image ghcr.io/profluent-ai/progen3:v0.1.0 that comes with certain libraries pre-installed to reduce installation time. This image still needs to be run on a machine with GPU device.

Within the container, follow the steps in the section Local Usage.

Interactive Usage

importtorchfromprogen3.modelingimportProGen3ForCausalLMfromprogen3.batch_preparerimportProGen3BatchPreparerfromprogen3.scorerimportProGen3Scorermodel=ProGen3ForCausalLM.from_pretrained("Profluent-Bio/progen3-3b", torch_dtype=torch.bfloat16)
model=model.eval().to("cuda:0")
batch_preparer=ProGen3BatchPreparer()
# Direct Usagesequence="MALWMRLLPLLALLALWGPDPAAAFVNQHLCGSHLVEALYLVCGERGFFYTPKTRREAEDLQVGQVELGGGPGAGSLQPLALEGSLQKRGIVEQCCTSICSLYQLENYCN"inputs=batch_preparer.get_batch_kwargs([sequence], device="cuda:0", reverse=False)
outputs=model(**inputs, return_dict=True)
print(outputs.logits)
# Usage with scorer (returns averaged log likelihood of forward and reverse direction)# Would suggest using Scoring CLI below if scoring very large number of sequencesscorer=ProGen3Scorer(model=model)
scores=scorer.score_batch(sequences=[sequence])
print(scores["log_likelihood"][0])

Scoring CLI

Prepare the fasta file (described in the next section) containing sequences to be scored. Then, run the following command (modifying the arguments as necessary):

torchrun --nproc-per-node=gpu -m progen3.tools.score \
--fasta-path sequences.fasta \
--output-path scores.csv \
--model-name Profluent-Bio/progen3-3b \
--fsdp

See an example in directory examples/protein_gym_zero_shot on how we use it to evaluate spearman score on protein gym assays.

bash examples/protein_gym_zero_shot/download_and_prepare.sh
# Run for individual assay
bash examples/protein_gym_zero_shot/run.sh A0A140D2T1_ZIKV_Sourisseau_2019 Profluent-Bio/progen3-3b
# Run for all assays and print aggregated spearman
bash examples/protein_gym_zero_shot/run.sh all Profluent-Bio/progen3-3b

Input File Formats

  • fasta-path : A path to fasta file

Each fasta file should be of form:

>seq_id_1
AASNNMETYR
>seq_id_2
MKKLPSDDEFGHKKL

where each sequence has a unique id and maximum length of 8190 residues. Once the scoring is complete, the output csv file will have following columns:

  • sequence_id : The sequence id corresponding to each sequence in the input fasta file
  • log_likelihood : The log likelihood of the sequence.
  • perplexity : The perplexity of the sequence.

Generation CLI

Prepare the prompt file (described in the next section). Then run the following command (modifying the arguments as necessary):

mkdir -p generation_outputs/
torchrun --nproc-per-node=gpu -m progen3.tools.generate \
--prompt-file examples/generations/prompts.csv \
--output-dir generation_outputs/ \
--model-name Profluent-Bio/progen3-3b \
--n-per-prompt 5000 \
--fsdp \
--temperature 0.85 \
--top-p 0.95

You will get two files per prompt in the generation_outputs directory after completion. First is {id}.gen.fasta which contains the raw tokens generated by the model for each generated sequence (not containing the prompt). The second is {id}.seq.fasta where the generations are combined with prompt properly to output a valid N-to-C terminal protein sequence. Note, the seq.fasta files only contain a subset of gen.fasta file generations since some of the generations in gen.fasta file may be invalid.

See an example in directory examples/generations of a prompt file for generations for Deaminase given forward and reverse prefixes and for infilling a section of PETase (again both in forward and reverse direction).

Prompt File Format

The prompt file should be a .csv file. Required columns: id, sequence, min_new_tokens, max_new_tokens. (min/max new tokens are added to the sequence length). id is the unique identifier for each prompt sequence.

sequence format: <1/2><prompt>

  • <1/2><prompt> : The prompt to generate from. Assuming you want to generate for original sequence AASNNMETYR, the prompt should be:
    • For CLM in N-to-C (forward) direction: residues from N direction (e.g. 1AAS)
    • For CLM in C-to-N (reverse) direction: residues from C direction (e.g. 2RYTE)
    • For unconditional generation, you can leave the prompt part empty and only provide the direction indicator. e.g. 1 for N-to-C (forward) direction and 2 for C-to-N (reverse) direction.

Info on GLM prompt preparation

Assume you have a original protein sequence of length N. Let's say you want to fill in the span from original residue from <start pos> to <end pos> (start inclusive, end exclusive, 0-indexed) but shorten/lengthen it to (on average) M residues. For the N-to-C (forward) direction, the prompt should be:

1<original protein sequence in N-to-C direction>[GLM]<start pos>-<end pos>-<M>

For the C-to-N (reverse) direction, the prompt should be:

2<original protein sequence in C-to-N direction>[GLM]<N - end pos>-<N - start pos>-<M>

where <N - end pos> and <N - start pos> is the start and end of the span from the opposite direction.

Note: you would want to set the max_new_tokens to the maximum length of the infill you want (this would be a hard constraint compared to <M> in the prompt format which is a soft length constraint).

Example: Consider you want to fill in the NNMET part of AASNNMETYR with, on average, 4 residues instead of 5. So the start pos is 3 and the end pos is 8.

If you want to fill in the NNMET part in N-to-C direction, the prompt should be:

1AASNNMETYR[GLM]3-8-4

If you want to fill in the NNMET part in C-to-N direction, the prompt should be:

2RYTEMNNSAA[GLM]2-7-4

LICENSE

The code in this repository in released under Apache 2.0 License.

The model weights are released under CC BY-NC-SA 4.0 License.

About

Public Release of ProGen3 family of Models

Resources

Stars

114 stars

Watchers

6 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

ProGen3

This repository contains code to access the ProGen3 family of models as well as cli interface to run these models for scoring and generation. The model weights are released under a non-commercial usage license. Please see the LICENSE section below.

Models Available

More info on ProGen3 Models can be found here. In the subsequent sections, you can provide a model name from following table where needed.

Model NameModel ParametersHuggingFace Link
Profluent-Bio/progen3-112m112MProfluent-Bio/progen3-112m
Profluent-Bio/progen3-219m219MProfluent-Bio/progen3-219m
Profluent-Bio/progen3-339m339MProfluent-Bio/progen3-339m
Profluent-Bio/progen3-762m762MProfluent-Bio/progen3-762m
Profluent-Bio/progen3-1b1BProfluent-Bio/progen3-1b
Profluent-Bio/progen3-3b3BProfluent-Bio/progen3-3b

Installation

Local Usage

ProGen3 family of models require atleast one GPU device (we have tested it only on a A100/H100 with atleast 40GB of VRAM; GPUs from prior generations may not work due to lack of support for bf16 precision and flash attention kernels) be available.

  1. Clone this repo and run bash setup.sh

Docker Usage

We also provide a docker image ghcr.io/profluent-ai/progen3:v0.1.0 that comes with certain libraries pre-installed to reduce installation time. This image still needs to be run on a machine with GPU device.

Within the container, follow the steps in the section Local Usage.

Interactive Usage

importtorchfromprogen3.modelingimportProGen3ForCausalLMfromprogen3.batch_preparerimportProGen3BatchPreparerfromprogen3.scorerimportProGen3Scorermodel=ProGen3ForCausalLM.from_pretrained("Profluent-Bio/progen3-3b", torch_dtype=torch.bfloat16)
model=model.eval().to("cuda:0")
batch_preparer=ProGen3BatchPreparer()
# Direct Usagesequence="MALWMRLLPLLALLALWGPDPAAAFVNQHLCGSHLVEALYLVCGERGFFYTPKTRREAEDLQVGQVELGGGPGAGSLQPLALEGSLQKRGIVEQCCTSICSLYQLENYCN"inputs=batch_preparer.get_batch_kwargs([sequence], device="cuda:0", reverse=False)
outputs=model(**inputs, return_dict=True)
print(outputs.logits)
# Usage with scorer (returns averaged log likelihood of forward and reverse direction)# Would suggest using Scoring CLI below if scoring very large number of sequencesscorer=ProGen3Scorer(model=model)
scores=scorer.score_batch(sequences=[sequence])
print(scores["log_likelihood"][0])

Scoring CLI

Prepare the fasta file (described in the next section) containing sequences to be scored. Then, run the following command (modifying the arguments as necessary):

torchrun --nproc-per-node=gpu -m progen3.tools.score \
--fasta-path sequences.fasta \
--output-path scores.csv \
--model-name Profluent-Bio/progen3-3b \
--fsdp

See an example in directory examples/protein_gym_zero_shot on how we use it to evaluate spearman score on protein gym assays.

bash examples/protein_gym_zero_shot/download_and_prepare.sh
# Run for individual assay
bash examples/protein_gym_zero_shot/run.sh A0A140D2T1_ZIKV_Sourisseau_2019 Profluent-Bio/progen3-3b
# Run for all assays and print aggregated spearman
bash examples/protein_gym_zero_shot/run.sh all Profluent-Bio/progen3-3b

Input File Formats

  • fasta-path : A path to fasta file

Each fasta file should be of form:

>seq_id_1
AASNNMETYR
>seq_id_2
MKKLPSDDEFGHKKL

where each sequence has a unique id and maximum length of 8190 residues. Once the scoring is complete, the output csv file will have following columns:

  • sequence_id : The sequence id corresponding to each sequence in the input fasta file
  • log_likelihood : The log likelihood of the sequence.
  • perplexity : The perplexity of the sequence.

Generation CLI

Prepare the prompt file (described in the next section). Then run the following command (modifying the arguments as necessary):

mkdir -p generation_outputs/
torchrun --nproc-per-node=gpu -m progen3.tools.generate \
--prompt-file examples/generations/prompts.csv \
--output-dir generation_outputs/ \
--model-name Profluent-Bio/progen3-3b \
--n-per-prompt 5000 \
--fsdp \
--temperature 0.85 \
--top-p 0.95

You will get two files per prompt in the generation_outputs directory after completion. First is {id}.gen.fasta which contains the raw tokens generated by the model for each generated sequence (not containing the prompt). The second is {id}.seq.fasta where the generations are combined with prompt properly to output a valid N-to-C terminal protein sequence. Note, the seq.fasta files only contain a subset of gen.fasta file generations since some of the generations in gen.fasta file may be invalid.

See an example in directory examples/generations of a prompt file for generations for Deaminase given forward and reverse prefixes and for infilling a section of PETase (again both in forward and reverse direction).

Prompt File Format

The prompt file should be a .csv file. Required columns: id, sequence, min_new_tokens, max_new_tokens. (min/max new tokens are added to the sequence length). id is the unique identifier for each prompt sequence.

sequence format: <1/2><prompt>

  • <1/2><prompt> : The prompt to generate from. Assuming you want to generate for original sequence AASNNMETYR, the prompt should be:
    • For CLM in N-to-C (forward) direction: residues from N direction (e.g. 1AAS)
    • For CLM in C-to-N (reverse) direction: residues from C direction (e.g. 2RYTE)
    • For unconditional generation, you can leave the prompt part empty and only provide the direction indicator. e.g. 1 for N-to-C (forward) direction and 2 for C-to-N (reverse) direction.

Info on GLM prompt preparation

Assume you have a original protein sequence of length N. Let's say you want to fill in the span from original residue from <start pos> to <end pos> (start inclusive, end exclusive, 0-indexed) but shorten/lengthen it to (on average) M residues. For the N-to-C (forward) direction, the prompt should be:

1<original protein sequence in N-to-C direction>[GLM]<start pos>-<end pos>-<M>

For the C-to-N (reverse) direction, the prompt should be:

2<original protein sequence in C-to-N direction>[GLM]<N - end pos>-<N - start pos>-<M>

where <N - end pos> and <N - start pos> is the start and end of the span from the opposite direction.

Note: you would want to set the max_new_tokens to the maximum length of the infill you want (this would be a hard constraint compared to <M> in the prompt format which is a soft length constraint).

Example: Consider you want to fill in the NNMET part of AASNNMETYR with, on average, 4 residues instead of 5. So the start pos is 3 and the end pos is 8.

If you want to fill in the NNMET part in N-to-C direction, the prompt should be:

1AASNNMETYR[GLM]3-8-4

If you want to fill in the NNMET part in C-to-N direction, the prompt should be:

2RYTEMNNSAA[GLM]2-7-4

LICENSE

The code in this repository in released under Apache 2.0 License.

The model weights are released under CC BY-NC-SA 4.0 License.

About

Public Release of ProGen3 family of Models

Resources

Stars

114 stars

Watchers

6 watching

Forks

Releases

Packages

Contributors

Languages