Repository files navigation

CatRange

CatRange predicts useful ranges for two enzyme kinetic parameters:

  • kcat — how quickly an enzyme converts substrate to product
  • KM — the substrate concentration associated with half-maximal reaction speed

You provide a protein sequence and a substrate SMILES string. The pipeline first uses CLEAN to check whether the sequence is enzyme-like, then runs CatRange only for rows that pass that screen.

Run CatRange in Google Colab

Open CatRange in Google Colab

RECOMMENDED FOR MOST USERS
No installation or coding required

Model weights: CatRange downloads its model weights automatically from Hugging Face. The model weights are not stored in this Git repository.

Google Colab: Quick Start

  1. Click the Open in Colab badge above.
  2. Sign in to Google if asked.
  3. In Colab, choose Runtime → Run all.
  4. Keep Demo selected for your first run.
  5. Review the results table and download inference_results.csv.

The notebook installs its own compatible software versions. The first run takes longer because it downloads the models; later runs in the same runtime reuse those files.

Choose How You Want to Run CatRange

MethodBest forSetup level
Google ColabFirst-time users, classes, and quick testsNone
Local Jupyter notebookWorking on your own Linux/WSL computerBasic
Source-code commandAutomation, scripts, and advanced useAdvanced

What You Need

Each input row needs two columns:

ColumnWhat to enterExample
sequenceA protein amino-acid sequence using one-letter codesMKT...
Isomeric SMILESThe substrate's isomeric SMILES stringCCO

You can start with inference/examples/demo_pairs.csv.

Supported input sizes:

  • Protein sequence: 9–1022 amino acids
  • Isomeric SMILES: 2–512 characters

How to Run Inference

Method 1: Google Colab (recommended)

Use this method if you want the simplest experience.

  1. Open the CatRange inference notebook in Colab.
  2. Choose an input mode:
    • Demo uses included example data.
    • Interactive asks for one or more sequence/SMILES pairs.
    • Bulk uploads a CSV.
    • Bulk-large processes a larger CSV in batches.
  3. Keep Mechanistic Mutation-Aware selected unless you are reproducing an older benchmark.
  4. Run the cells from top to bottom.
  5. Download inference_results.csv from the final cell.

Colab automatically performs these steps:

  1. Checks the input format and length limits.
  2. Runs the CLEAN enzyme screen.
  3. Predicts kcat and KM ranges for enzyme-like rows.
  4. Combines everything into one results table.

Method 2: Local Jupyter notebook

Use this method to run the same guided interface on your own computer. The local workflow currently requires Linux or Windows Subsystem for Linux (WSL). A GPU is helpful but not required.

The notebook runs locally, but the first setup still needs internet access to download software and model files. Cached files can be reused for later runs.

  1. Install Git, Python 3, and Jupyter.

  2. Clone the repository:

    git clone https://github.com/ssbio/CatRange.git
    cd CatRange
  3. Install the lightweight notebook launcher requirements:

    python3 -m pip install jupyter pandas
  4. Open the provided notebook:

    jupyter lab CatRange_Inference_Interface.ipynb
  5. Choose Demo for a first run, then run the cells from top to bottom.

The notebook creates isolated runtimes for CLEAN and CatRange, which prevents their machine-learning dependencies from interfering with each other.

Method 3: Source-code command

Use this method for repeatable scripts or batch jobs. It currently requires Linux or WSL and Python 3.12.

  1. Clone the repository:

    git clone https://github.com/ssbio/CatRange.git
    cd CatRange
  2. Create an environment and install the source inference requirements:

    python3.12 -m venv .venv
    source .venv/bin/activate
    python -m pip install --upgrade pip
    python -m pip install -r inference/requirements.txt
  3. Run the demo CSV:

    python inference/catrange_inference.py \
    --input inference/examples/demo_pairs.csv \
    --output inference_results.csv

That single command runs input validation → CLEAN → CatRange → merged results. There is no separate cleaning command to remember. The CLEAN environment, source, and pretrained files are downloaded automatically on the first run and cached in .clean_runtime/. CatRange model weights are downloaded from Hugging Face, not from this Git repository.

For all command options:

python inference/catrange_inference.py --help

Understanding the Results

The Colab notebook uses friendly column names; the source command uses compact machine-friendly names.

Colab columnSource columnMeaning
Predicted EC numberclean_top_ec_numberCLEAN's most likely EC number
clean_top_confidenceclean_top_confidenceCLEAN's confidence score for its top EC prediction
Classified as enzyme?clean_is_enzymeWhether the row passed the CLEAN enzyme screen
Pipeline notecatrange_statusWhether CatRange predicted the row or why it was skipped
Predicted kcat range (s^-1)kcat_pred_rangePredicted kcat range
Predicted KM range (M)km_pred_rangePredicted KM range

The CLEAN confidence is a model score, not an experimental measurement. CatRange reports ranges because enzyme measurements can vary substantially with experimental conditions.

For Researchers and Developers

CatRange combines ESM-C protein embeddings, ChemBERTa substrate embeddings, and XGBoost classification. The CatLog curated enzyme-kinetics data support model training, benchmarking, and manuscript analyses.

Repository layout

CatRange_Inference_Interface.ipynb Guided Colab/local inference notebook
inference/ End-to-end source inference and model files
catrange_model/ CatRange training and evaluation code
data/ CatLog/CatRange data and metadata
results/ CatRange and comparator benchmark outputs
benchmarks/retrained_comparators/ Comparator retraining scripts
ablation/ Feature-ablation scripts and results
figures/ Figure source and output files
manuscript/ Manuscript and supporting information
envs/ Reproducible environment definitions

Train CatRange

Create the research environments:

bash scripts/env/create_conda_envs.sh all

Run a manuscript configuration:

cd catrange_model
python3 -m pip install --no-deps -e .
PYTHONPATH=. python scripts/cv_train.py --config configs/kcat_esmc.yaml --device cuda

See envs/README.md and catrange_model/README.md for training, benchmarking, and reproducibility details.

Citation

Please cite the CatRange manuscript when using this code or data.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

CatRange

CatRange predicts useful ranges for two enzyme kinetic parameters:

  • kcat — how quickly an enzyme converts substrate to product
  • KM — the substrate concentration associated with half-maximal reaction speed

You provide a protein sequence and a substrate SMILES string. The pipeline first uses CLEAN to check whether the sequence is enzyme-like, then runs CatRange only for rows that pass that screen.

Run CatRange in Google Colab

Open CatRange in Google Colab

RECOMMENDED FOR MOST USERS
No installation or coding required

Model weights: CatRange downloads its model weights automatically from Hugging Face. The model weights are not stored in this Git repository.

Google Colab: Quick Start

  1. Click the Open in Colab badge above.
  2. Sign in to Google if asked.
  3. In Colab, choose Runtime → Run all.
  4. Keep Demo selected for your first run.
  5. Review the results table and download inference_results.csv.

The notebook installs its own compatible software versions. The first run takes longer because it downloads the models; later runs in the same runtime reuse those files.

Choose How You Want to Run CatRange

MethodBest forSetup level
Google ColabFirst-time users, classes, and quick testsNone
Local Jupyter notebookWorking on your own Linux/WSL computerBasic
Source-code commandAutomation, scripts, and advanced useAdvanced

What You Need

Each input row needs two columns:

ColumnWhat to enterExample
sequenceA protein amino-acid sequence using one-letter codesMKT...
Isomeric SMILESThe substrate's isomeric SMILES stringCCO

You can start with inference/examples/demo_pairs.csv.

Supported input sizes:

  • Protein sequence: 9–1022 amino acids
  • Isomeric SMILES: 2–512 characters

How to Run Inference

Method 1: Google Colab (recommended)

Use this method if you want the simplest experience.

  1. Open the CatRange inference notebook in Colab.
  2. Choose an input mode:
    • Demo uses included example data.
    • Interactive asks for one or more sequence/SMILES pairs.
    • Bulk uploads a CSV.
    • Bulk-large processes a larger CSV in batches.
  3. Keep Mechanistic Mutation-Aware selected unless you are reproducing an older benchmark.
  4. Run the cells from top to bottom.
  5. Download inference_results.csv from the final cell.

Colab automatically performs these steps:

  1. Checks the input format and length limits.
  2. Runs the CLEAN enzyme screen.
  3. Predicts kcat and KM ranges for enzyme-like rows.
  4. Combines everything into one results table.

Method 2: Local Jupyter notebook

Use this method to run the same guided interface on your own computer. The local workflow currently requires Linux or Windows Subsystem for Linux (WSL). A GPU is helpful but not required.

The notebook runs locally, but the first setup still needs internet access to download software and model files. Cached files can be reused for later runs.

  1. Install Git, Python 3, and Jupyter.

  2. Clone the repository:

    git clone https://github.com/ssbio/CatRange.git
    cd CatRange
  3. Install the lightweight notebook launcher requirements:

    python3 -m pip install jupyter pandas
  4. Open the provided notebook:

    jupyter lab CatRange_Inference_Interface.ipynb
  5. Choose Demo for a first run, then run the cells from top to bottom.

The notebook creates isolated runtimes for CLEAN and CatRange, which prevents their machine-learning dependencies from interfering with each other.

Method 3: Source-code command

Use this method for repeatable scripts or batch jobs. It currently requires Linux or WSL and Python 3.12.

  1. Clone the repository:

    git clone https://github.com/ssbio/CatRange.git
    cd CatRange
  2. Create an environment and install the source inference requirements:

    python3.12 -m venv .venv
    source .venv/bin/activate
    python -m pip install --upgrade pip
    python -m pip install -r inference/requirements.txt
  3. Run the demo CSV:

    python inference/catrange_inference.py \
    --input inference/examples/demo_pairs.csv \
    --output inference_results.csv

That single command runs input validation → CLEAN → CatRange → merged results. There is no separate cleaning command to remember. The CLEAN environment, source, and pretrained files are downloaded automatically on the first run and cached in .clean_runtime/. CatRange model weights are downloaded from Hugging Face, not from this Git repository.

For all command options:

python inference/catrange_inference.py --help

Understanding the Results

The Colab notebook uses friendly column names; the source command uses compact machine-friendly names.

Colab columnSource columnMeaning
Predicted EC numberclean_top_ec_numberCLEAN's most likely EC number
clean_top_confidenceclean_top_confidenceCLEAN's confidence score for its top EC prediction
Classified as enzyme?clean_is_enzymeWhether the row passed the CLEAN enzyme screen
Pipeline notecatrange_statusWhether CatRange predicted the row or why it was skipped
Predicted kcat range (s^-1)kcat_pred_rangePredicted kcat range
Predicted KM range (M)km_pred_rangePredicted KM range

The CLEAN confidence is a model score, not an experimental measurement. CatRange reports ranges because enzyme measurements can vary substantially with experimental conditions.

For Researchers and Developers

CatRange combines ESM-C protein embeddings, ChemBERTa substrate embeddings, and XGBoost classification. The CatLog curated enzyme-kinetics data support model training, benchmarking, and manuscript analyses.

Repository layout

CatRange_Inference_Interface.ipynb Guided Colab/local inference notebook
inference/ End-to-end source inference and model files
catrange_model/ CatRange training and evaluation code
data/ CatLog/CatRange data and metadata
results/ CatRange and comparator benchmark outputs
benchmarks/retrained_comparators/ Comparator retraining scripts
ablation/ Feature-ablation scripts and results
figures/ Figure source and output files
manuscript/ Manuscript and supporting information
envs/ Reproducible environment definitions

Train CatRange

Create the research environments:

bash scripts/env/create_conda_envs.sh all

Run a manuscript configuration:

cd catrange_model
python3 -m pip install --no-deps -e .
PYTHONPATH=. python scripts/cv_train.py --config configs/kcat_esmc.yaml --device cuda

See envs/README.md and catrange_model/README.md for training, benchmarking, and reproducibility details.

Citation

Please cite the CatRange manuscript when using this code or data.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

CatRange

CatRange predicts useful ranges for two enzyme kinetic parameters:

  • kcat — how quickly an enzyme converts substrate to product
  • KM — the substrate concentration associated with half-maximal reaction speed

You provide a protein sequence and a substrate SMILES string. The pipeline first uses CLEAN to check whether the sequence is enzyme-like, then runs CatRange only for rows that pass that screen.

Run CatRange in Google Colab

Open CatRange in Google Colab

RECOMMENDED FOR MOST USERS
No installation or coding required

Model weights: CatRange downloads its model weights automatically from Hugging Face. The model weights are not stored in this Git repository.

Google Colab: Quick Start

  1. Click the Open in Colab badge above.
  2. Sign in to Google if asked.
  3. In Colab, choose Runtime → Run all.
  4. Keep Demo selected for your first run.
  5. Review the results table and download inference_results.csv.

The notebook installs its own compatible software versions. The first run takes longer because it downloads the models; later runs in the same runtime reuse those files.

Choose How You Want to Run CatRange

MethodBest forSetup level
Google ColabFirst-time users, classes, and quick testsNone
Local Jupyter notebookWorking on your own Linux/WSL computerBasic
Source-code commandAutomation, scripts, and advanced useAdvanced

What You Need

Each input row needs two columns:

ColumnWhat to enterExample
sequenceA protein amino-acid sequence using one-letter codesMKT...
Isomeric SMILESThe substrate's isomeric SMILES stringCCO

You can start with inference/examples/demo_pairs.csv.

Supported input sizes:

  • Protein sequence: 9–1022 amino acids
  • Isomeric SMILES: 2–512 characters

How to Run Inference

Method 1: Google Colab (recommended)

Use this method if you want the simplest experience.

  1. Open the CatRange inference notebook in Colab.
  2. Choose an input mode:
    • Demo uses included example data.
    • Interactive asks for one or more sequence/SMILES pairs.
    • Bulk uploads a CSV.
    • Bulk-large processes a larger CSV in batches.
  3. Keep Mechanistic Mutation-Aware selected unless you are reproducing an older benchmark.
  4. Run the cells from top to bottom.
  5. Download inference_results.csv from the final cell.

Colab automatically performs these steps:

  1. Checks the input format and length limits.
  2. Runs the CLEAN enzyme screen.
  3. Predicts kcat and KM ranges for enzyme-like rows.
  4. Combines everything into one results table.

Method 2: Local Jupyter notebook

Use this method to run the same guided interface on your own computer. The local workflow currently requires Linux or Windows Subsystem for Linux (WSL). A GPU is helpful but not required.

The notebook runs locally, but the first setup still needs internet access to download software and model files. Cached files can be reused for later runs.

  1. Install Git, Python 3, and Jupyter.

  2. Clone the repository:

    git clone https://github.com/ssbio/CatRange.git
    cd CatRange
  3. Install the lightweight notebook launcher requirements:

    python3 -m pip install jupyter pandas
  4. Open the provided notebook:

    jupyter lab CatRange_Inference_Interface.ipynb
  5. Choose Demo for a first run, then run the cells from top to bottom.

The notebook creates isolated runtimes for CLEAN and CatRange, which prevents their machine-learning dependencies from interfering with each other.

Method 3: Source-code command

Use this method for repeatable scripts or batch jobs. It currently requires Linux or WSL and Python 3.12.

  1. Clone the repository:

    git clone https://github.com/ssbio/CatRange.git
    cd CatRange
  2. Create an environment and install the source inference requirements:

    python3.12 -m venv .venv
    source .venv/bin/activate
    python -m pip install --upgrade pip
    python -m pip install -r inference/requirements.txt
  3. Run the demo CSV:

    python inference/catrange_inference.py \
    --input inference/examples/demo_pairs.csv \
    --output inference_results.csv

That single command runs input validation → CLEAN → CatRange → merged results. There is no separate cleaning command to remember. The CLEAN environment, source, and pretrained files are downloaded automatically on the first run and cached in .clean_runtime/. CatRange model weights are downloaded from Hugging Face, not from this Git repository.

For all command options:

python inference/catrange_inference.py --help

Understanding the Results

The Colab notebook uses friendly column names; the source command uses compact machine-friendly names.

Colab columnSource columnMeaning
Predicted EC numberclean_top_ec_numberCLEAN's most likely EC number
clean_top_confidenceclean_top_confidenceCLEAN's confidence score for its top EC prediction
Classified as enzyme?clean_is_enzymeWhether the row passed the CLEAN enzyme screen
Pipeline notecatrange_statusWhether CatRange predicted the row or why it was skipped
Predicted kcat range (s^-1)kcat_pred_rangePredicted kcat range
Predicted KM range (M)km_pred_rangePredicted KM range

The CLEAN confidence is a model score, not an experimental measurement. CatRange reports ranges because enzyme measurements can vary substantially with experimental conditions.

For Researchers and Developers

CatRange combines ESM-C protein embeddings, ChemBERTa substrate embeddings, and XGBoost classification. The CatLog curated enzyme-kinetics data support model training, benchmarking, and manuscript analyses.

Repository layout

CatRange_Inference_Interface.ipynb Guided Colab/local inference notebook
inference/ End-to-end source inference and model files
catrange_model/ CatRange training and evaluation code
data/ CatLog/CatRange data and metadata
results/ CatRange and comparator benchmark outputs
benchmarks/retrained_comparators/ Comparator retraining scripts
ablation/ Feature-ablation scripts and results
figures/ Figure source and output files
manuscript/ Manuscript and supporting information
envs/ Reproducible environment definitions

Train CatRange

Create the research environments:

bash scripts/env/create_conda_envs.sh all

Run a manuscript configuration:

cd catrange_model
python3 -m pip install --no-deps -e .
PYTHONPATH=. python scripts/cv_train.py --config configs/kcat_esmc.yaml --device cuda

See envs/README.md and catrange_model/README.md for training, benchmarking, and reproducibility details.

Citation

Please cite the CatRange manuscript when using this code or data.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

CatRange

CatRange predicts useful ranges for two enzyme kinetic parameters:

  • kcat — how quickly an enzyme converts substrate to product
  • KM — the substrate concentration associated with half-maximal reaction speed

You provide a protein sequence and a substrate SMILES string. The pipeline first uses CLEAN to check whether the sequence is enzyme-like, then runs CatRange only for rows that pass that screen.

Run CatRange in Google Colab

Open CatRange in Google Colab

RECOMMENDED FOR MOST USERS
No installation or coding required

Model weights: CatRange downloads its model weights automatically from Hugging Face. The model weights are not stored in this Git repository.

Google Colab: Quick Start

  1. Click the Open in Colab badge above.
  2. Sign in to Google if asked.
  3. In Colab, choose Runtime → Run all.
  4. Keep Demo selected for your first run.
  5. Review the results table and download inference_results.csv.

The notebook installs its own compatible software versions. The first run takes longer because it downloads the models; later runs in the same runtime reuse those files.

Choose How You Want to Run CatRange

MethodBest forSetup level
Google ColabFirst-time users, classes, and quick testsNone
Local Jupyter notebookWorking on your own Linux/WSL computerBasic
Source-code commandAutomation, scripts, and advanced useAdvanced

What You Need

Each input row needs two columns:

ColumnWhat to enterExample
sequenceA protein amino-acid sequence using one-letter codesMKT...
Isomeric SMILESThe substrate's isomeric SMILES stringCCO

You can start with inference/examples/demo_pairs.csv.

Supported input sizes:

  • Protein sequence: 9–1022 amino acids
  • Isomeric SMILES: 2–512 characters

How to Run Inference

Method 1: Google Colab (recommended)

Use this method if you want the simplest experience.

  1. Open the CatRange inference notebook in Colab.
  2. Choose an input mode:
    • Demo uses included example data.
    • Interactive asks for one or more sequence/SMILES pairs.
    • Bulk uploads a CSV.
    • Bulk-large processes a larger CSV in batches.
  3. Keep Mechanistic Mutation-Aware selected unless you are reproducing an older benchmark.
  4. Run the cells from top to bottom.
  5. Download inference_results.csv from the final cell.

Colab automatically performs these steps:

  1. Checks the input format and length limits.
  2. Runs the CLEAN enzyme screen.
  3. Predicts kcat and KM ranges for enzyme-like rows.
  4. Combines everything into one results table.

Method 2: Local Jupyter notebook

Use this method to run the same guided interface on your own computer. The local workflow currently requires Linux or Windows Subsystem for Linux (WSL). A GPU is helpful but not required.

The notebook runs locally, but the first setup still needs internet access to download software and model files. Cached files can be reused for later runs.

  1. Install Git, Python 3, and Jupyter.

  2. Clone the repository:

    git clone https://github.com/ssbio/CatRange.git
    cd CatRange
  3. Install the lightweight notebook launcher requirements:

    python3 -m pip install jupyter pandas
  4. Open the provided notebook:

    jupyter lab CatRange_Inference_Interface.ipynb
  5. Choose Demo for a first run, then run the cells from top to bottom.

The notebook creates isolated runtimes for CLEAN and CatRange, which prevents their machine-learning dependencies from interfering with each other.

Method 3: Source-code command

Use this method for repeatable scripts or batch jobs. It currently requires Linux or WSL and Python 3.12.

  1. Clone the repository:

    git clone https://github.com/ssbio/CatRange.git
    cd CatRange
  2. Create an environment and install the source inference requirements:

    python3.12 -m venv .venv
    source .venv/bin/activate
    python -m pip install --upgrade pip
    python -m pip install -r inference/requirements.txt
  3. Run the demo CSV:

    python inference/catrange_inference.py \
    --input inference/examples/demo_pairs.csv \
    --output inference_results.csv

That single command runs input validation → CLEAN → CatRange → merged results. There is no separate cleaning command to remember. The CLEAN environment, source, and pretrained files are downloaded automatically on the first run and cached in .clean_runtime/. CatRange model weights are downloaded from Hugging Face, not from this Git repository.

For all command options:

python inference/catrange_inference.py --help

Understanding the Results

The Colab notebook uses friendly column names; the source command uses compact machine-friendly names.

Colab columnSource columnMeaning
Predicted EC numberclean_top_ec_numberCLEAN's most likely EC number
clean_top_confidenceclean_top_confidenceCLEAN's confidence score for its top EC prediction
Classified as enzyme?clean_is_enzymeWhether the row passed the CLEAN enzyme screen
Pipeline notecatrange_statusWhether CatRange predicted the row or why it was skipped
Predicted kcat range (s^-1)kcat_pred_rangePredicted kcat range
Predicted KM range (M)km_pred_rangePredicted KM range

The CLEAN confidence is a model score, not an experimental measurement. CatRange reports ranges because enzyme measurements can vary substantially with experimental conditions.

For Researchers and Developers

CatRange combines ESM-C protein embeddings, ChemBERTa substrate embeddings, and XGBoost classification. The CatLog curated enzyme-kinetics data support model training, benchmarking, and manuscript analyses.

Repository layout

CatRange_Inference_Interface.ipynb Guided Colab/local inference notebook
inference/ End-to-end source inference and model files
catrange_model/ CatRange training and evaluation code
data/ CatLog/CatRange data and metadata
results/ CatRange and comparator benchmark outputs
benchmarks/retrained_comparators/ Comparator retraining scripts
ablation/ Feature-ablation scripts and results
figures/ Figure source and output files
manuscript/ Manuscript and supporting information
envs/ Reproducible environment definitions

Train CatRange

Create the research environments:

bash scripts/env/create_conda_envs.sh all

Run a manuscript configuration:

cd catrange_model
python3 -m pip install --no-deps -e .
PYTHONPATH=. python scripts/cv_train.py --config configs/kcat_esmc.yaml --device cuda

See envs/README.md and catrange_model/README.md for training, benchmarking, and reproducibility details.

Citation

Please cite the CatRange manuscript when using this code or data.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

CatRange

CatRange predicts useful ranges for two enzyme kinetic parameters:

  • kcat — how quickly an enzyme converts substrate to product
  • KM — the substrate concentration associated with half-maximal reaction speed

You provide a protein sequence and a substrate SMILES string. The pipeline first uses CLEAN to check whether the sequence is enzyme-like, then runs CatRange only for rows that pass that screen.

Run CatRange in Google Colab

Open CatRange in Google Colab

RECOMMENDED FOR MOST USERS
No installation or coding required

Model weights: CatRange downloads its model weights automatically from Hugging Face. The model weights are not stored in this Git repository.

Google Colab: Quick Start

  1. Click the Open in Colab badge above.
  2. Sign in to Google if asked.
  3. In Colab, choose Runtime → Run all.
  4. Keep Demo selected for your first run.
  5. Review the results table and download inference_results.csv.

The notebook installs its own compatible software versions. The first run takes longer because it downloads the models; later runs in the same runtime reuse those files.

Choose How You Want to Run CatRange

MethodBest forSetup level
Google ColabFirst-time users, classes, and quick testsNone
Local Jupyter notebookWorking on your own Linux/WSL computerBasic
Source-code commandAutomation, scripts, and advanced useAdvanced

What You Need

Each input row needs two columns:

ColumnWhat to enterExample
sequenceA protein amino-acid sequence using one-letter codesMKT...
Isomeric SMILESThe substrate's isomeric SMILES stringCCO

You can start with inference/examples/demo_pairs.csv.

Supported input sizes:

  • Protein sequence: 9–1022 amino acids
  • Isomeric SMILES: 2–512 characters

How to Run Inference

Method 1: Google Colab (recommended)

Use this method if you want the simplest experience.

  1. Open the CatRange inference notebook in Colab.
  2. Choose an input mode:
    • Demo uses included example data.
    • Interactive asks for one or more sequence/SMILES pairs.
    • Bulk uploads a CSV.
    • Bulk-large processes a larger CSV in batches.
  3. Keep Mechanistic Mutation-Aware selected unless you are reproducing an older benchmark.
  4. Run the cells from top to bottom.
  5. Download inference_results.csv from the final cell.

Colab automatically performs these steps:

  1. Checks the input format and length limits.
  2. Runs the CLEAN enzyme screen.
  3. Predicts kcat and KM ranges for enzyme-like rows.
  4. Combines everything into one results table.

Method 2: Local Jupyter notebook

Use this method to run the same guided interface on your own computer. The local workflow currently requires Linux or Windows Subsystem for Linux (WSL). A GPU is helpful but not required.

The notebook runs locally, but the first setup still needs internet access to download software and model files. Cached files can be reused for later runs.

  1. Install Git, Python 3, and Jupyter.

  2. Clone the repository:

    git clone https://github.com/ssbio/CatRange.git
    cd CatRange
  3. Install the lightweight notebook launcher requirements:

    python3 -m pip install jupyter pandas
  4. Open the provided notebook:

    jupyter lab CatRange_Inference_Interface.ipynb
  5. Choose Demo for a first run, then run the cells from top to bottom.

The notebook creates isolated runtimes for CLEAN and CatRange, which prevents their machine-learning dependencies from interfering with each other.

Method 3: Source-code command

Use this method for repeatable scripts or batch jobs. It currently requires Linux or WSL and Python 3.12.

  1. Clone the repository:

    git clone https://github.com/ssbio/CatRange.git
    cd CatRange
  2. Create an environment and install the source inference requirements:

    python3.12 -m venv .venv
    source .venv/bin/activate
    python -m pip install --upgrade pip
    python -m pip install -r inference/requirements.txt
  3. Run the demo CSV:

    python inference/catrange_inference.py \
    --input inference/examples/demo_pairs.csv \
    --output inference_results.csv

That single command runs input validation → CLEAN → CatRange → merged results. There is no separate cleaning command to remember. The CLEAN environment, source, and pretrained files are downloaded automatically on the first run and cached in .clean_runtime/. CatRange model weights are downloaded from Hugging Face, not from this Git repository.

For all command options:

python inference/catrange_inference.py --help

Understanding the Results

The Colab notebook uses friendly column names; the source command uses compact machine-friendly names.

Colab columnSource columnMeaning
Predicted EC numberclean_top_ec_numberCLEAN's most likely EC number
clean_top_confidenceclean_top_confidenceCLEAN's confidence score for its top EC prediction
Classified as enzyme?clean_is_enzymeWhether the row passed the CLEAN enzyme screen
Pipeline notecatrange_statusWhether CatRange predicted the row or why it was skipped
Predicted kcat range (s^-1)kcat_pred_rangePredicted kcat range
Predicted KM range (M)km_pred_rangePredicted KM range

The CLEAN confidence is a model score, not an experimental measurement. CatRange reports ranges because enzyme measurements can vary substantially with experimental conditions.

For Researchers and Developers

CatRange combines ESM-C protein embeddings, ChemBERTa substrate embeddings, and XGBoost classification. The CatLog curated enzyme-kinetics data support model training, benchmarking, and manuscript analyses.

Repository layout

CatRange_Inference_Interface.ipynb Guided Colab/local inference notebook
inference/ End-to-end source inference and model files
catrange_model/ CatRange training and evaluation code
data/ CatLog/CatRange data and metadata
results/ CatRange and comparator benchmark outputs
benchmarks/retrained_comparators/ Comparator retraining scripts
ablation/ Feature-ablation scripts and results
figures/ Figure source and output files
manuscript/ Manuscript and supporting information
envs/ Reproducible environment definitions

Train CatRange

Create the research environments:

bash scripts/env/create_conda_envs.sh all

Run a manuscript configuration:

cd catrange_model
python3 -m pip install --no-deps -e .
PYTHONPATH=. python scripts/cv_train.py --config configs/kcat_esmc.yaml --device cuda

See envs/README.md and catrange_model/README.md for training, benchmarking, and reproducibility details.

Citation

Please cite the CatRange manuscript when using this code or data.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

CatRange

CatRange predicts useful ranges for two enzyme kinetic parameters:

  • kcat — how quickly an enzyme converts substrate to product
  • KM — the substrate concentration associated with half-maximal reaction speed

You provide a protein sequence and a substrate SMILES string. The pipeline first uses CLEAN to check whether the sequence is enzyme-like, then runs CatRange only for rows that pass that screen.

Run CatRange in Google Colab

Open CatRange in Google Colab

RECOMMENDED FOR MOST USERS
No installation or coding required

Model weights: CatRange downloads its model weights automatically from Hugging Face. The model weights are not stored in this Git repository.

Google Colab: Quick Start

  1. Click the Open in Colab badge above.
  2. Sign in to Google if asked.
  3. In Colab, choose Runtime → Run all.
  4. Keep Demo selected for your first run.
  5. Review the results table and download inference_results.csv.

The notebook installs its own compatible software versions. The first run takes longer because it downloads the models; later runs in the same runtime reuse those files.

Choose How You Want to Run CatRange

MethodBest forSetup level
Google ColabFirst-time users, classes, and quick testsNone
Local Jupyter notebookWorking on your own Linux/WSL computerBasic
Source-code commandAutomation, scripts, and advanced useAdvanced

What You Need

Each input row needs two columns:

ColumnWhat to enterExample
sequenceA protein amino-acid sequence using one-letter codesMKT...
Isomeric SMILESThe substrate's isomeric SMILES stringCCO

You can start with inference/examples/demo_pairs.csv.

Supported input sizes:

  • Protein sequence: 9–1022 amino acids
  • Isomeric SMILES: 2–512 characters

How to Run Inference

Method 1: Google Colab (recommended)

Use this method if you want the simplest experience.

  1. Open the CatRange inference notebook in Colab.
  2. Choose an input mode:
    • Demo uses included example data.
    • Interactive asks for one or more sequence/SMILES pairs.
    • Bulk uploads a CSV.
    • Bulk-large processes a larger CSV in batches.
  3. Keep Mechanistic Mutation-Aware selected unless you are reproducing an older benchmark.
  4. Run the cells from top to bottom.
  5. Download inference_results.csv from the final cell.

Colab automatically performs these steps:

  1. Checks the input format and length limits.
  2. Runs the CLEAN enzyme screen.
  3. Predicts kcat and KM ranges for enzyme-like rows.
  4. Combines everything into one results table.

Method 2: Local Jupyter notebook

Use this method to run the same guided interface on your own computer. The local workflow currently requires Linux or Windows Subsystem for Linux (WSL). A GPU is helpful but not required.

The notebook runs locally, but the first setup still needs internet access to download software and model files. Cached files can be reused for later runs.

  1. Install Git, Python 3, and Jupyter.

  2. Clone the repository:

    git clone https://github.com/ssbio/CatRange.git
    cd CatRange
  3. Install the lightweight notebook launcher requirements:

    python3 -m pip install jupyter pandas
  4. Open the provided notebook:

    jupyter lab CatRange_Inference_Interface.ipynb
  5. Choose Demo for a first run, then run the cells from top to bottom.

The notebook creates isolated runtimes for CLEAN and CatRange, which prevents their machine-learning dependencies from interfering with each other.

Method 3: Source-code command

Use this method for repeatable scripts or batch jobs. It currently requires Linux or WSL and Python 3.12.

  1. Clone the repository:

    git clone https://github.com/ssbio/CatRange.git
    cd CatRange
  2. Create an environment and install the source inference requirements:

    python3.12 -m venv .venv
    source .venv/bin/activate
    python -m pip install --upgrade pip
    python -m pip install -r inference/requirements.txt
  3. Run the demo CSV:

    python inference/catrange_inference.py \
    --input inference/examples/demo_pairs.csv \
    --output inference_results.csv

That single command runs input validation → CLEAN → CatRange → merged results. There is no separate cleaning command to remember. The CLEAN environment, source, and pretrained files are downloaded automatically on the first run and cached in .clean_runtime/. CatRange model weights are downloaded from Hugging Face, not from this Git repository.

For all command options:

python inference/catrange_inference.py --help

Understanding the Results

The Colab notebook uses friendly column names; the source command uses compact machine-friendly names.

Colab columnSource columnMeaning
Predicted EC numberclean_top_ec_numberCLEAN's most likely EC number
clean_top_confidenceclean_top_confidenceCLEAN's confidence score for its top EC prediction
Classified as enzyme?clean_is_enzymeWhether the row passed the CLEAN enzyme screen
Pipeline notecatrange_statusWhether CatRange predicted the row or why it was skipped
Predicted kcat range (s^-1)kcat_pred_rangePredicted kcat range
Predicted KM range (M)km_pred_rangePredicted KM range

The CLEAN confidence is a model score, not an experimental measurement. CatRange reports ranges because enzyme measurements can vary substantially with experimental conditions.

For Researchers and Developers

CatRange combines ESM-C protein embeddings, ChemBERTa substrate embeddings, and XGBoost classification. The CatLog curated enzyme-kinetics data support model training, benchmarking, and manuscript analyses.

Repository layout

CatRange_Inference_Interface.ipynb Guided Colab/local inference notebook
inference/ End-to-end source inference and model files
catrange_model/ CatRange training and evaluation code
data/ CatLog/CatRange data and metadata
results/ CatRange and comparator benchmark outputs
benchmarks/retrained_comparators/ Comparator retraining scripts
ablation/ Feature-ablation scripts and results
figures/ Figure source and output files
manuscript/ Manuscript and supporting information
envs/ Reproducible environment definitions

Train CatRange

Create the research environments:

bash scripts/env/create_conda_envs.sh all

Run a manuscript configuration:

cd catrange_model
python3 -m pip install --no-deps -e .
PYTHONPATH=. python scripts/cv_train.py --config configs/kcat_esmc.yaml --device cuda

See envs/README.md and catrange_model/README.md for training, benchmarking, and reproducibility details.

Citation

Please cite the CatRange manuscript when using this code or data.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

CatRange

CatRange predicts useful ranges for two enzyme kinetic parameters:

  • kcat — how quickly an enzyme converts substrate to product
  • KM — the substrate concentration associated with half-maximal reaction speed

You provide a protein sequence and a substrate SMILES string. The pipeline first uses CLEAN to check whether the sequence is enzyme-like, then runs CatRange only for rows that pass that screen.

Run CatRange in Google Colab

Open CatRange in Google Colab

RECOMMENDED FOR MOST USERS
No installation or coding required

Model weights: CatRange downloads its model weights automatically from Hugging Face. The model weights are not stored in this Git repository.

Google Colab: Quick Start

  1. Click the Open in Colab badge above.
  2. Sign in to Google if asked.
  3. In Colab, choose Runtime → Run all.
  4. Keep Demo selected for your first run.
  5. Review the results table and download inference_results.csv.

The notebook installs its own compatible software versions. The first run takes longer because it downloads the models; later runs in the same runtime reuse those files.

Choose How You Want to Run CatRange

MethodBest forSetup level
Google ColabFirst-time users, classes, and quick testsNone
Local Jupyter notebookWorking on your own Linux/WSL computerBasic
Source-code commandAutomation, scripts, and advanced useAdvanced

What You Need

Each input row needs two columns:

ColumnWhat to enterExample
sequenceA protein amino-acid sequence using one-letter codesMKT...
Isomeric SMILESThe substrate's isomeric SMILES stringCCO

You can start with inference/examples/demo_pairs.csv.

Supported input sizes:

  • Protein sequence: 9–1022 amino acids
  • Isomeric SMILES: 2–512 characters

How to Run Inference

Method 1: Google Colab (recommended)

Use this method if you want the simplest experience.

  1. Open the CatRange inference notebook in Colab.
  2. Choose an input mode:
    • Demo uses included example data.
    • Interactive asks for one or more sequence/SMILES pairs.
    • Bulk uploads a CSV.
    • Bulk-large processes a larger CSV in batches.
  3. Keep Mechanistic Mutation-Aware selected unless you are reproducing an older benchmark.
  4. Run the cells from top to bottom.
  5. Download inference_results.csv from the final cell.

Colab automatically performs these steps:

  1. Checks the input format and length limits.
  2. Runs the CLEAN enzyme screen.
  3. Predicts kcat and KM ranges for enzyme-like rows.
  4. Combines everything into one results table.

Method 2: Local Jupyter notebook

Use this method to run the same guided interface on your own computer. The local workflow currently requires Linux or Windows Subsystem for Linux (WSL). A GPU is helpful but not required.

The notebook runs locally, but the first setup still needs internet access to download software and model files. Cached files can be reused for later runs.

  1. Install Git, Python 3, and Jupyter.

  2. Clone the repository:

    git clone https://github.com/ssbio/CatRange.git
    cd CatRange
  3. Install the lightweight notebook launcher requirements:

    python3 -m pip install jupyter pandas
  4. Open the provided notebook:

    jupyter lab CatRange_Inference_Interface.ipynb
  5. Choose Demo for a first run, then run the cells from top to bottom.

The notebook creates isolated runtimes for CLEAN and CatRange, which prevents their machine-learning dependencies from interfering with each other.

Method 3: Source-code command

Use this method for repeatable scripts or batch jobs. It currently requires Linux or WSL and Python 3.12.

  1. Clone the repository:

    git clone https://github.com/ssbio/CatRange.git
    cd CatRange
  2. Create an environment and install the source inference requirements:

    python3.12 -m venv .venv
    source .venv/bin/activate
    python -m pip install --upgrade pip
    python -m pip install -r inference/requirements.txt
  3. Run the demo CSV:

    python inference/catrange_inference.py \
    --input inference/examples/demo_pairs.csv \
    --output inference_results.csv

That single command runs input validation → CLEAN → CatRange → merged results. There is no separate cleaning command to remember. The CLEAN environment, source, and pretrained files are downloaded automatically on the first run and cached in .clean_runtime/. CatRange model weights are downloaded from Hugging Face, not from this Git repository.

For all command options:

python inference/catrange_inference.py --help

Understanding the Results

The Colab notebook uses friendly column names; the source command uses compact machine-friendly names.

Colab columnSource columnMeaning
Predicted EC numberclean_top_ec_numberCLEAN's most likely EC number
clean_top_confidenceclean_top_confidenceCLEAN's confidence score for its top EC prediction
Classified as enzyme?clean_is_enzymeWhether the row passed the CLEAN enzyme screen
Pipeline notecatrange_statusWhether CatRange predicted the row or why it was skipped
Predicted kcat range (s^-1)kcat_pred_rangePredicted kcat range
Predicted KM range (M)km_pred_rangePredicted KM range

The CLEAN confidence is a model score, not an experimental measurement. CatRange reports ranges because enzyme measurements can vary substantially with experimental conditions.

For Researchers and Developers

CatRange combines ESM-C protein embeddings, ChemBERTa substrate embeddings, and XGBoost classification. The CatLog curated enzyme-kinetics data support model training, benchmarking, and manuscript analyses.

Repository layout

CatRange_Inference_Interface.ipynb Guided Colab/local inference notebook
inference/ End-to-end source inference and model files
catrange_model/ CatRange training and evaluation code
data/ CatLog/CatRange data and metadata
results/ CatRange and comparator benchmark outputs
benchmarks/retrained_comparators/ Comparator retraining scripts
ablation/ Feature-ablation scripts and results
figures/ Figure source and output files
manuscript/ Manuscript and supporting information
envs/ Reproducible environment definitions

Train CatRange

Create the research environments:

bash scripts/env/create_conda_envs.sh all

Run a manuscript configuration:

cd catrange_model
python3 -m pip install --no-deps -e .
PYTHONPATH=. python scripts/cv_train.py --config configs/kcat_esmc.yaml --device cuda

See envs/README.md and catrange_model/README.md for training, benchmarking, and reproducibility details.

Citation

Please cite the CatRange manuscript when using this code or data.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

CatRange

CatRange predicts useful ranges for two enzyme kinetic parameters:

  • kcat — how quickly an enzyme converts substrate to product
  • KM — the substrate concentration associated with half-maximal reaction speed

You provide a protein sequence and a substrate SMILES string. The pipeline first uses CLEAN to check whether the sequence is enzyme-like, then runs CatRange only for rows that pass that screen.

Run CatRange in Google Colab

Open CatRange in Google Colab

RECOMMENDED FOR MOST USERS
No installation or coding required

Model weights: CatRange downloads its model weights automatically from Hugging Face. The model weights are not stored in this Git repository.

Google Colab: Quick Start

  1. Click the Open in Colab badge above.
  2. Sign in to Google if asked.
  3. In Colab, choose Runtime → Run all.
  4. Keep Demo selected for your first run.
  5. Review the results table and download inference_results.csv.

The notebook installs its own compatible software versions. The first run takes longer because it downloads the models; later runs in the same runtime reuse those files.

Choose How You Want to Run CatRange

MethodBest forSetup level
Google ColabFirst-time users, classes, and quick testsNone
Local Jupyter notebookWorking on your own Linux/WSL computerBasic
Source-code commandAutomation, scripts, and advanced useAdvanced

What You Need

Each input row needs two columns:

ColumnWhat to enterExample
sequenceA protein amino-acid sequence using one-letter codesMKT...
Isomeric SMILESThe substrate's isomeric SMILES stringCCO

You can start with inference/examples/demo_pairs.csv.

Supported input sizes:

  • Protein sequence: 9–1022 amino acids
  • Isomeric SMILES: 2–512 characters

How to Run Inference

Method 1: Google Colab (recommended)

Use this method if you want the simplest experience.

  1. Open the CatRange inference notebook in Colab.
  2. Choose an input mode:
    • Demo uses included example data.
    • Interactive asks for one or more sequence/SMILES pairs.
    • Bulk uploads a CSV.
    • Bulk-large processes a larger CSV in batches.
  3. Keep Mechanistic Mutation-Aware selected unless you are reproducing an older benchmark.
  4. Run the cells from top to bottom.
  5. Download inference_results.csv from the final cell.

Colab automatically performs these steps:

  1. Checks the input format and length limits.
  2. Runs the CLEAN enzyme screen.
  3. Predicts kcat and KM ranges for enzyme-like rows.
  4. Combines everything into one results table.

Method 2: Local Jupyter notebook

Use this method to run the same guided interface on your own computer. The local workflow currently requires Linux or Windows Subsystem for Linux (WSL). A GPU is helpful but not required.

The notebook runs locally, but the first setup still needs internet access to download software and model files. Cached files can be reused for later runs.

  1. Install Git, Python 3, and Jupyter.

  2. Clone the repository:

    git clone https://github.com/ssbio/CatRange.git
    cd CatRange
  3. Install the lightweight notebook launcher requirements:

    python3 -m pip install jupyter pandas
  4. Open the provided notebook:

    jupyter lab CatRange_Inference_Interface.ipynb
  5. Choose Demo for a first run, then run the cells from top to bottom.

The notebook creates isolated runtimes for CLEAN and CatRange, which prevents their machine-learning dependencies from interfering with each other.

Method 3: Source-code command

Use this method for repeatable scripts or batch jobs. It currently requires Linux or WSL and Python 3.12.

  1. Clone the repository:

    git clone https://github.com/ssbio/CatRange.git
    cd CatRange
  2. Create an environment and install the source inference requirements:

    python3.12 -m venv .venv
    source .venv/bin/activate
    python -m pip install --upgrade pip
    python -m pip install -r inference/requirements.txt
  3. Run the demo CSV:

    python inference/catrange_inference.py \
    --input inference/examples/demo_pairs.csv \
    --output inference_results.csv

That single command runs input validation → CLEAN → CatRange → merged results. There is no separate cleaning command to remember. The CLEAN environment, source, and pretrained files are downloaded automatically on the first run and cached in .clean_runtime/. CatRange model weights are downloaded from Hugging Face, not from this Git repository.

For all command options:

python inference/catrange_inference.py --help

Understanding the Results

The Colab notebook uses friendly column names; the source command uses compact machine-friendly names.

Colab columnSource columnMeaning
Predicted EC numberclean_top_ec_numberCLEAN's most likely EC number
clean_top_confidenceclean_top_confidenceCLEAN's confidence score for its top EC prediction
Classified as enzyme?clean_is_enzymeWhether the row passed the CLEAN enzyme screen
Pipeline notecatrange_statusWhether CatRange predicted the row or why it was skipped
Predicted kcat range (s^-1)kcat_pred_rangePredicted kcat range
Predicted KM range (M)km_pred_rangePredicted KM range

The CLEAN confidence is a model score, not an experimental measurement. CatRange reports ranges because enzyme measurements can vary substantially with experimental conditions.

For Researchers and Developers

CatRange combines ESM-C protein embeddings, ChemBERTa substrate embeddings, and XGBoost classification. The CatLog curated enzyme-kinetics data support model training, benchmarking, and manuscript analyses.

Repository layout

CatRange_Inference_Interface.ipynb Guided Colab/local inference notebook
inference/ End-to-end source inference and model files
catrange_model/ CatRange training and evaluation code
data/ CatLog/CatRange data and metadata
results/ CatRange and comparator benchmark outputs
benchmarks/retrained_comparators/ Comparator retraining scripts
ablation/ Feature-ablation scripts and results
figures/ Figure source and output files
manuscript/ Manuscript and supporting information
envs/ Reproducible environment definitions

Train CatRange

Create the research environments:

bash scripts/env/create_conda_envs.sh all

Run a manuscript configuration:

cd catrange_model
python3 -m pip install --no-deps -e .
PYTHONPATH=. python scripts/cv_train.py --config configs/kcat_esmc.yaml --device cuda

See envs/README.md and catrange_model/README.md for training, benchmarking, and reproducibility details.

Citation

Please cite the CatRange manuscript when using this code or data.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages