Repository files navigation

#CatBench Vector Search Playground Cat Benchmarking at Scale, finally!

There are two separate Python apps in the app directory:

  • CatVector - a simple static vector heatmap visualization app (no database)
  • CatBench - a simple Python/Flask application using Postgres+pgvector similarity search queries (and joins to a regular TPCC schema)

Go to installation steps below

CatBench

You can test this app out yourself, installation steps are below.

Here are a few screenshots of the similarity search and recommendation engine app (for cats!) in action:

Cat similarity search outputCat similarity search queryCat recommendation engine outputCat recommendation engine query planCat recommendation engine query plan

Installation Steps

25000 cat/dog images are included in this repository. I have tested this on RHEL9 and Ubuntu 24.04 so far. You need to have python and pip installed in your OS for this. For installing Python packages locally with pip, you probably want to use a Python virtual environment (venv).

Interactive CatBench application that requires a Postgres database and loading data

Make sure that you have a Postgres database (with pgvector extension) running and accessible and change the psql commands below to include your username/password if you are not using a default local connection:

In the catbench repo root directory, run this to generate embedding vectors from the 25000 pet images (this uses PyTorch which automatically runs on CPUs if you don't have a GPU available).

git clone https://github.com/tanelpoder/catbench
cd catbench
pip install -r requirements-catbench.txt

NB! You may need to install Postgres and the PgVector extension and the python3-psycopg2 package using your OS package manager first, if pip doesn't successfully install psycopg2 on your Linux distro.

The next step generates vector embeddings for the 25000 pet photos included in this repository (using GPU's if cuda/NVIDIA GPUs are available, otherwise CPUs.

Process the 25000 pet photos and generate their embeddings for loading into postgres:

python app/catbench/scripts/generate_embeddings.py data/PetImages/Cat embeddings/cats.tsv
python app/catbench/scripts/generate_embeddings.py data/PetImages/Dog embeddings/dogs.tsv

This may take a while. Then load the vectors and other OLTP data into the database:

gunzip app/catbench/scripts/create_tpcc_tables.sql.gz
psql -f app/catbench/scripts/create_tpcc_tables.sql psql -f app/catbench/scripts/create_catbench_tables.sql psql -f app/catbench/scripts/create_recommendation_schema.sql 

If you're using a local Postgres instance that allows logging in as tpcc user without a password, no action needed. Otherwise open the catbench.py file to change your Postgres user/pass settings if you are not using a default local connection. And then run the app:

cd app/catbench
python3 catbench.py

You can now go to hostname:5000 and browse around:

CatBench app frontpage

Stress test

  • Check the app/catbench/scripts/ directory and run cat_loop.sh or cat_loop_wit_recall.sh scripts in there (the same for dogs). These shell scripts call similarly named .sql scripts under the hood, look inside them to see how they work. You can use similar patterns to construct your own stress test queries.
  • You currently need to change the "tpcc" to your database name (if you're not using "tpcc").
  • You can uncomment more psql lines to increase concurrency (and hit CTRL+C in terminal to cancel/kill all currently running psql loops`

Other

The data/PetImages directory is the Kaggle Cat/Dog dataset (total 25k images) originally released by Microsoft:

You don't need to separately download this file as it's already included in this repo (as permitted by Microsoft's CDLA license).

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

#CatBench Vector Search Playground Cat Benchmarking at Scale, finally!

There are two separate Python apps in the app directory:

  • CatVector - a simple static vector heatmap visualization app (no database)
  • CatBench - a simple Python/Flask application using Postgres+pgvector similarity search queries (and joins to a regular TPCC schema)

Go to installation steps below

CatBench

You can test this app out yourself, installation steps are below.

Here are a few screenshots of the similarity search and recommendation engine app (for cats!) in action:

Cat similarity search outputCat similarity search queryCat recommendation engine outputCat recommendation engine query planCat recommendation engine query plan

Installation Steps

25000 cat/dog images are included in this repository. I have tested this on RHEL9 and Ubuntu 24.04 so far. You need to have python and pip installed in your OS for this. For installing Python packages locally with pip, you probably want to use a Python virtual environment (venv).

Interactive CatBench application that requires a Postgres database and loading data

Make sure that you have a Postgres database (with pgvector extension) running and accessible and change the psql commands below to include your username/password if you are not using a default local connection:

In the catbench repo root directory, run this to generate embedding vectors from the 25000 pet images (this uses PyTorch which automatically runs on CPUs if you don't have a GPU available).

git clone https://github.com/tanelpoder/catbench
cd catbench
pip install -r requirements-catbench.txt

NB! You may need to install Postgres and the PgVector extension and the python3-psycopg2 package using your OS package manager first, if pip doesn't successfully install psycopg2 on your Linux distro.

The next step generates vector embeddings for the 25000 pet photos included in this repository (using GPU's if cuda/NVIDIA GPUs are available, otherwise CPUs.

Process the 25000 pet photos and generate their embeddings for loading into postgres:

python app/catbench/scripts/generate_embeddings.py data/PetImages/Cat embeddings/cats.tsv
python app/catbench/scripts/generate_embeddings.py data/PetImages/Dog embeddings/dogs.tsv

This may take a while. Then load the vectors and other OLTP data into the database:

gunzip app/catbench/scripts/create_tpcc_tables.sql.gz
psql -f app/catbench/scripts/create_tpcc_tables.sql psql -f app/catbench/scripts/create_catbench_tables.sql psql -f app/catbench/scripts/create_recommendation_schema.sql 

If you're using a local Postgres instance that allows logging in as tpcc user without a password, no action needed. Otherwise open the catbench.py file to change your Postgres user/pass settings if you are not using a default local connection. And then run the app:

cd app/catbench
python3 catbench.py

You can now go to hostname:5000 and browse around:

CatBench app frontpage

Stress test

  • Check the app/catbench/scripts/ directory and run cat_loop.sh or cat_loop_wit_recall.sh scripts in there (the same for dogs). These shell scripts call similarly named .sql scripts under the hood, look inside them to see how they work. You can use similar patterns to construct your own stress test queries.
  • You currently need to change the "tpcc" to your database name (if you're not using "tpcc").
  • You can uncomment more psql lines to increase concurrency (and hit CTRL+C in terminal to cancel/kill all currently running psql loops`

Other

The data/PetImages directory is the Kaggle Cat/Dog dataset (total 25k images) originally released by Microsoft:

You don't need to separately download this file as it's already included in this repo (as permitted by Microsoft's CDLA license).

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

#CatBench Vector Search Playground Cat Benchmarking at Scale, finally!

There are two separate Python apps in the app directory:

  • CatVector - a simple static vector heatmap visualization app (no database)
  • CatBench - a simple Python/Flask application using Postgres+pgvector similarity search queries (and joins to a regular TPCC schema)

Go to installation steps below

CatBench

You can test this app out yourself, installation steps are below.

Here are a few screenshots of the similarity search and recommendation engine app (for cats!) in action:

Cat similarity search outputCat similarity search queryCat recommendation engine outputCat recommendation engine query planCat recommendation engine query plan

Installation Steps

25000 cat/dog images are included in this repository. I have tested this on RHEL9 and Ubuntu 24.04 so far. You need to have python and pip installed in your OS for this. For installing Python packages locally with pip, you probably want to use a Python virtual environment (venv).

Interactive CatBench application that requires a Postgres database and loading data

Make sure that you have a Postgres database (with pgvector extension) running and accessible and change the psql commands below to include your username/password if you are not using a default local connection:

In the catbench repo root directory, run this to generate embedding vectors from the 25000 pet images (this uses PyTorch which automatically runs on CPUs if you don't have a GPU available).

git clone https://github.com/tanelpoder/catbench
cd catbench
pip install -r requirements-catbench.txt

NB! You may need to install Postgres and the PgVector extension and the python3-psycopg2 package using your OS package manager first, if pip doesn't successfully install psycopg2 on your Linux distro.

The next step generates vector embeddings for the 25000 pet photos included in this repository (using GPU's if cuda/NVIDIA GPUs are available, otherwise CPUs.

Process the 25000 pet photos and generate their embeddings for loading into postgres:

python app/catbench/scripts/generate_embeddings.py data/PetImages/Cat embeddings/cats.tsv
python app/catbench/scripts/generate_embeddings.py data/PetImages/Dog embeddings/dogs.tsv

This may take a while. Then load the vectors and other OLTP data into the database:

gunzip app/catbench/scripts/create_tpcc_tables.sql.gz
psql -f app/catbench/scripts/create_tpcc_tables.sql psql -f app/catbench/scripts/create_catbench_tables.sql psql -f app/catbench/scripts/create_recommendation_schema.sql 

If you're using a local Postgres instance that allows logging in as tpcc user without a password, no action needed. Otherwise open the catbench.py file to change your Postgres user/pass settings if you are not using a default local connection. And then run the app:

cd app/catbench
python3 catbench.py

You can now go to hostname:5000 and browse around:

CatBench app frontpage

Stress test

  • Check the app/catbench/scripts/ directory and run cat_loop.sh or cat_loop_wit_recall.sh scripts in there (the same for dogs). These shell scripts call similarly named .sql scripts under the hood, look inside them to see how they work. You can use similar patterns to construct your own stress test queries.
  • You currently need to change the "tpcc" to your database name (if you're not using "tpcc").
  • You can uncomment more psql lines to increase concurrency (and hit CTRL+C in terminal to cancel/kill all currently running psql loops`

Other

The data/PetImages directory is the Kaggle Cat/Dog dataset (total 25k images) originally released by Microsoft:

You don't need to separately download this file as it's already included in this repo (as permitted by Microsoft's CDLA license).

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

#CatBench Vector Search Playground Cat Benchmarking at Scale, finally!

There are two separate Python apps in the app directory:

  • CatVector - a simple static vector heatmap visualization app (no database)
  • CatBench - a simple Python/Flask application using Postgres+pgvector similarity search queries (and joins to a regular TPCC schema)

Go to installation steps below

CatBench

You can test this app out yourself, installation steps are below.

Here are a few screenshots of the similarity search and recommendation engine app (for cats!) in action:

Cat similarity search outputCat similarity search queryCat recommendation engine outputCat recommendation engine query planCat recommendation engine query plan

Installation Steps

25000 cat/dog images are included in this repository. I have tested this on RHEL9 and Ubuntu 24.04 so far. You need to have python and pip installed in your OS for this. For installing Python packages locally with pip, you probably want to use a Python virtual environment (venv).

Interactive CatBench application that requires a Postgres database and loading data

Make sure that you have a Postgres database (with pgvector extension) running and accessible and change the psql commands below to include your username/password if you are not using a default local connection:

In the catbench repo root directory, run this to generate embedding vectors from the 25000 pet images (this uses PyTorch which automatically runs on CPUs if you don't have a GPU available).

git clone https://github.com/tanelpoder/catbench
cd catbench
pip install -r requirements-catbench.txt

NB! You may need to install Postgres and the PgVector extension and the python3-psycopg2 package using your OS package manager first, if pip doesn't successfully install psycopg2 on your Linux distro.

The next step generates vector embeddings for the 25000 pet photos included in this repository (using GPU's if cuda/NVIDIA GPUs are available, otherwise CPUs.

Process the 25000 pet photos and generate their embeddings for loading into postgres:

python app/catbench/scripts/generate_embeddings.py data/PetImages/Cat embeddings/cats.tsv
python app/catbench/scripts/generate_embeddings.py data/PetImages/Dog embeddings/dogs.tsv

This may take a while. Then load the vectors and other OLTP data into the database:

gunzip app/catbench/scripts/create_tpcc_tables.sql.gz
psql -f app/catbench/scripts/create_tpcc_tables.sql psql -f app/catbench/scripts/create_catbench_tables.sql psql -f app/catbench/scripts/create_recommendation_schema.sql 

If you're using a local Postgres instance that allows logging in as tpcc user without a password, no action needed. Otherwise open the catbench.py file to change your Postgres user/pass settings if you are not using a default local connection. And then run the app:

cd app/catbench
python3 catbench.py

You can now go to hostname:5000 and browse around:

CatBench app frontpage

Stress test

  • Check the app/catbench/scripts/ directory and run cat_loop.sh or cat_loop_wit_recall.sh scripts in there (the same for dogs). These shell scripts call similarly named .sql scripts under the hood, look inside them to see how they work. You can use similar patterns to construct your own stress test queries.
  • You currently need to change the "tpcc" to your database name (if you're not using "tpcc").
  • You can uncomment more psql lines to increase concurrency (and hit CTRL+C in terminal to cancel/kill all currently running psql loops`

Other

The data/PetImages directory is the Kaggle Cat/Dog dataset (total 25k images) originally released by Microsoft:

You don't need to separately download this file as it's already included in this repo (as permitted by Microsoft's CDLA license).

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

#CatBench Vector Search Playground Cat Benchmarking at Scale, finally!

There are two separate Python apps in the app directory:

  • CatVector - a simple static vector heatmap visualization app (no database)
  • CatBench - a simple Python/Flask application using Postgres+pgvector similarity search queries (and joins to a regular TPCC schema)

Go to installation steps below

CatBench

You can test this app out yourself, installation steps are below.

Here are a few screenshots of the similarity search and recommendation engine app (for cats!) in action:

Cat similarity search outputCat similarity search queryCat recommendation engine outputCat recommendation engine query planCat recommendation engine query plan

Installation Steps

25000 cat/dog images are included in this repository. I have tested this on RHEL9 and Ubuntu 24.04 so far. You need to have python and pip installed in your OS for this. For installing Python packages locally with pip, you probably want to use a Python virtual environment (venv).

Interactive CatBench application that requires a Postgres database and loading data

Make sure that you have a Postgres database (with pgvector extension) running and accessible and change the psql commands below to include your username/password if you are not using a default local connection:

In the catbench repo root directory, run this to generate embedding vectors from the 25000 pet images (this uses PyTorch which automatically runs on CPUs if you don't have a GPU available).

git clone https://github.com/tanelpoder/catbench
cd catbench
pip install -r requirements-catbench.txt

NB! You may need to install Postgres and the PgVector extension and the python3-psycopg2 package using your OS package manager first, if pip doesn't successfully install psycopg2 on your Linux distro.

The next step generates vector embeddings for the 25000 pet photos included in this repository (using GPU's if cuda/NVIDIA GPUs are available, otherwise CPUs.

Process the 25000 pet photos and generate their embeddings for loading into postgres:

python app/catbench/scripts/generate_embeddings.py data/PetImages/Cat embeddings/cats.tsv
python app/catbench/scripts/generate_embeddings.py data/PetImages/Dog embeddings/dogs.tsv

This may take a while. Then load the vectors and other OLTP data into the database:

gunzip app/catbench/scripts/create_tpcc_tables.sql.gz
psql -f app/catbench/scripts/create_tpcc_tables.sql psql -f app/catbench/scripts/create_catbench_tables.sql psql -f app/catbench/scripts/create_recommendation_schema.sql 

If you're using a local Postgres instance that allows logging in as tpcc user without a password, no action needed. Otherwise open the catbench.py file to change your Postgres user/pass settings if you are not using a default local connection. And then run the app:

cd app/catbench
python3 catbench.py

You can now go to hostname:5000 and browse around:

CatBench app frontpage

Stress test

  • Check the app/catbench/scripts/ directory and run cat_loop.sh or cat_loop_wit_recall.sh scripts in there (the same for dogs). These shell scripts call similarly named .sql scripts under the hood, look inside them to see how they work. You can use similar patterns to construct your own stress test queries.
  • You currently need to change the "tpcc" to your database name (if you're not using "tpcc").
  • You can uncomment more psql lines to increase concurrency (and hit CTRL+C in terminal to cancel/kill all currently running psql loops`

Other

The data/PetImages directory is the Kaggle Cat/Dog dataset (total 25k images) originally released by Microsoft:

You don't need to separately download this file as it's already included in this repo (as permitted by Microsoft's CDLA license).

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

#CatBench Vector Search Playground Cat Benchmarking at Scale, finally!

There are two separate Python apps in the app directory:

  • CatVector - a simple static vector heatmap visualization app (no database)
  • CatBench - a simple Python/Flask application using Postgres+pgvector similarity search queries (and joins to a regular TPCC schema)

Go to installation steps below

CatBench

You can test this app out yourself, installation steps are below.

Here are a few screenshots of the similarity search and recommendation engine app (for cats!) in action:

Cat similarity search outputCat similarity search queryCat recommendation engine outputCat recommendation engine query planCat recommendation engine query plan

Installation Steps

25000 cat/dog images are included in this repository. I have tested this on RHEL9 and Ubuntu 24.04 so far. You need to have python and pip installed in your OS for this. For installing Python packages locally with pip, you probably want to use a Python virtual environment (venv).

Interactive CatBench application that requires a Postgres database and loading data

Make sure that you have a Postgres database (with pgvector extension) running and accessible and change the psql commands below to include your username/password if you are not using a default local connection:

In the catbench repo root directory, run this to generate embedding vectors from the 25000 pet images (this uses PyTorch which automatically runs on CPUs if you don't have a GPU available).

git clone https://github.com/tanelpoder/catbench
cd catbench
pip install -r requirements-catbench.txt

NB! You may need to install Postgres and the PgVector extension and the python3-psycopg2 package using your OS package manager first, if pip doesn't successfully install psycopg2 on your Linux distro.

The next step generates vector embeddings for the 25000 pet photos included in this repository (using GPU's if cuda/NVIDIA GPUs are available, otherwise CPUs.

Process the 25000 pet photos and generate their embeddings for loading into postgres:

python app/catbench/scripts/generate_embeddings.py data/PetImages/Cat embeddings/cats.tsv
python app/catbench/scripts/generate_embeddings.py data/PetImages/Dog embeddings/dogs.tsv

This may take a while. Then load the vectors and other OLTP data into the database:

gunzip app/catbench/scripts/create_tpcc_tables.sql.gz
psql -f app/catbench/scripts/create_tpcc_tables.sql psql -f app/catbench/scripts/create_catbench_tables.sql psql -f app/catbench/scripts/create_recommendation_schema.sql 

If you're using a local Postgres instance that allows logging in as tpcc user without a password, no action needed. Otherwise open the catbench.py file to change your Postgres user/pass settings if you are not using a default local connection. And then run the app:

cd app/catbench
python3 catbench.py

You can now go to hostname:5000 and browse around:

CatBench app frontpage

Stress test

  • Check the app/catbench/scripts/ directory and run cat_loop.sh or cat_loop_wit_recall.sh scripts in there (the same for dogs). These shell scripts call similarly named .sql scripts under the hood, look inside them to see how they work. You can use similar patterns to construct your own stress test queries.
  • You currently need to change the "tpcc" to your database name (if you're not using "tpcc").
  • You can uncomment more psql lines to increase concurrency (and hit CTRL+C in terminal to cancel/kill all currently running psql loops`

Other

The data/PetImages directory is the Kaggle Cat/Dog dataset (total 25k images) originally released by Microsoft:

You don't need to separately download this file as it's already included in this repo (as permitted by Microsoft's CDLA license).

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

#CatBench Vector Search Playground Cat Benchmarking at Scale, finally!

There are two separate Python apps in the app directory:

  • CatVector - a simple static vector heatmap visualization app (no database)
  • CatBench - a simple Python/Flask application using Postgres+pgvector similarity search queries (and joins to a regular TPCC schema)

Go to installation steps below

CatBench

You can test this app out yourself, installation steps are below.

Here are a few screenshots of the similarity search and recommendation engine app (for cats!) in action:

Cat similarity search outputCat similarity search queryCat recommendation engine outputCat recommendation engine query planCat recommendation engine query plan

Installation Steps

25000 cat/dog images are included in this repository. I have tested this on RHEL9 and Ubuntu 24.04 so far. You need to have python and pip installed in your OS for this. For installing Python packages locally with pip, you probably want to use a Python virtual environment (venv).

Interactive CatBench application that requires a Postgres database and loading data

Make sure that you have a Postgres database (with pgvector extension) running and accessible and change the psql commands below to include your username/password if you are not using a default local connection:

In the catbench repo root directory, run this to generate embedding vectors from the 25000 pet images (this uses PyTorch which automatically runs on CPUs if you don't have a GPU available).

git clone https://github.com/tanelpoder/catbench
cd catbench
pip install -r requirements-catbench.txt

NB! You may need to install Postgres and the PgVector extension and the python3-psycopg2 package using your OS package manager first, if pip doesn't successfully install psycopg2 on your Linux distro.

The next step generates vector embeddings for the 25000 pet photos included in this repository (using GPU's if cuda/NVIDIA GPUs are available, otherwise CPUs.

Process the 25000 pet photos and generate their embeddings for loading into postgres:

python app/catbench/scripts/generate_embeddings.py data/PetImages/Cat embeddings/cats.tsv
python app/catbench/scripts/generate_embeddings.py data/PetImages/Dog embeddings/dogs.tsv

This may take a while. Then load the vectors and other OLTP data into the database:

gunzip app/catbench/scripts/create_tpcc_tables.sql.gz
psql -f app/catbench/scripts/create_tpcc_tables.sql psql -f app/catbench/scripts/create_catbench_tables.sql psql -f app/catbench/scripts/create_recommendation_schema.sql 

If you're using a local Postgres instance that allows logging in as tpcc user without a password, no action needed. Otherwise open the catbench.py file to change your Postgres user/pass settings if you are not using a default local connection. And then run the app:

cd app/catbench
python3 catbench.py

You can now go to hostname:5000 and browse around:

CatBench app frontpage

Stress test

  • Check the app/catbench/scripts/ directory and run cat_loop.sh or cat_loop_wit_recall.sh scripts in there (the same for dogs). These shell scripts call similarly named .sql scripts under the hood, look inside them to see how they work. You can use similar patterns to construct your own stress test queries.
  • You currently need to change the "tpcc" to your database name (if you're not using "tpcc").
  • You can uncomment more psql lines to increase concurrency (and hit CTRL+C in terminal to cancel/kill all currently running psql loops`

Other

The data/PetImages directory is the Kaggle Cat/Dog dataset (total 25k images) originally released by Microsoft:

You don't need to separately download this file as it's already included in this repo (as permitted by Microsoft's CDLA license).

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

#CatBench Vector Search Playground Cat Benchmarking at Scale, finally!

There are two separate Python apps in the app directory:

  • CatVector - a simple static vector heatmap visualization app (no database)
  • CatBench - a simple Python/Flask application using Postgres+pgvector similarity search queries (and joins to a regular TPCC schema)

Go to installation steps below

CatBench

You can test this app out yourself, installation steps are below.

Here are a few screenshots of the similarity search and recommendation engine app (for cats!) in action:

Cat similarity search outputCat similarity search queryCat recommendation engine outputCat recommendation engine query planCat recommendation engine query plan

Installation Steps

25000 cat/dog images are included in this repository. I have tested this on RHEL9 and Ubuntu 24.04 so far. You need to have python and pip installed in your OS for this. For installing Python packages locally with pip, you probably want to use a Python virtual environment (venv).

Interactive CatBench application that requires a Postgres database and loading data

Make sure that you have a Postgres database (with pgvector extension) running and accessible and change the psql commands below to include your username/password if you are not using a default local connection:

In the catbench repo root directory, run this to generate embedding vectors from the 25000 pet images (this uses PyTorch which automatically runs on CPUs if you don't have a GPU available).

git clone https://github.com/tanelpoder/catbench
cd catbench
pip install -r requirements-catbench.txt

NB! You may need to install Postgres and the PgVector extension and the python3-psycopg2 package using your OS package manager first, if pip doesn't successfully install psycopg2 on your Linux distro.

The next step generates vector embeddings for the 25000 pet photos included in this repository (using GPU's if cuda/NVIDIA GPUs are available, otherwise CPUs.

Process the 25000 pet photos and generate their embeddings for loading into postgres:

python app/catbench/scripts/generate_embeddings.py data/PetImages/Cat embeddings/cats.tsv
python app/catbench/scripts/generate_embeddings.py data/PetImages/Dog embeddings/dogs.tsv

This may take a while. Then load the vectors and other OLTP data into the database:

gunzip app/catbench/scripts/create_tpcc_tables.sql.gz
psql -f app/catbench/scripts/create_tpcc_tables.sql psql -f app/catbench/scripts/create_catbench_tables.sql psql -f app/catbench/scripts/create_recommendation_schema.sql 

If you're using a local Postgres instance that allows logging in as tpcc user without a password, no action needed. Otherwise open the catbench.py file to change your Postgres user/pass settings if you are not using a default local connection. And then run the app:

cd app/catbench
python3 catbench.py

You can now go to hostname:5000 and browse around:

CatBench app frontpage

Stress test

  • Check the app/catbench/scripts/ directory and run cat_loop.sh or cat_loop_wit_recall.sh scripts in there (the same for dogs). These shell scripts call similarly named .sql scripts under the hood, look inside them to see how they work. You can use similar patterns to construct your own stress test queries.
  • You currently need to change the "tpcc" to your database name (if you're not using "tpcc").
  • You can uncomment more psql lines to increase concurrency (and hit CTRL+C in terminal to cancel/kill all currently running psql loops`

Other

The data/PetImages directory is the Kaggle Cat/Dog dataset (total 25k images) originally released by Microsoft:

You don't need to separately download this file as it's already included in this repo (as permitted by Microsoft's CDLA license).

Releases

Packages

Contributors

Languages