Repository files navigation

░█████████ ░███ ░██████ ░███████ ░████████ ░██ ░██ ░██░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░███████ ░█████████ ░█████████ ░██ █████ ░██ ░██ ░████████ ░██ ░██ ░██ ░██ ░██ ░██ ██ ░██ ░██ ░██ ░██ ░███████ ░██ ░██ ░██ ░██ ░██ ░███ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░█████░█ ░███████ ░█████████ ░███████ 

RAG Database Benchmarking Suite

This repository is designed for benchmarking popular Retrieval-Augmented Generation (RAG) databases. It enables you to:

  • Measure latency for reading and writing data to various vector databases.
  • Visualize metrics and compare performance using a built-in Streamlit dashboard and Grafana.
  • Easily extend the framework to support new database providers.
  • Modify and experiment with the codebase to suit your research or production needs.

Features

  • Streamlit Dashboard: Query databases, visualize results, and compare metrics interactively.
  • Extensible Structure: Add new database providers with minimal code changes.
  • Dockerized Services: Each database runs in its own Docker service for easy setup and isolation.
  • Metrics Visualization: Out-of-the-box Grafana dashboard for latency and token usage.

Quick Start

1. Clone & Install

Note: You need to have Docker installed

git clone <this-repo>cd rag

Set up LangChain and LangSmith

To visualize metrics, you will need to set up Langchain/Langsmith tracking. It is free and easy.

  1. Create Langchain Account
  2. Get API Key
  3. Create a project

Keep note of your API Key and Project name, you will need those later.

2. Environment Setup

Create a .env file in the project root:

# Guardian API related informationGUARDIAN_API_KEY=<your key here># Claude related informationANTHROPIC_API_KEY=<your key here># Langsmith related informationLANGSMITH_API_KEY=<your key here>LANGSMITH_ENDPOINT="https://api.smith.langchain.com"LANGSMITH_PROJECT=<name of your langsmith project>LANGSMITH_TRACING="true"LANGSMITH_API_KEY="<your-key>"DATABASE_TYPE="<name of database>"HOST="localhost" or shared IPv4 address 
  1. Install PostgreSQL; ensure you write down your login somewhere safe.

  2. build Docker image for PostgreSQL using 'docker-compose build --no-cache"

  3. Run Docker for PostgreSQL using docker-compose up

  4. When done, close Docker using docker-compose down

Running

Mac

In working directory, run chmod +x run.sh && ./run.sh.

Windows

Run run.bat.

Docker

Use the POST endpoint on the Docker container to upload articles to either database; this will upload 10 articles from the Guardian API.

If you're hosting a shared database, run docker compose up --build; this loads up the app with uvicorn at 0.0.0.0:8000.

Streamlit (Docker)

Run docker compose build(if not built previously) & docker compose up in rag/services/streamlit

Streamlit (Locally)

Ensure you've pip installed requirements in your local virtual env.

On Mac, run the script at dev/run.sh

On Windows, run the script at dev/run.bat (by clicking it in file explorer)

Local PostgreSQL

Run uvicorn services.postgres_controller:app --reload --port 8001.

(You will need to activate your venv & pip install -r requirements.txt)

Local Clickhouse

Run uvicorn services.clickhouse_controller:app --reload --port 8000.

(You will need to activate your venv & pip install -r requirements.txt)

Local Cassandra

Run uvicorn services.cassandra.cassandra_controller:app --reload.

Local Grafana

  1. cd into the llm folder.
  2. Run docker-compose up; this will create a Docker container for Grafana on port 3000.
  3. In your browser, open localhost:3000, and login using username admin and password admin.
  4. Add a new data source with the following parameters:
    1. Server address: the IP of your container.
    2. Server port: 8123
    3. Protocol: HTTP
    4. Skip TLS Verify: true
    5. Username: user
    6. Password: default
    7. Default database: guardian
    8. Default table: langchain_metrics
  5. Save and test the data source, and verify it connects successfully.
  6. Import dashboard.json from the llm folder.
  7. You should now have a local Grafana with metrics data; the metrics tab on the Streamlit application should now also work.

DEV

If you want to use the LLM and/or use the POST steps of the langchain pipeline, set the variables USE_POST and USE_LLM to "true" in your .env file.

Endpoints

GET

curl "http://localhost:8000/related-articles?query="

POST

curl -X POST "http://localhost:8000/upload-articles"

License

POSTGRES related information

POSTGRES_DB= POSTGRES_USER= POSTGRES_PASSWORD= POSTGRES_PORT=<port postgres will run on, normally 5432>


#### 3. Run Services
You have the option to choose between ClickHouse, PostgreSQL, and Cassandra as databases to use in benchmarking.
You can run any of them by the following (this example sets up ClickHouse):
```bash
cd services/clickhouse
docker compose up --build
cd services/streamlit
docker compose up --build

Your streamlit application should now be running.

4. Visualize Metrics with Grafana

  • The Streamlit service automatically creates a Grafana dashboard within the app, accessble at https://localhost:8501.

Benchmarking & Querying

  • Use the Streamlit dashboard to run queries against any configured database.
  • Latency and token usage metrics are automatically collected and visualized.
  • You can choose from Single Question, Multi Batch Question (simulating concurrent users), and Multi Batch Multi Questions (simulating multiple users with different questions).

Adding a New Database Provider

  1. Create a new folder under /services (e.g., /services/mydb).

  2. Add your Docker setup (Dockerfile, docker-compose.yml) and controller/DAO code.

  3. Extend the Database Enum in llm/llm_utils/langchain_pipeline.py:

    classDatabase(Enum):
    CLICKHOUSE= ("clickhouse", 8000)
    POSTGRES= ("postgres", 8001)
    CASSANDRA= ("cassandra", 8003)
    MYDB= ("mydb", <your-port>)
  4. Implement your controller to expose the required endpoints (/related-articles, /upload-articles).

  5. Update the dashboard (if needed) to include your new provider in the selection.


Customization & Extensibility

  • The codebase is modular—add new database backends by following the existing service structure.
  • The Streamlit dashboard auto-detects available databases from the Database Enum.
  • Metrics collection is built-in; extend or modify as needed for your research.

License

MIT


Happy benchmarking!


About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

░█████████ ░███ ░██████ ░███████ ░████████ ░██ ░██ ░██░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░███████ ░█████████ ░█████████ ░██ █████ ░██ ░██ ░████████ ░██ ░██ ░██ ░██ ░██ ░██ ██ ░██ ░██ ░██ ░██ ░███████ ░██ ░██ ░██ ░██ ░██ ░███ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░█████░█ ░███████ ░█████████ ░███████ 

RAG Database Benchmarking Suite

This repository is designed for benchmarking popular Retrieval-Augmented Generation (RAG) databases. It enables you to:

  • Measure latency for reading and writing data to various vector databases.
  • Visualize metrics and compare performance using a built-in Streamlit dashboard and Grafana.
  • Easily extend the framework to support new database providers.
  • Modify and experiment with the codebase to suit your research or production needs.

Features

  • Streamlit Dashboard: Query databases, visualize results, and compare metrics interactively.
  • Extensible Structure: Add new database providers with minimal code changes.
  • Dockerized Services: Each database runs in its own Docker service for easy setup and isolation.
  • Metrics Visualization: Out-of-the-box Grafana dashboard for latency and token usage.

Quick Start

1. Clone & Install

Note: You need to have Docker installed

git clone <this-repo>cd rag

Set up LangChain and LangSmith

To visualize metrics, you will need to set up Langchain/Langsmith tracking. It is free and easy.

  1. Create Langchain Account
  2. Get API Key
  3. Create a project

Keep note of your API Key and Project name, you will need those later.

2. Environment Setup

Create a .env file in the project root:

# Guardian API related informationGUARDIAN_API_KEY=<your key here># Claude related informationANTHROPIC_API_KEY=<your key here># Langsmith related informationLANGSMITH_API_KEY=<your key here>LANGSMITH_ENDPOINT="https://api.smith.langchain.com"LANGSMITH_PROJECT=<name of your langsmith project>LANGSMITH_TRACING="true"LANGSMITH_API_KEY="<your-key>"DATABASE_TYPE="<name of database>"HOST="localhost" or shared IPv4 address 
  1. Install PostgreSQL; ensure you write down your login somewhere safe.

  2. build Docker image for PostgreSQL using 'docker-compose build --no-cache"

  3. Run Docker for PostgreSQL using docker-compose up

  4. When done, close Docker using docker-compose down

Running

Mac

In working directory, run chmod +x run.sh && ./run.sh.

Windows

Run run.bat.

Docker

Use the POST endpoint on the Docker container to upload articles to either database; this will upload 10 articles from the Guardian API.

If you're hosting a shared database, run docker compose up --build; this loads up the app with uvicorn at 0.0.0.0:8000.

Streamlit (Docker)

Run docker compose build(if not built previously) & docker compose up in rag/services/streamlit

Streamlit (Locally)

Ensure you've pip installed requirements in your local virtual env.

On Mac, run the script at dev/run.sh

On Windows, run the script at dev/run.bat (by clicking it in file explorer)

Local PostgreSQL

Run uvicorn services.postgres_controller:app --reload --port 8001.

(You will need to activate your venv & pip install -r requirements.txt)

Local Clickhouse

Run uvicorn services.clickhouse_controller:app --reload --port 8000.

(You will need to activate your venv & pip install -r requirements.txt)

Local Cassandra

Run uvicorn services.cassandra.cassandra_controller:app --reload.

Local Grafana

  1. cd into the llm folder.
  2. Run docker-compose up; this will create a Docker container for Grafana on port 3000.
  3. In your browser, open localhost:3000, and login using username admin and password admin.
  4. Add a new data source with the following parameters:
    1. Server address: the IP of your container.
    2. Server port: 8123
    3. Protocol: HTTP
    4. Skip TLS Verify: true
    5. Username: user
    6. Password: default
    7. Default database: guardian
    8. Default table: langchain_metrics
  5. Save and test the data source, and verify it connects successfully.
  6. Import dashboard.json from the llm folder.
  7. You should now have a local Grafana with metrics data; the metrics tab on the Streamlit application should now also work.

DEV

If you want to use the LLM and/or use the POST steps of the langchain pipeline, set the variables USE_POST and USE_LLM to "true" in your .env file.

Endpoints

GET

curl "http://localhost:8000/related-articles?query="

POST

curl -X POST "http://localhost:8000/upload-articles"

License

POSTGRES related information

POSTGRES_DB= POSTGRES_USER= POSTGRES_PASSWORD= POSTGRES_PORT=<port postgres will run on, normally 5432>


#### 3. Run Services
You have the option to choose between ClickHouse, PostgreSQL, and Cassandra as databases to use in benchmarking.
You can run any of them by the following (this example sets up ClickHouse):
```bash
cd services/clickhouse
docker compose up --build
cd services/streamlit
docker compose up --build

Your streamlit application should now be running.

4. Visualize Metrics with Grafana

  • The Streamlit service automatically creates a Grafana dashboard within the app, accessble at https://localhost:8501.

Benchmarking & Querying

  • Use the Streamlit dashboard to run queries against any configured database.
  • Latency and token usage metrics are automatically collected and visualized.
  • You can choose from Single Question, Multi Batch Question (simulating concurrent users), and Multi Batch Multi Questions (simulating multiple users with different questions).

Adding a New Database Provider

  1. Create a new folder under /services (e.g., /services/mydb).

  2. Add your Docker setup (Dockerfile, docker-compose.yml) and controller/DAO code.

  3. Extend the Database Enum in llm/llm_utils/langchain_pipeline.py:

    classDatabase(Enum):
    CLICKHOUSE= ("clickhouse", 8000)
    POSTGRES= ("postgres", 8001)
    CASSANDRA= ("cassandra", 8003)
    MYDB= ("mydb", <your-port>)
  4. Implement your controller to expose the required endpoints (/related-articles, /upload-articles).

  5. Update the dashboard (if needed) to include your new provider in the selection.


Customization & Extensibility

  • The codebase is modular—add new database backends by following the existing service structure.
  • The Streamlit dashboard auto-detects available databases from the Database Enum.
  • Metrics collection is built-in; extend or modify as needed for your research.

License

MIT


Happy benchmarking!


About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

░█████████ ░███ ░██████ ░███████ ░████████ ░██ ░██ ░██░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░███████ ░█████████ ░█████████ ░██ █████ ░██ ░██ ░████████ ░██ ░██ ░██ ░██ ░██ ░██ ██ ░██ ░██ ░██ ░██ ░███████ ░██ ░██ ░██ ░██ ░██ ░███ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░█████░█ ░███████ ░█████████ ░███████ 

RAG Database Benchmarking Suite

This repository is designed for benchmarking popular Retrieval-Augmented Generation (RAG) databases. It enables you to:

  • Measure latency for reading and writing data to various vector databases.
  • Visualize metrics and compare performance using a built-in Streamlit dashboard and Grafana.
  • Easily extend the framework to support new database providers.
  • Modify and experiment with the codebase to suit your research or production needs.

Features

  • Streamlit Dashboard: Query databases, visualize results, and compare metrics interactively.
  • Extensible Structure: Add new database providers with minimal code changes.
  • Dockerized Services: Each database runs in its own Docker service for easy setup and isolation.
  • Metrics Visualization: Out-of-the-box Grafana dashboard for latency and token usage.

Quick Start

1. Clone & Install

Note: You need to have Docker installed

git clone <this-repo>cd rag

Set up LangChain and LangSmith

To visualize metrics, you will need to set up Langchain/Langsmith tracking. It is free and easy.

  1. Create Langchain Account
  2. Get API Key
  3. Create a project

Keep note of your API Key and Project name, you will need those later.

2. Environment Setup

Create a .env file in the project root:

# Guardian API related informationGUARDIAN_API_KEY=<your key here># Claude related informationANTHROPIC_API_KEY=<your key here># Langsmith related informationLANGSMITH_API_KEY=<your key here>LANGSMITH_ENDPOINT="https://api.smith.langchain.com"LANGSMITH_PROJECT=<name of your langsmith project>LANGSMITH_TRACING="true"LANGSMITH_API_KEY="<your-key>"DATABASE_TYPE="<name of database>"HOST="localhost" or shared IPv4 address 
  1. Install PostgreSQL; ensure you write down your login somewhere safe.

  2. build Docker image for PostgreSQL using 'docker-compose build --no-cache"

  3. Run Docker for PostgreSQL using docker-compose up

  4. When done, close Docker using docker-compose down

Running

Mac

In working directory, run chmod +x run.sh && ./run.sh.

Windows

Run run.bat.

Docker

Use the POST endpoint on the Docker container to upload articles to either database; this will upload 10 articles from the Guardian API.

If you're hosting a shared database, run docker compose up --build; this loads up the app with uvicorn at 0.0.0.0:8000.

Streamlit (Docker)

Run docker compose build(if not built previously) & docker compose up in rag/services/streamlit

Streamlit (Locally)

Ensure you've pip installed requirements in your local virtual env.

On Mac, run the script at dev/run.sh

On Windows, run the script at dev/run.bat (by clicking it in file explorer)

Local PostgreSQL

Run uvicorn services.postgres_controller:app --reload --port 8001.

(You will need to activate your venv & pip install -r requirements.txt)

Local Clickhouse

Run uvicorn services.clickhouse_controller:app --reload --port 8000.

(You will need to activate your venv & pip install -r requirements.txt)

Local Cassandra

Run uvicorn services.cassandra.cassandra_controller:app --reload.

Local Grafana

  1. cd into the llm folder.
  2. Run docker-compose up; this will create a Docker container for Grafana on port 3000.
  3. In your browser, open localhost:3000, and login using username admin and password admin.
  4. Add a new data source with the following parameters:
    1. Server address: the IP of your container.
    2. Server port: 8123
    3. Protocol: HTTP
    4. Skip TLS Verify: true
    5. Username: user
    6. Password: default
    7. Default database: guardian
    8. Default table: langchain_metrics
  5. Save and test the data source, and verify it connects successfully.
  6. Import dashboard.json from the llm folder.
  7. You should now have a local Grafana with metrics data; the metrics tab on the Streamlit application should now also work.

DEV

If you want to use the LLM and/or use the POST steps of the langchain pipeline, set the variables USE_POST and USE_LLM to "true" in your .env file.

Endpoints

GET

curl "http://localhost:8000/related-articles?query="

POST

curl -X POST "http://localhost:8000/upload-articles"

License

POSTGRES related information

POSTGRES_DB= POSTGRES_USER= POSTGRES_PASSWORD= POSTGRES_PORT=<port postgres will run on, normally 5432>


#### 3. Run Services
You have the option to choose between ClickHouse, PostgreSQL, and Cassandra as databases to use in benchmarking.
You can run any of them by the following (this example sets up ClickHouse):
```bash
cd services/clickhouse
docker compose up --build
cd services/streamlit
docker compose up --build

Your streamlit application should now be running.

4. Visualize Metrics with Grafana

  • The Streamlit service automatically creates a Grafana dashboard within the app, accessble at https://localhost:8501.

Benchmarking & Querying

  • Use the Streamlit dashboard to run queries against any configured database.
  • Latency and token usage metrics are automatically collected and visualized.
  • You can choose from Single Question, Multi Batch Question (simulating concurrent users), and Multi Batch Multi Questions (simulating multiple users with different questions).

Adding a New Database Provider

  1. Create a new folder under /services (e.g., /services/mydb).

  2. Add your Docker setup (Dockerfile, docker-compose.yml) and controller/DAO code.

  3. Extend the Database Enum in llm/llm_utils/langchain_pipeline.py:

    classDatabase(Enum):
    CLICKHOUSE= ("clickhouse", 8000)
    POSTGRES= ("postgres", 8001)
    CASSANDRA= ("cassandra", 8003)
    MYDB= ("mydb", <your-port>)
  4. Implement your controller to expose the required endpoints (/related-articles, /upload-articles).

  5. Update the dashboard (if needed) to include your new provider in the selection.


Customization & Extensibility

  • The codebase is modular—add new database backends by following the existing service structure.
  • The Streamlit dashboard auto-detects available databases from the Database Enum.
  • Metrics collection is built-in; extend or modify as needed for your research.

License

MIT


Happy benchmarking!


About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

░█████████ ░███ ░██████ ░███████ ░████████ ░██ ░██ ░██░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░███████ ░█████████ ░█████████ ░██ █████ ░██ ░██ ░████████ ░██ ░██ ░██ ░██ ░██ ░██ ██ ░██ ░██ ░██ ░██ ░███████ ░██ ░██ ░██ ░██ ░██ ░███ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░█████░█ ░███████ ░█████████ ░███████ 

RAG Database Benchmarking Suite

This repository is designed for benchmarking popular Retrieval-Augmented Generation (RAG) databases. It enables you to:

  • Measure latency for reading and writing data to various vector databases.
  • Visualize metrics and compare performance using a built-in Streamlit dashboard and Grafana.
  • Easily extend the framework to support new database providers.
  • Modify and experiment with the codebase to suit your research or production needs.

Features

  • Streamlit Dashboard: Query databases, visualize results, and compare metrics interactively.
  • Extensible Structure: Add new database providers with minimal code changes.
  • Dockerized Services: Each database runs in its own Docker service for easy setup and isolation.
  • Metrics Visualization: Out-of-the-box Grafana dashboard for latency and token usage.

Quick Start

1. Clone & Install

Note: You need to have Docker installed

git clone <this-repo>cd rag

Set up LangChain and LangSmith

To visualize metrics, you will need to set up Langchain/Langsmith tracking. It is free and easy.

  1. Create Langchain Account
  2. Get API Key
  3. Create a project

Keep note of your API Key and Project name, you will need those later.

2. Environment Setup

Create a .env file in the project root:

# Guardian API related informationGUARDIAN_API_KEY=<your key here># Claude related informationANTHROPIC_API_KEY=<your key here># Langsmith related informationLANGSMITH_API_KEY=<your key here>LANGSMITH_ENDPOINT="https://api.smith.langchain.com"LANGSMITH_PROJECT=<name of your langsmith project>LANGSMITH_TRACING="true"LANGSMITH_API_KEY="<your-key>"DATABASE_TYPE="<name of database>"HOST="localhost" or shared IPv4 address 
  1. Install PostgreSQL; ensure you write down your login somewhere safe.

  2. build Docker image for PostgreSQL using 'docker-compose build --no-cache"

  3. Run Docker for PostgreSQL using docker-compose up

  4. When done, close Docker using docker-compose down

Running

Mac

In working directory, run chmod +x run.sh && ./run.sh.

Windows

Run run.bat.

Docker

Use the POST endpoint on the Docker container to upload articles to either database; this will upload 10 articles from the Guardian API.

If you're hosting a shared database, run docker compose up --build; this loads up the app with uvicorn at 0.0.0.0:8000.

Streamlit (Docker)

Run docker compose build(if not built previously) & docker compose up in rag/services/streamlit

Streamlit (Locally)

Ensure you've pip installed requirements in your local virtual env.

On Mac, run the script at dev/run.sh

On Windows, run the script at dev/run.bat (by clicking it in file explorer)

Local PostgreSQL

Run uvicorn services.postgres_controller:app --reload --port 8001.

(You will need to activate your venv & pip install -r requirements.txt)

Local Clickhouse

Run uvicorn services.clickhouse_controller:app --reload --port 8000.

(You will need to activate your venv & pip install -r requirements.txt)

Local Cassandra

Run uvicorn services.cassandra.cassandra_controller:app --reload.

Local Grafana

  1. cd into the llm folder.
  2. Run docker-compose up; this will create a Docker container for Grafana on port 3000.
  3. In your browser, open localhost:3000, and login using username admin and password admin.
  4. Add a new data source with the following parameters:
    1. Server address: the IP of your container.
    2. Server port: 8123
    3. Protocol: HTTP
    4. Skip TLS Verify: true
    5. Username: user
    6. Password: default
    7. Default database: guardian
    8. Default table: langchain_metrics
  5. Save and test the data source, and verify it connects successfully.
  6. Import dashboard.json from the llm folder.
  7. You should now have a local Grafana with metrics data; the metrics tab on the Streamlit application should now also work.

DEV

If you want to use the LLM and/or use the POST steps of the langchain pipeline, set the variables USE_POST and USE_LLM to "true" in your .env file.

Endpoints

GET

curl "http://localhost:8000/related-articles?query="

POST

curl -X POST "http://localhost:8000/upload-articles"

License

POSTGRES related information

POSTGRES_DB= POSTGRES_USER= POSTGRES_PASSWORD= POSTGRES_PORT=<port postgres will run on, normally 5432>


#### 3. Run Services
You have the option to choose between ClickHouse, PostgreSQL, and Cassandra as databases to use in benchmarking.
You can run any of them by the following (this example sets up ClickHouse):
```bash
cd services/clickhouse
docker compose up --build
cd services/streamlit
docker compose up --build

Your streamlit application should now be running.

4. Visualize Metrics with Grafana

  • The Streamlit service automatically creates a Grafana dashboard within the app, accessble at https://localhost:8501.

Benchmarking & Querying

  • Use the Streamlit dashboard to run queries against any configured database.
  • Latency and token usage metrics are automatically collected and visualized.
  • You can choose from Single Question, Multi Batch Question (simulating concurrent users), and Multi Batch Multi Questions (simulating multiple users with different questions).

Adding a New Database Provider

  1. Create a new folder under /services (e.g., /services/mydb).

  2. Add your Docker setup (Dockerfile, docker-compose.yml) and controller/DAO code.

  3. Extend the Database Enum in llm/llm_utils/langchain_pipeline.py:

    classDatabase(Enum):
    CLICKHOUSE= ("clickhouse", 8000)
    POSTGRES= ("postgres", 8001)
    CASSANDRA= ("cassandra", 8003)
    MYDB= ("mydb", <your-port>)
  4. Implement your controller to expose the required endpoints (/related-articles, /upload-articles).

  5. Update the dashboard (if needed) to include your new provider in the selection.


Customization & Extensibility

  • The codebase is modular—add new database backends by following the existing service structure.
  • The Streamlit dashboard auto-detects available databases from the Database Enum.
  • Metrics collection is built-in; extend or modify as needed for your research.

License

MIT


Happy benchmarking!


About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

░█████████ ░███ ░██████ ░███████ ░████████ ░██ ░██ ░██░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░███████ ░█████████ ░█████████ ░██ █████ ░██ ░██ ░████████ ░██ ░██ ░██ ░██ ░██ ░██ ██ ░██ ░██ ░██ ░██ ░███████ ░██ ░██ ░██ ░██ ░██ ░███ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░█████░█ ░███████ ░█████████ ░███████ 

RAG Database Benchmarking Suite

This repository is designed for benchmarking popular Retrieval-Augmented Generation (RAG) databases. It enables you to:

  • Measure latency for reading and writing data to various vector databases.
  • Visualize metrics and compare performance using a built-in Streamlit dashboard and Grafana.
  • Easily extend the framework to support new database providers.
  • Modify and experiment with the codebase to suit your research or production needs.

Features

  • Streamlit Dashboard: Query databases, visualize results, and compare metrics interactively.
  • Extensible Structure: Add new database providers with minimal code changes.
  • Dockerized Services: Each database runs in its own Docker service for easy setup and isolation.
  • Metrics Visualization: Out-of-the-box Grafana dashboard for latency and token usage.

Quick Start

1. Clone & Install

Note: You need to have Docker installed

git clone <this-repo>cd rag

Set up LangChain and LangSmith

To visualize metrics, you will need to set up Langchain/Langsmith tracking. It is free and easy.

  1. Create Langchain Account
  2. Get API Key
  3. Create a project

Keep note of your API Key and Project name, you will need those later.

2. Environment Setup

Create a .env file in the project root:

# Guardian API related informationGUARDIAN_API_KEY=<your key here># Claude related informationANTHROPIC_API_KEY=<your key here># Langsmith related informationLANGSMITH_API_KEY=<your key here>LANGSMITH_ENDPOINT="https://api.smith.langchain.com"LANGSMITH_PROJECT=<name of your langsmith project>LANGSMITH_TRACING="true"LANGSMITH_API_KEY="<your-key>"DATABASE_TYPE="<name of database>"HOST="localhost" or shared IPv4 address 
  1. Install PostgreSQL; ensure you write down your login somewhere safe.

  2. build Docker image for PostgreSQL using 'docker-compose build --no-cache"

  3. Run Docker for PostgreSQL using docker-compose up

  4. When done, close Docker using docker-compose down

Running

Mac

In working directory, run chmod +x run.sh && ./run.sh.

Windows

Run run.bat.

Docker

Use the POST endpoint on the Docker container to upload articles to either database; this will upload 10 articles from the Guardian API.

If you're hosting a shared database, run docker compose up --build; this loads up the app with uvicorn at 0.0.0.0:8000.

Streamlit (Docker)

Run docker compose build(if not built previously) & docker compose up in rag/services/streamlit

Streamlit (Locally)

Ensure you've pip installed requirements in your local virtual env.

On Mac, run the script at dev/run.sh

On Windows, run the script at dev/run.bat (by clicking it in file explorer)

Local PostgreSQL

Run uvicorn services.postgres_controller:app --reload --port 8001.

(You will need to activate your venv & pip install -r requirements.txt)

Local Clickhouse

Run uvicorn services.clickhouse_controller:app --reload --port 8000.

(You will need to activate your venv & pip install -r requirements.txt)

Local Cassandra

Run uvicorn services.cassandra.cassandra_controller:app --reload.

Local Grafana

  1. cd into the llm folder.
  2. Run docker-compose up; this will create a Docker container for Grafana on port 3000.
  3. In your browser, open localhost:3000, and login using username admin and password admin.
  4. Add a new data source with the following parameters:
    1. Server address: the IP of your container.
    2. Server port: 8123
    3. Protocol: HTTP
    4. Skip TLS Verify: true
    5. Username: user
    6. Password: default
    7. Default database: guardian
    8. Default table: langchain_metrics
  5. Save and test the data source, and verify it connects successfully.
  6. Import dashboard.json from the llm folder.
  7. You should now have a local Grafana with metrics data; the metrics tab on the Streamlit application should now also work.

DEV

If you want to use the LLM and/or use the POST steps of the langchain pipeline, set the variables USE_POST and USE_LLM to "true" in your .env file.

Endpoints

GET

curl "http://localhost:8000/related-articles?query="

POST

curl -X POST "http://localhost:8000/upload-articles"

License

POSTGRES related information

POSTGRES_DB= POSTGRES_USER= POSTGRES_PASSWORD= POSTGRES_PORT=<port postgres will run on, normally 5432>


#### 3. Run Services
You have the option to choose between ClickHouse, PostgreSQL, and Cassandra as databases to use in benchmarking.
You can run any of them by the following (this example sets up ClickHouse):
```bash
cd services/clickhouse
docker compose up --build
cd services/streamlit
docker compose up --build

Your streamlit application should now be running.

4. Visualize Metrics with Grafana

  • The Streamlit service automatically creates a Grafana dashboard within the app, accessble at https://localhost:8501.

Benchmarking & Querying

  • Use the Streamlit dashboard to run queries against any configured database.
  • Latency and token usage metrics are automatically collected and visualized.
  • You can choose from Single Question, Multi Batch Question (simulating concurrent users), and Multi Batch Multi Questions (simulating multiple users with different questions).

Adding a New Database Provider

  1. Create a new folder under /services (e.g., /services/mydb).

  2. Add your Docker setup (Dockerfile, docker-compose.yml) and controller/DAO code.

  3. Extend the Database Enum in llm/llm_utils/langchain_pipeline.py:

    classDatabase(Enum):
    CLICKHOUSE= ("clickhouse", 8000)
    POSTGRES= ("postgres", 8001)
    CASSANDRA= ("cassandra", 8003)
    MYDB= ("mydb", <your-port>)
  4. Implement your controller to expose the required endpoints (/related-articles, /upload-articles).

  5. Update the dashboard (if needed) to include your new provider in the selection.


Customization & Extensibility

  • The codebase is modular—add new database backends by following the existing service structure.
  • The Streamlit dashboard auto-detects available databases from the Database Enum.
  • Metrics collection is built-in; extend or modify as needed for your research.

License

MIT


Happy benchmarking!


About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

░█████████ ░███ ░██████ ░███████ ░████████ ░██ ░██ ░██░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░███████ ░█████████ ░█████████ ░██ █████ ░██ ░██ ░████████ ░██ ░██ ░██ ░██ ░██ ░██ ██ ░██ ░██ ░██ ░██ ░███████ ░██ ░██ ░██ ░██ ░██ ░███ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░█████░█ ░███████ ░█████████ ░███████ 

RAG Database Benchmarking Suite

This repository is designed for benchmarking popular Retrieval-Augmented Generation (RAG) databases. It enables you to:

  • Measure latency for reading and writing data to various vector databases.
  • Visualize metrics and compare performance using a built-in Streamlit dashboard and Grafana.
  • Easily extend the framework to support new database providers.
  • Modify and experiment with the codebase to suit your research or production needs.

Features

  • Streamlit Dashboard: Query databases, visualize results, and compare metrics interactively.
  • Extensible Structure: Add new database providers with minimal code changes.
  • Dockerized Services: Each database runs in its own Docker service for easy setup and isolation.
  • Metrics Visualization: Out-of-the-box Grafana dashboard for latency and token usage.

Quick Start

1. Clone & Install

Note: You need to have Docker installed

git clone <this-repo>cd rag

Set up LangChain and LangSmith

To visualize metrics, you will need to set up Langchain/Langsmith tracking. It is free and easy.

  1. Create Langchain Account
  2. Get API Key
  3. Create a project

Keep note of your API Key and Project name, you will need those later.

2. Environment Setup

Create a .env file in the project root:

# Guardian API related informationGUARDIAN_API_KEY=<your key here># Claude related informationANTHROPIC_API_KEY=<your key here># Langsmith related informationLANGSMITH_API_KEY=<your key here>LANGSMITH_ENDPOINT="https://api.smith.langchain.com"LANGSMITH_PROJECT=<name of your langsmith project>LANGSMITH_TRACING="true"LANGSMITH_API_KEY="<your-key>"DATABASE_TYPE="<name of database>"HOST="localhost" or shared IPv4 address 
  1. Install PostgreSQL; ensure you write down your login somewhere safe.

  2. build Docker image for PostgreSQL using 'docker-compose build --no-cache"

  3. Run Docker for PostgreSQL using docker-compose up

  4. When done, close Docker using docker-compose down

Running

Mac

In working directory, run chmod +x run.sh && ./run.sh.

Windows

Run run.bat.

Docker

Use the POST endpoint on the Docker container to upload articles to either database; this will upload 10 articles from the Guardian API.

If you're hosting a shared database, run docker compose up --build; this loads up the app with uvicorn at 0.0.0.0:8000.

Streamlit (Docker)

Run docker compose build(if not built previously) & docker compose up in rag/services/streamlit

Streamlit (Locally)

Ensure you've pip installed requirements in your local virtual env.

On Mac, run the script at dev/run.sh

On Windows, run the script at dev/run.bat (by clicking it in file explorer)

Local PostgreSQL

Run uvicorn services.postgres_controller:app --reload --port 8001.

(You will need to activate your venv & pip install -r requirements.txt)

Local Clickhouse

Run uvicorn services.clickhouse_controller:app --reload --port 8000.

(You will need to activate your venv & pip install -r requirements.txt)

Local Cassandra

Run uvicorn services.cassandra.cassandra_controller:app --reload.

Local Grafana

  1. cd into the llm folder.
  2. Run docker-compose up; this will create a Docker container for Grafana on port 3000.
  3. In your browser, open localhost:3000, and login using username admin and password admin.
  4. Add a new data source with the following parameters:
    1. Server address: the IP of your container.
    2. Server port: 8123
    3. Protocol: HTTP
    4. Skip TLS Verify: true
    5. Username: user
    6. Password: default
    7. Default database: guardian
    8. Default table: langchain_metrics
  5. Save and test the data source, and verify it connects successfully.
  6. Import dashboard.json from the llm folder.
  7. You should now have a local Grafana with metrics data; the metrics tab on the Streamlit application should now also work.

DEV

If you want to use the LLM and/or use the POST steps of the langchain pipeline, set the variables USE_POST and USE_LLM to "true" in your .env file.

Endpoints

GET

curl "http://localhost:8000/related-articles?query="

POST

curl -X POST "http://localhost:8000/upload-articles"

License

POSTGRES related information

POSTGRES_DB= POSTGRES_USER= POSTGRES_PASSWORD= POSTGRES_PORT=<port postgres will run on, normally 5432>


#### 3. Run Services
You have the option to choose between ClickHouse, PostgreSQL, and Cassandra as databases to use in benchmarking.
You can run any of them by the following (this example sets up ClickHouse):
```bash
cd services/clickhouse
docker compose up --build
cd services/streamlit
docker compose up --build

Your streamlit application should now be running.

4. Visualize Metrics with Grafana

  • The Streamlit service automatically creates a Grafana dashboard within the app, accessble at https://localhost:8501.

Benchmarking & Querying

  • Use the Streamlit dashboard to run queries against any configured database.
  • Latency and token usage metrics are automatically collected and visualized.
  • You can choose from Single Question, Multi Batch Question (simulating concurrent users), and Multi Batch Multi Questions (simulating multiple users with different questions).

Adding a New Database Provider

  1. Create a new folder under /services (e.g., /services/mydb).

  2. Add your Docker setup (Dockerfile, docker-compose.yml) and controller/DAO code.

  3. Extend the Database Enum in llm/llm_utils/langchain_pipeline.py:

    classDatabase(Enum):
    CLICKHOUSE= ("clickhouse", 8000)
    POSTGRES= ("postgres", 8001)
    CASSANDRA= ("cassandra", 8003)
    MYDB= ("mydb", <your-port>)
  4. Implement your controller to expose the required endpoints (/related-articles, /upload-articles).

  5. Update the dashboard (if needed) to include your new provider in the selection.


Customization & Extensibility

  • The codebase is modular—add new database backends by following the existing service structure.
  • The Streamlit dashboard auto-detects available databases from the Database Enum.
  • Metrics collection is built-in; extend or modify as needed for your research.

License

MIT


Happy benchmarking!


About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

░█████████ ░███ ░██████ ░███████ ░████████ ░██ ░██ ░██░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░███████ ░█████████ ░█████████ ░██ █████ ░██ ░██ ░████████ ░██ ░██ ░██ ░██ ░██ ░██ ██ ░██ ░██ ░██ ░██ ░███████ ░██ ░██ ░██ ░██ ░██ ░███ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░█████░█ ░███████ ░█████████ ░███████ 

RAG Database Benchmarking Suite

This repository is designed for benchmarking popular Retrieval-Augmented Generation (RAG) databases. It enables you to:

  • Measure latency for reading and writing data to various vector databases.
  • Visualize metrics and compare performance using a built-in Streamlit dashboard and Grafana.
  • Easily extend the framework to support new database providers.
  • Modify and experiment with the codebase to suit your research or production needs.

Features

  • Streamlit Dashboard: Query databases, visualize results, and compare metrics interactively.
  • Extensible Structure: Add new database providers with minimal code changes.
  • Dockerized Services: Each database runs in its own Docker service for easy setup and isolation.
  • Metrics Visualization: Out-of-the-box Grafana dashboard for latency and token usage.

Quick Start

1. Clone & Install

Note: You need to have Docker installed

git clone <this-repo>cd rag

Set up LangChain and LangSmith

To visualize metrics, you will need to set up Langchain/Langsmith tracking. It is free and easy.

  1. Create Langchain Account
  2. Get API Key
  3. Create a project

Keep note of your API Key and Project name, you will need those later.

2. Environment Setup

Create a .env file in the project root:

# Guardian API related informationGUARDIAN_API_KEY=<your key here># Claude related informationANTHROPIC_API_KEY=<your key here># Langsmith related informationLANGSMITH_API_KEY=<your key here>LANGSMITH_ENDPOINT="https://api.smith.langchain.com"LANGSMITH_PROJECT=<name of your langsmith project>LANGSMITH_TRACING="true"LANGSMITH_API_KEY="<your-key>"DATABASE_TYPE="<name of database>"HOST="localhost" or shared IPv4 address 
  1. Install PostgreSQL; ensure you write down your login somewhere safe.

  2. build Docker image for PostgreSQL using 'docker-compose build --no-cache"

  3. Run Docker for PostgreSQL using docker-compose up

  4. When done, close Docker using docker-compose down

Running

Mac

In working directory, run chmod +x run.sh && ./run.sh.

Windows

Run run.bat.

Docker

Use the POST endpoint on the Docker container to upload articles to either database; this will upload 10 articles from the Guardian API.

If you're hosting a shared database, run docker compose up --build; this loads up the app with uvicorn at 0.0.0.0:8000.

Streamlit (Docker)

Run docker compose build(if not built previously) & docker compose up in rag/services/streamlit

Streamlit (Locally)

Ensure you've pip installed requirements in your local virtual env.

On Mac, run the script at dev/run.sh

On Windows, run the script at dev/run.bat (by clicking it in file explorer)

Local PostgreSQL

Run uvicorn services.postgres_controller:app --reload --port 8001.

(You will need to activate your venv & pip install -r requirements.txt)

Local Clickhouse

Run uvicorn services.clickhouse_controller:app --reload --port 8000.

(You will need to activate your venv & pip install -r requirements.txt)

Local Cassandra

Run uvicorn services.cassandra.cassandra_controller:app --reload.

Local Grafana

  1. cd into the llm folder.
  2. Run docker-compose up; this will create a Docker container for Grafana on port 3000.
  3. In your browser, open localhost:3000, and login using username admin and password admin.
  4. Add a new data source with the following parameters:
    1. Server address: the IP of your container.
    2. Server port: 8123
    3. Protocol: HTTP
    4. Skip TLS Verify: true
    5. Username: user
    6. Password: default
    7. Default database: guardian
    8. Default table: langchain_metrics
  5. Save and test the data source, and verify it connects successfully.
  6. Import dashboard.json from the llm folder.
  7. You should now have a local Grafana with metrics data; the metrics tab on the Streamlit application should now also work.

DEV

If you want to use the LLM and/or use the POST steps of the langchain pipeline, set the variables USE_POST and USE_LLM to "true" in your .env file.

Endpoints

GET

curl "http://localhost:8000/related-articles?query="

POST

curl -X POST "http://localhost:8000/upload-articles"

License

POSTGRES related information

POSTGRES_DB= POSTGRES_USER= POSTGRES_PASSWORD= POSTGRES_PORT=<port postgres will run on, normally 5432>


#### 3. Run Services
You have the option to choose between ClickHouse, PostgreSQL, and Cassandra as databases to use in benchmarking.
You can run any of them by the following (this example sets up ClickHouse):
```bash
cd services/clickhouse
docker compose up --build
cd services/streamlit
docker compose up --build

Your streamlit application should now be running.

4. Visualize Metrics with Grafana

  • The Streamlit service automatically creates a Grafana dashboard within the app, accessble at https://localhost:8501.

Benchmarking & Querying

  • Use the Streamlit dashboard to run queries against any configured database.
  • Latency and token usage metrics are automatically collected and visualized.
  • You can choose from Single Question, Multi Batch Question (simulating concurrent users), and Multi Batch Multi Questions (simulating multiple users with different questions).

Adding a New Database Provider

  1. Create a new folder under /services (e.g., /services/mydb).

  2. Add your Docker setup (Dockerfile, docker-compose.yml) and controller/DAO code.

  3. Extend the Database Enum in llm/llm_utils/langchain_pipeline.py:

    classDatabase(Enum):
    CLICKHOUSE= ("clickhouse", 8000)
    POSTGRES= ("postgres", 8001)
    CASSANDRA= ("cassandra", 8003)
    MYDB= ("mydb", <your-port>)
  4. Implement your controller to expose the required endpoints (/related-articles, /upload-articles).

  5. Update the dashboard (if needed) to include your new provider in the selection.


Customization & Extensibility

  • The codebase is modular—add new database backends by following the existing service structure.
  • The Streamlit dashboard auto-detects available databases from the Database Enum.
  • Metrics collection is built-in; extend or modify as needed for your research.

License

MIT


Happy benchmarking!


About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

░█████████ ░███ ░██████ ░███████ ░████████ ░██ ░██ ░██░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░███████ ░█████████ ░█████████ ░██ █████ ░██ ░██ ░████████ ░██ ░██ ░██ ░██ ░██ ░██ ██ ░██ ░██ ░██ ░██ ░███████ ░██ ░██ ░██ ░██ ░██ ░███ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░█████░█ ░███████ ░█████████ ░███████ 

RAG Database Benchmarking Suite

This repository is designed for benchmarking popular Retrieval-Augmented Generation (RAG) databases. It enables you to:

  • Measure latency for reading and writing data to various vector databases.
  • Visualize metrics and compare performance using a built-in Streamlit dashboard and Grafana.
  • Easily extend the framework to support new database providers.
  • Modify and experiment with the codebase to suit your research or production needs.

Features

  • Streamlit Dashboard: Query databases, visualize results, and compare metrics interactively.
  • Extensible Structure: Add new database providers with minimal code changes.
  • Dockerized Services: Each database runs in its own Docker service for easy setup and isolation.
  • Metrics Visualization: Out-of-the-box Grafana dashboard for latency and token usage.

Quick Start

1. Clone & Install

Note: You need to have Docker installed

git clone <this-repo>cd rag

Set up LangChain and LangSmith

To visualize metrics, you will need to set up Langchain/Langsmith tracking. It is free and easy.

  1. Create Langchain Account
  2. Get API Key
  3. Create a project

Keep note of your API Key and Project name, you will need those later.

2. Environment Setup

Create a .env file in the project root:

# Guardian API related informationGUARDIAN_API_KEY=<your key here># Claude related informationANTHROPIC_API_KEY=<your key here># Langsmith related informationLANGSMITH_API_KEY=<your key here>LANGSMITH_ENDPOINT="https://api.smith.langchain.com"LANGSMITH_PROJECT=<name of your langsmith project>LANGSMITH_TRACING="true"LANGSMITH_API_KEY="<your-key>"DATABASE_TYPE="<name of database>"HOST="localhost" or shared IPv4 address 
  1. Install PostgreSQL; ensure you write down your login somewhere safe.

  2. build Docker image for PostgreSQL using 'docker-compose build --no-cache"

  3. Run Docker for PostgreSQL using docker-compose up

  4. When done, close Docker using docker-compose down

Running

Mac

In working directory, run chmod +x run.sh && ./run.sh.

Windows

Run run.bat.

Docker

Use the POST endpoint on the Docker container to upload articles to either database; this will upload 10 articles from the Guardian API.

If you're hosting a shared database, run docker compose up --build; this loads up the app with uvicorn at 0.0.0.0:8000.

Streamlit (Docker)

Run docker compose build(if not built previously) & docker compose up in rag/services/streamlit

Streamlit (Locally)

Ensure you've pip installed requirements in your local virtual env.

On Mac, run the script at dev/run.sh

On Windows, run the script at dev/run.bat (by clicking it in file explorer)

Local PostgreSQL

Run uvicorn services.postgres_controller:app --reload --port 8001.

(You will need to activate your venv & pip install -r requirements.txt)

Local Clickhouse

Run uvicorn services.clickhouse_controller:app --reload --port 8000.

(You will need to activate your venv & pip install -r requirements.txt)

Local Cassandra

Run uvicorn services.cassandra.cassandra_controller:app --reload.

Local Grafana

  1. cd into the llm folder.
  2. Run docker-compose up; this will create a Docker container for Grafana on port 3000.
  3. In your browser, open localhost:3000, and login using username admin and password admin.
  4. Add a new data source with the following parameters:
    1. Server address: the IP of your container.
    2. Server port: 8123
    3. Protocol: HTTP
    4. Skip TLS Verify: true
    5. Username: user
    6. Password: default
    7. Default database: guardian
    8. Default table: langchain_metrics
  5. Save and test the data source, and verify it connects successfully.
  6. Import dashboard.json from the llm folder.
  7. You should now have a local Grafana with metrics data; the metrics tab on the Streamlit application should now also work.

DEV

If you want to use the LLM and/or use the POST steps of the langchain pipeline, set the variables USE_POST and USE_LLM to "true" in your .env file.

Endpoints

GET

curl "http://localhost:8000/related-articles?query="

POST

curl -X POST "http://localhost:8000/upload-articles"

License

POSTGRES related information

POSTGRES_DB= POSTGRES_USER= POSTGRES_PASSWORD= POSTGRES_PORT=<port postgres will run on, normally 5432>


#### 3. Run Services
You have the option to choose between ClickHouse, PostgreSQL, and Cassandra as databases to use in benchmarking.
You can run any of them by the following (this example sets up ClickHouse):
```bash
cd services/clickhouse
docker compose up --build
cd services/streamlit
docker compose up --build

Your streamlit application should now be running.

4. Visualize Metrics with Grafana

  • The Streamlit service automatically creates a Grafana dashboard within the app, accessble at https://localhost:8501.

Benchmarking & Querying

  • Use the Streamlit dashboard to run queries against any configured database.
  • Latency and token usage metrics are automatically collected and visualized.
  • You can choose from Single Question, Multi Batch Question (simulating concurrent users), and Multi Batch Multi Questions (simulating multiple users with different questions).

Adding a New Database Provider

  1. Create a new folder under /services (e.g., /services/mydb).

  2. Add your Docker setup (Dockerfile, docker-compose.yml) and controller/DAO code.

  3. Extend the Database Enum in llm/llm_utils/langchain_pipeline.py:

    classDatabase(Enum):
    CLICKHOUSE= ("clickhouse", 8000)
    POSTGRES= ("postgres", 8001)
    CASSANDRA= ("cassandra", 8003)
    MYDB= ("mydb", <your-port>)
  4. Implement your controller to expose the required endpoints (/related-articles, /upload-articles).

  5. Update the dashboard (if needed) to include your new provider in the selection.


Customization & Extensibility

  • The codebase is modular—add new database backends by following the existing service structure.
  • The Streamlit dashboard auto-detects available databases from the Database Enum.
  • Metrics collection is built-in; extend or modify as needed for your research.

License

MIT


Happy benchmarking!


About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages