Repository files navigation

SAGE: Personal RAG-enabled Information Retrieval Utility

SAGE is a complete RAG (Retrieval-Augmented Generation) management and interrogation system built with FastAPI and ChromaDB. It provides a powerful API for document management, vector storage, and LLM-powered chat functionality with support for multiple LLM providers.

Features

  • Document Management: Upload, process, and manage documents in vector collections
  • Document Tagging: Add tags to documents for better organization and filtering
  • Vector Database: ChromaDB integration for efficient semantic search
  • LLM Integration: Support for multiple LLM providers:
    • Ollama with local or remote models (including nomic-embed-text)
    • Anthropic with Claude models (including haiku)
    • OpenAI with GPT models (including 4o)
  • Conversational UI: Interactive chat interface with conversation history and context awareness
  • WebSocket Chat: Real-time, streaming chat interface with RAG capabilities
  • REST API: Comprehensive REST API for all operations
  • Document Processing: Support for various document types (PDF, DOCX, CSV, text)

Installation

Prerequisites

  • Python 3.8+
  • [Optional] Virtual environment (recommended)

Setup

  1. Clone the repository:

    jj git clone --colocate https://github.com/grepory/sage.git
    cd sage
  2. Create and activate a virtual environment (optional but recommended):

    python -m venv venv
    source venv/bin/activate # On Windows: venv\Scripts\activate
  3. Install dependencies:

    pip install -r requirements.txt
  4. Copy the .env.example file to .env in the project root and update it with your LLM provider configuration:

    cp .env.example .env
    # Then edit the .env file with your preferred text editor

    The .env.example file contains all the necessary environment variables with default values and explanatory comments. You only need to configure the providers you plan to use.

Running the Application

Start the server with:

python run_app.py

This will:

  • Start the server on the configured port (default 8000)
  • Display clear URLs for accessing the application
  • Automatically open your default web browser to the frontend

Alternatively, you can start the server manually with:

uvicorn app.main:app --reload --host 0.0.0.0 --port 8000

The API will be available at http://localhost:8000, and the API documentation at http://localhost:8000/docs.

Configuration

Sage can be configured through environment variables in the .env file:

Server Configuration

  • PORT=8000 - Port for the FastAPI server
  • ROOT_PATH=/sage - Root path prefix for reverse proxy setup (e.g., for hosting at http://example.com/sage/)

LLM Provider Configuration

Configure your preferred LLM providers in the .env file:

API Documentation

Collections API

Collections are used to organize documents in the vector database.

Create a Collection

POST /api/v1/collections/Content-Type: application/json
{
"name": "my_collection",
"description": "My first collection"
}

List Collections

GET /api/v1/collections/

Get Collection

GET /api/v1/collections/{collection_name}

Delete Collection

DELETE /api/v1/collections/{collection_name}

Documents API

Documents are the core content stored in the vector database.

Upload a Document

POST /api/v1/documents/uploadContent-Type: multipart/form-datafile: [file]collection_name: my_collectiontags: personal,house,importantadditional_metadata: {"source": "website", "author": "John Doe"}

The tags parameter is optional and should be a comma-separated list of tags to associate with the document.

Add Text Directly

POST /api/v1/documents/textContent-Type: application/json
{
"text": "This is some text to add to the vector database.",
"collection_name": "my_collection",
"tags": ["personal", "notes", "important"],
"metadata": {"source": "direct_input", "author": "Jane Smith"}
}

The tags parameter is optional and should be an array of strings representing tags to associate with the document.

Get Document

GET /api/v1/documents/{collection_name}/{document_id}

Delete Document

DELETE /api/v1/documents/{collection_name}/{document_id}

Query Documents

POST /api/v1/documents/queryContent-Type: application/json
{
"collection_name": "my_collection",
"query_text": "What is retrieval-augmented generation?",
"n_results": 5,
"where": {
"tags": {"$in": ["personal", "important"]}
}
}

The where parameter is optional and can be used to filter documents by metadata fields, including tags. In the example above, the query will only return documents that have either the "personal" or "important" tag. You can use other operators like $eq for exact match or $contains for substring match.

Chat API

The Chat API provides both REST and WebSocket interfaces for interacting with the LLM.

Model Selection

You can specify which LLM provider and model to use by setting the model parameter in your requests. The format is:

provider:model

For example:

  • ollama:llama2 - Use Ollama with the llama2 model
  • anthropic:claude-3-haiku-20240307 - Use Anthropic with the Claude 3 Haiku model
  • openai:gpt-4o - Use OpenAI with the GPT-4o model

If you don't specify a provider, the default provider from your configuration will be used. If you don't specify a model, the default model for the selected provider will be used.

REST Chat Endpoint

POST /api/v1/chat/Content-Type: application/json
{
"collection_name": "my_collection",
"query": "What is retrieval-augmented generation?",
"history": [
{"role": "user", "content": "Hello"},
{"role": "assistant", "content": "Hi there! How can I help you?"}
],
"model": "llama2"
}

WebSocket Chat

Connect to the WebSocket endpoint at /api/v1/chat/ws and send JSON messages with the following format:

{
"type": "chat",
"collection_name": "my_collection",
"query": "What is retrieval-augmented generation?",
"history": [
{"role": "user", "content": "Hello"},
{"role": "assistant", "content": "Hi there! How can I help you?"}
],
"model": "llama2"
}

The server will respond with messages in the following formats:

  1. Start message:
{
"type": "start",
"content": "Generating response..."
}
  1. Token messages (streaming):
{
"type": "token",
"content": "piece of text"
}
  1. Complete message:
{
"type": "complete",
"content": {
"answer": "The complete answer...",
"sources": [
{
"id": "document_id",
"text": "Snippet of the source document...",
"metadata": {"source": "document.pdf", "page": 1}
}
],
"history": [
{"role": "user", "content": "What is RAG?"},
{"role": "assistant", "content": "The complete answer..."}
]
}
}
  1. Error message:
{
"type": "error",
"content": "Error message"
}

Web Interface

Sage provides a responsive web interface that works seamlessly on desktop and mobile devices. The interface includes:

Document Upload

  • Upload files using the upload button
  • Add tags to organize your documents
  • View upload progress and status

Conversational Chat Interface

The chat interface allows you to have natural conversations with the AI about your documents:

  • Conversation History: The system maintains the full conversation history, allowing for contextual follow-up questions
  • Conversation Sidebar: Browse and manage your chat history with a collapsible sidebar
  • Enter Key Support: Press Enter to send your message (use Shift+Enter for new lines)
  • Visual Distinction: User and assistant messages are visually distinct for easy reading
  • Source Citations: The system displays the source documents used to generate responses
  • Mobile Responsive: Optimized interface for mobile devices with touch-friendly controls
  • Tag-based Filtering: Filter conversations and documents by tags with advanced search capabilities

To use the conversational interface:

  1. Select a a set of tags from the dropdown (or choose to search all tags)
  2. (Optional) Select a specific LLM model and provider
  3. Type your question in the message box
  4. Press Enter or click "Send Message"
  5. View the AI's response and the sources used
  6. Continue the conversation with follow-up questions

The system will maintain context between questions, allowing for a more natural conversation flow. On mobile devices, the conversation sidebar can be toggled using the sidebar button in the top navigation.

Examples

Python Client Example

Here's a simple Python client for interacting with the API:

importrequestsimportwebsocketsimportjsonimportasyncioBASE_URL="http://localhost:8000/api/v1"# Create a collectiondefcreate_collection(name, description=None):
response=requests.post(
f"{BASE_URL}/collections/",
json={"name": name, "description": description}
)
returnresponse.json()
# Upload a documentdefupload_document(file_path, collection_name, tags=None, metadata=None):
files= {"file": open(file_path, "rb")}
data= {"collection_name": collection_name}
iftags:
data["tags"] =",".join(tags)
ifmetadata:
data["additional_metadata"] =json.dumps(metadata)
response=requests.post(
f"{BASE_URL}/documents/upload",
files=files,
data=data
)
returnresponse.json()
# Query documentsdefquery_documents(collection_name, query_text, n_results=5, tags=None):
query_data= {
"collection_name": collection_name,
"query_text": query_text,
"n_results": n_results
}
# Add tag filtering if specifiediftags:
query_data["where"] = {"tags": {"$in": tags}}
response=requests.post(
f"{BASE_URL}/documents/query",
json=query_data
)
returnresponse.json()
# WebSocket chat clientasyncdefchat_websocket(query, tags=None, history=None, model=None, include_untagged=True):
uri=f"ws://localhost:8000/api/v1/chat/ws"asyncwithwebsockets.connect(uri) aswebsocket:
# Send chat messagemessage= {
"type": "chat",
"query": query,
"history": historyor [],
"include_untagged": include_untagged
}
iftags:
message["tags"] =tagsifmodel:
message["model"] =modelawaitwebsocket.send(json.dumps(message))
# Process responsesfull_response=""whileTrue:
response=awaitwebsocket.recv()
data=json.loads(response)
ifdata["type"] =="token":
# Print token as it arrivesprint(data["content"], end="", flush=True)
full_response+=data["content"]
elifdata["type"] =="complete":
# Chat completeprint("\n\nSources:")
fori, sourceinenumerate(data["content"]["sources"]):
print(f"{i+1}. {source['text']} (ID: {source['id']})")
returndata["content"]
elifdata["type"] =="error":
# Error occurredprint(f"\nError: {data['content']}")
returnNone# Example usageif__name__=="__main__":
# Create a collectioncreate_collection("my_docs", "My documents collection")
# Upload a document with tagsupload_document(
"path/to/document.pdf", "my_docs", tags=["house", "important", "insurance"],
metadata={"source": "local", "author": "Insurance Company"}
)
# Upload another document with different tagsupload_document(
"path/to/another_document.pdf", "my_docs", tags=["personal", "important", "legal"],
metadata={"source": "local", "author": "Law Firm"}
)
# Query documents without tag filteringresults=query_documents("my_docs", "What is RAG?")
# Query documents with tag filtering (only "house" tagged documents)house_results=query_documents("my_docs", "What is RAG?", tags=["house"])
# Query documents with multiple tag filtering (documents tagged as either "important" or "personal")important_results=query_documents("my_docs", "What is RAG?", tags=["important", "personal"])
# Chat with WebSocket - Single queryasyncio.run(chat_websocket("Explain RAG in simple terms."))
# Chat with WebSocket - Query with tag filteringasyncio.run(chat_websocket(
"What can you tell me about my house documents?", tags=["house", "important"]
))
# Chat with WebSocket - Conversation with historyasyncdefconversation_example():
# First queryprint("\n--- Starting conversation ---")
response=awaitchat_websocket("Explain RAG in simple terms.")
# Get the updated history from the responseconversation_history=response["history"]
# Follow-up query using the conversation history with tag filteringprint("\n--- Follow-up question with tag filtering ---")
response=awaitchat_websocket(
"What are the main benefits compared to traditional approaches?",
tags=["important", "personal"],
history=conversation_history
)
# Continue the conversation with another follow-upconversation_history=response["history"]
print("\n--- Another follow-up question ---")
awaitchat_websocket(
"Can you give me a simple code example?",
history=conversation_history
)
# Run the conversation exampleasyncio.run(conversation_example())

License

MIT

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

About

RAGU - RAG Utility for interrogating documents

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

SAGE: Personal RAG-enabled Information Retrieval Utility

SAGE is a complete RAG (Retrieval-Augmented Generation) management and interrogation system built with FastAPI and ChromaDB. It provides a powerful API for document management, vector storage, and LLM-powered chat functionality with support for multiple LLM providers.

Features

  • Document Management: Upload, process, and manage documents in vector collections
  • Document Tagging: Add tags to documents for better organization and filtering
  • Vector Database: ChromaDB integration for efficient semantic search
  • LLM Integration: Support for multiple LLM providers:
    • Ollama with local or remote models (including nomic-embed-text)
    • Anthropic with Claude models (including haiku)
    • OpenAI with GPT models (including 4o)
  • Conversational UI: Interactive chat interface with conversation history and context awareness
  • WebSocket Chat: Real-time, streaming chat interface with RAG capabilities
  • REST API: Comprehensive REST API for all operations
  • Document Processing: Support for various document types (PDF, DOCX, CSV, text)

Installation

Prerequisites

  • Python 3.8+
  • [Optional] Virtual environment (recommended)

Setup

  1. Clone the repository:

    jj git clone --colocate https://github.com/grepory/sage.git
    cd sage
  2. Create and activate a virtual environment (optional but recommended):

    python -m venv venv
    source venv/bin/activate # On Windows: venv\Scripts\activate
  3. Install dependencies:

    pip install -r requirements.txt
  4. Copy the .env.example file to .env in the project root and update it with your LLM provider configuration:

    cp .env.example .env
    # Then edit the .env file with your preferred text editor

    The .env.example file contains all the necessary environment variables with default values and explanatory comments. You only need to configure the providers you plan to use.

Running the Application

Start the server with:

python run_app.py

This will:

  • Start the server on the configured port (default 8000)
  • Display clear URLs for accessing the application
  • Automatically open your default web browser to the frontend

Alternatively, you can start the server manually with:

uvicorn app.main:app --reload --host 0.0.0.0 --port 8000

The API will be available at http://localhost:8000, and the API documentation at http://localhost:8000/docs.

Configuration

Sage can be configured through environment variables in the .env file:

Server Configuration

  • PORT=8000 - Port for the FastAPI server
  • ROOT_PATH=/sage - Root path prefix for reverse proxy setup (e.g., for hosting at http://example.com/sage/)

LLM Provider Configuration

Configure your preferred LLM providers in the .env file:

API Documentation

Collections API

Collections are used to organize documents in the vector database.

Create a Collection

POST /api/v1/collections/Content-Type: application/json
{
"name": "my_collection",
"description": "My first collection"
}

List Collections

GET /api/v1/collections/

Get Collection

GET /api/v1/collections/{collection_name}

Delete Collection

DELETE /api/v1/collections/{collection_name}

Documents API

Documents are the core content stored in the vector database.

Upload a Document

POST /api/v1/documents/uploadContent-Type: multipart/form-datafile: [file]collection_name: my_collectiontags: personal,house,importantadditional_metadata: {"source": "website", "author": "John Doe"}

The tags parameter is optional and should be a comma-separated list of tags to associate with the document.

Add Text Directly

POST /api/v1/documents/textContent-Type: application/json
{
"text": "This is some text to add to the vector database.",
"collection_name": "my_collection",
"tags": ["personal", "notes", "important"],
"metadata": {"source": "direct_input", "author": "Jane Smith"}
}

The tags parameter is optional and should be an array of strings representing tags to associate with the document.

Get Document

GET /api/v1/documents/{collection_name}/{document_id}

Delete Document

DELETE /api/v1/documents/{collection_name}/{document_id}

Query Documents

POST /api/v1/documents/queryContent-Type: application/json
{
"collection_name": "my_collection",
"query_text": "What is retrieval-augmented generation?",
"n_results": 5,
"where": {
"tags": {"$in": ["personal", "important"]}
}
}

The where parameter is optional and can be used to filter documents by metadata fields, including tags. In the example above, the query will only return documents that have either the "personal" or "important" tag. You can use other operators like $eq for exact match or $contains for substring match.

Chat API

The Chat API provides both REST and WebSocket interfaces for interacting with the LLM.

Model Selection

You can specify which LLM provider and model to use by setting the model parameter in your requests. The format is:

provider:model

For example:

  • ollama:llama2 - Use Ollama with the llama2 model
  • anthropic:claude-3-haiku-20240307 - Use Anthropic with the Claude 3 Haiku model
  • openai:gpt-4o - Use OpenAI with the GPT-4o model

If you don't specify a provider, the default provider from your configuration will be used. If you don't specify a model, the default model for the selected provider will be used.

REST Chat Endpoint

POST /api/v1/chat/Content-Type: application/json
{
"collection_name": "my_collection",
"query": "What is retrieval-augmented generation?",
"history": [
{"role": "user", "content": "Hello"},
{"role": "assistant", "content": "Hi there! How can I help you?"}
],
"model": "llama2"
}

WebSocket Chat

Connect to the WebSocket endpoint at /api/v1/chat/ws and send JSON messages with the following format:

{
"type": "chat",
"collection_name": "my_collection",
"query": "What is retrieval-augmented generation?",
"history": [
{"role": "user", "content": "Hello"},
{"role": "assistant", "content": "Hi there! How can I help you?"}
],
"model": "llama2"
}

The server will respond with messages in the following formats:

  1. Start message:
{
"type": "start",
"content": "Generating response..."
}
  1. Token messages (streaming):
{
"type": "token",
"content": "piece of text"
}
  1. Complete message:
{
"type": "complete",
"content": {
"answer": "The complete answer...",
"sources": [
{
"id": "document_id",
"text": "Snippet of the source document...",
"metadata": {"source": "document.pdf", "page": 1}
}
],
"history": [
{"role": "user", "content": "What is RAG?"},
{"role": "assistant", "content": "The complete answer..."}
]
}
}
  1. Error message:
{
"type": "error",
"content": "Error message"
}

Web Interface

Sage provides a responsive web interface that works seamlessly on desktop and mobile devices. The interface includes:

Document Upload

  • Upload files using the upload button
  • Add tags to organize your documents
  • View upload progress and status

Conversational Chat Interface

The chat interface allows you to have natural conversations with the AI about your documents:

  • Conversation History: The system maintains the full conversation history, allowing for contextual follow-up questions
  • Conversation Sidebar: Browse and manage your chat history with a collapsible sidebar
  • Enter Key Support: Press Enter to send your message (use Shift+Enter for new lines)
  • Visual Distinction: User and assistant messages are visually distinct for easy reading
  • Source Citations: The system displays the source documents used to generate responses
  • Mobile Responsive: Optimized interface for mobile devices with touch-friendly controls
  • Tag-based Filtering: Filter conversations and documents by tags with advanced search capabilities

To use the conversational interface:

  1. Select a a set of tags from the dropdown (or choose to search all tags)
  2. (Optional) Select a specific LLM model and provider
  3. Type your question in the message box
  4. Press Enter or click "Send Message"
  5. View the AI's response and the sources used
  6. Continue the conversation with follow-up questions

The system will maintain context between questions, allowing for a more natural conversation flow. On mobile devices, the conversation sidebar can be toggled using the sidebar button in the top navigation.

Examples

Python Client Example

Here's a simple Python client for interacting with the API:

importrequestsimportwebsocketsimportjsonimportasyncioBASE_URL="http://localhost:8000/api/v1"# Create a collectiondefcreate_collection(name, description=None):
response=requests.post(
f"{BASE_URL}/collections/",
json={"name": name, "description": description}
)
returnresponse.json()
# Upload a documentdefupload_document(file_path, collection_name, tags=None, metadata=None):
files= {"file": open(file_path, "rb")}
data= {"collection_name": collection_name}
iftags:
data["tags"] =",".join(tags)
ifmetadata:
data["additional_metadata"] =json.dumps(metadata)
response=requests.post(
f"{BASE_URL}/documents/upload",
files=files,
data=data
)
returnresponse.json()
# Query documentsdefquery_documents(collection_name, query_text, n_results=5, tags=None):
query_data= {
"collection_name": collection_name,
"query_text": query_text,
"n_results": n_results
}
# Add tag filtering if specifiediftags:
query_data["where"] = {"tags": {"$in": tags}}
response=requests.post(
f"{BASE_URL}/documents/query",
json=query_data
)
returnresponse.json()
# WebSocket chat clientasyncdefchat_websocket(query, tags=None, history=None, model=None, include_untagged=True):
uri=f"ws://localhost:8000/api/v1/chat/ws"asyncwithwebsockets.connect(uri) aswebsocket:
# Send chat messagemessage= {
"type": "chat",
"query": query,
"history": historyor [],
"include_untagged": include_untagged
}
iftags:
message["tags"] =tagsifmodel:
message["model"] =modelawaitwebsocket.send(json.dumps(message))
# Process responsesfull_response=""whileTrue:
response=awaitwebsocket.recv()
data=json.loads(response)
ifdata["type"] =="token":
# Print token as it arrivesprint(data["content"], end="", flush=True)
full_response+=data["content"]
elifdata["type"] =="complete":
# Chat completeprint("\n\nSources:")
fori, sourceinenumerate(data["content"]["sources"]):
print(f"{i+1}. {source['text']} (ID: {source['id']})")
returndata["content"]
elifdata["type"] =="error":
# Error occurredprint(f"\nError: {data['content']}")
returnNone# Example usageif__name__=="__main__":
# Create a collectioncreate_collection("my_docs", "My documents collection")
# Upload a document with tagsupload_document(
"path/to/document.pdf", "my_docs", tags=["house", "important", "insurance"],
metadata={"source": "local", "author": "Insurance Company"}
)
# Upload another document with different tagsupload_document(
"path/to/another_document.pdf", "my_docs", tags=["personal", "important", "legal"],
metadata={"source": "local", "author": "Law Firm"}
)
# Query documents without tag filteringresults=query_documents("my_docs", "What is RAG?")
# Query documents with tag filtering (only "house" tagged documents)house_results=query_documents("my_docs", "What is RAG?", tags=["house"])
# Query documents with multiple tag filtering (documents tagged as either "important" or "personal")important_results=query_documents("my_docs", "What is RAG?", tags=["important", "personal"])
# Chat with WebSocket - Single queryasyncio.run(chat_websocket("Explain RAG in simple terms."))
# Chat with WebSocket - Query with tag filteringasyncio.run(chat_websocket(
"What can you tell me about my house documents?", tags=["house", "important"]
))
# Chat with WebSocket - Conversation with historyasyncdefconversation_example():
# First queryprint("\n--- Starting conversation ---")
response=awaitchat_websocket("Explain RAG in simple terms.")
# Get the updated history from the responseconversation_history=response["history"]
# Follow-up query using the conversation history with tag filteringprint("\n--- Follow-up question with tag filtering ---")
response=awaitchat_websocket(
"What are the main benefits compared to traditional approaches?",
tags=["important", "personal"],
history=conversation_history
)
# Continue the conversation with another follow-upconversation_history=response["history"]
print("\n--- Another follow-up question ---")
awaitchat_websocket(
"Can you give me a simple code example?",
history=conversation_history
)
# Run the conversation exampleasyncio.run(conversation_example())

License

MIT

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

About

RAGU - RAG Utility for interrogating documents

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SAGE: Personal RAG-enabled Information Retrieval Utility

SAGE is a complete RAG (Retrieval-Augmented Generation) management and interrogation system built with FastAPI and ChromaDB. It provides a powerful API for document management, vector storage, and LLM-powered chat functionality with support for multiple LLM providers.

Features

  • Document Management: Upload, process, and manage documents in vector collections
  • Document Tagging: Add tags to documents for better organization and filtering
  • Vector Database: ChromaDB integration for efficient semantic search
  • LLM Integration: Support for multiple LLM providers:
    • Ollama with local or remote models (including nomic-embed-text)
    • Anthropic with Claude models (including haiku)
    • OpenAI with GPT models (including 4o)
  • Conversational UI: Interactive chat interface with conversation history and context awareness
  • WebSocket Chat: Real-time, streaming chat interface with RAG capabilities
  • REST API: Comprehensive REST API for all operations
  • Document Processing: Support for various document types (PDF, DOCX, CSV, text)

Installation

Prerequisites

  • Python 3.8+
  • [Optional] Virtual environment (recommended)

Setup

  1. Clone the repository:

    jj git clone --colocate https://github.com/grepory/sage.git
    cd sage
  2. Create and activate a virtual environment (optional but recommended):

    python -m venv venv
    source venv/bin/activate # On Windows: venv\Scripts\activate
  3. Install dependencies:

    pip install -r requirements.txt
  4. Copy the .env.example file to .env in the project root and update it with your LLM provider configuration:

    cp .env.example .env
    # Then edit the .env file with your preferred text editor

    The .env.example file contains all the necessary environment variables with default values and explanatory comments. You only need to configure the providers you plan to use.

Running the Application

Start the server with:

python run_app.py

This will:

  • Start the server on the configured port (default 8000)
  • Display clear URLs for accessing the application
  • Automatically open your default web browser to the frontend

Alternatively, you can start the server manually with:

uvicorn app.main:app --reload --host 0.0.0.0 --port 8000

The API will be available at http://localhost:8000, and the API documentation at http://localhost:8000/docs.

Configuration

Sage can be configured through environment variables in the .env file:

Server Configuration

  • PORT=8000 - Port for the FastAPI server
  • ROOT_PATH=/sage - Root path prefix for reverse proxy setup (e.g., for hosting at http://example.com/sage/)

LLM Provider Configuration

Configure your preferred LLM providers in the .env file:

API Documentation

Collections API

Collections are used to organize documents in the vector database.

Create a Collection

POST /api/v1/collections/Content-Type: application/json
{
"name": "my_collection",
"description": "My first collection"
}

List Collections

GET /api/v1/collections/

Get Collection

GET /api/v1/collections/{collection_name}

Delete Collection

DELETE /api/v1/collections/{collection_name}

Documents API

Documents are the core content stored in the vector database.

Upload a Document

POST /api/v1/documents/uploadContent-Type: multipart/form-datafile: [file]collection_name: my_collectiontags: personal,house,importantadditional_metadata: {"source": "website", "author": "John Doe"}

The tags parameter is optional and should be a comma-separated list of tags to associate with the document.

Add Text Directly

POST /api/v1/documents/textContent-Type: application/json
{
"text": "This is some text to add to the vector database.",
"collection_name": "my_collection",
"tags": ["personal", "notes", "important"],
"metadata": {"source": "direct_input", "author": "Jane Smith"}
}

The tags parameter is optional and should be an array of strings representing tags to associate with the document.

Get Document

GET /api/v1/documents/{collection_name}/{document_id}

Delete Document

DELETE /api/v1/documents/{collection_name}/{document_id}

Query Documents

POST /api/v1/documents/queryContent-Type: application/json
{
"collection_name": "my_collection",
"query_text": "What is retrieval-augmented generation?",
"n_results": 5,
"where": {
"tags": {"$in": ["personal", "important"]}
}
}

The where parameter is optional and can be used to filter documents by metadata fields, including tags. In the example above, the query will only return documents that have either the "personal" or "important" tag. You can use other operators like $eq for exact match or $contains for substring match.

Chat API

The Chat API provides both REST and WebSocket interfaces for interacting with the LLM.

Model Selection

You can specify which LLM provider and model to use by setting the model parameter in your requests. The format is:

provider:model

For example:

  • ollama:llama2 - Use Ollama with the llama2 model
  • anthropic:claude-3-haiku-20240307 - Use Anthropic with the Claude 3 Haiku model
  • openai:gpt-4o - Use OpenAI with the GPT-4o model

If you don't specify a provider, the default provider from your configuration will be used. If you don't specify a model, the default model for the selected provider will be used.

REST Chat Endpoint

POST /api/v1/chat/Content-Type: application/json
{
"collection_name": "my_collection",
"query": "What is retrieval-augmented generation?",
"history": [
{"role": "user", "content": "Hello"},
{"role": "assistant", "content": "Hi there! How can I help you?"}
],
"model": "llama2"
}

WebSocket Chat

Connect to the WebSocket endpoint at /api/v1/chat/ws and send JSON messages with the following format:

{
"type": "chat",
"collection_name": "my_collection",
"query": "What is retrieval-augmented generation?",
"history": [
{"role": "user", "content": "Hello"},
{"role": "assistant", "content": "Hi there! How can I help you?"}
],
"model": "llama2"
}

The server will respond with messages in the following formats:

  1. Start message:
{
"type": "start",
"content": "Generating response..."
}
  1. Token messages (streaming):
{
"type": "token",
"content": "piece of text"
}
  1. Complete message:
{
"type": "complete",
"content": {
"answer": "The complete answer...",
"sources": [
{
"id": "document_id",
"text": "Snippet of the source document...",
"metadata": {"source": "document.pdf", "page": 1}
}
],
"history": [
{"role": "user", "content": "What is RAG?"},
{"role": "assistant", "content": "The complete answer..."}
]
}
}
  1. Error message:
{
"type": "error",
"content": "Error message"
}

Web Interface

Sage provides a responsive web interface that works seamlessly on desktop and mobile devices. The interface includes:

Document Upload

  • Upload files using the upload button
  • Add tags to organize your documents
  • View upload progress and status

Conversational Chat Interface

The chat interface allows you to have natural conversations with the AI about your documents:

  • Conversation History: The system maintains the full conversation history, allowing for contextual follow-up questions
  • Conversation Sidebar: Browse and manage your chat history with a collapsible sidebar
  • Enter Key Support: Press Enter to send your message (use Shift+Enter for new lines)
  • Visual Distinction: User and assistant messages are visually distinct for easy reading
  • Source Citations: The system displays the source documents used to generate responses
  • Mobile Responsive: Optimized interface for mobile devices with touch-friendly controls
  • Tag-based Filtering: Filter conversations and documents by tags with advanced search capabilities

To use the conversational interface:

  1. Select a a set of tags from the dropdown (or choose to search all tags)
  2. (Optional) Select a specific LLM model and provider
  3. Type your question in the message box
  4. Press Enter or click "Send Message"
  5. View the AI's response and the sources used
  6. Continue the conversation with follow-up questions

The system will maintain context between questions, allowing for a more natural conversation flow. On mobile devices, the conversation sidebar can be toggled using the sidebar button in the top navigation.

Examples

Python Client Example

Here's a simple Python client for interacting with the API:

importrequestsimportwebsocketsimportjsonimportasyncioBASE_URL="http://localhost:8000/api/v1"# Create a collectiondefcreate_collection(name, description=None):
response=requests.post(
f"{BASE_URL}/collections/",
json={"name": name, "description": description}
)
returnresponse.json()
# Upload a documentdefupload_document(file_path, collection_name, tags=None, metadata=None):
files= {"file": open(file_path, "rb")}
data= {"collection_name": collection_name}
iftags:
data["tags"] =",".join(tags)
ifmetadata:
data["additional_metadata"] =json.dumps(metadata)
response=requests.post(
f"{BASE_URL}/documents/upload",
files=files,
data=data
)
returnresponse.json()
# Query documentsdefquery_documents(collection_name, query_text, n_results=5, tags=None):
query_data= {
"collection_name": collection_name,
"query_text": query_text,
"n_results": n_results
}
# Add tag filtering if specifiediftags:
query_data["where"] = {"tags": {"$in": tags}}
response=requests.post(
f"{BASE_URL}/documents/query",
json=query_data
)
returnresponse.json()
# WebSocket chat clientasyncdefchat_websocket(query, tags=None, history=None, model=None, include_untagged=True):
uri=f"ws://localhost:8000/api/v1/chat/ws"asyncwithwebsockets.connect(uri) aswebsocket:
# Send chat messagemessage= {
"type": "chat",
"query": query,
"history": historyor [],
"include_untagged": include_untagged
}
iftags:
message["tags"] =tagsifmodel:
message["model"] =modelawaitwebsocket.send(json.dumps(message))
# Process responsesfull_response=""whileTrue:
response=awaitwebsocket.recv()
data=json.loads(response)
ifdata["type"] =="token":
# Print token as it arrivesprint(data["content"], end="", flush=True)
full_response+=data["content"]
elifdata["type"] =="complete":
# Chat completeprint("\n\nSources:")
fori, sourceinenumerate(data["content"]["sources"]):
print(f"{i+1}. {source['text']} (ID: {source['id']})")
returndata["content"]
elifdata["type"] =="error":
# Error occurredprint(f"\nError: {data['content']}")
returnNone# Example usageif__name__=="__main__":
# Create a collectioncreate_collection("my_docs", "My documents collection")
# Upload a document with tagsupload_document(
"path/to/document.pdf", "my_docs", tags=["house", "important", "insurance"],
metadata={"source": "local", "author": "Insurance Company"}
)
# Upload another document with different tagsupload_document(
"path/to/another_document.pdf", "my_docs", tags=["personal", "important", "legal"],
metadata={"source": "local", "author": "Law Firm"}
)
# Query documents without tag filteringresults=query_documents("my_docs", "What is RAG?")
# Query documents with tag filtering (only "house" tagged documents)house_results=query_documents("my_docs", "What is RAG?", tags=["house"])
# Query documents with multiple tag filtering (documents tagged as either "important" or "personal")important_results=query_documents("my_docs", "What is RAG?", tags=["important", "personal"])
# Chat with WebSocket - Single queryasyncio.run(chat_websocket("Explain RAG in simple terms."))
# Chat with WebSocket - Query with tag filteringasyncio.run(chat_websocket(
"What can you tell me about my house documents?", tags=["house", "important"]
))
# Chat with WebSocket - Conversation with historyasyncdefconversation_example():
# First queryprint("\n--- Starting conversation ---")
response=awaitchat_websocket("Explain RAG in simple terms.")
# Get the updated history from the responseconversation_history=response["history"]
# Follow-up query using the conversation history with tag filteringprint("\n--- Follow-up question with tag filtering ---")
response=awaitchat_websocket(
"What are the main benefits compared to traditional approaches?",
tags=["important", "personal"],
history=conversation_history
)
# Continue the conversation with another follow-upconversation_history=response["history"]
print("\n--- Another follow-up question ---")
awaitchat_websocket(
"Can you give me a simple code example?",
history=conversation_history
)
# Run the conversation exampleasyncio.run(conversation_example())

License

MIT

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

About

RAGU - RAG Utility for interrogating documents

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SAGE: Personal RAG-enabled Information Retrieval Utility

SAGE is a complete RAG (Retrieval-Augmented Generation) management and interrogation system built with FastAPI and ChromaDB. It provides a powerful API for document management, vector storage, and LLM-powered chat functionality with support for multiple LLM providers.

Features

  • Document Management: Upload, process, and manage documents in vector collections
  • Document Tagging: Add tags to documents for better organization and filtering
  • Vector Database: ChromaDB integration for efficient semantic search
  • LLM Integration: Support for multiple LLM providers:
    • Ollama with local or remote models (including nomic-embed-text)
    • Anthropic with Claude models (including haiku)
    • OpenAI with GPT models (including 4o)
  • Conversational UI: Interactive chat interface with conversation history and context awareness
  • WebSocket Chat: Real-time, streaming chat interface with RAG capabilities
  • REST API: Comprehensive REST API for all operations
  • Document Processing: Support for various document types (PDF, DOCX, CSV, text)

Installation

Prerequisites

  • Python 3.8+
  • [Optional] Virtual environment (recommended)

Setup

  1. Clone the repository:

    jj git clone --colocate https://github.com/grepory/sage.git
    cd sage
  2. Create and activate a virtual environment (optional but recommended):

    python -m venv venv
    source venv/bin/activate # On Windows: venv\Scripts\activate
  3. Install dependencies:

    pip install -r requirements.txt
  4. Copy the .env.example file to .env in the project root and update it with your LLM provider configuration:

    cp .env.example .env
    # Then edit the .env file with your preferred text editor

    The .env.example file contains all the necessary environment variables with default values and explanatory comments. You only need to configure the providers you plan to use.

Running the Application

Start the server with:

python run_app.py

This will:

  • Start the server on the configured port (default 8000)
  • Display clear URLs for accessing the application
  • Automatically open your default web browser to the frontend

Alternatively, you can start the server manually with:

uvicorn app.main:app --reload --host 0.0.0.0 --port 8000

The API will be available at http://localhost:8000, and the API documentation at http://localhost:8000/docs.

Configuration

Sage can be configured through environment variables in the .env file:

Server Configuration

  • PORT=8000 - Port for the FastAPI server
  • ROOT_PATH=/sage - Root path prefix for reverse proxy setup (e.g., for hosting at http://example.com/sage/)

LLM Provider Configuration

Configure your preferred LLM providers in the .env file:

API Documentation

Collections API

Collections are used to organize documents in the vector database.

Create a Collection

POST /api/v1/collections/Content-Type: application/json
{
"name": "my_collection",
"description": "My first collection"
}

List Collections

GET /api/v1/collections/

Get Collection

GET /api/v1/collections/{collection_name}

Delete Collection

DELETE /api/v1/collections/{collection_name}

Documents API

Documents are the core content stored in the vector database.

Upload a Document

POST /api/v1/documents/uploadContent-Type: multipart/form-datafile: [file]collection_name: my_collectiontags: personal,house,importantadditional_metadata: {"source": "website", "author": "John Doe"}

The tags parameter is optional and should be a comma-separated list of tags to associate with the document.

Add Text Directly

POST /api/v1/documents/textContent-Type: application/json
{
"text": "This is some text to add to the vector database.",
"collection_name": "my_collection",
"tags": ["personal", "notes", "important"],
"metadata": {"source": "direct_input", "author": "Jane Smith"}
}

The tags parameter is optional and should be an array of strings representing tags to associate with the document.

Get Document

GET /api/v1/documents/{collection_name}/{document_id}

Delete Document

DELETE /api/v1/documents/{collection_name}/{document_id}

Query Documents

POST /api/v1/documents/queryContent-Type: application/json
{
"collection_name": "my_collection",
"query_text": "What is retrieval-augmented generation?",
"n_results": 5,
"where": {
"tags": {"$in": ["personal", "important"]}
}
}

The where parameter is optional and can be used to filter documents by metadata fields, including tags. In the example above, the query will only return documents that have either the "personal" or "important" tag. You can use other operators like $eq for exact match or $contains for substring match.

Chat API

The Chat API provides both REST and WebSocket interfaces for interacting with the LLM.

Model Selection

You can specify which LLM provider and model to use by setting the model parameter in your requests. The format is:

provider:model

For example:

  • ollama:llama2 - Use Ollama with the llama2 model
  • anthropic:claude-3-haiku-20240307 - Use Anthropic with the Claude 3 Haiku model
  • openai:gpt-4o - Use OpenAI with the GPT-4o model

If you don't specify a provider, the default provider from your configuration will be used. If you don't specify a model, the default model for the selected provider will be used.

REST Chat Endpoint

POST /api/v1/chat/Content-Type: application/json
{
"collection_name": "my_collection",
"query": "What is retrieval-augmented generation?",
"history": [
{"role": "user", "content": "Hello"},
{"role": "assistant", "content": "Hi there! How can I help you?"}
],
"model": "llama2"
}

WebSocket Chat

Connect to the WebSocket endpoint at /api/v1/chat/ws and send JSON messages with the following format:

{
"type": "chat",
"collection_name": "my_collection",
"query": "What is retrieval-augmented generation?",
"history": [
{"role": "user", "content": "Hello"},
{"role": "assistant", "content": "Hi there! How can I help you?"}
],
"model": "llama2"
}

The server will respond with messages in the following formats:

  1. Start message:
{
"type": "start",
"content": "Generating response..."
}
  1. Token messages (streaming):
{
"type": "token",
"content": "piece of text"
}
  1. Complete message:
{
"type": "complete",
"content": {
"answer": "The complete answer...",
"sources": [
{
"id": "document_id",
"text": "Snippet of the source document...",
"metadata": {"source": "document.pdf", "page": 1}
}
],
"history": [
{"role": "user", "content": "What is RAG?"},
{"role": "assistant", "content": "The complete answer..."}
]
}
}
  1. Error message:
{
"type": "error",
"content": "Error message"
}

Web Interface

Sage provides a responsive web interface that works seamlessly on desktop and mobile devices. The interface includes:

Document Upload

  • Upload files using the upload button
  • Add tags to organize your documents
  • View upload progress and status

Conversational Chat Interface

The chat interface allows you to have natural conversations with the AI about your documents:

  • Conversation History: The system maintains the full conversation history, allowing for contextual follow-up questions
  • Conversation Sidebar: Browse and manage your chat history with a collapsible sidebar
  • Enter Key Support: Press Enter to send your message (use Shift+Enter for new lines)
  • Visual Distinction: User and assistant messages are visually distinct for easy reading
  • Source Citations: The system displays the source documents used to generate responses
  • Mobile Responsive: Optimized interface for mobile devices with touch-friendly controls
  • Tag-based Filtering: Filter conversations and documents by tags with advanced search capabilities

To use the conversational interface:

  1. Select a a set of tags from the dropdown (or choose to search all tags)
  2. (Optional) Select a specific LLM model and provider
  3. Type your question in the message box
  4. Press Enter or click "Send Message"
  5. View the AI's response and the sources used
  6. Continue the conversation with follow-up questions

The system will maintain context between questions, allowing for a more natural conversation flow. On mobile devices, the conversation sidebar can be toggled using the sidebar button in the top navigation.

Examples

Python Client Example

Here's a simple Python client for interacting with the API:

importrequestsimportwebsocketsimportjsonimportasyncioBASE_URL="http://localhost:8000/api/v1"# Create a collectiondefcreate_collection(name, description=None):
response=requests.post(
f"{BASE_URL}/collections/",
json={"name": name, "description": description}
)
returnresponse.json()
# Upload a documentdefupload_document(file_path, collection_name, tags=None, metadata=None):
files= {"file": open(file_path, "rb")}
data= {"collection_name": collection_name}
iftags:
data["tags"] =",".join(tags)
ifmetadata:
data["additional_metadata"] =json.dumps(metadata)
response=requests.post(
f"{BASE_URL}/documents/upload",
files=files,
data=data
)
returnresponse.json()
# Query documentsdefquery_documents(collection_name, query_text, n_results=5, tags=None):
query_data= {
"collection_name": collection_name,
"query_text": query_text,
"n_results": n_results
}
# Add tag filtering if specifiediftags:
query_data["where"] = {"tags": {"$in": tags}}
response=requests.post(
f"{BASE_URL}/documents/query",
json=query_data
)
returnresponse.json()
# WebSocket chat clientasyncdefchat_websocket(query, tags=None, history=None, model=None, include_untagged=True):
uri=f"ws://localhost:8000/api/v1/chat/ws"asyncwithwebsockets.connect(uri) aswebsocket:
# Send chat messagemessage= {
"type": "chat",
"query": query,
"history": historyor [],
"include_untagged": include_untagged
}
iftags:
message["tags"] =tagsifmodel:
message["model"] =modelawaitwebsocket.send(json.dumps(message))
# Process responsesfull_response=""whileTrue:
response=awaitwebsocket.recv()
data=json.loads(response)
ifdata["type"] =="token":
# Print token as it arrivesprint(data["content"], end="", flush=True)
full_response+=data["content"]
elifdata["type"] =="complete":
# Chat completeprint("\n\nSources:")
fori, sourceinenumerate(data["content"]["sources"]):
print(f"{i+1}. {source['text']} (ID: {source['id']})")
returndata["content"]
elifdata["type"] =="error":
# Error occurredprint(f"\nError: {data['content']}")
returnNone# Example usageif__name__=="__main__":
# Create a collectioncreate_collection("my_docs", "My documents collection")
# Upload a document with tagsupload_document(
"path/to/document.pdf", "my_docs", tags=["house", "important", "insurance"],
metadata={"source": "local", "author": "Insurance Company"}
)
# Upload another document with different tagsupload_document(
"path/to/another_document.pdf", "my_docs", tags=["personal", "important", "legal"],
metadata={"source": "local", "author": "Law Firm"}
)
# Query documents without tag filteringresults=query_documents("my_docs", "What is RAG?")
# Query documents with tag filtering (only "house" tagged documents)house_results=query_documents("my_docs", "What is RAG?", tags=["house"])
# Query documents with multiple tag filtering (documents tagged as either "important" or "personal")important_results=query_documents("my_docs", "What is RAG?", tags=["important", "personal"])
# Chat with WebSocket - Single queryasyncio.run(chat_websocket("Explain RAG in simple terms."))
# Chat with WebSocket - Query with tag filteringasyncio.run(chat_websocket(
"What can you tell me about my house documents?", tags=["house", "important"]
))
# Chat with WebSocket - Conversation with historyasyncdefconversation_example():
# First queryprint("\n--- Starting conversation ---")
response=awaitchat_websocket("Explain RAG in simple terms.")
# Get the updated history from the responseconversation_history=response["history"]
# Follow-up query using the conversation history with tag filteringprint("\n--- Follow-up question with tag filtering ---")
response=awaitchat_websocket(
"What are the main benefits compared to traditional approaches?",
tags=["important", "personal"],
history=conversation_history
)
# Continue the conversation with another follow-upconversation_history=response["history"]
print("\n--- Another follow-up question ---")
awaitchat_websocket(
"Can you give me a simple code example?",
history=conversation_history
)
# Run the conversation exampleasyncio.run(conversation_example())

License

MIT

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

About

RAGU - RAG Utility for interrogating documents

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

SAGE: Personal RAG-enabled Information Retrieval Utility

SAGE is a complete RAG (Retrieval-Augmented Generation) management and interrogation system built with FastAPI and ChromaDB. It provides a powerful API for document management, vector storage, and LLM-powered chat functionality with support for multiple LLM providers.

Features

  • Document Management: Upload, process, and manage documents in vector collections
  • Document Tagging: Add tags to documents for better organization and filtering
  • Vector Database: ChromaDB integration for efficient semantic search
  • LLM Integration: Support for multiple LLM providers:
    • Ollama with local or remote models (including nomic-embed-text)
    • Anthropic with Claude models (including haiku)
    • OpenAI with GPT models (including 4o)
  • Conversational UI: Interactive chat interface with conversation history and context awareness
  • WebSocket Chat: Real-time, streaming chat interface with RAG capabilities
  • REST API: Comprehensive REST API for all operations
  • Document Processing: Support for various document types (PDF, DOCX, CSV, text)

Installation

Prerequisites

  • Python 3.8+
  • [Optional] Virtual environment (recommended)

Setup

  1. Clone the repository:

    jj git clone --colocate https://github.com/grepory/sage.git
    cd sage
  2. Create and activate a virtual environment (optional but recommended):

    python -m venv venv
    source venv/bin/activate # On Windows: venv\Scripts\activate
  3. Install dependencies:

    pip install -r requirements.txt
  4. Copy the .env.example file to .env in the project root and update it with your LLM provider configuration:

    cp .env.example .env
    # Then edit the .env file with your preferred text editor

    The .env.example file contains all the necessary environment variables with default values and explanatory comments. You only need to configure the providers you plan to use.

Running the Application

Start the server with:

python run_app.py

This will:

  • Start the server on the configured port (default 8000)
  • Display clear URLs for accessing the application
  • Automatically open your default web browser to the frontend

Alternatively, you can start the server manually with:

uvicorn app.main:app --reload --host 0.0.0.0 --port 8000

The API will be available at http://localhost:8000, and the API documentation at http://localhost:8000/docs.

Configuration

Sage can be configured through environment variables in the .env file:

Server Configuration

  • PORT=8000 - Port for the FastAPI server
  • ROOT_PATH=/sage - Root path prefix for reverse proxy setup (e.g., for hosting at http://example.com/sage/)

LLM Provider Configuration

Configure your preferred LLM providers in the .env file:

API Documentation

Collections API

Collections are used to organize documents in the vector database.

Create a Collection

POST /api/v1/collections/Content-Type: application/json
{
"name": "my_collection",
"description": "My first collection"
}

List Collections

GET /api/v1/collections/

Get Collection

GET /api/v1/collections/{collection_name}

Delete Collection

DELETE /api/v1/collections/{collection_name}

Documents API

Documents are the core content stored in the vector database.

Upload a Document

POST /api/v1/documents/uploadContent-Type: multipart/form-datafile: [file]collection_name: my_collectiontags: personal,house,importantadditional_metadata: {"source": "website", "author": "John Doe"}

The tags parameter is optional and should be a comma-separated list of tags to associate with the document.

Add Text Directly

POST /api/v1/documents/textContent-Type: application/json
{
"text": "This is some text to add to the vector database.",
"collection_name": "my_collection",
"tags": ["personal", "notes", "important"],
"metadata": {"source": "direct_input", "author": "Jane Smith"}
}

The tags parameter is optional and should be an array of strings representing tags to associate with the document.

Get Document

GET /api/v1/documents/{collection_name}/{document_id}

Delete Document

DELETE /api/v1/documents/{collection_name}/{document_id}

Query Documents

POST /api/v1/documents/queryContent-Type: application/json
{
"collection_name": "my_collection",
"query_text": "What is retrieval-augmented generation?",
"n_results": 5,
"where": {
"tags": {"$in": ["personal", "important"]}
}
}

The where parameter is optional and can be used to filter documents by metadata fields, including tags. In the example above, the query will only return documents that have either the "personal" or "important" tag. You can use other operators like $eq for exact match or $contains for substring match.

Chat API

The Chat API provides both REST and WebSocket interfaces for interacting with the LLM.

Model Selection

You can specify which LLM provider and model to use by setting the model parameter in your requests. The format is:

provider:model

For example:

  • ollama:llama2 - Use Ollama with the llama2 model
  • anthropic:claude-3-haiku-20240307 - Use Anthropic with the Claude 3 Haiku model
  • openai:gpt-4o - Use OpenAI with the GPT-4o model

If you don't specify a provider, the default provider from your configuration will be used. If you don't specify a model, the default model for the selected provider will be used.

REST Chat Endpoint

POST /api/v1/chat/Content-Type: application/json
{
"collection_name": "my_collection",
"query": "What is retrieval-augmented generation?",
"history": [
{"role": "user", "content": "Hello"},
{"role": "assistant", "content": "Hi there! How can I help you?"}
],
"model": "llama2"
}

WebSocket Chat

Connect to the WebSocket endpoint at /api/v1/chat/ws and send JSON messages with the following format:

{
"type": "chat",
"collection_name": "my_collection",
"query": "What is retrieval-augmented generation?",
"history": [
{"role": "user", "content": "Hello"},
{"role": "assistant", "content": "Hi there! How can I help you?"}
],
"model": "llama2"
}

The server will respond with messages in the following formats:

  1. Start message:
{
"type": "start",
"content": "Generating response..."
}
  1. Token messages (streaming):
{
"type": "token",
"content": "piece of text"
}
  1. Complete message:
{
"type": "complete",
"content": {
"answer": "The complete answer...",
"sources": [
{
"id": "document_id",
"text": "Snippet of the source document...",
"metadata": {"source": "document.pdf", "page": 1}
}
],
"history": [
{"role": "user", "content": "What is RAG?"},
{"role": "assistant", "content": "The complete answer..."}
]
}
}
  1. Error message:
{
"type": "error",
"content": "Error message"
}

Web Interface

Sage provides a responsive web interface that works seamlessly on desktop and mobile devices. The interface includes:

Document Upload

  • Upload files using the upload button
  • Add tags to organize your documents
  • View upload progress and status

Conversational Chat Interface

The chat interface allows you to have natural conversations with the AI about your documents:

  • Conversation History: The system maintains the full conversation history, allowing for contextual follow-up questions
  • Conversation Sidebar: Browse and manage your chat history with a collapsible sidebar
  • Enter Key Support: Press Enter to send your message (use Shift+Enter for new lines)
  • Visual Distinction: User and assistant messages are visually distinct for easy reading
  • Source Citations: The system displays the source documents used to generate responses
  • Mobile Responsive: Optimized interface for mobile devices with touch-friendly controls
  • Tag-based Filtering: Filter conversations and documents by tags with advanced search capabilities

To use the conversational interface:

  1. Select a a set of tags from the dropdown (or choose to search all tags)
  2. (Optional) Select a specific LLM model and provider
  3. Type your question in the message box
  4. Press Enter or click "Send Message"
  5. View the AI's response and the sources used
  6. Continue the conversation with follow-up questions

The system will maintain context between questions, allowing for a more natural conversation flow. On mobile devices, the conversation sidebar can be toggled using the sidebar button in the top navigation.

Examples

Python Client Example

Here's a simple Python client for interacting with the API:

importrequestsimportwebsocketsimportjsonimportasyncioBASE_URL="http://localhost:8000/api/v1"# Create a collectiondefcreate_collection(name, description=None):
response=requests.post(
f"{BASE_URL}/collections/",
json={"name": name, "description": description}
)
returnresponse.json()
# Upload a documentdefupload_document(file_path, collection_name, tags=None, metadata=None):
files= {"file": open(file_path, "rb")}
data= {"collection_name": collection_name}
iftags:
data["tags"] =",".join(tags)
ifmetadata:
data["additional_metadata"] =json.dumps(metadata)
response=requests.post(
f"{BASE_URL}/documents/upload",
files=files,
data=data
)
returnresponse.json()
# Query documentsdefquery_documents(collection_name, query_text, n_results=5, tags=None):
query_data= {
"collection_name": collection_name,
"query_text": query_text,
"n_results": n_results
}
# Add tag filtering if specifiediftags:
query_data["where"] = {"tags": {"$in": tags}}
response=requests.post(
f"{BASE_URL}/documents/query",
json=query_data
)
returnresponse.json()
# WebSocket chat clientasyncdefchat_websocket(query, tags=None, history=None, model=None, include_untagged=True):
uri=f"ws://localhost:8000/api/v1/chat/ws"asyncwithwebsockets.connect(uri) aswebsocket:
# Send chat messagemessage= {
"type": "chat",
"query": query,
"history": historyor [],
"include_untagged": include_untagged
}
iftags:
message["tags"] =tagsifmodel:
message["model"] =modelawaitwebsocket.send(json.dumps(message))
# Process responsesfull_response=""whileTrue:
response=awaitwebsocket.recv()
data=json.loads(response)
ifdata["type"] =="token":
# Print token as it arrivesprint(data["content"], end="", flush=True)
full_response+=data["content"]
elifdata["type"] =="complete":
# Chat completeprint("\n\nSources:")
fori, sourceinenumerate(data["content"]["sources"]):
print(f"{i+1}. {source['text']} (ID: {source['id']})")
returndata["content"]
elifdata["type"] =="error":
# Error occurredprint(f"\nError: {data['content']}")
returnNone# Example usageif__name__=="__main__":
# Create a collectioncreate_collection("my_docs", "My documents collection")
# Upload a document with tagsupload_document(
"path/to/document.pdf", "my_docs", tags=["house", "important", "insurance"],
metadata={"source": "local", "author": "Insurance Company"}
)
# Upload another document with different tagsupload_document(
"path/to/another_document.pdf", "my_docs", tags=["personal", "important", "legal"],
metadata={"source": "local", "author": "Law Firm"}
)
# Query documents without tag filteringresults=query_documents("my_docs", "What is RAG?")
# Query documents with tag filtering (only "house" tagged documents)house_results=query_documents("my_docs", "What is RAG?", tags=["house"])
# Query documents with multiple tag filtering (documents tagged as either "important" or "personal")important_results=query_documents("my_docs", "What is RAG?", tags=["important", "personal"])
# Chat with WebSocket - Single queryasyncio.run(chat_websocket("Explain RAG in simple terms."))
# Chat with WebSocket - Query with tag filteringasyncio.run(chat_websocket(
"What can you tell me about my house documents?", tags=["house", "important"]
))
# Chat with WebSocket - Conversation with historyasyncdefconversation_example():
# First queryprint("\n--- Starting conversation ---")
response=awaitchat_websocket("Explain RAG in simple terms.")
# Get the updated history from the responseconversation_history=response["history"]
# Follow-up query using the conversation history with tag filteringprint("\n--- Follow-up question with tag filtering ---")
response=awaitchat_websocket(
"What are the main benefits compared to traditional approaches?",
tags=["important", "personal"],
history=conversation_history
)
# Continue the conversation with another follow-upconversation_history=response["history"]
print("\n--- Another follow-up question ---")
awaitchat_websocket(
"Can you give me a simple code example?",
history=conversation_history
)
# Run the conversation exampleasyncio.run(conversation_example())

License

MIT

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

About

RAGU - RAG Utility for interrogating documents

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SAGE: Personal RAG-enabled Information Retrieval Utility

SAGE is a complete RAG (Retrieval-Augmented Generation) management and interrogation system built with FastAPI and ChromaDB. It provides a powerful API for document management, vector storage, and LLM-powered chat functionality with support for multiple LLM providers.

Features

  • Document Management: Upload, process, and manage documents in vector collections
  • Document Tagging: Add tags to documents for better organization and filtering
  • Vector Database: ChromaDB integration for efficient semantic search
  • LLM Integration: Support for multiple LLM providers:
    • Ollama with local or remote models (including nomic-embed-text)
    • Anthropic with Claude models (including haiku)
    • OpenAI with GPT models (including 4o)
  • Conversational UI: Interactive chat interface with conversation history and context awareness
  • WebSocket Chat: Real-time, streaming chat interface with RAG capabilities
  • REST API: Comprehensive REST API for all operations
  • Document Processing: Support for various document types (PDF, DOCX, CSV, text)

Installation

Prerequisites

  • Python 3.8+
  • [Optional] Virtual environment (recommended)

Setup

  1. Clone the repository:

    jj git clone --colocate https://github.com/grepory/sage.git
    cd sage
  2. Create and activate a virtual environment (optional but recommended):

    python -m venv venv
    source venv/bin/activate # On Windows: venv\Scripts\activate
  3. Install dependencies:

    pip install -r requirements.txt
  4. Copy the .env.example file to .env in the project root and update it with your LLM provider configuration:

    cp .env.example .env
    # Then edit the .env file with your preferred text editor

    The .env.example file contains all the necessary environment variables with default values and explanatory comments. You only need to configure the providers you plan to use.

Running the Application

Start the server with:

python run_app.py

This will:

  • Start the server on the configured port (default 8000)
  • Display clear URLs for accessing the application
  • Automatically open your default web browser to the frontend

Alternatively, you can start the server manually with:

uvicorn app.main:app --reload --host 0.0.0.0 --port 8000

The API will be available at http://localhost:8000, and the API documentation at http://localhost:8000/docs.

Configuration

Sage can be configured through environment variables in the .env file:

Server Configuration

  • PORT=8000 - Port for the FastAPI server
  • ROOT_PATH=/sage - Root path prefix for reverse proxy setup (e.g., for hosting at http://example.com/sage/)

LLM Provider Configuration

Configure your preferred LLM providers in the .env file:

API Documentation

Collections API

Collections are used to organize documents in the vector database.

Create a Collection

POST /api/v1/collections/Content-Type: application/json
{
"name": "my_collection",
"description": "My first collection"
}

List Collections

GET /api/v1/collections/

Get Collection

GET /api/v1/collections/{collection_name}

Delete Collection

DELETE /api/v1/collections/{collection_name}

Documents API

Documents are the core content stored in the vector database.

Upload a Document

POST /api/v1/documents/uploadContent-Type: multipart/form-datafile: [file]collection_name: my_collectiontags: personal,house,importantadditional_metadata: {"source": "website", "author": "John Doe"}

The tags parameter is optional and should be a comma-separated list of tags to associate with the document.

Add Text Directly

POST /api/v1/documents/textContent-Type: application/json
{
"text": "This is some text to add to the vector database.",
"collection_name": "my_collection",
"tags": ["personal", "notes", "important"],
"metadata": {"source": "direct_input", "author": "Jane Smith"}
}

The tags parameter is optional and should be an array of strings representing tags to associate with the document.

Get Document

GET /api/v1/documents/{collection_name}/{document_id}

Delete Document

DELETE /api/v1/documents/{collection_name}/{document_id}

Query Documents

POST /api/v1/documents/queryContent-Type: application/json
{
"collection_name": "my_collection",
"query_text": "What is retrieval-augmented generation?",
"n_results": 5,
"where": {
"tags": {"$in": ["personal", "important"]}
}
}

The where parameter is optional and can be used to filter documents by metadata fields, including tags. In the example above, the query will only return documents that have either the "personal" or "important" tag. You can use other operators like $eq for exact match or $contains for substring match.

Chat API

The Chat API provides both REST and WebSocket interfaces for interacting with the LLM.

Model Selection

You can specify which LLM provider and model to use by setting the model parameter in your requests. The format is:

provider:model

For example:

  • ollama:llama2 - Use Ollama with the llama2 model
  • anthropic:claude-3-haiku-20240307 - Use Anthropic with the Claude 3 Haiku model
  • openai:gpt-4o - Use OpenAI with the GPT-4o model

If you don't specify a provider, the default provider from your configuration will be used. If you don't specify a model, the default model for the selected provider will be used.

REST Chat Endpoint

POST /api/v1/chat/Content-Type: application/json
{
"collection_name": "my_collection",
"query": "What is retrieval-augmented generation?",
"history": [
{"role": "user", "content": "Hello"},
{"role": "assistant", "content": "Hi there! How can I help you?"}
],
"model": "llama2"
}

WebSocket Chat

Connect to the WebSocket endpoint at /api/v1/chat/ws and send JSON messages with the following format:

{
"type": "chat",
"collection_name": "my_collection",
"query": "What is retrieval-augmented generation?",
"history": [
{"role": "user", "content": "Hello"},
{"role": "assistant", "content": "Hi there! How can I help you?"}
],
"model": "llama2"
}

The server will respond with messages in the following formats:

  1. Start message:
{
"type": "start",
"content": "Generating response..."
}
  1. Token messages (streaming):
{
"type": "token",
"content": "piece of text"
}
  1. Complete message:
{
"type": "complete",
"content": {
"answer": "The complete answer...",
"sources": [
{
"id": "document_id",
"text": "Snippet of the source document...",
"metadata": {"source": "document.pdf", "page": 1}
}
],
"history": [
{"role": "user", "content": "What is RAG?"},
{"role": "assistant", "content": "The complete answer..."}
]
}
}
  1. Error message:
{
"type": "error",
"content": "Error message"
}

Web Interface

Sage provides a responsive web interface that works seamlessly on desktop and mobile devices. The interface includes:

Document Upload

  • Upload files using the upload button
  • Add tags to organize your documents
  • View upload progress and status

Conversational Chat Interface

The chat interface allows you to have natural conversations with the AI about your documents:

  • Conversation History: The system maintains the full conversation history, allowing for contextual follow-up questions
  • Conversation Sidebar: Browse and manage your chat history with a collapsible sidebar
  • Enter Key Support: Press Enter to send your message (use Shift+Enter for new lines)
  • Visual Distinction: User and assistant messages are visually distinct for easy reading
  • Source Citations: The system displays the source documents used to generate responses
  • Mobile Responsive: Optimized interface for mobile devices with touch-friendly controls
  • Tag-based Filtering: Filter conversations and documents by tags with advanced search capabilities

To use the conversational interface:

  1. Select a a set of tags from the dropdown (or choose to search all tags)
  2. (Optional) Select a specific LLM model and provider
  3. Type your question in the message box
  4. Press Enter or click "Send Message"
  5. View the AI's response and the sources used
  6. Continue the conversation with follow-up questions

The system will maintain context between questions, allowing for a more natural conversation flow. On mobile devices, the conversation sidebar can be toggled using the sidebar button in the top navigation.

Examples

Python Client Example

Here's a simple Python client for interacting with the API:

importrequestsimportwebsocketsimportjsonimportasyncioBASE_URL="http://localhost:8000/api/v1"# Create a collectiondefcreate_collection(name, description=None):
response=requests.post(
f"{BASE_URL}/collections/",
json={"name": name, "description": description}
)
returnresponse.json()
# Upload a documentdefupload_document(file_path, collection_name, tags=None, metadata=None):
files= {"file": open(file_path, "rb")}
data= {"collection_name": collection_name}
iftags:
data["tags"] =",".join(tags)
ifmetadata:
data["additional_metadata"] =json.dumps(metadata)
response=requests.post(
f"{BASE_URL}/documents/upload",
files=files,
data=data
)
returnresponse.json()
# Query documentsdefquery_documents(collection_name, query_text, n_results=5, tags=None):
query_data= {
"collection_name": collection_name,
"query_text": query_text,
"n_results": n_results
}
# Add tag filtering if specifiediftags:
query_data["where"] = {"tags": {"$in": tags}}
response=requests.post(
f"{BASE_URL}/documents/query",
json=query_data
)
returnresponse.json()
# WebSocket chat clientasyncdefchat_websocket(query, tags=None, history=None, model=None, include_untagged=True):
uri=f"ws://localhost:8000/api/v1/chat/ws"asyncwithwebsockets.connect(uri) aswebsocket:
# Send chat messagemessage= {
"type": "chat",
"query": query,
"history": historyor [],
"include_untagged": include_untagged
}
iftags:
message["tags"] =tagsifmodel:
message["model"] =modelawaitwebsocket.send(json.dumps(message))
# Process responsesfull_response=""whileTrue:
response=awaitwebsocket.recv()
data=json.loads(response)
ifdata["type"] =="token":
# Print token as it arrivesprint(data["content"], end="", flush=True)
full_response+=data["content"]
elifdata["type"] =="complete":
# Chat completeprint("\n\nSources:")
fori, sourceinenumerate(data["content"]["sources"]):
print(f"{i+1}. {source['text']} (ID: {source['id']})")
returndata["content"]
elifdata["type"] =="error":
# Error occurredprint(f"\nError: {data['content']}")
returnNone# Example usageif__name__=="__main__":
# Create a collectioncreate_collection("my_docs", "My documents collection")
# Upload a document with tagsupload_document(
"path/to/document.pdf", "my_docs", tags=["house", "important", "insurance"],
metadata={"source": "local", "author": "Insurance Company"}
)
# Upload another document with different tagsupload_document(
"path/to/another_document.pdf", "my_docs", tags=["personal", "important", "legal"],
metadata={"source": "local", "author": "Law Firm"}
)
# Query documents without tag filteringresults=query_documents("my_docs", "What is RAG?")
# Query documents with tag filtering (only "house" tagged documents)house_results=query_documents("my_docs", "What is RAG?", tags=["house"])
# Query documents with multiple tag filtering (documents tagged as either "important" or "personal")important_results=query_documents("my_docs", "What is RAG?", tags=["important", "personal"])
# Chat with WebSocket - Single queryasyncio.run(chat_websocket("Explain RAG in simple terms."))
# Chat with WebSocket - Query with tag filteringasyncio.run(chat_websocket(
"What can you tell me about my house documents?", tags=["house", "important"]
))
# Chat with WebSocket - Conversation with historyasyncdefconversation_example():
# First queryprint("\n--- Starting conversation ---")
response=awaitchat_websocket("Explain RAG in simple terms.")
# Get the updated history from the responseconversation_history=response["history"]
# Follow-up query using the conversation history with tag filteringprint("\n--- Follow-up question with tag filtering ---")
response=awaitchat_websocket(
"What are the main benefits compared to traditional approaches?",
tags=["important", "personal"],
history=conversation_history
)
# Continue the conversation with another follow-upconversation_history=response["history"]
print("\n--- Another follow-up question ---")
awaitchat_websocket(
"Can you give me a simple code example?",
history=conversation_history
)
# Run the conversation exampleasyncio.run(conversation_example())

License

MIT

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

About

RAGU - RAG Utility for interrogating documents

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SAGE: Personal RAG-enabled Information Retrieval Utility

SAGE is a complete RAG (Retrieval-Augmented Generation) management and interrogation system built with FastAPI and ChromaDB. It provides a powerful API for document management, vector storage, and LLM-powered chat functionality with support for multiple LLM providers.

Features

  • Document Management: Upload, process, and manage documents in vector collections
  • Document Tagging: Add tags to documents for better organization and filtering
  • Vector Database: ChromaDB integration for efficient semantic search
  • LLM Integration: Support for multiple LLM providers:
    • Ollama with local or remote models (including nomic-embed-text)
    • Anthropic with Claude models (including haiku)
    • OpenAI with GPT models (including 4o)
  • Conversational UI: Interactive chat interface with conversation history and context awareness
  • WebSocket Chat: Real-time, streaming chat interface with RAG capabilities
  • REST API: Comprehensive REST API for all operations
  • Document Processing: Support for various document types (PDF, DOCX, CSV, text)

Installation

Prerequisites

  • Python 3.8+
  • [Optional] Virtual environment (recommended)

Setup

  1. Clone the repository:

    jj git clone --colocate https://github.com/grepory/sage.git
    cd sage
  2. Create and activate a virtual environment (optional but recommended):

    python -m venv venv
    source venv/bin/activate # On Windows: venv\Scripts\activate
  3. Install dependencies:

    pip install -r requirements.txt
  4. Copy the .env.example file to .env in the project root and update it with your LLM provider configuration:

    cp .env.example .env
    # Then edit the .env file with your preferred text editor

    The .env.example file contains all the necessary environment variables with default values and explanatory comments. You only need to configure the providers you plan to use.

Running the Application

Start the server with:

python run_app.py

This will:

  • Start the server on the configured port (default 8000)
  • Display clear URLs for accessing the application
  • Automatically open your default web browser to the frontend

Alternatively, you can start the server manually with:

uvicorn app.main:app --reload --host 0.0.0.0 --port 8000

The API will be available at http://localhost:8000, and the API documentation at http://localhost:8000/docs.

Configuration

Sage can be configured through environment variables in the .env file:

Server Configuration

  • PORT=8000 - Port for the FastAPI server
  • ROOT_PATH=/sage - Root path prefix for reverse proxy setup (e.g., for hosting at http://example.com/sage/)

LLM Provider Configuration

Configure your preferred LLM providers in the .env file:

API Documentation

Collections API

Collections are used to organize documents in the vector database.

Create a Collection

POST /api/v1/collections/Content-Type: application/json
{
"name": "my_collection",
"description": "My first collection"
}

List Collections

GET /api/v1/collections/

Get Collection

GET /api/v1/collections/{collection_name}

Delete Collection

DELETE /api/v1/collections/{collection_name}

Documents API

Documents are the core content stored in the vector database.

Upload a Document

POST /api/v1/documents/uploadContent-Type: multipart/form-datafile: [file]collection_name: my_collectiontags: personal,house,importantadditional_metadata: {"source": "website", "author": "John Doe"}

The tags parameter is optional and should be a comma-separated list of tags to associate with the document.

Add Text Directly

POST /api/v1/documents/textContent-Type: application/json
{
"text": "This is some text to add to the vector database.",
"collection_name": "my_collection",
"tags": ["personal", "notes", "important"],
"metadata": {"source": "direct_input", "author": "Jane Smith"}
}

The tags parameter is optional and should be an array of strings representing tags to associate with the document.

Get Document

GET /api/v1/documents/{collection_name}/{document_id}

Delete Document

DELETE /api/v1/documents/{collection_name}/{document_id}

Query Documents

POST /api/v1/documents/queryContent-Type: application/json
{
"collection_name": "my_collection",
"query_text": "What is retrieval-augmented generation?",
"n_results": 5,
"where": {
"tags": {"$in": ["personal", "important"]}
}
}

The where parameter is optional and can be used to filter documents by metadata fields, including tags. In the example above, the query will only return documents that have either the "personal" or "important" tag. You can use other operators like $eq for exact match or $contains for substring match.

Chat API

The Chat API provides both REST and WebSocket interfaces for interacting with the LLM.

Model Selection

You can specify which LLM provider and model to use by setting the model parameter in your requests. The format is:

provider:model

For example:

  • ollama:llama2 - Use Ollama with the llama2 model
  • anthropic:claude-3-haiku-20240307 - Use Anthropic with the Claude 3 Haiku model
  • openai:gpt-4o - Use OpenAI with the GPT-4o model

If you don't specify a provider, the default provider from your configuration will be used. If you don't specify a model, the default model for the selected provider will be used.

REST Chat Endpoint

POST /api/v1/chat/Content-Type: application/json
{
"collection_name": "my_collection",
"query": "What is retrieval-augmented generation?",
"history": [
{"role": "user", "content": "Hello"},
{"role": "assistant", "content": "Hi there! How can I help you?"}
],
"model": "llama2"
}

WebSocket Chat

Connect to the WebSocket endpoint at /api/v1/chat/ws and send JSON messages with the following format:

{
"type": "chat",
"collection_name": "my_collection",
"query": "What is retrieval-augmented generation?",
"history": [
{"role": "user", "content": "Hello"},
{"role": "assistant", "content": "Hi there! How can I help you?"}
],
"model": "llama2"
}

The server will respond with messages in the following formats:

  1. Start message:
{
"type": "start",
"content": "Generating response..."
}
  1. Token messages (streaming):
{
"type": "token",
"content": "piece of text"
}
  1. Complete message:
{
"type": "complete",
"content": {
"answer": "The complete answer...",
"sources": [
{
"id": "document_id",
"text": "Snippet of the source document...",
"metadata": {"source": "document.pdf", "page": 1}
}
],
"history": [
{"role": "user", "content": "What is RAG?"},
{"role": "assistant", "content": "The complete answer..."}
]
}
}
  1. Error message:
{
"type": "error",
"content": "Error message"
}

Web Interface

Sage provides a responsive web interface that works seamlessly on desktop and mobile devices. The interface includes:

Document Upload

  • Upload files using the upload button
  • Add tags to organize your documents
  • View upload progress and status

Conversational Chat Interface

The chat interface allows you to have natural conversations with the AI about your documents:

  • Conversation History: The system maintains the full conversation history, allowing for contextual follow-up questions
  • Conversation Sidebar: Browse and manage your chat history with a collapsible sidebar
  • Enter Key Support: Press Enter to send your message (use Shift+Enter for new lines)
  • Visual Distinction: User and assistant messages are visually distinct for easy reading
  • Source Citations: The system displays the source documents used to generate responses
  • Mobile Responsive: Optimized interface for mobile devices with touch-friendly controls
  • Tag-based Filtering: Filter conversations and documents by tags with advanced search capabilities

To use the conversational interface:

  1. Select a a set of tags from the dropdown (or choose to search all tags)
  2. (Optional) Select a specific LLM model and provider
  3. Type your question in the message box
  4. Press Enter or click "Send Message"
  5. View the AI's response and the sources used
  6. Continue the conversation with follow-up questions

The system will maintain context between questions, allowing for a more natural conversation flow. On mobile devices, the conversation sidebar can be toggled using the sidebar button in the top navigation.

Examples

Python Client Example

Here's a simple Python client for interacting with the API:

importrequestsimportwebsocketsimportjsonimportasyncioBASE_URL="http://localhost:8000/api/v1"# Create a collectiondefcreate_collection(name, description=None):
response=requests.post(
f"{BASE_URL}/collections/",
json={"name": name, "description": description}
)
returnresponse.json()
# Upload a documentdefupload_document(file_path, collection_name, tags=None, metadata=None):
files= {"file": open(file_path, "rb")}
data= {"collection_name": collection_name}
iftags:
data["tags"] =",".join(tags)
ifmetadata:
data["additional_metadata"] =json.dumps(metadata)
response=requests.post(
f"{BASE_URL}/documents/upload",
files=files,
data=data
)
returnresponse.json()
# Query documentsdefquery_documents(collection_name, query_text, n_results=5, tags=None):
query_data= {
"collection_name": collection_name,
"query_text": query_text,
"n_results": n_results
}
# Add tag filtering if specifiediftags:
query_data["where"] = {"tags": {"$in": tags}}
response=requests.post(
f"{BASE_URL}/documents/query",
json=query_data
)
returnresponse.json()
# WebSocket chat clientasyncdefchat_websocket(query, tags=None, history=None, model=None, include_untagged=True):
uri=f"ws://localhost:8000/api/v1/chat/ws"asyncwithwebsockets.connect(uri) aswebsocket:
# Send chat messagemessage= {
"type": "chat",
"query": query,
"history": historyor [],
"include_untagged": include_untagged
}
iftags:
message["tags"] =tagsifmodel:
message["model"] =modelawaitwebsocket.send(json.dumps(message))
# Process responsesfull_response=""whileTrue:
response=awaitwebsocket.recv()
data=json.loads(response)
ifdata["type"] =="token":
# Print token as it arrivesprint(data["content"], end="", flush=True)
full_response+=data["content"]
elifdata["type"] =="complete":
# Chat completeprint("\n\nSources:")
fori, sourceinenumerate(data["content"]["sources"]):
print(f"{i+1}. {source['text']} (ID: {source['id']})")
returndata["content"]
elifdata["type"] =="error":
# Error occurredprint(f"\nError: {data['content']}")
returnNone# Example usageif__name__=="__main__":
# Create a collectioncreate_collection("my_docs", "My documents collection")
# Upload a document with tagsupload_document(
"path/to/document.pdf", "my_docs", tags=["house", "important", "insurance"],
metadata={"source": "local", "author": "Insurance Company"}
)
# Upload another document with different tagsupload_document(
"path/to/another_document.pdf", "my_docs", tags=["personal", "important", "legal"],
metadata={"source": "local", "author": "Law Firm"}
)
# Query documents without tag filteringresults=query_documents("my_docs", "What is RAG?")
# Query documents with tag filtering (only "house" tagged documents)house_results=query_documents("my_docs", "What is RAG?", tags=["house"])
# Query documents with multiple tag filtering (documents tagged as either "important" or "personal")important_results=query_documents("my_docs", "What is RAG?", tags=["important", "personal"])
# Chat with WebSocket - Single queryasyncio.run(chat_websocket("Explain RAG in simple terms."))
# Chat with WebSocket - Query with tag filteringasyncio.run(chat_websocket(
"What can you tell me about my house documents?", tags=["house", "important"]
))
# Chat with WebSocket - Conversation with historyasyncdefconversation_example():
# First queryprint("\n--- Starting conversation ---")
response=awaitchat_websocket("Explain RAG in simple terms.")
# Get the updated history from the responseconversation_history=response["history"]
# Follow-up query using the conversation history with tag filteringprint("\n--- Follow-up question with tag filtering ---")
response=awaitchat_websocket(
"What are the main benefits compared to traditional approaches?",
tags=["important", "personal"],
history=conversation_history
)
# Continue the conversation with another follow-upconversation_history=response["history"]
print("\n--- Another follow-up question ---")
awaitchat_websocket(
"Can you give me a simple code example?",
history=conversation_history
)
# Run the conversation exampleasyncio.run(conversation_example())

License

MIT

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

About

RAGU - RAG Utility for interrogating documents

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

SAGE: Personal RAG-enabled Information Retrieval Utility

SAGE is a complete RAG (Retrieval-Augmented Generation) management and interrogation system built with FastAPI and ChromaDB. It provides a powerful API for document management, vector storage, and LLM-powered chat functionality with support for multiple LLM providers.

Features

  • Document Management: Upload, process, and manage documents in vector collections
  • Document Tagging: Add tags to documents for better organization and filtering
  • Vector Database: ChromaDB integration for efficient semantic search
  • LLM Integration: Support for multiple LLM providers:
    • Ollama with local or remote models (including nomic-embed-text)
    • Anthropic with Claude models (including haiku)
    • OpenAI with GPT models (including 4o)
  • Conversational UI: Interactive chat interface with conversation history and context awareness
  • WebSocket Chat: Real-time, streaming chat interface with RAG capabilities
  • REST API: Comprehensive REST API for all operations
  • Document Processing: Support for various document types (PDF, DOCX, CSV, text)

Installation

Prerequisites

  • Python 3.8+
  • [Optional] Virtual environment (recommended)

Setup

  1. Clone the repository:

    jj git clone --colocate https://github.com/grepory/sage.git
    cd sage
  2. Create and activate a virtual environment (optional but recommended):

    python -m venv venv
    source venv/bin/activate # On Windows: venv\Scripts\activate
  3. Install dependencies:

    pip install -r requirements.txt
  4. Copy the .env.example file to .env in the project root and update it with your LLM provider configuration:

    cp .env.example .env
    # Then edit the .env file with your preferred text editor

    The .env.example file contains all the necessary environment variables with default values and explanatory comments. You only need to configure the providers you plan to use.

Running the Application

Start the server with:

python run_app.py

This will:

  • Start the server on the configured port (default 8000)
  • Display clear URLs for accessing the application
  • Automatically open your default web browser to the frontend

Alternatively, you can start the server manually with:

uvicorn app.main:app --reload --host 0.0.0.0 --port 8000

The API will be available at http://localhost:8000, and the API documentation at http://localhost:8000/docs.

Configuration

Sage can be configured through environment variables in the .env file:

Server Configuration

  • PORT=8000 - Port for the FastAPI server
  • ROOT_PATH=/sage - Root path prefix for reverse proxy setup (e.g., for hosting at http://example.com/sage/)

LLM Provider Configuration

Configure your preferred LLM providers in the .env file:

API Documentation

Collections API

Collections are used to organize documents in the vector database.

Create a Collection

POST /api/v1/collections/Content-Type: application/json
{
"name": "my_collection",
"description": "My first collection"
}

List Collections

GET /api/v1/collections/

Get Collection

GET /api/v1/collections/{collection_name}

Delete Collection

DELETE /api/v1/collections/{collection_name}

Documents API

Documents are the core content stored in the vector database.

Upload a Document

POST /api/v1/documents/uploadContent-Type: multipart/form-datafile: [file]collection_name: my_collectiontags: personal,house,importantadditional_metadata: {"source": "website", "author": "John Doe"}

The tags parameter is optional and should be a comma-separated list of tags to associate with the document.

Add Text Directly

POST /api/v1/documents/textContent-Type: application/json
{
"text": "This is some text to add to the vector database.",
"collection_name": "my_collection",
"tags": ["personal", "notes", "important"],
"metadata": {"source": "direct_input", "author": "Jane Smith"}
}

The tags parameter is optional and should be an array of strings representing tags to associate with the document.

Get Document

GET /api/v1/documents/{collection_name}/{document_id}

Delete Document

DELETE /api/v1/documents/{collection_name}/{document_id}

Query Documents

POST /api/v1/documents/queryContent-Type: application/json
{
"collection_name": "my_collection",
"query_text": "What is retrieval-augmented generation?",
"n_results": 5,
"where": {
"tags": {"$in": ["personal", "important"]}
}
}

The where parameter is optional and can be used to filter documents by metadata fields, including tags. In the example above, the query will only return documents that have either the "personal" or "important" tag. You can use other operators like $eq for exact match or $contains for substring match.

Chat API

The Chat API provides both REST and WebSocket interfaces for interacting with the LLM.

Model Selection

You can specify which LLM provider and model to use by setting the model parameter in your requests. The format is:

provider:model

For example:

  • ollama:llama2 - Use Ollama with the llama2 model
  • anthropic:claude-3-haiku-20240307 - Use Anthropic with the Claude 3 Haiku model
  • openai:gpt-4o - Use OpenAI with the GPT-4o model

If you don't specify a provider, the default provider from your configuration will be used. If you don't specify a model, the default model for the selected provider will be used.

REST Chat Endpoint

POST /api/v1/chat/Content-Type: application/json
{
"collection_name": "my_collection",
"query": "What is retrieval-augmented generation?",
"history": [
{"role": "user", "content": "Hello"},
{"role": "assistant", "content": "Hi there! How can I help you?"}
],
"model": "llama2"
}

WebSocket Chat

Connect to the WebSocket endpoint at /api/v1/chat/ws and send JSON messages with the following format:

{
"type": "chat",
"collection_name": "my_collection",
"query": "What is retrieval-augmented generation?",
"history": [
{"role": "user", "content": "Hello"},
{"role": "assistant", "content": "Hi there! How can I help you?"}
],
"model": "llama2"
}

The server will respond with messages in the following formats:

  1. Start message:
{
"type": "start",
"content": "Generating response..."
}
  1. Token messages (streaming):
{
"type": "token",
"content": "piece of text"
}
  1. Complete message:
{
"type": "complete",
"content": {
"answer": "The complete answer...",
"sources": [
{
"id": "document_id",
"text": "Snippet of the source document...",
"metadata": {"source": "document.pdf", "page": 1}
}
],
"history": [
{"role": "user", "content": "What is RAG?"},
{"role": "assistant", "content": "The complete answer..."}
]
}
}
  1. Error message:
{
"type": "error",
"content": "Error message"
}

Web Interface

Sage provides a responsive web interface that works seamlessly on desktop and mobile devices. The interface includes:

Document Upload

  • Upload files using the upload button
  • Add tags to organize your documents
  • View upload progress and status

Conversational Chat Interface

The chat interface allows you to have natural conversations with the AI about your documents:

  • Conversation History: The system maintains the full conversation history, allowing for contextual follow-up questions
  • Conversation Sidebar: Browse and manage your chat history with a collapsible sidebar
  • Enter Key Support: Press Enter to send your message (use Shift+Enter for new lines)
  • Visual Distinction: User and assistant messages are visually distinct for easy reading
  • Source Citations: The system displays the source documents used to generate responses
  • Mobile Responsive: Optimized interface for mobile devices with touch-friendly controls
  • Tag-based Filtering: Filter conversations and documents by tags with advanced search capabilities

To use the conversational interface:

  1. Select a a set of tags from the dropdown (or choose to search all tags)
  2. (Optional) Select a specific LLM model and provider
  3. Type your question in the message box
  4. Press Enter or click "Send Message"
  5. View the AI's response and the sources used
  6. Continue the conversation with follow-up questions

The system will maintain context between questions, allowing for a more natural conversation flow. On mobile devices, the conversation sidebar can be toggled using the sidebar button in the top navigation.

Examples

Python Client Example

Here's a simple Python client for interacting with the API:

importrequestsimportwebsocketsimportjsonimportasyncioBASE_URL="http://localhost:8000/api/v1"# Create a collectiondefcreate_collection(name, description=None):
response=requests.post(
f"{BASE_URL}/collections/",
json={"name": name, "description": description}
)
returnresponse.json()
# Upload a documentdefupload_document(file_path, collection_name, tags=None, metadata=None):
files= {"file": open(file_path, "rb")}
data= {"collection_name": collection_name}
iftags:
data["tags"] =",".join(tags)
ifmetadata:
data["additional_metadata"] =json.dumps(metadata)
response=requests.post(
f"{BASE_URL}/documents/upload",
files=files,
data=data
)
returnresponse.json()
# Query documentsdefquery_documents(collection_name, query_text, n_results=5, tags=None):
query_data= {
"collection_name": collection_name,
"query_text": query_text,
"n_results": n_results
}
# Add tag filtering if specifiediftags:
query_data["where"] = {"tags": {"$in": tags}}
response=requests.post(
f"{BASE_URL}/documents/query",
json=query_data
)
returnresponse.json()
# WebSocket chat clientasyncdefchat_websocket(query, tags=None, history=None, model=None, include_untagged=True):
uri=f"ws://localhost:8000/api/v1/chat/ws"asyncwithwebsockets.connect(uri) aswebsocket:
# Send chat messagemessage= {
"type": "chat",
"query": query,
"history": historyor [],
"include_untagged": include_untagged
}
iftags:
message["tags"] =tagsifmodel:
message["model"] =modelawaitwebsocket.send(json.dumps(message))
# Process responsesfull_response=""whileTrue:
response=awaitwebsocket.recv()
data=json.loads(response)
ifdata["type"] =="token":
# Print token as it arrivesprint(data["content"], end="", flush=True)
full_response+=data["content"]
elifdata["type"] =="complete":
# Chat completeprint("\n\nSources:")
fori, sourceinenumerate(data["content"]["sources"]):
print(f"{i+1}. {source['text']} (ID: {source['id']})")
returndata["content"]
elifdata["type"] =="error":
# Error occurredprint(f"\nError: {data['content']}")
returnNone# Example usageif__name__=="__main__":
# Create a collectioncreate_collection("my_docs", "My documents collection")
# Upload a document with tagsupload_document(
"path/to/document.pdf", "my_docs", tags=["house", "important", "insurance"],
metadata={"source": "local", "author": "Insurance Company"}
)
# Upload another document with different tagsupload_document(
"path/to/another_document.pdf", "my_docs", tags=["personal", "important", "legal"],
metadata={"source": "local", "author": "Law Firm"}
)
# Query documents without tag filteringresults=query_documents("my_docs", "What is RAG?")
# Query documents with tag filtering (only "house" tagged documents)house_results=query_documents("my_docs", "What is RAG?", tags=["house"])
# Query documents with multiple tag filtering (documents tagged as either "important" or "personal")important_results=query_documents("my_docs", "What is RAG?", tags=["important", "personal"])
# Chat with WebSocket - Single queryasyncio.run(chat_websocket("Explain RAG in simple terms."))
# Chat with WebSocket - Query with tag filteringasyncio.run(chat_websocket(
"What can you tell me about my house documents?", tags=["house", "important"]
))
# Chat with WebSocket - Conversation with historyasyncdefconversation_example():
# First queryprint("\n--- Starting conversation ---")
response=awaitchat_websocket("Explain RAG in simple terms.")
# Get the updated history from the responseconversation_history=response["history"]
# Follow-up query using the conversation history with tag filteringprint("\n--- Follow-up question with tag filtering ---")
response=awaitchat_websocket(
"What are the main benefits compared to traditional approaches?",
tags=["important", "personal"],
history=conversation_history
)
# Continue the conversation with another follow-upconversation_history=response["history"]
print("\n--- Another follow-up question ---")
awaitchat_websocket(
"Can you give me a simple code example?",
history=conversation_history
)
# Run the conversation exampleasyncio.run(conversation_example())

License

MIT

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

About

RAGU - RAG Utility for interrogating documents

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages