Repository files navigation

🚀 RAG API

Multi-tenant Multimodal Document Intelligent Retrieval System

Enterprise-grade RAG service built on RAG-Anything and LightRAG

CIPythonFastAPILightRAGDockerLicense

English | 简体中文

FeaturesQuick StartArchitectureAPI DocumentationDeployment


📖 Introduction

RAG API is an enterprise-grade Retrieval-Augmented Generation (RAG) service that combines the powerful document parsing capabilities of RAG-Anything with the efficient knowledge graph retrieval technology of LightRAG, providing intelligent Q&A capabilities for your documents.

🎯 Key Highlights

  • 🏢 Multi-tenant Isolation - Complete tenant data isolation for enterprise multi-tenant scenarios
  • 🎨 Multimodal Parsing - Support for PDF, Word, images and more, with full OCR, tables, and formulas coverage
  • High-performance Retrieval - Knowledge graph-based hybrid retrieval with 6-15 second query response
  • 🔄 Flexible Deployment - Support for production and development modes with one-click switching
  • 📦 Ready to Use - One-click Docker deployment, service starts in 3 minutes
  • 🎛️ Multiple Parsing Engines - DeepSeek-OCR (Remote API) + MinerU (Local/Remote API) + Docling (Fast)
  • 🎨 RAG-Anything VLM Enhancement - Three modes (off/selective/full) for deep chart understanding
  • 💾 Task Persistence - Redis storage support, tasks recoverable after container restart/instance rebuild

✨ Features

📄 Document Processing

  • Multiple Format Support

    • PDF, Word, Excel, PPT
    • PNG, JPG, WebP images
    • TXT, Markdown text
  • Intelligent Parsing

    • Plain text (.txt, .md) → Direct insertion (ultra-fast ~1s, skip parser)
    • OCR text recognition
    • Structured table extraction
    • Mathematical formula recognition
    • Layout analysis
  • RAG-Anything VLM Enhancement 🆕

    • off - Markdown only (fastest)
    • selective - Selective processing of important charts
    • full - Complete context enhancement processing
    • Smart filtering: with titles, large size, first page content
    • ⚠️Only supports remote MinerU mode, local mode uses RAG-Anything native methods
  • Batch Processing

    • Up to 100 files per batch
    • Async task queue
    • Real-time progress tracking

🔍 Intelligent Retrieval

  • Multi-mode Query

    • naive - Vector retrieval (fastest)
    • local - Local graph
    • global - Global graph
    • hybrid - Hybrid retrieval
    • mix - Full retrieval (most accurate)
  • Knowledge Graph

    • Automatic entity extraction
    • Relationship reasoning
    • Semantic understanding
    • Context enhancement
  • External Storage

    • DragonflyDB (KV storage + task storage)
    • Qdrant (vector storage)
    • Memgraph (graph database)
    • Task persistence (Redis mode)

🏗️ Architecture

System Architecture Diagram

graph TB
subgraph "Client Layer"
Client[Client Application]
WebUI[Web Interface]
end
subgraph "API Gateway Layer"
FastAPI[FastAPI Service]
Auth[Tenant Authentication]
end
subgraph "Business Logic Layer"
TenantMgr[Tenant Manager]
TaskQueue[Task Queue]
subgraph "Document Processing"
DeepSeekOCR[DeepSeek-OCR<br/>Fast OCR 80% cases]
MinerU[MinerU Parser<br/>Complex multimodal]
Docling[Docling Parser<br/>Fast lightweight]
FileRouter[Smart Router<br/>Complexity scoring]
end
subgraph "RAG Engine"
LightRAG[LightRAG Instance Pool<br/>LRU Cache 50]
KG[Knowledge Graph Engine]
Vector[Vector Retrieval Engine]
end
end
subgraph "Storage Layer"
DragonflyDB[(DragonflyDB<br/>KV Storage)]
Qdrant[(Qdrant<br/>Vector Database)]
Memgraph[(Memgraph<br/>Graph Database)]
Local[(Local Files<br/>Temp Storage)]
end
subgraph "External Services"
LLM[LLM<br/>Entity Extraction/Generation]
Embedding[Embedding<br/>Vectorization]
Rerank[Rerank<br/>Reranking]
end
Client --> FastAPI
WebUI --> FastAPI
FastAPI --> Auth
Auth --> TenantMgr
TenantMgr --> TaskQueue
TenantMgr --> LightRAG
TaskQueue --> FileRouter
FileRouter --> DeepSeekOCR
FileRouter --> MinerU
FileRouter --> Docling
DeepSeekOCR --> LightRAG
MinerU --> LightRAG
Docling --> LightRAG
LightRAG --> KG
LightRAG --> Vector
KG --> DragonflyDB
KG --> Memgraph
Vector --> Qdrant
LightRAG --> Local
LightRAG --> LLM
LightRAG --> Embedding
Vector --> Rerank
style FastAPI fill:#00C7B7
style LightRAG fill:#FF6B6B
style DeepSeekOCR fill:#5DADE2
style MinerU fill:#4ECDC4
style Docling fill:#95E1D3
style TenantMgr fill:#F38181
Loading

Multi-tenant Architecture

graph TB
subgraph "Tenant A"
A_Config[Tenant A Config<br/>Independent API Key]
A_Instance[LightRAG Instance A<br/>Dedicated LLM/Embedding]
A_Data[(Tenant A Data<br/>Fully Isolated)]
A_Config --> A_Instance
A_Instance --> A_Data
end
subgraph "Tenant B"
B_Config[Tenant B Config<br/>Independent API Key]
B_Instance[LightRAG Instance B<br/>Dedicated LLM/Embedding]
B_Data[(Tenant B Data<br/>Fully Isolated)]
B_Config --> B_Instance
B_Instance --> B_Data
end
subgraph "Tenant C"
C_Config[Using Global Config]
C_Instance[LightRAG Instance C<br/>Shared LLM/Embedding]
C_Data[(Tenant C Data<br/>Fully Isolated)]
C_Config --> C_Instance
C_Instance --> C_Data
end
Pool[Instance Pool Manager<br/>LRU Cache + Config Isolation]
Global[Global Config<br/>Default API Key]
Pool --> A_Instance
Pool --> B_Instance
Pool --> C_Instance
C_Config -.fallback.-> Global
style Pool fill:#F38181
style Global fill:#95E1D3
style A_Config fill:#FFD93D
style B_Config fill:#FFD93D
style C_Config fill:#E8E8E8
Loading

Core Technology Stack

🔧 Frameworks & Runtime

  • FastAPI 0.115+
  • Python 3.11+
  • Uvicorn
  • Docker & Docker Compose

🧠 AI & RAG

  • LightRAG 1.4.9.4
  • RAG-Anything
  • MinerU (PDF-Extract-Kit)
  • Docling

💾 Storage & Database

  • DragonflyDB(Redis compatible)
  • Qdrant(Vector Database)
  • Memgraph(Graph Database)
  • Local filesystem

🚀 Quick Start

Option 1: One-click Deployment (Recommended)

Suitable for production and testing environments:

# 1. Clone the project
git clone https://github.com/BukeLy/rag-api.git
cd rag-api
# 2. Configure environment variables
cp env.example .env
nano .env # Fill in your API keys# 3. Run deployment script
chmod +x deploy.sh
./deploy.sh
# Select deployment mode:# 1) Production Mode - Standard container deployment# 2) Development Mode - Code hot-reload# 4. Verify service
curl http://localhost:8000/

Access Swagger Documentation:http://localhost:8000/docs

Option 2: Docker Compose

Production Mode

# Configure environment variables
cp env.example .env
nano .env
# Start services
docker compose -f docker-compose.yml up -d
# View logs
docker compose -f docker-compose.yml logs -f

Development Mode (Code Hot-reload)

# Start development environment
docker compose -f docker-compose.dev.yml up -d
# Or use quick script
./scripts/dev.sh
# Code changes will auto-reload without restart

Option 3: Local Development

# Install uv (Python package manager)
curl -LsSf https://astral.sh/uv/install.sh | sh
# Install dependencies
uv sync
# Configure environment variables
cp env.example .env
nano .env
# Start services
uv run uvicorn main:app --host 0.0.0.0 --port 8000 --reload

Environment Variable Configuration

Minimum configuration (required):

# LLM Configuration (Function-oriented naming)
LLM_API_KEY=your_llm_api_key
LLM_BASE_URL=https://ark.cn-beijing.volces.com/api/v3
LLM_MODEL=ep-xxx-xxx
# LLM_REQUESTS_PER_MINUTE=800 # Rate limit (optional)# LLM_TOKENS_PER_MINUTE=40000 # Rate limit (optional)# LLM_MAX_ASYNC=8 # [Optional, expert mode] Manual concurrency control# # Auto-calculated when unset: min(RPM, TPM/3500) = 11# Embedding Configuration (Function-oriented naming)
EMBEDDING_API_KEY=your_embedding_api_key
EMBEDDING_BASE_URL=https://api.siliconflow.cn/v1
EMBEDDING_MODEL=Qwen/Qwen3-Embedding-0.6B
EMBEDDING_DIM=1024
# EMBEDDING_MAX_ASYNC=32 # [Optional, expert mode] Auto-calculated when unset: 800# MinerU Mode (Remote recommended)
MINERU_MODE=remote
MINERU_API_TOKEN=your_token
MINERU_HTTP_TIMEOUT=60 # MinerU download timeout (seconds, default 60)
FILE_SERVICE_BASE_URL=http://your-ip:8000
# VLM Chart Enhancement Configuration 🆕# ⚠️ Note: Only effective in MINERU_MODE=remote
RAG_VLM_MODE=off # off / selective / full
RAG_IMPORTANCE_THRESHOLD=0.5 # Importance threshold (selective mode)
RAG_CONTEXT_WINDOW=2 # Context window (full mode)
RAG_CONTEXT_MODE=page # page / chunk
RAG_MAX_CONTEXT_TOKENS=3000 # Max context tokens# Task Storage Configuration 🆕
TASK_STORE_STORAGE=redis # memory / redis (production recommends redis)# Document Insert Verification Configuration 🆕
DOC_INSERT_VERIFICATION_TIMEOUT=300 # Verification timeout (seconds, default 5 minutes)
DOC_INSERT_VERIFICATION_POLL_INTERVAL=0.5 # Poll interval (seconds, default 500ms)# Model Call Timeout Configuration 🆕
MODEL_CALL_TIMEOUT=90 # Model call max timeout (seconds, default 90)

⚡ Auto Concurrency Calculation:

  • LLM: When LLM_MAX_ASYNC is unset, auto-calculated as min(RPM, TPM/3500) ≈ 11
  • Embedding: When EMBEDDING_MAX_ASYNC is unset, auto-calculated as min(RPM, TPM/500) ≈ 800
  • Rerank: When RERANK_MAX_ASYNC is unset, auto-calculated as min(RPM, TPM/500) ≈ 800

✅ Recommended: Don't set *_MAX_ASYNC, let the system auto-calculate to completely avoid 429 errors

See env.example for complete configuration.


📚 API Documentation

Core Endpoints

1️⃣ Upload Document

# Single file upload (default mode)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc1" \
-F "file=@document.pdf" \
-F "parser=auto"# VLM chart enhancement mode 🆕# off: Markdown only (fastest, default)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc2&vlm_mode=off" \
-F "file=@document.pdf"# selective: Selective processing of important charts (balance performance and quality)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc3&vlm_mode=selective" \
-F "file=@document.pdf"# full: Complete RAG-Anything processing (highest quality, context enhancement enabled)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc4&vlm_mode=full" \
-F "file=@document.pdf"# Response
{
"task_id": "task-xxx-xxx",
"doc_id": "doc1",
"filename": "document.pdf",
"vlm_mode": "off",
"status": "pending"
}

2️⃣ Batch Upload

curl -X POST "http://localhost:8000/batch?tenant_id=your_tenant" \
-F "files=@doc1.pdf" \
-F "files=@doc2.docx" \
-F "files=@image.png"# Response
{
"batch_id": "batch-xxx-xxx",
"total_files": 3,
"accepted_files": 3,
"tasks": [...]
}

3️⃣ Intelligent Query (Query API v2.0)

New Advanced Features:

  • Conversation History: Support for multi-turn conversation context
  • Custom Prompts: Customize response style
  • Response Format Control: paragraph/list/json
  • Keyword Precision Retrieval: hl_keywords/ll_keywords
  • Streaming Output: Real-time generation viewing
# Basic query
curl -X POST "http://localhost:8000/query?tenant_id=your_tenant" \
-H "Content-Type: application/json" \
-d '{ "query": "What are the core viewpoints in the document?", "mode": "hybrid" }'# Advanced query (multi-turn dialogue + custom prompt)
curl -X POST "http://localhost:8000/query?tenant_id=your_tenant" \
-H "Content-Type: application/json" \
-d '{ "query": "Can you elaborate on the second point?", "mode": "hybrid", "conversation_history": [ {"role": "user", "content": "What are the key points?"}, {"role": "assistant", "content": "There are mainly three points..."} ], "user_prompt": "Please answer in professional academic language", "response_type": "list" }'# Streaming query (SSE)
curl -N -X POST "http://localhost:8000/query/stream?tenant_id=your_tenant" \
-H "Content-Type: application/json" \
-d '{ "query": "What are the core viewpoints in the document?", "mode": "hybrid" }'# Response (real-time streaming output)
data: {"chunk": "Based on", "done": false}
data: {"chunk": "document content", "done": false}
data: {"done": true}

4️⃣ Task Status Query

curl "http://localhost:8000/task/task-xxx-xxx?tenant_id=your_tenant"# Response
{
"task_id": "task-xxx-xxx",
"status": "completed",
"progress": 100,
"result": {...}
}

5️⃣ Tenant Management

# Get tenant statistics
curl "http://localhost:8000/tenants/stats?tenant_id=your_tenant"# Clear tenant cache
curl -X DELETE "http://localhost:8000/tenants/cache?tenant_id=your_tenant"# View instance pool status (admin)
curl "http://localhost:8000/tenants/pool/stats"

VLM Mode Comparison 🆕

ModeSpeedQualityResource UsageUse Case
off⚡⚡⚡⚡⚡⭐⭐⭐Very LowPlain text documents, fast batch processing
selective⚡⚡⚡⚡⭐⭐⭐⭐LowDocuments with key charts (recommended)
full⚡⚡⭐⭐⭐⭐⭐HighChart-intensive research reports, papers

Processing Time Estimate (20-page PDF example):

  • off: ~10 seconds(Markdown only)
  • selective: ~30 seconds(5-10 important charts)
  • full: ~120 seconds(complete context processing)

Query Mode Comparison

ModeSpeedAccuracyUse Case
naive⚡⚡⚡⚡⚡⭐⭐⭐Simple Q&A, fast retrieval
local⚡⚡⚡⚡⭐⭐⭐⭐Local entity relationship queries
global⚡⚡⚡⭐⭐⭐⭐Global knowledge graph reasoning
hybrid⚡⚡⚡⭐⭐⭐⭐⭐Hybrid retrieval (recommended)
mix⚡⚡⭐⭐⭐⭐⭐Complex questions, deep analysis

Query API v2.0 Advanced Parameters

ParameterTypeDescriptionExample
conversation_historyList[Dict]Multi-turn conversation context[{"role": "user", "content": "..."}]
user_promptstrCustom prompt"Please answer in professional academic language"
response_typestrResponse format"paragraph", "list", "json"
hl_keywordsList[str]High priority keywords["artificial intelligence", "machine learning"]
ll_keywordsList[str]Low priority keywords["application", "case study"]
only_need_contextboolReturn context only (debug)true
max_entity_tokensintEntity token limit6000

Complete API documentation:http://localhost:8000/docs


🎯 Usage Examples

Python SDK

importrequests# ConfigurationBASE_URL="http://localhost:8000"TENANT_ID="your_tenant"# Upload documentwithopen("document.pdf", "rb") asf:
response=requests.post(
f"{BASE_URL}/insert",
params={"tenant_id": TENANT_ID, "doc_id": "doc1"},
files={"file": f}
)
task_id=response.json()["task_id"]
print(f"Task ID: {task_id}")
# Queryresponse=requests.post(
f"{BASE_URL}/query",
params={"tenant_id": TENANT_ID},
json={
"query": "What is the main content of the document?",
"mode": "hybrid",
"top_k": 10
}
)
result=response.json()
print(f"Answer: {result['answer']}")

Complete cURL Example

# 1. Upload PDF document
TASK_ID=$(curl -X POST "http://localhost:8000/insert?tenant_id=demo&doc_id=report" \ -F "file=@report.pdf"| jq -r '.task_id')echo"Task ID: $TASK_ID"# 2. Wait for processing completionwhiletrue;do
STATUS=$(curl -s "http://localhost:8000/task/$TASK_ID?tenant_id=demo"| jq -r '.status')echo"Status: $STATUS"if [ "$STATUS"="completed" ] || [ "$STATUS"="failed" ];thenbreakfi
sleep 2
done# 3. Query document content
curl -X POST "http://localhost:8000/query?tenant_id=demo" \
-H "Content-Type: application/json" \
-d '{ "query": "What are the main conclusions of this report?", "mode": "hybrid" }'| jq '.answer'

🛠️ Deployment

System Requirements

Minimum Configuration:

  • CPU: 2 cores
  • RAM: 4GB
  • Disk: 40GB SSD
  • OS: Ubuntu 20.04+ / Debian 11+ / CentOS 8+

Recommended Configuration (Production):

  • CPU: 4 cores
  • RAM: 8GB
  • Disk: 100GB SSD
  • OS: Ubuntu 22.04 LTS

Server Deployment

Quick Deployment on Aliyun/Tencent Cloud

# SSH login to server
ssh root@your-server-ip
# Clone project
git clone https://github.com/BukeLy/rag-api.git
cd rag-api
# Run one-click deployment script
chmod +x deploy.sh
./deploy.sh
# The script will automatically:# 1. Install Docker and Docker Compose# 2. Configure environment variables# 3. Optimize system parameters# 4. Start services# 5. Verify health status

External Storage Configuration

Supports DragonflyDB + Qdrant + Memgraph external storage (enabled by default):

# Configure in .env
USE_EXTERNAL_STORAGE=true
# DragonflyDB configuration (KV Storage)
KV_STORAGE=RedisKVStorage
REDIS_URI=redis://dragonflydb:6379/0
# Qdrant configuration (vector storage)
VECTOR_STORAGE=QdrantVectorDBStorage
QDRANT_URL=http://qdrant:6333
# Memgraph configuration (graph storage)
GRAPH_STORAGE=MemgraphStorage
MEMGRAPH_URI=bolt://memgraph:7687
MEMGRAPH_USERNAME=
MEMGRAPH_PASSWORD=

See External Storage Deployment Documentation

Docker Compose Configuration

The project provides two configuration files:

FilePurposeFeatures
docker-compose.ymlProduction modeCode packaged in image, optimal performance
docker-compose.dev.ymlDevelopment modeCode mounted externally, supports hot-reload

Select configuration file:

# Production mode
docker compose -f docker-compose.yml up -d
# Development mode
docker compose -f docker-compose.dev.yml up -d

Performance Optimization

Tuning Parameters

Configure in .env:

# ⚡ Concurrency Control (Recommended: use auto-calculation)# LLM_MAX_ASYNC=8 # [Expert mode] Manually specify LLM concurrency# # Auto-calculated when unset: min(RPM, TPM/3500) ≈ 11# EMBEDDING_MAX_ASYNC=32 # [Expert mode] Manually specify Embedding concurrency# # Auto-calculated when unset: min(RPM, TPM/500) ≈ 800# RERANK_MAX_ASYNC=16 # [Expert mode] Manually specify Rerank concurrency# # Auto-calculated when unset: min(RPM, TPM/500) ≈ 800# Retrieval count (affects query quality and speed)
TOP_K=20 # Entity/relationship retrieval count
CHUNK_TOP_K=10 # Text chunk retrieval count# Document processing concurrency
DOCUMENT_PROCESSING_CONCURRENCY=10 # Remote mode can be set high, local mode set to 1

🎯 Concurrency Configuration Recommendations:

  • Recommended: Don't set *_MAX_ASYNC, let the system auto-calculate based on TPM/RPM
  • Expert mode: If manual control needed, can set LLM_MAX_ASYNC and other parameters
  • Advantage: Auto-calculation completely avoids 429 errors (TPM limit reached)

Mode Selection

  • MinerU Remote Mode (Recommended): High concurrency, resource-efficient
  • MinerU Local Mode: Requires GPU, high memory usage
  • Docling Mode: Fast and lightweight, suitable for simple documents

🏢 Multi-tenant Usage

Tenant Isolation

Each tenant has:

  • ✅ Independent LightRAG instance
  • ✅ Isolated data storage space
  • ✅ Independent vector index
  • ✅ Dedicated knowledge graph
  • Independent service configuration (LLM, Embedding, Rerank, DeepSeek-OCR, MinerU)🆕

Tenant Configuration Management 🆕

Each tenant can independently configure 5 services with hot-reload support:

# 1️⃣ Configure independent DeepSeek-OCR API key for Tenant A
curl -X PUT "http://localhost:8000/tenants/tenant_a/config" \
-H "Content-Type: application/json" \
-d '{ "ds_ocr_config": { "api_key": "sk-tenant-a-ds-ocr-key", "base_url": "https://api.siliconflow.cn/v1", "model": "deepseek-ai/DeepSeek-OCR", "timeout": 90 } }'# 2️⃣ Configure independent MinerU API token for Tenant B
curl -X PUT "http://localhost:8000/tenants/tenant_b/config" \
-H "Content-Type: application/json" \
-d '{ "mineru_config": { "api_token": "tenant-b-mineru-token", "base_url": "https://mineru.net", "model_version": "vlm" } }'# 3️⃣ Configure multiple services simultaneously (LLM + Embedding + DeepSeek-OCR)
curl -X PUT "http://localhost:8000/tenants/tenant_c/config" \
-H "Content-Type: application/json" \
-d '{ "llm_config": { "api_key": "sk-tenant-c-llm-key", "model": "gpt-4" }, "embedding_config": { "api_key": "sk-tenant-c-embedding-key", "model": "Qwen/Qwen3-Embedding-0.6B", "dim": 1024 }, "ds_ocr_config": { "api_key": "sk-tenant-c-ds-ocr-key" } }'# 4️⃣ Query tenant configuration (API key auto-masked)
curl "http://localhost:8000/tenants/tenant_a/config"# Response example
{
"tenant_id": "tenant_a",
"ds_ocr_config": {
"api_key": "sk-***-key", // Auto-masked
"timeout": 90
},
"merged_config": {
"llm": {...}, // Using Global Config
"embedding": {...}, // Using Global Config
"rerank": {...}, // Using Global Config
"ds_ocr": {...}, // Using tenant config
"mineru": {...} // Using Global Config
}
}
# 5️⃣ Refresh config cache (config hot-reload)
curl -X POST "http://localhost:8000/tenants/tenant_a/config/refresh"# 6️⃣ Delete tenant config (restore to global config)
curl -X DELETE "http://localhost:8000/tenants/tenant_a/config"

Supported Configuration Items:

ServiceConfig FieldDescription
LLMllm_configModel, API key, base_url, etc.
Embeddingembedding_configModel, API key, dimension, etc.
Rerankrerank_configModel, API key, etc.
DeepSeek-OCRds_ocr_configAPI key, timeout, mode, etc.
MinerUmineru_configAPI token, version, timeout, etc.

Configuration Priority: Tenant config > Global config

Use Cases:

  • 🔐 Multi-tenant SaaS: Each tenant uses their own API key
  • 💰 Pay-per-use: Track tenant usage through independent API keys
  • 🎯 Differentiated Services: Different tenants use different models (GPT-4 vs GPT-3.5)
  • 🧪 A/B Testing: Compare different models/parameters

Usage

All APIs require tenant_id parameter:

# Tenant A upload document
curl -X POST "http://localhost:8000/insert?tenant_id=tenant_a&doc_id=doc1" \
-F "file=@doc.pdf"# Tenant B upload document (fully isolated)
curl -X POST "http://localhost:8000/insert?tenant_id=tenant_b&doc_id=doc1" \
-F "file=@doc.pdf"# Tenant A query (can only query own documents)
curl -X POST "http://localhost:8000/query?tenant_id=tenant_a" \
-H "Content-Type: application/json" \
-d '{"query": "document content", "mode": "hybrid"}'

Instance Pool Management

  • Capacity: Cache up to 50 tenant instances
  • Strategy: LRU (Least Recently Used) automatic cleanup
  • Config Isolation: Each tenant can use independent LLM, Embedding, parser configuration

📊 Monitoring & Maintenance

Common Commands

# View service status
docker compose ps
# View real-time logs
docker compose logs -f
# Restart services
docker compose restart
# Stop services
docker compose down
# View resource usage
docker stats
# Clean Docker resources
docker system prune -f

Maintenance Scripts

# Monitor service health
./scripts/monitor.sh
# Backup data
./scripts/backup.sh
# Update services
./scripts/update.sh
# Performance testing
./scripts/test_concurrent_perf.sh
# Performance monitoring
./scripts/monitor_performance.sh

Health Checks

# Complete health check (recommended)
./scripts/health_check.sh
./scripts/health_check.sh --verbose # verbose output# API health check
curl http://localhost:8000/
# Tenant statistics
curl "http://localhost:8000/tenants/stats?tenant_id=your_tenant"# Instance pool status
curl "http://localhost:8000/tenants/pool/stats"

🗂️ Project Structure

rag-api/
├── main.py # FastAPI application entry
├── api/ # API route modules
│ ├── __init__.py # Route aggregation
│ ├── insert.py # Document upload (single/batch)
│ ├── query.py # Intelligent query
│ ├── task.py # Task status query
│ ├── tenant.py # Tenant management
│ ├── files.py # File service
│ ├── models.py # Pydantic models
│ └── task_store.py # Task storage
├── src/ # Core business logic
│ ├── rag.py # LightRAG lifecycle management
│ ├── multi_tenant.py # Multi-tenant instance manager
│ ├── tenant_deps.py # Tenant dependency injection
│ ├── logger.py # Unified logging
│ ├── metrics.py # Performance metrics
│ ├── file_url_service.py # Temporary file service
│ ├── mineru_client.py # MinerU client
│ └── mineru_result_processor.py # Result processing
├── docs/ # Documentation
│ ├── ARCHITECTURE.md # Architecture design documentation
│ ├── USAGE.md # Detailed usage guide
│ ├── DEPLOY_MODES.md # Deployment mode description
│ ├── PR_WORKFLOW.md # PR workflow
│ └── ...
├── scripts/ # Maintenance scripts
│ ├── dev.sh # Development mode quick start
│ ├── monitor.sh # Service monitoring
│ ├── backup.sh # Data backup
│ ├── update.sh # Service update
│ └── ...
├── deploy.sh # One-click deployment script
├── docker-compose.yml # Production mode configuration
├── docker-compose.dev.yml # Development mode configuration
├── Dockerfile # Production image
├── Dockerfile.dev # Development image
├── pyproject.toml # Project dependencies
├── uv.lock # Dependency lock
├── env.example # Environment variable template
├── CLAUDE.md # Claude AI guide
└── README.md # This documentation

🐛 Troubleshooting

Common Issues

Q1: What to do if service fails to start?
# View detailed logs
docker compose logs
# Check port usage
netstat -tulpn | grep 8000
# Check Docker status
docker ps -a
Q2: multimodal_processed error?

Note: This issue has been fixed in LightRAG 1.4.9.4+. If you encounter this error, your version is outdated.

Solution:

# Option 1: Upgrade to latest version (recommended)# Modify LightRAG version in pyproject.toml# lightrag = "^1.4.9.4"# Rebuild image
docker compose down
docker compose up -d --build
# Option 2: Clean old data (temporary solution)
rm -rf ./rag_local_storage
docker compose restart
Q3: File upload returns 400 error?

Check:

  • File format supported (PDF, DOCX, PNG, JPG, etc.)
  • File size exceeds 100MB
  • File is empty
# View supported formats
curl http://localhost:8000/docs
Q3.5: Embedding dimension error?

If you encounter dimension-related errors, need to clean data and rebuild:

# Stop services
docker compose down
# Delete all volumes (clear database)
docker volume rm rag-api_dragonflydb_data rag-api_qdrant_data rag-api_memgraph_data
# Modify EMBEDDING_DIM in .env
EMBEDDING_DIM=1024 # or 4096, must match the model# Restart
docker compose up -d
Q4: Query is very slow (>30 seconds)?

Optimization suggestions:

  1. Use naive or hybrid mode instead of mix
  2. Increase MAX_ASYNC parameter (in .env)
  3. Reduce TOP_K and CHUNK_TOP_K
  4. Enable Reranker
# Modify .env
MAX_ASYNC=8
TOP_K=20
CHUNK_TOP_K=10
Q5: Out of memory (OOM)?

If using local MinerU:

# Switch to remote mode# Modify in .env
MINERU_MODE=remote
MINERU_API_TOKEN=your_token
# Or limit concurrency
DOCUMENT_PROCESSING_CONCURRENCY=1
Q6: Tasks lost after container restart?

Problem Symptoms:

  • Cannot query previous task status after container restart
  • Tasks disappear after tenant instance evicted by LRU

Solution: Enable Redis task storage

# Modify .env
TASK_STORE_STORAGE=redis
# Restart services
docker compose restart
# Verify
docker compose logs api | grep TaskStore
# Should see: ✅ TaskStore: Redis connection successful

Configuration Description:

  • memory mode: In-memory storage, data lost after restart (default, suitable for development)
  • redis mode: Persistent storage, supports container restart and instance rebuild (production recommended)

TTL Strategy (Redis mode auto-cleanup):

  • completed tasks: 24 hours
  • failed tasks: 24 hours
  • pending/processing tasks: 6 hours
Q7: VLM mode processing failed?

Check Items:

  1. vision_model_func not configured

    • Check logs:vision_model_func not found, fallback to off mode
    • Ensure LLM API is configured in .env
  2. Image file does not exist

    • Check logs:Image file not found: xxx
    • Possibly corrupted MinerU ZIP or extraction failed
  3. Timeout error

    • full mode may timeout on large files
    • Suggestion: Use selective mode first, or increase VLM_TIMEOUT
# Modify .env
VLM_TIMEOUT=300 # Increase to 5 minutes
RAG_VLM_MODE=selective # downgrade to selective

Debugging Tips:

# View detailed logs
docker compose logs -f | grep VLM
# Test single file
curl -X POST 'http://localhost:8000/insert?tenant_id=test&doc_id=test&vlm_mode=off' \
-F 'file=@test.pdf'

Performance Tuning Recommendations

ScenarioMAX_ASYNCTOP_KCHUNK_TOP_KMINERU_MODE
Fast response8105remote
Balanced mode82010remote
High accuracy46020remote
Resource limited42010remote

📖 Documentation


🤝 Contributing

We welcome all forms of contribution!

How to Contribute

  1. Fork the project
git clone https://github.com/BukeLy/rag-api.git
cd rag-api
  1. Create feature branch
git checkout -b feature/your-feature-name
  1. Development and Testing
# Install dependencies
uv sync
# Run tests
uv run pytest
# Code formatting
uv run black .
uv run isort .
  1. Submit code
git add .
git commit -m "feat: Add new feature"
git push origin feature/your-feature-name
  1. Create Pull Request

Create a PR on GitHub with detailed description of your changes.

Commit Conventions

Use semantic commit messages:

  • feat: New feature
  • fix: Bug fix
  • docs: Documentation update
  • style: Code formatting
  • refactor: Code refactoring
  • perf: Performance optimization
  • test: Testing
  • chore: Build/tools

See PR Workflow Documentation


📄 License

This project is licensed under the MIT License. See the LICENSE file for details.


🙏 Acknowledgments

This project is built on the following excellent open source projects:

  • LightRAG - Efficient knowledge graph RAG framework
  • RAG-Anything - Multimodal document parsing
  • MinerU - Powerful PDF parsing tool
  • Docling - Lightweight document parsing
  • FastAPI - Modern Python web framework

Special thanks to all contributors and users for their support! 🎉


📬 Contact Us


⭐ If this project helps you, please give it a Star!

Made with ❤️ by BukeLy

© 2025 RAG API. All rights reserved.

About

Multi-tenant RAG API powered by LightRAG/RAG-Anything. Auto-selects best parser (DeepSeek-OCR/MinerU/Docling) via complexity scoring

Topics

Resources

Stars

64 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

🚀 RAG API

Multi-tenant Multimodal Document Intelligent Retrieval System

Enterprise-grade RAG service built on RAG-Anything and LightRAG

CIPythonFastAPILightRAGDockerLicense

English | 简体中文

FeaturesQuick StartArchitectureAPI DocumentationDeployment


📖 Introduction

RAG API is an enterprise-grade Retrieval-Augmented Generation (RAG) service that combines the powerful document parsing capabilities of RAG-Anything with the efficient knowledge graph retrieval technology of LightRAG, providing intelligent Q&A capabilities for your documents.

🎯 Key Highlights

  • 🏢 Multi-tenant Isolation - Complete tenant data isolation for enterprise multi-tenant scenarios
  • 🎨 Multimodal Parsing - Support for PDF, Word, images and more, with full OCR, tables, and formulas coverage
  • High-performance Retrieval - Knowledge graph-based hybrid retrieval with 6-15 second query response
  • 🔄 Flexible Deployment - Support for production and development modes with one-click switching
  • 📦 Ready to Use - One-click Docker deployment, service starts in 3 minutes
  • 🎛️ Multiple Parsing Engines - DeepSeek-OCR (Remote API) + MinerU (Local/Remote API) + Docling (Fast)
  • 🎨 RAG-Anything VLM Enhancement - Three modes (off/selective/full) for deep chart understanding
  • 💾 Task Persistence - Redis storage support, tasks recoverable after container restart/instance rebuild

✨ Features

📄 Document Processing

  • Multiple Format Support

    • PDF, Word, Excel, PPT
    • PNG, JPG, WebP images
    • TXT, Markdown text
  • Intelligent Parsing

    • Plain text (.txt, .md) → Direct insertion (ultra-fast ~1s, skip parser)
    • OCR text recognition
    • Structured table extraction
    • Mathematical formula recognition
    • Layout analysis
  • RAG-Anything VLM Enhancement 🆕

    • off - Markdown only (fastest)
    • selective - Selective processing of important charts
    • full - Complete context enhancement processing
    • Smart filtering: with titles, large size, first page content
    • ⚠️Only supports remote MinerU mode, local mode uses RAG-Anything native methods
  • Batch Processing

    • Up to 100 files per batch
    • Async task queue
    • Real-time progress tracking

🔍 Intelligent Retrieval

  • Multi-mode Query

    • naive - Vector retrieval (fastest)
    • local - Local graph
    • global - Global graph
    • hybrid - Hybrid retrieval
    • mix - Full retrieval (most accurate)
  • Knowledge Graph

    • Automatic entity extraction
    • Relationship reasoning
    • Semantic understanding
    • Context enhancement
  • External Storage

    • DragonflyDB (KV storage + task storage)
    • Qdrant (vector storage)
    • Memgraph (graph database)
    • Task persistence (Redis mode)

🏗️ Architecture

System Architecture Diagram

graph TB
subgraph "Client Layer"
Client[Client Application]
WebUI[Web Interface]
end
subgraph "API Gateway Layer"
FastAPI[FastAPI Service]
Auth[Tenant Authentication]
end
subgraph "Business Logic Layer"
TenantMgr[Tenant Manager]
TaskQueue[Task Queue]
subgraph "Document Processing"
DeepSeekOCR[DeepSeek-OCR<br/>Fast OCR 80% cases]
MinerU[MinerU Parser<br/>Complex multimodal]
Docling[Docling Parser<br/>Fast lightweight]
FileRouter[Smart Router<br/>Complexity scoring]
end
subgraph "RAG Engine"
LightRAG[LightRAG Instance Pool<br/>LRU Cache 50]
KG[Knowledge Graph Engine]
Vector[Vector Retrieval Engine]
end
end
subgraph "Storage Layer"
DragonflyDB[(DragonflyDB<br/>KV Storage)]
Qdrant[(Qdrant<br/>Vector Database)]
Memgraph[(Memgraph<br/>Graph Database)]
Local[(Local Files<br/>Temp Storage)]
end
subgraph "External Services"
LLM[LLM<br/>Entity Extraction/Generation]
Embedding[Embedding<br/>Vectorization]
Rerank[Rerank<br/>Reranking]
end
Client --> FastAPI
WebUI --> FastAPI
FastAPI --> Auth
Auth --> TenantMgr
TenantMgr --> TaskQueue
TenantMgr --> LightRAG
TaskQueue --> FileRouter
FileRouter --> DeepSeekOCR
FileRouter --> MinerU
FileRouter --> Docling
DeepSeekOCR --> LightRAG
MinerU --> LightRAG
Docling --> LightRAG
LightRAG --> KG
LightRAG --> Vector
KG --> DragonflyDB
KG --> Memgraph
Vector --> Qdrant
LightRAG --> Local
LightRAG --> LLM
LightRAG --> Embedding
Vector --> Rerank
style FastAPI fill:#00C7B7
style LightRAG fill:#FF6B6B
style DeepSeekOCR fill:#5DADE2
style MinerU fill:#4ECDC4
style Docling fill:#95E1D3
style TenantMgr fill:#F38181
Loading

Multi-tenant Architecture

graph TB
subgraph "Tenant A"
A_Config[Tenant A Config<br/>Independent API Key]
A_Instance[LightRAG Instance A<br/>Dedicated LLM/Embedding]
A_Data[(Tenant A Data<br/>Fully Isolated)]
A_Config --> A_Instance
A_Instance --> A_Data
end
subgraph "Tenant B"
B_Config[Tenant B Config<br/>Independent API Key]
B_Instance[LightRAG Instance B<br/>Dedicated LLM/Embedding]
B_Data[(Tenant B Data<br/>Fully Isolated)]
B_Config --> B_Instance
B_Instance --> B_Data
end
subgraph "Tenant C"
C_Config[Using Global Config]
C_Instance[LightRAG Instance C<br/>Shared LLM/Embedding]
C_Data[(Tenant C Data<br/>Fully Isolated)]
C_Config --> C_Instance
C_Instance --> C_Data
end
Pool[Instance Pool Manager<br/>LRU Cache + Config Isolation]
Global[Global Config<br/>Default API Key]
Pool --> A_Instance
Pool --> B_Instance
Pool --> C_Instance
C_Config -.fallback.-> Global
style Pool fill:#F38181
style Global fill:#95E1D3
style A_Config fill:#FFD93D
style B_Config fill:#FFD93D
style C_Config fill:#E8E8E8
Loading

Core Technology Stack

🔧 Frameworks & Runtime

  • FastAPI 0.115+
  • Python 3.11+
  • Uvicorn
  • Docker & Docker Compose

🧠 AI & RAG

  • LightRAG 1.4.9.4
  • RAG-Anything
  • MinerU (PDF-Extract-Kit)
  • Docling

💾 Storage & Database

  • DragonflyDB(Redis compatible)
  • Qdrant(Vector Database)
  • Memgraph(Graph Database)
  • Local filesystem

🚀 Quick Start

Option 1: One-click Deployment (Recommended)

Suitable for production and testing environments:

# 1. Clone the project
git clone https://github.com/BukeLy/rag-api.git
cd rag-api
# 2. Configure environment variables
cp env.example .env
nano .env # Fill in your API keys# 3. Run deployment script
chmod +x deploy.sh
./deploy.sh
# Select deployment mode:# 1) Production Mode - Standard container deployment# 2) Development Mode - Code hot-reload# 4. Verify service
curl http://localhost:8000/

Access Swagger Documentation:http://localhost:8000/docs

Option 2: Docker Compose

Production Mode

# Configure environment variables
cp env.example .env
nano .env
# Start services
docker compose -f docker-compose.yml up -d
# View logs
docker compose -f docker-compose.yml logs -f

Development Mode (Code Hot-reload)

# Start development environment
docker compose -f docker-compose.dev.yml up -d
# Or use quick script
./scripts/dev.sh
# Code changes will auto-reload without restart

Option 3: Local Development

# Install uv (Python package manager)
curl -LsSf https://astral.sh/uv/install.sh | sh
# Install dependencies
uv sync
# Configure environment variables
cp env.example .env
nano .env
# Start services
uv run uvicorn main:app --host 0.0.0.0 --port 8000 --reload

Environment Variable Configuration

Minimum configuration (required):

# LLM Configuration (Function-oriented naming)
LLM_API_KEY=your_llm_api_key
LLM_BASE_URL=https://ark.cn-beijing.volces.com/api/v3
LLM_MODEL=ep-xxx-xxx
# LLM_REQUESTS_PER_MINUTE=800 # Rate limit (optional)# LLM_TOKENS_PER_MINUTE=40000 # Rate limit (optional)# LLM_MAX_ASYNC=8 # [Optional, expert mode] Manual concurrency control# # Auto-calculated when unset: min(RPM, TPM/3500) = 11# Embedding Configuration (Function-oriented naming)
EMBEDDING_API_KEY=your_embedding_api_key
EMBEDDING_BASE_URL=https://api.siliconflow.cn/v1
EMBEDDING_MODEL=Qwen/Qwen3-Embedding-0.6B
EMBEDDING_DIM=1024
# EMBEDDING_MAX_ASYNC=32 # [Optional, expert mode] Auto-calculated when unset: 800# MinerU Mode (Remote recommended)
MINERU_MODE=remote
MINERU_API_TOKEN=your_token
MINERU_HTTP_TIMEOUT=60 # MinerU download timeout (seconds, default 60)
FILE_SERVICE_BASE_URL=http://your-ip:8000
# VLM Chart Enhancement Configuration 🆕# ⚠️ Note: Only effective in MINERU_MODE=remote
RAG_VLM_MODE=off # off / selective / full
RAG_IMPORTANCE_THRESHOLD=0.5 # Importance threshold (selective mode)
RAG_CONTEXT_WINDOW=2 # Context window (full mode)
RAG_CONTEXT_MODE=page # page / chunk
RAG_MAX_CONTEXT_TOKENS=3000 # Max context tokens# Task Storage Configuration 🆕
TASK_STORE_STORAGE=redis # memory / redis (production recommends redis)# Document Insert Verification Configuration 🆕
DOC_INSERT_VERIFICATION_TIMEOUT=300 # Verification timeout (seconds, default 5 minutes)
DOC_INSERT_VERIFICATION_POLL_INTERVAL=0.5 # Poll interval (seconds, default 500ms)# Model Call Timeout Configuration 🆕
MODEL_CALL_TIMEOUT=90 # Model call max timeout (seconds, default 90)

⚡ Auto Concurrency Calculation:

  • LLM: When LLM_MAX_ASYNC is unset, auto-calculated as min(RPM, TPM/3500) ≈ 11
  • Embedding: When EMBEDDING_MAX_ASYNC is unset, auto-calculated as min(RPM, TPM/500) ≈ 800
  • Rerank: When RERANK_MAX_ASYNC is unset, auto-calculated as min(RPM, TPM/500) ≈ 800

✅ Recommended: Don't set *_MAX_ASYNC, let the system auto-calculate to completely avoid 429 errors

See env.example for complete configuration.


📚 API Documentation

Core Endpoints

1️⃣ Upload Document

# Single file upload (default mode)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc1" \
-F "file=@document.pdf" \
-F "parser=auto"# VLM chart enhancement mode 🆕# off: Markdown only (fastest, default)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc2&vlm_mode=off" \
-F "file=@document.pdf"# selective: Selective processing of important charts (balance performance and quality)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc3&vlm_mode=selective" \
-F "file=@document.pdf"# full: Complete RAG-Anything processing (highest quality, context enhancement enabled)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc4&vlm_mode=full" \
-F "file=@document.pdf"# Response
{
"task_id": "task-xxx-xxx",
"doc_id": "doc1",
"filename": "document.pdf",
"vlm_mode": "off",
"status": "pending"
}

2️⃣ Batch Upload

curl -X POST "http://localhost:8000/batch?tenant_id=your_tenant" \
-F "files=@doc1.pdf" \
-F "files=@doc2.docx" \
-F "files=@image.png"# Response
{
"batch_id": "batch-xxx-xxx",
"total_files": 3,
"accepted_files": 3,
"tasks": [...]
}

3️⃣ Intelligent Query (Query API v2.0)

New Advanced Features:

  • Conversation History: Support for multi-turn conversation context
  • Custom Prompts: Customize response style
  • Response Format Control: paragraph/list/json
  • Keyword Precision Retrieval: hl_keywords/ll_keywords
  • Streaming Output: Real-time generation viewing
# Basic query
curl -X POST "http://localhost:8000/query?tenant_id=your_tenant" \
-H "Content-Type: application/json" \
-d '{ "query": "What are the core viewpoints in the document?", "mode": "hybrid" }'# Advanced query (multi-turn dialogue + custom prompt)
curl -X POST "http://localhost:8000/query?tenant_id=your_tenant" \
-H "Content-Type: application/json" \
-d '{ "query": "Can you elaborate on the second point?", "mode": "hybrid", "conversation_history": [ {"role": "user", "content": "What are the key points?"}, {"role": "assistant", "content": "There are mainly three points..."} ], "user_prompt": "Please answer in professional academic language", "response_type": "list" }'# Streaming query (SSE)
curl -N -X POST "http://localhost:8000/query/stream?tenant_id=your_tenant" \
-H "Content-Type: application/json" \
-d '{ "query": "What are the core viewpoints in the document?", "mode": "hybrid" }'# Response (real-time streaming output)
data: {"chunk": "Based on", "done": false}
data: {"chunk": "document content", "done": false}
data: {"done": true}

4️⃣ Task Status Query

curl "http://localhost:8000/task/task-xxx-xxx?tenant_id=your_tenant"# Response
{
"task_id": "task-xxx-xxx",
"status": "completed",
"progress": 100,
"result": {...}
}

5️⃣ Tenant Management

# Get tenant statistics
curl "http://localhost:8000/tenants/stats?tenant_id=your_tenant"# Clear tenant cache
curl -X DELETE "http://localhost:8000/tenants/cache?tenant_id=your_tenant"# View instance pool status (admin)
curl "http://localhost:8000/tenants/pool/stats"

VLM Mode Comparison 🆕

ModeSpeedQualityResource UsageUse Case
off⚡⚡⚡⚡⚡⭐⭐⭐Very LowPlain text documents, fast batch processing
selective⚡⚡⚡⚡⭐⭐⭐⭐LowDocuments with key charts (recommended)
full⚡⚡⭐⭐⭐⭐⭐HighChart-intensive research reports, papers

Processing Time Estimate (20-page PDF example):

  • off: ~10 seconds(Markdown only)
  • selective: ~30 seconds(5-10 important charts)
  • full: ~120 seconds(complete context processing)

Query Mode Comparison

ModeSpeedAccuracyUse Case
naive⚡⚡⚡⚡⚡⭐⭐⭐Simple Q&A, fast retrieval
local⚡⚡⚡⚡⭐⭐⭐⭐Local entity relationship queries
global⚡⚡⚡⭐⭐⭐⭐Global knowledge graph reasoning
hybrid⚡⚡⚡⭐⭐⭐⭐⭐Hybrid retrieval (recommended)
mix⚡⚡⭐⭐⭐⭐⭐Complex questions, deep analysis

Query API v2.0 Advanced Parameters

ParameterTypeDescriptionExample
conversation_historyList[Dict]Multi-turn conversation context[{"role": "user", "content": "..."}]
user_promptstrCustom prompt"Please answer in professional academic language"
response_typestrResponse format"paragraph", "list", "json"
hl_keywordsList[str]High priority keywords["artificial intelligence", "machine learning"]
ll_keywordsList[str]Low priority keywords["application", "case study"]
only_need_contextboolReturn context only (debug)true
max_entity_tokensintEntity token limit6000

Complete API documentation:http://localhost:8000/docs


🎯 Usage Examples

Python SDK

importrequests# ConfigurationBASE_URL="http://localhost:8000"TENANT_ID="your_tenant"# Upload documentwithopen("document.pdf", "rb") asf:
response=requests.post(
f"{BASE_URL}/insert",
params={"tenant_id": TENANT_ID, "doc_id": "doc1"},
files={"file": f}
)
task_id=response.json()["task_id"]
print(f"Task ID: {task_id}")
# Queryresponse=requests.post(
f"{BASE_URL}/query",
params={"tenant_id": TENANT_ID},
json={
"query": "What is the main content of the document?",
"mode": "hybrid",
"top_k": 10
}
)
result=response.json()
print(f"Answer: {result['answer']}")

Complete cURL Example

# 1. Upload PDF document
TASK_ID=$(curl -X POST "http://localhost:8000/insert?tenant_id=demo&doc_id=report" \ -F "file=@report.pdf"| jq -r '.task_id')echo"Task ID: $TASK_ID"# 2. Wait for processing completionwhiletrue;do
STATUS=$(curl -s "http://localhost:8000/task/$TASK_ID?tenant_id=demo"| jq -r '.status')echo"Status: $STATUS"if [ "$STATUS"="completed" ] || [ "$STATUS"="failed" ];thenbreakfi
sleep 2
done# 3. Query document content
curl -X POST "http://localhost:8000/query?tenant_id=demo" \
-H "Content-Type: application/json" \
-d '{ "query": "What are the main conclusions of this report?", "mode": "hybrid" }'| jq '.answer'

🛠️ Deployment

System Requirements

Minimum Configuration:

  • CPU: 2 cores
  • RAM: 4GB
  • Disk: 40GB SSD
  • OS: Ubuntu 20.04+ / Debian 11+ / CentOS 8+

Recommended Configuration (Production):

  • CPU: 4 cores
  • RAM: 8GB
  • Disk: 100GB SSD
  • OS: Ubuntu 22.04 LTS

Server Deployment

Quick Deployment on Aliyun/Tencent Cloud

# SSH login to server
ssh root@your-server-ip
# Clone project
git clone https://github.com/BukeLy/rag-api.git
cd rag-api
# Run one-click deployment script
chmod +x deploy.sh
./deploy.sh
# The script will automatically:# 1. Install Docker and Docker Compose# 2. Configure environment variables# 3. Optimize system parameters# 4. Start services# 5. Verify health status

External Storage Configuration

Supports DragonflyDB + Qdrant + Memgraph external storage (enabled by default):

# Configure in .env
USE_EXTERNAL_STORAGE=true
# DragonflyDB configuration (KV Storage)
KV_STORAGE=RedisKVStorage
REDIS_URI=redis://dragonflydb:6379/0
# Qdrant configuration (vector storage)
VECTOR_STORAGE=QdrantVectorDBStorage
QDRANT_URL=http://qdrant:6333
# Memgraph configuration (graph storage)
GRAPH_STORAGE=MemgraphStorage
MEMGRAPH_URI=bolt://memgraph:7687
MEMGRAPH_USERNAME=
MEMGRAPH_PASSWORD=

See External Storage Deployment Documentation

Docker Compose Configuration

The project provides two configuration files:

FilePurposeFeatures
docker-compose.ymlProduction modeCode packaged in image, optimal performance
docker-compose.dev.ymlDevelopment modeCode mounted externally, supports hot-reload

Select configuration file:

# Production mode
docker compose -f docker-compose.yml up -d
# Development mode
docker compose -f docker-compose.dev.yml up -d

Performance Optimization

Tuning Parameters

Configure in .env:

# ⚡ Concurrency Control (Recommended: use auto-calculation)# LLM_MAX_ASYNC=8 # [Expert mode] Manually specify LLM concurrency# # Auto-calculated when unset: min(RPM, TPM/3500) ≈ 11# EMBEDDING_MAX_ASYNC=32 # [Expert mode] Manually specify Embedding concurrency# # Auto-calculated when unset: min(RPM, TPM/500) ≈ 800# RERANK_MAX_ASYNC=16 # [Expert mode] Manually specify Rerank concurrency# # Auto-calculated when unset: min(RPM, TPM/500) ≈ 800# Retrieval count (affects query quality and speed)
TOP_K=20 # Entity/relationship retrieval count
CHUNK_TOP_K=10 # Text chunk retrieval count# Document processing concurrency
DOCUMENT_PROCESSING_CONCURRENCY=10 # Remote mode can be set high, local mode set to 1

🎯 Concurrency Configuration Recommendations:

  • Recommended: Don't set *_MAX_ASYNC, let the system auto-calculate based on TPM/RPM
  • Expert mode: If manual control needed, can set LLM_MAX_ASYNC and other parameters
  • Advantage: Auto-calculation completely avoids 429 errors (TPM limit reached)

Mode Selection

  • MinerU Remote Mode (Recommended): High concurrency, resource-efficient
  • MinerU Local Mode: Requires GPU, high memory usage
  • Docling Mode: Fast and lightweight, suitable for simple documents

🏢 Multi-tenant Usage

Tenant Isolation

Each tenant has:

  • ✅ Independent LightRAG instance
  • ✅ Isolated data storage space
  • ✅ Independent vector index
  • ✅ Dedicated knowledge graph
  • Independent service configuration (LLM, Embedding, Rerank, DeepSeek-OCR, MinerU)🆕

Tenant Configuration Management 🆕

Each tenant can independently configure 5 services with hot-reload support:

# 1️⃣ Configure independent DeepSeek-OCR API key for Tenant A
curl -X PUT "http://localhost:8000/tenants/tenant_a/config" \
-H "Content-Type: application/json" \
-d '{ "ds_ocr_config": { "api_key": "sk-tenant-a-ds-ocr-key", "base_url": "https://api.siliconflow.cn/v1", "model": "deepseek-ai/DeepSeek-OCR", "timeout": 90 } }'# 2️⃣ Configure independent MinerU API token for Tenant B
curl -X PUT "http://localhost:8000/tenants/tenant_b/config" \
-H "Content-Type: application/json" \
-d '{ "mineru_config": { "api_token": "tenant-b-mineru-token", "base_url": "https://mineru.net", "model_version": "vlm" } }'# 3️⃣ Configure multiple services simultaneously (LLM + Embedding + DeepSeek-OCR)
curl -X PUT "http://localhost:8000/tenants/tenant_c/config" \
-H "Content-Type: application/json" \
-d '{ "llm_config": { "api_key": "sk-tenant-c-llm-key", "model": "gpt-4" }, "embedding_config": { "api_key": "sk-tenant-c-embedding-key", "model": "Qwen/Qwen3-Embedding-0.6B", "dim": 1024 }, "ds_ocr_config": { "api_key": "sk-tenant-c-ds-ocr-key" } }'# 4️⃣ Query tenant configuration (API key auto-masked)
curl "http://localhost:8000/tenants/tenant_a/config"# Response example
{
"tenant_id": "tenant_a",
"ds_ocr_config": {
"api_key": "sk-***-key", // Auto-masked
"timeout": 90
},
"merged_config": {
"llm": {...}, // Using Global Config
"embedding": {...}, // Using Global Config
"rerank": {...}, // Using Global Config
"ds_ocr": {...}, // Using tenant config
"mineru": {...} // Using Global Config
}
}
# 5️⃣ Refresh config cache (config hot-reload)
curl -X POST "http://localhost:8000/tenants/tenant_a/config/refresh"# 6️⃣ Delete tenant config (restore to global config)
curl -X DELETE "http://localhost:8000/tenants/tenant_a/config"

Supported Configuration Items:

ServiceConfig FieldDescription
LLMllm_configModel, API key, base_url, etc.
Embeddingembedding_configModel, API key, dimension, etc.
Rerankrerank_configModel, API key, etc.
DeepSeek-OCRds_ocr_configAPI key, timeout, mode, etc.
MinerUmineru_configAPI token, version, timeout, etc.

Configuration Priority: Tenant config > Global config

Use Cases:

  • 🔐 Multi-tenant SaaS: Each tenant uses their own API key
  • 💰 Pay-per-use: Track tenant usage through independent API keys
  • 🎯 Differentiated Services: Different tenants use different models (GPT-4 vs GPT-3.5)
  • 🧪 A/B Testing: Compare different models/parameters

Usage

All APIs require tenant_id parameter:

# Tenant A upload document
curl -X POST "http://localhost:8000/insert?tenant_id=tenant_a&doc_id=doc1" \
-F "file=@doc.pdf"# Tenant B upload document (fully isolated)
curl -X POST "http://localhost:8000/insert?tenant_id=tenant_b&doc_id=doc1" \
-F "file=@doc.pdf"# Tenant A query (can only query own documents)
curl -X POST "http://localhost:8000/query?tenant_id=tenant_a" \
-H "Content-Type: application/json" \
-d '{"query": "document content", "mode": "hybrid"}'

Instance Pool Management

  • Capacity: Cache up to 50 tenant instances
  • Strategy: LRU (Least Recently Used) automatic cleanup
  • Config Isolation: Each tenant can use independent LLM, Embedding, parser configuration

📊 Monitoring & Maintenance

Common Commands

# View service status
docker compose ps
# View real-time logs
docker compose logs -f
# Restart services
docker compose restart
# Stop services
docker compose down
# View resource usage
docker stats
# Clean Docker resources
docker system prune -f

Maintenance Scripts

# Monitor service health
./scripts/monitor.sh
# Backup data
./scripts/backup.sh
# Update services
./scripts/update.sh
# Performance testing
./scripts/test_concurrent_perf.sh
# Performance monitoring
./scripts/monitor_performance.sh

Health Checks

# Complete health check (recommended)
./scripts/health_check.sh
./scripts/health_check.sh --verbose # verbose output# API health check
curl http://localhost:8000/
# Tenant statistics
curl "http://localhost:8000/tenants/stats?tenant_id=your_tenant"# Instance pool status
curl "http://localhost:8000/tenants/pool/stats"

🗂️ Project Structure

rag-api/
├── main.py # FastAPI application entry
├── api/ # API route modules
│ ├── __init__.py # Route aggregation
│ ├── insert.py # Document upload (single/batch)
│ ├── query.py # Intelligent query
│ ├── task.py # Task status query
│ ├── tenant.py # Tenant management
│ ├── files.py # File service
│ ├── models.py # Pydantic models
│ └── task_store.py # Task storage
├── src/ # Core business logic
│ ├── rag.py # LightRAG lifecycle management
│ ├── multi_tenant.py # Multi-tenant instance manager
│ ├── tenant_deps.py # Tenant dependency injection
│ ├── logger.py # Unified logging
│ ├── metrics.py # Performance metrics
│ ├── file_url_service.py # Temporary file service
│ ├── mineru_client.py # MinerU client
│ └── mineru_result_processor.py # Result processing
├── docs/ # Documentation
│ ├── ARCHITECTURE.md # Architecture design documentation
│ ├── USAGE.md # Detailed usage guide
│ ├── DEPLOY_MODES.md # Deployment mode description
│ ├── PR_WORKFLOW.md # PR workflow
│ └── ...
├── scripts/ # Maintenance scripts
│ ├── dev.sh # Development mode quick start
│ ├── monitor.sh # Service monitoring
│ ├── backup.sh # Data backup
│ ├── update.sh # Service update
│ └── ...
├── deploy.sh # One-click deployment script
├── docker-compose.yml # Production mode configuration
├── docker-compose.dev.yml # Development mode configuration
├── Dockerfile # Production image
├── Dockerfile.dev # Development image
├── pyproject.toml # Project dependencies
├── uv.lock # Dependency lock
├── env.example # Environment variable template
├── CLAUDE.md # Claude AI guide
└── README.md # This documentation

🐛 Troubleshooting

Common Issues

Q1: What to do if service fails to start?
# View detailed logs
docker compose logs
# Check port usage
netstat -tulpn | grep 8000
# Check Docker status
docker ps -a
Q2: multimodal_processed error?

Note: This issue has been fixed in LightRAG 1.4.9.4+. If you encounter this error, your version is outdated.

Solution:

# Option 1: Upgrade to latest version (recommended)# Modify LightRAG version in pyproject.toml# lightrag = "^1.4.9.4"# Rebuild image
docker compose down
docker compose up -d --build
# Option 2: Clean old data (temporary solution)
rm -rf ./rag_local_storage
docker compose restart
Q3: File upload returns 400 error?

Check:

  • File format supported (PDF, DOCX, PNG, JPG, etc.)
  • File size exceeds 100MB
  • File is empty
# View supported formats
curl http://localhost:8000/docs
Q3.5: Embedding dimension error?

If you encounter dimension-related errors, need to clean data and rebuild:

# Stop services
docker compose down
# Delete all volumes (clear database)
docker volume rm rag-api_dragonflydb_data rag-api_qdrant_data rag-api_memgraph_data
# Modify EMBEDDING_DIM in .env
EMBEDDING_DIM=1024 # or 4096, must match the model# Restart
docker compose up -d
Q4: Query is very slow (>30 seconds)?

Optimization suggestions:

  1. Use naive or hybrid mode instead of mix
  2. Increase MAX_ASYNC parameter (in .env)
  3. Reduce TOP_K and CHUNK_TOP_K
  4. Enable Reranker
# Modify .env
MAX_ASYNC=8
TOP_K=20
CHUNK_TOP_K=10
Q5: Out of memory (OOM)?

If using local MinerU:

# Switch to remote mode# Modify in .env
MINERU_MODE=remote
MINERU_API_TOKEN=your_token
# Or limit concurrency
DOCUMENT_PROCESSING_CONCURRENCY=1
Q6: Tasks lost after container restart?

Problem Symptoms:

  • Cannot query previous task status after container restart
  • Tasks disappear after tenant instance evicted by LRU

Solution: Enable Redis task storage

# Modify .env
TASK_STORE_STORAGE=redis
# Restart services
docker compose restart
# Verify
docker compose logs api | grep TaskStore
# Should see: ✅ TaskStore: Redis connection successful

Configuration Description:

  • memory mode: In-memory storage, data lost after restart (default, suitable for development)
  • redis mode: Persistent storage, supports container restart and instance rebuild (production recommended)

TTL Strategy (Redis mode auto-cleanup):

  • completed tasks: 24 hours
  • failed tasks: 24 hours
  • pending/processing tasks: 6 hours
Q7: VLM mode processing failed?

Check Items:

  1. vision_model_func not configured

    • Check logs:vision_model_func not found, fallback to off mode
    • Ensure LLM API is configured in .env
  2. Image file does not exist

    • Check logs:Image file not found: xxx
    • Possibly corrupted MinerU ZIP or extraction failed
  3. Timeout error

    • full mode may timeout on large files
    • Suggestion: Use selective mode first, or increase VLM_TIMEOUT
# Modify .env
VLM_TIMEOUT=300 # Increase to 5 minutes
RAG_VLM_MODE=selective # downgrade to selective

Debugging Tips:

# View detailed logs
docker compose logs -f | grep VLM
# Test single file
curl -X POST 'http://localhost:8000/insert?tenant_id=test&doc_id=test&vlm_mode=off' \
-F 'file=@test.pdf'

Performance Tuning Recommendations

ScenarioMAX_ASYNCTOP_KCHUNK_TOP_KMINERU_MODE
Fast response8105remote
Balanced mode82010remote
High accuracy46020remote
Resource limited42010remote

📖 Documentation


🤝 Contributing

We welcome all forms of contribution!

How to Contribute

  1. Fork the project
git clone https://github.com/BukeLy/rag-api.git
cd rag-api
  1. Create feature branch
git checkout -b feature/your-feature-name
  1. Development and Testing
# Install dependencies
uv sync
# Run tests
uv run pytest
# Code formatting
uv run black .
uv run isort .
  1. Submit code
git add .
git commit -m "feat: Add new feature"
git push origin feature/your-feature-name
  1. Create Pull Request

Create a PR on GitHub with detailed description of your changes.

Commit Conventions

Use semantic commit messages:

  • feat: New feature
  • fix: Bug fix
  • docs: Documentation update
  • style: Code formatting
  • refactor: Code refactoring
  • perf: Performance optimization
  • test: Testing
  • chore: Build/tools

See PR Workflow Documentation


📄 License

This project is licensed under the MIT License. See the LICENSE file for details.


🙏 Acknowledgments

This project is built on the following excellent open source projects:

  • LightRAG - Efficient knowledge graph RAG framework
  • RAG-Anything - Multimodal document parsing
  • MinerU - Powerful PDF parsing tool
  • Docling - Lightweight document parsing
  • FastAPI - Modern Python web framework

Special thanks to all contributors and users for their support! 🎉


📬 Contact Us


⭐ If this project helps you, please give it a Star!

Made with ❤️ by BukeLy

© 2025 RAG API. All rights reserved.

About

Multi-tenant RAG API powered by LightRAG/RAG-Anything. Auto-selects best parser (DeepSeek-OCR/MinerU/Docling) via complexity scoring

Topics

Resources

Stars

64 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

🚀 RAG API

Multi-tenant Multimodal Document Intelligent Retrieval System

Enterprise-grade RAG service built on RAG-Anything and LightRAG

CIPythonFastAPILightRAGDockerLicense

English | 简体中文

FeaturesQuick StartArchitectureAPI DocumentationDeployment


📖 Introduction

RAG API is an enterprise-grade Retrieval-Augmented Generation (RAG) service that combines the powerful document parsing capabilities of RAG-Anything with the efficient knowledge graph retrieval technology of LightRAG, providing intelligent Q&A capabilities for your documents.

🎯 Key Highlights

  • 🏢 Multi-tenant Isolation - Complete tenant data isolation for enterprise multi-tenant scenarios
  • 🎨 Multimodal Parsing - Support for PDF, Word, images and more, with full OCR, tables, and formulas coverage
  • High-performance Retrieval - Knowledge graph-based hybrid retrieval with 6-15 second query response
  • 🔄 Flexible Deployment - Support for production and development modes with one-click switching
  • 📦 Ready to Use - One-click Docker deployment, service starts in 3 minutes
  • 🎛️ Multiple Parsing Engines - DeepSeek-OCR (Remote API) + MinerU (Local/Remote API) + Docling (Fast)
  • 🎨 RAG-Anything VLM Enhancement - Three modes (off/selective/full) for deep chart understanding
  • 💾 Task Persistence - Redis storage support, tasks recoverable after container restart/instance rebuild

✨ Features

📄 Document Processing

  • Multiple Format Support

    • PDF, Word, Excel, PPT
    • PNG, JPG, WebP images
    • TXT, Markdown text
  • Intelligent Parsing

    • Plain text (.txt, .md) → Direct insertion (ultra-fast ~1s, skip parser)
    • OCR text recognition
    • Structured table extraction
    • Mathematical formula recognition
    • Layout analysis
  • RAG-Anything VLM Enhancement 🆕

    • off - Markdown only (fastest)
    • selective - Selective processing of important charts
    • full - Complete context enhancement processing
    • Smart filtering: with titles, large size, first page content
    • ⚠️Only supports remote MinerU mode, local mode uses RAG-Anything native methods
  • Batch Processing

    • Up to 100 files per batch
    • Async task queue
    • Real-time progress tracking

🔍 Intelligent Retrieval

  • Multi-mode Query

    • naive - Vector retrieval (fastest)
    • local - Local graph
    • global - Global graph
    • hybrid - Hybrid retrieval
    • mix - Full retrieval (most accurate)
  • Knowledge Graph

    • Automatic entity extraction
    • Relationship reasoning
    • Semantic understanding
    • Context enhancement
  • External Storage

    • DragonflyDB (KV storage + task storage)
    • Qdrant (vector storage)
    • Memgraph (graph database)
    • Task persistence (Redis mode)

🏗️ Architecture

System Architecture Diagram

graph TB
subgraph "Client Layer"
Client[Client Application]
WebUI[Web Interface]
end
subgraph "API Gateway Layer"
FastAPI[FastAPI Service]
Auth[Tenant Authentication]
end
subgraph "Business Logic Layer"
TenantMgr[Tenant Manager]
TaskQueue[Task Queue]
subgraph "Document Processing"
DeepSeekOCR[DeepSeek-OCR<br/>Fast OCR 80% cases]
MinerU[MinerU Parser<br/>Complex multimodal]
Docling[Docling Parser<br/>Fast lightweight]
FileRouter[Smart Router<br/>Complexity scoring]
end
subgraph "RAG Engine"
LightRAG[LightRAG Instance Pool<br/>LRU Cache 50]
KG[Knowledge Graph Engine]
Vector[Vector Retrieval Engine]
end
end
subgraph "Storage Layer"
DragonflyDB[(DragonflyDB<br/>KV Storage)]
Qdrant[(Qdrant<br/>Vector Database)]
Memgraph[(Memgraph<br/>Graph Database)]
Local[(Local Files<br/>Temp Storage)]
end
subgraph "External Services"
LLM[LLM<br/>Entity Extraction/Generation]
Embedding[Embedding<br/>Vectorization]
Rerank[Rerank<br/>Reranking]
end
Client --> FastAPI
WebUI --> FastAPI
FastAPI --> Auth
Auth --> TenantMgr
TenantMgr --> TaskQueue
TenantMgr --> LightRAG
TaskQueue --> FileRouter
FileRouter --> DeepSeekOCR
FileRouter --> MinerU
FileRouter --> Docling
DeepSeekOCR --> LightRAG
MinerU --> LightRAG
Docling --> LightRAG
LightRAG --> KG
LightRAG --> Vector
KG --> DragonflyDB
KG --> Memgraph
Vector --> Qdrant
LightRAG --> Local
LightRAG --> LLM
LightRAG --> Embedding
Vector --> Rerank
style FastAPI fill:#00C7B7
style LightRAG fill:#FF6B6B
style DeepSeekOCR fill:#5DADE2
style MinerU fill:#4ECDC4
style Docling fill:#95E1D3
style TenantMgr fill:#F38181
Loading

Multi-tenant Architecture

graph TB
subgraph "Tenant A"
A_Config[Tenant A Config<br/>Independent API Key]
A_Instance[LightRAG Instance A<br/>Dedicated LLM/Embedding]
A_Data[(Tenant A Data<br/>Fully Isolated)]
A_Config --> A_Instance
A_Instance --> A_Data
end
subgraph "Tenant B"
B_Config[Tenant B Config<br/>Independent API Key]
B_Instance[LightRAG Instance B<br/>Dedicated LLM/Embedding]
B_Data[(Tenant B Data<br/>Fully Isolated)]
B_Config --> B_Instance
B_Instance --> B_Data
end
subgraph "Tenant C"
C_Config[Using Global Config]
C_Instance[LightRAG Instance C<br/>Shared LLM/Embedding]
C_Data[(Tenant C Data<br/>Fully Isolated)]
C_Config --> C_Instance
C_Instance --> C_Data
end
Pool[Instance Pool Manager<br/>LRU Cache + Config Isolation]
Global[Global Config<br/>Default API Key]
Pool --> A_Instance
Pool --> B_Instance
Pool --> C_Instance
C_Config -.fallback.-> Global
style Pool fill:#F38181
style Global fill:#95E1D3
style A_Config fill:#FFD93D
style B_Config fill:#FFD93D
style C_Config fill:#E8E8E8
Loading

Core Technology Stack

🔧 Frameworks & Runtime

  • FastAPI 0.115+
  • Python 3.11+
  • Uvicorn
  • Docker & Docker Compose

🧠 AI & RAG

  • LightRAG 1.4.9.4
  • RAG-Anything
  • MinerU (PDF-Extract-Kit)
  • Docling

💾 Storage & Database

  • DragonflyDB(Redis compatible)
  • Qdrant(Vector Database)
  • Memgraph(Graph Database)
  • Local filesystem

🚀 Quick Start

Option 1: One-click Deployment (Recommended)

Suitable for production and testing environments:

# 1. Clone the project
git clone https://github.com/BukeLy/rag-api.git
cd rag-api
# 2. Configure environment variables
cp env.example .env
nano .env # Fill in your API keys# 3. Run deployment script
chmod +x deploy.sh
./deploy.sh
# Select deployment mode:# 1) Production Mode - Standard container deployment# 2) Development Mode - Code hot-reload# 4. Verify service
curl http://localhost:8000/

Access Swagger Documentation:http://localhost:8000/docs

Option 2: Docker Compose

Production Mode

# Configure environment variables
cp env.example .env
nano .env
# Start services
docker compose -f docker-compose.yml up -d
# View logs
docker compose -f docker-compose.yml logs -f

Development Mode (Code Hot-reload)

# Start development environment
docker compose -f docker-compose.dev.yml up -d
# Or use quick script
./scripts/dev.sh
# Code changes will auto-reload without restart

Option 3: Local Development

# Install uv (Python package manager)
curl -LsSf https://astral.sh/uv/install.sh | sh
# Install dependencies
uv sync
# Configure environment variables
cp env.example .env
nano .env
# Start services
uv run uvicorn main:app --host 0.0.0.0 --port 8000 --reload

Environment Variable Configuration

Minimum configuration (required):

# LLM Configuration (Function-oriented naming)
LLM_API_KEY=your_llm_api_key
LLM_BASE_URL=https://ark.cn-beijing.volces.com/api/v3
LLM_MODEL=ep-xxx-xxx
# LLM_REQUESTS_PER_MINUTE=800 # Rate limit (optional)# LLM_TOKENS_PER_MINUTE=40000 # Rate limit (optional)# LLM_MAX_ASYNC=8 # [Optional, expert mode] Manual concurrency control# # Auto-calculated when unset: min(RPM, TPM/3500) = 11# Embedding Configuration (Function-oriented naming)
EMBEDDING_API_KEY=your_embedding_api_key
EMBEDDING_BASE_URL=https://api.siliconflow.cn/v1
EMBEDDING_MODEL=Qwen/Qwen3-Embedding-0.6B
EMBEDDING_DIM=1024
# EMBEDDING_MAX_ASYNC=32 # [Optional, expert mode] Auto-calculated when unset: 800# MinerU Mode (Remote recommended)
MINERU_MODE=remote
MINERU_API_TOKEN=your_token
MINERU_HTTP_TIMEOUT=60 # MinerU download timeout (seconds, default 60)
FILE_SERVICE_BASE_URL=http://your-ip:8000
# VLM Chart Enhancement Configuration 🆕# ⚠️ Note: Only effective in MINERU_MODE=remote
RAG_VLM_MODE=off # off / selective / full
RAG_IMPORTANCE_THRESHOLD=0.5 # Importance threshold (selective mode)
RAG_CONTEXT_WINDOW=2 # Context window (full mode)
RAG_CONTEXT_MODE=page # page / chunk
RAG_MAX_CONTEXT_TOKENS=3000 # Max context tokens# Task Storage Configuration 🆕
TASK_STORE_STORAGE=redis # memory / redis (production recommends redis)# Document Insert Verification Configuration 🆕
DOC_INSERT_VERIFICATION_TIMEOUT=300 # Verification timeout (seconds, default 5 minutes)
DOC_INSERT_VERIFICATION_POLL_INTERVAL=0.5 # Poll interval (seconds, default 500ms)# Model Call Timeout Configuration 🆕
MODEL_CALL_TIMEOUT=90 # Model call max timeout (seconds, default 90)

⚡ Auto Concurrency Calculation:

  • LLM: When LLM_MAX_ASYNC is unset, auto-calculated as min(RPM, TPM/3500) ≈ 11
  • Embedding: When EMBEDDING_MAX_ASYNC is unset, auto-calculated as min(RPM, TPM/500) ≈ 800
  • Rerank: When RERANK_MAX_ASYNC is unset, auto-calculated as min(RPM, TPM/500) ≈ 800

✅ Recommended: Don't set *_MAX_ASYNC, let the system auto-calculate to completely avoid 429 errors

See env.example for complete configuration.


📚 API Documentation

Core Endpoints

1️⃣ Upload Document

# Single file upload (default mode)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc1" \
-F "file=@document.pdf" \
-F "parser=auto"# VLM chart enhancement mode 🆕# off: Markdown only (fastest, default)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc2&vlm_mode=off" \
-F "file=@document.pdf"# selective: Selective processing of important charts (balance performance and quality)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc3&vlm_mode=selective" \
-F "file=@document.pdf"# full: Complete RAG-Anything processing (highest quality, context enhancement enabled)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc4&vlm_mode=full" \
-F "file=@document.pdf"# Response
{
"task_id": "task-xxx-xxx",
"doc_id": "doc1",
"filename": "document.pdf",
"vlm_mode": "off",
"status": "pending"
}

2️⃣ Batch Upload

curl -X POST "http://localhost:8000/batch?tenant_id=your_tenant" \
-F "files=@doc1.pdf" \
-F "files=@doc2.docx" \
-F "files=@image.png"# Response
{
"batch_id": "batch-xxx-xxx",
"total_files": 3,
"accepted_files": 3,
"tasks": [...]
}

3️⃣ Intelligent Query (Query API v2.0)

New Advanced Features:

  • Conversation History: Support for multi-turn conversation context
  • Custom Prompts: Customize response style
  • Response Format Control: paragraph/list/json
  • Keyword Precision Retrieval: hl_keywords/ll_keywords
  • Streaming Output: Real-time generation viewing
# Basic query
curl -X POST "http://localhost:8000/query?tenant_id=your_tenant" \
-H "Content-Type: application/json" \
-d '{ "query": "What are the core viewpoints in the document?", "mode": "hybrid" }'# Advanced query (multi-turn dialogue + custom prompt)
curl -X POST "http://localhost:8000/query?tenant_id=your_tenant" \
-H "Content-Type: application/json" \
-d '{ "query": "Can you elaborate on the second point?", "mode": "hybrid", "conversation_history": [ {"role": "user", "content": "What are the key points?"}, {"role": "assistant", "content": "There are mainly three points..."} ], "user_prompt": "Please answer in professional academic language", "response_type": "list" }'# Streaming query (SSE)
curl -N -X POST "http://localhost:8000/query/stream?tenant_id=your_tenant" \
-H "Content-Type: application/json" \
-d '{ "query": "What are the core viewpoints in the document?", "mode": "hybrid" }'# Response (real-time streaming output)
data: {"chunk": "Based on", "done": false}
data: {"chunk": "document content", "done": false}
data: {"done": true}

4️⃣ Task Status Query

curl "http://localhost:8000/task/task-xxx-xxx?tenant_id=your_tenant"# Response
{
"task_id": "task-xxx-xxx",
"status": "completed",
"progress": 100,
"result": {...}
}

5️⃣ Tenant Management

# Get tenant statistics
curl "http://localhost:8000/tenants/stats?tenant_id=your_tenant"# Clear tenant cache
curl -X DELETE "http://localhost:8000/tenants/cache?tenant_id=your_tenant"# View instance pool status (admin)
curl "http://localhost:8000/tenants/pool/stats"

VLM Mode Comparison 🆕

ModeSpeedQualityResource UsageUse Case
off⚡⚡⚡⚡⚡⭐⭐⭐Very LowPlain text documents, fast batch processing
selective⚡⚡⚡⚡⭐⭐⭐⭐LowDocuments with key charts (recommended)
full⚡⚡⭐⭐⭐⭐⭐HighChart-intensive research reports, papers

Processing Time Estimate (20-page PDF example):

  • off: ~10 seconds(Markdown only)
  • selective: ~30 seconds(5-10 important charts)
  • full: ~120 seconds(complete context processing)

Query Mode Comparison

ModeSpeedAccuracyUse Case
naive⚡⚡⚡⚡⚡⭐⭐⭐Simple Q&A, fast retrieval
local⚡⚡⚡⚡⭐⭐⭐⭐Local entity relationship queries
global⚡⚡⚡⭐⭐⭐⭐Global knowledge graph reasoning
hybrid⚡⚡⚡⭐⭐⭐⭐⭐Hybrid retrieval (recommended)
mix⚡⚡⭐⭐⭐⭐⭐Complex questions, deep analysis

Query API v2.0 Advanced Parameters

ParameterTypeDescriptionExample
conversation_historyList[Dict]Multi-turn conversation context[{"role": "user", "content": "..."}]
user_promptstrCustom prompt"Please answer in professional academic language"
response_typestrResponse format"paragraph", "list", "json"
hl_keywordsList[str]High priority keywords["artificial intelligence", "machine learning"]
ll_keywordsList[str]Low priority keywords["application", "case study"]
only_need_contextboolReturn context only (debug)true
max_entity_tokensintEntity token limit6000

Complete API documentation:http://localhost:8000/docs


🎯 Usage Examples

Python SDK

importrequests# ConfigurationBASE_URL="http://localhost:8000"TENANT_ID="your_tenant"# Upload documentwithopen("document.pdf", "rb") asf:
response=requests.post(
f"{BASE_URL}/insert",
params={"tenant_id": TENANT_ID, "doc_id": "doc1"},
files={"file": f}
)
task_id=response.json()["task_id"]
print(f"Task ID: {task_id}")
# Queryresponse=requests.post(
f"{BASE_URL}/query",
params={"tenant_id": TENANT_ID},
json={
"query": "What is the main content of the document?",
"mode": "hybrid",
"top_k": 10
}
)
result=response.json()
print(f"Answer: {result['answer']}")

Complete cURL Example

# 1. Upload PDF document
TASK_ID=$(curl -X POST "http://localhost:8000/insert?tenant_id=demo&doc_id=report" \ -F "file=@report.pdf"| jq -r '.task_id')echo"Task ID: $TASK_ID"# 2. Wait for processing completionwhiletrue;do
STATUS=$(curl -s "http://localhost:8000/task/$TASK_ID?tenant_id=demo"| jq -r '.status')echo"Status: $STATUS"if [ "$STATUS"="completed" ] || [ "$STATUS"="failed" ];thenbreakfi
sleep 2
done# 3. Query document content
curl -X POST "http://localhost:8000/query?tenant_id=demo" \
-H "Content-Type: application/json" \
-d '{ "query": "What are the main conclusions of this report?", "mode": "hybrid" }'| jq '.answer'

🛠️ Deployment

System Requirements

Minimum Configuration:

  • CPU: 2 cores
  • RAM: 4GB
  • Disk: 40GB SSD
  • OS: Ubuntu 20.04+ / Debian 11+ / CentOS 8+

Recommended Configuration (Production):

  • CPU: 4 cores
  • RAM: 8GB
  • Disk: 100GB SSD
  • OS: Ubuntu 22.04 LTS

Server Deployment

Quick Deployment on Aliyun/Tencent Cloud

# SSH login to server
ssh root@your-server-ip
# Clone project
git clone https://github.com/BukeLy/rag-api.git
cd rag-api
# Run one-click deployment script
chmod +x deploy.sh
./deploy.sh
# The script will automatically:# 1. Install Docker and Docker Compose# 2. Configure environment variables# 3. Optimize system parameters# 4. Start services# 5. Verify health status

External Storage Configuration

Supports DragonflyDB + Qdrant + Memgraph external storage (enabled by default):

# Configure in .env
USE_EXTERNAL_STORAGE=true
# DragonflyDB configuration (KV Storage)
KV_STORAGE=RedisKVStorage
REDIS_URI=redis://dragonflydb:6379/0
# Qdrant configuration (vector storage)
VECTOR_STORAGE=QdrantVectorDBStorage
QDRANT_URL=http://qdrant:6333
# Memgraph configuration (graph storage)
GRAPH_STORAGE=MemgraphStorage
MEMGRAPH_URI=bolt://memgraph:7687
MEMGRAPH_USERNAME=
MEMGRAPH_PASSWORD=

See External Storage Deployment Documentation

Docker Compose Configuration

The project provides two configuration files:

FilePurposeFeatures
docker-compose.ymlProduction modeCode packaged in image, optimal performance
docker-compose.dev.ymlDevelopment modeCode mounted externally, supports hot-reload

Select configuration file:

# Production mode
docker compose -f docker-compose.yml up -d
# Development mode
docker compose -f docker-compose.dev.yml up -d

Performance Optimization

Tuning Parameters

Configure in .env:

# ⚡ Concurrency Control (Recommended: use auto-calculation)# LLM_MAX_ASYNC=8 # [Expert mode] Manually specify LLM concurrency# # Auto-calculated when unset: min(RPM, TPM/3500) ≈ 11# EMBEDDING_MAX_ASYNC=32 # [Expert mode] Manually specify Embedding concurrency# # Auto-calculated when unset: min(RPM, TPM/500) ≈ 800# RERANK_MAX_ASYNC=16 # [Expert mode] Manually specify Rerank concurrency# # Auto-calculated when unset: min(RPM, TPM/500) ≈ 800# Retrieval count (affects query quality and speed)
TOP_K=20 # Entity/relationship retrieval count
CHUNK_TOP_K=10 # Text chunk retrieval count# Document processing concurrency
DOCUMENT_PROCESSING_CONCURRENCY=10 # Remote mode can be set high, local mode set to 1

🎯 Concurrency Configuration Recommendations:

  • Recommended: Don't set *_MAX_ASYNC, let the system auto-calculate based on TPM/RPM
  • Expert mode: If manual control needed, can set LLM_MAX_ASYNC and other parameters
  • Advantage: Auto-calculation completely avoids 429 errors (TPM limit reached)

Mode Selection

  • MinerU Remote Mode (Recommended): High concurrency, resource-efficient
  • MinerU Local Mode: Requires GPU, high memory usage
  • Docling Mode: Fast and lightweight, suitable for simple documents

🏢 Multi-tenant Usage

Tenant Isolation

Each tenant has:

  • ✅ Independent LightRAG instance
  • ✅ Isolated data storage space
  • ✅ Independent vector index
  • ✅ Dedicated knowledge graph
  • Independent service configuration (LLM, Embedding, Rerank, DeepSeek-OCR, MinerU)🆕

Tenant Configuration Management 🆕

Each tenant can independently configure 5 services with hot-reload support:

# 1️⃣ Configure independent DeepSeek-OCR API key for Tenant A
curl -X PUT "http://localhost:8000/tenants/tenant_a/config" \
-H "Content-Type: application/json" \
-d '{ "ds_ocr_config": { "api_key": "sk-tenant-a-ds-ocr-key", "base_url": "https://api.siliconflow.cn/v1", "model": "deepseek-ai/DeepSeek-OCR", "timeout": 90 } }'# 2️⃣ Configure independent MinerU API token for Tenant B
curl -X PUT "http://localhost:8000/tenants/tenant_b/config" \
-H "Content-Type: application/json" \
-d '{ "mineru_config": { "api_token": "tenant-b-mineru-token", "base_url": "https://mineru.net", "model_version": "vlm" } }'# 3️⃣ Configure multiple services simultaneously (LLM + Embedding + DeepSeek-OCR)
curl -X PUT "http://localhost:8000/tenants/tenant_c/config" \
-H "Content-Type: application/json" \
-d '{ "llm_config": { "api_key": "sk-tenant-c-llm-key", "model": "gpt-4" }, "embedding_config": { "api_key": "sk-tenant-c-embedding-key", "model": "Qwen/Qwen3-Embedding-0.6B", "dim": 1024 }, "ds_ocr_config": { "api_key": "sk-tenant-c-ds-ocr-key" } }'# 4️⃣ Query tenant configuration (API key auto-masked)
curl "http://localhost:8000/tenants/tenant_a/config"# Response example
{
"tenant_id": "tenant_a",
"ds_ocr_config": {
"api_key": "sk-***-key", // Auto-masked
"timeout": 90
},
"merged_config": {
"llm": {...}, // Using Global Config
"embedding": {...}, // Using Global Config
"rerank": {...}, // Using Global Config
"ds_ocr": {...}, // Using tenant config
"mineru": {...} // Using Global Config
}
}
# 5️⃣ Refresh config cache (config hot-reload)
curl -X POST "http://localhost:8000/tenants/tenant_a/config/refresh"# 6️⃣ Delete tenant config (restore to global config)
curl -X DELETE "http://localhost:8000/tenants/tenant_a/config"

Supported Configuration Items:

ServiceConfig FieldDescription
LLMllm_configModel, API key, base_url, etc.
Embeddingembedding_configModel, API key, dimension, etc.
Rerankrerank_configModel, API key, etc.
DeepSeek-OCRds_ocr_configAPI key, timeout, mode, etc.
MinerUmineru_configAPI token, version, timeout, etc.

Configuration Priority: Tenant config > Global config

Use Cases:

  • 🔐 Multi-tenant SaaS: Each tenant uses their own API key
  • 💰 Pay-per-use: Track tenant usage through independent API keys
  • 🎯 Differentiated Services: Different tenants use different models (GPT-4 vs GPT-3.5)
  • 🧪 A/B Testing: Compare different models/parameters

Usage

All APIs require tenant_id parameter:

# Tenant A upload document
curl -X POST "http://localhost:8000/insert?tenant_id=tenant_a&doc_id=doc1" \
-F "file=@doc.pdf"# Tenant B upload document (fully isolated)
curl -X POST "http://localhost:8000/insert?tenant_id=tenant_b&doc_id=doc1" \
-F "file=@doc.pdf"# Tenant A query (can only query own documents)
curl -X POST "http://localhost:8000/query?tenant_id=tenant_a" \
-H "Content-Type: application/json" \
-d '{"query": "document content", "mode": "hybrid"}'

Instance Pool Management

  • Capacity: Cache up to 50 tenant instances
  • Strategy: LRU (Least Recently Used) automatic cleanup
  • Config Isolation: Each tenant can use independent LLM, Embedding, parser configuration

📊 Monitoring & Maintenance

Common Commands

# View service status
docker compose ps
# View real-time logs
docker compose logs -f
# Restart services
docker compose restart
# Stop services
docker compose down
# View resource usage
docker stats
# Clean Docker resources
docker system prune -f

Maintenance Scripts

# Monitor service health
./scripts/monitor.sh
# Backup data
./scripts/backup.sh
# Update services
./scripts/update.sh
# Performance testing
./scripts/test_concurrent_perf.sh
# Performance monitoring
./scripts/monitor_performance.sh

Health Checks

# Complete health check (recommended)
./scripts/health_check.sh
./scripts/health_check.sh --verbose # verbose output# API health check
curl http://localhost:8000/
# Tenant statistics
curl "http://localhost:8000/tenants/stats?tenant_id=your_tenant"# Instance pool status
curl "http://localhost:8000/tenants/pool/stats"

🗂️ Project Structure

rag-api/
├── main.py # FastAPI application entry
├── api/ # API route modules
│ ├── __init__.py # Route aggregation
│ ├── insert.py # Document upload (single/batch)
│ ├── query.py # Intelligent query
│ ├── task.py # Task status query
│ ├── tenant.py # Tenant management
│ ├── files.py # File service
│ ├── models.py # Pydantic models
│ └── task_store.py # Task storage
├── src/ # Core business logic
│ ├── rag.py # LightRAG lifecycle management
│ ├── multi_tenant.py # Multi-tenant instance manager
│ ├── tenant_deps.py # Tenant dependency injection
│ ├── logger.py # Unified logging
│ ├── metrics.py # Performance metrics
│ ├── file_url_service.py # Temporary file service
│ ├── mineru_client.py # MinerU client
│ └── mineru_result_processor.py # Result processing
├── docs/ # Documentation
│ ├── ARCHITECTURE.md # Architecture design documentation
│ ├── USAGE.md # Detailed usage guide
│ ├── DEPLOY_MODES.md # Deployment mode description
│ ├── PR_WORKFLOW.md # PR workflow
│ └── ...
├── scripts/ # Maintenance scripts
│ ├── dev.sh # Development mode quick start
│ ├── monitor.sh # Service monitoring
│ ├── backup.sh # Data backup
│ ├── update.sh # Service update
│ └── ...
├── deploy.sh # One-click deployment script
├── docker-compose.yml # Production mode configuration
├── docker-compose.dev.yml # Development mode configuration
├── Dockerfile # Production image
├── Dockerfile.dev # Development image
├── pyproject.toml # Project dependencies
├── uv.lock # Dependency lock
├── env.example # Environment variable template
├── CLAUDE.md # Claude AI guide
└── README.md # This documentation

🐛 Troubleshooting

Common Issues

Q1: What to do if service fails to start?
# View detailed logs
docker compose logs
# Check port usage
netstat -tulpn | grep 8000
# Check Docker status
docker ps -a
Q2: multimodal_processed error?

Note: This issue has been fixed in LightRAG 1.4.9.4+. If you encounter this error, your version is outdated.

Solution:

# Option 1: Upgrade to latest version (recommended)# Modify LightRAG version in pyproject.toml# lightrag = "^1.4.9.4"# Rebuild image
docker compose down
docker compose up -d --build
# Option 2: Clean old data (temporary solution)
rm -rf ./rag_local_storage
docker compose restart
Q3: File upload returns 400 error?

Check:

  • File format supported (PDF, DOCX, PNG, JPG, etc.)
  • File size exceeds 100MB
  • File is empty
# View supported formats
curl http://localhost:8000/docs
Q3.5: Embedding dimension error?

If you encounter dimension-related errors, need to clean data and rebuild:

# Stop services
docker compose down
# Delete all volumes (clear database)
docker volume rm rag-api_dragonflydb_data rag-api_qdrant_data rag-api_memgraph_data
# Modify EMBEDDING_DIM in .env
EMBEDDING_DIM=1024 # or 4096, must match the model# Restart
docker compose up -d
Q4: Query is very slow (>30 seconds)?

Optimization suggestions:

  1. Use naive or hybrid mode instead of mix
  2. Increase MAX_ASYNC parameter (in .env)
  3. Reduce TOP_K and CHUNK_TOP_K
  4. Enable Reranker
# Modify .env
MAX_ASYNC=8
TOP_K=20
CHUNK_TOP_K=10
Q5: Out of memory (OOM)?

If using local MinerU:

# Switch to remote mode# Modify in .env
MINERU_MODE=remote
MINERU_API_TOKEN=your_token
# Or limit concurrency
DOCUMENT_PROCESSING_CONCURRENCY=1
Q6: Tasks lost after container restart?

Problem Symptoms:

  • Cannot query previous task status after container restart
  • Tasks disappear after tenant instance evicted by LRU

Solution: Enable Redis task storage

# Modify .env
TASK_STORE_STORAGE=redis
# Restart services
docker compose restart
# Verify
docker compose logs api | grep TaskStore
# Should see: ✅ TaskStore: Redis connection successful

Configuration Description:

  • memory mode: In-memory storage, data lost after restart (default, suitable for development)
  • redis mode: Persistent storage, supports container restart and instance rebuild (production recommended)

TTL Strategy (Redis mode auto-cleanup):

  • completed tasks: 24 hours
  • failed tasks: 24 hours
  • pending/processing tasks: 6 hours
Q7: VLM mode processing failed?

Check Items:

  1. vision_model_func not configured

    • Check logs:vision_model_func not found, fallback to off mode
    • Ensure LLM API is configured in .env
  2. Image file does not exist

    • Check logs:Image file not found: xxx
    • Possibly corrupted MinerU ZIP or extraction failed
  3. Timeout error

    • full mode may timeout on large files
    • Suggestion: Use selective mode first, or increase VLM_TIMEOUT
# Modify .env
VLM_TIMEOUT=300 # Increase to 5 minutes
RAG_VLM_MODE=selective # downgrade to selective

Debugging Tips:

# View detailed logs
docker compose logs -f | grep VLM
# Test single file
curl -X POST 'http://localhost:8000/insert?tenant_id=test&doc_id=test&vlm_mode=off' \
-F 'file=@test.pdf'

Performance Tuning Recommendations

ScenarioMAX_ASYNCTOP_KCHUNK_TOP_KMINERU_MODE
Fast response8105remote
Balanced mode82010remote
High accuracy46020remote
Resource limited42010remote

📖 Documentation


🤝 Contributing

We welcome all forms of contribution!

How to Contribute

  1. Fork the project
git clone https://github.com/BukeLy/rag-api.git
cd rag-api
  1. Create feature branch
git checkout -b feature/your-feature-name
  1. Development and Testing
# Install dependencies
uv sync
# Run tests
uv run pytest
# Code formatting
uv run black .
uv run isort .
  1. Submit code
git add .
git commit -m "feat: Add new feature"
git push origin feature/your-feature-name
  1. Create Pull Request

Create a PR on GitHub with detailed description of your changes.

Commit Conventions

Use semantic commit messages:

  • feat: New feature
  • fix: Bug fix
  • docs: Documentation update
  • style: Code formatting
  • refactor: Code refactoring
  • perf: Performance optimization
  • test: Testing
  • chore: Build/tools

See PR Workflow Documentation


📄 License

This project is licensed under the MIT License. See the LICENSE file for details.


🙏 Acknowledgments

This project is built on the following excellent open source projects:

  • LightRAG - Efficient knowledge graph RAG framework
  • RAG-Anything - Multimodal document parsing
  • MinerU - Powerful PDF parsing tool
  • Docling - Lightweight document parsing
  • FastAPI - Modern Python web framework

Special thanks to all contributors and users for their support! 🎉


📬 Contact Us


⭐ If this project helps you, please give it a Star!

Made with ❤️ by BukeLy

© 2025 RAG API. All rights reserved.

About

Multi-tenant RAG API powered by LightRAG/RAG-Anything. Auto-selects best parser (DeepSeek-OCR/MinerU/Docling) via complexity scoring

Topics

Resources

Stars

64 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

🚀 RAG API

Multi-tenant Multimodal Document Intelligent Retrieval System

Enterprise-grade RAG service built on RAG-Anything and LightRAG

CIPythonFastAPILightRAGDockerLicense

English | 简体中文

FeaturesQuick StartArchitectureAPI DocumentationDeployment


📖 Introduction

RAG API is an enterprise-grade Retrieval-Augmented Generation (RAG) service that combines the powerful document parsing capabilities of RAG-Anything with the efficient knowledge graph retrieval technology of LightRAG, providing intelligent Q&A capabilities for your documents.

🎯 Key Highlights

  • 🏢 Multi-tenant Isolation - Complete tenant data isolation for enterprise multi-tenant scenarios
  • 🎨 Multimodal Parsing - Support for PDF, Word, images and more, with full OCR, tables, and formulas coverage
  • High-performance Retrieval - Knowledge graph-based hybrid retrieval with 6-15 second query response
  • 🔄 Flexible Deployment - Support for production and development modes with one-click switching
  • 📦 Ready to Use - One-click Docker deployment, service starts in 3 minutes
  • 🎛️ Multiple Parsing Engines - DeepSeek-OCR (Remote API) + MinerU (Local/Remote API) + Docling (Fast)
  • 🎨 RAG-Anything VLM Enhancement - Three modes (off/selective/full) for deep chart understanding
  • 💾 Task Persistence - Redis storage support, tasks recoverable after container restart/instance rebuild

✨ Features

📄 Document Processing

  • Multiple Format Support

    • PDF, Word, Excel, PPT
    • PNG, JPG, WebP images
    • TXT, Markdown text
  • Intelligent Parsing

    • Plain text (.txt, .md) → Direct insertion (ultra-fast ~1s, skip parser)
    • OCR text recognition
    • Structured table extraction
    • Mathematical formula recognition
    • Layout analysis
  • RAG-Anything VLM Enhancement 🆕

    • off - Markdown only (fastest)
    • selective - Selective processing of important charts
    • full - Complete context enhancement processing
    • Smart filtering: with titles, large size, first page content
    • ⚠️Only supports remote MinerU mode, local mode uses RAG-Anything native methods
  • Batch Processing

    • Up to 100 files per batch
    • Async task queue
    • Real-time progress tracking

🔍 Intelligent Retrieval

  • Multi-mode Query

    • naive - Vector retrieval (fastest)
    • local - Local graph
    • global - Global graph
    • hybrid - Hybrid retrieval
    • mix - Full retrieval (most accurate)
  • Knowledge Graph

    • Automatic entity extraction
    • Relationship reasoning
    • Semantic understanding
    • Context enhancement
  • External Storage

    • DragonflyDB (KV storage + task storage)
    • Qdrant (vector storage)
    • Memgraph (graph database)
    • Task persistence (Redis mode)

🏗️ Architecture

System Architecture Diagram

graph TB
subgraph "Client Layer"
Client[Client Application]
WebUI[Web Interface]
end
subgraph "API Gateway Layer"
FastAPI[FastAPI Service]
Auth[Tenant Authentication]
end
subgraph "Business Logic Layer"
TenantMgr[Tenant Manager]
TaskQueue[Task Queue]
subgraph "Document Processing"
DeepSeekOCR[DeepSeek-OCR<br/>Fast OCR 80% cases]
MinerU[MinerU Parser<br/>Complex multimodal]
Docling[Docling Parser<br/>Fast lightweight]
FileRouter[Smart Router<br/>Complexity scoring]
end
subgraph "RAG Engine"
LightRAG[LightRAG Instance Pool<br/>LRU Cache 50]
KG[Knowledge Graph Engine]
Vector[Vector Retrieval Engine]
end
end
subgraph "Storage Layer"
DragonflyDB[(DragonflyDB<br/>KV Storage)]
Qdrant[(Qdrant<br/>Vector Database)]
Memgraph[(Memgraph<br/>Graph Database)]
Local[(Local Files<br/>Temp Storage)]
end
subgraph "External Services"
LLM[LLM<br/>Entity Extraction/Generation]
Embedding[Embedding<br/>Vectorization]
Rerank[Rerank<br/>Reranking]
end
Client --> FastAPI
WebUI --> FastAPI
FastAPI --> Auth
Auth --> TenantMgr
TenantMgr --> TaskQueue
TenantMgr --> LightRAG
TaskQueue --> FileRouter
FileRouter --> DeepSeekOCR
FileRouter --> MinerU
FileRouter --> Docling
DeepSeekOCR --> LightRAG
MinerU --> LightRAG
Docling --> LightRAG
LightRAG --> KG
LightRAG --> Vector
KG --> DragonflyDB
KG --> Memgraph
Vector --> Qdrant
LightRAG --> Local
LightRAG --> LLM
LightRAG --> Embedding
Vector --> Rerank
style FastAPI fill:#00C7B7
style LightRAG fill:#FF6B6B
style DeepSeekOCR fill:#5DADE2
style MinerU fill:#4ECDC4
style Docling fill:#95E1D3
style TenantMgr fill:#F38181
Loading

Multi-tenant Architecture

graph TB
subgraph "Tenant A"
A_Config[Tenant A Config<br/>Independent API Key]
A_Instance[LightRAG Instance A<br/>Dedicated LLM/Embedding]
A_Data[(Tenant A Data<br/>Fully Isolated)]
A_Config --> A_Instance
A_Instance --> A_Data
end
subgraph "Tenant B"
B_Config[Tenant B Config<br/>Independent API Key]
B_Instance[LightRAG Instance B<br/>Dedicated LLM/Embedding]
B_Data[(Tenant B Data<br/>Fully Isolated)]
B_Config --> B_Instance
B_Instance --> B_Data
end
subgraph "Tenant C"
C_Config[Using Global Config]
C_Instance[LightRAG Instance C<br/>Shared LLM/Embedding]
C_Data[(Tenant C Data<br/>Fully Isolated)]
C_Config --> C_Instance
C_Instance --> C_Data
end
Pool[Instance Pool Manager<br/>LRU Cache + Config Isolation]
Global[Global Config<br/>Default API Key]
Pool --> A_Instance
Pool --> B_Instance
Pool --> C_Instance
C_Config -.fallback.-> Global
style Pool fill:#F38181
style Global fill:#95E1D3
style A_Config fill:#FFD93D
style B_Config fill:#FFD93D
style C_Config fill:#E8E8E8
Loading

Core Technology Stack

🔧 Frameworks & Runtime

  • FastAPI 0.115+
  • Python 3.11+
  • Uvicorn
  • Docker & Docker Compose

🧠 AI & RAG

  • LightRAG 1.4.9.4
  • RAG-Anything
  • MinerU (PDF-Extract-Kit)
  • Docling

💾 Storage & Database

  • DragonflyDB(Redis compatible)
  • Qdrant(Vector Database)
  • Memgraph(Graph Database)
  • Local filesystem

🚀 Quick Start

Option 1: One-click Deployment (Recommended)

Suitable for production and testing environments:

# 1. Clone the project
git clone https://github.com/BukeLy/rag-api.git
cd rag-api
# 2. Configure environment variables
cp env.example .env
nano .env # Fill in your API keys# 3. Run deployment script
chmod +x deploy.sh
./deploy.sh
# Select deployment mode:# 1) Production Mode - Standard container deployment# 2) Development Mode - Code hot-reload# 4. Verify service
curl http://localhost:8000/

Access Swagger Documentation:http://localhost:8000/docs

Option 2: Docker Compose

Production Mode

# Configure environment variables
cp env.example .env
nano .env
# Start services
docker compose -f docker-compose.yml up -d
# View logs
docker compose -f docker-compose.yml logs -f

Development Mode (Code Hot-reload)

# Start development environment
docker compose -f docker-compose.dev.yml up -d
# Or use quick script
./scripts/dev.sh
# Code changes will auto-reload without restart

Option 3: Local Development

# Install uv (Python package manager)
curl -LsSf https://astral.sh/uv/install.sh | sh
# Install dependencies
uv sync
# Configure environment variables
cp env.example .env
nano .env
# Start services
uv run uvicorn main:app --host 0.0.0.0 --port 8000 --reload

Environment Variable Configuration

Minimum configuration (required):

# LLM Configuration (Function-oriented naming)
LLM_API_KEY=your_llm_api_key
LLM_BASE_URL=https://ark.cn-beijing.volces.com/api/v3
LLM_MODEL=ep-xxx-xxx
# LLM_REQUESTS_PER_MINUTE=800 # Rate limit (optional)# LLM_TOKENS_PER_MINUTE=40000 # Rate limit (optional)# LLM_MAX_ASYNC=8 # [Optional, expert mode] Manual concurrency control# # Auto-calculated when unset: min(RPM, TPM/3500) = 11# Embedding Configuration (Function-oriented naming)
EMBEDDING_API_KEY=your_embedding_api_key
EMBEDDING_BASE_URL=https://api.siliconflow.cn/v1
EMBEDDING_MODEL=Qwen/Qwen3-Embedding-0.6B
EMBEDDING_DIM=1024
# EMBEDDING_MAX_ASYNC=32 # [Optional, expert mode] Auto-calculated when unset: 800# MinerU Mode (Remote recommended)
MINERU_MODE=remote
MINERU_API_TOKEN=your_token
MINERU_HTTP_TIMEOUT=60 # MinerU download timeout (seconds, default 60)
FILE_SERVICE_BASE_URL=http://your-ip:8000
# VLM Chart Enhancement Configuration 🆕# ⚠️ Note: Only effective in MINERU_MODE=remote
RAG_VLM_MODE=off # off / selective / full
RAG_IMPORTANCE_THRESHOLD=0.5 # Importance threshold (selective mode)
RAG_CONTEXT_WINDOW=2 # Context window (full mode)
RAG_CONTEXT_MODE=page # page / chunk
RAG_MAX_CONTEXT_TOKENS=3000 # Max context tokens# Task Storage Configuration 🆕
TASK_STORE_STORAGE=redis # memory / redis (production recommends redis)# Document Insert Verification Configuration 🆕
DOC_INSERT_VERIFICATION_TIMEOUT=300 # Verification timeout (seconds, default 5 minutes)
DOC_INSERT_VERIFICATION_POLL_INTERVAL=0.5 # Poll interval (seconds, default 500ms)# Model Call Timeout Configuration 🆕
MODEL_CALL_TIMEOUT=90 # Model call max timeout (seconds, default 90)

⚡ Auto Concurrency Calculation:

  • LLM: When LLM_MAX_ASYNC is unset, auto-calculated as min(RPM, TPM/3500) ≈ 11
  • Embedding: When EMBEDDING_MAX_ASYNC is unset, auto-calculated as min(RPM, TPM/500) ≈ 800
  • Rerank: When RERANK_MAX_ASYNC is unset, auto-calculated as min(RPM, TPM/500) ≈ 800

✅ Recommended: Don't set *_MAX_ASYNC, let the system auto-calculate to completely avoid 429 errors

See env.example for complete configuration.


📚 API Documentation

Core Endpoints

1️⃣ Upload Document

# Single file upload (default mode)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc1" \
-F "file=@document.pdf" \
-F "parser=auto"# VLM chart enhancement mode 🆕# off: Markdown only (fastest, default)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc2&vlm_mode=off" \
-F "file=@document.pdf"# selective: Selective processing of important charts (balance performance and quality)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc3&vlm_mode=selective" \
-F "file=@document.pdf"# full: Complete RAG-Anything processing (highest quality, context enhancement enabled)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc4&vlm_mode=full" \
-F "file=@document.pdf"# Response
{
"task_id": "task-xxx-xxx",
"doc_id": "doc1",
"filename": "document.pdf",
"vlm_mode": "off",
"status": "pending"
}

2️⃣ Batch Upload

curl -X POST "http://localhost:8000/batch?tenant_id=your_tenant" \
-F "files=@doc1.pdf" \
-F "files=@doc2.docx" \
-F "files=@image.png"# Response
{
"batch_id": "batch-xxx-xxx",
"total_files": 3,
"accepted_files": 3,
"tasks": [...]
}

3️⃣ Intelligent Query (Query API v2.0)

New Advanced Features:

  • Conversation History: Support for multi-turn conversation context
  • Custom Prompts: Customize response style
  • Response Format Control: paragraph/list/json
  • Keyword Precision Retrieval: hl_keywords/ll_keywords
  • Streaming Output: Real-time generation viewing
# Basic query
curl -X POST "http://localhost:8000/query?tenant_id=your_tenant" \
-H "Content-Type: application/json" \
-d '{ "query": "What are the core viewpoints in the document?", "mode": "hybrid" }'# Advanced query (multi-turn dialogue + custom prompt)
curl -X POST "http://localhost:8000/query?tenant_id=your_tenant" \
-H "Content-Type: application/json" \
-d '{ "query": "Can you elaborate on the second point?", "mode": "hybrid", "conversation_history": [ {"role": "user", "content": "What are the key points?"}, {"role": "assistant", "content": "There are mainly three points..."} ], "user_prompt": "Please answer in professional academic language", "response_type": "list" }'# Streaming query (SSE)
curl -N -X POST "http://localhost:8000/query/stream?tenant_id=your_tenant" \
-H "Content-Type: application/json" \
-d '{ "query": "What are the core viewpoints in the document?", "mode": "hybrid" }'# Response (real-time streaming output)
data: {"chunk": "Based on", "done": false}
data: {"chunk": "document content", "done": false}
data: {"done": true}

4️⃣ Task Status Query

curl "http://localhost:8000/task/task-xxx-xxx?tenant_id=your_tenant"# Response
{
"task_id": "task-xxx-xxx",
"status": "completed",
"progress": 100,
"result": {...}
}

5️⃣ Tenant Management

# Get tenant statistics
curl "http://localhost:8000/tenants/stats?tenant_id=your_tenant"# Clear tenant cache
curl -X DELETE "http://localhost:8000/tenants/cache?tenant_id=your_tenant"# View instance pool status (admin)
curl "http://localhost:8000/tenants/pool/stats"

VLM Mode Comparison 🆕

ModeSpeedQualityResource UsageUse Case
off⚡⚡⚡⚡⚡⭐⭐⭐Very LowPlain text documents, fast batch processing
selective⚡⚡⚡⚡⭐⭐⭐⭐LowDocuments with key charts (recommended)
full⚡⚡⭐⭐⭐⭐⭐HighChart-intensive research reports, papers

Processing Time Estimate (20-page PDF example):

  • off: ~10 seconds(Markdown only)
  • selective: ~30 seconds(5-10 important charts)
  • full: ~120 seconds(complete context processing)

Query Mode Comparison

ModeSpeedAccuracyUse Case
naive⚡⚡⚡⚡⚡⭐⭐⭐Simple Q&A, fast retrieval
local⚡⚡⚡⚡⭐⭐⭐⭐Local entity relationship queries
global⚡⚡⚡⭐⭐⭐⭐Global knowledge graph reasoning
hybrid⚡⚡⚡⭐⭐⭐⭐⭐Hybrid retrieval (recommended)
mix⚡⚡⭐⭐⭐⭐⭐Complex questions, deep analysis

Query API v2.0 Advanced Parameters

ParameterTypeDescriptionExample
conversation_historyList[Dict]Multi-turn conversation context[{"role": "user", "content": "..."}]
user_promptstrCustom prompt"Please answer in professional academic language"
response_typestrResponse format"paragraph", "list", "json"
hl_keywordsList[str]High priority keywords["artificial intelligence", "machine learning"]
ll_keywordsList[str]Low priority keywords["application", "case study"]
only_need_contextboolReturn context only (debug)true
max_entity_tokensintEntity token limit6000

Complete API documentation:http://localhost:8000/docs


🎯 Usage Examples

Python SDK

importrequests# ConfigurationBASE_URL="http://localhost:8000"TENANT_ID="your_tenant"# Upload documentwithopen("document.pdf", "rb") asf:
response=requests.post(
f"{BASE_URL}/insert",
params={"tenant_id": TENANT_ID, "doc_id": "doc1"},
files={"file": f}
)
task_id=response.json()["task_id"]
print(f"Task ID: {task_id}")
# Queryresponse=requests.post(
f"{BASE_URL}/query",
params={"tenant_id": TENANT_ID},
json={
"query": "What is the main content of the document?",
"mode": "hybrid",
"top_k": 10
}
)
result=response.json()
print(f"Answer: {result['answer']}")

Complete cURL Example

# 1. Upload PDF document
TASK_ID=$(curl -X POST "http://localhost:8000/insert?tenant_id=demo&doc_id=report" \ -F "file=@report.pdf"| jq -r '.task_id')echo"Task ID: $TASK_ID"# 2. Wait for processing completionwhiletrue;do
STATUS=$(curl -s "http://localhost:8000/task/$TASK_ID?tenant_id=demo"| jq -r '.status')echo"Status: $STATUS"if [ "$STATUS"="completed" ] || [ "$STATUS"="failed" ];thenbreakfi
sleep 2
done# 3. Query document content
curl -X POST "http://localhost:8000/query?tenant_id=demo" \
-H "Content-Type: application/json" \
-d '{ "query": "What are the main conclusions of this report?", "mode": "hybrid" }'| jq '.answer'

🛠️ Deployment

System Requirements

Minimum Configuration:

  • CPU: 2 cores
  • RAM: 4GB
  • Disk: 40GB SSD
  • OS: Ubuntu 20.04+ / Debian 11+ / CentOS 8+

Recommended Configuration (Production):

  • CPU: 4 cores
  • RAM: 8GB
  • Disk: 100GB SSD
  • OS: Ubuntu 22.04 LTS

Server Deployment

Quick Deployment on Aliyun/Tencent Cloud

# SSH login to server
ssh root@your-server-ip
# Clone project
git clone https://github.com/BukeLy/rag-api.git
cd rag-api
# Run one-click deployment script
chmod +x deploy.sh
./deploy.sh
# The script will automatically:# 1. Install Docker and Docker Compose# 2. Configure environment variables# 3. Optimize system parameters# 4. Start services# 5. Verify health status

External Storage Configuration

Supports DragonflyDB + Qdrant + Memgraph external storage (enabled by default):

# Configure in .env
USE_EXTERNAL_STORAGE=true
# DragonflyDB configuration (KV Storage)
KV_STORAGE=RedisKVStorage
REDIS_URI=redis://dragonflydb:6379/0
# Qdrant configuration (vector storage)
VECTOR_STORAGE=QdrantVectorDBStorage
QDRANT_URL=http://qdrant:6333
# Memgraph configuration (graph storage)
GRAPH_STORAGE=MemgraphStorage
MEMGRAPH_URI=bolt://memgraph:7687
MEMGRAPH_USERNAME=
MEMGRAPH_PASSWORD=

See External Storage Deployment Documentation

Docker Compose Configuration

The project provides two configuration files:

FilePurposeFeatures
docker-compose.ymlProduction modeCode packaged in image, optimal performance
docker-compose.dev.ymlDevelopment modeCode mounted externally, supports hot-reload

Select configuration file:

# Production mode
docker compose -f docker-compose.yml up -d
# Development mode
docker compose -f docker-compose.dev.yml up -d

Performance Optimization

Tuning Parameters

Configure in .env:

# ⚡ Concurrency Control (Recommended: use auto-calculation)# LLM_MAX_ASYNC=8 # [Expert mode] Manually specify LLM concurrency# # Auto-calculated when unset: min(RPM, TPM/3500) ≈ 11# EMBEDDING_MAX_ASYNC=32 # [Expert mode] Manually specify Embedding concurrency# # Auto-calculated when unset: min(RPM, TPM/500) ≈ 800# RERANK_MAX_ASYNC=16 # [Expert mode] Manually specify Rerank concurrency# # Auto-calculated when unset: min(RPM, TPM/500) ≈ 800# Retrieval count (affects query quality and speed)
TOP_K=20 # Entity/relationship retrieval count
CHUNK_TOP_K=10 # Text chunk retrieval count# Document processing concurrency
DOCUMENT_PROCESSING_CONCURRENCY=10 # Remote mode can be set high, local mode set to 1

🎯 Concurrency Configuration Recommendations:

  • Recommended: Don't set *_MAX_ASYNC, let the system auto-calculate based on TPM/RPM
  • Expert mode: If manual control needed, can set LLM_MAX_ASYNC and other parameters
  • Advantage: Auto-calculation completely avoids 429 errors (TPM limit reached)

Mode Selection

  • MinerU Remote Mode (Recommended): High concurrency, resource-efficient
  • MinerU Local Mode: Requires GPU, high memory usage
  • Docling Mode: Fast and lightweight, suitable for simple documents

🏢 Multi-tenant Usage

Tenant Isolation

Each tenant has:

  • ✅ Independent LightRAG instance
  • ✅ Isolated data storage space
  • ✅ Independent vector index
  • ✅ Dedicated knowledge graph
  • Independent service configuration (LLM, Embedding, Rerank, DeepSeek-OCR, MinerU)🆕

Tenant Configuration Management 🆕

Each tenant can independently configure 5 services with hot-reload support:

# 1️⃣ Configure independent DeepSeek-OCR API key for Tenant A
curl -X PUT "http://localhost:8000/tenants/tenant_a/config" \
-H "Content-Type: application/json" \
-d '{ "ds_ocr_config": { "api_key": "sk-tenant-a-ds-ocr-key", "base_url": "https://api.siliconflow.cn/v1", "model": "deepseek-ai/DeepSeek-OCR", "timeout": 90 } }'# 2️⃣ Configure independent MinerU API token for Tenant B
curl -X PUT "http://localhost:8000/tenants/tenant_b/config" \
-H "Content-Type: application/json" \
-d '{ "mineru_config": { "api_token": "tenant-b-mineru-token", "base_url": "https://mineru.net", "model_version": "vlm" } }'# 3️⃣ Configure multiple services simultaneously (LLM + Embedding + DeepSeek-OCR)
curl -X PUT "http://localhost:8000/tenants/tenant_c/config" \
-H "Content-Type: application/json" \
-d '{ "llm_config": { "api_key": "sk-tenant-c-llm-key", "model": "gpt-4" }, "embedding_config": { "api_key": "sk-tenant-c-embedding-key", "model": "Qwen/Qwen3-Embedding-0.6B", "dim": 1024 }, "ds_ocr_config": { "api_key": "sk-tenant-c-ds-ocr-key" } }'# 4️⃣ Query tenant configuration (API key auto-masked)
curl "http://localhost:8000/tenants/tenant_a/config"# Response example
{
"tenant_id": "tenant_a",
"ds_ocr_config": {
"api_key": "sk-***-key", // Auto-masked
"timeout": 90
},
"merged_config": {
"llm": {...}, // Using Global Config
"embedding": {...}, // Using Global Config
"rerank": {...}, // Using Global Config
"ds_ocr": {...}, // Using tenant config
"mineru": {...} // Using Global Config
}
}
# 5️⃣ Refresh config cache (config hot-reload)
curl -X POST "http://localhost:8000/tenants/tenant_a/config/refresh"# 6️⃣ Delete tenant config (restore to global config)
curl -X DELETE "http://localhost:8000/tenants/tenant_a/config"

Supported Configuration Items:

ServiceConfig FieldDescription
LLMllm_configModel, API key, base_url, etc.
Embeddingembedding_configModel, API key, dimension, etc.
Rerankrerank_configModel, API key, etc.
DeepSeek-OCRds_ocr_configAPI key, timeout, mode, etc.
MinerUmineru_configAPI token, version, timeout, etc.

Configuration Priority: Tenant config > Global config

Use Cases:

  • 🔐 Multi-tenant SaaS: Each tenant uses their own API key
  • 💰 Pay-per-use: Track tenant usage through independent API keys
  • 🎯 Differentiated Services: Different tenants use different models (GPT-4 vs GPT-3.5)
  • 🧪 A/B Testing: Compare different models/parameters

Usage

All APIs require tenant_id parameter:

# Tenant A upload document
curl -X POST "http://localhost:8000/insert?tenant_id=tenant_a&doc_id=doc1" \
-F "file=@doc.pdf"# Tenant B upload document (fully isolated)
curl -X POST "http://localhost:8000/insert?tenant_id=tenant_b&doc_id=doc1" \
-F "file=@doc.pdf"# Tenant A query (can only query own documents)
curl -X POST "http://localhost:8000/query?tenant_id=tenant_a" \
-H "Content-Type: application/json" \
-d '{"query": "document content", "mode": "hybrid"}'

Instance Pool Management

  • Capacity: Cache up to 50 tenant instances
  • Strategy: LRU (Least Recently Used) automatic cleanup
  • Config Isolation: Each tenant can use independent LLM, Embedding, parser configuration

📊 Monitoring & Maintenance

Common Commands

# View service status
docker compose ps
# View real-time logs
docker compose logs -f
# Restart services
docker compose restart
# Stop services
docker compose down
# View resource usage
docker stats
# Clean Docker resources
docker system prune -f

Maintenance Scripts

# Monitor service health
./scripts/monitor.sh
# Backup data
./scripts/backup.sh
# Update services
./scripts/update.sh
# Performance testing
./scripts/test_concurrent_perf.sh
# Performance monitoring
./scripts/monitor_performance.sh

Health Checks

# Complete health check (recommended)
./scripts/health_check.sh
./scripts/health_check.sh --verbose # verbose output# API health check
curl http://localhost:8000/
# Tenant statistics
curl "http://localhost:8000/tenants/stats?tenant_id=your_tenant"# Instance pool status
curl "http://localhost:8000/tenants/pool/stats"

🗂️ Project Structure

rag-api/
├── main.py # FastAPI application entry
├── api/ # API route modules
│ ├── __init__.py # Route aggregation
│ ├── insert.py # Document upload (single/batch)
│ ├── query.py # Intelligent query
│ ├── task.py # Task status query
│ ├── tenant.py # Tenant management
│ ├── files.py # File service
│ ├── models.py # Pydantic models
│ └── task_store.py # Task storage
├── src/ # Core business logic
│ ├── rag.py # LightRAG lifecycle management
│ ├── multi_tenant.py # Multi-tenant instance manager
│ ├── tenant_deps.py # Tenant dependency injection
│ ├── logger.py # Unified logging
│ ├── metrics.py # Performance metrics
│ ├── file_url_service.py # Temporary file service
│ ├── mineru_client.py # MinerU client
│ └── mineru_result_processor.py # Result processing
├── docs/ # Documentation
│ ├── ARCHITECTURE.md # Architecture design documentation
│ ├── USAGE.md # Detailed usage guide
│ ├── DEPLOY_MODES.md # Deployment mode description
│ ├── PR_WORKFLOW.md # PR workflow
│ └── ...
├── scripts/ # Maintenance scripts
│ ├── dev.sh # Development mode quick start
│ ├── monitor.sh # Service monitoring
│ ├── backup.sh # Data backup
│ ├── update.sh # Service update
│ └── ...
├── deploy.sh # One-click deployment script
├── docker-compose.yml # Production mode configuration
├── docker-compose.dev.yml # Development mode configuration
├── Dockerfile # Production image
├── Dockerfile.dev # Development image
├── pyproject.toml # Project dependencies
├── uv.lock # Dependency lock
├── env.example # Environment variable template
├── CLAUDE.md # Claude AI guide
└── README.md # This documentation

🐛 Troubleshooting

Common Issues

Q1: What to do if service fails to start?
# View detailed logs
docker compose logs
# Check port usage
netstat -tulpn | grep 8000
# Check Docker status
docker ps -a
Q2: multimodal_processed error?

Note: This issue has been fixed in LightRAG 1.4.9.4+. If you encounter this error, your version is outdated.

Solution:

# Option 1: Upgrade to latest version (recommended)# Modify LightRAG version in pyproject.toml# lightrag = "^1.4.9.4"# Rebuild image
docker compose down
docker compose up -d --build
# Option 2: Clean old data (temporary solution)
rm -rf ./rag_local_storage
docker compose restart
Q3: File upload returns 400 error?

Check:

  • File format supported (PDF, DOCX, PNG, JPG, etc.)
  • File size exceeds 100MB
  • File is empty
# View supported formats
curl http://localhost:8000/docs
Q3.5: Embedding dimension error?

If you encounter dimension-related errors, need to clean data and rebuild:

# Stop services
docker compose down
# Delete all volumes (clear database)
docker volume rm rag-api_dragonflydb_data rag-api_qdrant_data rag-api_memgraph_data
# Modify EMBEDDING_DIM in .env
EMBEDDING_DIM=1024 # or 4096, must match the model# Restart
docker compose up -d
Q4: Query is very slow (>30 seconds)?

Optimization suggestions:

  1. Use naive or hybrid mode instead of mix
  2. Increase MAX_ASYNC parameter (in .env)
  3. Reduce TOP_K and CHUNK_TOP_K
  4. Enable Reranker
# Modify .env
MAX_ASYNC=8
TOP_K=20
CHUNK_TOP_K=10
Q5: Out of memory (OOM)?

If using local MinerU:

# Switch to remote mode# Modify in .env
MINERU_MODE=remote
MINERU_API_TOKEN=your_token
# Or limit concurrency
DOCUMENT_PROCESSING_CONCURRENCY=1
Q6: Tasks lost after container restart?

Problem Symptoms:

  • Cannot query previous task status after container restart
  • Tasks disappear after tenant instance evicted by LRU

Solution: Enable Redis task storage

# Modify .env
TASK_STORE_STORAGE=redis
# Restart services
docker compose restart
# Verify
docker compose logs api | grep TaskStore
# Should see: ✅ TaskStore: Redis connection successful

Configuration Description:

  • memory mode: In-memory storage, data lost after restart (default, suitable for development)
  • redis mode: Persistent storage, supports container restart and instance rebuild (production recommended)

TTL Strategy (Redis mode auto-cleanup):

  • completed tasks: 24 hours
  • failed tasks: 24 hours
  • pending/processing tasks: 6 hours
Q7: VLM mode processing failed?

Check Items:

  1. vision_model_func not configured

    • Check logs:vision_model_func not found, fallback to off mode
    • Ensure LLM API is configured in .env
  2. Image file does not exist

    • Check logs:Image file not found: xxx
    • Possibly corrupted MinerU ZIP or extraction failed
  3. Timeout error

    • full mode may timeout on large files
    • Suggestion: Use selective mode first, or increase VLM_TIMEOUT
# Modify .env
VLM_TIMEOUT=300 # Increase to 5 minutes
RAG_VLM_MODE=selective # downgrade to selective

Debugging Tips:

# View detailed logs
docker compose logs -f | grep VLM
# Test single file
curl -X POST 'http://localhost:8000/insert?tenant_id=test&doc_id=test&vlm_mode=off' \
-F 'file=@test.pdf'

Performance Tuning Recommendations

ScenarioMAX_ASYNCTOP_KCHUNK_TOP_KMINERU_MODE
Fast response8105remote
Balanced mode82010remote
High accuracy46020remote
Resource limited42010remote

📖 Documentation


🤝 Contributing

We welcome all forms of contribution!

How to Contribute

  1. Fork the project
git clone https://github.com/BukeLy/rag-api.git
cd rag-api
  1. Create feature branch
git checkout -b feature/your-feature-name
  1. Development and Testing
# Install dependencies
uv sync
# Run tests
uv run pytest
# Code formatting
uv run black .
uv run isort .
  1. Submit code
git add .
git commit -m "feat: Add new feature"
git push origin feature/your-feature-name
  1. Create Pull Request

Create a PR on GitHub with detailed description of your changes.

Commit Conventions

Use semantic commit messages:

  • feat: New feature
  • fix: Bug fix
  • docs: Documentation update
  • style: Code formatting
  • refactor: Code refactoring
  • perf: Performance optimization
  • test: Testing
  • chore: Build/tools

See PR Workflow Documentation


📄 License

This project is licensed under the MIT License. See the LICENSE file for details.


🙏 Acknowledgments

This project is built on the following excellent open source projects:

  • LightRAG - Efficient knowledge graph RAG framework
  • RAG-Anything - Multimodal document parsing
  • MinerU - Powerful PDF parsing tool
  • Docling - Lightweight document parsing
  • FastAPI - Modern Python web framework

Special thanks to all contributors and users for their support! 🎉


📬 Contact Us


⭐ If this project helps you, please give it a Star!

Made with ❤️ by BukeLy

© 2025 RAG API. All rights reserved.

About

Multi-tenant RAG API powered by LightRAG/RAG-Anything. Auto-selects best parser (DeepSeek-OCR/MinerU/Docling) via complexity scoring

Topics

Resources

Stars

64 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

🚀 RAG API

Multi-tenant Multimodal Document Intelligent Retrieval System

Enterprise-grade RAG service built on RAG-Anything and LightRAG

CIPythonFastAPILightRAGDockerLicense

English | 简体中文

FeaturesQuick StartArchitectureAPI DocumentationDeployment


📖 Introduction

RAG API is an enterprise-grade Retrieval-Augmented Generation (RAG) service that combines the powerful document parsing capabilities of RAG-Anything with the efficient knowledge graph retrieval technology of LightRAG, providing intelligent Q&A capabilities for your documents.

🎯 Key Highlights

  • 🏢 Multi-tenant Isolation - Complete tenant data isolation for enterprise multi-tenant scenarios
  • 🎨 Multimodal Parsing - Support for PDF, Word, images and more, with full OCR, tables, and formulas coverage
  • High-performance Retrieval - Knowledge graph-based hybrid retrieval with 6-15 second query response
  • 🔄 Flexible Deployment - Support for production and development modes with one-click switching
  • 📦 Ready to Use - One-click Docker deployment, service starts in 3 minutes
  • 🎛️ Multiple Parsing Engines - DeepSeek-OCR (Remote API) + MinerU (Local/Remote API) + Docling (Fast)
  • 🎨 RAG-Anything VLM Enhancement - Three modes (off/selective/full) for deep chart understanding
  • 💾 Task Persistence - Redis storage support, tasks recoverable after container restart/instance rebuild

✨ Features

📄 Document Processing

  • Multiple Format Support

    • PDF, Word, Excel, PPT
    • PNG, JPG, WebP images
    • TXT, Markdown text
  • Intelligent Parsing

    • Plain text (.txt, .md) → Direct insertion (ultra-fast ~1s, skip parser)
    • OCR text recognition
    • Structured table extraction
    • Mathematical formula recognition
    • Layout analysis
  • RAG-Anything VLM Enhancement 🆕

    • off - Markdown only (fastest)
    • selective - Selective processing of important charts
    • full - Complete context enhancement processing
    • Smart filtering: with titles, large size, first page content
    • ⚠️Only supports remote MinerU mode, local mode uses RAG-Anything native methods
  • Batch Processing

    • Up to 100 files per batch
    • Async task queue
    • Real-time progress tracking

🔍 Intelligent Retrieval

  • Multi-mode Query

    • naive - Vector retrieval (fastest)
    • local - Local graph
    • global - Global graph
    • hybrid - Hybrid retrieval
    • mix - Full retrieval (most accurate)
  • Knowledge Graph

    • Automatic entity extraction
    • Relationship reasoning
    • Semantic understanding
    • Context enhancement
  • External Storage

    • DragonflyDB (KV storage + task storage)
    • Qdrant (vector storage)
    • Memgraph (graph database)
    • Task persistence (Redis mode)

🏗️ Architecture

System Architecture Diagram

graph TB
subgraph "Client Layer"
Client[Client Application]
WebUI[Web Interface]
end
subgraph "API Gateway Layer"
FastAPI[FastAPI Service]
Auth[Tenant Authentication]
end
subgraph "Business Logic Layer"
TenantMgr[Tenant Manager]
TaskQueue[Task Queue]
subgraph "Document Processing"
DeepSeekOCR[DeepSeek-OCR<br/>Fast OCR 80% cases]
MinerU[MinerU Parser<br/>Complex multimodal]
Docling[Docling Parser<br/>Fast lightweight]
FileRouter[Smart Router<br/>Complexity scoring]
end
subgraph "RAG Engine"
LightRAG[LightRAG Instance Pool<br/>LRU Cache 50]
KG[Knowledge Graph Engine]
Vector[Vector Retrieval Engine]
end
end
subgraph "Storage Layer"
DragonflyDB[(DragonflyDB<br/>KV Storage)]
Qdrant[(Qdrant<br/>Vector Database)]
Memgraph[(Memgraph<br/>Graph Database)]
Local[(Local Files<br/>Temp Storage)]
end
subgraph "External Services"
LLM[LLM<br/>Entity Extraction/Generation]
Embedding[Embedding<br/>Vectorization]
Rerank[Rerank<br/>Reranking]
end
Client --> FastAPI
WebUI --> FastAPI
FastAPI --> Auth
Auth --> TenantMgr
TenantMgr --> TaskQueue
TenantMgr --> LightRAG
TaskQueue --> FileRouter
FileRouter --> DeepSeekOCR
FileRouter --> MinerU
FileRouter --> Docling
DeepSeekOCR --> LightRAG
MinerU --> LightRAG
Docling --> LightRAG
LightRAG --> KG
LightRAG --> Vector
KG --> DragonflyDB
KG --> Memgraph
Vector --> Qdrant
LightRAG --> Local
LightRAG --> LLM
LightRAG --> Embedding
Vector --> Rerank
style FastAPI fill:#00C7B7
style LightRAG fill:#FF6B6B
style DeepSeekOCR fill:#5DADE2
style MinerU fill:#4ECDC4
style Docling fill:#95E1D3
style TenantMgr fill:#F38181
Loading

Multi-tenant Architecture

graph TB
subgraph "Tenant A"
A_Config[Tenant A Config<br/>Independent API Key]
A_Instance[LightRAG Instance A<br/>Dedicated LLM/Embedding]
A_Data[(Tenant A Data<br/>Fully Isolated)]
A_Config --> A_Instance
A_Instance --> A_Data
end
subgraph "Tenant B"
B_Config[Tenant B Config<br/>Independent API Key]
B_Instance[LightRAG Instance B<br/>Dedicated LLM/Embedding]
B_Data[(Tenant B Data<br/>Fully Isolated)]
B_Config --> B_Instance
B_Instance --> B_Data
end
subgraph "Tenant C"
C_Config[Using Global Config]
C_Instance[LightRAG Instance C<br/>Shared LLM/Embedding]
C_Data[(Tenant C Data<br/>Fully Isolated)]
C_Config --> C_Instance
C_Instance --> C_Data
end
Pool[Instance Pool Manager<br/>LRU Cache + Config Isolation]
Global[Global Config<br/>Default API Key]
Pool --> A_Instance
Pool --> B_Instance
Pool --> C_Instance
C_Config -.fallback.-> Global
style Pool fill:#F38181
style Global fill:#95E1D3
style A_Config fill:#FFD93D
style B_Config fill:#FFD93D
style C_Config fill:#E8E8E8
Loading

Core Technology Stack

🔧 Frameworks & Runtime

  • FastAPI 0.115+
  • Python 3.11+
  • Uvicorn
  • Docker & Docker Compose

🧠 AI & RAG

  • LightRAG 1.4.9.4
  • RAG-Anything
  • MinerU (PDF-Extract-Kit)
  • Docling

💾 Storage & Database

  • DragonflyDB(Redis compatible)
  • Qdrant(Vector Database)
  • Memgraph(Graph Database)
  • Local filesystem

🚀 Quick Start

Option 1: One-click Deployment (Recommended)

Suitable for production and testing environments:

# 1. Clone the project
git clone https://github.com/BukeLy/rag-api.git
cd rag-api
# 2. Configure environment variables
cp env.example .env
nano .env # Fill in your API keys# 3. Run deployment script
chmod +x deploy.sh
./deploy.sh
# Select deployment mode:# 1) Production Mode - Standard container deployment# 2) Development Mode - Code hot-reload# 4. Verify service
curl http://localhost:8000/

Access Swagger Documentation:http://localhost:8000/docs

Option 2: Docker Compose

Production Mode

# Configure environment variables
cp env.example .env
nano .env
# Start services
docker compose -f docker-compose.yml up -d
# View logs
docker compose -f docker-compose.yml logs -f

Development Mode (Code Hot-reload)

# Start development environment
docker compose -f docker-compose.dev.yml up -d
# Or use quick script
./scripts/dev.sh
# Code changes will auto-reload without restart

Option 3: Local Development

# Install uv (Python package manager)
curl -LsSf https://astral.sh/uv/install.sh | sh
# Install dependencies
uv sync
# Configure environment variables
cp env.example .env
nano .env
# Start services
uv run uvicorn main:app --host 0.0.0.0 --port 8000 --reload

Environment Variable Configuration

Minimum configuration (required):

# LLM Configuration (Function-oriented naming)
LLM_API_KEY=your_llm_api_key
LLM_BASE_URL=https://ark.cn-beijing.volces.com/api/v3
LLM_MODEL=ep-xxx-xxx
# LLM_REQUESTS_PER_MINUTE=800 # Rate limit (optional)# LLM_TOKENS_PER_MINUTE=40000 # Rate limit (optional)# LLM_MAX_ASYNC=8 # [Optional, expert mode] Manual concurrency control# # Auto-calculated when unset: min(RPM, TPM/3500) = 11# Embedding Configuration (Function-oriented naming)
EMBEDDING_API_KEY=your_embedding_api_key
EMBEDDING_BASE_URL=https://api.siliconflow.cn/v1
EMBEDDING_MODEL=Qwen/Qwen3-Embedding-0.6B
EMBEDDING_DIM=1024
# EMBEDDING_MAX_ASYNC=32 # [Optional, expert mode] Auto-calculated when unset: 800# MinerU Mode (Remote recommended)
MINERU_MODE=remote
MINERU_API_TOKEN=your_token
MINERU_HTTP_TIMEOUT=60 # MinerU download timeout (seconds, default 60)
FILE_SERVICE_BASE_URL=http://your-ip:8000
# VLM Chart Enhancement Configuration 🆕# ⚠️ Note: Only effective in MINERU_MODE=remote
RAG_VLM_MODE=off # off / selective / full
RAG_IMPORTANCE_THRESHOLD=0.5 # Importance threshold (selective mode)
RAG_CONTEXT_WINDOW=2 # Context window (full mode)
RAG_CONTEXT_MODE=page # page / chunk
RAG_MAX_CONTEXT_TOKENS=3000 # Max context tokens# Task Storage Configuration 🆕
TASK_STORE_STORAGE=redis # memory / redis (production recommends redis)# Document Insert Verification Configuration 🆕
DOC_INSERT_VERIFICATION_TIMEOUT=300 # Verification timeout (seconds, default 5 minutes)
DOC_INSERT_VERIFICATION_POLL_INTERVAL=0.5 # Poll interval (seconds, default 500ms)# Model Call Timeout Configuration 🆕
MODEL_CALL_TIMEOUT=90 # Model call max timeout (seconds, default 90)

⚡ Auto Concurrency Calculation:

  • LLM: When LLM_MAX_ASYNC is unset, auto-calculated as min(RPM, TPM/3500) ≈ 11
  • Embedding: When EMBEDDING_MAX_ASYNC is unset, auto-calculated as min(RPM, TPM/500) ≈ 800
  • Rerank: When RERANK_MAX_ASYNC is unset, auto-calculated as min(RPM, TPM/500) ≈ 800

✅ Recommended: Don't set *_MAX_ASYNC, let the system auto-calculate to completely avoid 429 errors

See env.example for complete configuration.


📚 API Documentation

Core Endpoints

1️⃣ Upload Document

# Single file upload (default mode)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc1" \
-F "file=@document.pdf" \
-F "parser=auto"# VLM chart enhancement mode 🆕# off: Markdown only (fastest, default)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc2&vlm_mode=off" \
-F "file=@document.pdf"# selective: Selective processing of important charts (balance performance and quality)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc3&vlm_mode=selective" \
-F "file=@document.pdf"# full: Complete RAG-Anything processing (highest quality, context enhancement enabled)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc4&vlm_mode=full" \
-F "file=@document.pdf"# Response
{
"task_id": "task-xxx-xxx",
"doc_id": "doc1",
"filename": "document.pdf",
"vlm_mode": "off",
"status": "pending"
}

2️⃣ Batch Upload

curl -X POST "http://localhost:8000/batch?tenant_id=your_tenant" \
-F "files=@doc1.pdf" \
-F "files=@doc2.docx" \
-F "files=@image.png"# Response
{
"batch_id": "batch-xxx-xxx",
"total_files": 3,
"accepted_files": 3,
"tasks": [...]
}

3️⃣ Intelligent Query (Query API v2.0)

New Advanced Features:

  • Conversation History: Support for multi-turn conversation context
  • Custom Prompts: Customize response style
  • Response Format Control: paragraph/list/json
  • Keyword Precision Retrieval: hl_keywords/ll_keywords
  • Streaming Output: Real-time generation viewing
# Basic query
curl -X POST "http://localhost:8000/query?tenant_id=your_tenant" \
-H "Content-Type: application/json" \
-d '{ "query": "What are the core viewpoints in the document?", "mode": "hybrid" }'# Advanced query (multi-turn dialogue + custom prompt)
curl -X POST "http://localhost:8000/query?tenant_id=your_tenant" \
-H "Content-Type: application/json" \
-d '{ "query": "Can you elaborate on the second point?", "mode": "hybrid", "conversation_history": [ {"role": "user", "content": "What are the key points?"}, {"role": "assistant", "content": "There are mainly three points..."} ], "user_prompt": "Please answer in professional academic language", "response_type": "list" }'# Streaming query (SSE)
curl -N -X POST "http://localhost:8000/query/stream?tenant_id=your_tenant" \
-H "Content-Type: application/json" \
-d '{ "query": "What are the core viewpoints in the document?", "mode": "hybrid" }'# Response (real-time streaming output)
data: {"chunk": "Based on", "done": false}
data: {"chunk": "document content", "done": false}
data: {"done": true}

4️⃣ Task Status Query

curl "http://localhost:8000/task/task-xxx-xxx?tenant_id=your_tenant"# Response
{
"task_id": "task-xxx-xxx",
"status": "completed",
"progress": 100,
"result": {...}
}

5️⃣ Tenant Management

# Get tenant statistics
curl "http://localhost:8000/tenants/stats?tenant_id=your_tenant"# Clear tenant cache
curl -X DELETE "http://localhost:8000/tenants/cache?tenant_id=your_tenant"# View instance pool status (admin)
curl "http://localhost:8000/tenants/pool/stats"

VLM Mode Comparison 🆕

ModeSpeedQualityResource UsageUse Case
off⚡⚡⚡⚡⚡⭐⭐⭐Very LowPlain text documents, fast batch processing
selective⚡⚡⚡⚡⭐⭐⭐⭐LowDocuments with key charts (recommended)
full⚡⚡⭐⭐⭐⭐⭐HighChart-intensive research reports, papers

Processing Time Estimate (20-page PDF example):

  • off: ~10 seconds(Markdown only)
  • selective: ~30 seconds(5-10 important charts)
  • full: ~120 seconds(complete context processing)

Query Mode Comparison

ModeSpeedAccuracyUse Case
naive⚡⚡⚡⚡⚡⭐⭐⭐Simple Q&A, fast retrieval
local⚡⚡⚡⚡⭐⭐⭐⭐Local entity relationship queries
global⚡⚡⚡⭐⭐⭐⭐Global knowledge graph reasoning
hybrid⚡⚡⚡⭐⭐⭐⭐⭐Hybrid retrieval (recommended)
mix⚡⚡⭐⭐⭐⭐⭐Complex questions, deep analysis

Query API v2.0 Advanced Parameters

ParameterTypeDescriptionExample
conversation_historyList[Dict]Multi-turn conversation context[{"role": "user", "content": "..."}]
user_promptstrCustom prompt"Please answer in professional academic language"
response_typestrResponse format"paragraph", "list", "json"
hl_keywordsList[str]High priority keywords["artificial intelligence", "machine learning"]
ll_keywordsList[str]Low priority keywords["application", "case study"]
only_need_contextboolReturn context only (debug)true
max_entity_tokensintEntity token limit6000

Complete API documentation:http://localhost:8000/docs


🎯 Usage Examples

Python SDK

importrequests# ConfigurationBASE_URL="http://localhost:8000"TENANT_ID="your_tenant"# Upload documentwithopen("document.pdf", "rb") asf:
response=requests.post(
f"{BASE_URL}/insert",
params={"tenant_id": TENANT_ID, "doc_id": "doc1"},
files={"file": f}
)
task_id=response.json()["task_id"]
print(f"Task ID: {task_id}")
# Queryresponse=requests.post(
f"{BASE_URL}/query",
params={"tenant_id": TENANT_ID},
json={
"query": "What is the main content of the document?",
"mode": "hybrid",
"top_k": 10
}
)
result=response.json()
print(f"Answer: {result['answer']}")

Complete cURL Example

# 1. Upload PDF document
TASK_ID=$(curl -X POST "http://localhost:8000/insert?tenant_id=demo&doc_id=report" \ -F "file=@report.pdf"| jq -r '.task_id')echo"Task ID: $TASK_ID"# 2. Wait for processing completionwhiletrue;do
STATUS=$(curl -s "http://localhost:8000/task/$TASK_ID?tenant_id=demo"| jq -r '.status')echo"Status: $STATUS"if [ "$STATUS"="completed" ] || [ "$STATUS"="failed" ];thenbreakfi
sleep 2
done# 3. Query document content
curl -X POST "http://localhost:8000/query?tenant_id=demo" \
-H "Content-Type: application/json" \
-d '{ "query": "What are the main conclusions of this report?", "mode": "hybrid" }'| jq '.answer'

🛠️ Deployment

System Requirements

Minimum Configuration:

  • CPU: 2 cores
  • RAM: 4GB
  • Disk: 40GB SSD
  • OS: Ubuntu 20.04+ / Debian 11+ / CentOS 8+

Recommended Configuration (Production):

  • CPU: 4 cores
  • RAM: 8GB
  • Disk: 100GB SSD
  • OS: Ubuntu 22.04 LTS

Server Deployment

Quick Deployment on Aliyun/Tencent Cloud

# SSH login to server
ssh root@your-server-ip
# Clone project
git clone https://github.com/BukeLy/rag-api.git
cd rag-api
# Run one-click deployment script
chmod +x deploy.sh
./deploy.sh
# The script will automatically:# 1. Install Docker and Docker Compose# 2. Configure environment variables# 3. Optimize system parameters# 4. Start services# 5. Verify health status

External Storage Configuration

Supports DragonflyDB + Qdrant + Memgraph external storage (enabled by default):

# Configure in .env
USE_EXTERNAL_STORAGE=true
# DragonflyDB configuration (KV Storage)
KV_STORAGE=RedisKVStorage
REDIS_URI=redis://dragonflydb:6379/0
# Qdrant configuration (vector storage)
VECTOR_STORAGE=QdrantVectorDBStorage
QDRANT_URL=http://qdrant:6333
# Memgraph configuration (graph storage)
GRAPH_STORAGE=MemgraphStorage
MEMGRAPH_URI=bolt://memgraph:7687
MEMGRAPH_USERNAME=
MEMGRAPH_PASSWORD=

See External Storage Deployment Documentation

Docker Compose Configuration

The project provides two configuration files:

FilePurposeFeatures
docker-compose.ymlProduction modeCode packaged in image, optimal performance
docker-compose.dev.ymlDevelopment modeCode mounted externally, supports hot-reload

Select configuration file:

# Production mode
docker compose -f docker-compose.yml up -d
# Development mode
docker compose -f docker-compose.dev.yml up -d

Performance Optimization

Tuning Parameters

Configure in .env:

# ⚡ Concurrency Control (Recommended: use auto-calculation)# LLM_MAX_ASYNC=8 # [Expert mode] Manually specify LLM concurrency# # Auto-calculated when unset: min(RPM, TPM/3500) ≈ 11# EMBEDDING_MAX_ASYNC=32 # [Expert mode] Manually specify Embedding concurrency# # Auto-calculated when unset: min(RPM, TPM/500) ≈ 800# RERANK_MAX_ASYNC=16 # [Expert mode] Manually specify Rerank concurrency# # Auto-calculated when unset: min(RPM, TPM/500) ≈ 800# Retrieval count (affects query quality and speed)
TOP_K=20 # Entity/relationship retrieval count
CHUNK_TOP_K=10 # Text chunk retrieval count# Document processing concurrency
DOCUMENT_PROCESSING_CONCURRENCY=10 # Remote mode can be set high, local mode set to 1

🎯 Concurrency Configuration Recommendations:

  • Recommended: Don't set *_MAX_ASYNC, let the system auto-calculate based on TPM/RPM
  • Expert mode: If manual control needed, can set LLM_MAX_ASYNC and other parameters
  • Advantage: Auto-calculation completely avoids 429 errors (TPM limit reached)

Mode Selection

  • MinerU Remote Mode (Recommended): High concurrency, resource-efficient
  • MinerU Local Mode: Requires GPU, high memory usage
  • Docling Mode: Fast and lightweight, suitable for simple documents

🏢 Multi-tenant Usage

Tenant Isolation

Each tenant has:

  • ✅ Independent LightRAG instance
  • ✅ Isolated data storage space
  • ✅ Independent vector index
  • ✅ Dedicated knowledge graph
  • Independent service configuration (LLM, Embedding, Rerank, DeepSeek-OCR, MinerU)🆕

Tenant Configuration Management 🆕

Each tenant can independently configure 5 services with hot-reload support:

# 1️⃣ Configure independent DeepSeek-OCR API key for Tenant A
curl -X PUT "http://localhost:8000/tenants/tenant_a/config" \
-H "Content-Type: application/json" \
-d '{ "ds_ocr_config": { "api_key": "sk-tenant-a-ds-ocr-key", "base_url": "https://api.siliconflow.cn/v1", "model": "deepseek-ai/DeepSeek-OCR", "timeout": 90 } }'# 2️⃣ Configure independent MinerU API token for Tenant B
curl -X PUT "http://localhost:8000/tenants/tenant_b/config" \
-H "Content-Type: application/json" \
-d '{ "mineru_config": { "api_token": "tenant-b-mineru-token", "base_url": "https://mineru.net", "model_version": "vlm" } }'# 3️⃣ Configure multiple services simultaneously (LLM + Embedding + DeepSeek-OCR)
curl -X PUT "http://localhost:8000/tenants/tenant_c/config" \
-H "Content-Type: application/json" \
-d '{ "llm_config": { "api_key": "sk-tenant-c-llm-key", "model": "gpt-4" }, "embedding_config": { "api_key": "sk-tenant-c-embedding-key", "model": "Qwen/Qwen3-Embedding-0.6B", "dim": 1024 }, "ds_ocr_config": { "api_key": "sk-tenant-c-ds-ocr-key" } }'# 4️⃣ Query tenant configuration (API key auto-masked)
curl "http://localhost:8000/tenants/tenant_a/config"# Response example
{
"tenant_id": "tenant_a",
"ds_ocr_config": {
"api_key": "sk-***-key", // Auto-masked
"timeout": 90
},
"merged_config": {
"llm": {...}, // Using Global Config
"embedding": {...}, // Using Global Config
"rerank": {...}, // Using Global Config
"ds_ocr": {...}, // Using tenant config
"mineru": {...} // Using Global Config
}
}
# 5️⃣ Refresh config cache (config hot-reload)
curl -X POST "http://localhost:8000/tenants/tenant_a/config/refresh"# 6️⃣ Delete tenant config (restore to global config)
curl -X DELETE "http://localhost:8000/tenants/tenant_a/config"

Supported Configuration Items:

ServiceConfig FieldDescription
LLMllm_configModel, API key, base_url, etc.
Embeddingembedding_configModel, API key, dimension, etc.
Rerankrerank_configModel, API key, etc.
DeepSeek-OCRds_ocr_configAPI key, timeout, mode, etc.
MinerUmineru_configAPI token, version, timeout, etc.

Configuration Priority: Tenant config > Global config

Use Cases:

  • 🔐 Multi-tenant SaaS: Each tenant uses their own API key
  • 💰 Pay-per-use: Track tenant usage through independent API keys
  • 🎯 Differentiated Services: Different tenants use different models (GPT-4 vs GPT-3.5)
  • 🧪 A/B Testing: Compare different models/parameters

Usage

All APIs require tenant_id parameter:

# Tenant A upload document
curl -X POST "http://localhost:8000/insert?tenant_id=tenant_a&doc_id=doc1" \
-F "file=@doc.pdf"# Tenant B upload document (fully isolated)
curl -X POST "http://localhost:8000/insert?tenant_id=tenant_b&doc_id=doc1" \
-F "file=@doc.pdf"# Tenant A query (can only query own documents)
curl -X POST "http://localhost:8000/query?tenant_id=tenant_a" \
-H "Content-Type: application/json" \
-d '{"query": "document content", "mode": "hybrid"}'

Instance Pool Management

  • Capacity: Cache up to 50 tenant instances
  • Strategy: LRU (Least Recently Used) automatic cleanup
  • Config Isolation: Each tenant can use independent LLM, Embedding, parser configuration

📊 Monitoring & Maintenance

Common Commands

# View service status
docker compose ps
# View real-time logs
docker compose logs -f
# Restart services
docker compose restart
# Stop services
docker compose down
# View resource usage
docker stats
# Clean Docker resources
docker system prune -f

Maintenance Scripts

# Monitor service health
./scripts/monitor.sh
# Backup data
./scripts/backup.sh
# Update services
./scripts/update.sh
# Performance testing
./scripts/test_concurrent_perf.sh
# Performance monitoring
./scripts/monitor_performance.sh

Health Checks

# Complete health check (recommended)
./scripts/health_check.sh
./scripts/health_check.sh --verbose # verbose output# API health check
curl http://localhost:8000/
# Tenant statistics
curl "http://localhost:8000/tenants/stats?tenant_id=your_tenant"# Instance pool status
curl "http://localhost:8000/tenants/pool/stats"

🗂️ Project Structure

rag-api/
├── main.py # FastAPI application entry
├── api/ # API route modules
│ ├── __init__.py # Route aggregation
│ ├── insert.py # Document upload (single/batch)
│ ├── query.py # Intelligent query
│ ├── task.py # Task status query
│ ├── tenant.py # Tenant management
│ ├── files.py # File service
│ ├── models.py # Pydantic models
│ └── task_store.py # Task storage
├── src/ # Core business logic
│ ├── rag.py # LightRAG lifecycle management
│ ├── multi_tenant.py # Multi-tenant instance manager
│ ├── tenant_deps.py # Tenant dependency injection
│ ├── logger.py # Unified logging
│ ├── metrics.py # Performance metrics
│ ├── file_url_service.py # Temporary file service
│ ├── mineru_client.py # MinerU client
│ └── mineru_result_processor.py # Result processing
├── docs/ # Documentation
│ ├── ARCHITECTURE.md # Architecture design documentation
│ ├── USAGE.md # Detailed usage guide
│ ├── DEPLOY_MODES.md # Deployment mode description
│ ├── PR_WORKFLOW.md # PR workflow
│ └── ...
├── scripts/ # Maintenance scripts
│ ├── dev.sh # Development mode quick start
│ ├── monitor.sh # Service monitoring
│ ├── backup.sh # Data backup
│ ├── update.sh # Service update
│ └── ...
├── deploy.sh # One-click deployment script
├── docker-compose.yml # Production mode configuration
├── docker-compose.dev.yml # Development mode configuration
├── Dockerfile # Production image
├── Dockerfile.dev # Development image
├── pyproject.toml # Project dependencies
├── uv.lock # Dependency lock
├── env.example # Environment variable template
├── CLAUDE.md # Claude AI guide
└── README.md # This documentation

🐛 Troubleshooting

Common Issues

Q1: What to do if service fails to start?
# View detailed logs
docker compose logs
# Check port usage
netstat -tulpn | grep 8000
# Check Docker status
docker ps -a
Q2: multimodal_processed error?

Note: This issue has been fixed in LightRAG 1.4.9.4+. If you encounter this error, your version is outdated.

Solution:

# Option 1: Upgrade to latest version (recommended)# Modify LightRAG version in pyproject.toml# lightrag = "^1.4.9.4"# Rebuild image
docker compose down
docker compose up -d --build
# Option 2: Clean old data (temporary solution)
rm -rf ./rag_local_storage
docker compose restart
Q3: File upload returns 400 error?

Check:

  • File format supported (PDF, DOCX, PNG, JPG, etc.)
  • File size exceeds 100MB
  • File is empty
# View supported formats
curl http://localhost:8000/docs
Q3.5: Embedding dimension error?

If you encounter dimension-related errors, need to clean data and rebuild:

# Stop services
docker compose down
# Delete all volumes (clear database)
docker volume rm rag-api_dragonflydb_data rag-api_qdrant_data rag-api_memgraph_data
# Modify EMBEDDING_DIM in .env
EMBEDDING_DIM=1024 # or 4096, must match the model# Restart
docker compose up -d
Q4: Query is very slow (>30 seconds)?

Optimization suggestions:

  1. Use naive or hybrid mode instead of mix
  2. Increase MAX_ASYNC parameter (in .env)
  3. Reduce TOP_K and CHUNK_TOP_K
  4. Enable Reranker
# Modify .env
MAX_ASYNC=8
TOP_K=20
CHUNK_TOP_K=10
Q5: Out of memory (OOM)?

If using local MinerU:

# Switch to remote mode# Modify in .env
MINERU_MODE=remote
MINERU_API_TOKEN=your_token
# Or limit concurrency
DOCUMENT_PROCESSING_CONCURRENCY=1
Q6: Tasks lost after container restart?

Problem Symptoms:

  • Cannot query previous task status after container restart
  • Tasks disappear after tenant instance evicted by LRU

Solution: Enable Redis task storage

# Modify .env
TASK_STORE_STORAGE=redis
# Restart services
docker compose restart
# Verify
docker compose logs api | grep TaskStore
# Should see: ✅ TaskStore: Redis connection successful

Configuration Description:

  • memory mode: In-memory storage, data lost after restart (default, suitable for development)
  • redis mode: Persistent storage, supports container restart and instance rebuild (production recommended)

TTL Strategy (Redis mode auto-cleanup):

  • completed tasks: 24 hours
  • failed tasks: 24 hours
  • pending/processing tasks: 6 hours
Q7: VLM mode processing failed?

Check Items:

  1. vision_model_func not configured

    • Check logs:vision_model_func not found, fallback to off mode
    • Ensure LLM API is configured in .env
  2. Image file does not exist

    • Check logs:Image file not found: xxx
    • Possibly corrupted MinerU ZIP or extraction failed
  3. Timeout error

    • full mode may timeout on large files
    • Suggestion: Use selective mode first, or increase VLM_TIMEOUT
# Modify .env
VLM_TIMEOUT=300 # Increase to 5 minutes
RAG_VLM_MODE=selective # downgrade to selective

Debugging Tips:

# View detailed logs
docker compose logs -f | grep VLM
# Test single file
curl -X POST 'http://localhost:8000/insert?tenant_id=test&doc_id=test&vlm_mode=off' \
-F 'file=@test.pdf'

Performance Tuning Recommendations

ScenarioMAX_ASYNCTOP_KCHUNK_TOP_KMINERU_MODE
Fast response8105remote
Balanced mode82010remote
High accuracy46020remote
Resource limited42010remote

📖 Documentation


🤝 Contributing

We welcome all forms of contribution!

How to Contribute

  1. Fork the project
git clone https://github.com/BukeLy/rag-api.git
cd rag-api
  1. Create feature branch
git checkout -b feature/your-feature-name
  1. Development and Testing
# Install dependencies
uv sync
# Run tests
uv run pytest
# Code formatting
uv run black .
uv run isort .
  1. Submit code
git add .
git commit -m "feat: Add new feature"
git push origin feature/your-feature-name
  1. Create Pull Request

Create a PR on GitHub with detailed description of your changes.

Commit Conventions

Use semantic commit messages:

  • feat: New feature
  • fix: Bug fix
  • docs: Documentation update
  • style: Code formatting
  • refactor: Code refactoring
  • perf: Performance optimization
  • test: Testing
  • chore: Build/tools

See PR Workflow Documentation


📄 License

This project is licensed under the MIT License. See the LICENSE file for details.


🙏 Acknowledgments

This project is built on the following excellent open source projects:

  • LightRAG - Efficient knowledge graph RAG framework
  • RAG-Anything - Multimodal document parsing
  • MinerU - Powerful PDF parsing tool
  • Docling - Lightweight document parsing
  • FastAPI - Modern Python web framework

Special thanks to all contributors and users for their support! 🎉


📬 Contact Us


⭐ If this project helps you, please give it a Star!

Made with ❤️ by BukeLy

© 2025 RAG API. All rights reserved.

About

Multi-tenant RAG API powered by LightRAG/RAG-Anything. Auto-selects best parser (DeepSeek-OCR/MinerU/Docling) via complexity scoring

Topics

Resources

Stars

64 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

🚀 RAG API

Multi-tenant Multimodal Document Intelligent Retrieval System

Enterprise-grade RAG service built on RAG-Anything and LightRAG

CIPythonFastAPILightRAGDockerLicense

English | 简体中文

FeaturesQuick StartArchitectureAPI DocumentationDeployment


📖 Introduction

RAG API is an enterprise-grade Retrieval-Augmented Generation (RAG) service that combines the powerful document parsing capabilities of RAG-Anything with the efficient knowledge graph retrieval technology of LightRAG, providing intelligent Q&A capabilities for your documents.

🎯 Key Highlights

  • 🏢 Multi-tenant Isolation - Complete tenant data isolation for enterprise multi-tenant scenarios
  • 🎨 Multimodal Parsing - Support for PDF, Word, images and more, with full OCR, tables, and formulas coverage
  • High-performance Retrieval - Knowledge graph-based hybrid retrieval with 6-15 second query response
  • 🔄 Flexible Deployment - Support for production and development modes with one-click switching
  • 📦 Ready to Use - One-click Docker deployment, service starts in 3 minutes
  • 🎛️ Multiple Parsing Engines - DeepSeek-OCR (Remote API) + MinerU (Local/Remote API) + Docling (Fast)
  • 🎨 RAG-Anything VLM Enhancement - Three modes (off/selective/full) for deep chart understanding
  • 💾 Task Persistence - Redis storage support, tasks recoverable after container restart/instance rebuild

✨ Features

📄 Document Processing

  • Multiple Format Support

    • PDF, Word, Excel, PPT
    • PNG, JPG, WebP images
    • TXT, Markdown text
  • Intelligent Parsing

    • Plain text (.txt, .md) → Direct insertion (ultra-fast ~1s, skip parser)
    • OCR text recognition
    • Structured table extraction
    • Mathematical formula recognition
    • Layout analysis
  • RAG-Anything VLM Enhancement 🆕

    • off - Markdown only (fastest)
    • selective - Selective processing of important charts
    • full - Complete context enhancement processing
    • Smart filtering: with titles, large size, first page content
    • ⚠️Only supports remote MinerU mode, local mode uses RAG-Anything native methods
  • Batch Processing

    • Up to 100 files per batch
    • Async task queue
    • Real-time progress tracking

🔍 Intelligent Retrieval

  • Multi-mode Query

    • naive - Vector retrieval (fastest)
    • local - Local graph
    • global - Global graph
    • hybrid - Hybrid retrieval
    • mix - Full retrieval (most accurate)
  • Knowledge Graph

    • Automatic entity extraction
    • Relationship reasoning
    • Semantic understanding
    • Context enhancement
  • External Storage

    • DragonflyDB (KV storage + task storage)
    • Qdrant (vector storage)
    • Memgraph (graph database)
    • Task persistence (Redis mode)

🏗️ Architecture

System Architecture Diagram

graph TB
subgraph "Client Layer"
Client[Client Application]
WebUI[Web Interface]
end
subgraph "API Gateway Layer"
FastAPI[FastAPI Service]
Auth[Tenant Authentication]
end
subgraph "Business Logic Layer"
TenantMgr[Tenant Manager]
TaskQueue[Task Queue]
subgraph "Document Processing"
DeepSeekOCR[DeepSeek-OCR<br/>Fast OCR 80% cases]
MinerU[MinerU Parser<br/>Complex multimodal]
Docling[Docling Parser<br/>Fast lightweight]
FileRouter[Smart Router<br/>Complexity scoring]
end
subgraph "RAG Engine"
LightRAG[LightRAG Instance Pool<br/>LRU Cache 50]
KG[Knowledge Graph Engine]
Vector[Vector Retrieval Engine]
end
end
subgraph "Storage Layer"
DragonflyDB[(DragonflyDB<br/>KV Storage)]
Qdrant[(Qdrant<br/>Vector Database)]
Memgraph[(Memgraph<br/>Graph Database)]
Local[(Local Files<br/>Temp Storage)]
end
subgraph "External Services"
LLM[LLM<br/>Entity Extraction/Generation]
Embedding[Embedding<br/>Vectorization]
Rerank[Rerank<br/>Reranking]
end
Client --> FastAPI
WebUI --> FastAPI
FastAPI --> Auth
Auth --> TenantMgr
TenantMgr --> TaskQueue
TenantMgr --> LightRAG
TaskQueue --> FileRouter
FileRouter --> DeepSeekOCR
FileRouter --> MinerU
FileRouter --> Docling
DeepSeekOCR --> LightRAG
MinerU --> LightRAG
Docling --> LightRAG
LightRAG --> KG
LightRAG --> Vector
KG --> DragonflyDB
KG --> Memgraph
Vector --> Qdrant
LightRAG --> Local
LightRAG --> LLM
LightRAG --> Embedding
Vector --> Rerank
style FastAPI fill:#00C7B7
style LightRAG fill:#FF6B6B
style DeepSeekOCR fill:#5DADE2
style MinerU fill:#4ECDC4
style Docling fill:#95E1D3
style TenantMgr fill:#F38181
Loading

Multi-tenant Architecture

graph TB
subgraph "Tenant A"
A_Config[Tenant A Config<br/>Independent API Key]
A_Instance[LightRAG Instance A<br/>Dedicated LLM/Embedding]
A_Data[(Tenant A Data<br/>Fully Isolated)]
A_Config --> A_Instance
A_Instance --> A_Data
end
subgraph "Tenant B"
B_Config[Tenant B Config<br/>Independent API Key]
B_Instance[LightRAG Instance B<br/>Dedicated LLM/Embedding]
B_Data[(Tenant B Data<br/>Fully Isolated)]
B_Config --> B_Instance
B_Instance --> B_Data
end
subgraph "Tenant C"
C_Config[Using Global Config]
C_Instance[LightRAG Instance C<br/>Shared LLM/Embedding]
C_Data[(Tenant C Data<br/>Fully Isolated)]
C_Config --> C_Instance
C_Instance --> C_Data
end
Pool[Instance Pool Manager<br/>LRU Cache + Config Isolation]
Global[Global Config<br/>Default API Key]
Pool --> A_Instance
Pool --> B_Instance
Pool --> C_Instance
C_Config -.fallback.-> Global
style Pool fill:#F38181
style Global fill:#95E1D3
style A_Config fill:#FFD93D
style B_Config fill:#FFD93D
style C_Config fill:#E8E8E8
Loading

Core Technology Stack

🔧 Frameworks & Runtime

  • FastAPI 0.115+
  • Python 3.11+
  • Uvicorn
  • Docker & Docker Compose

🧠 AI & RAG

  • LightRAG 1.4.9.4
  • RAG-Anything
  • MinerU (PDF-Extract-Kit)
  • Docling

💾 Storage & Database

  • DragonflyDB(Redis compatible)
  • Qdrant(Vector Database)
  • Memgraph(Graph Database)
  • Local filesystem

🚀 Quick Start

Option 1: One-click Deployment (Recommended)

Suitable for production and testing environments:

# 1. Clone the project
git clone https://github.com/BukeLy/rag-api.git
cd rag-api
# 2. Configure environment variables
cp env.example .env
nano .env # Fill in your API keys# 3. Run deployment script
chmod +x deploy.sh
./deploy.sh
# Select deployment mode:# 1) Production Mode - Standard container deployment# 2) Development Mode - Code hot-reload# 4. Verify service
curl http://localhost:8000/

Access Swagger Documentation:http://localhost:8000/docs

Option 2: Docker Compose

Production Mode

# Configure environment variables
cp env.example .env
nano .env
# Start services
docker compose -f docker-compose.yml up -d
# View logs
docker compose -f docker-compose.yml logs -f

Development Mode (Code Hot-reload)

# Start development environment
docker compose -f docker-compose.dev.yml up -d
# Or use quick script
./scripts/dev.sh
# Code changes will auto-reload without restart

Option 3: Local Development

# Install uv (Python package manager)
curl -LsSf https://astral.sh/uv/install.sh | sh
# Install dependencies
uv sync
# Configure environment variables
cp env.example .env
nano .env
# Start services
uv run uvicorn main:app --host 0.0.0.0 --port 8000 --reload

Environment Variable Configuration

Minimum configuration (required):

# LLM Configuration (Function-oriented naming)
LLM_API_KEY=your_llm_api_key
LLM_BASE_URL=https://ark.cn-beijing.volces.com/api/v3
LLM_MODEL=ep-xxx-xxx
# LLM_REQUESTS_PER_MINUTE=800 # Rate limit (optional)# LLM_TOKENS_PER_MINUTE=40000 # Rate limit (optional)# LLM_MAX_ASYNC=8 # [Optional, expert mode] Manual concurrency control# # Auto-calculated when unset: min(RPM, TPM/3500) = 11# Embedding Configuration (Function-oriented naming)
EMBEDDING_API_KEY=your_embedding_api_key
EMBEDDING_BASE_URL=https://api.siliconflow.cn/v1
EMBEDDING_MODEL=Qwen/Qwen3-Embedding-0.6B
EMBEDDING_DIM=1024
# EMBEDDING_MAX_ASYNC=32 # [Optional, expert mode] Auto-calculated when unset: 800# MinerU Mode (Remote recommended)
MINERU_MODE=remote
MINERU_API_TOKEN=your_token
MINERU_HTTP_TIMEOUT=60 # MinerU download timeout (seconds, default 60)
FILE_SERVICE_BASE_URL=http://your-ip:8000
# VLM Chart Enhancement Configuration 🆕# ⚠️ Note: Only effective in MINERU_MODE=remote
RAG_VLM_MODE=off # off / selective / full
RAG_IMPORTANCE_THRESHOLD=0.5 # Importance threshold (selective mode)
RAG_CONTEXT_WINDOW=2 # Context window (full mode)
RAG_CONTEXT_MODE=page # page / chunk
RAG_MAX_CONTEXT_TOKENS=3000 # Max context tokens# Task Storage Configuration 🆕
TASK_STORE_STORAGE=redis # memory / redis (production recommends redis)# Document Insert Verification Configuration 🆕
DOC_INSERT_VERIFICATION_TIMEOUT=300 # Verification timeout (seconds, default 5 minutes)
DOC_INSERT_VERIFICATION_POLL_INTERVAL=0.5 # Poll interval (seconds, default 500ms)# Model Call Timeout Configuration 🆕
MODEL_CALL_TIMEOUT=90 # Model call max timeout (seconds, default 90)

⚡ Auto Concurrency Calculation:

  • LLM: When LLM_MAX_ASYNC is unset, auto-calculated as min(RPM, TPM/3500) ≈ 11
  • Embedding: When EMBEDDING_MAX_ASYNC is unset, auto-calculated as min(RPM, TPM/500) ≈ 800
  • Rerank: When RERANK_MAX_ASYNC is unset, auto-calculated as min(RPM, TPM/500) ≈ 800

✅ Recommended: Don't set *_MAX_ASYNC, let the system auto-calculate to completely avoid 429 errors

See env.example for complete configuration.


📚 API Documentation

Core Endpoints

1️⃣ Upload Document

# Single file upload (default mode)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc1" \
-F "file=@document.pdf" \
-F "parser=auto"# VLM chart enhancement mode 🆕# off: Markdown only (fastest, default)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc2&vlm_mode=off" \
-F "file=@document.pdf"# selective: Selective processing of important charts (balance performance and quality)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc3&vlm_mode=selective" \
-F "file=@document.pdf"# full: Complete RAG-Anything processing (highest quality, context enhancement enabled)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc4&vlm_mode=full" \
-F "file=@document.pdf"# Response
{
"task_id": "task-xxx-xxx",
"doc_id": "doc1",
"filename": "document.pdf",
"vlm_mode": "off",
"status": "pending"
}

2️⃣ Batch Upload

curl -X POST "http://localhost:8000/batch?tenant_id=your_tenant" \
-F "files=@doc1.pdf" \
-F "files=@doc2.docx" \
-F "files=@image.png"# Response
{
"batch_id": "batch-xxx-xxx",
"total_files": 3,
"accepted_files": 3,
"tasks": [...]
}

3️⃣ Intelligent Query (Query API v2.0)

New Advanced Features:

  • Conversation History: Support for multi-turn conversation context
  • Custom Prompts: Customize response style
  • Response Format Control: paragraph/list/json
  • Keyword Precision Retrieval: hl_keywords/ll_keywords
  • Streaming Output: Real-time generation viewing
# Basic query
curl -X POST "http://localhost:8000/query?tenant_id=your_tenant" \
-H "Content-Type: application/json" \
-d '{ "query": "What are the core viewpoints in the document?", "mode": "hybrid" }'# Advanced query (multi-turn dialogue + custom prompt)
curl -X POST "http://localhost:8000/query?tenant_id=your_tenant" \
-H "Content-Type: application/json" \
-d '{ "query": "Can you elaborate on the second point?", "mode": "hybrid", "conversation_history": [ {"role": "user", "content": "What are the key points?"}, {"role": "assistant", "content": "There are mainly three points..."} ], "user_prompt": "Please answer in professional academic language", "response_type": "list" }'# Streaming query (SSE)
curl -N -X POST "http://localhost:8000/query/stream?tenant_id=your_tenant" \
-H "Content-Type: application/json" \
-d '{ "query": "What are the core viewpoints in the document?", "mode": "hybrid" }'# Response (real-time streaming output)
data: {"chunk": "Based on", "done": false}
data: {"chunk": "document content", "done": false}
data: {"done": true}

4️⃣ Task Status Query

curl "http://localhost:8000/task/task-xxx-xxx?tenant_id=your_tenant"# Response
{
"task_id": "task-xxx-xxx",
"status": "completed",
"progress": 100,
"result": {...}
}

5️⃣ Tenant Management

# Get tenant statistics
curl "http://localhost:8000/tenants/stats?tenant_id=your_tenant"# Clear tenant cache
curl -X DELETE "http://localhost:8000/tenants/cache?tenant_id=your_tenant"# View instance pool status (admin)
curl "http://localhost:8000/tenants/pool/stats"

VLM Mode Comparison 🆕

ModeSpeedQualityResource UsageUse Case
off⚡⚡⚡⚡⚡⭐⭐⭐Very LowPlain text documents, fast batch processing
selective⚡⚡⚡⚡⭐⭐⭐⭐LowDocuments with key charts (recommended)
full⚡⚡⭐⭐⭐⭐⭐HighChart-intensive research reports, papers

Processing Time Estimate (20-page PDF example):

  • off: ~10 seconds(Markdown only)
  • selective: ~30 seconds(5-10 important charts)
  • full: ~120 seconds(complete context processing)

Query Mode Comparison

ModeSpeedAccuracyUse Case
naive⚡⚡⚡⚡⚡⭐⭐⭐Simple Q&A, fast retrieval
local⚡⚡⚡⚡⭐⭐⭐⭐Local entity relationship queries
global⚡⚡⚡⭐⭐⭐⭐Global knowledge graph reasoning
hybrid⚡⚡⚡⭐⭐⭐⭐⭐Hybrid retrieval (recommended)
mix⚡⚡⭐⭐⭐⭐⭐Complex questions, deep analysis

Query API v2.0 Advanced Parameters

ParameterTypeDescriptionExample
conversation_historyList[Dict]Multi-turn conversation context[{"role": "user", "content": "..."}]
user_promptstrCustom prompt"Please answer in professional academic language"
response_typestrResponse format"paragraph", "list", "json"
hl_keywordsList[str]High priority keywords["artificial intelligence", "machine learning"]
ll_keywordsList[str]Low priority keywords["application", "case study"]
only_need_contextboolReturn context only (debug)true
max_entity_tokensintEntity token limit6000

Complete API documentation:http://localhost:8000/docs


🎯 Usage Examples

Python SDK

importrequests# ConfigurationBASE_URL="http://localhost:8000"TENANT_ID="your_tenant"# Upload documentwithopen("document.pdf", "rb") asf:
response=requests.post(
f"{BASE_URL}/insert",
params={"tenant_id": TENANT_ID, "doc_id": "doc1"},
files={"file": f}
)
task_id=response.json()["task_id"]
print(f"Task ID: {task_id}")
# Queryresponse=requests.post(
f"{BASE_URL}/query",
params={"tenant_id": TENANT_ID},
json={
"query": "What is the main content of the document?",
"mode": "hybrid",
"top_k": 10
}
)
result=response.json()
print(f"Answer: {result['answer']}")

Complete cURL Example

# 1. Upload PDF document
TASK_ID=$(curl -X POST "http://localhost:8000/insert?tenant_id=demo&doc_id=report" \ -F "file=@report.pdf"| jq -r '.task_id')echo"Task ID: $TASK_ID"# 2. Wait for processing completionwhiletrue;do
STATUS=$(curl -s "http://localhost:8000/task/$TASK_ID?tenant_id=demo"| jq -r '.status')echo"Status: $STATUS"if [ "$STATUS"="completed" ] || [ "$STATUS"="failed" ];thenbreakfi
sleep 2
done# 3. Query document content
curl -X POST "http://localhost:8000/query?tenant_id=demo" \
-H "Content-Type: application/json" \
-d '{ "query": "What are the main conclusions of this report?", "mode": "hybrid" }'| jq '.answer'

🛠️ Deployment

System Requirements

Minimum Configuration:

  • CPU: 2 cores
  • RAM: 4GB
  • Disk: 40GB SSD
  • OS: Ubuntu 20.04+ / Debian 11+ / CentOS 8+

Recommended Configuration (Production):

  • CPU: 4 cores
  • RAM: 8GB
  • Disk: 100GB SSD
  • OS: Ubuntu 22.04 LTS

Server Deployment

Quick Deployment on Aliyun/Tencent Cloud

# SSH login to server
ssh root@your-server-ip
# Clone project
git clone https://github.com/BukeLy/rag-api.git
cd rag-api
# Run one-click deployment script
chmod +x deploy.sh
./deploy.sh
# The script will automatically:# 1. Install Docker and Docker Compose# 2. Configure environment variables# 3. Optimize system parameters# 4. Start services# 5. Verify health status

External Storage Configuration

Supports DragonflyDB + Qdrant + Memgraph external storage (enabled by default):

# Configure in .env
USE_EXTERNAL_STORAGE=true
# DragonflyDB configuration (KV Storage)
KV_STORAGE=RedisKVStorage
REDIS_URI=redis://dragonflydb:6379/0
# Qdrant configuration (vector storage)
VECTOR_STORAGE=QdrantVectorDBStorage
QDRANT_URL=http://qdrant:6333
# Memgraph configuration (graph storage)
GRAPH_STORAGE=MemgraphStorage
MEMGRAPH_URI=bolt://memgraph:7687
MEMGRAPH_USERNAME=
MEMGRAPH_PASSWORD=

See External Storage Deployment Documentation

Docker Compose Configuration

The project provides two configuration files:

FilePurposeFeatures
docker-compose.ymlProduction modeCode packaged in image, optimal performance
docker-compose.dev.ymlDevelopment modeCode mounted externally, supports hot-reload

Select configuration file:

# Production mode
docker compose -f docker-compose.yml up -d
# Development mode
docker compose -f docker-compose.dev.yml up -d

Performance Optimization

Tuning Parameters

Configure in .env:

# ⚡ Concurrency Control (Recommended: use auto-calculation)# LLM_MAX_ASYNC=8 # [Expert mode] Manually specify LLM concurrency# # Auto-calculated when unset: min(RPM, TPM/3500) ≈ 11# EMBEDDING_MAX_ASYNC=32 # [Expert mode] Manually specify Embedding concurrency# # Auto-calculated when unset: min(RPM, TPM/500) ≈ 800# RERANK_MAX_ASYNC=16 # [Expert mode] Manually specify Rerank concurrency# # Auto-calculated when unset: min(RPM, TPM/500) ≈ 800# Retrieval count (affects query quality and speed)
TOP_K=20 # Entity/relationship retrieval count
CHUNK_TOP_K=10 # Text chunk retrieval count# Document processing concurrency
DOCUMENT_PROCESSING_CONCURRENCY=10 # Remote mode can be set high, local mode set to 1

🎯 Concurrency Configuration Recommendations:

  • Recommended: Don't set *_MAX_ASYNC, let the system auto-calculate based on TPM/RPM
  • Expert mode: If manual control needed, can set LLM_MAX_ASYNC and other parameters
  • Advantage: Auto-calculation completely avoids 429 errors (TPM limit reached)

Mode Selection

  • MinerU Remote Mode (Recommended): High concurrency, resource-efficient
  • MinerU Local Mode: Requires GPU, high memory usage
  • Docling Mode: Fast and lightweight, suitable for simple documents

🏢 Multi-tenant Usage

Tenant Isolation

Each tenant has:

  • ✅ Independent LightRAG instance
  • ✅ Isolated data storage space
  • ✅ Independent vector index
  • ✅ Dedicated knowledge graph
  • Independent service configuration (LLM, Embedding, Rerank, DeepSeek-OCR, MinerU)🆕

Tenant Configuration Management 🆕

Each tenant can independently configure 5 services with hot-reload support:

# 1️⃣ Configure independent DeepSeek-OCR API key for Tenant A
curl -X PUT "http://localhost:8000/tenants/tenant_a/config" \
-H "Content-Type: application/json" \
-d '{ "ds_ocr_config": { "api_key": "sk-tenant-a-ds-ocr-key", "base_url": "https://api.siliconflow.cn/v1", "model": "deepseek-ai/DeepSeek-OCR", "timeout": 90 } }'# 2️⃣ Configure independent MinerU API token for Tenant B
curl -X PUT "http://localhost:8000/tenants/tenant_b/config" \
-H "Content-Type: application/json" \
-d '{ "mineru_config": { "api_token": "tenant-b-mineru-token", "base_url": "https://mineru.net", "model_version": "vlm" } }'# 3️⃣ Configure multiple services simultaneously (LLM + Embedding + DeepSeek-OCR)
curl -X PUT "http://localhost:8000/tenants/tenant_c/config" \
-H "Content-Type: application/json" \
-d '{ "llm_config": { "api_key": "sk-tenant-c-llm-key", "model": "gpt-4" }, "embedding_config": { "api_key": "sk-tenant-c-embedding-key", "model": "Qwen/Qwen3-Embedding-0.6B", "dim": 1024 }, "ds_ocr_config": { "api_key": "sk-tenant-c-ds-ocr-key" } }'# 4️⃣ Query tenant configuration (API key auto-masked)
curl "http://localhost:8000/tenants/tenant_a/config"# Response example
{
"tenant_id": "tenant_a",
"ds_ocr_config": {
"api_key": "sk-***-key", // Auto-masked
"timeout": 90
},
"merged_config": {
"llm": {...}, // Using Global Config
"embedding": {...}, // Using Global Config
"rerank": {...}, // Using Global Config
"ds_ocr": {...}, // Using tenant config
"mineru": {...} // Using Global Config
}
}
# 5️⃣ Refresh config cache (config hot-reload)
curl -X POST "http://localhost:8000/tenants/tenant_a/config/refresh"# 6️⃣ Delete tenant config (restore to global config)
curl -X DELETE "http://localhost:8000/tenants/tenant_a/config"

Supported Configuration Items:

ServiceConfig FieldDescription
LLMllm_configModel, API key, base_url, etc.
Embeddingembedding_configModel, API key, dimension, etc.
Rerankrerank_configModel, API key, etc.
DeepSeek-OCRds_ocr_configAPI key, timeout, mode, etc.
MinerUmineru_configAPI token, version, timeout, etc.

Configuration Priority: Tenant config > Global config

Use Cases:

  • 🔐 Multi-tenant SaaS: Each tenant uses their own API key
  • 💰 Pay-per-use: Track tenant usage through independent API keys
  • 🎯 Differentiated Services: Different tenants use different models (GPT-4 vs GPT-3.5)
  • 🧪 A/B Testing: Compare different models/parameters

Usage

All APIs require tenant_id parameter:

# Tenant A upload document
curl -X POST "http://localhost:8000/insert?tenant_id=tenant_a&doc_id=doc1" \
-F "file=@doc.pdf"# Tenant B upload document (fully isolated)
curl -X POST "http://localhost:8000/insert?tenant_id=tenant_b&doc_id=doc1" \
-F "file=@doc.pdf"# Tenant A query (can only query own documents)
curl -X POST "http://localhost:8000/query?tenant_id=tenant_a" \
-H "Content-Type: application/json" \
-d '{"query": "document content", "mode": "hybrid"}'

Instance Pool Management

  • Capacity: Cache up to 50 tenant instances
  • Strategy: LRU (Least Recently Used) automatic cleanup
  • Config Isolation: Each tenant can use independent LLM, Embedding, parser configuration

📊 Monitoring & Maintenance

Common Commands

# View service status
docker compose ps
# View real-time logs
docker compose logs -f
# Restart services
docker compose restart
# Stop services
docker compose down
# View resource usage
docker stats
# Clean Docker resources
docker system prune -f

Maintenance Scripts

# Monitor service health
./scripts/monitor.sh
# Backup data
./scripts/backup.sh
# Update services
./scripts/update.sh
# Performance testing
./scripts/test_concurrent_perf.sh
# Performance monitoring
./scripts/monitor_performance.sh

Health Checks

# Complete health check (recommended)
./scripts/health_check.sh
./scripts/health_check.sh --verbose # verbose output# API health check
curl http://localhost:8000/
# Tenant statistics
curl "http://localhost:8000/tenants/stats?tenant_id=your_tenant"# Instance pool status
curl "http://localhost:8000/tenants/pool/stats"

🗂️ Project Structure

rag-api/
├── main.py # FastAPI application entry
├── api/ # API route modules
│ ├── __init__.py # Route aggregation
│ ├── insert.py # Document upload (single/batch)
│ ├── query.py # Intelligent query
│ ├── task.py # Task status query
│ ├── tenant.py # Tenant management
│ ├── files.py # File service
│ ├── models.py # Pydantic models
│ └── task_store.py # Task storage
├── src/ # Core business logic
│ ├── rag.py # LightRAG lifecycle management
│ ├── multi_tenant.py # Multi-tenant instance manager
│ ├── tenant_deps.py # Tenant dependency injection
│ ├── logger.py # Unified logging
│ ├── metrics.py # Performance metrics
│ ├── file_url_service.py # Temporary file service
│ ├── mineru_client.py # MinerU client
│ └── mineru_result_processor.py # Result processing
├── docs/ # Documentation
│ ├── ARCHITECTURE.md # Architecture design documentation
│ ├── USAGE.md # Detailed usage guide
│ ├── DEPLOY_MODES.md # Deployment mode description
│ ├── PR_WORKFLOW.md # PR workflow
│ └── ...
├── scripts/ # Maintenance scripts
│ ├── dev.sh # Development mode quick start
│ ├── monitor.sh # Service monitoring
│ ├── backup.sh # Data backup
│ ├── update.sh # Service update
│ └── ...
├── deploy.sh # One-click deployment script
├── docker-compose.yml # Production mode configuration
├── docker-compose.dev.yml # Development mode configuration
├── Dockerfile # Production image
├── Dockerfile.dev # Development image
├── pyproject.toml # Project dependencies
├── uv.lock # Dependency lock
├── env.example # Environment variable template
├── CLAUDE.md # Claude AI guide
└── README.md # This documentation

🐛 Troubleshooting

Common Issues

Q1: What to do if service fails to start?
# View detailed logs
docker compose logs
# Check port usage
netstat -tulpn | grep 8000
# Check Docker status
docker ps -a
Q2: multimodal_processed error?

Note: This issue has been fixed in LightRAG 1.4.9.4+. If you encounter this error, your version is outdated.

Solution:

# Option 1: Upgrade to latest version (recommended)# Modify LightRAG version in pyproject.toml# lightrag = "^1.4.9.4"# Rebuild image
docker compose down
docker compose up -d --build
# Option 2: Clean old data (temporary solution)
rm -rf ./rag_local_storage
docker compose restart
Q3: File upload returns 400 error?

Check:

  • File format supported (PDF, DOCX, PNG, JPG, etc.)
  • File size exceeds 100MB
  • File is empty
# View supported formats
curl http://localhost:8000/docs
Q3.5: Embedding dimension error?

If you encounter dimension-related errors, need to clean data and rebuild:

# Stop services
docker compose down
# Delete all volumes (clear database)
docker volume rm rag-api_dragonflydb_data rag-api_qdrant_data rag-api_memgraph_data
# Modify EMBEDDING_DIM in .env
EMBEDDING_DIM=1024 # or 4096, must match the model# Restart
docker compose up -d
Q4: Query is very slow (>30 seconds)?

Optimization suggestions:

  1. Use naive or hybrid mode instead of mix
  2. Increase MAX_ASYNC parameter (in .env)
  3. Reduce TOP_K and CHUNK_TOP_K
  4. Enable Reranker
# Modify .env
MAX_ASYNC=8
TOP_K=20
CHUNK_TOP_K=10
Q5: Out of memory (OOM)?

If using local MinerU:

# Switch to remote mode# Modify in .env
MINERU_MODE=remote
MINERU_API_TOKEN=your_token
# Or limit concurrency
DOCUMENT_PROCESSING_CONCURRENCY=1
Q6: Tasks lost after container restart?

Problem Symptoms:

  • Cannot query previous task status after container restart
  • Tasks disappear after tenant instance evicted by LRU

Solution: Enable Redis task storage

# Modify .env
TASK_STORE_STORAGE=redis
# Restart services
docker compose restart
# Verify
docker compose logs api | grep TaskStore
# Should see: ✅ TaskStore: Redis connection successful

Configuration Description:

  • memory mode: In-memory storage, data lost after restart (default, suitable for development)
  • redis mode: Persistent storage, supports container restart and instance rebuild (production recommended)

TTL Strategy (Redis mode auto-cleanup):

  • completed tasks: 24 hours
  • failed tasks: 24 hours
  • pending/processing tasks: 6 hours
Q7: VLM mode processing failed?

Check Items:

  1. vision_model_func not configured

    • Check logs:vision_model_func not found, fallback to off mode
    • Ensure LLM API is configured in .env
  2. Image file does not exist

    • Check logs:Image file not found: xxx
    • Possibly corrupted MinerU ZIP or extraction failed
  3. Timeout error

    • full mode may timeout on large files
    • Suggestion: Use selective mode first, or increase VLM_TIMEOUT
# Modify .env
VLM_TIMEOUT=300 # Increase to 5 minutes
RAG_VLM_MODE=selective # downgrade to selective

Debugging Tips:

# View detailed logs
docker compose logs -f | grep VLM
# Test single file
curl -X POST 'http://localhost:8000/insert?tenant_id=test&doc_id=test&vlm_mode=off' \
-F 'file=@test.pdf'

Performance Tuning Recommendations

ScenarioMAX_ASYNCTOP_KCHUNK_TOP_KMINERU_MODE
Fast response8105remote
Balanced mode82010remote
High accuracy46020remote
Resource limited42010remote

📖 Documentation


🤝 Contributing

We welcome all forms of contribution!

How to Contribute

  1. Fork the project
git clone https://github.com/BukeLy/rag-api.git
cd rag-api
  1. Create feature branch
git checkout -b feature/your-feature-name
  1. Development and Testing
# Install dependencies
uv sync
# Run tests
uv run pytest
# Code formatting
uv run black .
uv run isort .
  1. Submit code
git add .
git commit -m "feat: Add new feature"
git push origin feature/your-feature-name
  1. Create Pull Request

Create a PR on GitHub with detailed description of your changes.

Commit Conventions

Use semantic commit messages:

  • feat: New feature
  • fix: Bug fix
  • docs: Documentation update
  • style: Code formatting
  • refactor: Code refactoring
  • perf: Performance optimization
  • test: Testing
  • chore: Build/tools

See PR Workflow Documentation


📄 License

This project is licensed under the MIT License. See the LICENSE file for details.


🙏 Acknowledgments

This project is built on the following excellent open source projects:

  • LightRAG - Efficient knowledge graph RAG framework
  • RAG-Anything - Multimodal document parsing
  • MinerU - Powerful PDF parsing tool
  • Docling - Lightweight document parsing
  • FastAPI - Modern Python web framework

Special thanks to all contributors and users for their support! 🎉


📬 Contact Us


⭐ If this project helps you, please give it a Star!

Made with ❤️ by BukeLy

© 2025 RAG API. All rights reserved.

About

Multi-tenant RAG API powered by LightRAG/RAG-Anything. Auto-selects best parser (DeepSeek-OCR/MinerU/Docling) via complexity scoring

Topics

Resources

Stars

64 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

🚀 RAG API

Multi-tenant Multimodal Document Intelligent Retrieval System

Enterprise-grade RAG service built on RAG-Anything and LightRAG

CIPythonFastAPILightRAGDockerLicense

English | 简体中文

FeaturesQuick StartArchitectureAPI DocumentationDeployment


📖 Introduction

RAG API is an enterprise-grade Retrieval-Augmented Generation (RAG) service that combines the powerful document parsing capabilities of RAG-Anything with the efficient knowledge graph retrieval technology of LightRAG, providing intelligent Q&A capabilities for your documents.

🎯 Key Highlights

  • 🏢 Multi-tenant Isolation - Complete tenant data isolation for enterprise multi-tenant scenarios
  • 🎨 Multimodal Parsing - Support for PDF, Word, images and more, with full OCR, tables, and formulas coverage
  • High-performance Retrieval - Knowledge graph-based hybrid retrieval with 6-15 second query response
  • 🔄 Flexible Deployment - Support for production and development modes with one-click switching
  • 📦 Ready to Use - One-click Docker deployment, service starts in 3 minutes
  • 🎛️ Multiple Parsing Engines - DeepSeek-OCR (Remote API) + MinerU (Local/Remote API) + Docling (Fast)
  • 🎨 RAG-Anything VLM Enhancement - Three modes (off/selective/full) for deep chart understanding
  • 💾 Task Persistence - Redis storage support, tasks recoverable after container restart/instance rebuild

✨ Features

📄 Document Processing

  • Multiple Format Support

    • PDF, Word, Excel, PPT
    • PNG, JPG, WebP images
    • TXT, Markdown text
  • Intelligent Parsing

    • Plain text (.txt, .md) → Direct insertion (ultra-fast ~1s, skip parser)
    • OCR text recognition
    • Structured table extraction
    • Mathematical formula recognition
    • Layout analysis
  • RAG-Anything VLM Enhancement 🆕

    • off - Markdown only (fastest)
    • selective - Selective processing of important charts
    • full - Complete context enhancement processing
    • Smart filtering: with titles, large size, first page content
    • ⚠️Only supports remote MinerU mode, local mode uses RAG-Anything native methods
  • Batch Processing

    • Up to 100 files per batch
    • Async task queue
    • Real-time progress tracking

🔍 Intelligent Retrieval

  • Multi-mode Query

    • naive - Vector retrieval (fastest)
    • local - Local graph
    • global - Global graph
    • hybrid - Hybrid retrieval
    • mix - Full retrieval (most accurate)
  • Knowledge Graph

    • Automatic entity extraction
    • Relationship reasoning
    • Semantic understanding
    • Context enhancement
  • External Storage

    • DragonflyDB (KV storage + task storage)
    • Qdrant (vector storage)
    • Memgraph (graph database)
    • Task persistence (Redis mode)

🏗️ Architecture

System Architecture Diagram

graph TB
subgraph "Client Layer"
Client[Client Application]
WebUI[Web Interface]
end
subgraph "API Gateway Layer"
FastAPI[FastAPI Service]
Auth[Tenant Authentication]
end
subgraph "Business Logic Layer"
TenantMgr[Tenant Manager]
TaskQueue[Task Queue]
subgraph "Document Processing"
DeepSeekOCR[DeepSeek-OCR<br/>Fast OCR 80% cases]
MinerU[MinerU Parser<br/>Complex multimodal]
Docling[Docling Parser<br/>Fast lightweight]
FileRouter[Smart Router<br/>Complexity scoring]
end
subgraph "RAG Engine"
LightRAG[LightRAG Instance Pool<br/>LRU Cache 50]
KG[Knowledge Graph Engine]
Vector[Vector Retrieval Engine]
end
end
subgraph "Storage Layer"
DragonflyDB[(DragonflyDB<br/>KV Storage)]
Qdrant[(Qdrant<br/>Vector Database)]
Memgraph[(Memgraph<br/>Graph Database)]
Local[(Local Files<br/>Temp Storage)]
end
subgraph "External Services"
LLM[LLM<br/>Entity Extraction/Generation]
Embedding[Embedding<br/>Vectorization]
Rerank[Rerank<br/>Reranking]
end
Client --> FastAPI
WebUI --> FastAPI
FastAPI --> Auth
Auth --> TenantMgr
TenantMgr --> TaskQueue
TenantMgr --> LightRAG
TaskQueue --> FileRouter
FileRouter --> DeepSeekOCR
FileRouter --> MinerU
FileRouter --> Docling
DeepSeekOCR --> LightRAG
MinerU --> LightRAG
Docling --> LightRAG
LightRAG --> KG
LightRAG --> Vector
KG --> DragonflyDB
KG --> Memgraph
Vector --> Qdrant
LightRAG --> Local
LightRAG --> LLM
LightRAG --> Embedding
Vector --> Rerank
style FastAPI fill:#00C7B7
style LightRAG fill:#FF6B6B
style DeepSeekOCR fill:#5DADE2
style MinerU fill:#4ECDC4
style Docling fill:#95E1D3
style TenantMgr fill:#F38181
Loading

Multi-tenant Architecture

graph TB
subgraph "Tenant A"
A_Config[Tenant A Config<br/>Independent API Key]
A_Instance[LightRAG Instance A<br/>Dedicated LLM/Embedding]
A_Data[(Tenant A Data<br/>Fully Isolated)]
A_Config --> A_Instance
A_Instance --> A_Data
end
subgraph "Tenant B"
B_Config[Tenant B Config<br/>Independent API Key]
B_Instance[LightRAG Instance B<br/>Dedicated LLM/Embedding]
B_Data[(Tenant B Data<br/>Fully Isolated)]
B_Config --> B_Instance
B_Instance --> B_Data
end
subgraph "Tenant C"
C_Config[Using Global Config]
C_Instance[LightRAG Instance C<br/>Shared LLM/Embedding]
C_Data[(Tenant C Data<br/>Fully Isolated)]
C_Config --> C_Instance
C_Instance --> C_Data
end
Pool[Instance Pool Manager<br/>LRU Cache + Config Isolation]
Global[Global Config<br/>Default API Key]
Pool --> A_Instance
Pool --> B_Instance
Pool --> C_Instance
C_Config -.fallback.-> Global
style Pool fill:#F38181
style Global fill:#95E1D3
style A_Config fill:#FFD93D
style B_Config fill:#FFD93D
style C_Config fill:#E8E8E8
Loading

Core Technology Stack

🔧 Frameworks & Runtime

  • FastAPI 0.115+
  • Python 3.11+
  • Uvicorn
  • Docker & Docker Compose

🧠 AI & RAG

  • LightRAG 1.4.9.4
  • RAG-Anything
  • MinerU (PDF-Extract-Kit)
  • Docling

💾 Storage & Database

  • DragonflyDB(Redis compatible)
  • Qdrant(Vector Database)
  • Memgraph(Graph Database)
  • Local filesystem

🚀 Quick Start

Option 1: One-click Deployment (Recommended)

Suitable for production and testing environments:

# 1. Clone the project
git clone https://github.com/BukeLy/rag-api.git
cd rag-api
# 2. Configure environment variables
cp env.example .env
nano .env # Fill in your API keys# 3. Run deployment script
chmod +x deploy.sh
./deploy.sh
# Select deployment mode:# 1) Production Mode - Standard container deployment# 2) Development Mode - Code hot-reload# 4. Verify service
curl http://localhost:8000/

Access Swagger Documentation:http://localhost:8000/docs

Option 2: Docker Compose

Production Mode

# Configure environment variables
cp env.example .env
nano .env
# Start services
docker compose -f docker-compose.yml up -d
# View logs
docker compose -f docker-compose.yml logs -f

Development Mode (Code Hot-reload)

# Start development environment
docker compose -f docker-compose.dev.yml up -d
# Or use quick script
./scripts/dev.sh
# Code changes will auto-reload without restart

Option 3: Local Development

# Install uv (Python package manager)
curl -LsSf https://astral.sh/uv/install.sh | sh
# Install dependencies
uv sync
# Configure environment variables
cp env.example .env
nano .env
# Start services
uv run uvicorn main:app --host 0.0.0.0 --port 8000 --reload

Environment Variable Configuration

Minimum configuration (required):

# LLM Configuration (Function-oriented naming)
LLM_API_KEY=your_llm_api_key
LLM_BASE_URL=https://ark.cn-beijing.volces.com/api/v3
LLM_MODEL=ep-xxx-xxx
# LLM_REQUESTS_PER_MINUTE=800 # Rate limit (optional)# LLM_TOKENS_PER_MINUTE=40000 # Rate limit (optional)# LLM_MAX_ASYNC=8 # [Optional, expert mode] Manual concurrency control# # Auto-calculated when unset: min(RPM, TPM/3500) = 11# Embedding Configuration (Function-oriented naming)
EMBEDDING_API_KEY=your_embedding_api_key
EMBEDDING_BASE_URL=https://api.siliconflow.cn/v1
EMBEDDING_MODEL=Qwen/Qwen3-Embedding-0.6B
EMBEDDING_DIM=1024
# EMBEDDING_MAX_ASYNC=32 # [Optional, expert mode] Auto-calculated when unset: 800# MinerU Mode (Remote recommended)
MINERU_MODE=remote
MINERU_API_TOKEN=your_token
MINERU_HTTP_TIMEOUT=60 # MinerU download timeout (seconds, default 60)
FILE_SERVICE_BASE_URL=http://your-ip:8000
# VLM Chart Enhancement Configuration 🆕# ⚠️ Note: Only effective in MINERU_MODE=remote
RAG_VLM_MODE=off # off / selective / full
RAG_IMPORTANCE_THRESHOLD=0.5 # Importance threshold (selective mode)
RAG_CONTEXT_WINDOW=2 # Context window (full mode)
RAG_CONTEXT_MODE=page # page / chunk
RAG_MAX_CONTEXT_TOKENS=3000 # Max context tokens# Task Storage Configuration 🆕
TASK_STORE_STORAGE=redis # memory / redis (production recommends redis)# Document Insert Verification Configuration 🆕
DOC_INSERT_VERIFICATION_TIMEOUT=300 # Verification timeout (seconds, default 5 minutes)
DOC_INSERT_VERIFICATION_POLL_INTERVAL=0.5 # Poll interval (seconds, default 500ms)# Model Call Timeout Configuration 🆕
MODEL_CALL_TIMEOUT=90 # Model call max timeout (seconds, default 90)

⚡ Auto Concurrency Calculation:

  • LLM: When LLM_MAX_ASYNC is unset, auto-calculated as min(RPM, TPM/3500) ≈ 11
  • Embedding: When EMBEDDING_MAX_ASYNC is unset, auto-calculated as min(RPM, TPM/500) ≈ 800
  • Rerank: When RERANK_MAX_ASYNC is unset, auto-calculated as min(RPM, TPM/500) ≈ 800

✅ Recommended: Don't set *_MAX_ASYNC, let the system auto-calculate to completely avoid 429 errors

See env.example for complete configuration.


📚 API Documentation

Core Endpoints

1️⃣ Upload Document

# Single file upload (default mode)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc1" \
-F "file=@document.pdf" \
-F "parser=auto"# VLM chart enhancement mode 🆕# off: Markdown only (fastest, default)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc2&vlm_mode=off" \
-F "file=@document.pdf"# selective: Selective processing of important charts (balance performance and quality)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc3&vlm_mode=selective" \
-F "file=@document.pdf"# full: Complete RAG-Anything processing (highest quality, context enhancement enabled)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc4&vlm_mode=full" \
-F "file=@document.pdf"# Response
{
"task_id": "task-xxx-xxx",
"doc_id": "doc1",
"filename": "document.pdf",
"vlm_mode": "off",
"status": "pending"
}

2️⃣ Batch Upload

curl -X POST "http://localhost:8000/batch?tenant_id=your_tenant" \
-F "files=@doc1.pdf" \
-F "files=@doc2.docx" \
-F "files=@image.png"# Response
{
"batch_id": "batch-xxx-xxx",
"total_files": 3,
"accepted_files": 3,
"tasks": [...]
}

3️⃣ Intelligent Query (Query API v2.0)

New Advanced Features:

  • Conversation History: Support for multi-turn conversation context
  • Custom Prompts: Customize response style
  • Response Format Control: paragraph/list/json
  • Keyword Precision Retrieval: hl_keywords/ll_keywords
  • Streaming Output: Real-time generation viewing
# Basic query
curl -X POST "http://localhost:8000/query?tenant_id=your_tenant" \
-H "Content-Type: application/json" \
-d '{ "query": "What are the core viewpoints in the document?", "mode": "hybrid" }'# Advanced query (multi-turn dialogue + custom prompt)
curl -X POST "http://localhost:8000/query?tenant_id=your_tenant" \
-H "Content-Type: application/json" \
-d '{ "query": "Can you elaborate on the second point?", "mode": "hybrid", "conversation_history": [ {"role": "user", "content": "What are the key points?"}, {"role": "assistant", "content": "There are mainly three points..."} ], "user_prompt": "Please answer in professional academic language", "response_type": "list" }'# Streaming query (SSE)
curl -N -X POST "http://localhost:8000/query/stream?tenant_id=your_tenant" \
-H "Content-Type: application/json" \
-d '{ "query": "What are the core viewpoints in the document?", "mode": "hybrid" }'# Response (real-time streaming output)
data: {"chunk": "Based on", "done": false}
data: {"chunk": "document content", "done": false}
data: {"done": true}

4️⃣ Task Status Query

curl "http://localhost:8000/task/task-xxx-xxx?tenant_id=your_tenant"# Response
{
"task_id": "task-xxx-xxx",
"status": "completed",
"progress": 100,
"result": {...}
}

5️⃣ Tenant Management

# Get tenant statistics
curl "http://localhost:8000/tenants/stats?tenant_id=your_tenant"# Clear tenant cache
curl -X DELETE "http://localhost:8000/tenants/cache?tenant_id=your_tenant"# View instance pool status (admin)
curl "http://localhost:8000/tenants/pool/stats"

VLM Mode Comparison 🆕

ModeSpeedQualityResource UsageUse Case
off⚡⚡⚡⚡⚡⭐⭐⭐Very LowPlain text documents, fast batch processing
selective⚡⚡⚡⚡⭐⭐⭐⭐LowDocuments with key charts (recommended)
full⚡⚡⭐⭐⭐⭐⭐HighChart-intensive research reports, papers

Processing Time Estimate (20-page PDF example):

  • off: ~10 seconds(Markdown only)
  • selective: ~30 seconds(5-10 important charts)
  • full: ~120 seconds(complete context processing)

Query Mode Comparison

ModeSpeedAccuracyUse Case
naive⚡⚡⚡⚡⚡⭐⭐⭐Simple Q&A, fast retrieval
local⚡⚡⚡⚡⭐⭐⭐⭐Local entity relationship queries
global⚡⚡⚡⭐⭐⭐⭐Global knowledge graph reasoning
hybrid⚡⚡⚡⭐⭐⭐⭐⭐Hybrid retrieval (recommended)
mix⚡⚡⭐⭐⭐⭐⭐Complex questions, deep analysis

Query API v2.0 Advanced Parameters

ParameterTypeDescriptionExample
conversation_historyList[Dict]Multi-turn conversation context[{"role": "user", "content": "..."}]
user_promptstrCustom prompt"Please answer in professional academic language"
response_typestrResponse format"paragraph", "list", "json"
hl_keywordsList[str]High priority keywords["artificial intelligence", "machine learning"]
ll_keywordsList[str]Low priority keywords["application", "case study"]
only_need_contextboolReturn context only (debug)true
max_entity_tokensintEntity token limit6000

Complete API documentation:http://localhost:8000/docs


🎯 Usage Examples

Python SDK

importrequests# ConfigurationBASE_URL="http://localhost:8000"TENANT_ID="your_tenant"# Upload documentwithopen("document.pdf", "rb") asf:
response=requests.post(
f"{BASE_URL}/insert",
params={"tenant_id": TENANT_ID, "doc_id": "doc1"},
files={"file": f}
)
task_id=response.json()["task_id"]
print(f"Task ID: {task_id}")
# Queryresponse=requests.post(
f"{BASE_URL}/query",
params={"tenant_id": TENANT_ID},
json={
"query": "What is the main content of the document?",
"mode": "hybrid",
"top_k": 10
}
)
result=response.json()
print(f"Answer: {result['answer']}")

Complete cURL Example

# 1. Upload PDF document
TASK_ID=$(curl -X POST "http://localhost:8000/insert?tenant_id=demo&doc_id=report" \ -F "file=@report.pdf"| jq -r '.task_id')echo"Task ID: $TASK_ID"# 2. Wait for processing completionwhiletrue;do
STATUS=$(curl -s "http://localhost:8000/task/$TASK_ID?tenant_id=demo"| jq -r '.status')echo"Status: $STATUS"if [ "$STATUS"="completed" ] || [ "$STATUS"="failed" ];thenbreakfi
sleep 2
done# 3. Query document content
curl -X POST "http://localhost:8000/query?tenant_id=demo" \
-H "Content-Type: application/json" \
-d '{ "query": "What are the main conclusions of this report?", "mode": "hybrid" }'| jq '.answer'

🛠️ Deployment

System Requirements

Minimum Configuration:

  • CPU: 2 cores
  • RAM: 4GB
  • Disk: 40GB SSD
  • OS: Ubuntu 20.04+ / Debian 11+ / CentOS 8+

Recommended Configuration (Production):

  • CPU: 4 cores
  • RAM: 8GB
  • Disk: 100GB SSD
  • OS: Ubuntu 22.04 LTS

Server Deployment

Quick Deployment on Aliyun/Tencent Cloud

# SSH login to server
ssh root@your-server-ip
# Clone project
git clone https://github.com/BukeLy/rag-api.git
cd rag-api
# Run one-click deployment script
chmod +x deploy.sh
./deploy.sh
# The script will automatically:# 1. Install Docker and Docker Compose# 2. Configure environment variables# 3. Optimize system parameters# 4. Start services# 5. Verify health status

External Storage Configuration

Supports DragonflyDB + Qdrant + Memgraph external storage (enabled by default):

# Configure in .env
USE_EXTERNAL_STORAGE=true
# DragonflyDB configuration (KV Storage)
KV_STORAGE=RedisKVStorage
REDIS_URI=redis://dragonflydb:6379/0
# Qdrant configuration (vector storage)
VECTOR_STORAGE=QdrantVectorDBStorage
QDRANT_URL=http://qdrant:6333
# Memgraph configuration (graph storage)
GRAPH_STORAGE=MemgraphStorage
MEMGRAPH_URI=bolt://memgraph:7687
MEMGRAPH_USERNAME=
MEMGRAPH_PASSWORD=

See External Storage Deployment Documentation

Docker Compose Configuration

The project provides two configuration files:

FilePurposeFeatures
docker-compose.ymlProduction modeCode packaged in image, optimal performance
docker-compose.dev.ymlDevelopment modeCode mounted externally, supports hot-reload

Select configuration file:

# Production mode
docker compose -f docker-compose.yml up -d
# Development mode
docker compose -f docker-compose.dev.yml up -d

Performance Optimization

Tuning Parameters

Configure in .env:

# ⚡ Concurrency Control (Recommended: use auto-calculation)# LLM_MAX_ASYNC=8 # [Expert mode] Manually specify LLM concurrency# # Auto-calculated when unset: min(RPM, TPM/3500) ≈ 11# EMBEDDING_MAX_ASYNC=32 # [Expert mode] Manually specify Embedding concurrency# # Auto-calculated when unset: min(RPM, TPM/500) ≈ 800# RERANK_MAX_ASYNC=16 # [Expert mode] Manually specify Rerank concurrency# # Auto-calculated when unset: min(RPM, TPM/500) ≈ 800# Retrieval count (affects query quality and speed)
TOP_K=20 # Entity/relationship retrieval count
CHUNK_TOP_K=10 # Text chunk retrieval count# Document processing concurrency
DOCUMENT_PROCESSING_CONCURRENCY=10 # Remote mode can be set high, local mode set to 1

🎯 Concurrency Configuration Recommendations:

  • Recommended: Don't set *_MAX_ASYNC, let the system auto-calculate based on TPM/RPM
  • Expert mode: If manual control needed, can set LLM_MAX_ASYNC and other parameters
  • Advantage: Auto-calculation completely avoids 429 errors (TPM limit reached)

Mode Selection

  • MinerU Remote Mode (Recommended): High concurrency, resource-efficient
  • MinerU Local Mode: Requires GPU, high memory usage
  • Docling Mode: Fast and lightweight, suitable for simple documents

🏢 Multi-tenant Usage

Tenant Isolation

Each tenant has:

  • ✅ Independent LightRAG instance
  • ✅ Isolated data storage space
  • ✅ Independent vector index
  • ✅ Dedicated knowledge graph
  • Independent service configuration (LLM, Embedding, Rerank, DeepSeek-OCR, MinerU)🆕

Tenant Configuration Management 🆕

Each tenant can independently configure 5 services with hot-reload support:

# 1️⃣ Configure independent DeepSeek-OCR API key for Tenant A
curl -X PUT "http://localhost:8000/tenants/tenant_a/config" \
-H "Content-Type: application/json" \
-d '{ "ds_ocr_config": { "api_key": "sk-tenant-a-ds-ocr-key", "base_url": "https://api.siliconflow.cn/v1", "model": "deepseek-ai/DeepSeek-OCR", "timeout": 90 } }'# 2️⃣ Configure independent MinerU API token for Tenant B
curl -X PUT "http://localhost:8000/tenants/tenant_b/config" \
-H "Content-Type: application/json" \
-d '{ "mineru_config": { "api_token": "tenant-b-mineru-token", "base_url": "https://mineru.net", "model_version": "vlm" } }'# 3️⃣ Configure multiple services simultaneously (LLM + Embedding + DeepSeek-OCR)
curl -X PUT "http://localhost:8000/tenants/tenant_c/config" \
-H "Content-Type: application/json" \
-d '{ "llm_config": { "api_key": "sk-tenant-c-llm-key", "model": "gpt-4" }, "embedding_config": { "api_key": "sk-tenant-c-embedding-key", "model": "Qwen/Qwen3-Embedding-0.6B", "dim": 1024 }, "ds_ocr_config": { "api_key": "sk-tenant-c-ds-ocr-key" } }'# 4️⃣ Query tenant configuration (API key auto-masked)
curl "http://localhost:8000/tenants/tenant_a/config"# Response example
{
"tenant_id": "tenant_a",
"ds_ocr_config": {
"api_key": "sk-***-key", // Auto-masked
"timeout": 90
},
"merged_config": {
"llm": {...}, // Using Global Config
"embedding": {...}, // Using Global Config
"rerank": {...}, // Using Global Config
"ds_ocr": {...}, // Using tenant config
"mineru": {...} // Using Global Config
}
}
# 5️⃣ Refresh config cache (config hot-reload)
curl -X POST "http://localhost:8000/tenants/tenant_a/config/refresh"# 6️⃣ Delete tenant config (restore to global config)
curl -X DELETE "http://localhost:8000/tenants/tenant_a/config"

Supported Configuration Items:

ServiceConfig FieldDescription
LLMllm_configModel, API key, base_url, etc.
Embeddingembedding_configModel, API key, dimension, etc.
Rerankrerank_configModel, API key, etc.
DeepSeek-OCRds_ocr_configAPI key, timeout, mode, etc.
MinerUmineru_configAPI token, version, timeout, etc.

Configuration Priority: Tenant config > Global config

Use Cases:

  • 🔐 Multi-tenant SaaS: Each tenant uses their own API key
  • 💰 Pay-per-use: Track tenant usage through independent API keys
  • 🎯 Differentiated Services: Different tenants use different models (GPT-4 vs GPT-3.5)
  • 🧪 A/B Testing: Compare different models/parameters

Usage

All APIs require tenant_id parameter:

# Tenant A upload document
curl -X POST "http://localhost:8000/insert?tenant_id=tenant_a&doc_id=doc1" \
-F "file=@doc.pdf"# Tenant B upload document (fully isolated)
curl -X POST "http://localhost:8000/insert?tenant_id=tenant_b&doc_id=doc1" \
-F "file=@doc.pdf"# Tenant A query (can only query own documents)
curl -X POST "http://localhost:8000/query?tenant_id=tenant_a" \
-H "Content-Type: application/json" \
-d '{"query": "document content", "mode": "hybrid"}'

Instance Pool Management

  • Capacity: Cache up to 50 tenant instances
  • Strategy: LRU (Least Recently Used) automatic cleanup
  • Config Isolation: Each tenant can use independent LLM, Embedding, parser configuration

📊 Monitoring & Maintenance

Common Commands

# View service status
docker compose ps
# View real-time logs
docker compose logs -f
# Restart services
docker compose restart
# Stop services
docker compose down
# View resource usage
docker stats
# Clean Docker resources
docker system prune -f

Maintenance Scripts

# Monitor service health
./scripts/monitor.sh
# Backup data
./scripts/backup.sh
# Update services
./scripts/update.sh
# Performance testing
./scripts/test_concurrent_perf.sh
# Performance monitoring
./scripts/monitor_performance.sh

Health Checks

# Complete health check (recommended)
./scripts/health_check.sh
./scripts/health_check.sh --verbose # verbose output# API health check
curl http://localhost:8000/
# Tenant statistics
curl "http://localhost:8000/tenants/stats?tenant_id=your_tenant"# Instance pool status
curl "http://localhost:8000/tenants/pool/stats"

🗂️ Project Structure

rag-api/
├── main.py # FastAPI application entry
├── api/ # API route modules
│ ├── __init__.py # Route aggregation
│ ├── insert.py # Document upload (single/batch)
│ ├── query.py # Intelligent query
│ ├── task.py # Task status query
│ ├── tenant.py # Tenant management
│ ├── files.py # File service
│ ├── models.py # Pydantic models
│ └── task_store.py # Task storage
├── src/ # Core business logic
│ ├── rag.py # LightRAG lifecycle management
│ ├── multi_tenant.py # Multi-tenant instance manager
│ ├── tenant_deps.py # Tenant dependency injection
│ ├── logger.py # Unified logging
│ ├── metrics.py # Performance metrics
│ ├── file_url_service.py # Temporary file service
│ ├── mineru_client.py # MinerU client
│ └── mineru_result_processor.py # Result processing
├── docs/ # Documentation
│ ├── ARCHITECTURE.md # Architecture design documentation
│ ├── USAGE.md # Detailed usage guide
│ ├── DEPLOY_MODES.md # Deployment mode description
│ ├── PR_WORKFLOW.md # PR workflow
│ └── ...
├── scripts/ # Maintenance scripts
│ ├── dev.sh # Development mode quick start
│ ├── monitor.sh # Service monitoring
│ ├── backup.sh # Data backup
│ ├── update.sh # Service update
│ └── ...
├── deploy.sh # One-click deployment script
├── docker-compose.yml # Production mode configuration
├── docker-compose.dev.yml # Development mode configuration
├── Dockerfile # Production image
├── Dockerfile.dev # Development image
├── pyproject.toml # Project dependencies
├── uv.lock # Dependency lock
├── env.example # Environment variable template
├── CLAUDE.md # Claude AI guide
└── README.md # This documentation

🐛 Troubleshooting

Common Issues

Q1: What to do if service fails to start?
# View detailed logs
docker compose logs
# Check port usage
netstat -tulpn | grep 8000
# Check Docker status
docker ps -a
Q2: multimodal_processed error?

Note: This issue has been fixed in LightRAG 1.4.9.4+. If you encounter this error, your version is outdated.

Solution:

# Option 1: Upgrade to latest version (recommended)# Modify LightRAG version in pyproject.toml# lightrag = "^1.4.9.4"# Rebuild image
docker compose down
docker compose up -d --build
# Option 2: Clean old data (temporary solution)
rm -rf ./rag_local_storage
docker compose restart
Q3: File upload returns 400 error?

Check:

  • File format supported (PDF, DOCX, PNG, JPG, etc.)
  • File size exceeds 100MB
  • File is empty
# View supported formats
curl http://localhost:8000/docs
Q3.5: Embedding dimension error?

If you encounter dimension-related errors, need to clean data and rebuild:

# Stop services
docker compose down
# Delete all volumes (clear database)
docker volume rm rag-api_dragonflydb_data rag-api_qdrant_data rag-api_memgraph_data
# Modify EMBEDDING_DIM in .env
EMBEDDING_DIM=1024 # or 4096, must match the model# Restart
docker compose up -d
Q4: Query is very slow (>30 seconds)?

Optimization suggestions:

  1. Use naive or hybrid mode instead of mix
  2. Increase MAX_ASYNC parameter (in .env)
  3. Reduce TOP_K and CHUNK_TOP_K
  4. Enable Reranker
# Modify .env
MAX_ASYNC=8
TOP_K=20
CHUNK_TOP_K=10
Q5: Out of memory (OOM)?

If using local MinerU:

# Switch to remote mode# Modify in .env
MINERU_MODE=remote
MINERU_API_TOKEN=your_token
# Or limit concurrency
DOCUMENT_PROCESSING_CONCURRENCY=1
Q6: Tasks lost after container restart?

Problem Symptoms:

  • Cannot query previous task status after container restart
  • Tasks disappear after tenant instance evicted by LRU

Solution: Enable Redis task storage

# Modify .env
TASK_STORE_STORAGE=redis
# Restart services
docker compose restart
# Verify
docker compose logs api | grep TaskStore
# Should see: ✅ TaskStore: Redis connection successful

Configuration Description:

  • memory mode: In-memory storage, data lost after restart (default, suitable for development)
  • redis mode: Persistent storage, supports container restart and instance rebuild (production recommended)

TTL Strategy (Redis mode auto-cleanup):

  • completed tasks: 24 hours
  • failed tasks: 24 hours
  • pending/processing tasks: 6 hours
Q7: VLM mode processing failed?

Check Items:

  1. vision_model_func not configured

    • Check logs:vision_model_func not found, fallback to off mode
    • Ensure LLM API is configured in .env
  2. Image file does not exist

    • Check logs:Image file not found: xxx
    • Possibly corrupted MinerU ZIP or extraction failed
  3. Timeout error

    • full mode may timeout on large files
    • Suggestion: Use selective mode first, or increase VLM_TIMEOUT
# Modify .env
VLM_TIMEOUT=300 # Increase to 5 minutes
RAG_VLM_MODE=selective # downgrade to selective

Debugging Tips:

# View detailed logs
docker compose logs -f | grep VLM
# Test single file
curl -X POST 'http://localhost:8000/insert?tenant_id=test&doc_id=test&vlm_mode=off' \
-F 'file=@test.pdf'

Performance Tuning Recommendations

ScenarioMAX_ASYNCTOP_KCHUNK_TOP_KMINERU_MODE
Fast response8105remote
Balanced mode82010remote
High accuracy46020remote
Resource limited42010remote

📖 Documentation


🤝 Contributing

We welcome all forms of contribution!

How to Contribute

  1. Fork the project
git clone https://github.com/BukeLy/rag-api.git
cd rag-api
  1. Create feature branch
git checkout -b feature/your-feature-name
  1. Development and Testing
# Install dependencies
uv sync
# Run tests
uv run pytest
# Code formatting
uv run black .
uv run isort .
  1. Submit code
git add .
git commit -m "feat: Add new feature"
git push origin feature/your-feature-name
  1. Create Pull Request

Create a PR on GitHub with detailed description of your changes.

Commit Conventions

Use semantic commit messages:

  • feat: New feature
  • fix: Bug fix
  • docs: Documentation update
  • style: Code formatting
  • refactor: Code refactoring
  • perf: Performance optimization
  • test: Testing
  • chore: Build/tools

See PR Workflow Documentation


📄 License

This project is licensed under the MIT License. See the LICENSE file for details.


🙏 Acknowledgments

This project is built on the following excellent open source projects:

  • LightRAG - Efficient knowledge graph RAG framework
  • RAG-Anything - Multimodal document parsing
  • MinerU - Powerful PDF parsing tool
  • Docling - Lightweight document parsing
  • FastAPI - Modern Python web framework

Special thanks to all contributors and users for their support! 🎉


📬 Contact Us


⭐ If this project helps you, please give it a Star!

Made with ❤️ by BukeLy

© 2025 RAG API. All rights reserved.

About

Multi-tenant RAG API powered by LightRAG/RAG-Anything. Auto-selects best parser (DeepSeek-OCR/MinerU/Docling) via complexity scoring

Topics

Resources

Stars

64 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

🚀 RAG API

Multi-tenant Multimodal Document Intelligent Retrieval System

Enterprise-grade RAG service built on RAG-Anything and LightRAG

CIPythonFastAPILightRAGDockerLicense

English | 简体中文

FeaturesQuick StartArchitectureAPI DocumentationDeployment


📖 Introduction

RAG API is an enterprise-grade Retrieval-Augmented Generation (RAG) service that combines the powerful document parsing capabilities of RAG-Anything with the efficient knowledge graph retrieval technology of LightRAG, providing intelligent Q&A capabilities for your documents.

🎯 Key Highlights

  • 🏢 Multi-tenant Isolation - Complete tenant data isolation for enterprise multi-tenant scenarios
  • 🎨 Multimodal Parsing - Support for PDF, Word, images and more, with full OCR, tables, and formulas coverage
  • High-performance Retrieval - Knowledge graph-based hybrid retrieval with 6-15 second query response
  • 🔄 Flexible Deployment - Support for production and development modes with one-click switching
  • 📦 Ready to Use - One-click Docker deployment, service starts in 3 minutes
  • 🎛️ Multiple Parsing Engines - DeepSeek-OCR (Remote API) + MinerU (Local/Remote API) + Docling (Fast)
  • 🎨 RAG-Anything VLM Enhancement - Three modes (off/selective/full) for deep chart understanding
  • 💾 Task Persistence - Redis storage support, tasks recoverable after container restart/instance rebuild

✨ Features

📄 Document Processing

  • Multiple Format Support

    • PDF, Word, Excel, PPT
    • PNG, JPG, WebP images
    • TXT, Markdown text
  • Intelligent Parsing

    • Plain text (.txt, .md) → Direct insertion (ultra-fast ~1s, skip parser)
    • OCR text recognition
    • Structured table extraction
    • Mathematical formula recognition
    • Layout analysis
  • RAG-Anything VLM Enhancement 🆕

    • off - Markdown only (fastest)
    • selective - Selective processing of important charts
    • full - Complete context enhancement processing
    • Smart filtering: with titles, large size, first page content
    • ⚠️Only supports remote MinerU mode, local mode uses RAG-Anything native methods
  • Batch Processing

    • Up to 100 files per batch
    • Async task queue
    • Real-time progress tracking

🔍 Intelligent Retrieval

  • Multi-mode Query

    • naive - Vector retrieval (fastest)
    • local - Local graph
    • global - Global graph
    • hybrid - Hybrid retrieval
    • mix - Full retrieval (most accurate)
  • Knowledge Graph

    • Automatic entity extraction
    • Relationship reasoning
    • Semantic understanding
    • Context enhancement
  • External Storage

    • DragonflyDB (KV storage + task storage)
    • Qdrant (vector storage)
    • Memgraph (graph database)
    • Task persistence (Redis mode)

🏗️ Architecture

System Architecture Diagram

graph TB
subgraph "Client Layer"
Client[Client Application]
WebUI[Web Interface]
end
subgraph "API Gateway Layer"
FastAPI[FastAPI Service]
Auth[Tenant Authentication]
end
subgraph "Business Logic Layer"
TenantMgr[Tenant Manager]
TaskQueue[Task Queue]
subgraph "Document Processing"
DeepSeekOCR[DeepSeek-OCR<br/>Fast OCR 80% cases]
MinerU[MinerU Parser<br/>Complex multimodal]
Docling[Docling Parser<br/>Fast lightweight]
FileRouter[Smart Router<br/>Complexity scoring]
end
subgraph "RAG Engine"
LightRAG[LightRAG Instance Pool<br/>LRU Cache 50]
KG[Knowledge Graph Engine]
Vector[Vector Retrieval Engine]
end
end
subgraph "Storage Layer"
DragonflyDB[(DragonflyDB<br/>KV Storage)]
Qdrant[(Qdrant<br/>Vector Database)]
Memgraph[(Memgraph<br/>Graph Database)]
Local[(Local Files<br/>Temp Storage)]
end
subgraph "External Services"
LLM[LLM<br/>Entity Extraction/Generation]
Embedding[Embedding<br/>Vectorization]
Rerank[Rerank<br/>Reranking]
end
Client --> FastAPI
WebUI --> FastAPI
FastAPI --> Auth
Auth --> TenantMgr
TenantMgr --> TaskQueue
TenantMgr --> LightRAG
TaskQueue --> FileRouter
FileRouter --> DeepSeekOCR
FileRouter --> MinerU
FileRouter --> Docling
DeepSeekOCR --> LightRAG
MinerU --> LightRAG
Docling --> LightRAG
LightRAG --> KG
LightRAG --> Vector
KG --> DragonflyDB
KG --> Memgraph
Vector --> Qdrant
LightRAG --> Local
LightRAG --> LLM
LightRAG --> Embedding
Vector --> Rerank
style FastAPI fill:#00C7B7
style LightRAG fill:#FF6B6B
style DeepSeekOCR fill:#5DADE2
style MinerU fill:#4ECDC4
style Docling fill:#95E1D3
style TenantMgr fill:#F38181
Loading

Multi-tenant Architecture

graph TB
subgraph "Tenant A"
A_Config[Tenant A Config<br/>Independent API Key]
A_Instance[LightRAG Instance A<br/>Dedicated LLM/Embedding]
A_Data[(Tenant A Data<br/>Fully Isolated)]
A_Config --> A_Instance
A_Instance --> A_Data
end
subgraph "Tenant B"
B_Config[Tenant B Config<br/>Independent API Key]
B_Instance[LightRAG Instance B<br/>Dedicated LLM/Embedding]
B_Data[(Tenant B Data<br/>Fully Isolated)]
B_Config --> B_Instance
B_Instance --> B_Data
end
subgraph "Tenant C"
C_Config[Using Global Config]
C_Instance[LightRAG Instance C<br/>Shared LLM/Embedding]
C_Data[(Tenant C Data<br/>Fully Isolated)]
C_Config --> C_Instance
C_Instance --> C_Data
end
Pool[Instance Pool Manager<br/>LRU Cache + Config Isolation]
Global[Global Config<br/>Default API Key]
Pool --> A_Instance
Pool --> B_Instance
Pool --> C_Instance
C_Config -.fallback.-> Global
style Pool fill:#F38181
style Global fill:#95E1D3
style A_Config fill:#FFD93D
style B_Config fill:#FFD93D
style C_Config fill:#E8E8E8
Loading

Core Technology Stack

🔧 Frameworks & Runtime

  • FastAPI 0.115+
  • Python 3.11+
  • Uvicorn
  • Docker & Docker Compose

🧠 AI & RAG

  • LightRAG 1.4.9.4
  • RAG-Anything
  • MinerU (PDF-Extract-Kit)
  • Docling

💾 Storage & Database

  • DragonflyDB(Redis compatible)
  • Qdrant(Vector Database)
  • Memgraph(Graph Database)
  • Local filesystem

🚀 Quick Start

Option 1: One-click Deployment (Recommended)

Suitable for production and testing environments:

# 1. Clone the project
git clone https://github.com/BukeLy/rag-api.git
cd rag-api
# 2. Configure environment variables
cp env.example .env
nano .env # Fill in your API keys# 3. Run deployment script
chmod +x deploy.sh
./deploy.sh
# Select deployment mode:# 1) Production Mode - Standard container deployment# 2) Development Mode - Code hot-reload# 4. Verify service
curl http://localhost:8000/

Access Swagger Documentation:http://localhost:8000/docs

Option 2: Docker Compose

Production Mode

# Configure environment variables
cp env.example .env
nano .env
# Start services
docker compose -f docker-compose.yml up -d
# View logs
docker compose -f docker-compose.yml logs -f

Development Mode (Code Hot-reload)

# Start development environment
docker compose -f docker-compose.dev.yml up -d
# Or use quick script
./scripts/dev.sh
# Code changes will auto-reload without restart

Option 3: Local Development

# Install uv (Python package manager)
curl -LsSf https://astral.sh/uv/install.sh | sh
# Install dependencies
uv sync
# Configure environment variables
cp env.example .env
nano .env
# Start services
uv run uvicorn main:app --host 0.0.0.0 --port 8000 --reload

Environment Variable Configuration

Minimum configuration (required):

# LLM Configuration (Function-oriented naming)
LLM_API_KEY=your_llm_api_key
LLM_BASE_URL=https://ark.cn-beijing.volces.com/api/v3
LLM_MODEL=ep-xxx-xxx
# LLM_REQUESTS_PER_MINUTE=800 # Rate limit (optional)# LLM_TOKENS_PER_MINUTE=40000 # Rate limit (optional)# LLM_MAX_ASYNC=8 # [Optional, expert mode] Manual concurrency control# # Auto-calculated when unset: min(RPM, TPM/3500) = 11# Embedding Configuration (Function-oriented naming)
EMBEDDING_API_KEY=your_embedding_api_key
EMBEDDING_BASE_URL=https://api.siliconflow.cn/v1
EMBEDDING_MODEL=Qwen/Qwen3-Embedding-0.6B
EMBEDDING_DIM=1024
# EMBEDDING_MAX_ASYNC=32 # [Optional, expert mode] Auto-calculated when unset: 800# MinerU Mode (Remote recommended)
MINERU_MODE=remote
MINERU_API_TOKEN=your_token
MINERU_HTTP_TIMEOUT=60 # MinerU download timeout (seconds, default 60)
FILE_SERVICE_BASE_URL=http://your-ip:8000
# VLM Chart Enhancement Configuration 🆕# ⚠️ Note: Only effective in MINERU_MODE=remote
RAG_VLM_MODE=off # off / selective / full
RAG_IMPORTANCE_THRESHOLD=0.5 # Importance threshold (selective mode)
RAG_CONTEXT_WINDOW=2 # Context window (full mode)
RAG_CONTEXT_MODE=page # page / chunk
RAG_MAX_CONTEXT_TOKENS=3000 # Max context tokens# Task Storage Configuration 🆕
TASK_STORE_STORAGE=redis # memory / redis (production recommends redis)# Document Insert Verification Configuration 🆕
DOC_INSERT_VERIFICATION_TIMEOUT=300 # Verification timeout (seconds, default 5 minutes)
DOC_INSERT_VERIFICATION_POLL_INTERVAL=0.5 # Poll interval (seconds, default 500ms)# Model Call Timeout Configuration 🆕
MODEL_CALL_TIMEOUT=90 # Model call max timeout (seconds, default 90)

⚡ Auto Concurrency Calculation:

  • LLM: When LLM_MAX_ASYNC is unset, auto-calculated as min(RPM, TPM/3500) ≈ 11
  • Embedding: When EMBEDDING_MAX_ASYNC is unset, auto-calculated as min(RPM, TPM/500) ≈ 800
  • Rerank: When RERANK_MAX_ASYNC is unset, auto-calculated as min(RPM, TPM/500) ≈ 800

✅ Recommended: Don't set *_MAX_ASYNC, let the system auto-calculate to completely avoid 429 errors

See env.example for complete configuration.


📚 API Documentation

Core Endpoints

1️⃣ Upload Document

# Single file upload (default mode)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc1" \
-F "file=@document.pdf" \
-F "parser=auto"# VLM chart enhancement mode 🆕# off: Markdown only (fastest, default)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc2&vlm_mode=off" \
-F "file=@document.pdf"# selective: Selective processing of important charts (balance performance and quality)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc3&vlm_mode=selective" \
-F "file=@document.pdf"# full: Complete RAG-Anything processing (highest quality, context enhancement enabled)
curl -X POST "http://localhost:8000/insert?tenant_id=your_tenant&doc_id=doc4&vlm_mode=full" \
-F "file=@document.pdf"# Response
{
"task_id": "task-xxx-xxx",
"doc_id": "doc1",
"filename": "document.pdf",
"vlm_mode": "off",
"status": "pending"
}

2️⃣ Batch Upload

curl -X POST "http://localhost:8000/batch?tenant_id=your_tenant" \
-F "files=@doc1.pdf" \
-F "files=@doc2.docx" \
-F "files=@image.png"# Response
{
"batch_id": "batch-xxx-xxx",
"total_files": 3,
"accepted_files": 3,
"tasks": [...]
}

3️⃣ Intelligent Query (Query API v2.0)

New Advanced Features:

  • Conversation History: Support for multi-turn conversation context
  • Custom Prompts: Customize response style
  • Response Format Control: paragraph/list/json
  • Keyword Precision Retrieval: hl_keywords/ll_keywords
  • Streaming Output: Real-time generation viewing
# Basic query
curl -X POST "http://localhost:8000/query?tenant_id=your_tenant" \
-H "Content-Type: application/json" \
-d '{ "query": "What are the core viewpoints in the document?", "mode": "hybrid" }'# Advanced query (multi-turn dialogue + custom prompt)
curl -X POST "http://localhost:8000/query?tenant_id=your_tenant" \
-H "Content-Type: application/json" \
-d '{ "query": "Can you elaborate on the second point?", "mode": "hybrid", "conversation_history": [ {"role": "user", "content": "What are the key points?"}, {"role": "assistant", "content": "There are mainly three points..."} ], "user_prompt": "Please answer in professional academic language", "response_type": "list" }'# Streaming query (SSE)
curl -N -X POST "http://localhost:8000/query/stream?tenant_id=your_tenant" \
-H "Content-Type: application/json" \
-d '{ "query": "What are the core viewpoints in the document?", "mode": "hybrid" }'# Response (real-time streaming output)
data: {"chunk": "Based on", "done": false}
data: {"chunk": "document content", "done": false}
data: {"done": true}

4️⃣ Task Status Query

curl "http://localhost:8000/task/task-xxx-xxx?tenant_id=your_tenant"# Response
{
"task_id": "task-xxx-xxx",
"status": "completed",
"progress": 100,
"result": {...}
}

5️⃣ Tenant Management

# Get tenant statistics
curl "http://localhost:8000/tenants/stats?tenant_id=your_tenant"# Clear tenant cache
curl -X DELETE "http://localhost:8000/tenants/cache?tenant_id=your_tenant"# View instance pool status (admin)
curl "http://localhost:8000/tenants/pool/stats"

VLM Mode Comparison 🆕

ModeSpeedQualityResource UsageUse Case
off⚡⚡⚡⚡⚡⭐⭐⭐Very LowPlain text documents, fast batch processing
selective⚡⚡⚡⚡⭐⭐⭐⭐LowDocuments with key charts (recommended)
full⚡⚡⭐⭐⭐⭐⭐HighChart-intensive research reports, papers

Processing Time Estimate (20-page PDF example):

  • off: ~10 seconds(Markdown only)
  • selective: ~30 seconds(5-10 important charts)
  • full: ~120 seconds(complete context processing)

Query Mode Comparison

ModeSpeedAccuracyUse Case
naive⚡⚡⚡⚡⚡⭐⭐⭐Simple Q&A, fast retrieval
local⚡⚡⚡⚡⭐⭐⭐⭐Local entity relationship queries
global⚡⚡⚡⭐⭐⭐⭐Global knowledge graph reasoning
hybrid⚡⚡⚡⭐⭐⭐⭐⭐Hybrid retrieval (recommended)
mix⚡⚡⭐⭐⭐⭐⭐Complex questions, deep analysis

Query API v2.0 Advanced Parameters

ParameterTypeDescriptionExample
conversation_historyList[Dict]Multi-turn conversation context[{"role": "user", "content": "..."}]
user_promptstrCustom prompt"Please answer in professional academic language"
response_typestrResponse format"paragraph", "list", "json"
hl_keywordsList[str]High priority keywords["artificial intelligence", "machine learning"]
ll_keywordsList[str]Low priority keywords["application", "case study"]
only_need_contextboolReturn context only (debug)true
max_entity_tokensintEntity token limit6000

Complete API documentation:http://localhost:8000/docs


🎯 Usage Examples

Python SDK

importrequests# ConfigurationBASE_URL="http://localhost:8000"TENANT_ID="your_tenant"# Upload documentwithopen("document.pdf", "rb") asf:
response=requests.post(
f"{BASE_URL}/insert",
params={"tenant_id": TENANT_ID, "doc_id": "doc1"},
files={"file": f}
)
task_id=response.json()["task_id"]
print(f"Task ID: {task_id}")
# Queryresponse=requests.post(
f"{BASE_URL}/query",
params={"tenant_id": TENANT_ID},
json={
"query": "What is the main content of the document?",
"mode": "hybrid",
"top_k": 10
}
)
result=response.json()
print(f"Answer: {result['answer']}")

Complete cURL Example

# 1. Upload PDF document
TASK_ID=$(curl -X POST "http://localhost:8000/insert?tenant_id=demo&doc_id=report" \ -F "file=@report.pdf"| jq -r '.task_id')echo"Task ID: $TASK_ID"# 2. Wait for processing completionwhiletrue;do
STATUS=$(curl -s "http://localhost:8000/task/$TASK_ID?tenant_id=demo"| jq -r '.status')echo"Status: $STATUS"if [ "$STATUS"="completed" ] || [ "$STATUS"="failed" ];thenbreakfi
sleep 2
done# 3. Query document content
curl -X POST "http://localhost:8000/query?tenant_id=demo" \
-H "Content-Type: application/json" \
-d '{ "query": "What are the main conclusions of this report?", "mode": "hybrid" }'| jq '.answer'

🛠️ Deployment

System Requirements

Minimum Configuration:

  • CPU: 2 cores
  • RAM: 4GB
  • Disk: 40GB SSD
  • OS: Ubuntu 20.04+ / Debian 11+ / CentOS 8+

Recommended Configuration (Production):

  • CPU: 4 cores
  • RAM: 8GB
  • Disk: 100GB SSD
  • OS: Ubuntu 22.04 LTS

Server Deployment

Quick Deployment on Aliyun/Tencent Cloud

# SSH login to server
ssh root@your-server-ip
# Clone project
git clone https://github.com/BukeLy/rag-api.git
cd rag-api
# Run one-click deployment script
chmod +x deploy.sh
./deploy.sh
# The script will automatically:# 1. Install Docker and Docker Compose# 2. Configure environment variables# 3. Optimize system parameters# 4. Start services# 5. Verify health status

External Storage Configuration

Supports DragonflyDB + Qdrant + Memgraph external storage (enabled by default):

# Configure in .env
USE_EXTERNAL_STORAGE=true
# DragonflyDB configuration (KV Storage)
KV_STORAGE=RedisKVStorage
REDIS_URI=redis://dragonflydb:6379/0
# Qdrant configuration (vector storage)
VECTOR_STORAGE=QdrantVectorDBStorage
QDRANT_URL=http://qdrant:6333
# Memgraph configuration (graph storage)
GRAPH_STORAGE=MemgraphStorage
MEMGRAPH_URI=bolt://memgraph:7687
MEMGRAPH_USERNAME=
MEMGRAPH_PASSWORD=

See External Storage Deployment Documentation

Docker Compose Configuration

The project provides two configuration files:

FilePurposeFeatures
docker-compose.ymlProduction modeCode packaged in image, optimal performance
docker-compose.dev.ymlDevelopment modeCode mounted externally, supports hot-reload

Select configuration file:

# Production mode
docker compose -f docker-compose.yml up -d
# Development mode
docker compose -f docker-compose.dev.yml up -d

Performance Optimization

Tuning Parameters

Configure in .env:

# ⚡ Concurrency Control (Recommended: use auto-calculation)# LLM_MAX_ASYNC=8 # [Expert mode] Manually specify LLM concurrency# # Auto-calculated when unset: min(RPM, TPM/3500) ≈ 11# EMBEDDING_MAX_ASYNC=32 # [Expert mode] Manually specify Embedding concurrency# # Auto-calculated when unset: min(RPM, TPM/500) ≈ 800# RERANK_MAX_ASYNC=16 # [Expert mode] Manually specify Rerank concurrency# # Auto-calculated when unset: min(RPM, TPM/500) ≈ 800# Retrieval count (affects query quality and speed)
TOP_K=20 # Entity/relationship retrieval count
CHUNK_TOP_K=10 # Text chunk retrieval count# Document processing concurrency
DOCUMENT_PROCESSING_CONCURRENCY=10 # Remote mode can be set high, local mode set to 1

🎯 Concurrency Configuration Recommendations:

  • Recommended: Don't set *_MAX_ASYNC, let the system auto-calculate based on TPM/RPM
  • Expert mode: If manual control needed, can set LLM_MAX_ASYNC and other parameters
  • Advantage: Auto-calculation completely avoids 429 errors (TPM limit reached)

Mode Selection

  • MinerU Remote Mode (Recommended): High concurrency, resource-efficient
  • MinerU Local Mode: Requires GPU, high memory usage
  • Docling Mode: Fast and lightweight, suitable for simple documents

🏢 Multi-tenant Usage

Tenant Isolation

Each tenant has:

  • ✅ Independent LightRAG instance
  • ✅ Isolated data storage space
  • ✅ Independent vector index
  • ✅ Dedicated knowledge graph
  • Independent service configuration (LLM, Embedding, Rerank, DeepSeek-OCR, MinerU)🆕

Tenant Configuration Management 🆕

Each tenant can independently configure 5 services with hot-reload support:

# 1️⃣ Configure independent DeepSeek-OCR API key for Tenant A
curl -X PUT "http://localhost:8000/tenants/tenant_a/config" \
-H "Content-Type: application/json" \
-d '{ "ds_ocr_config": { "api_key": "sk-tenant-a-ds-ocr-key", "base_url": "https://api.siliconflow.cn/v1", "model": "deepseek-ai/DeepSeek-OCR", "timeout": 90 } }'# 2️⃣ Configure independent MinerU API token for Tenant B
curl -X PUT "http://localhost:8000/tenants/tenant_b/config" \
-H "Content-Type: application/json" \
-d '{ "mineru_config": { "api_token": "tenant-b-mineru-token", "base_url": "https://mineru.net", "model_version": "vlm" } }'# 3️⃣ Configure multiple services simultaneously (LLM + Embedding + DeepSeek-OCR)
curl -X PUT "http://localhost:8000/tenants/tenant_c/config" \
-H "Content-Type: application/json" \
-d '{ "llm_config": { "api_key": "sk-tenant-c-llm-key", "model": "gpt-4" }, "embedding_config": { "api_key": "sk-tenant-c-embedding-key", "model": "Qwen/Qwen3-Embedding-0.6B", "dim": 1024 }, "ds_ocr_config": { "api_key": "sk-tenant-c-ds-ocr-key" } }'# 4️⃣ Query tenant configuration (API key auto-masked)
curl "http://localhost:8000/tenants/tenant_a/config"# Response example
{
"tenant_id": "tenant_a",
"ds_ocr_config": {
"api_key": "sk-***-key", // Auto-masked
"timeout": 90
},
"merged_config": {
"llm": {...}, // Using Global Config
"embedding": {...}, // Using Global Config
"rerank": {...}, // Using Global Config
"ds_ocr": {...}, // Using tenant config
"mineru": {...} // Using Global Config
}
}
# 5️⃣ Refresh config cache (config hot-reload)
curl -X POST "http://localhost:8000/tenants/tenant_a/config/refresh"# 6️⃣ Delete tenant config (restore to global config)
curl -X DELETE "http://localhost:8000/tenants/tenant_a/config"

Supported Configuration Items:

ServiceConfig FieldDescription
LLMllm_configModel, API key, base_url, etc.
Embeddingembedding_configModel, API key, dimension, etc.
Rerankrerank_configModel, API key, etc.
DeepSeek-OCRds_ocr_configAPI key, timeout, mode, etc.
MinerUmineru_configAPI token, version, timeout, etc.

Configuration Priority: Tenant config > Global config

Use Cases:

  • 🔐 Multi-tenant SaaS: Each tenant uses their own API key
  • 💰 Pay-per-use: Track tenant usage through independent API keys
  • 🎯 Differentiated Services: Different tenants use different models (GPT-4 vs GPT-3.5)
  • 🧪 A/B Testing: Compare different models/parameters

Usage

All APIs require tenant_id parameter:

# Tenant A upload document
curl -X POST "http://localhost:8000/insert?tenant_id=tenant_a&doc_id=doc1" \
-F "file=@doc.pdf"# Tenant B upload document (fully isolated)
curl -X POST "http://localhost:8000/insert?tenant_id=tenant_b&doc_id=doc1" \
-F "file=@doc.pdf"# Tenant A query (can only query own documents)
curl -X POST "http://localhost:8000/query?tenant_id=tenant_a" \
-H "Content-Type: application/json" \
-d '{"query": "document content", "mode": "hybrid"}'

Instance Pool Management

  • Capacity: Cache up to 50 tenant instances
  • Strategy: LRU (Least Recently Used) automatic cleanup
  • Config Isolation: Each tenant can use independent LLM, Embedding, parser configuration

📊 Monitoring & Maintenance

Common Commands

# View service status
docker compose ps
# View real-time logs
docker compose logs -f
# Restart services
docker compose restart
# Stop services
docker compose down
# View resource usage
docker stats
# Clean Docker resources
docker system prune -f

Maintenance Scripts

# Monitor service health
./scripts/monitor.sh
# Backup data
./scripts/backup.sh
# Update services
./scripts/update.sh
# Performance testing
./scripts/test_concurrent_perf.sh
# Performance monitoring
./scripts/monitor_performance.sh

Health Checks

# Complete health check (recommended)
./scripts/health_check.sh
./scripts/health_check.sh --verbose # verbose output# API health check
curl http://localhost:8000/
# Tenant statistics
curl "http://localhost:8000/tenants/stats?tenant_id=your_tenant"# Instance pool status
curl "http://localhost:8000/tenants/pool/stats"

🗂️ Project Structure

rag-api/
├── main.py # FastAPI application entry
├── api/ # API route modules
│ ├── __init__.py # Route aggregation
│ ├── insert.py # Document upload (single/batch)
│ ├── query.py # Intelligent query
│ ├── task.py # Task status query
│ ├── tenant.py # Tenant management
│ ├── files.py # File service
│ ├── models.py # Pydantic models
│ └── task_store.py # Task storage
├── src/ # Core business logic
│ ├── rag.py # LightRAG lifecycle management
│ ├── multi_tenant.py # Multi-tenant instance manager
│ ├── tenant_deps.py # Tenant dependency injection
│ ├── logger.py # Unified logging
│ ├── metrics.py # Performance metrics
│ ├── file_url_service.py # Temporary file service
│ ├── mineru_client.py # MinerU client
│ └── mineru_result_processor.py # Result processing
├── docs/ # Documentation
│ ├── ARCHITECTURE.md # Architecture design documentation
│ ├── USAGE.md # Detailed usage guide
│ ├── DEPLOY_MODES.md # Deployment mode description
│ ├── PR_WORKFLOW.md # PR workflow
│ └── ...
├── scripts/ # Maintenance scripts
│ ├── dev.sh # Development mode quick start
│ ├── monitor.sh # Service monitoring
│ ├── backup.sh # Data backup
│ ├── update.sh # Service update
│ └── ...
├── deploy.sh # One-click deployment script
├── docker-compose.yml # Production mode configuration
├── docker-compose.dev.yml # Development mode configuration
├── Dockerfile # Production image
├── Dockerfile.dev # Development image
├── pyproject.toml # Project dependencies
├── uv.lock # Dependency lock
├── env.example # Environment variable template
├── CLAUDE.md # Claude AI guide
└── README.md # This documentation

🐛 Troubleshooting

Common Issues

Q1: What to do if service fails to start?
# View detailed logs
docker compose logs
# Check port usage
netstat -tulpn | grep 8000
# Check Docker status
docker ps -a
Q2: multimodal_processed error?

Note: This issue has been fixed in LightRAG 1.4.9.4+. If you encounter this error, your version is outdated.

Solution:

# Option 1: Upgrade to latest version (recommended)# Modify LightRAG version in pyproject.toml# lightrag = "^1.4.9.4"# Rebuild image
docker compose down
docker compose up -d --build
# Option 2: Clean old data (temporary solution)
rm -rf ./rag_local_storage
docker compose restart
Q3: File upload returns 400 error?

Check:

  • File format supported (PDF, DOCX, PNG, JPG, etc.)
  • File size exceeds 100MB
  • File is empty
# View supported formats
curl http://localhost:8000/docs
Q3.5: Embedding dimension error?

If you encounter dimension-related errors, need to clean data and rebuild:

# Stop services
docker compose down
# Delete all volumes (clear database)
docker volume rm rag-api_dragonflydb_data rag-api_qdrant_data rag-api_memgraph_data
# Modify EMBEDDING_DIM in .env
EMBEDDING_DIM=1024 # or 4096, must match the model# Restart
docker compose up -d
Q4: Query is very slow (>30 seconds)?

Optimization suggestions:

  1. Use naive or hybrid mode instead of mix
  2. Increase MAX_ASYNC parameter (in .env)
  3. Reduce TOP_K and CHUNK_TOP_K
  4. Enable Reranker
# Modify .env
MAX_ASYNC=8
TOP_K=20
CHUNK_TOP_K=10
Q5: Out of memory (OOM)?

If using local MinerU:

# Switch to remote mode# Modify in .env
MINERU_MODE=remote
MINERU_API_TOKEN=your_token
# Or limit concurrency
DOCUMENT_PROCESSING_CONCURRENCY=1
Q6: Tasks lost after container restart?

Problem Symptoms:

  • Cannot query previous task status after container restart
  • Tasks disappear after tenant instance evicted by LRU

Solution: Enable Redis task storage

# Modify .env
TASK_STORE_STORAGE=redis
# Restart services
docker compose restart
# Verify
docker compose logs api | grep TaskStore
# Should see: ✅ TaskStore: Redis connection successful

Configuration Description:

  • memory mode: In-memory storage, data lost after restart (default, suitable for development)
  • redis mode: Persistent storage, supports container restart and instance rebuild (production recommended)

TTL Strategy (Redis mode auto-cleanup):

  • completed tasks: 24 hours
  • failed tasks: 24 hours
  • pending/processing tasks: 6 hours
Q7: VLM mode processing failed?

Check Items:

  1. vision_model_func not configured

    • Check logs:vision_model_func not found, fallback to off mode
    • Ensure LLM API is configured in .env
  2. Image file does not exist

    • Check logs:Image file not found: xxx
    • Possibly corrupted MinerU ZIP or extraction failed
  3. Timeout error

    • full mode may timeout on large files
    • Suggestion: Use selective mode first, or increase VLM_TIMEOUT
# Modify .env
VLM_TIMEOUT=300 # Increase to 5 minutes
RAG_VLM_MODE=selective # downgrade to selective

Debugging Tips:

# View detailed logs
docker compose logs -f | grep VLM
# Test single file
curl -X POST 'http://localhost:8000/insert?tenant_id=test&doc_id=test&vlm_mode=off' \
-F 'file=@test.pdf'

Performance Tuning Recommendations

ScenarioMAX_ASYNCTOP_KCHUNK_TOP_KMINERU_MODE
Fast response8105remote
Balanced mode82010remote
High accuracy46020remote
Resource limited42010remote

📖 Documentation


🤝 Contributing

We welcome all forms of contribution!

How to Contribute

  1. Fork the project
git clone https://github.com/BukeLy/rag-api.git
cd rag-api
  1. Create feature branch
git checkout -b feature/your-feature-name
  1. Development and Testing
# Install dependencies
uv sync
# Run tests
uv run pytest
# Code formatting
uv run black .
uv run isort .
  1. Submit code
git add .
git commit -m "feat: Add new feature"
git push origin feature/your-feature-name
  1. Create Pull Request

Create a PR on GitHub with detailed description of your changes.

Commit Conventions

Use semantic commit messages:

  • feat: New feature
  • fix: Bug fix
  • docs: Documentation update
  • style: Code formatting
  • refactor: Code refactoring
  • perf: Performance optimization
  • test: Testing
  • chore: Build/tools

See PR Workflow Documentation


📄 License

This project is licensed under the MIT License. See the LICENSE file for details.


🙏 Acknowledgments

This project is built on the following excellent open source projects:

  • LightRAG - Efficient knowledge graph RAG framework
  • RAG-Anything - Multimodal document parsing
  • MinerU - Powerful PDF parsing tool
  • Docling - Lightweight document parsing
  • FastAPI - Modern Python web framework

Special thanks to all contributors and users for their support! 🎉


📬 Contact Us


⭐ If this project helps you, please give it a Star!

Made with ❤️ by BukeLy

© 2025 RAG API. All rights reserved.

About

Multi-tenant RAG API powered by LightRAG/RAG-Anything. Auto-selects best parser (DeepSeek-OCR/MinerU/Docling) via complexity scoring

Topics

Resources

Stars

64 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages