Skip to content

Repository files navigation

💰 LLMscope - Track Your LLM Costs in Real-Time

Stop guessing what your ChatGPT and Claude API calls cost. Track them in 60 seconds with 3 lines of code.

LicenseDocker ReadyPRs Welcome

LLMscope Dashboard

🎯 Quick Access

After running docker-compose up -d:


🚨 The Problem

LLM API costs are spiraling out of control. You don't know:

  • 💸 How much you're spending on each model
  • 📈 Which providers are most expensive
  • 🔄 If there are cheaper alternatives
  • 📊 Your usage patterns over time

✅ The Solution

LLMscope gives you complete visibility and control over your LLM costs:

  • Real-time cost tracking - See costs as they happen
  • Cost breakdown - By provider, model, and time period
  • Smart recommendations - Get suggestions for cheaper models
  • Usage analytics - Track token usage and patterns
  • Self-hosted - Keep your data private

🚀 Quick Start (60 Seconds)

Step 1: Deploy LLMscope

Using Docker (Recommended):

git clone https://github.com/Blb3D/LLMscope.git
cd LLMscope
docker-compose up -d

Visit http://localhost:8081 - You'll see the dashboard (empty until you track your first API call).

Step 2: Track YOUR First LLM API Call

After making any OpenAI, Anthropic, or other LLM API call, add 3 lines to log the cost:

Example: Track OpenAI GPT-4 Call

importopenaiimportrequests# Your normal OpenAI API callresponse=openai.ChatCompletion.create(
model="gpt-4",
messages=[{"role": "user", "content": "Hello!"}]
)
# Add these 3 lines to track cost:requests.post('http://localhost:8000/api/usage', json={
'provider': 'openai',
'model': 'gpt-4',
'prompt_tokens': response['usage']['prompt_tokens'],
'completion_tokens': response['usage']['completion_tokens']
})

That's it! Refresh the dashboard to see your real costs.

Step 3 (Optional): Generate Demo Data for Testing

Want to test the dashboard before integrating your real API calls?

cd backend
python generate_demo_data.py

This creates 100 sample API calls to preview the dashboard features.


📊 Real-World Integration Examples

OpenAI Integration

Track every OpenAI API call:

importopenaiimportrequestsdeftrack_openai_usage(response):
"""Helper function to track OpenAI costs"""requests.post('http://localhost:8000/api/usage', json={
'provider': 'openai',
'model': response['model'],
'prompt_tokens': response['usage']['prompt_tokens'],
'completion_tokens': response['usage']['completion_tokens']
})
# Use it after any OpenAI call:response=openai.ChatCompletion.create(
model="gpt-4-turbo",
messages=[{"role": "user", "content": "Explain quantum computing"}]
)
track_openai_usage(response)

Anthropic Claude Integration

Track Claude API calls:

importanthropicimportrequestsdeftrack_anthropic_usage(model, response):
"""Helper function to track Anthropic costs"""requests.post('http://localhost:8000/api/usage', json={
'provider': 'anthropic',
'model': model,
'prompt_tokens': response.usage.input_tokens,
'completion_tokens': response.usage.output_tokens
})
# Use it after Claude API calls:client=anthropic.Anthropic(api_key="your-key")
message=client.messages.create(
model="claude-3-sonnet-20240229",
max_tokens=1024,
messages=[{"role": "user", "content": "Hello, Claude!"}]
)
track_anthropic_usage("claude-3-sonnet", message)

LangChain Integration

Automatic tracking with LangChain callback:

fromlangchain.callbacks.baseimportBaseCallbackHandlerfromlangchain.chat_modelsimportChatOpenAIimportrequestsclassLLMscopeCallback(BaseCallbackHandler):
defon_llm_end(self, response, **kwargs):
ifhasattr(response, 'llm_output') andresponse.llm_output:
usage=response.llm_output.get('token_usage', {})
requests.post('http://localhost:8000/api/usage', json={
'provider': 'openai',
'model': response.llm_output.get('model_name', 'gpt-3.5-turbo'),
'prompt_tokens': usage.get('prompt_tokens', 0),
'completion_tokens': usage.get('completion_tokens', 0)
})
# Use with LangChain:llm=ChatOpenAI(callbacks=[LLMscopeCallback()])
result=llm.predict("What is the capital of France?")

Google Gemini Integration

Track Gemini API calls:

importgoogle.generativeaiasgenaiimportrequestsdeftrack_gemini_usage(model_name, response):
"""Helper function to track Google Gemini costs"""requests.post('http://localhost:8000/api/usage', json={
'provider': 'google',
'model': model_name,
'prompt_tokens': response.usage_metadata.prompt_token_count,
'completion_tokens': response.usage_metadata.candidates_token_count
})
# Use it after Gemini calls:genai.configure(api_key="your-key")
model=genai.GenerativeModel('gemini-pro')
response=model.generate_content("Write a poem about AI")
track_gemini_usage('gemini-pro', response)

🔧 Manual Setup (Without Docker)

Backend:

cd backend
pip install -r requirements.txt
# Seed the database with LLM pricing data
python seed_pricing.py
# Start the backend
python app.py

Frontend:

cd frontend
npm install
npm run dev

Visit http://localhost:8081


📊 Features

💰 Real-Time Cost Tracking

Track every LLM API call with automatic cost calculation. Monitor spending across 63+ models from OpenAI, Anthropic, Google, Cohere, Together AI, Mistral, Groq, and more.

  • Instant cost visibility - See exactly what each request costs
  • Provider comparison - Compare costs across different LLM providers
  • Token usage analytics - Track prompt and completion tokens
  • Auto-refresh - Dashboard updates every 5 seconds

💡 Smart Cost Optimization

Cost Recommendations

Get intelligent recommendations for cheaper model alternatives:

  • Cheapest models first - Groq's Llama-3-8B at $0.000065/1K tokens
  • Side-by-side pricing - Compare input/output costs instantly
  • Recent usage history - Track your last 100 API calls
  • Save money automatically - Identify where you're overspending

🔐 Privacy-First & Self-Hosted

  • 100% local - Your data never leaves your infrastructure
  • No external dependencies - Runs entirely on Docker
  • Open source - Audit every line of code

🔌 API Reference

Log API Usage (POST)

Endpoint:POST http://localhost:8000/api/usage

Request Body:

{
"provider": "openai",
"model": "gpt-4",
"prompt_tokens": 100,
"completion_tokens": 50
}

Response:

{
"status": "logged",
"cost_usd": 0.006,
"timestamp": "2025-11-02T10:30:00Z"
}

Get Cost Summary (GET)

Endpoint:GET http://localhost:8000/api/costs/summary

Response:

{
"total_cost": 15.23,
"total_requests": 1250,
"by_provider": {
"openai": 12.45,
"anthropic": 2.78
},
"by_model": {
"gpt-4": 10.20,
"gpt-3.5-turbo": 2.25,
"claude-3-sonnet": 2.78
}
}

Get Model Recommendations (GET)

Endpoint:GET http://localhost:8000/api/recommendations

Returns a list of LLM models sorted by cost (cheapest first) with pricing details.

Interactive API Docs

Visit http://localhost:8000/docs for full interactive API documentation.


⚙️ Configuration

Create a .env file:

DATABASE_PATH=./data/llmscope.dbLLMSCOPE_API_KEY=your-secret-key

🏗️ Architecture

┌─────────────────┐
│ Frontend │ React + Vite
│ (Port 8081) │ Cost dashboard
└────────┬────────┘
│
│ HTTP/REST
│
┌────────▼────────┐
│ Backend API │ FastAPI + SQLite
│ (Port 8000) │ Cost tracking
└─────────────────┘

📦 Tech Stack

Backend:

  • FastAPI
  • SQLite
  • Python 3.9+

Frontend:

  • React
  • Vite
  • Tailwind CSS

Deployment:

  • Docker
  • Docker Compose

🗄️ Database Schema

api_usage

Tracks all API calls with token counts and costs

model_pricing

Stores pricing data for different LLM models

settings

Application configuration


🛠️ Development

# Backend testscd backend
pytest
# Frontend developmentcd frontend
npm run dev
# Production build
docker-compose -f docker-compose.prod.yml up -d

🗺️ Roadmap

Supported Providers: OpenAI, Anthropic, Google, Cohere, Together AI, Mistral, Groq, Ollama, and 60+ models

Coming Soon:

  • Cost alerts and budget thresholds
  • Export to CSV/PDF
  • More provider integrations (request yours in Issues!)

Want to influence the roadmap? Open an issue or start a discussion!


🤝 Contributing

Contributions are welcome! Please feel free to submit a Pull Request.


📄 License

Business Source License 1.1

✅ Free for Non-Commercial Use

  • Self-hosting for personal use, education, and research
  • Modify and redistribute for non-commercial purposes
  • Full source code access - no restrictions on reading the code

💼 Commercial Use Requires License

Commercial use includes:

  • Using LLMscope to monitor production LLM deployments in a business
  • Offering LLMscope as a hosted/managed service to customers
  • Incorporating LLMscope into a commercial product

Need a commercial license? Contact: bbaker@blb3dprinting.com

🔓 Future: Converts to MIT License

On October 29, 2028 (3 years from first publication), this license automatically converts to MIT - making it fully open source forever.


Why BSL? We want LLMscope to be freely available for individuals and small teams, while ensuring companies using it commercially contribute back. This allows us to keep developing new features like SPC analysis, AI copilot, and enhanced reporting.


💬 Support


Built with ❤️ for the LLM community

About

Real-time Statistical Process Control (SPC) monitoring for LLMs

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - BLB3DPrinting/LLMscope: Real-time Statistical Process Control (SPC) monitoring for LLMs · GitHub
Skip to content

Repository files navigation

💰 LLMscope - Track Your LLM Costs in Real-Time

Stop guessing what your ChatGPT and Claude API calls cost. Track them in 60 seconds with 3 lines of code.

LicenseDocker ReadyPRs Welcome

LLMscope Dashboard

🎯 Quick Access

After running docker-compose up -d:


🚨 The Problem

LLM API costs are spiraling out of control. You don't know:

  • 💸 How much you're spending on each model
  • 📈 Which providers are most expensive
  • 🔄 If there are cheaper alternatives
  • 📊 Your usage patterns over time

✅ The Solution

LLMscope gives you complete visibility and control over your LLM costs:

  • Real-time cost tracking - See costs as they happen
  • Cost breakdown - By provider, model, and time period
  • Smart recommendations - Get suggestions for cheaper models
  • Usage analytics - Track token usage and patterns
  • Self-hosted - Keep your data private

🚀 Quick Start (60 Seconds)

Step 1: Deploy LLMscope

Using Docker (Recommended):

git clone https://github.com/Blb3D/LLMscope.git
cd LLMscope
docker-compose up -d

Visit http://localhost:8081 - You'll see the dashboard (empty until you track your first API call).

Step 2: Track YOUR First LLM API Call

After making any OpenAI, Anthropic, or other LLM API call, add 3 lines to log the cost:

Example: Track OpenAI GPT-4 Call

importopenaiimportrequests# Your normal OpenAI API callresponse=openai.ChatCompletion.create(
model="gpt-4",
messages=[{"role": "user", "content": "Hello!"}]
)
# Add these 3 lines to track cost:requests.post('http://localhost:8000/api/usage', json={
'provider': 'openai',
'model': 'gpt-4',
'prompt_tokens': response['usage']['prompt_tokens'],
'completion_tokens': response['usage']['completion_tokens']
})

That's it! Refresh the dashboard to see your real costs.

Step 3 (Optional): Generate Demo Data for Testing

Want to test the dashboard before integrating your real API calls?

cd backend
python generate_demo_data.py

This creates 100 sample API calls to preview the dashboard features.


📊 Real-World Integration Examples

OpenAI Integration

Track every OpenAI API call:

importopenaiimportrequestsdeftrack_openai_usage(response):
"""Helper function to track OpenAI costs"""requests.post('http://localhost:8000/api/usage', json={
'provider': 'openai',
'model': response['model'],
'prompt_tokens': response['usage']['prompt_tokens'],
'completion_tokens': response['usage']['completion_tokens']
})
# Use it after any OpenAI call:response=openai.ChatCompletion.create(
model="gpt-4-turbo",
messages=[{"role": "user", "content": "Explain quantum computing"}]
)
track_openai_usage(response)

Anthropic Claude Integration

Track Claude API calls:

importanthropicimportrequestsdeftrack_anthropic_usage(model, response):
"""Helper function to track Anthropic costs"""requests.post('http://localhost:8000/api/usage', json={
'provider': 'anthropic',
'model': model,
'prompt_tokens': response.usage.input_tokens,
'completion_tokens': response.usage.output_tokens
})
# Use it after Claude API calls:client=anthropic.Anthropic(api_key="your-key")
message=client.messages.create(
model="claude-3-sonnet-20240229",
max_tokens=1024,
messages=[{"role": "user", "content": "Hello, Claude!"}]
)
track_anthropic_usage("claude-3-sonnet", message)

LangChain Integration

Automatic tracking with LangChain callback:

fromlangchain.callbacks.baseimportBaseCallbackHandlerfromlangchain.chat_modelsimportChatOpenAIimportrequestsclassLLMscopeCallback(BaseCallbackHandler):
defon_llm_end(self, response, **kwargs):
ifhasattr(response, 'llm_output') andresponse.llm_output:
usage=response.llm_output.get('token_usage', {})
requests.post('http://localhost:8000/api/usage', json={
'provider': 'openai',
'model': response.llm_output.get('model_name', 'gpt-3.5-turbo'),
'prompt_tokens': usage.get('prompt_tokens', 0),
'completion_tokens': usage.get('completion_tokens', 0)
})
# Use with LangChain:llm=ChatOpenAI(callbacks=[LLMscopeCallback()])
result=llm.predict("What is the capital of France?")

Google Gemini Integration

Track Gemini API calls:

importgoogle.generativeaiasgenaiimportrequestsdeftrack_gemini_usage(model_name, response):
"""Helper function to track Google Gemini costs"""requests.post('http://localhost:8000/api/usage', json={
'provider': 'google',
'model': model_name,
'prompt_tokens': response.usage_metadata.prompt_token_count,
'completion_tokens': response.usage_metadata.candidates_token_count
})
# Use it after Gemini calls:genai.configure(api_key="your-key")
model=genai.GenerativeModel('gemini-pro')
response=model.generate_content("Write a poem about AI")
track_gemini_usage('gemini-pro', response)

🔧 Manual Setup (Without Docker)

Backend:

cd backend
pip install -r requirements.txt
# Seed the database with LLM pricing data
python seed_pricing.py
# Start the backend
python app.py

Frontend:

cd frontend
npm install
npm run dev

Visit http://localhost:8081


📊 Features

💰 Real-Time Cost Tracking

Track every LLM API call with automatic cost calculation. Monitor spending across 63+ models from OpenAI, Anthropic, Google, Cohere, Together AI, Mistral, Groq, and more.

  • Instant cost visibility - See exactly what each request costs
  • Provider comparison - Compare costs across different LLM providers
  • Token usage analytics - Track prompt and completion tokens
  • Auto-refresh - Dashboard updates every 5 seconds

💡 Smart Cost Optimization

Cost Recommendations

Get intelligent recommendations for cheaper model alternatives:

  • Cheapest models first - Groq's Llama-3-8B at $0.000065/1K tokens
  • Side-by-side pricing - Compare input/output costs instantly
  • Recent usage history - Track your last 100 API calls
  • Save money automatically - Identify where you're overspending

🔐 Privacy-First & Self-Hosted

  • 100% local - Your data never leaves your infrastructure
  • No external dependencies - Runs entirely on Docker
  • Open source - Audit every line of code

🔌 API Reference

Log API Usage (POST)

Endpoint:POST http://localhost:8000/api/usage

Request Body:

{
"provider": "openai",
"model": "gpt-4",
"prompt_tokens": 100,
"completion_tokens": 50
}

Response:

{
"status": "logged",
"cost_usd": 0.006,
"timestamp": "2025-11-02T10:30:00Z"
}

Get Cost Summary (GET)

Endpoint:GET http://localhost:8000/api/costs/summary

Response:

{
"total_cost": 15.23,
"total_requests": 1250,
"by_provider": {
"openai": 12.45,
"anthropic": 2.78
},
"by_model": {
"gpt-4": 10.20,
"gpt-3.5-turbo": 2.25,
"claude-3-sonnet": 2.78
}
}

Get Model Recommendations (GET)

Endpoint:GET http://localhost:8000/api/recommendations

Returns a list of LLM models sorted by cost (cheapest first) with pricing details.

Interactive API Docs

Visit http://localhost:8000/docs for full interactive API documentation.


⚙️ Configuration

Create a .env file:

DATABASE_PATH=./data/llmscope.dbLLMSCOPE_API_KEY=your-secret-key

🏗️ Architecture

┌─────────────────┐
│ Frontend │ React + Vite
│ (Port 8081) │ Cost dashboard
└────────┬────────┘
│
│ HTTP/REST
│
┌────────▼────────┐
│ Backend API │ FastAPI + SQLite
│ (Port 8000) │ Cost tracking
└─────────────────┘

📦 Tech Stack

Backend:

  • FastAPI
  • SQLite
  • Python 3.9+

Frontend:

  • React
  • Vite
  • Tailwind CSS

Deployment:

  • Docker
  • Docker Compose

🗄️ Database Schema

api_usage

Tracks all API calls with token counts and costs

model_pricing

Stores pricing data for different LLM models

settings

Application configuration


🛠️ Development

# Backend testscd backend
pytest
# Frontend developmentcd frontend
npm run dev
# Production build
docker-compose -f docker-compose.prod.yml up -d

🗺️ Roadmap

Supported Providers: OpenAI, Anthropic, Google, Cohere, Together AI, Mistral, Groq, Ollama, and 60+ models

Coming Soon:

  • Cost alerts and budget thresholds
  • Export to CSV/PDF
  • More provider integrations (request yours in Issues!)

Want to influence the roadmap? Open an issue or start a discussion!


🤝 Contributing

Contributions are welcome! Please feel free to submit a Pull Request.


📄 License

Business Source License 1.1

✅ Free for Non-Commercial Use

  • Self-hosting for personal use, education, and research
  • Modify and redistribute for non-commercial purposes
  • Full source code access - no restrictions on reading the code

💼 Commercial Use Requires License

Commercial use includes:

  • Using LLMscope to monitor production LLM deployments in a business
  • Offering LLMscope as a hosted/managed service to customers
  • Incorporating LLMscope into a commercial product

Need a commercial license? Contact: bbaker@blb3dprinting.com

🔓 Future: Converts to MIT License

On October 29, 2028 (3 years from first publication), this license automatically converts to MIT - making it fully open source forever.


Why BSL? We want LLMscope to be freely available for individuals and small teams, while ensuring companies using it commercially contribute back. This allows us to keep developing new features like SPC analysis, AI copilot, and enhanced reporting.


💬 Support


Built with ❤️ for the LLM community

About

Real-time Statistical Process Control (SPC) monitoring for LLMs

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - BLB3DPrinting/LLMscope: Real-time Statistical Process Control (SPC) monitoring for LLMs · GitHub
Skip to content

Repository files navigation

💰 LLMscope - Track Your LLM Costs in Real-Time

Stop guessing what your ChatGPT and Claude API calls cost. Track them in 60 seconds with 3 lines of code.

LicenseDocker ReadyPRs Welcome

LLMscope Dashboard

🎯 Quick Access

After running docker-compose up -d:


🚨 The Problem

LLM API costs are spiraling out of control. You don't know:

  • 💸 How much you're spending on each model
  • 📈 Which providers are most expensive
  • 🔄 If there are cheaper alternatives
  • 📊 Your usage patterns over time

✅ The Solution

LLMscope gives you complete visibility and control over your LLM costs:

  • Real-time cost tracking - See costs as they happen
  • Cost breakdown - By provider, model, and time period
  • Smart recommendations - Get suggestions for cheaper models
  • Usage analytics - Track token usage and patterns
  • Self-hosted - Keep your data private

🚀 Quick Start (60 Seconds)

Step 1: Deploy LLMscope

Using Docker (Recommended):

git clone https://github.com/Blb3D/LLMscope.git
cd LLMscope
docker-compose up -d

Visit http://localhost:8081 - You'll see the dashboard (empty until you track your first API call).

Step 2: Track YOUR First LLM API Call

After making any OpenAI, Anthropic, or other LLM API call, add 3 lines to log the cost:

Example: Track OpenAI GPT-4 Call

importopenaiimportrequests# Your normal OpenAI API callresponse=openai.ChatCompletion.create(
model="gpt-4",
messages=[{"role": "user", "content": "Hello!"}]
)
# Add these 3 lines to track cost:requests.post('http://localhost:8000/api/usage', json={
'provider': 'openai',
'model': 'gpt-4',
'prompt_tokens': response['usage']['prompt_tokens'],
'completion_tokens': response['usage']['completion_tokens']
})

That's it! Refresh the dashboard to see your real costs.

Step 3 (Optional): Generate Demo Data for Testing

Want to test the dashboard before integrating your real API calls?

cd backend
python generate_demo_data.py

This creates 100 sample API calls to preview the dashboard features.


📊 Real-World Integration Examples

OpenAI Integration

Track every OpenAI API call:

importopenaiimportrequestsdeftrack_openai_usage(response):
"""Helper function to track OpenAI costs"""requests.post('http://localhost:8000/api/usage', json={
'provider': 'openai',
'model': response['model'],
'prompt_tokens': response['usage']['prompt_tokens'],
'completion_tokens': response['usage']['completion_tokens']
})
# Use it after any OpenAI call:response=openai.ChatCompletion.create(
model="gpt-4-turbo",
messages=[{"role": "user", "content": "Explain quantum computing"}]
)
track_openai_usage(response)

Anthropic Claude Integration

Track Claude API calls:

importanthropicimportrequestsdeftrack_anthropic_usage(model, response):
"""Helper function to track Anthropic costs"""requests.post('http://localhost:8000/api/usage', json={
'provider': 'anthropic',
'model': model,
'prompt_tokens': response.usage.input_tokens,
'completion_tokens': response.usage.output_tokens
})
# Use it after Claude API calls:client=anthropic.Anthropic(api_key="your-key")
message=client.messages.create(
model="claude-3-sonnet-20240229",
max_tokens=1024,
messages=[{"role": "user", "content": "Hello, Claude!"}]
)
track_anthropic_usage("claude-3-sonnet", message)

LangChain Integration

Automatic tracking with LangChain callback:

fromlangchain.callbacks.baseimportBaseCallbackHandlerfromlangchain.chat_modelsimportChatOpenAIimportrequestsclassLLMscopeCallback(BaseCallbackHandler):
defon_llm_end(self, response, **kwargs):
ifhasattr(response, 'llm_output') andresponse.llm_output:
usage=response.llm_output.get('token_usage', {})
requests.post('http://localhost:8000/api/usage', json={
'provider': 'openai',
'model': response.llm_output.get('model_name', 'gpt-3.5-turbo'),
'prompt_tokens': usage.get('prompt_tokens', 0),
'completion_tokens': usage.get('completion_tokens', 0)
})
# Use with LangChain:llm=ChatOpenAI(callbacks=[LLMscopeCallback()])
result=llm.predict("What is the capital of France?")

Google Gemini Integration

Track Gemini API calls:

importgoogle.generativeaiasgenaiimportrequestsdeftrack_gemini_usage(model_name, response):
"""Helper function to track Google Gemini costs"""requests.post('http://localhost:8000/api/usage', json={
'provider': 'google',
'model': model_name,
'prompt_tokens': response.usage_metadata.prompt_token_count,
'completion_tokens': response.usage_metadata.candidates_token_count
})
# Use it after Gemini calls:genai.configure(api_key="your-key")
model=genai.GenerativeModel('gemini-pro')
response=model.generate_content("Write a poem about AI")
track_gemini_usage('gemini-pro', response)

🔧 Manual Setup (Without Docker)

Backend:

cd backend
pip install -r requirements.txt
# Seed the database with LLM pricing data
python seed_pricing.py
# Start the backend
python app.py

Frontend:

cd frontend
npm install
npm run dev

Visit http://localhost:8081


📊 Features

💰 Real-Time Cost Tracking

Track every LLM API call with automatic cost calculation. Monitor spending across 63+ models from OpenAI, Anthropic, Google, Cohere, Together AI, Mistral, Groq, and more.

  • Instant cost visibility - See exactly what each request costs
  • Provider comparison - Compare costs across different LLM providers
  • Token usage analytics - Track prompt and completion tokens
  • Auto-refresh - Dashboard updates every 5 seconds

💡 Smart Cost Optimization

Cost Recommendations

Get intelligent recommendations for cheaper model alternatives:

  • Cheapest models first - Groq's Llama-3-8B at $0.000065/1K tokens
  • Side-by-side pricing - Compare input/output costs instantly
  • Recent usage history - Track your last 100 API calls
  • Save money automatically - Identify where you're overspending

🔐 Privacy-First & Self-Hosted

  • 100% local - Your data never leaves your infrastructure
  • No external dependencies - Runs entirely on Docker
  • Open source - Audit every line of code

🔌 API Reference

Log API Usage (POST)

Endpoint:POST http://localhost:8000/api/usage

Request Body:

{
"provider": "openai",
"model": "gpt-4",
"prompt_tokens": 100,
"completion_tokens": 50
}

Response:

{
"status": "logged",
"cost_usd": 0.006,
"timestamp": "2025-11-02T10:30:00Z"
}

Get Cost Summary (GET)

Endpoint:GET http://localhost:8000/api/costs/summary

Response:

{
"total_cost": 15.23,
"total_requests": 1250,
"by_provider": {
"openai": 12.45,
"anthropic": 2.78
},
"by_model": {
"gpt-4": 10.20,
"gpt-3.5-turbo": 2.25,
"claude-3-sonnet": 2.78
}
}

Get Model Recommendations (GET)

Endpoint:GET http://localhost:8000/api/recommendations

Returns a list of LLM models sorted by cost (cheapest first) with pricing details.

Interactive API Docs

Visit http://localhost:8000/docs for full interactive API documentation.


⚙️ Configuration

Create a .env file:

DATABASE_PATH=./data/llmscope.dbLLMSCOPE_API_KEY=your-secret-key

🏗️ Architecture

┌─────────────────┐
│ Frontend │ React + Vite
│ (Port 8081) │ Cost dashboard
└────────┬────────┘
│
│ HTTP/REST
│
┌────────▼────────┐
│ Backend API │ FastAPI + SQLite
│ (Port 8000) │ Cost tracking
└─────────────────┘

📦 Tech Stack

Backend:

  • FastAPI
  • SQLite
  • Python 3.9+

Frontend:

  • React
  • Vite
  • Tailwind CSS

Deployment:

  • Docker
  • Docker Compose

🗄️ Database Schema

api_usage

Tracks all API calls with token counts and costs

model_pricing

Stores pricing data for different LLM models

settings

Application configuration


🛠️ Development

# Backend testscd backend
pytest
# Frontend developmentcd frontend
npm run dev
# Production build
docker-compose -f docker-compose.prod.yml up -d

🗺️ Roadmap

Supported Providers: OpenAI, Anthropic, Google, Cohere, Together AI, Mistral, Groq, Ollama, and 60+ models

Coming Soon:

  • Cost alerts and budget thresholds
  • Export to CSV/PDF
  • More provider integrations (request yours in Issues!)

Want to influence the roadmap? Open an issue or start a discussion!


🤝 Contributing

Contributions are welcome! Please feel free to submit a Pull Request.


📄 License

Business Source License 1.1

✅ Free for Non-Commercial Use

  • Self-hosting for personal use, education, and research
  • Modify and redistribute for non-commercial purposes
  • Full source code access - no restrictions on reading the code

💼 Commercial Use Requires License

Commercial use includes:

  • Using LLMscope to monitor production LLM deployments in a business
  • Offering LLMscope as a hosted/managed service to customers
  • Incorporating LLMscope into a commercial product

Need a commercial license? Contact: bbaker@blb3dprinting.com

🔓 Future: Converts to MIT License

On October 29, 2028 (3 years from first publication), this license automatically converts to MIT - making it fully open source forever.


Why BSL? We want LLMscope to be freely available for individuals and small teams, while ensuring companies using it commercially contribute back. This allows us to keep developing new features like SPC analysis, AI copilot, and enhanced reporting.


💬 Support


Built with ❤️ for the LLM community

About

Real-time Statistical Process Control (SPC) monitoring for LLMs

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - BLB3DPrinting/LLMscope: Real-time Statistical Process Control (SPC) monitoring for LLMs · GitHub
Skip to content

Repository files navigation

💰 LLMscope - Track Your LLM Costs in Real-Time

Stop guessing what your ChatGPT and Claude API calls cost. Track them in 60 seconds with 3 lines of code.

LicenseDocker ReadyPRs Welcome

LLMscope Dashboard

🎯 Quick Access

After running docker-compose up -d:


🚨 The Problem

LLM API costs are spiraling out of control. You don't know:

  • 💸 How much you're spending on each model
  • 📈 Which providers are most expensive
  • 🔄 If there are cheaper alternatives
  • 📊 Your usage patterns over time

✅ The Solution

LLMscope gives you complete visibility and control over your LLM costs:

  • Real-time cost tracking - See costs as they happen
  • Cost breakdown - By provider, model, and time period
  • Smart recommendations - Get suggestions for cheaper models
  • Usage analytics - Track token usage and patterns
  • Self-hosted - Keep your data private

🚀 Quick Start (60 Seconds)

Step 1: Deploy LLMscope

Using Docker (Recommended):

git clone https://github.com/Blb3D/LLMscope.git
cd LLMscope
docker-compose up -d

Visit http://localhost:8081 - You'll see the dashboard (empty until you track your first API call).

Step 2: Track YOUR First LLM API Call

After making any OpenAI, Anthropic, or other LLM API call, add 3 lines to log the cost:

Example: Track OpenAI GPT-4 Call

importopenaiimportrequests# Your normal OpenAI API callresponse=openai.ChatCompletion.create(
model="gpt-4",
messages=[{"role": "user", "content": "Hello!"}]
)
# Add these 3 lines to track cost:requests.post('http://localhost:8000/api/usage', json={
'provider': 'openai',
'model': 'gpt-4',
'prompt_tokens': response['usage']['prompt_tokens'],
'completion_tokens': response['usage']['completion_tokens']
})

That's it! Refresh the dashboard to see your real costs.

Step 3 (Optional): Generate Demo Data for Testing

Want to test the dashboard before integrating your real API calls?

cd backend
python generate_demo_data.py

This creates 100 sample API calls to preview the dashboard features.


📊 Real-World Integration Examples

OpenAI Integration

Track every OpenAI API call:

importopenaiimportrequestsdeftrack_openai_usage(response):
"""Helper function to track OpenAI costs"""requests.post('http://localhost:8000/api/usage', json={
'provider': 'openai',
'model': response['model'],
'prompt_tokens': response['usage']['prompt_tokens'],
'completion_tokens': response['usage']['completion_tokens']
})
# Use it after any OpenAI call:response=openai.ChatCompletion.create(
model="gpt-4-turbo",
messages=[{"role": "user", "content": "Explain quantum computing"}]
)
track_openai_usage(response)

Anthropic Claude Integration

Track Claude API calls:

importanthropicimportrequestsdeftrack_anthropic_usage(model, response):
"""Helper function to track Anthropic costs"""requests.post('http://localhost:8000/api/usage', json={
'provider': 'anthropic',
'model': model,
'prompt_tokens': response.usage.input_tokens,
'completion_tokens': response.usage.output_tokens
})
# Use it after Claude API calls:client=anthropic.Anthropic(api_key="your-key")
message=client.messages.create(
model="claude-3-sonnet-20240229",
max_tokens=1024,
messages=[{"role": "user", "content": "Hello, Claude!"}]
)
track_anthropic_usage("claude-3-sonnet", message)

LangChain Integration

Automatic tracking with LangChain callback:

fromlangchain.callbacks.baseimportBaseCallbackHandlerfromlangchain.chat_modelsimportChatOpenAIimportrequestsclassLLMscopeCallback(BaseCallbackHandler):
defon_llm_end(self, response, **kwargs):
ifhasattr(response, 'llm_output') andresponse.llm_output:
usage=response.llm_output.get('token_usage', {})
requests.post('http://localhost:8000/api/usage', json={
'provider': 'openai',
'model': response.llm_output.get('model_name', 'gpt-3.5-turbo'),
'prompt_tokens': usage.get('prompt_tokens', 0),
'completion_tokens': usage.get('completion_tokens', 0)
})
# Use with LangChain:llm=ChatOpenAI(callbacks=[LLMscopeCallback()])
result=llm.predict("What is the capital of France?")

Google Gemini Integration

Track Gemini API calls:

importgoogle.generativeaiasgenaiimportrequestsdeftrack_gemini_usage(model_name, response):
"""Helper function to track Google Gemini costs"""requests.post('http://localhost:8000/api/usage', json={
'provider': 'google',
'model': model_name,
'prompt_tokens': response.usage_metadata.prompt_token_count,
'completion_tokens': response.usage_metadata.candidates_token_count
})
# Use it after Gemini calls:genai.configure(api_key="your-key")
model=genai.GenerativeModel('gemini-pro')
response=model.generate_content("Write a poem about AI")
track_gemini_usage('gemini-pro', response)

🔧 Manual Setup (Without Docker)

Backend:

cd backend
pip install -r requirements.txt
# Seed the database with LLM pricing data
python seed_pricing.py
# Start the backend
python app.py

Frontend:

cd frontend
npm install
npm run dev

Visit http://localhost:8081


📊 Features

💰 Real-Time Cost Tracking

Track every LLM API call with automatic cost calculation. Monitor spending across 63+ models from OpenAI, Anthropic, Google, Cohere, Together AI, Mistral, Groq, and more.

  • Instant cost visibility - See exactly what each request costs
  • Provider comparison - Compare costs across different LLM providers
  • Token usage analytics - Track prompt and completion tokens
  • Auto-refresh - Dashboard updates every 5 seconds

💡 Smart Cost Optimization

Cost Recommendations

Get intelligent recommendations for cheaper model alternatives:

  • Cheapest models first - Groq's Llama-3-8B at $0.000065/1K tokens
  • Side-by-side pricing - Compare input/output costs instantly
  • Recent usage history - Track your last 100 API calls
  • Save money automatically - Identify where you're overspending

🔐 Privacy-First & Self-Hosted

  • 100% local - Your data never leaves your infrastructure
  • No external dependencies - Runs entirely on Docker
  • Open source - Audit every line of code

🔌 API Reference

Log API Usage (POST)

Endpoint:POST http://localhost:8000/api/usage

Request Body:

{
"provider": "openai",
"model": "gpt-4",
"prompt_tokens": 100,
"completion_tokens": 50
}

Response:

{
"status": "logged",
"cost_usd": 0.006,
"timestamp": "2025-11-02T10:30:00Z"
}

Get Cost Summary (GET)

Endpoint:GET http://localhost:8000/api/costs/summary

Response:

{
"total_cost": 15.23,
"total_requests": 1250,
"by_provider": {
"openai": 12.45,
"anthropic": 2.78
},
"by_model": {
"gpt-4": 10.20,
"gpt-3.5-turbo": 2.25,
"claude-3-sonnet": 2.78
}
}

Get Model Recommendations (GET)

Endpoint:GET http://localhost:8000/api/recommendations

Returns a list of LLM models sorted by cost (cheapest first) with pricing details.

Interactive API Docs

Visit http://localhost:8000/docs for full interactive API documentation.


⚙️ Configuration

Create a .env file:

DATABASE_PATH=./data/llmscope.dbLLMSCOPE_API_KEY=your-secret-key

🏗️ Architecture

┌─────────────────┐
│ Frontend │ React + Vite
│ (Port 8081) │ Cost dashboard
└────────┬────────┘
│
│ HTTP/REST
│
┌────────▼────────┐
│ Backend API │ FastAPI + SQLite
│ (Port 8000) │ Cost tracking
└─────────────────┘

📦 Tech Stack

Backend:

  • FastAPI
  • SQLite
  • Python 3.9+

Frontend:

  • React
  • Vite
  • Tailwind CSS

Deployment:

  • Docker
  • Docker Compose

🗄️ Database Schema

api_usage

Tracks all API calls with token counts and costs

model_pricing

Stores pricing data for different LLM models

settings

Application configuration


🛠️ Development

# Backend testscd backend
pytest
# Frontend developmentcd frontend
npm run dev
# Production build
docker-compose -f docker-compose.prod.yml up -d

🗺️ Roadmap

Supported Providers: OpenAI, Anthropic, Google, Cohere, Together AI, Mistral, Groq, Ollama, and 60+ models

Coming Soon:

  • Cost alerts and budget thresholds
  • Export to CSV/PDF
  • More provider integrations (request yours in Issues!)

Want to influence the roadmap? Open an issue or start a discussion!


🤝 Contributing

Contributions are welcome! Please feel free to submit a Pull Request.


📄 License

Business Source License 1.1

✅ Free for Non-Commercial Use

  • Self-hosting for personal use, education, and research
  • Modify and redistribute for non-commercial purposes
  • Full source code access - no restrictions on reading the code

💼 Commercial Use Requires License

Commercial use includes:

  • Using LLMscope to monitor production LLM deployments in a business
  • Offering LLMscope as a hosted/managed service to customers
  • Incorporating LLMscope into a commercial product

Need a commercial license? Contact: bbaker@blb3dprinting.com

🔓 Future: Converts to MIT License

On October 29, 2028 (3 years from first publication), this license automatically converts to MIT - making it fully open source forever.


Why BSL? We want LLMscope to be freely available for individuals and small teams, while ensuring companies using it commercially contribute back. This allows us to keep developing new features like SPC analysis, AI copilot, and enhanced reporting.


💬 Support


Built with ❤️ for the LLM community

About

Real-time Statistical Process Control (SPC) monitoring for LLMs

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - BLB3DPrinting/LLMscope: Real-time Statistical Process Control (SPC) monitoring for LLMs · GitHub
Skip to content

Repository files navigation

💰 LLMscope - Track Your LLM Costs in Real-Time

Stop guessing what your ChatGPT and Claude API calls cost. Track them in 60 seconds with 3 lines of code.

LicenseDocker ReadyPRs Welcome

LLMscope Dashboard

🎯 Quick Access

After running docker-compose up -d:


🚨 The Problem

LLM API costs are spiraling out of control. You don't know:

  • 💸 How much you're spending on each model
  • 📈 Which providers are most expensive
  • 🔄 If there are cheaper alternatives
  • 📊 Your usage patterns over time

✅ The Solution

LLMscope gives you complete visibility and control over your LLM costs:

  • Real-time cost tracking - See costs as they happen
  • Cost breakdown - By provider, model, and time period
  • Smart recommendations - Get suggestions for cheaper models
  • Usage analytics - Track token usage and patterns
  • Self-hosted - Keep your data private

🚀 Quick Start (60 Seconds)

Step 1: Deploy LLMscope

Using Docker (Recommended):

git clone https://github.com/Blb3D/LLMscope.git
cd LLMscope
docker-compose up -d

Visit http://localhost:8081 - You'll see the dashboard (empty until you track your first API call).

Step 2: Track YOUR First LLM API Call

After making any OpenAI, Anthropic, or other LLM API call, add 3 lines to log the cost:

Example: Track OpenAI GPT-4 Call

importopenaiimportrequests# Your normal OpenAI API callresponse=openai.ChatCompletion.create(
model="gpt-4",
messages=[{"role": "user", "content": "Hello!"}]
)
# Add these 3 lines to track cost:requests.post('http://localhost:8000/api/usage', json={
'provider': 'openai',
'model': 'gpt-4',
'prompt_tokens': response['usage']['prompt_tokens'],
'completion_tokens': response['usage']['completion_tokens']
})

That's it! Refresh the dashboard to see your real costs.

Step 3 (Optional): Generate Demo Data for Testing

Want to test the dashboard before integrating your real API calls?

cd backend
python generate_demo_data.py

This creates 100 sample API calls to preview the dashboard features.


📊 Real-World Integration Examples

OpenAI Integration

Track every OpenAI API call:

importopenaiimportrequestsdeftrack_openai_usage(response):
"""Helper function to track OpenAI costs"""requests.post('http://localhost:8000/api/usage', json={
'provider': 'openai',
'model': response['model'],
'prompt_tokens': response['usage']['prompt_tokens'],
'completion_tokens': response['usage']['completion_tokens']
})
# Use it after any OpenAI call:response=openai.ChatCompletion.create(
model="gpt-4-turbo",
messages=[{"role": "user", "content": "Explain quantum computing"}]
)
track_openai_usage(response)

Anthropic Claude Integration

Track Claude API calls:

importanthropicimportrequestsdeftrack_anthropic_usage(model, response):
"""Helper function to track Anthropic costs"""requests.post('http://localhost:8000/api/usage', json={
'provider': 'anthropic',
'model': model,
'prompt_tokens': response.usage.input_tokens,
'completion_tokens': response.usage.output_tokens
})
# Use it after Claude API calls:client=anthropic.Anthropic(api_key="your-key")
message=client.messages.create(
model="claude-3-sonnet-20240229",
max_tokens=1024,
messages=[{"role": "user", "content": "Hello, Claude!"}]
)
track_anthropic_usage("claude-3-sonnet", message)

LangChain Integration

Automatic tracking with LangChain callback:

fromlangchain.callbacks.baseimportBaseCallbackHandlerfromlangchain.chat_modelsimportChatOpenAIimportrequestsclassLLMscopeCallback(BaseCallbackHandler):
defon_llm_end(self, response, **kwargs):
ifhasattr(response, 'llm_output') andresponse.llm_output:
usage=response.llm_output.get('token_usage', {})
requests.post('http://localhost:8000/api/usage', json={
'provider': 'openai',
'model': response.llm_output.get('model_name', 'gpt-3.5-turbo'),
'prompt_tokens': usage.get('prompt_tokens', 0),
'completion_tokens': usage.get('completion_tokens', 0)
})
# Use with LangChain:llm=ChatOpenAI(callbacks=[LLMscopeCallback()])
result=llm.predict("What is the capital of France?")

Google Gemini Integration

Track Gemini API calls:

importgoogle.generativeaiasgenaiimportrequestsdeftrack_gemini_usage(model_name, response):
"""Helper function to track Google Gemini costs"""requests.post('http://localhost:8000/api/usage', json={
'provider': 'google',
'model': model_name,
'prompt_tokens': response.usage_metadata.prompt_token_count,
'completion_tokens': response.usage_metadata.candidates_token_count
})
# Use it after Gemini calls:genai.configure(api_key="your-key")
model=genai.GenerativeModel('gemini-pro')
response=model.generate_content("Write a poem about AI")
track_gemini_usage('gemini-pro', response)

🔧 Manual Setup (Without Docker)

Backend:

cd backend
pip install -r requirements.txt
# Seed the database with LLM pricing data
python seed_pricing.py
# Start the backend
python app.py

Frontend:

cd frontend
npm install
npm run dev

Visit http://localhost:8081


📊 Features

💰 Real-Time Cost Tracking

Track every LLM API call with automatic cost calculation. Monitor spending across 63+ models from OpenAI, Anthropic, Google, Cohere, Together AI, Mistral, Groq, and more.

  • Instant cost visibility - See exactly what each request costs
  • Provider comparison - Compare costs across different LLM providers
  • Token usage analytics - Track prompt and completion tokens
  • Auto-refresh - Dashboard updates every 5 seconds

💡 Smart Cost Optimization

Cost Recommendations

Get intelligent recommendations for cheaper model alternatives:

  • Cheapest models first - Groq's Llama-3-8B at $0.000065/1K tokens
  • Side-by-side pricing - Compare input/output costs instantly
  • Recent usage history - Track your last 100 API calls
  • Save money automatically - Identify where you're overspending

🔐 Privacy-First & Self-Hosted

  • 100% local - Your data never leaves your infrastructure
  • No external dependencies - Runs entirely on Docker
  • Open source - Audit every line of code

🔌 API Reference

Log API Usage (POST)

Endpoint:POST http://localhost:8000/api/usage

Request Body:

{
"provider": "openai",
"model": "gpt-4",
"prompt_tokens": 100,
"completion_tokens": 50
}

Response:

{
"status": "logged",
"cost_usd": 0.006,
"timestamp": "2025-11-02T10:30:00Z"
}

Get Cost Summary (GET)

Endpoint:GET http://localhost:8000/api/costs/summary

Response:

{
"total_cost": 15.23,
"total_requests": 1250,
"by_provider": {
"openai": 12.45,
"anthropic": 2.78
},
"by_model": {
"gpt-4": 10.20,
"gpt-3.5-turbo": 2.25,
"claude-3-sonnet": 2.78
}
}

Get Model Recommendations (GET)

Endpoint:GET http://localhost:8000/api/recommendations

Returns a list of LLM models sorted by cost (cheapest first) with pricing details.

Interactive API Docs

Visit http://localhost:8000/docs for full interactive API documentation.


⚙️ Configuration

Create a .env file:

DATABASE_PATH=./data/llmscope.dbLLMSCOPE_API_KEY=your-secret-key

🏗️ Architecture

┌─────────────────┐
│ Frontend │ React + Vite
│ (Port 8081) │ Cost dashboard
└────────┬────────┘
│
│ HTTP/REST
│
┌────────▼────────┐
│ Backend API │ FastAPI + SQLite
│ (Port 8000) │ Cost tracking
└─────────────────┘

📦 Tech Stack

Backend:

  • FastAPI
  • SQLite
  • Python 3.9+

Frontend:

  • React
  • Vite
  • Tailwind CSS

Deployment:

  • Docker
  • Docker Compose

🗄️ Database Schema

api_usage

Tracks all API calls with token counts and costs

model_pricing

Stores pricing data for different LLM models

settings

Application configuration


🛠️ Development

# Backend testscd backend
pytest
# Frontend developmentcd frontend
npm run dev
# Production build
docker-compose -f docker-compose.prod.yml up -d

🗺️ Roadmap

Supported Providers: OpenAI, Anthropic, Google, Cohere, Together AI, Mistral, Groq, Ollama, and 60+ models

Coming Soon:

  • Cost alerts and budget thresholds
  • Export to CSV/PDF
  • More provider integrations (request yours in Issues!)

Want to influence the roadmap? Open an issue or start a discussion!


🤝 Contributing

Contributions are welcome! Please feel free to submit a Pull Request.


📄 License

Business Source License 1.1

✅ Free for Non-Commercial Use

  • Self-hosting for personal use, education, and research
  • Modify and redistribute for non-commercial purposes
  • Full source code access - no restrictions on reading the code

💼 Commercial Use Requires License

Commercial use includes:

  • Using LLMscope to monitor production LLM deployments in a business
  • Offering LLMscope as a hosted/managed service to customers
  • Incorporating LLMscope into a commercial product

Need a commercial license? Contact: bbaker@blb3dprinting.com

🔓 Future: Converts to MIT License

On October 29, 2028 (3 years from first publication), this license automatically converts to MIT - making it fully open source forever.


Why BSL? We want LLMscope to be freely available for individuals and small teams, while ensuring companies using it commercially contribute back. This allows us to keep developing new features like SPC analysis, AI copilot, and enhanced reporting.


💬 Support


Built with ❤️ for the LLM community

About

Real-time Statistical Process Control (SPC) monitoring for LLMs

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - BLB3DPrinting/LLMscope: Real-time Statistical Process Control (SPC) monitoring for LLMs · GitHub
Skip to content

Repository files navigation

💰 LLMscope - Track Your LLM Costs in Real-Time

Stop guessing what your ChatGPT and Claude API calls cost. Track them in 60 seconds with 3 lines of code.

LicenseDocker ReadyPRs Welcome

LLMscope Dashboard

🎯 Quick Access

After running docker-compose up -d:


🚨 The Problem

LLM API costs are spiraling out of control. You don't know:

  • 💸 How much you're spending on each model
  • 📈 Which providers are most expensive
  • 🔄 If there are cheaper alternatives
  • 📊 Your usage patterns over time

✅ The Solution

LLMscope gives you complete visibility and control over your LLM costs:

  • Real-time cost tracking - See costs as they happen
  • Cost breakdown - By provider, model, and time period
  • Smart recommendations - Get suggestions for cheaper models
  • Usage analytics - Track token usage and patterns
  • Self-hosted - Keep your data private

🚀 Quick Start (60 Seconds)

Step 1: Deploy LLMscope

Using Docker (Recommended):

git clone https://github.com/Blb3D/LLMscope.git
cd LLMscope
docker-compose up -d

Visit http://localhost:8081 - You'll see the dashboard (empty until you track your first API call).

Step 2: Track YOUR First LLM API Call

After making any OpenAI, Anthropic, or other LLM API call, add 3 lines to log the cost:

Example: Track OpenAI GPT-4 Call

importopenaiimportrequests# Your normal OpenAI API callresponse=openai.ChatCompletion.create(
model="gpt-4",
messages=[{"role": "user", "content": "Hello!"}]
)
# Add these 3 lines to track cost:requests.post('http://localhost:8000/api/usage', json={
'provider': 'openai',
'model': 'gpt-4',
'prompt_tokens': response['usage']['prompt_tokens'],
'completion_tokens': response['usage']['completion_tokens']
})

That's it! Refresh the dashboard to see your real costs.

Step 3 (Optional): Generate Demo Data for Testing

Want to test the dashboard before integrating your real API calls?

cd backend
python generate_demo_data.py

This creates 100 sample API calls to preview the dashboard features.


📊 Real-World Integration Examples

OpenAI Integration

Track every OpenAI API call:

importopenaiimportrequestsdeftrack_openai_usage(response):
"""Helper function to track OpenAI costs"""requests.post('http://localhost:8000/api/usage', json={
'provider': 'openai',
'model': response['model'],
'prompt_tokens': response['usage']['prompt_tokens'],
'completion_tokens': response['usage']['completion_tokens']
})
# Use it after any OpenAI call:response=openai.ChatCompletion.create(
model="gpt-4-turbo",
messages=[{"role": "user", "content": "Explain quantum computing"}]
)
track_openai_usage(response)

Anthropic Claude Integration

Track Claude API calls:

importanthropicimportrequestsdeftrack_anthropic_usage(model, response):
"""Helper function to track Anthropic costs"""requests.post('http://localhost:8000/api/usage', json={
'provider': 'anthropic',
'model': model,
'prompt_tokens': response.usage.input_tokens,
'completion_tokens': response.usage.output_tokens
})
# Use it after Claude API calls:client=anthropic.Anthropic(api_key="your-key")
message=client.messages.create(
model="claude-3-sonnet-20240229",
max_tokens=1024,
messages=[{"role": "user", "content": "Hello, Claude!"}]
)
track_anthropic_usage("claude-3-sonnet", message)

LangChain Integration

Automatic tracking with LangChain callback:

fromlangchain.callbacks.baseimportBaseCallbackHandlerfromlangchain.chat_modelsimportChatOpenAIimportrequestsclassLLMscopeCallback(BaseCallbackHandler):
defon_llm_end(self, response, **kwargs):
ifhasattr(response, 'llm_output') andresponse.llm_output:
usage=response.llm_output.get('token_usage', {})
requests.post('http://localhost:8000/api/usage', json={
'provider': 'openai',
'model': response.llm_output.get('model_name', 'gpt-3.5-turbo'),
'prompt_tokens': usage.get('prompt_tokens', 0),
'completion_tokens': usage.get('completion_tokens', 0)
})
# Use with LangChain:llm=ChatOpenAI(callbacks=[LLMscopeCallback()])
result=llm.predict("What is the capital of France?")

Google Gemini Integration

Track Gemini API calls:

importgoogle.generativeaiasgenaiimportrequestsdeftrack_gemini_usage(model_name, response):
"""Helper function to track Google Gemini costs"""requests.post('http://localhost:8000/api/usage', json={
'provider': 'google',
'model': model_name,
'prompt_tokens': response.usage_metadata.prompt_token_count,
'completion_tokens': response.usage_metadata.candidates_token_count
})
# Use it after Gemini calls:genai.configure(api_key="your-key")
model=genai.GenerativeModel('gemini-pro')
response=model.generate_content("Write a poem about AI")
track_gemini_usage('gemini-pro', response)

🔧 Manual Setup (Without Docker)

Backend:

cd backend
pip install -r requirements.txt
# Seed the database with LLM pricing data
python seed_pricing.py
# Start the backend
python app.py

Frontend:

cd frontend
npm install
npm run dev

Visit http://localhost:8081


📊 Features

💰 Real-Time Cost Tracking

Track every LLM API call with automatic cost calculation. Monitor spending across 63+ models from OpenAI, Anthropic, Google, Cohere, Together AI, Mistral, Groq, and more.

  • Instant cost visibility - See exactly what each request costs
  • Provider comparison - Compare costs across different LLM providers
  • Token usage analytics - Track prompt and completion tokens
  • Auto-refresh - Dashboard updates every 5 seconds

💡 Smart Cost Optimization

Cost Recommendations

Get intelligent recommendations for cheaper model alternatives:

  • Cheapest models first - Groq's Llama-3-8B at $0.000065/1K tokens
  • Side-by-side pricing - Compare input/output costs instantly
  • Recent usage history - Track your last 100 API calls
  • Save money automatically - Identify where you're overspending

🔐 Privacy-First & Self-Hosted

  • 100% local - Your data never leaves your infrastructure
  • No external dependencies - Runs entirely on Docker
  • Open source - Audit every line of code

🔌 API Reference

Log API Usage (POST)

Endpoint:POST http://localhost:8000/api/usage

Request Body:

{
"provider": "openai",
"model": "gpt-4",
"prompt_tokens": 100,
"completion_tokens": 50
}

Response:

{
"status": "logged",
"cost_usd": 0.006,
"timestamp": "2025-11-02T10:30:00Z"
}

Get Cost Summary (GET)

Endpoint:GET http://localhost:8000/api/costs/summary

Response:

{
"total_cost": 15.23,
"total_requests": 1250,
"by_provider": {
"openai": 12.45,
"anthropic": 2.78
},
"by_model": {
"gpt-4": 10.20,
"gpt-3.5-turbo": 2.25,
"claude-3-sonnet": 2.78
}
}

Get Model Recommendations (GET)

Endpoint:GET http://localhost:8000/api/recommendations

Returns a list of LLM models sorted by cost (cheapest first) with pricing details.

Interactive API Docs

Visit http://localhost:8000/docs for full interactive API documentation.


⚙️ Configuration

Create a .env file:

DATABASE_PATH=./data/llmscope.dbLLMSCOPE_API_KEY=your-secret-key

🏗️ Architecture

┌─────────────────┐
│ Frontend │ React + Vite
│ (Port 8081) │ Cost dashboard
└────────┬────────┘
│
│ HTTP/REST
│
┌────────▼────────┐
│ Backend API │ FastAPI + SQLite
│ (Port 8000) │ Cost tracking
└─────────────────┘

📦 Tech Stack

Backend:

  • FastAPI
  • SQLite
  • Python 3.9+

Frontend:

  • React
  • Vite
  • Tailwind CSS

Deployment:

  • Docker
  • Docker Compose

🗄️ Database Schema

api_usage

Tracks all API calls with token counts and costs

model_pricing

Stores pricing data for different LLM models

settings

Application configuration


🛠️ Development

# Backend testscd backend
pytest
# Frontend developmentcd frontend
npm run dev
# Production build
docker-compose -f docker-compose.prod.yml up -d

🗺️ Roadmap

Supported Providers: OpenAI, Anthropic, Google, Cohere, Together AI, Mistral, Groq, Ollama, and 60+ models

Coming Soon:

  • Cost alerts and budget thresholds
  • Export to CSV/PDF
  • More provider integrations (request yours in Issues!)

Want to influence the roadmap? Open an issue or start a discussion!


🤝 Contributing

Contributions are welcome! Please feel free to submit a Pull Request.


📄 License

Business Source License 1.1

✅ Free for Non-Commercial Use

  • Self-hosting for personal use, education, and research
  • Modify and redistribute for non-commercial purposes
  • Full source code access - no restrictions on reading the code

💼 Commercial Use Requires License

Commercial use includes:

  • Using LLMscope to monitor production LLM deployments in a business
  • Offering LLMscope as a hosted/managed service to customers
  • Incorporating LLMscope into a commercial product

Need a commercial license? Contact: bbaker@blb3dprinting.com

🔓 Future: Converts to MIT License

On October 29, 2028 (3 years from first publication), this license automatically converts to MIT - making it fully open source forever.


Why BSL? We want LLMscope to be freely available for individuals and small teams, while ensuring companies using it commercially contribute back. This allows us to keep developing new features like SPC analysis, AI copilot, and enhanced reporting.


💬 Support


Built with ❤️ for the LLM community

About

Real-time Statistical Process Control (SPC) monitoring for LLMs

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - BLB3DPrinting/LLMscope: Real-time Statistical Process Control (SPC) monitoring for LLMs · GitHub
Skip to content

Repository files navigation

💰 LLMscope - Track Your LLM Costs in Real-Time

Stop guessing what your ChatGPT and Claude API calls cost. Track them in 60 seconds with 3 lines of code.

LicenseDocker ReadyPRs Welcome

LLMscope Dashboard

🎯 Quick Access

After running docker-compose up -d:


🚨 The Problem

LLM API costs are spiraling out of control. You don't know:

  • 💸 How much you're spending on each model
  • 📈 Which providers are most expensive
  • 🔄 If there are cheaper alternatives
  • 📊 Your usage patterns over time

✅ The Solution

LLMscope gives you complete visibility and control over your LLM costs:

  • Real-time cost tracking - See costs as they happen
  • Cost breakdown - By provider, model, and time period
  • Smart recommendations - Get suggestions for cheaper models
  • Usage analytics - Track token usage and patterns
  • Self-hosted - Keep your data private

🚀 Quick Start (60 Seconds)

Step 1: Deploy LLMscope

Using Docker (Recommended):

git clone https://github.com/Blb3D/LLMscope.git
cd LLMscope
docker-compose up -d

Visit http://localhost:8081 - You'll see the dashboard (empty until you track your first API call).

Step 2: Track YOUR First LLM API Call

After making any OpenAI, Anthropic, or other LLM API call, add 3 lines to log the cost:

Example: Track OpenAI GPT-4 Call

importopenaiimportrequests# Your normal OpenAI API callresponse=openai.ChatCompletion.create(
model="gpt-4",
messages=[{"role": "user", "content": "Hello!"}]
)
# Add these 3 lines to track cost:requests.post('http://localhost:8000/api/usage', json={
'provider': 'openai',
'model': 'gpt-4',
'prompt_tokens': response['usage']['prompt_tokens'],
'completion_tokens': response['usage']['completion_tokens']
})

That's it! Refresh the dashboard to see your real costs.

Step 3 (Optional): Generate Demo Data for Testing

Want to test the dashboard before integrating your real API calls?

cd backend
python generate_demo_data.py

This creates 100 sample API calls to preview the dashboard features.


📊 Real-World Integration Examples

OpenAI Integration

Track every OpenAI API call:

importopenaiimportrequestsdeftrack_openai_usage(response):
"""Helper function to track OpenAI costs"""requests.post('http://localhost:8000/api/usage', json={
'provider': 'openai',
'model': response['model'],
'prompt_tokens': response['usage']['prompt_tokens'],
'completion_tokens': response['usage']['completion_tokens']
})
# Use it after any OpenAI call:response=openai.ChatCompletion.create(
model="gpt-4-turbo",
messages=[{"role": "user", "content": "Explain quantum computing"}]
)
track_openai_usage(response)

Anthropic Claude Integration

Track Claude API calls:

importanthropicimportrequestsdeftrack_anthropic_usage(model, response):
"""Helper function to track Anthropic costs"""requests.post('http://localhost:8000/api/usage', json={
'provider': 'anthropic',
'model': model,
'prompt_tokens': response.usage.input_tokens,
'completion_tokens': response.usage.output_tokens
})
# Use it after Claude API calls:client=anthropic.Anthropic(api_key="your-key")
message=client.messages.create(
model="claude-3-sonnet-20240229",
max_tokens=1024,
messages=[{"role": "user", "content": "Hello, Claude!"}]
)
track_anthropic_usage("claude-3-sonnet", message)

LangChain Integration

Automatic tracking with LangChain callback:

fromlangchain.callbacks.baseimportBaseCallbackHandlerfromlangchain.chat_modelsimportChatOpenAIimportrequestsclassLLMscopeCallback(BaseCallbackHandler):
defon_llm_end(self, response, **kwargs):
ifhasattr(response, 'llm_output') andresponse.llm_output:
usage=response.llm_output.get('token_usage', {})
requests.post('http://localhost:8000/api/usage', json={
'provider': 'openai',
'model': response.llm_output.get('model_name', 'gpt-3.5-turbo'),
'prompt_tokens': usage.get('prompt_tokens', 0),
'completion_tokens': usage.get('completion_tokens', 0)
})
# Use with LangChain:llm=ChatOpenAI(callbacks=[LLMscopeCallback()])
result=llm.predict("What is the capital of France?")

Google Gemini Integration

Track Gemini API calls:

importgoogle.generativeaiasgenaiimportrequestsdeftrack_gemini_usage(model_name, response):
"""Helper function to track Google Gemini costs"""requests.post('http://localhost:8000/api/usage', json={
'provider': 'google',
'model': model_name,
'prompt_tokens': response.usage_metadata.prompt_token_count,
'completion_tokens': response.usage_metadata.candidates_token_count
})
# Use it after Gemini calls:genai.configure(api_key="your-key")
model=genai.GenerativeModel('gemini-pro')
response=model.generate_content("Write a poem about AI")
track_gemini_usage('gemini-pro', response)

🔧 Manual Setup (Without Docker)

Backend:

cd backend
pip install -r requirements.txt
# Seed the database with LLM pricing data
python seed_pricing.py
# Start the backend
python app.py

Frontend:

cd frontend
npm install
npm run dev

Visit http://localhost:8081


📊 Features

💰 Real-Time Cost Tracking

Track every LLM API call with automatic cost calculation. Monitor spending across 63+ models from OpenAI, Anthropic, Google, Cohere, Together AI, Mistral, Groq, and more.

  • Instant cost visibility - See exactly what each request costs
  • Provider comparison - Compare costs across different LLM providers
  • Token usage analytics - Track prompt and completion tokens
  • Auto-refresh - Dashboard updates every 5 seconds

💡 Smart Cost Optimization

Cost Recommendations

Get intelligent recommendations for cheaper model alternatives:

  • Cheapest models first - Groq's Llama-3-8B at $0.000065/1K tokens
  • Side-by-side pricing - Compare input/output costs instantly
  • Recent usage history - Track your last 100 API calls
  • Save money automatically - Identify where you're overspending

🔐 Privacy-First & Self-Hosted

  • 100% local - Your data never leaves your infrastructure
  • No external dependencies - Runs entirely on Docker
  • Open source - Audit every line of code

🔌 API Reference

Log API Usage (POST)

Endpoint:POST http://localhost:8000/api/usage

Request Body:

{
"provider": "openai",
"model": "gpt-4",
"prompt_tokens": 100,
"completion_tokens": 50
}

Response:

{
"status": "logged",
"cost_usd": 0.006,
"timestamp": "2025-11-02T10:30:00Z"
}

Get Cost Summary (GET)

Endpoint:GET http://localhost:8000/api/costs/summary

Response:

{
"total_cost": 15.23,
"total_requests": 1250,
"by_provider": {
"openai": 12.45,
"anthropic": 2.78
},
"by_model": {
"gpt-4": 10.20,
"gpt-3.5-turbo": 2.25,
"claude-3-sonnet": 2.78
}
}

Get Model Recommendations (GET)

Endpoint:GET http://localhost:8000/api/recommendations

Returns a list of LLM models sorted by cost (cheapest first) with pricing details.

Interactive API Docs

Visit http://localhost:8000/docs for full interactive API documentation.


⚙️ Configuration

Create a .env file:

DATABASE_PATH=./data/llmscope.dbLLMSCOPE_API_KEY=your-secret-key

🏗️ Architecture

┌─────────────────┐
│ Frontend │ React + Vite
│ (Port 8081) │ Cost dashboard
└────────┬────────┘
│
│ HTTP/REST
│
┌────────▼────────┐
│ Backend API │ FastAPI + SQLite
│ (Port 8000) │ Cost tracking
└─────────────────┘

📦 Tech Stack

Backend:

  • FastAPI
  • SQLite
  • Python 3.9+

Frontend:

  • React
  • Vite
  • Tailwind CSS

Deployment:

  • Docker
  • Docker Compose

🗄️ Database Schema

api_usage

Tracks all API calls with token counts and costs

model_pricing

Stores pricing data for different LLM models

settings

Application configuration


🛠️ Development

# Backend testscd backend
pytest
# Frontend developmentcd frontend
npm run dev
# Production build
docker-compose -f docker-compose.prod.yml up -d

🗺️ Roadmap

Supported Providers: OpenAI, Anthropic, Google, Cohere, Together AI, Mistral, Groq, Ollama, and 60+ models

Coming Soon:

  • Cost alerts and budget thresholds
  • Export to CSV/PDF
  • More provider integrations (request yours in Issues!)

Want to influence the roadmap? Open an issue or start a discussion!


🤝 Contributing

Contributions are welcome! Please feel free to submit a Pull Request.


📄 License

Business Source License 1.1

✅ Free for Non-Commercial Use

  • Self-hosting for personal use, education, and research
  • Modify and redistribute for non-commercial purposes
  • Full source code access - no restrictions on reading the code

💼 Commercial Use Requires License

Commercial use includes:

  • Using LLMscope to monitor production LLM deployments in a business
  • Offering LLMscope as a hosted/managed service to customers
  • Incorporating LLMscope into a commercial product

Need a commercial license? Contact: bbaker@blb3dprinting.com

🔓 Future: Converts to MIT License

On October 29, 2028 (3 years from first publication), this license automatically converts to MIT - making it fully open source forever.


Why BSL? We want LLMscope to be freely available for individuals and small teams, while ensuring companies using it commercially contribute back. This allows us to keep developing new features like SPC analysis, AI copilot, and enhanced reporting.


💬 Support


Built with ❤️ for the LLM community

About

Real-time Statistical Process Control (SPC) monitoring for LLMs

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - BLB3DPrinting/LLMscope: Real-time Statistical Process Control (SPC) monitoring for LLMs · GitHub
Skip to content

Repository files navigation

💰 LLMscope - Track Your LLM Costs in Real-Time

Stop guessing what your ChatGPT and Claude API calls cost. Track them in 60 seconds with 3 lines of code.

LicenseDocker ReadyPRs Welcome

LLMscope Dashboard

🎯 Quick Access

After running docker-compose up -d:


🚨 The Problem

LLM API costs are spiraling out of control. You don't know:

  • 💸 How much you're spending on each model
  • 📈 Which providers are most expensive
  • 🔄 If there are cheaper alternatives
  • 📊 Your usage patterns over time

✅ The Solution

LLMscope gives you complete visibility and control over your LLM costs:

  • Real-time cost tracking - See costs as they happen
  • Cost breakdown - By provider, model, and time period
  • Smart recommendations - Get suggestions for cheaper models
  • Usage analytics - Track token usage and patterns
  • Self-hosted - Keep your data private

🚀 Quick Start (60 Seconds)

Step 1: Deploy LLMscope

Using Docker (Recommended):

git clone https://github.com/Blb3D/LLMscope.git
cd LLMscope
docker-compose up -d

Visit http://localhost:8081 - You'll see the dashboard (empty until you track your first API call).

Step 2: Track YOUR First LLM API Call

After making any OpenAI, Anthropic, or other LLM API call, add 3 lines to log the cost:

Example: Track OpenAI GPT-4 Call

importopenaiimportrequests# Your normal OpenAI API callresponse=openai.ChatCompletion.create(
model="gpt-4",
messages=[{"role": "user", "content": "Hello!"}]
)
# Add these 3 lines to track cost:requests.post('http://localhost:8000/api/usage', json={
'provider': 'openai',
'model': 'gpt-4',
'prompt_tokens': response['usage']['prompt_tokens'],
'completion_tokens': response['usage']['completion_tokens']
})

That's it! Refresh the dashboard to see your real costs.

Step 3 (Optional): Generate Demo Data for Testing

Want to test the dashboard before integrating your real API calls?

cd backend
python generate_demo_data.py

This creates 100 sample API calls to preview the dashboard features.


📊 Real-World Integration Examples

OpenAI Integration

Track every OpenAI API call:

importopenaiimportrequestsdeftrack_openai_usage(response):
"""Helper function to track OpenAI costs"""requests.post('http://localhost:8000/api/usage', json={
'provider': 'openai',
'model': response['model'],
'prompt_tokens': response['usage']['prompt_tokens'],
'completion_tokens': response['usage']['completion_tokens']
})
# Use it after any OpenAI call:response=openai.ChatCompletion.create(
model="gpt-4-turbo",
messages=[{"role": "user", "content": "Explain quantum computing"}]
)
track_openai_usage(response)

Anthropic Claude Integration

Track Claude API calls:

importanthropicimportrequestsdeftrack_anthropic_usage(model, response):
"""Helper function to track Anthropic costs"""requests.post('http://localhost:8000/api/usage', json={
'provider': 'anthropic',
'model': model,
'prompt_tokens': response.usage.input_tokens,
'completion_tokens': response.usage.output_tokens
})
# Use it after Claude API calls:client=anthropic.Anthropic(api_key="your-key")
message=client.messages.create(
model="claude-3-sonnet-20240229",
max_tokens=1024,
messages=[{"role": "user", "content": "Hello, Claude!"}]
)
track_anthropic_usage("claude-3-sonnet", message)

LangChain Integration

Automatic tracking with LangChain callback:

fromlangchain.callbacks.baseimportBaseCallbackHandlerfromlangchain.chat_modelsimportChatOpenAIimportrequestsclassLLMscopeCallback(BaseCallbackHandler):
defon_llm_end(self, response, **kwargs):
ifhasattr(response, 'llm_output') andresponse.llm_output:
usage=response.llm_output.get('token_usage', {})
requests.post('http://localhost:8000/api/usage', json={
'provider': 'openai',
'model': response.llm_output.get('model_name', 'gpt-3.5-turbo'),
'prompt_tokens': usage.get('prompt_tokens', 0),
'completion_tokens': usage.get('completion_tokens', 0)
})
# Use with LangChain:llm=ChatOpenAI(callbacks=[LLMscopeCallback()])
result=llm.predict("What is the capital of France?")

Google Gemini Integration

Track Gemini API calls:

importgoogle.generativeaiasgenaiimportrequestsdeftrack_gemini_usage(model_name, response):
"""Helper function to track Google Gemini costs"""requests.post('http://localhost:8000/api/usage', json={
'provider': 'google',
'model': model_name,
'prompt_tokens': response.usage_metadata.prompt_token_count,
'completion_tokens': response.usage_metadata.candidates_token_count
})
# Use it after Gemini calls:genai.configure(api_key="your-key")
model=genai.GenerativeModel('gemini-pro')
response=model.generate_content("Write a poem about AI")
track_gemini_usage('gemini-pro', response)

🔧 Manual Setup (Without Docker)

Backend:

cd backend
pip install -r requirements.txt
# Seed the database with LLM pricing data
python seed_pricing.py
# Start the backend
python app.py

Frontend:

cd frontend
npm install
npm run dev

Visit http://localhost:8081


📊 Features

💰 Real-Time Cost Tracking

Track every LLM API call with automatic cost calculation. Monitor spending across 63+ models from OpenAI, Anthropic, Google, Cohere, Together AI, Mistral, Groq, and more.

  • Instant cost visibility - See exactly what each request costs
  • Provider comparison - Compare costs across different LLM providers
  • Token usage analytics - Track prompt and completion tokens
  • Auto-refresh - Dashboard updates every 5 seconds

💡 Smart Cost Optimization

Cost Recommendations

Get intelligent recommendations for cheaper model alternatives:

  • Cheapest models first - Groq's Llama-3-8B at $0.000065/1K tokens
  • Side-by-side pricing - Compare input/output costs instantly
  • Recent usage history - Track your last 100 API calls
  • Save money automatically - Identify where you're overspending

🔐 Privacy-First & Self-Hosted

  • 100% local - Your data never leaves your infrastructure
  • No external dependencies - Runs entirely on Docker
  • Open source - Audit every line of code

🔌 API Reference

Log API Usage (POST)

Endpoint:POST http://localhost:8000/api/usage

Request Body:

{
"provider": "openai",
"model": "gpt-4",
"prompt_tokens": 100,
"completion_tokens": 50
}

Response:

{
"status": "logged",
"cost_usd": 0.006,
"timestamp": "2025-11-02T10:30:00Z"
}

Get Cost Summary (GET)

Endpoint:GET http://localhost:8000/api/costs/summary

Response:

{
"total_cost": 15.23,
"total_requests": 1250,
"by_provider": {
"openai": 12.45,
"anthropic": 2.78
},
"by_model": {
"gpt-4": 10.20,
"gpt-3.5-turbo": 2.25,
"claude-3-sonnet": 2.78
}
}

Get Model Recommendations (GET)

Endpoint:GET http://localhost:8000/api/recommendations

Returns a list of LLM models sorted by cost (cheapest first) with pricing details.

Interactive API Docs

Visit http://localhost:8000/docs for full interactive API documentation.


⚙️ Configuration

Create a .env file:

DATABASE_PATH=./data/llmscope.dbLLMSCOPE_API_KEY=your-secret-key

🏗️ Architecture

┌─────────────────┐
│ Frontend │ React + Vite
│ (Port 8081) │ Cost dashboard
└────────┬────────┘
│
│ HTTP/REST
│
┌────────▼────────┐
│ Backend API │ FastAPI + SQLite
│ (Port 8000) │ Cost tracking
└─────────────────┘

📦 Tech Stack

Backend:

  • FastAPI
  • SQLite
  • Python 3.9+

Frontend:

  • React
  • Vite
  • Tailwind CSS

Deployment:

  • Docker
  • Docker Compose

🗄️ Database Schema

api_usage

Tracks all API calls with token counts and costs

model_pricing

Stores pricing data for different LLM models

settings

Application configuration


🛠️ Development

# Backend testscd backend
pytest
# Frontend developmentcd frontend
npm run dev
# Production build
docker-compose -f docker-compose.prod.yml up -d

🗺️ Roadmap

Supported Providers: OpenAI, Anthropic, Google, Cohere, Together AI, Mistral, Groq, Ollama, and 60+ models

Coming Soon:

  • Cost alerts and budget thresholds
  • Export to CSV/PDF
  • More provider integrations (request yours in Issues!)

Want to influence the roadmap? Open an issue or start a discussion!


🤝 Contributing

Contributions are welcome! Please feel free to submit a Pull Request.


📄 License

Business Source License 1.1

✅ Free for Non-Commercial Use

  • Self-hosting for personal use, education, and research
  • Modify and redistribute for non-commercial purposes
  • Full source code access - no restrictions on reading the code

💼 Commercial Use Requires License

Commercial use includes:

  • Using LLMscope to monitor production LLM deployments in a business
  • Offering LLMscope as a hosted/managed service to customers
  • Incorporating LLMscope into a commercial product

Need a commercial license? Contact: bbaker@blb3dprinting.com

🔓 Future: Converts to MIT License

On October 29, 2028 (3 years from first publication), this license automatically converts to MIT - making it fully open source forever.


Why BSL? We want LLMscope to be freely available for individuals and small teams, while ensuring companies using it commercially contribute back. This allows us to keep developing new features like SPC analysis, AI copilot, and enhanced reporting.


💬 Support


Built with ❤️ for the LLM community

About

Real-time Statistical Process Control (SPC) monitoring for LLMs

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages