High-level FastAPI service to help with RAG (Retrieval Augmented Generation) pipelines. It manages approved embedding models, provides an API to preload/unload models, compute embeddings with optional Redis caching, and exposes basic monitoring hooks
docker embedded cpu async docker-compose asynchronous cuda torch pytorch reranking uv sentence-embeddings reranker fastapi huggingface rorm sentence-transformers huggingface-transformers
-
Updated
Nov 23, 2025 - Python