A FastAPI-based LLM serving layer with pluggable inference backends, custom routing, and GCP deployment support.
-
Updated
Apr 14, 2026 - Python
A FastAPI-based LLM serving layer with pluggable inference backends, custom routing, and GCP deployment support.
To associate your repository with the multiple-llm topic, visit your repo's landing page and select "manage topics."