KV-cache-aware intelligent routing for self-hosted and hybrid LLM fleets. Route requests using model quality, latency, cost, policy, and live GPU state.
multi-armed-banditrouting-controllersmlopsfleetskv-cacheopenai-apipii-redactionllmllmopsvllmai-gatewaysemantic-routingllm-gatewayllm-routingself-hosted-llm
-
Updated
Mar 20, 2026 - Python