Run GGUF through llama.cpp and SafeTensors through vLLM behind one OpenAI-compatible endpoint. Your coding tools select a model; the switchboard manages the local runtime, process, and resident-model change.
openaillama-cppvllmlocal-llmllm-inferencellm-serveggufgguf-modelsvllm-serveopenai-compatiblegguf-model-supportmodel-swapllm-servicegguf-managervllm-servermodel-swappingllm-serving-systemsgguf-modelgguf-runnermodels-switcher
-
Updated
Sep 6, 2026 - Rust