Serverless-GPU LLM serving: scale-to-zero with fast GPU snapshot/restore (cuda-checkpoint), multi-tenant packing, and an OpenAI-compatible API — built on vLLM.
serverlessgpucudaloraautoscalingscale-to-zeroopenai-apillm-servingvllmllm-inferencesnapshot-restorecuda-checkpoint
-
Updated
Jul 5, 2026 - Python