Multi-backend LLM serving and training platform — vLLM/Triton/Ray Serve/KServe/BentoML behind one contract, Kueue/Karpenter GPU orchestration, Ray Train/FSDP/DeepSpeed with LoRA/PEFT and DVC, MLflow/W&B tracking, and a tool-grounded LangGraph advisor — CI-validated without real GPU cost.
kubernetesbedrockloradistributed-trainingfine-tuningpeftdvcsagemakermlflowbentomlexperiment-trackingtriton-inference-serverai-agentkservekarpenterray-trainray-servevllmlanggraphkueue
-
Updated
Jul 27, 2026 - Python