Skip to content
View hiteshsahu's full-sized avatar
📈
Trying to make a difference
📈
Trying to make a difference

Block or report hiteshsahu

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
hiteshsahu/README.md

👋 Hi, I'm Hitesh Sahu

🚀 AI Infrastructure Engineer • GPU Systems • Cloud Native • Open Source

🌍 View Portfolio:https://hiteshsahu.com

I build the infrastructure behind modern AI systems.

My work focuses on GPU clusters, Slurm, Kubernetes, distributed systems, observability, and * developer tooling* that makes AI workloads easier to run, benchmark, and debug.

I'm currently building an open-source ecosystem for AI infrastructure, including local Slurm clusters, GPU observability, HPC developer tools, and LLM benchmarking.


🏆 Certifications

🎓 View all on Credly →

AI Infrastructure Ecosystem

These projects complement each other and solve a piece of the puzzle in the AI workflow from training, inference, deployment to monitoring

ProjectDescriptionTech
🏋️ Model Gym

A fitness center for AI models. Import, export, optimize, benchmark, and report LLM inference performance across engines, runtimes, and hardware platforms.Next.js • TypeScript • Python • PyTorch • Hugging Face
🦆 RAG Factory

Transforms chaotic PDFs, documents, websites, databases, and APIs into trusted answers using embeddings, retrieval, reranking, and large language models.Python • FastAPI • LangChain • Vector Databases • OpenAI • Ollama
🐸 NVIDIA SuperPod

GPU Infrastructure Lab for building an AI supercomputer from commodity GPU servers. Explore multi-node training, networking, storage, scheduling, observability, and large-scale AI infrastructure.Go • Kubernetes • Slurm • NVIDIA GPUs • InfiniBand • Prometheus
🏴‍☠️ GhostFleet

Simulates a 1,000-node / 8,000-GPU Kubernetes cluster on a laptop using KWOK, loads it with ClusterLoader2 and a custom GPU scheduling workload, and measures control-plane behavior against upstream scalability SLOs.Go • Kubernetes • KWOK • ClusterLoader2 • Prometheus • Docker
֎ GPU Lens

Drop-in GPU + scheduler observability for clusters you already have. Get instant visibility into GPU health, utilization, memory, temperatures, ECC errors, XID faults, scheduler activity, and queue health.Go • Prometheus • Grafana • DCGM Exporter • Kubernetes • Slurm
🦝🐾 Squint

A GPU-aware Slurm monitor for your terminal. Read-only, zero-config, and runs anywhere. Visualize jobs, nodes, GPUs, queue health, and pending reasons through a fast terminal UI.Go • Bubble Tea • Lip Gloss • Slurm • TUI
🐪 Caravan

Spin up a complete local Slurm cluster with a single command. Develop, test, and submit HPC and AI workloads on Docker or Podman with GPU scheduling, making local experimentation fast and reproducible.Go • Slurm • Docker • Podman • Cobra • HPC
🛜 GPU-Fabric-Bench

Reproducible RDMA fabric benchmarking suite for NCCL GPU collective communications on AWS EFA. Maps InfiniBand concepts to cloud-native HPC networking and visualizes latency, bandwidth, topology, and scaling behavior.NCCL • AWS EFA • RDMA • MPI • NVIDIA GPUs • Python

📊 Open Source Activity

GitHub StatsStreak

📈 Contribution Graph


⚙️ Tech Stack

🤖 AI & Machine Learning

PyTorchHugging FaceLangChainOllamaOpenAI

⚡ GPU Computing & HPC

CUDANVIDIASlurmNCCLRDMAInfiniBand

🛠 Backend & Distributed Systems

JavaSpring BootQuarkusFastAPIKafkaPostgreSQL

☸️ Cloud Native

GoKubernetesHelmTerraformDockerPodmanAWS

📊 Observability

PrometheusGrafanaOpenTelemetryDCGM

🎨 Frontend

ReactNext.jsTypeScript


GitHubLinkedInEmailX

🌍 hiteshsahu.com • 📚 Stack Overflow (42k+)

```

Pinned Loading

  1. ECommerce-App-AndroidECommerce-App-AndroidPublic

    E-Commerce App for Android with Material Design Pattern

    Java 597 476

  2. Nvidia-Super-PodNvidia-Super-PodPublic

    Custom self hosted AWS GPU cluster with Ansible and Kubernates for ML workload

    HCL 1

  3. Model-GymModel-GymPublic

    GPU Pipeline for Model Import + Inference + Benchmark + Dashboard + Export

    Python 1

  4. caravancaravanPublic

    A CLI for GPU Slurm - stand up a SLURM cluster in one command, built to run your workloads on it.

    Go 1

  5. ghostfleetghostfleetPublic

    Kubernetes scalability testing powered by KWOK and ClusterLoader2.

    HTML 1

  6. RAG-FactoryRAG-FactoryPublic

    RAG Factory 🦆 (powered by the Raginator-3000) transforms chaotic PDFs, docs, websites, and APIs into trusted answers using embeddings, retrieval, reranking, and LLMs.

    Python 1