Skip to content
View punitvara's full-sized avatar

Highlights

  • Pro

Block or report punitvara

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
punitvara/README.md

👋 About Me

Hi there, this is Punit Vara. Please feel free to contact me on LinkedIn

Software/ML systems engineer with experience across backend engineering, distributed data systems, cloud/MLOps, and LLM infrastructure — increasingly specializing in GPU-accelerated inference, distributed AI systems, and high-performance ML infrastructure. I like understanding systems from first principles rather than treating frameworks as black boxes.

🌇 Experience

🔬 Focus Areas

LLM Inference & Serving (KV cache, quantization, llama.cpp) · GPU Systems & Collective Communication (NCCL, All-Reduce) · Distributed Data Systems (Kafka, Spark, Trino) · ML Infrastructure & MLOps (Kubernetes, Terraform, MLflow)

🌱 Currently Exploring

GPU/NCCL communication primitives and distributed inference performance bottlenecks

📫 Contact

Pinned Loading

  1. perplexity-search-extensionperplexity-search-extensionPublic

    JavaScript 1

  2. image-search-appimage-search-appPublic

    Python

  3. ml-practice-notebookml-practice-notebookPublic

    Jupyter Notebook

  4. machine_learningmachine_learningPublic

    Forked from omairaasim/machine_learning

    Python

  5. PVKeyboardCleaner-PVKeyboardCleaner-Public

    Full-screen macOS utility that blocks all keystrokes system-wide so you can safely clean your keyboard — no accidental typing, launches, or shortcuts.

    Swift

  6. vllm-project/production-stackvllm-project/production-stackPublic

    vLLM’s reference system for K8S-native cluster-wide deployment with community-driven performance optimization

    Python 2.6k 482