I build efficient AI systems, from GPU kernels and LLM serving to grounded, observable agent products.
Bengaluru, India · Open to AI Engineer, AI Inference Engineer, LLM Engineer, and GenAI roles
| Fast Inference | Agentic Systems | Reliable ML Products |
|---|---|---|
| ROCm, MI300X/MI355X, FP8, MXFP4, Triton, vLLM, profiling | RAG, LangGraph, Google ADK, MCP, multi-agent workflows | Evaluation, grounding, FastAPI, Docker, Cloud Run, Vertex AI |
I focus on the engineering details that make AI useful in the real world: latency, cost, evaluation, reliability, and deployment.
An AI assistant for AMD ROCm developers. It reads profiler output or training metrics, explains likely GPU bottlenecks, and suggests concrete optimizations.
Explore:Live demo · README · SFT training · Benchmarking
LoRADPOMI300XROCmvLLMGradio
A low-level inference case study covering MXFP4 GEMM, MoE MXFP4, and Mixed MLA decode. It documents quantization-aware dispatch, runtime-path tuning, metadata reuse, and benchmark-driven iteration while preserving correctness.
Explore:README · MXFP4 GEMM · MoE MXFP4 · Mixed MLA · Benchmark summary
FP8MXFP4GEMMKernel OptimizationLatency Benchmarking
An OpenEnv benchmark for agents working with stale or conflicting knowledge. The agent detects hallucinations, identifies outdated sources, repairs the knowledge base, and verifies corrected answers.
Explore:Live demo · API docs · README · Inference · Task suite
RAG EvaluationOpenEnvGroundingAI Safety
An API-first assistant that turns a natural-language goal into a structured workflow. Agents retrieve context, plan work, create tasks and notes, schedule calendar events, and return clean results through FastAPI.
Explore:README · Agent workflow · MCP tools
FastAPIGeminiGoogle ADKMCPAlloyDB
- FlashAttention, KV-cache optimization, Triton, and vLLM internals
- Distributed inference, serving systems, and ML systems design
- Evaluation methods for grounded and reliable AI agents
Building efficient, scalable AI systems, from GPU kernels to intelligent agents.

