ML Systems / AI Infra
I work on inference optimization, GPU computing, and training infrastructure.
- SGLang — Reviewer & Code Owner for SGLang Diffusion, working on inference performance and resource efficiency.
- FlashInfer — Contributing to GPU communication optimization.
- NeMo RL — Contributing to reinforcement learning infrastructure.
I've also contributed to inference optimization in cache-dit and LightX2V.
|
Author Efficient YOLOv8 deployment with TensorRT in Python and C++. |
Author Efficient GPU communication for sequence parallelism. |
Languages & GPU
Frameworks
Tools & Platforms




