Skip to content
View triple-mu's full-sized avatar
💭
Contributing for the open source community.
💭
Contributing for the open source community.

Block or report triple-mu

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
triple-mu/README.md

triple-mu — ML Systems Engineer — Training/Inference Infra

ML Systems / AI Infra
I work on inference optimization, GPU computing, and training infrastructure.

Profile views GitHub followers Email: gpu@163.com

Open Source

I've also contributed to inference optimization in cache-dit and LightX2V.

Selected Projects

Author

Efficient YOLOv8 deployment with TensorRT in Python and C++.

Author

Efficient GPU communication for sequence parallelism.

Tech Stack

Languages & GPU

python cplusplus cuda java html JavaScript

Frameworks

pytorch oneflow numpy

Tools & Platforms

cmake pycharm clion jupyter git github github-actions vscode latex visualstudio macOS linux ubuntu windows redmi

Earlier Contributions

MMYOLO, YOLOv6, YOLOv7, and yolort.

📊 GitHub Stats & Productive Time

Stats Productive Time

Popular repositories Loading

  1. YOLOv8-TensorRT YOLOv8-TensorRT Public

    YOLOv8 using TensorRT accelerate !

    Python 1.8k 299

  2. ncnn-examples ncnn-examples Public

    Learning ncnn with some examples

    Python 74 11

  3. fast-ulysses fast-ulysses Public

    Ulysses sequence-parallel all-to-all as a torch custom op, moved by the GPU copy engines into torch symmetric memory. Zero SM usage; 1.66-2.17x over torch.distributed on NVLink.

    Python 48 7

  4. TensorRT2ONNX TensorRT2ONNX Public

    A tool convert TensorRT engine/plan to a fake onnx

    Python 41 5

  5. yolov8 yolov8 Public

    Forked from ultralytics/ultralytics

    YOLOv8 🚀 in PyTorch > ONNX > CoreML > TFLite

    Python 22 7

  6. Qwen-Image-TensorRT Qwen-Image-TensorRT Public

    Qwen-Image's DiT inference with TensorRT-10

    Python 21 2