Skip to content
View wu840407's full-sized avatar
💭
AI infra & speech-AI engineer · self-hosted LLM · ASR · PyTorch/CUDA
💭
AI infra & speech-AI engineer · self-hosted LLM · ASR · PyTorch/CUDA

Block or report wu840407

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
wu840407/README.md

Hi, I'm Roger (ChengRung Wu) 👋

AI infrastructure engineer — I make large language models run fast on hardware that shouldn't be able to run them.

9 years building and deploying mission-critical systems. Currently focused on LLM inference optimization: quantization, tensor parallelism, and serving 30B-class models fully offline on constrained GPUs.


🚀 Featured Projects

ProjectWhat it is
enterprise-airgapped-llmProduction self-hosted LLM platform for air-gapped environments — 30B MoE (AWQ 4-bit) across dual Turing GPUs via tensor parallelism, 80–120 tok/s, sub-500 ms TTFT, zero cloud dependency.
YaYan-AIOffline multi-dialect speech-intelligence system — 22 Chinese dialects + 40 languages, speaker diarization, character-level timestamps, LLM-assisted correction.
yolov5-parallelMulti-GPU parallelized YOLOv5 training and inference pipeline.
MultiBodyCuboidsM.S. thesis — multi-body motion segmentation for arbitrary numbers of disordered 3D point sets.
openclaw-agentsAutonomous LLM agents — OODA decision support, scheduled briefings, human-in-the-loop safety gates.

🛠 Tech Stack

LLM InferencevLLM · quantization (AWQ / INT4) · tensor parallelism · KV-cache tuning · FlashInfer · TensorRT-LLM

ML & LanguagesPyTorch · HuggingFace · CUDA · Python · C++

InfrastructureDocker · Linux · Kafka · Elasticsearch · PostgreSQL · AWS

🎓 Background

M.S. CS @ NYCU · AWS Solutions Architect – Associate · CEH · CompTIA Security+ / Network+ · HITCON white-hat · CVE research

📫 Contact

wu840407@gmail.com — open to AI infrastructure / LLM inference / ML systems roles (2027)

GitHub stats

Pinned Loading

  1. YaYan-AIYaYan-AIPublic

    Fully-offline multi-dialect speech-intelligence system — 22 Chinese dialects + 40 languages, speaker diarization, character-level timestamps, LLM-assisted correction.

    Python 1

  2. enterprise-airgapped-llmenterprise-airgapped-llmPublic

    Production reference architecture for self-hosted LLM in air-gapped enterprise environments. Dell R740 + Turing GPUs + vLLM + Qwen3-Coder + AD/LDAPS.

    Shell

  3. openclaw-agentsopenclaw-agentsPublic

    Four standalone autonomous LLM agents — OODA strategy support, scheduled briefings, content pipelines with human-in-the-loop safety gates.

    Python

  4. MultiBodyCuboidsMultiBodyCuboidsPublic

    M.S. thesis — multi-body motion-segmentation architecture for arbitrary numbers of disordered 3D point sets (PyTorch, C++).

    Python

  5. yolov5-parallelyolov5-parallelPublic

    Multi-GPU parallelized YOLOv5 training and inference pipeline — data-parallel scaling and throughput tuning.

    Python 2

  6. wu840407.github.iowu840407.github.ioPublic

    Personal site — projects, writing, and contact.

    HTML