Skip to content

All

    Repositories list

    • Python
      Apache License 2.0
      1000Updated Aug 27, 2026Aug 27, 2026
    • Community maintained hardware plugin for vLLM on Ascend
      C++
      Apache License 2.0
      2.2k000Updated Aug 25, 2026Aug 25, 2026
    • spec-ptc

      Public
      Speculative programmatic tool calling (sPTC) for harnesses like RLM, CodeAct, etc.
      Python
      MIT License
      16000Updated Aug 24, 2026Aug 24, 2026
    • A Distributed Attention Towards Linear Scalability for Ultra-Long Context, Heterogeneous Data Training
      Python
      Apache License 2.0
      68000Updated Aug 24, 2026Aug 24, 2026
    • Code for SOSP paper
      HTML
      MIT License
      1000Updated Aug 12, 2026Aug 12, 2026
    • Python
      0000Updated Aug 10, 2026Aug 10, 2026
    • C++
      MIT License
      2000Updated Jul 30, 2026Jul 30, 2026
    • MoonEP

      Public
      MoonEP: A Perfectly Balanced Expert Parallelism Library via Dynamic Redundant Experts
      Python
      MIT License
      129000Updated Jul 28, 2026Jul 28, 2026
    • [Accepted to SOSP 2026] Fast Deterministic LLM Inference
      Python
      Apache License 2.0
      5000Updated Jul 24, 2026Jul 24, 2026
    • C++
      Apache License 2.0
      1000Updated Jul 24, 2026Jul 24, 2026
    • 26SC-Maze

      Public
      A distributed framework for LLM agents
      Python
      MIT License
      26000Updated Jul 24, 2026Jul 24, 2026
    • UltraEP

      Public
      Production-ready expert load balancing library
      Cuda
      MIT License
      17000Updated Jul 17, 2026Jul 17, 2026
    • High-Throughput Batch Inference
      C++
      Apache License 2.0
      1000Updated Jul 14, 2026Jul 14, 2026
    • labs-molt

      Public
      Python
      Apache License 2.0
      100000Updated Jul 14, 2026Jul 14, 2026
    • Repo for ECHO: Efficient KV Cache Offloading with Lossless Prefetching for Serving Native Sparse Attention LLMs (OSDI'26)
      Python
      6000Updated Jul 12, 2026Jul 12, 2026
    • Repo for ECHO: Efficient KV Cache Offloading with Lossless Prefetching for Serving Native Sparse Attention LLMs (OSDI'26)
      Python
      6100Updated Jul 12, 2026Jul 12, 2026
    • A prototype inference engine implementing StreamEP.
      Python
      Apache License 2.0
      1000Updated Jul 12, 2026Jul 12, 2026
    • [OSDI' 26] Efficient LLM Serving on Commodity GPU Clusters with Data-Reduced Cross-Instance Orchestration
      Python
      Apache License 2.0
      1000Updated Jul 5, 2026Jul 5, 2026
    • An efficient service for transparent GPU multiplexing with VRAM oversubscription
      Rust
      Apache License 2.0
      6000Updated Jun 28, 2026Jun 28, 2026
    • dmuon

      Public
      Python
      Other
      3000Updated Jun 24, 2026Jun 24, 2026
    • TraceLab

      Public
      An open toolkit and public dataset hub for collecting, sanitizing, analyzing, and visualizing agent traces.
      Python
      Apache License 2.0
      17000Updated Jun 21, 2026Jun 21, 2026
    • Python
      3000Updated Jun 14, 2026Jun 14, 2026
    • uw-piper

      Public
      A programmable distributed training system for PyTorch
      Python
      MIT License
      5000Updated Jun 10, 2026Jun 10, 2026
    • Python
      Other
      3100Updated Jun 4, 2026Jun 4, 2026
    • An ultra-fast, distributed Safetensors loader
      C++
      Apache License 2.0
      16000Updated May 27, 2026May 27, 2026
    • BlitzScale Router - Distributed LLM Inference Router (Rust)
      Rust
      1000Updated May 25, 2026May 25, 2026
    • Python
      MIT License
      1000Updated May 17, 2026May 17, 2026
    • OpenTela is a decentralized compute fabric for running machine learning applications.
      Jupyter Notebook
      Apache License 2.0
      12000Updated May 14, 2026May 14, 2026
    • Python
      Apache License 2.0
      9000Updated May 13, 2026May 13, 2026
    • Python
      1000Updated May 12, 2026May 12, 2026
    ProTip! When viewing an organization's repositories, you can use the props. filter to filter by custom property.