Skip to content
View liyanboSustech's full-sized avatar
  • Shenzhen Guangdong, China
  • 14:58 (UTC +08:00)

Block or report liyanboSustech

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Popular repositories Loading

  1. Diff-cache Diff-cachePublic

    Forked from xdit-project/xDiT

    xDiT: A Scalable Inference Engine for Diffusion Transformers (DiTs) with Massive Parallelism

    Python 1

  2. llama.cpp llama.cppPublic

    Forked from ggml-org/llama.cpp

    LLM inference in C/C++

    C++

  3. tensorrtx tensorrtxPublic

    Forked from wang-xinyu/tensorrtx

    Implementation of popular deep learning networks with TensorRT network definition API

    C++

  4. InfiniGen InfiniGenPublic

    Forked from snu-comparch/InfiniGen

    InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Management (OSDI'24)

    Python

  5. H2O H2OPublic

    Forked from FMInference/H2O

    [NeurIPS'23] H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models.

    Python

  6. prompt-cache prompt-cachePublic

    Forked from yale-sys/prompt-cache

    Modular and structured prompt caching for low-latency LLM inference

    Python