Skip to content
View yangulei's full-sized avatar
  • Intel
  • Shanghai, China

Block or report yangulei

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Popular repositories Loading

  1. vllm-fork vllm-forkPublic

    Forked from HabanaAI/vllm-fork

    A high-throughput and memory-efficient inference and serving engine for LLMs

    Python 4 2

  2. vllm-xpu-breakdown vllm-xpu-breakdownPublic

    Profile and visualize vLLM inference op dispatch on Intel XPU — supports 65+ models across LLM, VL, diffusion, audio, and embedding architectures

    Python 3 3

  3. vllm-hpu-extension vllm-hpu-extensionPublic

    Forked from HabanaAI/vllm-hpu-extension

    Python 1

  4. vllm-gaudi vllm-gaudiPublic

    Forked from vllm-project/vllm-gaudi

    Community maintained hardware plugin for vLLM on Intel Gaudi

    Python 1

  5. MXNet2Caffe MXNet2CaffePublic

    Forked from GarrickLin/MXNet2Caffe

    Convert MXNet model to Caffe model

    Python

  6. TensorRT TensorRTPublic

    Forked from NVIDIA/TensorRT

    TensorRT is a C++ library for high performance inference on NVIDIA GPUs and deep learning accelerators.

    C++