Skip to content
#

int4-quantization

Here are 5 public repositories matching this topic...

Language:Python
Filter by language

Three hand-written Triton kernels for LLM inference (fused RMSNorm plus residual, online softmax, INT4 g128 GEMV) benchmarked on NVIDIA Blackwell against PyTorch eager and torch.compile, with every raw CUDA-event sample, measured device ceiling, and Nsight Compute report committed and CI-verified.

  • Updated Aug 12, 2026
  • Python

Add this topic to your repo

To associate your repository with the int4-quantization topic, visit your repo's landing page and select "manage topics."

Learn more