An extremely simple video meeting, integrated whiteboard, chat and screen sharing
-
Updated
Apr 9, 2021 - Go
An extremely simple video meeting, integrated whiteboard, chat and screen sharing
A from-scratch implementation of Llama-3.2-1B in PyTorch, decode-latency benchmarks on three GPUs (T4, L4, A100), three weight-only quantization methods (RTN, GPTQ, AWQ) measured against both, and a packed int4 format with a fused Triton GEMV so the quantized weights are actually 4 bits in HBM.
To associate your repository with the rtn topic, visit your repo's landing page and select "manage topics."