#
awq-quantization
Here are 3 public repositories matching this topic...
Embeddable Python Engine for VibeVoice TTS with AWQ-INT4 quantization (50% VRAM, 2x Faster RTF) of VibeVoice Large 7B model
-
Updated
May 7, 2026 - Python
scaleDown AI: Enterprise Model Quantization Platform Slash inference costs by 70%. Deploy LLMs anywhere. A microservices-based orchestration engine for isolating and automating incompatible AI optimization workflows.
-
Updated
Jan 28, 2026 - Python
Add this topic to your repo
To associate your repository with the awq-quantization topic, visit your repo's landing page and select "manage topics."