Repository files navigation

QuEST: Stable Training of LLMs with 1-Bit Weights and Activations

arXivYouTube

image

QuEST is a Quantization-Aware Training (QAT) method that allows stable training of 1-bit Llama-type LLMs, and makes 4-bit training Pareto-optimal.

[UPDATE 30.05.25]: Kernels code updated for ease of installation and benchmarking.

[UPDATE 01.05.25]: QuEST has been accepted to ICML2025!

Quickstart

Create a conda environment and install dependencies (we recommend Python 3.10):

conda create -n env python=3.10
conda activate env
pip install -r requirements.txt

Run a simple training on the SlimPajama 6B dataset:

python ./src/main.py

The above command trains a 213.34M parameters model with the Llama-style architecture. We recommend to use the --compile flag that speeds up training noticeably (up to 20% in our setup).

Abstract

One approach to reducing the massive costs of large language models (LLMs) is the use of quantized or sparse representations for training or deployment. While post-training compression methods are very popular, the question of obtaining even more accurate compressed models by directly training over such representations, i.e., Quantization-Aware Training (QAT), is still open: for example, a recent study (arXiv:2411.04330v2) put the "optimal" bit-width at which models can be trained using QAT, while staying accuracy-competitive with standard FP16/BF16 precision, at 8-bits weights and activations.

We advance this state-of-the-art via a new method called QuEST, which is Pareto-competitive with FP16, i.e., it provides better accuracy at lower model size, while training models with weights and activations in 4-bits or less. Moreover, QuEST allows stable training with 1-bit weights and activations. QuEST achieves this by improving two key aspects of QAT methods: (1) accurate and fast quantization of the (continuous) distributions of weights and activations via Hadamard normalization and MSE-optimal fitting; (2) a new trust gradient estimator based on the idea of explicitly minimizing the error between the noisy gradient computed over quantized states and the "true" (but unknown) full-precision gradient. Experiments on Llama-type architectures show that QuEST induces stable scaling laws across the entire range of hardware-supported precisions, and can be extended to sparse representations. We provide GPU kernel support showing that models produced by QuEST can be executed efficiently. Our code is available at this https URL.

Quantization

See train.sh for an example on how to run quantized training.

INT4 Inference Kernels

We provide Triton/CUDA kernels for INT4 Inference on NVIDIA Ampere GPUs. The code is particularly optimized for RTX 4090.

Install

Dependencies

  • cmake
  • C++ compiler (GCC/clang/...)
  • nvcc

Instructions

See gemm-questREADME.

Testing the Models

This notebook and similar provide a way to test that the trained model generate coherent text and are, indeed, serveable in quantized formats.

HellaSWAG Benchmark

To run the HellaSWAG benchmark, please run the following command:

python ./src/eval_hswag.py --model_name "<MODEL_OUTPUT_NAME>" --ckpts_dir "./ckpts"

Source Code

This repository is based on the epfml/schedules-and-scaling repository for their "Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations" paper. We thank the original creators for making public and open-licenced.

Cite This Work

@misc{panferov2025queststabletrainingllms,
title={QuEST: Stable Training of LLMs with 1-Bit Weights and Activations}, author={Andrei Panferov and Jiale Chen and Soroush Tabesh and Roberto L. Castro and Mahdi Nikdan and Dan Alistarh},
year={2025},
eprint={2502.05003},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2502.05003}, }

About

Work in progress.

Resources

Stars

82 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all \u003cpre\u003e\u003ccode\u003e blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks"); } } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); } })(); (function(){ try { var __m = "github.com"; var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

QuEST: Stable Training of LLMs with 1-Bit Weights and Activations

arXivYouTube

image

QuEST is a Quantization-Aware Training (QAT) method that allows stable training of 1-bit Llama-type LLMs, and makes 4-bit training Pareto-optimal.

[UPDATE 30.05.25]: Kernels code updated for ease of installation and benchmarking.

[UPDATE 01.05.25]: QuEST has been accepted to ICML2025!

Quickstart

Create a conda environment and install dependencies (we recommend Python 3.10):

conda create -n env python=3.10
conda activate env
pip install -r requirements.txt

Run a simple training on the SlimPajama 6B dataset:

python ./src/main.py

The above command trains a 213.34M parameters model with the Llama-style architecture. We recommend to use the --compile flag that speeds up training noticeably (up to 20% in our setup).

Abstract

One approach to reducing the massive costs of large language models (LLMs) is the use of quantized or sparse representations for training or deployment. While post-training compression methods are very popular, the question of obtaining even more accurate compressed models by directly training over such representations, i.e., Quantization-Aware Training (QAT), is still open: for example, a recent study (arXiv:2411.04330v2) put the "optimal" bit-width at which models can be trained using QAT, while staying accuracy-competitive with standard FP16/BF16 precision, at 8-bits weights and activations.

We advance this state-of-the-art via a new method called QuEST, which is Pareto-competitive with FP16, i.e., it provides better accuracy at lower model size, while training models with weights and activations in 4-bits or less. Moreover, QuEST allows stable training with 1-bit weights and activations. QuEST achieves this by improving two key aspects of QAT methods: (1) accurate and fast quantization of the (continuous) distributions of weights and activations via Hadamard normalization and MSE-optimal fitting; (2) a new trust gradient estimator based on the idea of explicitly minimizing the error between the noisy gradient computed over quantized states and the "true" (but unknown) full-precision gradient. Experiments on Llama-type architectures show that QuEST induces stable scaling laws across the entire range of hardware-supported precisions, and can be extended to sparse representations. We provide GPU kernel support showing that models produced by QuEST can be executed efficiently. Our code is available at this https URL.

Quantization

See train.sh for an example on how to run quantized training.

INT4 Inference Kernels

We provide Triton/CUDA kernels for INT4 Inference on NVIDIA Ampere GPUs. The code is particularly optimized for RTX 4090.

Install

Dependencies

  • cmake
  • C++ compiler (GCC/clang/...)
  • nvcc

Instructions

See gemm-questREADME.

Testing the Models

This notebook and similar provide a way to test that the trained model generate coherent text and are, indeed, serveable in quantized formats.

HellaSWAG Benchmark

To run the HellaSWAG benchmark, please run the following command:

python ./src/eval_hswag.py --model_name "<MODEL_OUTPUT_NAME>" --ckpts_dir "./ckpts"

Source Code

This repository is based on the epfml/schedules-and-scaling repository for their "Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations" paper. We thank the original creators for making public and open-licenced.

Cite This Work

@misc{panferov2025queststabletrainingllms,
title={QuEST: Stable Training of LLMs with 1-Bit Weights and Activations}, author={Andrei Panferov and Jiale Chen and Soroush Tabesh and Roberto L. Castro and Mahdi Nikdan and Dan Alistarh},
year={2025},
eprint={2502.05003},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2502.05003}, }

About

Work in progress.

Resources

Stars

82 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

QuEST: Stable Training of LLMs with 1-Bit Weights and Activations

arXivYouTube

image

QuEST is a Quantization-Aware Training (QAT) method that allows stable training of 1-bit Llama-type LLMs, and makes 4-bit training Pareto-optimal.

[UPDATE 30.05.25]: Kernels code updated for ease of installation and benchmarking.

[UPDATE 01.05.25]: QuEST has been accepted to ICML2025!

Quickstart

Create a conda environment and install dependencies (we recommend Python 3.10):

conda create -n env python=3.10
conda activate env
pip install -r requirements.txt

Run a simple training on the SlimPajama 6B dataset:

python ./src/main.py

The above command trains a 213.34M parameters model with the Llama-style architecture. We recommend to use the --compile flag that speeds up training noticeably (up to 20% in our setup).

Abstract

One approach to reducing the massive costs of large language models (LLMs) is the use of quantized or sparse representations for training or deployment. While post-training compression methods are very popular, the question of obtaining even more accurate compressed models by directly training over such representations, i.e., Quantization-Aware Training (QAT), is still open: for example, a recent study (arXiv:2411.04330v2) put the "optimal" bit-width at which models can be trained using QAT, while staying accuracy-competitive with standard FP16/BF16 precision, at 8-bits weights and activations.

We advance this state-of-the-art via a new method called QuEST, which is Pareto-competitive with FP16, i.e., it provides better accuracy at lower model size, while training models with weights and activations in 4-bits or less. Moreover, QuEST allows stable training with 1-bit weights and activations. QuEST achieves this by improving two key aspects of QAT methods: (1) accurate and fast quantization of the (continuous) distributions of weights and activations via Hadamard normalization and MSE-optimal fitting; (2) a new trust gradient estimator based on the idea of explicitly minimizing the error between the noisy gradient computed over quantized states and the "true" (but unknown) full-precision gradient. Experiments on Llama-type architectures show that QuEST induces stable scaling laws across the entire range of hardware-supported precisions, and can be extended to sparse representations. We provide GPU kernel support showing that models produced by QuEST can be executed efficiently. Our code is available at this https URL.

Quantization

See train.sh for an example on how to run quantized training.

INT4 Inference Kernels

We provide Triton/CUDA kernels for INT4 Inference on NVIDIA Ampere GPUs. The code is particularly optimized for RTX 4090.

Install

Dependencies

  • cmake
  • C++ compiler (GCC/clang/...)
  • nvcc

Instructions

See gemm-questREADME.

Testing the Models

This notebook and similar provide a way to test that the trained model generate coherent text and are, indeed, serveable in quantized formats.

HellaSWAG Benchmark

To run the HellaSWAG benchmark, please run the following command:

python ./src/eval_hswag.py --model_name "<MODEL_OUTPUT_NAME>" --ckpts_dir "./ckpts"

Source Code

This repository is based on the epfml/schedules-and-scaling repository for their "Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations" paper. We thank the original creators for making public and open-licenced.

Cite This Work

@misc{panferov2025queststabletrainingllms,
title={QuEST: Stable Training of LLMs with 1-Bit Weights and Activations}, author={Andrei Panferov and Jiale Chen and Soroush Tabesh and Roberto L. Castro and Mahdi Nikdan and Dan Alistarh},
year={2025},
eprint={2502.05003},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2502.05003}, }

About

Work in progress.

Resources

Stars

82 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length \u003e 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

QuEST: Stable Training of LLMs with 1-Bit Weights and Activations

arXivYouTube

image

QuEST is a Quantization-Aware Training (QAT) method that allows stable training of 1-bit Llama-type LLMs, and makes 4-bit training Pareto-optimal.

[UPDATE 30.05.25]: Kernels code updated for ease of installation and benchmarking.

[UPDATE 01.05.25]: QuEST has been accepted to ICML2025!

Quickstart

Create a conda environment and install dependencies (we recommend Python 3.10):

conda create -n env python=3.10
conda activate env
pip install -r requirements.txt

Run a simple training on the SlimPajama 6B dataset:

python ./src/main.py

The above command trains a 213.34M parameters model with the Llama-style architecture. We recommend to use the --compile flag that speeds up training noticeably (up to 20% in our setup).

Abstract

One approach to reducing the massive costs of large language models (LLMs) is the use of quantized or sparse representations for training or deployment. While post-training compression methods are very popular, the question of obtaining even more accurate compressed models by directly training over such representations, i.e., Quantization-Aware Training (QAT), is still open: for example, a recent study (arXiv:2411.04330v2) put the "optimal" bit-width at which models can be trained using QAT, while staying accuracy-competitive with standard FP16/BF16 precision, at 8-bits weights and activations.

We advance this state-of-the-art via a new method called QuEST, which is Pareto-competitive with FP16, i.e., it provides better accuracy at lower model size, while training models with weights and activations in 4-bits or less. Moreover, QuEST allows stable training with 1-bit weights and activations. QuEST achieves this by improving two key aspects of QAT methods: (1) accurate and fast quantization of the (continuous) distributions of weights and activations via Hadamard normalization and MSE-optimal fitting; (2) a new trust gradient estimator based on the idea of explicitly minimizing the error between the noisy gradient computed over quantized states and the "true" (but unknown) full-precision gradient. Experiments on Llama-type architectures show that QuEST induces stable scaling laws across the entire range of hardware-supported precisions, and can be extended to sparse representations. We provide GPU kernel support showing that models produced by QuEST can be executed efficiently. Our code is available at this https URL.

Quantization

See train.sh for an example on how to run quantized training.

INT4 Inference Kernels

We provide Triton/CUDA kernels for INT4 Inference on NVIDIA Ampere GPUs. The code is particularly optimized for RTX 4090.

Install

Dependencies

  • cmake
  • C++ compiler (GCC/clang/...)
  • nvcc

Instructions

See gemm-questREADME.

Testing the Models

This notebook and similar provide a way to test that the trained model generate coherent text and are, indeed, serveable in quantized formats.

HellaSWAG Benchmark

To run the HellaSWAG benchmark, please run the following command:

python ./src/eval_hswag.py --model_name "<MODEL_OUTPUT_NAME>" --ckpts_dir "./ckpts"

Source Code

This repository is based on the epfml/schedules-and-scaling repository for their "Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations" paper. We thank the original creators for making public and open-licenced.

Cite This Work

@misc{panferov2025queststabletrainingllms,
title={QuEST: Stable Training of LLMs with 1-Bit Weights and Activations}, author={Andrei Panferov and Jiale Chen and Soroush Tabesh and Roberto L. Castro and Mahdi Nikdan and Dan Alistarh},
year={2025},
eprint={2502.05003},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2502.05003}, }

About

Work in progress.

Resources

Stars

82 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

QuEST: Stable Training of LLMs with 1-Bit Weights and Activations

arXivYouTube

image

QuEST is a Quantization-Aware Training (QAT) method that allows stable training of 1-bit Llama-type LLMs, and makes 4-bit training Pareto-optimal.

[UPDATE 30.05.25]: Kernels code updated for ease of installation and benchmarking.

[UPDATE 01.05.25]: QuEST has been accepted to ICML2025!

Quickstart

Create a conda environment and install dependencies (we recommend Python 3.10):

conda create -n env python=3.10
conda activate env
pip install -r requirements.txt

Run a simple training on the SlimPajama 6B dataset:

python ./src/main.py

The above command trains a 213.34M parameters model with the Llama-style architecture. We recommend to use the --compile flag that speeds up training noticeably (up to 20% in our setup).

Abstract

One approach to reducing the massive costs of large language models (LLMs) is the use of quantized or sparse representations for training or deployment. While post-training compression methods are very popular, the question of obtaining even more accurate compressed models by directly training over such representations, i.e., Quantization-Aware Training (QAT), is still open: for example, a recent study (arXiv:2411.04330v2) put the "optimal" bit-width at which models can be trained using QAT, while staying accuracy-competitive with standard FP16/BF16 precision, at 8-bits weights and activations.

We advance this state-of-the-art via a new method called QuEST, which is Pareto-competitive with FP16, i.e., it provides better accuracy at lower model size, while training models with weights and activations in 4-bits or less. Moreover, QuEST allows stable training with 1-bit weights and activations. QuEST achieves this by improving two key aspects of QAT methods: (1) accurate and fast quantization of the (continuous) distributions of weights and activations via Hadamard normalization and MSE-optimal fitting; (2) a new trust gradient estimator based on the idea of explicitly minimizing the error between the noisy gradient computed over quantized states and the "true" (but unknown) full-precision gradient. Experiments on Llama-type architectures show that QuEST induces stable scaling laws across the entire range of hardware-supported precisions, and can be extended to sparse representations. We provide GPU kernel support showing that models produced by QuEST can be executed efficiently. Our code is available at this https URL.

Quantization

See train.sh for an example on how to run quantized training.

INT4 Inference Kernels

We provide Triton/CUDA kernels for INT4 Inference on NVIDIA Ampere GPUs. The code is particularly optimized for RTX 4090.

Install

Dependencies

  • cmake
  • C++ compiler (GCC/clang/...)
  • nvcc

Instructions

See gemm-questREADME.

Testing the Models

This notebook and similar provide a way to test that the trained model generate coherent text and are, indeed, serveable in quantized formats.

HellaSWAG Benchmark

To run the HellaSWAG benchmark, please run the following command:

python ./src/eval_hswag.py --model_name "<MODEL_OUTPUT_NAME>" --ckpts_dir "./ckpts"

Source Code

This repository is based on the epfml/schedules-and-scaling repository for their "Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations" paper. We thank the original creators for making public and open-licenced.

Cite This Work

@misc{panferov2025queststabletrainingllms,
title={QuEST: Stable Training of LLMs with 1-Bit Weights and Activations}, author={Andrei Panferov and Jiale Chen and Soroush Tabesh and Roberto L. Castro and Mahdi Nikdan and Dan Alistarh},
year={2025},
eprint={2502.05003},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2502.05003}, }

About

Work in progress.

Resources

Stars

82 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

QuEST: Stable Training of LLMs with 1-Bit Weights and Activations

arXivYouTube

image

QuEST is a Quantization-Aware Training (QAT) method that allows stable training of 1-bit Llama-type LLMs, and makes 4-bit training Pareto-optimal.

[UPDATE 30.05.25]: Kernels code updated for ease of installation and benchmarking.

[UPDATE 01.05.25]: QuEST has been accepted to ICML2025!

Quickstart

Create a conda environment and install dependencies (we recommend Python 3.10):

conda create -n env python=3.10
conda activate env
pip install -r requirements.txt

Run a simple training on the SlimPajama 6B dataset:

python ./src/main.py

The above command trains a 213.34M parameters model with the Llama-style architecture. We recommend to use the --compile flag that speeds up training noticeably (up to 20% in our setup).

Abstract

One approach to reducing the massive costs of large language models (LLMs) is the use of quantized or sparse representations for training or deployment. While post-training compression methods are very popular, the question of obtaining even more accurate compressed models by directly training over such representations, i.e., Quantization-Aware Training (QAT), is still open: for example, a recent study (arXiv:2411.04330v2) put the "optimal" bit-width at which models can be trained using QAT, while staying accuracy-competitive with standard FP16/BF16 precision, at 8-bits weights and activations.

We advance this state-of-the-art via a new method called QuEST, which is Pareto-competitive with FP16, i.e., it provides better accuracy at lower model size, while training models with weights and activations in 4-bits or less. Moreover, QuEST allows stable training with 1-bit weights and activations. QuEST achieves this by improving two key aspects of QAT methods: (1) accurate and fast quantization of the (continuous) distributions of weights and activations via Hadamard normalization and MSE-optimal fitting; (2) a new trust gradient estimator based on the idea of explicitly minimizing the error between the noisy gradient computed over quantized states and the "true" (but unknown) full-precision gradient. Experiments on Llama-type architectures show that QuEST induces stable scaling laws across the entire range of hardware-supported precisions, and can be extended to sparse representations. We provide GPU kernel support showing that models produced by QuEST can be executed efficiently. Our code is available at this https URL.

Quantization

See train.sh for an example on how to run quantized training.

INT4 Inference Kernels

We provide Triton/CUDA kernels for INT4 Inference on NVIDIA Ampere GPUs. The code is particularly optimized for RTX 4090.

Install

Dependencies

  • cmake
  • C++ compiler (GCC/clang/...)
  • nvcc

Instructions

See gemm-questREADME.

Testing the Models

This notebook and similar provide a way to test that the trained model generate coherent text and are, indeed, serveable in quantized formats.

HellaSWAG Benchmark

To run the HellaSWAG benchmark, please run the following command:

python ./src/eval_hswag.py --model_name "<MODEL_OUTPUT_NAME>" --ckpts_dir "./ckpts"

Source Code

This repository is based on the epfml/schedules-and-scaling repository for their "Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations" paper. We thank the original creators for making public and open-licenced.

Cite This Work

@misc{panferov2025queststabletrainingllms,
title={QuEST: Stable Training of LLMs with 1-Bit Weights and Activations}, author={Andrei Panferov and Jiale Chen and Soroush Tabesh and Roberto L. Castro and Mahdi Nikdan and Dan Alistarh},
year={2025},
eprint={2502.05003},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2502.05003}, }

About

Work in progress.

Resources

Stars

82 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

QuEST: Stable Training of LLMs with 1-Bit Weights and Activations

arXivYouTube

image

QuEST is a Quantization-Aware Training (QAT) method that allows stable training of 1-bit Llama-type LLMs, and makes 4-bit training Pareto-optimal.

[UPDATE 30.05.25]: Kernels code updated for ease of installation and benchmarking.

[UPDATE 01.05.25]: QuEST has been accepted to ICML2025!

Quickstart

Create a conda environment and install dependencies (we recommend Python 3.10):

conda create -n env python=3.10
conda activate env
pip install -r requirements.txt

Run a simple training on the SlimPajama 6B dataset:

python ./src/main.py

The above command trains a 213.34M parameters model with the Llama-style architecture. We recommend to use the --compile flag that speeds up training noticeably (up to 20% in our setup).

Abstract

One approach to reducing the massive costs of large language models (LLMs) is the use of quantized or sparse representations for training or deployment. While post-training compression methods are very popular, the question of obtaining even more accurate compressed models by directly training over such representations, i.e., Quantization-Aware Training (QAT), is still open: for example, a recent study (arXiv:2411.04330v2) put the "optimal" bit-width at which models can be trained using QAT, while staying accuracy-competitive with standard FP16/BF16 precision, at 8-bits weights and activations.

We advance this state-of-the-art via a new method called QuEST, which is Pareto-competitive with FP16, i.e., it provides better accuracy at lower model size, while training models with weights and activations in 4-bits or less. Moreover, QuEST allows stable training with 1-bit weights and activations. QuEST achieves this by improving two key aspects of QAT methods: (1) accurate and fast quantization of the (continuous) distributions of weights and activations via Hadamard normalization and MSE-optimal fitting; (2) a new trust gradient estimator based on the idea of explicitly minimizing the error between the noisy gradient computed over quantized states and the "true" (but unknown) full-precision gradient. Experiments on Llama-type architectures show that QuEST induces stable scaling laws across the entire range of hardware-supported precisions, and can be extended to sparse representations. We provide GPU kernel support showing that models produced by QuEST can be executed efficiently. Our code is available at this https URL.

Quantization

See train.sh for an example on how to run quantized training.

INT4 Inference Kernels

We provide Triton/CUDA kernels for INT4 Inference on NVIDIA Ampere GPUs. The code is particularly optimized for RTX 4090.

Install

Dependencies

  • cmake
  • C++ compiler (GCC/clang/...)
  • nvcc

Instructions

See gemm-questREADME.

Testing the Models

This notebook and similar provide a way to test that the trained model generate coherent text and are, indeed, serveable in quantized formats.

HellaSWAG Benchmark

To run the HellaSWAG benchmark, please run the following command:

python ./src/eval_hswag.py --model_name "<MODEL_OUTPUT_NAME>" --ckpts_dir "./ckpts"

Source Code

This repository is based on the epfml/schedules-and-scaling repository for their "Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations" paper. We thank the original creators for making public and open-licenced.

Cite This Work

@misc{panferov2025queststabletrainingllms,
title={QuEST: Stable Training of LLMs with 1-Bit Weights and Activations}, author={Andrei Panferov and Jiale Chen and Soroush Tabesh and Roberto L. Castro and Mahdi Nikdan and Dan Alistarh},
year={2025},
eprint={2502.05003},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2502.05003}, }

About

Work in progress.

Resources

Stars

82 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

QuEST: Stable Training of LLMs with 1-Bit Weights and Activations

arXivYouTube

image

QuEST is a Quantization-Aware Training (QAT) method that allows stable training of 1-bit Llama-type LLMs, and makes 4-bit training Pareto-optimal.

[UPDATE 30.05.25]: Kernels code updated for ease of installation and benchmarking.

[UPDATE 01.05.25]: QuEST has been accepted to ICML2025!

Quickstart

Create a conda environment and install dependencies (we recommend Python 3.10):

conda create -n env python=3.10
conda activate env
pip install -r requirements.txt

Run a simple training on the SlimPajama 6B dataset:

python ./src/main.py

The above command trains a 213.34M parameters model with the Llama-style architecture. We recommend to use the --compile flag that speeds up training noticeably (up to 20% in our setup).

Abstract

One approach to reducing the massive costs of large language models (LLMs) is the use of quantized or sparse representations for training or deployment. While post-training compression methods are very popular, the question of obtaining even more accurate compressed models by directly training over such representations, i.e., Quantization-Aware Training (QAT), is still open: for example, a recent study (arXiv:2411.04330v2) put the "optimal" bit-width at which models can be trained using QAT, while staying accuracy-competitive with standard FP16/BF16 precision, at 8-bits weights and activations.

We advance this state-of-the-art via a new method called QuEST, which is Pareto-competitive with FP16, i.e., it provides better accuracy at lower model size, while training models with weights and activations in 4-bits or less. Moreover, QuEST allows stable training with 1-bit weights and activations. QuEST achieves this by improving two key aspects of QAT methods: (1) accurate and fast quantization of the (continuous) distributions of weights and activations via Hadamard normalization and MSE-optimal fitting; (2) a new trust gradient estimator based on the idea of explicitly minimizing the error between the noisy gradient computed over quantized states and the "true" (but unknown) full-precision gradient. Experiments on Llama-type architectures show that QuEST induces stable scaling laws across the entire range of hardware-supported precisions, and can be extended to sparse representations. We provide GPU kernel support showing that models produced by QuEST can be executed efficiently. Our code is available at this https URL.

Quantization

See train.sh for an example on how to run quantized training.

INT4 Inference Kernels

We provide Triton/CUDA kernels for INT4 Inference on NVIDIA Ampere GPUs. The code is particularly optimized for RTX 4090.

Install

Dependencies

  • cmake
  • C++ compiler (GCC/clang/...)
  • nvcc

Instructions

See gemm-questREADME.

Testing the Models

This notebook and similar provide a way to test that the trained model generate coherent text and are, indeed, serveable in quantized formats.

HellaSWAG Benchmark

To run the HellaSWAG benchmark, please run the following command:

python ./src/eval_hswag.py --model_name "<MODEL_OUTPUT_NAME>" --ckpts_dir "./ckpts"

Source Code

This repository is based on the epfml/schedules-and-scaling repository for their "Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations" paper. We thank the original creators for making public and open-licenced.

Cite This Work

@misc{panferov2025queststabletrainingllms,
title={QuEST: Stable Training of LLMs with 1-Bit Weights and Activations}, author={Andrei Panferov and Jiale Chen and Soroush Tabesh and Roberto L. Castro and Mahdi Nikdan and Dan Alistarh},
year={2025},
eprint={2502.05003},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2502.05003}, }

About

Work in progress.

Resources

Stars

82 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages