Repository files navigation

vit.cpp

Inference Vision Transformer (ViT) in plain C/C++ using ggml without any extra dependencies

Description

This project presents a standalone implementation of the well known Vision Transformer (ViT) model family, used in a broad spectrum of applications and SOTA models like Large Multimodal Models(LMM). The primary goal is to develop a C/C++ inference engine tailored for ViT models, utilizing ggml to enhance performance, particularly on edge devices. Designed to be both lightweight and self-contained, this implementation can be run across diverse platforms.

Table of Contents
  1. Description
  2. Features
  3. Vision Transformer Architecture
  4. Quick Example
  5. Convert PyTorch to GGUF
  6. Build
  7. Run
  8. Benchmark against PyTorch
  9. Quantization
  10. To-Do List

Features

  • Dependency-free and lightweight inference thanks to ggml.
  • 4-bit, 5-bit and 8-bit quantization support.
  • Support for timm ViTs with different variants out of the box.

An important aspect of using vit.cpp is that it has short startup times compared to common DL frameworks, which makes it suitable for serverless deployments where the cold start is an issue.

Vision Transformer architecture

The implemented architecture is based on the original Vision Transformer from:

Vision Transformer overview

ViT architecture. Taken from the original paper.

Quick example

Details

example input

See output
 $ ./bin/vit -t 4 -m ../ggml-model-f16.gguf -i ../assets/magpie.jpeg -k 5
main: seed = 1701176263
main: n_threads = 4 / 8
vit_model_load: loading model from '../ggml-model-f16.gguf' - please wait
vit_model_load: hidden_size = 192
vit_model_load: num_hidden_layers = 12
vit_model_load: num_attention_heads = 3
vit_model_load: patch_size = 16
vit_model_load: img_size = 224
vit_model_load: num_classes = 1000
vit_model_load: ftype = 1
vit_model_load: qntvr = 0
operator(): ggml ctx size = 11.13 MB
vit_model_load: ................... done
vit_model_load: model size = 11.04 MB / num tensors = 152
main: loaded image '../assets/magpie.jpeg' (500 x 470)
vit_image_preprocess: scale = 2.232143
processed, out dims : (224 x 224)

> magpie : 0.87 > goose : 0.02 > toucan : 0.01 > drake : 0.01 > king penguin, Aptenodytes patagonica : 0.01

main: model load time = 17.92 ms main: processing time = 146.96 ms main: total time = 164.88 ms

Convert PyTorch to GGUF

# clone the repo recursively
git clone --recurse-submodules https://github.com/staghado/vit.cpp.git
cd vit.cpp
# install torch and timm
pip install torch timm
# list available models if needed; note that not all models are supported
python convert-pth-to-ggml.py --list
# convert the weights to gguf : vit tiny with patch size of 16 and an image # size of 384 pre-trained on ImageNet21k and fine-tuned on ImageNet1k
python convert-pth-to-ggml.py --model_name vit_tiny_patch16_384.augreg_in21k_ft_in1k --ftype 1

Note: You can also download the converted weights from Hugging Face directly.

wget https://huggingface.co/staghado/vit.cpp/blob/main/tiny-ggml-model-f16.gguf

Build

Simple build

# build ggml and vit mkdir build && cd build
cmake .. && make -j4
# run inference
./bin/vit -t 4 -m ../ggml-model-f16.gguf -i ../assets/tench.jpg

The optimal number of threads to use depends on many factors and more is not always better. Usually using a number of threads equal to the number of available physical cores gives the best performance in terms of speed.

Per device optimizations

Generate per-device instructions that work best for the given machine rather than using general CPU instructions. This can be done by specifying -march=native in the compiler flags.

  • Multi-threading and vectorization
  • Loop transformations(unrolling)

For AMD host processors

You can use a specialized compiler released by AMD to make full use of your specific processor's architecture. Read more here : AMD Optimizing C/C++ and Fortran Compilers (AOCC)

You can follow the given instructions to install the AOCC compiler.

Note : For my AMD Ryzen 7 3700U, the improvements were not very significant but for more recent processors there could be a gain in using a specialized compiler.

Using OpenMP

Additionally compile with OpenMP by specifying the -fopenmp flag to the compiler in the CMakeLists file, allowing multithreaded runs. Make sure to also enable multiple threads when running, e.g.:

OMP_NUM_THREADS=4 ./bin/vit -t 4 -m ../ggml-model-f16.bin -i ../assets/tench.jpg

Run

usage: ./bin/vit [options]
options:
-h, --help show this help message and exit
-s SEED, --seed SEED RNG seed (default: -1)
-t N, --threads N number of threads to use during computation (default: 4)
-m FNAME, --model FNAME model path (default: ../ggml-model-f16.bin)
-i FNAME, --inp FNAME input file (default: ../assets/tench.jpg)
-k N, --topk N top k classes to print (default: 5)
-e FLOAT, --epsilon epsilon (default: 0.000001)

Benchmark against PyTorch

First experiments on Apple M1 show inference speedups(up to 6x faster for base model) compared to native PyTorch inference.

ViT inference

You can efficiently run ViT inference on the CPU. Memory requirements and inference speed on AMD Ryzen 7 3700U(4 cores, 8 threads) for both native PyTorch and vit.cpp. Using 4 threads gives better results for my machine. The reported results of inference speed correspond to 10 runs averages for both PyTorch and vit.cpp.

ModelMax Mem(PyTorch)Max MemSpeed(PyTorch)Speed
tiny~780 MB~20 MB431 ms120 ms
small~965 MB~52 MB780 ms463 ms
base~1.61 GB~179 MB2393 ms1441 ms
large~3.86 GB~597 MB8151 ms4892 ms

Note: The models used are of the form vit_{size}_patch16_224.augreg_in21k_ft_in1k.

Benchmark on your machine

In order to test the inference speed on your machine, you can run the following scripts:

chmod +x scripts/benchmark.*
# install memory_profiler & threadpoolctl
pip install memory_profiler threadpoolctl
# run the benchmark of PyTorch
python scripts/benchmark.py
# run the benchmark of vit.cpp for non-qunatized model
./scripts/benchmark.sh
# to run the benchamrk for qunatized models; 4 threads and quantize flag
./scripts/benchmark.sh 4 1

Both scripts use 4 threads by default. In Python, the threadpoolctl library is used to limit the number of threads used by PyTorch.

Quantization

vit.cpp supports many quantization strategies from ggml such as q4_0, q4_1, q5_0, q5_1 and q8_0 types. You can quantize a model in F32 (the patch embedding is in F16) to one of these types by using the ./bin/quantize binary.

usage: ./bin/quantize /path/to/ggml-model-f32.gguf /path/to/ggml-model-quantized.gguf type type = 2 - q4_0 type = 3 - q4_1 type = 6 - q5_0 type = 7 - q5_1 type = 8 - q8_0 

For example, you can run the following to convert the model to q5_1:

./bin/quantize ../tiny-ggml-model-f16.gguf ../tiny-ggml-model-f16-quant.gguf 7

Then you can use tiny-ggml-model-f16-quant.gguf just like the model in F16.

Results

Here are the benchmarks for the different models and quantizations on my machine: For accurate estimation of run times, these benchmarks were run 100 times each.

ModelQuantizationSpeed (ms)Mem (MB)
tinyq4_0105 ms12 MB
tinyq4_197 ms12 MB
tinyq5_0116 ms13 MB
tinyq5_1112 ms13 MB
tinyq8_090 ms15 MB
smallq4_0240 ms23 MB
smallq4_1224 ms24 MB
smallq5_0288 ms25 MB
smallq5_1277 ms27 MB
smallq8_0228 ms33 MB
baseq4_0704 ms61 MB
baseq4_1626 ms66 MB
baseq5_0851 ms71 MB
baseq5_1806 ms76 MB
baseq8_0659 ms102 MB
largeq4_02189 ms181 MB
largeq4_11919 ms199 MB
largeq5_02676 ms217 MB
largeq5_12547 ms235 MB
largeq8_01994 ms325 MB

To-Do List

  • Evaluate performance on ImageNet1k:

    Run evaluation on ImageNet1k test set and analyze the performance of different quantization schemes.

This project was highly inspired by the following projects:

Star History

Star History Chart

About

Inference Vision Transformer (ViT) in plain C/C++ with ggml

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

vit.cpp

Inference Vision Transformer (ViT) in plain C/C++ using ggml without any extra dependencies

Description

This project presents a standalone implementation of the well known Vision Transformer (ViT) model family, used in a broad spectrum of applications and SOTA models like Large Multimodal Models(LMM). The primary goal is to develop a C/C++ inference engine tailored for ViT models, utilizing ggml to enhance performance, particularly on edge devices. Designed to be both lightweight and self-contained, this implementation can be run across diverse platforms.

Table of Contents
  1. Description
  2. Features
  3. Vision Transformer Architecture
  4. Quick Example
  5. Convert PyTorch to GGUF
  6. Build
  7. Run
  8. Benchmark against PyTorch
  9. Quantization
  10. To-Do List

Features

  • Dependency-free and lightweight inference thanks to ggml.
  • 4-bit, 5-bit and 8-bit quantization support.
  • Support for timm ViTs with different variants out of the box.

An important aspect of using vit.cpp is that it has short startup times compared to common DL frameworks, which makes it suitable for serverless deployments where the cold start is an issue.

Vision Transformer architecture

The implemented architecture is based on the original Vision Transformer from:

Vision Transformer overview

ViT architecture. Taken from the original paper.

Quick example

Details

example input

See output
 $ ./bin/vit -t 4 -m ../ggml-model-f16.gguf -i ../assets/magpie.jpeg -k 5
main: seed = 1701176263
main: n_threads = 4 / 8
vit_model_load: loading model from '../ggml-model-f16.gguf' - please wait
vit_model_load: hidden_size = 192
vit_model_load: num_hidden_layers = 12
vit_model_load: num_attention_heads = 3
vit_model_load: patch_size = 16
vit_model_load: img_size = 224
vit_model_load: num_classes = 1000
vit_model_load: ftype = 1
vit_model_load: qntvr = 0
operator(): ggml ctx size = 11.13 MB
vit_model_load: ................... done
vit_model_load: model size = 11.04 MB / num tensors = 152
main: loaded image '../assets/magpie.jpeg' (500 x 470)
vit_image_preprocess: scale = 2.232143
processed, out dims : (224 x 224)

> magpie : 0.87 > goose : 0.02 > toucan : 0.01 > drake : 0.01 > king penguin, Aptenodytes patagonica : 0.01

main: model load time = 17.92 ms main: processing time = 146.96 ms main: total time = 164.88 ms

Convert PyTorch to GGUF

# clone the repo recursively
git clone --recurse-submodules https://github.com/staghado/vit.cpp.git
cd vit.cpp
# install torch and timm
pip install torch timm
# list available models if needed; note that not all models are supported
python convert-pth-to-ggml.py --list
# convert the weights to gguf : vit tiny with patch size of 16 and an image # size of 384 pre-trained on ImageNet21k and fine-tuned on ImageNet1k
python convert-pth-to-ggml.py --model_name vit_tiny_patch16_384.augreg_in21k_ft_in1k --ftype 1

Note: You can also download the converted weights from Hugging Face directly.

wget https://huggingface.co/staghado/vit.cpp/blob/main/tiny-ggml-model-f16.gguf

Build

Simple build

# build ggml and vit mkdir build && cd build
cmake .. && make -j4
# run inference
./bin/vit -t 4 -m ../ggml-model-f16.gguf -i ../assets/tench.jpg

The optimal number of threads to use depends on many factors and more is not always better. Usually using a number of threads equal to the number of available physical cores gives the best performance in terms of speed.

Per device optimizations

Generate per-device instructions that work best for the given machine rather than using general CPU instructions. This can be done by specifying -march=native in the compiler flags.

  • Multi-threading and vectorization
  • Loop transformations(unrolling)

For AMD host processors

You can use a specialized compiler released by AMD to make full use of your specific processor's architecture. Read more here : AMD Optimizing C/C++ and Fortran Compilers (AOCC)

You can follow the given instructions to install the AOCC compiler.

Note : For my AMD Ryzen 7 3700U, the improvements were not very significant but for more recent processors there could be a gain in using a specialized compiler.

Using OpenMP

Additionally compile with OpenMP by specifying the -fopenmp flag to the compiler in the CMakeLists file, allowing multithreaded runs. Make sure to also enable multiple threads when running, e.g.:

OMP_NUM_THREADS=4 ./bin/vit -t 4 -m ../ggml-model-f16.bin -i ../assets/tench.jpg

Run

usage: ./bin/vit [options]
options:
-h, --help show this help message and exit
-s SEED, --seed SEED RNG seed (default: -1)
-t N, --threads N number of threads to use during computation (default: 4)
-m FNAME, --model FNAME model path (default: ../ggml-model-f16.bin)
-i FNAME, --inp FNAME input file (default: ../assets/tench.jpg)
-k N, --topk N top k classes to print (default: 5)
-e FLOAT, --epsilon epsilon (default: 0.000001)

Benchmark against PyTorch

First experiments on Apple M1 show inference speedups(up to 6x faster for base model) compared to native PyTorch inference.

ViT inference

You can efficiently run ViT inference on the CPU. Memory requirements and inference speed on AMD Ryzen 7 3700U(4 cores, 8 threads) for both native PyTorch and vit.cpp. Using 4 threads gives better results for my machine. The reported results of inference speed correspond to 10 runs averages for both PyTorch and vit.cpp.

ModelMax Mem(PyTorch)Max MemSpeed(PyTorch)Speed
tiny~780 MB~20 MB431 ms120 ms
small~965 MB~52 MB780 ms463 ms
base~1.61 GB~179 MB2393 ms1441 ms
large~3.86 GB~597 MB8151 ms4892 ms

Note: The models used are of the form vit_{size}_patch16_224.augreg_in21k_ft_in1k.

Benchmark on your machine

In order to test the inference speed on your machine, you can run the following scripts:

chmod +x scripts/benchmark.*
# install memory_profiler & threadpoolctl
pip install memory_profiler threadpoolctl
# run the benchmark of PyTorch
python scripts/benchmark.py
# run the benchmark of vit.cpp for non-qunatized model
./scripts/benchmark.sh
# to run the benchamrk for qunatized models; 4 threads and quantize flag
./scripts/benchmark.sh 4 1

Both scripts use 4 threads by default. In Python, the threadpoolctl library is used to limit the number of threads used by PyTorch.

Quantization

vit.cpp supports many quantization strategies from ggml such as q4_0, q4_1, q5_0, q5_1 and q8_0 types. You can quantize a model in F32 (the patch embedding is in F16) to one of these types by using the ./bin/quantize binary.

usage: ./bin/quantize /path/to/ggml-model-f32.gguf /path/to/ggml-model-quantized.gguf type type = 2 - q4_0 type = 3 - q4_1 type = 6 - q5_0 type = 7 - q5_1 type = 8 - q8_0 

For example, you can run the following to convert the model to q5_1:

./bin/quantize ../tiny-ggml-model-f16.gguf ../tiny-ggml-model-f16-quant.gguf 7

Then you can use tiny-ggml-model-f16-quant.gguf just like the model in F16.

Results

Here are the benchmarks for the different models and quantizations on my machine: For accurate estimation of run times, these benchmarks were run 100 times each.

ModelQuantizationSpeed (ms)Mem (MB)
tinyq4_0105 ms12 MB
tinyq4_197 ms12 MB
tinyq5_0116 ms13 MB
tinyq5_1112 ms13 MB
tinyq8_090 ms15 MB
smallq4_0240 ms23 MB
smallq4_1224 ms24 MB
smallq5_0288 ms25 MB
smallq5_1277 ms27 MB
smallq8_0228 ms33 MB
baseq4_0704 ms61 MB
baseq4_1626 ms66 MB
baseq5_0851 ms71 MB
baseq5_1806 ms76 MB
baseq8_0659 ms102 MB
largeq4_02189 ms181 MB
largeq4_11919 ms199 MB
largeq5_02676 ms217 MB
largeq5_12547 ms235 MB
largeq8_01994 ms325 MB

To-Do List

  • Evaluate performance on ImageNet1k:

    Run evaluation on ImageNet1k test set and analyze the performance of different quantization schemes.

This project was highly inspired by the following projects:

Star History

Star History Chart

About

Inference Vision Transformer (ViT) in plain C/C++ with ggml

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

vit.cpp

Inference Vision Transformer (ViT) in plain C/C++ using ggml without any extra dependencies

Description

This project presents a standalone implementation of the well known Vision Transformer (ViT) model family, used in a broad spectrum of applications and SOTA models like Large Multimodal Models(LMM). The primary goal is to develop a C/C++ inference engine tailored for ViT models, utilizing ggml to enhance performance, particularly on edge devices. Designed to be both lightweight and self-contained, this implementation can be run across diverse platforms.

Table of Contents
  1. Description
  2. Features
  3. Vision Transformer Architecture
  4. Quick Example
  5. Convert PyTorch to GGUF
  6. Build
  7. Run
  8. Benchmark against PyTorch
  9. Quantization
  10. To-Do List

Features

  • Dependency-free and lightweight inference thanks to ggml.
  • 4-bit, 5-bit and 8-bit quantization support.
  • Support for timm ViTs with different variants out of the box.

An important aspect of using vit.cpp is that it has short startup times compared to common DL frameworks, which makes it suitable for serverless deployments where the cold start is an issue.

Vision Transformer architecture

The implemented architecture is based on the original Vision Transformer from:

Vision Transformer overview

ViT architecture. Taken from the original paper.

Quick example

Details

example input

See output
 $ ./bin/vit -t 4 -m ../ggml-model-f16.gguf -i ../assets/magpie.jpeg -k 5
main: seed = 1701176263
main: n_threads = 4 / 8
vit_model_load: loading model from '../ggml-model-f16.gguf' - please wait
vit_model_load: hidden_size = 192
vit_model_load: num_hidden_layers = 12
vit_model_load: num_attention_heads = 3
vit_model_load: patch_size = 16
vit_model_load: img_size = 224
vit_model_load: num_classes = 1000
vit_model_load: ftype = 1
vit_model_load: qntvr = 0
operator(): ggml ctx size = 11.13 MB
vit_model_load: ................... done
vit_model_load: model size = 11.04 MB / num tensors = 152
main: loaded image '../assets/magpie.jpeg' (500 x 470)
vit_image_preprocess: scale = 2.232143
processed, out dims : (224 x 224)

> magpie : 0.87 > goose : 0.02 > toucan : 0.01 > drake : 0.01 > king penguin, Aptenodytes patagonica : 0.01

main: model load time = 17.92 ms main: processing time = 146.96 ms main: total time = 164.88 ms

Convert PyTorch to GGUF

# clone the repo recursively
git clone --recurse-submodules https://github.com/staghado/vit.cpp.git
cd vit.cpp
# install torch and timm
pip install torch timm
# list available models if needed; note that not all models are supported
python convert-pth-to-ggml.py --list
# convert the weights to gguf : vit tiny with patch size of 16 and an image # size of 384 pre-trained on ImageNet21k and fine-tuned on ImageNet1k
python convert-pth-to-ggml.py --model_name vit_tiny_patch16_384.augreg_in21k_ft_in1k --ftype 1

Note: You can also download the converted weights from Hugging Face directly.

wget https://huggingface.co/staghado/vit.cpp/blob/main/tiny-ggml-model-f16.gguf

Build

Simple build

# build ggml and vit mkdir build && cd build
cmake .. && make -j4
# run inference
./bin/vit -t 4 -m ../ggml-model-f16.gguf -i ../assets/tench.jpg

The optimal number of threads to use depends on many factors and more is not always better. Usually using a number of threads equal to the number of available physical cores gives the best performance in terms of speed.

Per device optimizations

Generate per-device instructions that work best for the given machine rather than using general CPU instructions. This can be done by specifying -march=native in the compiler flags.

  • Multi-threading and vectorization
  • Loop transformations(unrolling)

For AMD host processors

You can use a specialized compiler released by AMD to make full use of your specific processor's architecture. Read more here : AMD Optimizing C/C++ and Fortran Compilers (AOCC)

You can follow the given instructions to install the AOCC compiler.

Note : For my AMD Ryzen 7 3700U, the improvements were not very significant but for more recent processors there could be a gain in using a specialized compiler.

Using OpenMP

Additionally compile with OpenMP by specifying the -fopenmp flag to the compiler in the CMakeLists file, allowing multithreaded runs. Make sure to also enable multiple threads when running, e.g.:

OMP_NUM_THREADS=4 ./bin/vit -t 4 -m ../ggml-model-f16.bin -i ../assets/tench.jpg

Run

usage: ./bin/vit [options]
options:
-h, --help show this help message and exit
-s SEED, --seed SEED RNG seed (default: -1)
-t N, --threads N number of threads to use during computation (default: 4)
-m FNAME, --model FNAME model path (default: ../ggml-model-f16.bin)
-i FNAME, --inp FNAME input file (default: ../assets/tench.jpg)
-k N, --topk N top k classes to print (default: 5)
-e FLOAT, --epsilon epsilon (default: 0.000001)

Benchmark against PyTorch

First experiments on Apple M1 show inference speedups(up to 6x faster for base model) compared to native PyTorch inference.

ViT inference

You can efficiently run ViT inference on the CPU. Memory requirements and inference speed on AMD Ryzen 7 3700U(4 cores, 8 threads) for both native PyTorch and vit.cpp. Using 4 threads gives better results for my machine. The reported results of inference speed correspond to 10 runs averages for both PyTorch and vit.cpp.

ModelMax Mem(PyTorch)Max MemSpeed(PyTorch)Speed
tiny~780 MB~20 MB431 ms120 ms
small~965 MB~52 MB780 ms463 ms
base~1.61 GB~179 MB2393 ms1441 ms
large~3.86 GB~597 MB8151 ms4892 ms

Note: The models used are of the form vit_{size}_patch16_224.augreg_in21k_ft_in1k.

Benchmark on your machine

In order to test the inference speed on your machine, you can run the following scripts:

chmod +x scripts/benchmark.*
# install memory_profiler & threadpoolctl
pip install memory_profiler threadpoolctl
# run the benchmark of PyTorch
python scripts/benchmark.py
# run the benchmark of vit.cpp for non-qunatized model
./scripts/benchmark.sh
# to run the benchamrk for qunatized models; 4 threads and quantize flag
./scripts/benchmark.sh 4 1

Both scripts use 4 threads by default. In Python, the threadpoolctl library is used to limit the number of threads used by PyTorch.

Quantization

vit.cpp supports many quantization strategies from ggml such as q4_0, q4_1, q5_0, q5_1 and q8_0 types. You can quantize a model in F32 (the patch embedding is in F16) to one of these types by using the ./bin/quantize binary.

usage: ./bin/quantize /path/to/ggml-model-f32.gguf /path/to/ggml-model-quantized.gguf type type = 2 - q4_0 type = 3 - q4_1 type = 6 - q5_0 type = 7 - q5_1 type = 8 - q8_0 

For example, you can run the following to convert the model to q5_1:

./bin/quantize ../tiny-ggml-model-f16.gguf ../tiny-ggml-model-f16-quant.gguf 7

Then you can use tiny-ggml-model-f16-quant.gguf just like the model in F16.

Results

Here are the benchmarks for the different models and quantizations on my machine: For accurate estimation of run times, these benchmarks were run 100 times each.

ModelQuantizationSpeed (ms)Mem (MB)
tinyq4_0105 ms12 MB
tinyq4_197 ms12 MB
tinyq5_0116 ms13 MB
tinyq5_1112 ms13 MB
tinyq8_090 ms15 MB
smallq4_0240 ms23 MB
smallq4_1224 ms24 MB
smallq5_0288 ms25 MB
smallq5_1277 ms27 MB
smallq8_0228 ms33 MB
baseq4_0704 ms61 MB
baseq4_1626 ms66 MB
baseq5_0851 ms71 MB
baseq5_1806 ms76 MB
baseq8_0659 ms102 MB
largeq4_02189 ms181 MB
largeq4_11919 ms199 MB
largeq5_02676 ms217 MB
largeq5_12547 ms235 MB
largeq8_01994 ms325 MB

To-Do List

  • Evaluate performance on ImageNet1k:

    Run evaluation on ImageNet1k test set and analyze the performance of different quantization schemes.

This project was highly inspired by the following projects:

Star History

Star History Chart

About

Inference Vision Transformer (ViT) in plain C/C++ with ggml

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

vit.cpp

Inference Vision Transformer (ViT) in plain C/C++ using ggml without any extra dependencies

Description

This project presents a standalone implementation of the well known Vision Transformer (ViT) model family, used in a broad spectrum of applications and SOTA models like Large Multimodal Models(LMM). The primary goal is to develop a C/C++ inference engine tailored for ViT models, utilizing ggml to enhance performance, particularly on edge devices. Designed to be both lightweight and self-contained, this implementation can be run across diverse platforms.

Table of Contents
  1. Description
  2. Features
  3. Vision Transformer Architecture
  4. Quick Example
  5. Convert PyTorch to GGUF
  6. Build
  7. Run
  8. Benchmark against PyTorch
  9. Quantization
  10. To-Do List

Features

  • Dependency-free and lightweight inference thanks to ggml.
  • 4-bit, 5-bit and 8-bit quantization support.
  • Support for timm ViTs with different variants out of the box.

An important aspect of using vit.cpp is that it has short startup times compared to common DL frameworks, which makes it suitable for serverless deployments where the cold start is an issue.

Vision Transformer architecture

The implemented architecture is based on the original Vision Transformer from:

Vision Transformer overview

ViT architecture. Taken from the original paper.

Quick example

Details

example input

See output
 $ ./bin/vit -t 4 -m ../ggml-model-f16.gguf -i ../assets/magpie.jpeg -k 5
main: seed = 1701176263
main: n_threads = 4 / 8
vit_model_load: loading model from '../ggml-model-f16.gguf' - please wait
vit_model_load: hidden_size = 192
vit_model_load: num_hidden_layers = 12
vit_model_load: num_attention_heads = 3
vit_model_load: patch_size = 16
vit_model_load: img_size = 224
vit_model_load: num_classes = 1000
vit_model_load: ftype = 1
vit_model_load: qntvr = 0
operator(): ggml ctx size = 11.13 MB
vit_model_load: ................... done
vit_model_load: model size = 11.04 MB / num tensors = 152
main: loaded image '../assets/magpie.jpeg' (500 x 470)
vit_image_preprocess: scale = 2.232143
processed, out dims : (224 x 224)

> magpie : 0.87 > goose : 0.02 > toucan : 0.01 > drake : 0.01 > king penguin, Aptenodytes patagonica : 0.01

main: model load time = 17.92 ms main: processing time = 146.96 ms main: total time = 164.88 ms

Convert PyTorch to GGUF

# clone the repo recursively
git clone --recurse-submodules https://github.com/staghado/vit.cpp.git
cd vit.cpp
# install torch and timm
pip install torch timm
# list available models if needed; note that not all models are supported
python convert-pth-to-ggml.py --list
# convert the weights to gguf : vit tiny with patch size of 16 and an image # size of 384 pre-trained on ImageNet21k and fine-tuned on ImageNet1k
python convert-pth-to-ggml.py --model_name vit_tiny_patch16_384.augreg_in21k_ft_in1k --ftype 1

Note: You can also download the converted weights from Hugging Face directly.

wget https://huggingface.co/staghado/vit.cpp/blob/main/tiny-ggml-model-f16.gguf

Build

Simple build

# build ggml and vit mkdir build && cd build
cmake .. && make -j4
# run inference
./bin/vit -t 4 -m ../ggml-model-f16.gguf -i ../assets/tench.jpg

The optimal number of threads to use depends on many factors and more is not always better. Usually using a number of threads equal to the number of available physical cores gives the best performance in terms of speed.

Per device optimizations

Generate per-device instructions that work best for the given machine rather than using general CPU instructions. This can be done by specifying -march=native in the compiler flags.

  • Multi-threading and vectorization
  • Loop transformations(unrolling)

For AMD host processors

You can use a specialized compiler released by AMD to make full use of your specific processor's architecture. Read more here : AMD Optimizing C/C++ and Fortran Compilers (AOCC)

You can follow the given instructions to install the AOCC compiler.

Note : For my AMD Ryzen 7 3700U, the improvements were not very significant but for more recent processors there could be a gain in using a specialized compiler.

Using OpenMP

Additionally compile with OpenMP by specifying the -fopenmp flag to the compiler in the CMakeLists file, allowing multithreaded runs. Make sure to also enable multiple threads when running, e.g.:

OMP_NUM_THREADS=4 ./bin/vit -t 4 -m ../ggml-model-f16.bin -i ../assets/tench.jpg

Run

usage: ./bin/vit [options]
options:
-h, --help show this help message and exit
-s SEED, --seed SEED RNG seed (default: -1)
-t N, --threads N number of threads to use during computation (default: 4)
-m FNAME, --model FNAME model path (default: ../ggml-model-f16.bin)
-i FNAME, --inp FNAME input file (default: ../assets/tench.jpg)
-k N, --topk N top k classes to print (default: 5)
-e FLOAT, --epsilon epsilon (default: 0.000001)

Benchmark against PyTorch

First experiments on Apple M1 show inference speedups(up to 6x faster for base model) compared to native PyTorch inference.

ViT inference

You can efficiently run ViT inference on the CPU. Memory requirements and inference speed on AMD Ryzen 7 3700U(4 cores, 8 threads) for both native PyTorch and vit.cpp. Using 4 threads gives better results for my machine. The reported results of inference speed correspond to 10 runs averages for both PyTorch and vit.cpp.

ModelMax Mem(PyTorch)Max MemSpeed(PyTorch)Speed
tiny~780 MB~20 MB431 ms120 ms
small~965 MB~52 MB780 ms463 ms
base~1.61 GB~179 MB2393 ms1441 ms
large~3.86 GB~597 MB8151 ms4892 ms

Note: The models used are of the form vit_{size}_patch16_224.augreg_in21k_ft_in1k.

Benchmark on your machine

In order to test the inference speed on your machine, you can run the following scripts:

chmod +x scripts/benchmark.*
# install memory_profiler & threadpoolctl
pip install memory_profiler threadpoolctl
# run the benchmark of PyTorch
python scripts/benchmark.py
# run the benchmark of vit.cpp for non-qunatized model
./scripts/benchmark.sh
# to run the benchamrk for qunatized models; 4 threads and quantize flag
./scripts/benchmark.sh 4 1

Both scripts use 4 threads by default. In Python, the threadpoolctl library is used to limit the number of threads used by PyTorch.

Quantization

vit.cpp supports many quantization strategies from ggml such as q4_0, q4_1, q5_0, q5_1 and q8_0 types. You can quantize a model in F32 (the patch embedding is in F16) to one of these types by using the ./bin/quantize binary.

usage: ./bin/quantize /path/to/ggml-model-f32.gguf /path/to/ggml-model-quantized.gguf type type = 2 - q4_0 type = 3 - q4_1 type = 6 - q5_0 type = 7 - q5_1 type = 8 - q8_0 

For example, you can run the following to convert the model to q5_1:

./bin/quantize ../tiny-ggml-model-f16.gguf ../tiny-ggml-model-f16-quant.gguf 7

Then you can use tiny-ggml-model-f16-quant.gguf just like the model in F16.

Results

Here are the benchmarks for the different models and quantizations on my machine: For accurate estimation of run times, these benchmarks were run 100 times each.

ModelQuantizationSpeed (ms)Mem (MB)
tinyq4_0105 ms12 MB
tinyq4_197 ms12 MB
tinyq5_0116 ms13 MB
tinyq5_1112 ms13 MB
tinyq8_090 ms15 MB
smallq4_0240 ms23 MB
smallq4_1224 ms24 MB
smallq5_0288 ms25 MB
smallq5_1277 ms27 MB
smallq8_0228 ms33 MB
baseq4_0704 ms61 MB
baseq4_1626 ms66 MB
baseq5_0851 ms71 MB
baseq5_1806 ms76 MB
baseq8_0659 ms102 MB
largeq4_02189 ms181 MB
largeq4_11919 ms199 MB
largeq5_02676 ms217 MB
largeq5_12547 ms235 MB
largeq8_01994 ms325 MB

To-Do List

  • Evaluate performance on ImageNet1k:

    Run evaluation on ImageNet1k test set and analyze the performance of different quantization schemes.

This project was highly inspired by the following projects:

Star History

Star History Chart

About

Inference Vision Transformer (ViT) in plain C/C++ with ggml

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

vit.cpp

Inference Vision Transformer (ViT) in plain C/C++ using ggml without any extra dependencies

Description

This project presents a standalone implementation of the well known Vision Transformer (ViT) model family, used in a broad spectrum of applications and SOTA models like Large Multimodal Models(LMM). The primary goal is to develop a C/C++ inference engine tailored for ViT models, utilizing ggml to enhance performance, particularly on edge devices. Designed to be both lightweight and self-contained, this implementation can be run across diverse platforms.

Table of Contents
  1. Description
  2. Features
  3. Vision Transformer Architecture
  4. Quick Example
  5. Convert PyTorch to GGUF
  6. Build
  7. Run
  8. Benchmark against PyTorch
  9. Quantization
  10. To-Do List

Features

  • Dependency-free and lightweight inference thanks to ggml.
  • 4-bit, 5-bit and 8-bit quantization support.
  • Support for timm ViTs with different variants out of the box.

An important aspect of using vit.cpp is that it has short startup times compared to common DL frameworks, which makes it suitable for serverless deployments where the cold start is an issue.

Vision Transformer architecture

The implemented architecture is based on the original Vision Transformer from:

Vision Transformer overview

ViT architecture. Taken from the original paper.

Quick example

Details

example input

See output
 $ ./bin/vit -t 4 -m ../ggml-model-f16.gguf -i ../assets/magpie.jpeg -k 5
main: seed = 1701176263
main: n_threads = 4 / 8
vit_model_load: loading model from '../ggml-model-f16.gguf' - please wait
vit_model_load: hidden_size = 192
vit_model_load: num_hidden_layers = 12
vit_model_load: num_attention_heads = 3
vit_model_load: patch_size = 16
vit_model_load: img_size = 224
vit_model_load: num_classes = 1000
vit_model_load: ftype = 1
vit_model_load: qntvr = 0
operator(): ggml ctx size = 11.13 MB
vit_model_load: ................... done
vit_model_load: model size = 11.04 MB / num tensors = 152
main: loaded image '../assets/magpie.jpeg' (500 x 470)
vit_image_preprocess: scale = 2.232143
processed, out dims : (224 x 224)

> magpie : 0.87 > goose : 0.02 > toucan : 0.01 > drake : 0.01 > king penguin, Aptenodytes patagonica : 0.01

main: model load time = 17.92 ms main: processing time = 146.96 ms main: total time = 164.88 ms

Convert PyTorch to GGUF

# clone the repo recursively
git clone --recurse-submodules https://github.com/staghado/vit.cpp.git
cd vit.cpp
# install torch and timm
pip install torch timm
# list available models if needed; note that not all models are supported
python convert-pth-to-ggml.py --list
# convert the weights to gguf : vit tiny with patch size of 16 and an image # size of 384 pre-trained on ImageNet21k and fine-tuned on ImageNet1k
python convert-pth-to-ggml.py --model_name vit_tiny_patch16_384.augreg_in21k_ft_in1k --ftype 1

Note: You can also download the converted weights from Hugging Face directly.

wget https://huggingface.co/staghado/vit.cpp/blob/main/tiny-ggml-model-f16.gguf

Build

Simple build

# build ggml and vit mkdir build && cd build
cmake .. && make -j4
# run inference
./bin/vit -t 4 -m ../ggml-model-f16.gguf -i ../assets/tench.jpg

The optimal number of threads to use depends on many factors and more is not always better. Usually using a number of threads equal to the number of available physical cores gives the best performance in terms of speed.

Per device optimizations

Generate per-device instructions that work best for the given machine rather than using general CPU instructions. This can be done by specifying -march=native in the compiler flags.

  • Multi-threading and vectorization
  • Loop transformations(unrolling)

For AMD host processors

You can use a specialized compiler released by AMD to make full use of your specific processor's architecture. Read more here : AMD Optimizing C/C++ and Fortran Compilers (AOCC)

You can follow the given instructions to install the AOCC compiler.

Note : For my AMD Ryzen 7 3700U, the improvements were not very significant but for more recent processors there could be a gain in using a specialized compiler.

Using OpenMP

Additionally compile with OpenMP by specifying the -fopenmp flag to the compiler in the CMakeLists file, allowing multithreaded runs. Make sure to also enable multiple threads when running, e.g.:

OMP_NUM_THREADS=4 ./bin/vit -t 4 -m ../ggml-model-f16.bin -i ../assets/tench.jpg

Run

usage: ./bin/vit [options]
options:
-h, --help show this help message and exit
-s SEED, --seed SEED RNG seed (default: -1)
-t N, --threads N number of threads to use during computation (default: 4)
-m FNAME, --model FNAME model path (default: ../ggml-model-f16.bin)
-i FNAME, --inp FNAME input file (default: ../assets/tench.jpg)
-k N, --topk N top k classes to print (default: 5)
-e FLOAT, --epsilon epsilon (default: 0.000001)

Benchmark against PyTorch

First experiments on Apple M1 show inference speedups(up to 6x faster for base model) compared to native PyTorch inference.

ViT inference

You can efficiently run ViT inference on the CPU. Memory requirements and inference speed on AMD Ryzen 7 3700U(4 cores, 8 threads) for both native PyTorch and vit.cpp. Using 4 threads gives better results for my machine. The reported results of inference speed correspond to 10 runs averages for both PyTorch and vit.cpp.

ModelMax Mem(PyTorch)Max MemSpeed(PyTorch)Speed
tiny~780 MB~20 MB431 ms120 ms
small~965 MB~52 MB780 ms463 ms
base~1.61 GB~179 MB2393 ms1441 ms
large~3.86 GB~597 MB8151 ms4892 ms

Note: The models used are of the form vit_{size}_patch16_224.augreg_in21k_ft_in1k.

Benchmark on your machine

In order to test the inference speed on your machine, you can run the following scripts:

chmod +x scripts/benchmark.*
# install memory_profiler & threadpoolctl
pip install memory_profiler threadpoolctl
# run the benchmark of PyTorch
python scripts/benchmark.py
# run the benchmark of vit.cpp for non-qunatized model
./scripts/benchmark.sh
# to run the benchamrk for qunatized models; 4 threads and quantize flag
./scripts/benchmark.sh 4 1

Both scripts use 4 threads by default. In Python, the threadpoolctl library is used to limit the number of threads used by PyTorch.

Quantization

vit.cpp supports many quantization strategies from ggml such as q4_0, q4_1, q5_0, q5_1 and q8_0 types. You can quantize a model in F32 (the patch embedding is in F16) to one of these types by using the ./bin/quantize binary.

usage: ./bin/quantize /path/to/ggml-model-f32.gguf /path/to/ggml-model-quantized.gguf type type = 2 - q4_0 type = 3 - q4_1 type = 6 - q5_0 type = 7 - q5_1 type = 8 - q8_0 

For example, you can run the following to convert the model to q5_1:

./bin/quantize ../tiny-ggml-model-f16.gguf ../tiny-ggml-model-f16-quant.gguf 7

Then you can use tiny-ggml-model-f16-quant.gguf just like the model in F16.

Results

Here are the benchmarks for the different models and quantizations on my machine: For accurate estimation of run times, these benchmarks were run 100 times each.

ModelQuantizationSpeed (ms)Mem (MB)
tinyq4_0105 ms12 MB
tinyq4_197 ms12 MB
tinyq5_0116 ms13 MB
tinyq5_1112 ms13 MB
tinyq8_090 ms15 MB
smallq4_0240 ms23 MB
smallq4_1224 ms24 MB
smallq5_0288 ms25 MB
smallq5_1277 ms27 MB
smallq8_0228 ms33 MB
baseq4_0704 ms61 MB
baseq4_1626 ms66 MB
baseq5_0851 ms71 MB
baseq5_1806 ms76 MB
baseq8_0659 ms102 MB
largeq4_02189 ms181 MB
largeq4_11919 ms199 MB
largeq5_02676 ms217 MB
largeq5_12547 ms235 MB
largeq8_01994 ms325 MB

To-Do List

  • Evaluate performance on ImageNet1k:

    Run evaluation on ImageNet1k test set and analyze the performance of different quantization schemes.

This project was highly inspired by the following projects:

Star History

Star History Chart

About

Inference Vision Transformer (ViT) in plain C/C++ with ggml

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

vit.cpp

Inference Vision Transformer (ViT) in plain C/C++ using ggml without any extra dependencies

Description

This project presents a standalone implementation of the well known Vision Transformer (ViT) model family, used in a broad spectrum of applications and SOTA models like Large Multimodal Models(LMM). The primary goal is to develop a C/C++ inference engine tailored for ViT models, utilizing ggml to enhance performance, particularly on edge devices. Designed to be both lightweight and self-contained, this implementation can be run across diverse platforms.

Table of Contents
  1. Description
  2. Features
  3. Vision Transformer Architecture
  4. Quick Example
  5. Convert PyTorch to GGUF
  6. Build
  7. Run
  8. Benchmark against PyTorch
  9. Quantization
  10. To-Do List

Features

  • Dependency-free and lightweight inference thanks to ggml.
  • 4-bit, 5-bit and 8-bit quantization support.
  • Support for timm ViTs with different variants out of the box.

An important aspect of using vit.cpp is that it has short startup times compared to common DL frameworks, which makes it suitable for serverless deployments where the cold start is an issue.

Vision Transformer architecture

The implemented architecture is based on the original Vision Transformer from:

Vision Transformer overview

ViT architecture. Taken from the original paper.

Quick example

Details

example input

See output
 $ ./bin/vit -t 4 -m ../ggml-model-f16.gguf -i ../assets/magpie.jpeg -k 5
main: seed = 1701176263
main: n_threads = 4 / 8
vit_model_load: loading model from '../ggml-model-f16.gguf' - please wait
vit_model_load: hidden_size = 192
vit_model_load: num_hidden_layers = 12
vit_model_load: num_attention_heads = 3
vit_model_load: patch_size = 16
vit_model_load: img_size = 224
vit_model_load: num_classes = 1000
vit_model_load: ftype = 1
vit_model_load: qntvr = 0
operator(): ggml ctx size = 11.13 MB
vit_model_load: ................... done
vit_model_load: model size = 11.04 MB / num tensors = 152
main: loaded image '../assets/magpie.jpeg' (500 x 470)
vit_image_preprocess: scale = 2.232143
processed, out dims : (224 x 224)

> magpie : 0.87 > goose : 0.02 > toucan : 0.01 > drake : 0.01 > king penguin, Aptenodytes patagonica : 0.01

main: model load time = 17.92 ms main: processing time = 146.96 ms main: total time = 164.88 ms

Convert PyTorch to GGUF

# clone the repo recursively
git clone --recurse-submodules https://github.com/staghado/vit.cpp.git
cd vit.cpp
# install torch and timm
pip install torch timm
# list available models if needed; note that not all models are supported
python convert-pth-to-ggml.py --list
# convert the weights to gguf : vit tiny with patch size of 16 and an image # size of 384 pre-trained on ImageNet21k and fine-tuned on ImageNet1k
python convert-pth-to-ggml.py --model_name vit_tiny_patch16_384.augreg_in21k_ft_in1k --ftype 1

Note: You can also download the converted weights from Hugging Face directly.

wget https://huggingface.co/staghado/vit.cpp/blob/main/tiny-ggml-model-f16.gguf

Build

Simple build

# build ggml and vit mkdir build && cd build
cmake .. && make -j4
# run inference
./bin/vit -t 4 -m ../ggml-model-f16.gguf -i ../assets/tench.jpg

The optimal number of threads to use depends on many factors and more is not always better. Usually using a number of threads equal to the number of available physical cores gives the best performance in terms of speed.

Per device optimizations

Generate per-device instructions that work best for the given machine rather than using general CPU instructions. This can be done by specifying -march=native in the compiler flags.

  • Multi-threading and vectorization
  • Loop transformations(unrolling)

For AMD host processors

You can use a specialized compiler released by AMD to make full use of your specific processor's architecture. Read more here : AMD Optimizing C/C++ and Fortran Compilers (AOCC)

You can follow the given instructions to install the AOCC compiler.

Note : For my AMD Ryzen 7 3700U, the improvements were not very significant but for more recent processors there could be a gain in using a specialized compiler.

Using OpenMP

Additionally compile with OpenMP by specifying the -fopenmp flag to the compiler in the CMakeLists file, allowing multithreaded runs. Make sure to also enable multiple threads when running, e.g.:

OMP_NUM_THREADS=4 ./bin/vit -t 4 -m ../ggml-model-f16.bin -i ../assets/tench.jpg

Run

usage: ./bin/vit [options]
options:
-h, --help show this help message and exit
-s SEED, --seed SEED RNG seed (default: -1)
-t N, --threads N number of threads to use during computation (default: 4)
-m FNAME, --model FNAME model path (default: ../ggml-model-f16.bin)
-i FNAME, --inp FNAME input file (default: ../assets/tench.jpg)
-k N, --topk N top k classes to print (default: 5)
-e FLOAT, --epsilon epsilon (default: 0.000001)

Benchmark against PyTorch

First experiments on Apple M1 show inference speedups(up to 6x faster for base model) compared to native PyTorch inference.

ViT inference

You can efficiently run ViT inference on the CPU. Memory requirements and inference speed on AMD Ryzen 7 3700U(4 cores, 8 threads) for both native PyTorch and vit.cpp. Using 4 threads gives better results for my machine. The reported results of inference speed correspond to 10 runs averages for both PyTorch and vit.cpp.

ModelMax Mem(PyTorch)Max MemSpeed(PyTorch)Speed
tiny~780 MB~20 MB431 ms120 ms
small~965 MB~52 MB780 ms463 ms
base~1.61 GB~179 MB2393 ms1441 ms
large~3.86 GB~597 MB8151 ms4892 ms

Note: The models used are of the form vit_{size}_patch16_224.augreg_in21k_ft_in1k.

Benchmark on your machine

In order to test the inference speed on your machine, you can run the following scripts:

chmod +x scripts/benchmark.*
# install memory_profiler & threadpoolctl
pip install memory_profiler threadpoolctl
# run the benchmark of PyTorch
python scripts/benchmark.py
# run the benchmark of vit.cpp for non-qunatized model
./scripts/benchmark.sh
# to run the benchamrk for qunatized models; 4 threads and quantize flag
./scripts/benchmark.sh 4 1

Both scripts use 4 threads by default. In Python, the threadpoolctl library is used to limit the number of threads used by PyTorch.

Quantization

vit.cpp supports many quantization strategies from ggml such as q4_0, q4_1, q5_0, q5_1 and q8_0 types. You can quantize a model in F32 (the patch embedding is in F16) to one of these types by using the ./bin/quantize binary.

usage: ./bin/quantize /path/to/ggml-model-f32.gguf /path/to/ggml-model-quantized.gguf type type = 2 - q4_0 type = 3 - q4_1 type = 6 - q5_0 type = 7 - q5_1 type = 8 - q8_0 

For example, you can run the following to convert the model to q5_1:

./bin/quantize ../tiny-ggml-model-f16.gguf ../tiny-ggml-model-f16-quant.gguf 7

Then you can use tiny-ggml-model-f16-quant.gguf just like the model in F16.

Results

Here are the benchmarks for the different models and quantizations on my machine: For accurate estimation of run times, these benchmarks were run 100 times each.

ModelQuantizationSpeed (ms)Mem (MB)
tinyq4_0105 ms12 MB
tinyq4_197 ms12 MB
tinyq5_0116 ms13 MB
tinyq5_1112 ms13 MB
tinyq8_090 ms15 MB
smallq4_0240 ms23 MB
smallq4_1224 ms24 MB
smallq5_0288 ms25 MB
smallq5_1277 ms27 MB
smallq8_0228 ms33 MB
baseq4_0704 ms61 MB
baseq4_1626 ms66 MB
baseq5_0851 ms71 MB
baseq5_1806 ms76 MB
baseq8_0659 ms102 MB
largeq4_02189 ms181 MB
largeq4_11919 ms199 MB
largeq5_02676 ms217 MB
largeq5_12547 ms235 MB
largeq8_01994 ms325 MB

To-Do List

  • Evaluate performance on ImageNet1k:

    Run evaluation on ImageNet1k test set and analyze the performance of different quantization schemes.

This project was highly inspired by the following projects:

Star History

Star History Chart

About

Inference Vision Transformer (ViT) in plain C/C++ with ggml

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

vit.cpp

Inference Vision Transformer (ViT) in plain C/C++ using ggml without any extra dependencies

Description

This project presents a standalone implementation of the well known Vision Transformer (ViT) model family, used in a broad spectrum of applications and SOTA models like Large Multimodal Models(LMM). The primary goal is to develop a C/C++ inference engine tailored for ViT models, utilizing ggml to enhance performance, particularly on edge devices. Designed to be both lightweight and self-contained, this implementation can be run across diverse platforms.

Table of Contents
  1. Description
  2. Features
  3. Vision Transformer Architecture
  4. Quick Example
  5. Convert PyTorch to GGUF
  6. Build
  7. Run
  8. Benchmark against PyTorch
  9. Quantization
  10. To-Do List

Features

  • Dependency-free and lightweight inference thanks to ggml.
  • 4-bit, 5-bit and 8-bit quantization support.
  • Support for timm ViTs with different variants out of the box.

An important aspect of using vit.cpp is that it has short startup times compared to common DL frameworks, which makes it suitable for serverless deployments where the cold start is an issue.

Vision Transformer architecture

The implemented architecture is based on the original Vision Transformer from:

Vision Transformer overview

ViT architecture. Taken from the original paper.

Quick example

Details

example input

See output
 $ ./bin/vit -t 4 -m ../ggml-model-f16.gguf -i ../assets/magpie.jpeg -k 5
main: seed = 1701176263
main: n_threads = 4 / 8
vit_model_load: loading model from '../ggml-model-f16.gguf' - please wait
vit_model_load: hidden_size = 192
vit_model_load: num_hidden_layers = 12
vit_model_load: num_attention_heads = 3
vit_model_load: patch_size = 16
vit_model_load: img_size = 224
vit_model_load: num_classes = 1000
vit_model_load: ftype = 1
vit_model_load: qntvr = 0
operator(): ggml ctx size = 11.13 MB
vit_model_load: ................... done
vit_model_load: model size = 11.04 MB / num tensors = 152
main: loaded image '../assets/magpie.jpeg' (500 x 470)
vit_image_preprocess: scale = 2.232143
processed, out dims : (224 x 224)

> magpie : 0.87 > goose : 0.02 > toucan : 0.01 > drake : 0.01 > king penguin, Aptenodytes patagonica : 0.01

main: model load time = 17.92 ms main: processing time = 146.96 ms main: total time = 164.88 ms

Convert PyTorch to GGUF

# clone the repo recursively
git clone --recurse-submodules https://github.com/staghado/vit.cpp.git
cd vit.cpp
# install torch and timm
pip install torch timm
# list available models if needed; note that not all models are supported
python convert-pth-to-ggml.py --list
# convert the weights to gguf : vit tiny with patch size of 16 and an image # size of 384 pre-trained on ImageNet21k and fine-tuned on ImageNet1k
python convert-pth-to-ggml.py --model_name vit_tiny_patch16_384.augreg_in21k_ft_in1k --ftype 1

Note: You can also download the converted weights from Hugging Face directly.

wget https://huggingface.co/staghado/vit.cpp/blob/main/tiny-ggml-model-f16.gguf

Build

Simple build

# build ggml and vit mkdir build && cd build
cmake .. && make -j4
# run inference
./bin/vit -t 4 -m ../ggml-model-f16.gguf -i ../assets/tench.jpg

The optimal number of threads to use depends on many factors and more is not always better. Usually using a number of threads equal to the number of available physical cores gives the best performance in terms of speed.

Per device optimizations

Generate per-device instructions that work best for the given machine rather than using general CPU instructions. This can be done by specifying -march=native in the compiler flags.

  • Multi-threading and vectorization
  • Loop transformations(unrolling)

For AMD host processors

You can use a specialized compiler released by AMD to make full use of your specific processor's architecture. Read more here : AMD Optimizing C/C++ and Fortran Compilers (AOCC)

You can follow the given instructions to install the AOCC compiler.

Note : For my AMD Ryzen 7 3700U, the improvements were not very significant but for more recent processors there could be a gain in using a specialized compiler.

Using OpenMP

Additionally compile with OpenMP by specifying the -fopenmp flag to the compiler in the CMakeLists file, allowing multithreaded runs. Make sure to also enable multiple threads when running, e.g.:

OMP_NUM_THREADS=4 ./bin/vit -t 4 -m ../ggml-model-f16.bin -i ../assets/tench.jpg

Run

usage: ./bin/vit [options]
options:
-h, --help show this help message and exit
-s SEED, --seed SEED RNG seed (default: -1)
-t N, --threads N number of threads to use during computation (default: 4)
-m FNAME, --model FNAME model path (default: ../ggml-model-f16.bin)
-i FNAME, --inp FNAME input file (default: ../assets/tench.jpg)
-k N, --topk N top k classes to print (default: 5)
-e FLOAT, --epsilon epsilon (default: 0.000001)

Benchmark against PyTorch

First experiments on Apple M1 show inference speedups(up to 6x faster for base model) compared to native PyTorch inference.

ViT inference

You can efficiently run ViT inference on the CPU. Memory requirements and inference speed on AMD Ryzen 7 3700U(4 cores, 8 threads) for both native PyTorch and vit.cpp. Using 4 threads gives better results for my machine. The reported results of inference speed correspond to 10 runs averages for both PyTorch and vit.cpp.

ModelMax Mem(PyTorch)Max MemSpeed(PyTorch)Speed
tiny~780 MB~20 MB431 ms120 ms
small~965 MB~52 MB780 ms463 ms
base~1.61 GB~179 MB2393 ms1441 ms
large~3.86 GB~597 MB8151 ms4892 ms

Note: The models used are of the form vit_{size}_patch16_224.augreg_in21k_ft_in1k.

Benchmark on your machine

In order to test the inference speed on your machine, you can run the following scripts:

chmod +x scripts/benchmark.*
# install memory_profiler & threadpoolctl
pip install memory_profiler threadpoolctl
# run the benchmark of PyTorch
python scripts/benchmark.py
# run the benchmark of vit.cpp for non-qunatized model
./scripts/benchmark.sh
# to run the benchamrk for qunatized models; 4 threads and quantize flag
./scripts/benchmark.sh 4 1

Both scripts use 4 threads by default. In Python, the threadpoolctl library is used to limit the number of threads used by PyTorch.

Quantization

vit.cpp supports many quantization strategies from ggml such as q4_0, q4_1, q5_0, q5_1 and q8_0 types. You can quantize a model in F32 (the patch embedding is in F16) to one of these types by using the ./bin/quantize binary.

usage: ./bin/quantize /path/to/ggml-model-f32.gguf /path/to/ggml-model-quantized.gguf type type = 2 - q4_0 type = 3 - q4_1 type = 6 - q5_0 type = 7 - q5_1 type = 8 - q8_0 

For example, you can run the following to convert the model to q5_1:

./bin/quantize ../tiny-ggml-model-f16.gguf ../tiny-ggml-model-f16-quant.gguf 7

Then you can use tiny-ggml-model-f16-quant.gguf just like the model in F16.

Results

Here are the benchmarks for the different models and quantizations on my machine: For accurate estimation of run times, these benchmarks were run 100 times each.

ModelQuantizationSpeed (ms)Mem (MB)
tinyq4_0105 ms12 MB
tinyq4_197 ms12 MB
tinyq5_0116 ms13 MB
tinyq5_1112 ms13 MB
tinyq8_090 ms15 MB
smallq4_0240 ms23 MB
smallq4_1224 ms24 MB
smallq5_0288 ms25 MB
smallq5_1277 ms27 MB
smallq8_0228 ms33 MB
baseq4_0704 ms61 MB
baseq4_1626 ms66 MB
baseq5_0851 ms71 MB
baseq5_1806 ms76 MB
baseq8_0659 ms102 MB
largeq4_02189 ms181 MB
largeq4_11919 ms199 MB
largeq5_02676 ms217 MB
largeq5_12547 ms235 MB
largeq8_01994 ms325 MB

To-Do List

  • Evaluate performance on ImageNet1k:

    Run evaluation on ImageNet1k test set and analyze the performance of different quantization schemes.

This project was highly inspired by the following projects:

Star History

Star History Chart

About

Inference Vision Transformer (ViT) in plain C/C++ with ggml

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

vit.cpp

Inference Vision Transformer (ViT) in plain C/C++ using ggml without any extra dependencies

Description

This project presents a standalone implementation of the well known Vision Transformer (ViT) model family, used in a broad spectrum of applications and SOTA models like Large Multimodal Models(LMM). The primary goal is to develop a C/C++ inference engine tailored for ViT models, utilizing ggml to enhance performance, particularly on edge devices. Designed to be both lightweight and self-contained, this implementation can be run across diverse platforms.

Table of Contents
  1. Description
  2. Features
  3. Vision Transformer Architecture
  4. Quick Example
  5. Convert PyTorch to GGUF
  6. Build
  7. Run
  8. Benchmark against PyTorch
  9. Quantization
  10. To-Do List

Features

  • Dependency-free and lightweight inference thanks to ggml.
  • 4-bit, 5-bit and 8-bit quantization support.
  • Support for timm ViTs with different variants out of the box.

An important aspect of using vit.cpp is that it has short startup times compared to common DL frameworks, which makes it suitable for serverless deployments where the cold start is an issue.

Vision Transformer architecture

The implemented architecture is based on the original Vision Transformer from:

Vision Transformer overview

ViT architecture. Taken from the original paper.

Quick example

Details

example input

See output
 $ ./bin/vit -t 4 -m ../ggml-model-f16.gguf -i ../assets/magpie.jpeg -k 5
main: seed = 1701176263
main: n_threads = 4 / 8
vit_model_load: loading model from '../ggml-model-f16.gguf' - please wait
vit_model_load: hidden_size = 192
vit_model_load: num_hidden_layers = 12
vit_model_load: num_attention_heads = 3
vit_model_load: patch_size = 16
vit_model_load: img_size = 224
vit_model_load: num_classes = 1000
vit_model_load: ftype = 1
vit_model_load: qntvr = 0
operator(): ggml ctx size = 11.13 MB
vit_model_load: ................... done
vit_model_load: model size = 11.04 MB / num tensors = 152
main: loaded image '../assets/magpie.jpeg' (500 x 470)
vit_image_preprocess: scale = 2.232143
processed, out dims : (224 x 224)

> magpie : 0.87 > goose : 0.02 > toucan : 0.01 > drake : 0.01 > king penguin, Aptenodytes patagonica : 0.01

main: model load time = 17.92 ms main: processing time = 146.96 ms main: total time = 164.88 ms

Convert PyTorch to GGUF

# clone the repo recursively
git clone --recurse-submodules https://github.com/staghado/vit.cpp.git
cd vit.cpp
# install torch and timm
pip install torch timm
# list available models if needed; note that not all models are supported
python convert-pth-to-ggml.py --list
# convert the weights to gguf : vit tiny with patch size of 16 and an image # size of 384 pre-trained on ImageNet21k and fine-tuned on ImageNet1k
python convert-pth-to-ggml.py --model_name vit_tiny_patch16_384.augreg_in21k_ft_in1k --ftype 1

Note: You can also download the converted weights from Hugging Face directly.

wget https://huggingface.co/staghado/vit.cpp/blob/main/tiny-ggml-model-f16.gguf

Build

Simple build

# build ggml and vit mkdir build && cd build
cmake .. && make -j4
# run inference
./bin/vit -t 4 -m ../ggml-model-f16.gguf -i ../assets/tench.jpg

The optimal number of threads to use depends on many factors and more is not always better. Usually using a number of threads equal to the number of available physical cores gives the best performance in terms of speed.

Per device optimizations

Generate per-device instructions that work best for the given machine rather than using general CPU instructions. This can be done by specifying -march=native in the compiler flags.

  • Multi-threading and vectorization
  • Loop transformations(unrolling)

For AMD host processors

You can use a specialized compiler released by AMD to make full use of your specific processor's architecture. Read more here : AMD Optimizing C/C++ and Fortran Compilers (AOCC)

You can follow the given instructions to install the AOCC compiler.

Note : For my AMD Ryzen 7 3700U, the improvements were not very significant but for more recent processors there could be a gain in using a specialized compiler.

Using OpenMP

Additionally compile with OpenMP by specifying the -fopenmp flag to the compiler in the CMakeLists file, allowing multithreaded runs. Make sure to also enable multiple threads when running, e.g.:

OMP_NUM_THREADS=4 ./bin/vit -t 4 -m ../ggml-model-f16.bin -i ../assets/tench.jpg

Run

usage: ./bin/vit [options]
options:
-h, --help show this help message and exit
-s SEED, --seed SEED RNG seed (default: -1)
-t N, --threads N number of threads to use during computation (default: 4)
-m FNAME, --model FNAME model path (default: ../ggml-model-f16.bin)
-i FNAME, --inp FNAME input file (default: ../assets/tench.jpg)
-k N, --topk N top k classes to print (default: 5)
-e FLOAT, --epsilon epsilon (default: 0.000001)

Benchmark against PyTorch

First experiments on Apple M1 show inference speedups(up to 6x faster for base model) compared to native PyTorch inference.

ViT inference

You can efficiently run ViT inference on the CPU. Memory requirements and inference speed on AMD Ryzen 7 3700U(4 cores, 8 threads) for both native PyTorch and vit.cpp. Using 4 threads gives better results for my machine. The reported results of inference speed correspond to 10 runs averages for both PyTorch and vit.cpp.

ModelMax Mem(PyTorch)Max MemSpeed(PyTorch)Speed
tiny~780 MB~20 MB431 ms120 ms
small~965 MB~52 MB780 ms463 ms
base~1.61 GB~179 MB2393 ms1441 ms
large~3.86 GB~597 MB8151 ms4892 ms

Note: The models used are of the form vit_{size}_patch16_224.augreg_in21k_ft_in1k.

Benchmark on your machine

In order to test the inference speed on your machine, you can run the following scripts:

chmod +x scripts/benchmark.*
# install memory_profiler & threadpoolctl
pip install memory_profiler threadpoolctl
# run the benchmark of PyTorch
python scripts/benchmark.py
# run the benchmark of vit.cpp for non-qunatized model
./scripts/benchmark.sh
# to run the benchamrk for qunatized models; 4 threads and quantize flag
./scripts/benchmark.sh 4 1

Both scripts use 4 threads by default. In Python, the threadpoolctl library is used to limit the number of threads used by PyTorch.

Quantization

vit.cpp supports many quantization strategies from ggml such as q4_0, q4_1, q5_0, q5_1 and q8_0 types. You can quantize a model in F32 (the patch embedding is in F16) to one of these types by using the ./bin/quantize binary.

usage: ./bin/quantize /path/to/ggml-model-f32.gguf /path/to/ggml-model-quantized.gguf type type = 2 - q4_0 type = 3 - q4_1 type = 6 - q5_0 type = 7 - q5_1 type = 8 - q8_0 

For example, you can run the following to convert the model to q5_1:

./bin/quantize ../tiny-ggml-model-f16.gguf ../tiny-ggml-model-f16-quant.gguf 7

Then you can use tiny-ggml-model-f16-quant.gguf just like the model in F16.

Results

Here are the benchmarks for the different models and quantizations on my machine: For accurate estimation of run times, these benchmarks were run 100 times each.

ModelQuantizationSpeed (ms)Mem (MB)
tinyq4_0105 ms12 MB
tinyq4_197 ms12 MB
tinyq5_0116 ms13 MB
tinyq5_1112 ms13 MB
tinyq8_090 ms15 MB
smallq4_0240 ms23 MB
smallq4_1224 ms24 MB
smallq5_0288 ms25 MB
smallq5_1277 ms27 MB
smallq8_0228 ms33 MB
baseq4_0704 ms61 MB
baseq4_1626 ms66 MB
baseq5_0851 ms71 MB
baseq5_1806 ms76 MB
baseq8_0659 ms102 MB
largeq4_02189 ms181 MB
largeq4_11919 ms199 MB
largeq5_02676 ms217 MB
largeq5_12547 ms235 MB
largeq8_01994 ms325 MB

To-Do List

  • Evaluate performance on ImageNet1k:

    Run evaluation on ImageNet1k test set and analyze the performance of different quantization schemes.

This project was highly inspired by the following projects:

Star History

Star History Chart

About

Inference Vision Transformer (ViT) in plain C/C++ with ggml

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages