Repository files navigation

vision.cpp

Computer Vision ML inference in C++

  • Self-contained C++ library
  • Efficient inference on consumer CPU and GPUs (NVIDIA, AMD, Intel)
  • Lightweight deployment on many platforms (Windows, Linux, MacOS)
  • Growing number of supported models behind a simple API
  • Modular design for full control and implementing your own models

Based on ggml similar to the llama.cpp project.

Features

ModelTaskBackends
MobileSAMPromptable segmentationCPU, Vulkan
BiRefNetDichotomous segmentationCPU, Vulkan
Depth-AnythingDepth estimationCPU, Vulkan
MI-GANInpaintingCPU, Vulkan
ESRGANSuper-resolutionCPU, Vulkan
Implement a model [Guide]

Backbones: SWIN (v1), DINO (v2), TinyViT

Get Started

Get the library and executables:

Example: Select an object in an image

Let's use MobileSAM to generate a segmentation mask of the plushy on the right by passing in a box describing its approximate location.

Example image showing box prompt at pixel location (420, 120) - (650, 430), and the output mask

You can download the model and input image here: MobileSAM-F16.gguf | input.jpg

CLI

Find the vision-cli executable in the bin folder and run it to generate the mask:

vision-cli -m MobileSAM-F16.gguf -i input.jpg -p 420 120 650 430 -o mask.png

Pass --composite output.png to composite input and mask. Use --help for more options.

API

#include<visp/vision.h>usingnamespacevisp;voidmain() {
backend_device cpu = backend_init(backend_type::cpu);
sam_model sam = sam_load_model("MobileSAM-F16.gguf", cpu);
image_data input_image = image_load("input.jpg");
sam_encode(sam, input_image);
image_data object_mask = sam_compute(sam, box_2d{{420, 120}, {650, 320}});
image_save(object_mask, "mask.png");
}

This shows the high-level API. Internally it is composed of multiple smaller functions that handle model loading, pre-processing inputs, transferring data to backend devices, post-processing output, etc. These can be used as building blocks for flexible functions which integrate with your existing data sources and infrastructure.

Models

MobileSAM

example-sam

Model download | Paper (arXiv) | Repository (GitHub) | Segment-Anything-Model | License: Apache-2

vision-cli sam -m MobileSAM-F16.gguf -i input.png -p 300 200 -o mask.png --composite comp.png

BiRefNet

example-birefnet

Model download | Paper (arXiv) | Repository (GitHub) | License: MIT

vision-cli birefnet -m BiRefNet-lite-F16.gguf -i input.png -o mask.png --composite comp.png

Depth-Anything V2

example-depth-anything

Model download | Paper (arXiv) | Repository (GitHub) | License: Apache-2 / CC-BY-NC-4

vision-cli depth-anything -m Depth-Anything-V2-Small-F16.gguf -i input.png -o depth.png

MI-GAN

example-migan

Model download | Paper (thecvf.com) | Repository (GitHub) | License: MIT

vision-cli migan -m MIGAN-512-places2-F16.gguf -i image.png mask.png -o output.png

Real-ESRGAN

example-esrgan

Model download | Paper (arXiv) | Repository (GitHub) | License: BSD-3-Clause

vision-cli esrgan -m ESRGAN-4x-foolhardy_Remacri-F16.gguf -i input.png -o output.png

Converting models

Models need to be converted to GGUF before they can be used. This will also rearrange or precompute tensors for more optimal inference.

To convert a model, install uv and run:

uv run scripts/convert.py <arch> MyModel.pth

where <arch> is one of sam, birefnet, esrgan, ....

This will create models/MyModel.gguf. See convert.py --help for more options.

Building

Building requires CMake and a compiler with C++20 support.

Get the sources

git clone https://github.com/Acly/vision.cpp.git --recursive
cd vision.cpp

Configure and build

cmake . -B build
cmake --build build --config Release

Vulkan (Optional)

Building with Vulkan GPU support requires the Vulkan SDK to be installed.

cmake . -B build -D VISP_VULKAN=ON

Tests (Optional)

Build with -DVISP_TESTS=ON. Run all C++ tests with the following command:

cd build
ctest -C Release

Some tests require a Python environment. It can be set up with uv:

# Setup venv and install dependencies (once only)
uv sync --dev
# Run python tests
uv run pytest

Performance

Performance optimization is an ongoing process. The aim is to be in the same ballpark as other frameworks for inference speed, but with:

  • much faster initialization and model loading time (<100 ms)
  • lower memory overhead
  • tiny deployment size (<5 MB for CPU, +30 MB for GPU)

Inference speed

  • CPU: AMD Ryzen 5 5600X (6 cores)
  • GPU: NVIDIA GeForce RTX 4070

MobileSAM, 1024x1024

vision.cppPyTorchONNX Runtime
cpuf32669 ms601 ms805 ms
gpuf1619 ms16 ms

BiRefNet, 1024x1024

Modelvision.cppPyTorchONNX Runtime
Fullcpuf3216333 ms18290 ms
Fullgpuf16208 ms190 ms
Litecpuf324505 ms10900 ms6978 ms
Litegpuf1685 ms84 ms

Depth-Anything, 518x714

Modelvision.cppPyTorch
Smallgpuf1611 ms10 ms
Basegpuf1624 ms22 ms

MI-GAN, 512x512

Modelvision.cppPyTorch
512-places2cpuf32523 ms637 ms
512-places2gpuf1621 ms17 ms

Setup

  • vision.cpp: using vision-bench, GPU via Vulkan, eg. vision-bench -m sam
  • PyTorch: v2.7.1+cu128, eager eval, GPU via CUDA, average n iterations after warm-up

Dependencies (integrated)

  • ggml - ML tensor library | MIT
  • stb-image - Image load/save/resize | Public Domain
  • fmt - String formatting (only if compiler doesn't support <format>) | MIT

About

Computer Vision ML inference in C++

Topics

Resources

Contributing

Stars

53 stars

Watchers

5 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

vision.cpp

Computer Vision ML inference in C++

  • Self-contained C++ library
  • Efficient inference on consumer CPU and GPUs (NVIDIA, AMD, Intel)
  • Lightweight deployment on many platforms (Windows, Linux, MacOS)
  • Growing number of supported models behind a simple API
  • Modular design for full control and implementing your own models

Based on ggml similar to the llama.cpp project.

Features

ModelTaskBackends
MobileSAMPromptable segmentationCPU, Vulkan
BiRefNetDichotomous segmentationCPU, Vulkan
Depth-AnythingDepth estimationCPU, Vulkan
MI-GANInpaintingCPU, Vulkan
ESRGANSuper-resolutionCPU, Vulkan
Implement a model [Guide]

Backbones: SWIN (v1), DINO (v2), TinyViT

Get Started

Get the library and executables:

Example: Select an object in an image

Let's use MobileSAM to generate a segmentation mask of the plushy on the right by passing in a box describing its approximate location.

Example image showing box prompt at pixel location (420, 120) - (650, 430), and the output mask

You can download the model and input image here: MobileSAM-F16.gguf | input.jpg

CLI

Find the vision-cli executable in the bin folder and run it to generate the mask:

vision-cli -m MobileSAM-F16.gguf -i input.jpg -p 420 120 650 430 -o mask.png

Pass --composite output.png to composite input and mask. Use --help for more options.

API

#include<visp/vision.h>usingnamespacevisp;voidmain() {
backend_device cpu = backend_init(backend_type::cpu);
sam_model sam = sam_load_model("MobileSAM-F16.gguf", cpu);
image_data input_image = image_load("input.jpg");
sam_encode(sam, input_image);
image_data object_mask = sam_compute(sam, box_2d{{420, 120}, {650, 320}});
image_save(object_mask, "mask.png");
}

This shows the high-level API. Internally it is composed of multiple smaller functions that handle model loading, pre-processing inputs, transferring data to backend devices, post-processing output, etc. These can be used as building blocks for flexible functions which integrate with your existing data sources and infrastructure.

Models

MobileSAM

example-sam

Model download | Paper (arXiv) | Repository (GitHub) | Segment-Anything-Model | License: Apache-2

vision-cli sam -m MobileSAM-F16.gguf -i input.png -p 300 200 -o mask.png --composite comp.png

BiRefNet

example-birefnet

Model download | Paper (arXiv) | Repository (GitHub) | License: MIT

vision-cli birefnet -m BiRefNet-lite-F16.gguf -i input.png -o mask.png --composite comp.png

Depth-Anything V2

example-depth-anything

Model download | Paper (arXiv) | Repository (GitHub) | License: Apache-2 / CC-BY-NC-4

vision-cli depth-anything -m Depth-Anything-V2-Small-F16.gguf -i input.png -o depth.png

MI-GAN

example-migan

Model download | Paper (thecvf.com) | Repository (GitHub) | License: MIT

vision-cli migan -m MIGAN-512-places2-F16.gguf -i image.png mask.png -o output.png

Real-ESRGAN

example-esrgan

Model download | Paper (arXiv) | Repository (GitHub) | License: BSD-3-Clause

vision-cli esrgan -m ESRGAN-4x-foolhardy_Remacri-F16.gguf -i input.png -o output.png

Converting models

Models need to be converted to GGUF before they can be used. This will also rearrange or precompute tensors for more optimal inference.

To convert a model, install uv and run:

uv run scripts/convert.py <arch> MyModel.pth

where <arch> is one of sam, birefnet, esrgan, ....

This will create models/MyModel.gguf. See convert.py --help for more options.

Building

Building requires CMake and a compiler with C++20 support.

Get the sources

git clone https://github.com/Acly/vision.cpp.git --recursive
cd vision.cpp

Configure and build

cmake . -B build
cmake --build build --config Release

Vulkan (Optional)

Building with Vulkan GPU support requires the Vulkan SDK to be installed.

cmake . -B build -D VISP_VULKAN=ON

Tests (Optional)

Build with -DVISP_TESTS=ON. Run all C++ tests with the following command:

cd build
ctest -C Release

Some tests require a Python environment. It can be set up with uv:

# Setup venv and install dependencies (once only)
uv sync --dev
# Run python tests
uv run pytest

Performance

Performance optimization is an ongoing process. The aim is to be in the same ballpark as other frameworks for inference speed, but with:

  • much faster initialization and model loading time (<100 ms)
  • lower memory overhead
  • tiny deployment size (<5 MB for CPU, +30 MB for GPU)

Inference speed

  • CPU: AMD Ryzen 5 5600X (6 cores)
  • GPU: NVIDIA GeForce RTX 4070

MobileSAM, 1024x1024

vision.cppPyTorchONNX Runtime
cpuf32669 ms601 ms805 ms
gpuf1619 ms16 ms

BiRefNet, 1024x1024

Modelvision.cppPyTorchONNX Runtime
Fullcpuf3216333 ms18290 ms
Fullgpuf16208 ms190 ms
Litecpuf324505 ms10900 ms6978 ms
Litegpuf1685 ms84 ms

Depth-Anything, 518x714

Modelvision.cppPyTorch
Smallgpuf1611 ms10 ms
Basegpuf1624 ms22 ms

MI-GAN, 512x512

Modelvision.cppPyTorch
512-places2cpuf32523 ms637 ms
512-places2gpuf1621 ms17 ms

Setup

  • vision.cpp: using vision-bench, GPU via Vulkan, eg. vision-bench -m sam
  • PyTorch: v2.7.1+cu128, eager eval, GPU via CUDA, average n iterations after warm-up

Dependencies (integrated)

  • ggml - ML tensor library | MIT
  • stb-image - Image load/save/resize | Public Domain
  • fmt - String formatting (only if compiler doesn't support <format>) | MIT

About

Computer Vision ML inference in C++

Topics

Resources

Contributing

Stars

53 stars

Watchers

5 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

vision.cpp

Computer Vision ML inference in C++

  • Self-contained C++ library
  • Efficient inference on consumer CPU and GPUs (NVIDIA, AMD, Intel)
  • Lightweight deployment on many platforms (Windows, Linux, MacOS)
  • Growing number of supported models behind a simple API
  • Modular design for full control and implementing your own models

Based on ggml similar to the llama.cpp project.

Features

ModelTaskBackends
MobileSAMPromptable segmentationCPU, Vulkan
BiRefNetDichotomous segmentationCPU, Vulkan
Depth-AnythingDepth estimationCPU, Vulkan
MI-GANInpaintingCPU, Vulkan
ESRGANSuper-resolutionCPU, Vulkan
Implement a model [Guide]

Backbones: SWIN (v1), DINO (v2), TinyViT

Get Started

Get the library and executables:

Example: Select an object in an image

Let's use MobileSAM to generate a segmentation mask of the plushy on the right by passing in a box describing its approximate location.

Example image showing box prompt at pixel location (420, 120) - (650, 430), and the output mask

You can download the model and input image here: MobileSAM-F16.gguf | input.jpg

CLI

Find the vision-cli executable in the bin folder and run it to generate the mask:

vision-cli -m MobileSAM-F16.gguf -i input.jpg -p 420 120 650 430 -o mask.png

Pass --composite output.png to composite input and mask. Use --help for more options.

API

#include<visp/vision.h>usingnamespacevisp;voidmain() {
backend_device cpu = backend_init(backend_type::cpu);
sam_model sam = sam_load_model("MobileSAM-F16.gguf", cpu);
image_data input_image = image_load("input.jpg");
sam_encode(sam, input_image);
image_data object_mask = sam_compute(sam, box_2d{{420, 120}, {650, 320}});
image_save(object_mask, "mask.png");
}

This shows the high-level API. Internally it is composed of multiple smaller functions that handle model loading, pre-processing inputs, transferring data to backend devices, post-processing output, etc. These can be used as building blocks for flexible functions which integrate with your existing data sources and infrastructure.

Models

MobileSAM

example-sam

Model download | Paper (arXiv) | Repository (GitHub) | Segment-Anything-Model | License: Apache-2

vision-cli sam -m MobileSAM-F16.gguf -i input.png -p 300 200 -o mask.png --composite comp.png

BiRefNet

example-birefnet

Model download | Paper (arXiv) | Repository (GitHub) | License: MIT

vision-cli birefnet -m BiRefNet-lite-F16.gguf -i input.png -o mask.png --composite comp.png

Depth-Anything V2

example-depth-anything

Model download | Paper (arXiv) | Repository (GitHub) | License: Apache-2 / CC-BY-NC-4

vision-cli depth-anything -m Depth-Anything-V2-Small-F16.gguf -i input.png -o depth.png

MI-GAN

example-migan

Model download | Paper (thecvf.com) | Repository (GitHub) | License: MIT

vision-cli migan -m MIGAN-512-places2-F16.gguf -i image.png mask.png -o output.png

Real-ESRGAN

example-esrgan

Model download | Paper (arXiv) | Repository (GitHub) | License: BSD-3-Clause

vision-cli esrgan -m ESRGAN-4x-foolhardy_Remacri-F16.gguf -i input.png -o output.png

Converting models

Models need to be converted to GGUF before they can be used. This will also rearrange or precompute tensors for more optimal inference.

To convert a model, install uv and run:

uv run scripts/convert.py <arch> MyModel.pth

where <arch> is one of sam, birefnet, esrgan, ....

This will create models/MyModel.gguf. See convert.py --help for more options.

Building

Building requires CMake and a compiler with C++20 support.

Get the sources

git clone https://github.com/Acly/vision.cpp.git --recursive
cd vision.cpp

Configure and build

cmake . -B build
cmake --build build --config Release

Vulkan (Optional)

Building with Vulkan GPU support requires the Vulkan SDK to be installed.

cmake . -B build -D VISP_VULKAN=ON

Tests (Optional)

Build with -DVISP_TESTS=ON. Run all C++ tests with the following command:

cd build
ctest -C Release

Some tests require a Python environment. It can be set up with uv:

# Setup venv and install dependencies (once only)
uv sync --dev
# Run python tests
uv run pytest

Performance

Performance optimization is an ongoing process. The aim is to be in the same ballpark as other frameworks for inference speed, but with:

  • much faster initialization and model loading time (<100 ms)
  • lower memory overhead
  • tiny deployment size (<5 MB for CPU, +30 MB for GPU)

Inference speed

  • CPU: AMD Ryzen 5 5600X (6 cores)
  • GPU: NVIDIA GeForce RTX 4070

MobileSAM, 1024x1024

vision.cppPyTorchONNX Runtime
cpuf32669 ms601 ms805 ms
gpuf1619 ms16 ms

BiRefNet, 1024x1024

Modelvision.cppPyTorchONNX Runtime
Fullcpuf3216333 ms18290 ms
Fullgpuf16208 ms190 ms
Litecpuf324505 ms10900 ms6978 ms
Litegpuf1685 ms84 ms

Depth-Anything, 518x714

Modelvision.cppPyTorch
Smallgpuf1611 ms10 ms
Basegpuf1624 ms22 ms

MI-GAN, 512x512

Modelvision.cppPyTorch
512-places2cpuf32523 ms637 ms
512-places2gpuf1621 ms17 ms

Setup

  • vision.cpp: using vision-bench, GPU via Vulkan, eg. vision-bench -m sam
  • PyTorch: v2.7.1+cu128, eager eval, GPU via CUDA, average n iterations after warm-up

Dependencies (integrated)

  • ggml - ML tensor library | MIT
  • stb-image - Image load/save/resize | Public Domain
  • fmt - String formatting (only if compiler doesn't support <format>) | MIT

About

Computer Vision ML inference in C++

Topics

Resources

Contributing

Stars

53 stars

Watchers

5 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

vision.cpp

Computer Vision ML inference in C++

  • Self-contained C++ library
  • Efficient inference on consumer CPU and GPUs (NVIDIA, AMD, Intel)
  • Lightweight deployment on many platforms (Windows, Linux, MacOS)
  • Growing number of supported models behind a simple API
  • Modular design for full control and implementing your own models

Based on ggml similar to the llama.cpp project.

Features

ModelTaskBackends
MobileSAMPromptable segmentationCPU, Vulkan
BiRefNetDichotomous segmentationCPU, Vulkan
Depth-AnythingDepth estimationCPU, Vulkan
MI-GANInpaintingCPU, Vulkan
ESRGANSuper-resolutionCPU, Vulkan
Implement a model [Guide]

Backbones: SWIN (v1), DINO (v2), TinyViT

Get Started

Get the library and executables:

Example: Select an object in an image

Let's use MobileSAM to generate a segmentation mask of the plushy on the right by passing in a box describing its approximate location.

Example image showing box prompt at pixel location (420, 120) - (650, 430), and the output mask

You can download the model and input image here: MobileSAM-F16.gguf | input.jpg

CLI

Find the vision-cli executable in the bin folder and run it to generate the mask:

vision-cli -m MobileSAM-F16.gguf -i input.jpg -p 420 120 650 430 -o mask.png

Pass --composite output.png to composite input and mask. Use --help for more options.

API

#include<visp/vision.h>usingnamespacevisp;voidmain() {
backend_device cpu = backend_init(backend_type::cpu);
sam_model sam = sam_load_model("MobileSAM-F16.gguf", cpu);
image_data input_image = image_load("input.jpg");
sam_encode(sam, input_image);
image_data object_mask = sam_compute(sam, box_2d{{420, 120}, {650, 320}});
image_save(object_mask, "mask.png");
}

This shows the high-level API. Internally it is composed of multiple smaller functions that handle model loading, pre-processing inputs, transferring data to backend devices, post-processing output, etc. These can be used as building blocks for flexible functions which integrate with your existing data sources and infrastructure.

Models

MobileSAM

example-sam

Model download | Paper (arXiv) | Repository (GitHub) | Segment-Anything-Model | License: Apache-2

vision-cli sam -m MobileSAM-F16.gguf -i input.png -p 300 200 -o mask.png --composite comp.png

BiRefNet

example-birefnet

Model download | Paper (arXiv) | Repository (GitHub) | License: MIT

vision-cli birefnet -m BiRefNet-lite-F16.gguf -i input.png -o mask.png --composite comp.png

Depth-Anything V2

example-depth-anything

Model download | Paper (arXiv) | Repository (GitHub) | License: Apache-2 / CC-BY-NC-4

vision-cli depth-anything -m Depth-Anything-V2-Small-F16.gguf -i input.png -o depth.png

MI-GAN

example-migan

Model download | Paper (thecvf.com) | Repository (GitHub) | License: MIT

vision-cli migan -m MIGAN-512-places2-F16.gguf -i image.png mask.png -o output.png

Real-ESRGAN

example-esrgan

Model download | Paper (arXiv) | Repository (GitHub) | License: BSD-3-Clause

vision-cli esrgan -m ESRGAN-4x-foolhardy_Remacri-F16.gguf -i input.png -o output.png

Converting models

Models need to be converted to GGUF before they can be used. This will also rearrange or precompute tensors for more optimal inference.

To convert a model, install uv and run:

uv run scripts/convert.py <arch> MyModel.pth

where <arch> is one of sam, birefnet, esrgan, ....

This will create models/MyModel.gguf. See convert.py --help for more options.

Building

Building requires CMake and a compiler with C++20 support.

Get the sources

git clone https://github.com/Acly/vision.cpp.git --recursive
cd vision.cpp

Configure and build

cmake . -B build
cmake --build build --config Release

Vulkan (Optional)

Building with Vulkan GPU support requires the Vulkan SDK to be installed.

cmake . -B build -D VISP_VULKAN=ON

Tests (Optional)

Build with -DVISP_TESTS=ON. Run all C++ tests with the following command:

cd build
ctest -C Release

Some tests require a Python environment. It can be set up with uv:

# Setup venv and install dependencies (once only)
uv sync --dev
# Run python tests
uv run pytest

Performance

Performance optimization is an ongoing process. The aim is to be in the same ballpark as other frameworks for inference speed, but with:

  • much faster initialization and model loading time (<100 ms)
  • lower memory overhead
  • tiny deployment size (<5 MB for CPU, +30 MB for GPU)

Inference speed

  • CPU: AMD Ryzen 5 5600X (6 cores)
  • GPU: NVIDIA GeForce RTX 4070

MobileSAM, 1024x1024

vision.cppPyTorchONNX Runtime
cpuf32669 ms601 ms805 ms
gpuf1619 ms16 ms

BiRefNet, 1024x1024

Modelvision.cppPyTorchONNX Runtime
Fullcpuf3216333 ms18290 ms
Fullgpuf16208 ms190 ms
Litecpuf324505 ms10900 ms6978 ms
Litegpuf1685 ms84 ms

Depth-Anything, 518x714

Modelvision.cppPyTorch
Smallgpuf1611 ms10 ms
Basegpuf1624 ms22 ms

MI-GAN, 512x512

Modelvision.cppPyTorch
512-places2cpuf32523 ms637 ms
512-places2gpuf1621 ms17 ms

Setup

  • vision.cpp: using vision-bench, GPU via Vulkan, eg. vision-bench -m sam
  • PyTorch: v2.7.1+cu128, eager eval, GPU via CUDA, average n iterations after warm-up

Dependencies (integrated)

  • ggml - ML tensor library | MIT
  • stb-image - Image load/save/resize | Public Domain
  • fmt - String formatting (only if compiler doesn't support <format>) | MIT

About

Computer Vision ML inference in C++

Topics

Resources

Contributing

Stars

53 stars

Watchers

5 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

vision.cpp

Computer Vision ML inference in C++

  • Self-contained C++ library
  • Efficient inference on consumer CPU and GPUs (NVIDIA, AMD, Intel)
  • Lightweight deployment on many platforms (Windows, Linux, MacOS)
  • Growing number of supported models behind a simple API
  • Modular design for full control and implementing your own models

Based on ggml similar to the llama.cpp project.

Features

ModelTaskBackends
MobileSAMPromptable segmentationCPU, Vulkan
BiRefNetDichotomous segmentationCPU, Vulkan
Depth-AnythingDepth estimationCPU, Vulkan
MI-GANInpaintingCPU, Vulkan
ESRGANSuper-resolutionCPU, Vulkan
Implement a model [Guide]

Backbones: SWIN (v1), DINO (v2), TinyViT

Get Started

Get the library and executables:

Example: Select an object in an image

Let's use MobileSAM to generate a segmentation mask of the plushy on the right by passing in a box describing its approximate location.

Example image showing box prompt at pixel location (420, 120) - (650, 430), and the output mask

You can download the model and input image here: MobileSAM-F16.gguf | input.jpg

CLI

Find the vision-cli executable in the bin folder and run it to generate the mask:

vision-cli -m MobileSAM-F16.gguf -i input.jpg -p 420 120 650 430 -o mask.png

Pass --composite output.png to composite input and mask. Use --help for more options.

API

#include<visp/vision.h>usingnamespacevisp;voidmain() {
backend_device cpu = backend_init(backend_type::cpu);
sam_model sam = sam_load_model("MobileSAM-F16.gguf", cpu);
image_data input_image = image_load("input.jpg");
sam_encode(sam, input_image);
image_data object_mask = sam_compute(sam, box_2d{{420, 120}, {650, 320}});
image_save(object_mask, "mask.png");
}

This shows the high-level API. Internally it is composed of multiple smaller functions that handle model loading, pre-processing inputs, transferring data to backend devices, post-processing output, etc. These can be used as building blocks for flexible functions which integrate with your existing data sources and infrastructure.

Models

MobileSAM

example-sam

Model download | Paper (arXiv) | Repository (GitHub) | Segment-Anything-Model | License: Apache-2

vision-cli sam -m MobileSAM-F16.gguf -i input.png -p 300 200 -o mask.png --composite comp.png

BiRefNet

example-birefnet

Model download | Paper (arXiv) | Repository (GitHub) | License: MIT

vision-cli birefnet -m BiRefNet-lite-F16.gguf -i input.png -o mask.png --composite comp.png

Depth-Anything V2

example-depth-anything

Model download | Paper (arXiv) | Repository (GitHub) | License: Apache-2 / CC-BY-NC-4

vision-cli depth-anything -m Depth-Anything-V2-Small-F16.gguf -i input.png -o depth.png

MI-GAN

example-migan

Model download | Paper (thecvf.com) | Repository (GitHub) | License: MIT

vision-cli migan -m MIGAN-512-places2-F16.gguf -i image.png mask.png -o output.png

Real-ESRGAN

example-esrgan

Model download | Paper (arXiv) | Repository (GitHub) | License: BSD-3-Clause

vision-cli esrgan -m ESRGAN-4x-foolhardy_Remacri-F16.gguf -i input.png -o output.png

Converting models

Models need to be converted to GGUF before they can be used. This will also rearrange or precompute tensors for more optimal inference.

To convert a model, install uv and run:

uv run scripts/convert.py <arch> MyModel.pth

where <arch> is one of sam, birefnet, esrgan, ....

This will create models/MyModel.gguf. See convert.py --help for more options.

Building

Building requires CMake and a compiler with C++20 support.

Get the sources

git clone https://github.com/Acly/vision.cpp.git --recursive
cd vision.cpp

Configure and build

cmake . -B build
cmake --build build --config Release

Vulkan (Optional)

Building with Vulkan GPU support requires the Vulkan SDK to be installed.

cmake . -B build -D VISP_VULKAN=ON

Tests (Optional)

Build with -DVISP_TESTS=ON. Run all C++ tests with the following command:

cd build
ctest -C Release

Some tests require a Python environment. It can be set up with uv:

# Setup venv and install dependencies (once only)
uv sync --dev
# Run python tests
uv run pytest

Performance

Performance optimization is an ongoing process. The aim is to be in the same ballpark as other frameworks for inference speed, but with:

  • much faster initialization and model loading time (<100 ms)
  • lower memory overhead
  • tiny deployment size (<5 MB for CPU, +30 MB for GPU)

Inference speed

  • CPU: AMD Ryzen 5 5600X (6 cores)
  • GPU: NVIDIA GeForce RTX 4070

MobileSAM, 1024x1024

vision.cppPyTorchONNX Runtime
cpuf32669 ms601 ms805 ms
gpuf1619 ms16 ms

BiRefNet, 1024x1024

Modelvision.cppPyTorchONNX Runtime
Fullcpuf3216333 ms18290 ms
Fullgpuf16208 ms190 ms
Litecpuf324505 ms10900 ms6978 ms
Litegpuf1685 ms84 ms

Depth-Anything, 518x714

Modelvision.cppPyTorch
Smallgpuf1611 ms10 ms
Basegpuf1624 ms22 ms

MI-GAN, 512x512

Modelvision.cppPyTorch
512-places2cpuf32523 ms637 ms
512-places2gpuf1621 ms17 ms

Setup

  • vision.cpp: using vision-bench, GPU via Vulkan, eg. vision-bench -m sam
  • PyTorch: v2.7.1+cu128, eager eval, GPU via CUDA, average n iterations after warm-up

Dependencies (integrated)

  • ggml - ML tensor library | MIT
  • stb-image - Image load/save/resize | Public Domain
  • fmt - String formatting (only if compiler doesn't support <format>) | MIT

About

Computer Vision ML inference in C++

Topics

Resources

Contributing

Stars

53 stars

Watchers

5 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

vision.cpp

Computer Vision ML inference in C++

  • Self-contained C++ library
  • Efficient inference on consumer CPU and GPUs (NVIDIA, AMD, Intel)
  • Lightweight deployment on many platforms (Windows, Linux, MacOS)
  • Growing number of supported models behind a simple API
  • Modular design for full control and implementing your own models

Based on ggml similar to the llama.cpp project.

Features

ModelTaskBackends
MobileSAMPromptable segmentationCPU, Vulkan
BiRefNetDichotomous segmentationCPU, Vulkan
Depth-AnythingDepth estimationCPU, Vulkan
MI-GANInpaintingCPU, Vulkan
ESRGANSuper-resolutionCPU, Vulkan
Implement a model [Guide]

Backbones: SWIN (v1), DINO (v2), TinyViT

Get Started

Get the library and executables:

Example: Select an object in an image

Let's use MobileSAM to generate a segmentation mask of the plushy on the right by passing in a box describing its approximate location.

Example image showing box prompt at pixel location (420, 120) - (650, 430), and the output mask

You can download the model and input image here: MobileSAM-F16.gguf | input.jpg

CLI

Find the vision-cli executable in the bin folder and run it to generate the mask:

vision-cli -m MobileSAM-F16.gguf -i input.jpg -p 420 120 650 430 -o mask.png

Pass --composite output.png to composite input and mask. Use --help for more options.

API

#include<visp/vision.h>usingnamespacevisp;voidmain() {
backend_device cpu = backend_init(backend_type::cpu);
sam_model sam = sam_load_model("MobileSAM-F16.gguf", cpu);
image_data input_image = image_load("input.jpg");
sam_encode(sam, input_image);
image_data object_mask = sam_compute(sam, box_2d{{420, 120}, {650, 320}});
image_save(object_mask, "mask.png");
}

This shows the high-level API. Internally it is composed of multiple smaller functions that handle model loading, pre-processing inputs, transferring data to backend devices, post-processing output, etc. These can be used as building blocks for flexible functions which integrate with your existing data sources and infrastructure.

Models

MobileSAM

example-sam

Model download | Paper (arXiv) | Repository (GitHub) | Segment-Anything-Model | License: Apache-2

vision-cli sam -m MobileSAM-F16.gguf -i input.png -p 300 200 -o mask.png --composite comp.png

BiRefNet

example-birefnet

Model download | Paper (arXiv) | Repository (GitHub) | License: MIT

vision-cli birefnet -m BiRefNet-lite-F16.gguf -i input.png -o mask.png --composite comp.png

Depth-Anything V2

example-depth-anything

Model download | Paper (arXiv) | Repository (GitHub) | License: Apache-2 / CC-BY-NC-4

vision-cli depth-anything -m Depth-Anything-V2-Small-F16.gguf -i input.png -o depth.png

MI-GAN

example-migan

Model download | Paper (thecvf.com) | Repository (GitHub) | License: MIT

vision-cli migan -m MIGAN-512-places2-F16.gguf -i image.png mask.png -o output.png

Real-ESRGAN

example-esrgan

Model download | Paper (arXiv) | Repository (GitHub) | License: BSD-3-Clause

vision-cli esrgan -m ESRGAN-4x-foolhardy_Remacri-F16.gguf -i input.png -o output.png

Converting models

Models need to be converted to GGUF before they can be used. This will also rearrange or precompute tensors for more optimal inference.

To convert a model, install uv and run:

uv run scripts/convert.py <arch> MyModel.pth

where <arch> is one of sam, birefnet, esrgan, ....

This will create models/MyModel.gguf. See convert.py --help for more options.

Building

Building requires CMake and a compiler with C++20 support.

Get the sources

git clone https://github.com/Acly/vision.cpp.git --recursive
cd vision.cpp

Configure and build

cmake . -B build
cmake --build build --config Release

Vulkan (Optional)

Building with Vulkan GPU support requires the Vulkan SDK to be installed.

cmake . -B build -D VISP_VULKAN=ON

Tests (Optional)

Build with -DVISP_TESTS=ON. Run all C++ tests with the following command:

cd build
ctest -C Release

Some tests require a Python environment. It can be set up with uv:

# Setup venv and install dependencies (once only)
uv sync --dev
# Run python tests
uv run pytest

Performance

Performance optimization is an ongoing process. The aim is to be in the same ballpark as other frameworks for inference speed, but with:

  • much faster initialization and model loading time (<100 ms)
  • lower memory overhead
  • tiny deployment size (<5 MB for CPU, +30 MB for GPU)

Inference speed

  • CPU: AMD Ryzen 5 5600X (6 cores)
  • GPU: NVIDIA GeForce RTX 4070

MobileSAM, 1024x1024

vision.cppPyTorchONNX Runtime
cpuf32669 ms601 ms805 ms
gpuf1619 ms16 ms

BiRefNet, 1024x1024

Modelvision.cppPyTorchONNX Runtime
Fullcpuf3216333 ms18290 ms
Fullgpuf16208 ms190 ms
Litecpuf324505 ms10900 ms6978 ms
Litegpuf1685 ms84 ms

Depth-Anything, 518x714

Modelvision.cppPyTorch
Smallgpuf1611 ms10 ms
Basegpuf1624 ms22 ms

MI-GAN, 512x512

Modelvision.cppPyTorch
512-places2cpuf32523 ms637 ms
512-places2gpuf1621 ms17 ms

Setup

  • vision.cpp: using vision-bench, GPU via Vulkan, eg. vision-bench -m sam
  • PyTorch: v2.7.1+cu128, eager eval, GPU via CUDA, average n iterations after warm-up

Dependencies (integrated)

  • ggml - ML tensor library | MIT
  • stb-image - Image load/save/resize | Public Domain
  • fmt - String formatting (only if compiler doesn't support <format>) | MIT

About

Computer Vision ML inference in C++

Topics

Resources

Contributing

Stars

53 stars

Watchers

5 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

vision.cpp

Computer Vision ML inference in C++

  • Self-contained C++ library
  • Efficient inference on consumer CPU and GPUs (NVIDIA, AMD, Intel)
  • Lightweight deployment on many platforms (Windows, Linux, MacOS)
  • Growing number of supported models behind a simple API
  • Modular design for full control and implementing your own models

Based on ggml similar to the llama.cpp project.

Features

ModelTaskBackends
MobileSAMPromptable segmentationCPU, Vulkan
BiRefNetDichotomous segmentationCPU, Vulkan
Depth-AnythingDepth estimationCPU, Vulkan
MI-GANInpaintingCPU, Vulkan
ESRGANSuper-resolutionCPU, Vulkan
Implement a model [Guide]

Backbones: SWIN (v1), DINO (v2), TinyViT

Get Started

Get the library and executables:

Example: Select an object in an image

Let's use MobileSAM to generate a segmentation mask of the plushy on the right by passing in a box describing its approximate location.

Example image showing box prompt at pixel location (420, 120) - (650, 430), and the output mask

You can download the model and input image here: MobileSAM-F16.gguf | input.jpg

CLI

Find the vision-cli executable in the bin folder and run it to generate the mask:

vision-cli -m MobileSAM-F16.gguf -i input.jpg -p 420 120 650 430 -o mask.png

Pass --composite output.png to composite input and mask. Use --help for more options.

API

#include<visp/vision.h>usingnamespacevisp;voidmain() {
backend_device cpu = backend_init(backend_type::cpu);
sam_model sam = sam_load_model("MobileSAM-F16.gguf", cpu);
image_data input_image = image_load("input.jpg");
sam_encode(sam, input_image);
image_data object_mask = sam_compute(sam, box_2d{{420, 120}, {650, 320}});
image_save(object_mask, "mask.png");
}

This shows the high-level API. Internally it is composed of multiple smaller functions that handle model loading, pre-processing inputs, transferring data to backend devices, post-processing output, etc. These can be used as building blocks for flexible functions which integrate with your existing data sources and infrastructure.

Models

MobileSAM

example-sam

Model download | Paper (arXiv) | Repository (GitHub) | Segment-Anything-Model | License: Apache-2

vision-cli sam -m MobileSAM-F16.gguf -i input.png -p 300 200 -o mask.png --composite comp.png

BiRefNet

example-birefnet

Model download | Paper (arXiv) | Repository (GitHub) | License: MIT

vision-cli birefnet -m BiRefNet-lite-F16.gguf -i input.png -o mask.png --composite comp.png

Depth-Anything V2

example-depth-anything

Model download | Paper (arXiv) | Repository (GitHub) | License: Apache-2 / CC-BY-NC-4

vision-cli depth-anything -m Depth-Anything-V2-Small-F16.gguf -i input.png -o depth.png

MI-GAN

example-migan

Model download | Paper (thecvf.com) | Repository (GitHub) | License: MIT

vision-cli migan -m MIGAN-512-places2-F16.gguf -i image.png mask.png -o output.png

Real-ESRGAN

example-esrgan

Model download | Paper (arXiv) | Repository (GitHub) | License: BSD-3-Clause

vision-cli esrgan -m ESRGAN-4x-foolhardy_Remacri-F16.gguf -i input.png -o output.png

Converting models

Models need to be converted to GGUF before they can be used. This will also rearrange or precompute tensors for more optimal inference.

To convert a model, install uv and run:

uv run scripts/convert.py <arch> MyModel.pth

where <arch> is one of sam, birefnet, esrgan, ....

This will create models/MyModel.gguf. See convert.py --help for more options.

Building

Building requires CMake and a compiler with C++20 support.

Get the sources

git clone https://github.com/Acly/vision.cpp.git --recursive
cd vision.cpp

Configure and build

cmake . -B build
cmake --build build --config Release

Vulkan (Optional)

Building with Vulkan GPU support requires the Vulkan SDK to be installed.

cmake . -B build -D VISP_VULKAN=ON

Tests (Optional)

Build with -DVISP_TESTS=ON. Run all C++ tests with the following command:

cd build
ctest -C Release

Some tests require a Python environment. It can be set up with uv:

# Setup venv and install dependencies (once only)
uv sync --dev
# Run python tests
uv run pytest

Performance

Performance optimization is an ongoing process. The aim is to be in the same ballpark as other frameworks for inference speed, but with:

  • much faster initialization and model loading time (<100 ms)
  • lower memory overhead
  • tiny deployment size (<5 MB for CPU, +30 MB for GPU)

Inference speed

  • CPU: AMD Ryzen 5 5600X (6 cores)
  • GPU: NVIDIA GeForce RTX 4070

MobileSAM, 1024x1024

vision.cppPyTorchONNX Runtime
cpuf32669 ms601 ms805 ms
gpuf1619 ms16 ms

BiRefNet, 1024x1024

Modelvision.cppPyTorchONNX Runtime
Fullcpuf3216333 ms18290 ms
Fullgpuf16208 ms190 ms
Litecpuf324505 ms10900 ms6978 ms
Litegpuf1685 ms84 ms

Depth-Anything, 518x714

Modelvision.cppPyTorch
Smallgpuf1611 ms10 ms
Basegpuf1624 ms22 ms

MI-GAN, 512x512

Modelvision.cppPyTorch
512-places2cpuf32523 ms637 ms
512-places2gpuf1621 ms17 ms

Setup

  • vision.cpp: using vision-bench, GPU via Vulkan, eg. vision-bench -m sam
  • PyTorch: v2.7.1+cu128, eager eval, GPU via CUDA, average n iterations after warm-up

Dependencies (integrated)

  • ggml - ML tensor library | MIT
  • stb-image - Image load/save/resize | Public Domain
  • fmt - String formatting (only if compiler doesn't support <format>) | MIT

About

Computer Vision ML inference in C++

Topics

Resources

Contributing

Stars

53 stars

Watchers

5 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

vision.cpp

Computer Vision ML inference in C++

  • Self-contained C++ library
  • Efficient inference on consumer CPU and GPUs (NVIDIA, AMD, Intel)
  • Lightweight deployment on many platforms (Windows, Linux, MacOS)
  • Growing number of supported models behind a simple API
  • Modular design for full control and implementing your own models

Based on ggml similar to the llama.cpp project.

Features

ModelTaskBackends
MobileSAMPromptable segmentationCPU, Vulkan
BiRefNetDichotomous segmentationCPU, Vulkan
Depth-AnythingDepth estimationCPU, Vulkan
MI-GANInpaintingCPU, Vulkan
ESRGANSuper-resolutionCPU, Vulkan
Implement a model [Guide]

Backbones: SWIN (v1), DINO (v2), TinyViT

Get Started

Get the library and executables:

Example: Select an object in an image

Let's use MobileSAM to generate a segmentation mask of the plushy on the right by passing in a box describing its approximate location.

Example image showing box prompt at pixel location (420, 120) - (650, 430), and the output mask

You can download the model and input image here: MobileSAM-F16.gguf | input.jpg

CLI

Find the vision-cli executable in the bin folder and run it to generate the mask:

vision-cli -m MobileSAM-F16.gguf -i input.jpg -p 420 120 650 430 -o mask.png

Pass --composite output.png to composite input and mask. Use --help for more options.

API

#include<visp/vision.h>usingnamespacevisp;voidmain() {
backend_device cpu = backend_init(backend_type::cpu);
sam_model sam = sam_load_model("MobileSAM-F16.gguf", cpu);
image_data input_image = image_load("input.jpg");
sam_encode(sam, input_image);
image_data object_mask = sam_compute(sam, box_2d{{420, 120}, {650, 320}});
image_save(object_mask, "mask.png");
}

This shows the high-level API. Internally it is composed of multiple smaller functions that handle model loading, pre-processing inputs, transferring data to backend devices, post-processing output, etc. These can be used as building blocks for flexible functions which integrate with your existing data sources and infrastructure.

Models

MobileSAM

example-sam

Model download | Paper (arXiv) | Repository (GitHub) | Segment-Anything-Model | License: Apache-2

vision-cli sam -m MobileSAM-F16.gguf -i input.png -p 300 200 -o mask.png --composite comp.png

BiRefNet

example-birefnet

Model download | Paper (arXiv) | Repository (GitHub) | License: MIT

vision-cli birefnet -m BiRefNet-lite-F16.gguf -i input.png -o mask.png --composite comp.png

Depth-Anything V2

example-depth-anything

Model download | Paper (arXiv) | Repository (GitHub) | License: Apache-2 / CC-BY-NC-4

vision-cli depth-anything -m Depth-Anything-V2-Small-F16.gguf -i input.png -o depth.png

MI-GAN

example-migan

Model download | Paper (thecvf.com) | Repository (GitHub) | License: MIT

vision-cli migan -m MIGAN-512-places2-F16.gguf -i image.png mask.png -o output.png

Real-ESRGAN

example-esrgan

Model download | Paper (arXiv) | Repository (GitHub) | License: BSD-3-Clause

vision-cli esrgan -m ESRGAN-4x-foolhardy_Remacri-F16.gguf -i input.png -o output.png

Converting models

Models need to be converted to GGUF before they can be used. This will also rearrange or precompute tensors for more optimal inference.

To convert a model, install uv and run:

uv run scripts/convert.py <arch> MyModel.pth

where <arch> is one of sam, birefnet, esrgan, ....

This will create models/MyModel.gguf. See convert.py --help for more options.

Building

Building requires CMake and a compiler with C++20 support.

Get the sources

git clone https://github.com/Acly/vision.cpp.git --recursive
cd vision.cpp

Configure and build

cmake . -B build
cmake --build build --config Release

Vulkan (Optional)

Building with Vulkan GPU support requires the Vulkan SDK to be installed.

cmake . -B build -D VISP_VULKAN=ON

Tests (Optional)

Build with -DVISP_TESTS=ON. Run all C++ tests with the following command:

cd build
ctest -C Release

Some tests require a Python environment. It can be set up with uv:

# Setup venv and install dependencies (once only)
uv sync --dev
# Run python tests
uv run pytest

Performance

Performance optimization is an ongoing process. The aim is to be in the same ballpark as other frameworks for inference speed, but with:

  • much faster initialization and model loading time (<100 ms)
  • lower memory overhead
  • tiny deployment size (<5 MB for CPU, +30 MB for GPU)

Inference speed

  • CPU: AMD Ryzen 5 5600X (6 cores)
  • GPU: NVIDIA GeForce RTX 4070

MobileSAM, 1024x1024

vision.cppPyTorchONNX Runtime
cpuf32669 ms601 ms805 ms
gpuf1619 ms16 ms

BiRefNet, 1024x1024

Modelvision.cppPyTorchONNX Runtime
Fullcpuf3216333 ms18290 ms
Fullgpuf16208 ms190 ms
Litecpuf324505 ms10900 ms6978 ms
Litegpuf1685 ms84 ms

Depth-Anything, 518x714

Modelvision.cppPyTorch
Smallgpuf1611 ms10 ms
Basegpuf1624 ms22 ms

MI-GAN, 512x512

Modelvision.cppPyTorch
512-places2cpuf32523 ms637 ms
512-places2gpuf1621 ms17 ms

Setup

  • vision.cpp: using vision-bench, GPU via Vulkan, eg. vision-bench -m sam
  • PyTorch: v2.7.1+cu128, eager eval, GPU via CUDA, average n iterations after warm-up

Dependencies (integrated)

  • ggml - ML tensor library | MIT
  • stb-image - Image load/save/resize | Public Domain
  • fmt - String formatting (only if compiler doesn't support <format>) | MIT

About

Computer Vision ML inference in C++

Topics

Resources

Contributing

Stars

53 stars

Watchers

5 watching

Forks

Releases

Packages

Contributors

Languages