Repository files navigation

🏭 FastVision: Fast Industrial Vision Platform (C++)

High-throughput, low-latency industrial vision platform based on TensorRT 10 and CUDA 11

C++17CUDATensorRTOpenCVImGuiWindows

English | 中文

📄 Paper🐍 Python Prototype


📸 System Demo

System Demo
Real-time three-screen display: Original Capture (Left) | Anomaly Heatmap (Middle) | Defect Overlay (Right)

✨ Core Highlights

🚀 Extreme Performance

Pure CUDA operator post-processing + OpenGL zero-copy rendering
Inference latency < 15ms (80+ FPS)

🎯 Anomalib Compatible

Fully compatible with Anomalib ecosystem
Seamless loading of standard models, supporting pixel-level defect segmentation and localization

🖥️ Interactive Panel

Modern console built with ImGui
Support for real-time threshold adjustment and model hot-switching

📦 Ready to Use

Provides Windows one-click installer (Setup.exe)
No need to configure Python/CUDA environment, just double-click to run


⚡ Quick Start

👥 I'm an End User

Don't want to code, just want to run the software?

  1. Download the latest release: 👉 Download (Release)
  2. Need help? Check the documentation: 📖 User Manual

👨‍💻 I'm a Developer

Want to modify the source code or develop secondary applications?

Please refer to the compilation and build guide: 🛠️ Developer Environment Setup


📌 Performance Benchmark

We conducted comprehensive tests on an RTX 3060 platform and achieved more than 2x throughput improvement compared to the original implementation.

Optimization StagePreprocessing (ms)Inference (ms)Postprocessing (ms)Total Latency (ms)Throughput (FPS)
Original C++ Implementation1117533~30
+ CUDA Preprocessing417526~38
+ Zero-copy Postprocessing417324~42
+ FP16 Quantization48315~67
+ Rendering Pipeline Optimization483< 15> 80

🏗️ System Architecture

The system adopts a modular design, achieving efficient interoperation between computation (CUDA) and display (OpenGL).

  • 🧠 Inference Core
    • Encapsulates TensorRT 10, supports FP32/FP16 dynamic precision.
    • Multi-threaded pipeline design, separating input IO from GPU computation.
  • ⚡ CUDA Acceleration Layer
    • Preprocessor: Color space conversion, normalization, Resize (NPP).
    • Postprocessor: Anomaly map generation, threshold segmentation, heatmap rendering (Custom Kernels).
  • 🎨 Visualization & Interaction
    • Dashboard: Control panel based on ImGui.
    • Renderer: Uses CUDA-OpenGL Interop to directly map VRAM textures, eliminating CPU-GPU bandwidth bottleneck.

📅 Changelog

v1.0.0 - 2026-01-28: Architecture Refactoring and Performance Optimization (Click to expand)
  • New Modular Architecture: Refactored configuration management, rendering, inference engine, and other independent modules.
  • Rendering Optimization: Implemented CUDA-OpenGL interoperability, eliminating Host-to-Device copy overhead.
  • Pipeline Enhancement: Multi-threaded Pipeline, support for FP16 acceleration.
  • UI Upgrade: Integrated ImGui for interactive parameter adjustment.

For complete records, please refer to CHANGELOG.md


🤝 Acknowledgements & Feedback

If this project helps your research or work, please give it a ⭐ Star on GitHub!

If you find any bugs or have improvement suggestions, please submit an Issue or Pull Request.

📚 Reference

If you find this project useful in your research or work, please cite this project:

@article{liao2026limr,
author = {Shaowei Liao and Wenyong Yu and Shaolin Liao},
title = {Lightweight Masked Reconstruction for Real-Time Sensor-Driven
Anomaly Detection in Industrial IoT},
journal = {IEEE Internet of Things Journal},
year = {2026},
doi = {10.1109/JIOT.2026.3712733}
}

📜 License

This repository is licensed under the Apache-2.0 License.

About

Real-time Industrial Anomaly Defect Inference Detection implemented by cpp(实时工业缺陷检测cpp)

Topics

Resources

Stars

7 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

🏭 FastVision: Fast Industrial Vision Platform (C++)

High-throughput, low-latency industrial vision platform based on TensorRT 10 and CUDA 11

C++17CUDATensorRTOpenCVImGuiWindows

English | 中文

📄 Paper🐍 Python Prototype


📸 System Demo

System Demo
Real-time three-screen display: Original Capture (Left) | Anomaly Heatmap (Middle) | Defect Overlay (Right)

✨ Core Highlights

🚀 Extreme Performance

Pure CUDA operator post-processing + OpenGL zero-copy rendering
Inference latency < 15ms (80+ FPS)

🎯 Anomalib Compatible

Fully compatible with Anomalib ecosystem
Seamless loading of standard models, supporting pixel-level defect segmentation and localization

🖥️ Interactive Panel

Modern console built with ImGui
Support for real-time threshold adjustment and model hot-switching

📦 Ready to Use

Provides Windows one-click installer (Setup.exe)
No need to configure Python/CUDA environment, just double-click to run


⚡ Quick Start

👥 I'm an End User

Don't want to code, just want to run the software?

  1. Download the latest release: 👉 Download (Release)
  2. Need help? Check the documentation: 📖 User Manual

👨‍💻 I'm a Developer

Want to modify the source code or develop secondary applications?

Please refer to the compilation and build guide: 🛠️ Developer Environment Setup


📌 Performance Benchmark

We conducted comprehensive tests on an RTX 3060 platform and achieved more than 2x throughput improvement compared to the original implementation.

Optimization StagePreprocessing (ms)Inference (ms)Postprocessing (ms)Total Latency (ms)Throughput (FPS)
Original C++ Implementation1117533~30
+ CUDA Preprocessing417526~38
+ Zero-copy Postprocessing417324~42
+ FP16 Quantization48315~67
+ Rendering Pipeline Optimization483< 15> 80

🏗️ System Architecture

The system adopts a modular design, achieving efficient interoperation between computation (CUDA) and display (OpenGL).

  • 🧠 Inference Core
    • Encapsulates TensorRT 10, supports FP32/FP16 dynamic precision.
    • Multi-threaded pipeline design, separating input IO from GPU computation.
  • ⚡ CUDA Acceleration Layer
    • Preprocessor: Color space conversion, normalization, Resize (NPP).
    • Postprocessor: Anomaly map generation, threshold segmentation, heatmap rendering (Custom Kernels).
  • 🎨 Visualization & Interaction
    • Dashboard: Control panel based on ImGui.
    • Renderer: Uses CUDA-OpenGL Interop to directly map VRAM textures, eliminating CPU-GPU bandwidth bottleneck.

📅 Changelog

v1.0.0 - 2026-01-28: Architecture Refactoring and Performance Optimization (Click to expand)
  • New Modular Architecture: Refactored configuration management, rendering, inference engine, and other independent modules.
  • Rendering Optimization: Implemented CUDA-OpenGL interoperability, eliminating Host-to-Device copy overhead.
  • Pipeline Enhancement: Multi-threaded Pipeline, support for FP16 acceleration.
  • UI Upgrade: Integrated ImGui for interactive parameter adjustment.

For complete records, please refer to CHANGELOG.md


🤝 Acknowledgements & Feedback

If this project helps your research or work, please give it a ⭐ Star on GitHub!

If you find any bugs or have improvement suggestions, please submit an Issue or Pull Request.

📚 Reference

If you find this project useful in your research or work, please cite this project:

@article{liao2026limr,
author = {Shaowei Liao and Wenyong Yu and Shaolin Liao},
title = {Lightweight Masked Reconstruction for Real-Time Sensor-Driven
Anomaly Detection in Industrial IoT},
journal = {IEEE Internet of Things Journal},
year = {2026},
doi = {10.1109/JIOT.2026.3712733}
}

📜 License

This repository is licensed under the Apache-2.0 License.

About

Real-time Industrial Anomaly Defect Inference Detection implemented by cpp(实时工业缺陷检测cpp)

Topics

Resources

Stars

7 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

🏭 FastVision: Fast Industrial Vision Platform (C++)

High-throughput, low-latency industrial vision platform based on TensorRT 10 and CUDA 11

C++17CUDATensorRTOpenCVImGuiWindows

English | 中文

📄 Paper🐍 Python Prototype


📸 System Demo

System Demo
Real-time three-screen display: Original Capture (Left) | Anomaly Heatmap (Middle) | Defect Overlay (Right)

✨ Core Highlights

🚀 Extreme Performance

Pure CUDA operator post-processing + OpenGL zero-copy rendering
Inference latency < 15ms (80+ FPS)

🎯 Anomalib Compatible

Fully compatible with Anomalib ecosystem
Seamless loading of standard models, supporting pixel-level defect segmentation and localization

🖥️ Interactive Panel

Modern console built with ImGui
Support for real-time threshold adjustment and model hot-switching

📦 Ready to Use

Provides Windows one-click installer (Setup.exe)
No need to configure Python/CUDA environment, just double-click to run


⚡ Quick Start

👥 I'm an End User

Don't want to code, just want to run the software?

  1. Download the latest release: 👉 Download (Release)
  2. Need help? Check the documentation: 📖 User Manual

👨‍💻 I'm a Developer

Want to modify the source code or develop secondary applications?

Please refer to the compilation and build guide: 🛠️ Developer Environment Setup


📌 Performance Benchmark

We conducted comprehensive tests on an RTX 3060 platform and achieved more than 2x throughput improvement compared to the original implementation.

Optimization StagePreprocessing (ms)Inference (ms)Postprocessing (ms)Total Latency (ms)Throughput (FPS)
Original C++ Implementation1117533~30
+ CUDA Preprocessing417526~38
+ Zero-copy Postprocessing417324~42
+ FP16 Quantization48315~67
+ Rendering Pipeline Optimization483< 15> 80

🏗️ System Architecture

The system adopts a modular design, achieving efficient interoperation between computation (CUDA) and display (OpenGL).

  • 🧠 Inference Core
    • Encapsulates TensorRT 10, supports FP32/FP16 dynamic precision.
    • Multi-threaded pipeline design, separating input IO from GPU computation.
  • ⚡ CUDA Acceleration Layer
    • Preprocessor: Color space conversion, normalization, Resize (NPP).
    • Postprocessor: Anomaly map generation, threshold segmentation, heatmap rendering (Custom Kernels).
  • 🎨 Visualization & Interaction
    • Dashboard: Control panel based on ImGui.
    • Renderer: Uses CUDA-OpenGL Interop to directly map VRAM textures, eliminating CPU-GPU bandwidth bottleneck.

📅 Changelog

v1.0.0 - 2026-01-28: Architecture Refactoring and Performance Optimization (Click to expand)
  • New Modular Architecture: Refactored configuration management, rendering, inference engine, and other independent modules.
  • Rendering Optimization: Implemented CUDA-OpenGL interoperability, eliminating Host-to-Device copy overhead.
  • Pipeline Enhancement: Multi-threaded Pipeline, support for FP16 acceleration.
  • UI Upgrade: Integrated ImGui for interactive parameter adjustment.

For complete records, please refer to CHANGELOG.md


🤝 Acknowledgements & Feedback

If this project helps your research or work, please give it a ⭐ Star on GitHub!

If you find any bugs or have improvement suggestions, please submit an Issue or Pull Request.

📚 Reference

If you find this project useful in your research or work, please cite this project:

@article{liao2026limr,
author = {Shaowei Liao and Wenyong Yu and Shaolin Liao},
title = {Lightweight Masked Reconstruction for Real-Time Sensor-Driven
Anomaly Detection in Industrial IoT},
journal = {IEEE Internet of Things Journal},
year = {2026},
doi = {10.1109/JIOT.2026.3712733}
}

📜 License

This repository is licensed under the Apache-2.0 License.

About

Real-time Industrial Anomaly Defect Inference Detection implemented by cpp(实时工业缺陷检测cpp)

Topics

Resources

Stars

7 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

🏭 FastVision: Fast Industrial Vision Platform (C++)

High-throughput, low-latency industrial vision platform based on TensorRT 10 and CUDA 11

C++17CUDATensorRTOpenCVImGuiWindows

English | 中文

📄 Paper🐍 Python Prototype


📸 System Demo

System Demo
Real-time three-screen display: Original Capture (Left) | Anomaly Heatmap (Middle) | Defect Overlay (Right)

✨ Core Highlights

🚀 Extreme Performance

Pure CUDA operator post-processing + OpenGL zero-copy rendering
Inference latency < 15ms (80+ FPS)

🎯 Anomalib Compatible

Fully compatible with Anomalib ecosystem
Seamless loading of standard models, supporting pixel-level defect segmentation and localization

🖥️ Interactive Panel

Modern console built with ImGui
Support for real-time threshold adjustment and model hot-switching

📦 Ready to Use

Provides Windows one-click installer (Setup.exe)
No need to configure Python/CUDA environment, just double-click to run


⚡ Quick Start

👥 I'm an End User

Don't want to code, just want to run the software?

  1. Download the latest release: 👉 Download (Release)
  2. Need help? Check the documentation: 📖 User Manual

👨‍💻 I'm a Developer

Want to modify the source code or develop secondary applications?

Please refer to the compilation and build guide: 🛠️ Developer Environment Setup


📌 Performance Benchmark

We conducted comprehensive tests on an RTX 3060 platform and achieved more than 2x throughput improvement compared to the original implementation.

Optimization StagePreprocessing (ms)Inference (ms)Postprocessing (ms)Total Latency (ms)Throughput (FPS)
Original C++ Implementation1117533~30
+ CUDA Preprocessing417526~38
+ Zero-copy Postprocessing417324~42
+ FP16 Quantization48315~67
+ Rendering Pipeline Optimization483< 15> 80

🏗️ System Architecture

The system adopts a modular design, achieving efficient interoperation between computation (CUDA) and display (OpenGL).

  • 🧠 Inference Core
    • Encapsulates TensorRT 10, supports FP32/FP16 dynamic precision.
    • Multi-threaded pipeline design, separating input IO from GPU computation.
  • ⚡ CUDA Acceleration Layer
    • Preprocessor: Color space conversion, normalization, Resize (NPP).
    • Postprocessor: Anomaly map generation, threshold segmentation, heatmap rendering (Custom Kernels).
  • 🎨 Visualization & Interaction
    • Dashboard: Control panel based on ImGui.
    • Renderer: Uses CUDA-OpenGL Interop to directly map VRAM textures, eliminating CPU-GPU bandwidth bottleneck.

📅 Changelog

v1.0.0 - 2026-01-28: Architecture Refactoring and Performance Optimization (Click to expand)
  • New Modular Architecture: Refactored configuration management, rendering, inference engine, and other independent modules.
  • Rendering Optimization: Implemented CUDA-OpenGL interoperability, eliminating Host-to-Device copy overhead.
  • Pipeline Enhancement: Multi-threaded Pipeline, support for FP16 acceleration.
  • UI Upgrade: Integrated ImGui for interactive parameter adjustment.

For complete records, please refer to CHANGELOG.md


🤝 Acknowledgements & Feedback

If this project helps your research or work, please give it a ⭐ Star on GitHub!

If you find any bugs or have improvement suggestions, please submit an Issue or Pull Request.

📚 Reference

If you find this project useful in your research or work, please cite this project:

@article{liao2026limr,
author = {Shaowei Liao and Wenyong Yu and Shaolin Liao},
title = {Lightweight Masked Reconstruction for Real-Time Sensor-Driven
Anomaly Detection in Industrial IoT},
journal = {IEEE Internet of Things Journal},
year = {2026},
doi = {10.1109/JIOT.2026.3712733}
}

📜 License

This repository is licensed under the Apache-2.0 License.

About

Real-time Industrial Anomaly Defect Inference Detection implemented by cpp(实时工业缺陷检测cpp)

Topics

Resources

Stars

7 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

🏭 FastVision: Fast Industrial Vision Platform (C++)

High-throughput, low-latency industrial vision platform based on TensorRT 10 and CUDA 11

C++17CUDATensorRTOpenCVImGuiWindows

English | 中文

📄 Paper🐍 Python Prototype


📸 System Demo

System Demo
Real-time three-screen display: Original Capture (Left) | Anomaly Heatmap (Middle) | Defect Overlay (Right)

✨ Core Highlights

🚀 Extreme Performance

Pure CUDA operator post-processing + OpenGL zero-copy rendering
Inference latency < 15ms (80+ FPS)

🎯 Anomalib Compatible

Fully compatible with Anomalib ecosystem
Seamless loading of standard models, supporting pixel-level defect segmentation and localization

🖥️ Interactive Panel

Modern console built with ImGui
Support for real-time threshold adjustment and model hot-switching

📦 Ready to Use

Provides Windows one-click installer (Setup.exe)
No need to configure Python/CUDA environment, just double-click to run


⚡ Quick Start

👥 I'm an End User

Don't want to code, just want to run the software?

  1. Download the latest release: 👉 Download (Release)
  2. Need help? Check the documentation: 📖 User Manual

👨‍💻 I'm a Developer

Want to modify the source code or develop secondary applications?

Please refer to the compilation and build guide: 🛠️ Developer Environment Setup


📌 Performance Benchmark

We conducted comprehensive tests on an RTX 3060 platform and achieved more than 2x throughput improvement compared to the original implementation.

Optimization StagePreprocessing (ms)Inference (ms)Postprocessing (ms)Total Latency (ms)Throughput (FPS)
Original C++ Implementation1117533~30
+ CUDA Preprocessing417526~38
+ Zero-copy Postprocessing417324~42
+ FP16 Quantization48315~67
+ Rendering Pipeline Optimization483< 15> 80

🏗️ System Architecture

The system adopts a modular design, achieving efficient interoperation between computation (CUDA) and display (OpenGL).

  • 🧠 Inference Core
    • Encapsulates TensorRT 10, supports FP32/FP16 dynamic precision.
    • Multi-threaded pipeline design, separating input IO from GPU computation.
  • ⚡ CUDA Acceleration Layer
    • Preprocessor: Color space conversion, normalization, Resize (NPP).
    • Postprocessor: Anomaly map generation, threshold segmentation, heatmap rendering (Custom Kernels).
  • 🎨 Visualization & Interaction
    • Dashboard: Control panel based on ImGui.
    • Renderer: Uses CUDA-OpenGL Interop to directly map VRAM textures, eliminating CPU-GPU bandwidth bottleneck.

📅 Changelog

v1.0.0 - 2026-01-28: Architecture Refactoring and Performance Optimization (Click to expand)
  • New Modular Architecture: Refactored configuration management, rendering, inference engine, and other independent modules.
  • Rendering Optimization: Implemented CUDA-OpenGL interoperability, eliminating Host-to-Device copy overhead.
  • Pipeline Enhancement: Multi-threaded Pipeline, support for FP16 acceleration.
  • UI Upgrade: Integrated ImGui for interactive parameter adjustment.

For complete records, please refer to CHANGELOG.md


🤝 Acknowledgements & Feedback

If this project helps your research or work, please give it a ⭐ Star on GitHub!

If you find any bugs or have improvement suggestions, please submit an Issue or Pull Request.

📚 Reference

If you find this project useful in your research or work, please cite this project:

@article{liao2026limr,
author = {Shaowei Liao and Wenyong Yu and Shaolin Liao},
title = {Lightweight Masked Reconstruction for Real-Time Sensor-Driven
Anomaly Detection in Industrial IoT},
journal = {IEEE Internet of Things Journal},
year = {2026},
doi = {10.1109/JIOT.2026.3712733}
}

📜 License

This repository is licensed under the Apache-2.0 License.

About

Real-time Industrial Anomaly Defect Inference Detection implemented by cpp(实时工业缺陷检测cpp)

Topics

Resources

Stars

7 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

🏭 FastVision: Fast Industrial Vision Platform (C++)

High-throughput, low-latency industrial vision platform based on TensorRT 10 and CUDA 11

C++17CUDATensorRTOpenCVImGuiWindows

English | 中文

📄 Paper🐍 Python Prototype


📸 System Demo

System Demo
Real-time three-screen display: Original Capture (Left) | Anomaly Heatmap (Middle) | Defect Overlay (Right)

✨ Core Highlights

🚀 Extreme Performance

Pure CUDA operator post-processing + OpenGL zero-copy rendering
Inference latency < 15ms (80+ FPS)

🎯 Anomalib Compatible

Fully compatible with Anomalib ecosystem
Seamless loading of standard models, supporting pixel-level defect segmentation and localization

🖥️ Interactive Panel

Modern console built with ImGui
Support for real-time threshold adjustment and model hot-switching

📦 Ready to Use

Provides Windows one-click installer (Setup.exe)
No need to configure Python/CUDA environment, just double-click to run


⚡ Quick Start

👥 I'm an End User

Don't want to code, just want to run the software?

  1. Download the latest release: 👉 Download (Release)
  2. Need help? Check the documentation: 📖 User Manual

👨‍💻 I'm a Developer

Want to modify the source code or develop secondary applications?

Please refer to the compilation and build guide: 🛠️ Developer Environment Setup


📌 Performance Benchmark

We conducted comprehensive tests on an RTX 3060 platform and achieved more than 2x throughput improvement compared to the original implementation.

Optimization StagePreprocessing (ms)Inference (ms)Postprocessing (ms)Total Latency (ms)Throughput (FPS)
Original C++ Implementation1117533~30
+ CUDA Preprocessing417526~38
+ Zero-copy Postprocessing417324~42
+ FP16 Quantization48315~67
+ Rendering Pipeline Optimization483< 15> 80

🏗️ System Architecture

The system adopts a modular design, achieving efficient interoperation between computation (CUDA) and display (OpenGL).

  • 🧠 Inference Core
    • Encapsulates TensorRT 10, supports FP32/FP16 dynamic precision.
    • Multi-threaded pipeline design, separating input IO from GPU computation.
  • ⚡ CUDA Acceleration Layer
    • Preprocessor: Color space conversion, normalization, Resize (NPP).
    • Postprocessor: Anomaly map generation, threshold segmentation, heatmap rendering (Custom Kernels).
  • 🎨 Visualization & Interaction
    • Dashboard: Control panel based on ImGui.
    • Renderer: Uses CUDA-OpenGL Interop to directly map VRAM textures, eliminating CPU-GPU bandwidth bottleneck.

📅 Changelog

v1.0.0 - 2026-01-28: Architecture Refactoring and Performance Optimization (Click to expand)
  • New Modular Architecture: Refactored configuration management, rendering, inference engine, and other independent modules.
  • Rendering Optimization: Implemented CUDA-OpenGL interoperability, eliminating Host-to-Device copy overhead.
  • Pipeline Enhancement: Multi-threaded Pipeline, support for FP16 acceleration.
  • UI Upgrade: Integrated ImGui for interactive parameter adjustment.

For complete records, please refer to CHANGELOG.md


🤝 Acknowledgements & Feedback

If this project helps your research or work, please give it a ⭐ Star on GitHub!

If you find any bugs or have improvement suggestions, please submit an Issue or Pull Request.

📚 Reference

If you find this project useful in your research or work, please cite this project:

@article{liao2026limr,
author = {Shaowei Liao and Wenyong Yu and Shaolin Liao},
title = {Lightweight Masked Reconstruction for Real-Time Sensor-Driven
Anomaly Detection in Industrial IoT},
journal = {IEEE Internet of Things Journal},
year = {2026},
doi = {10.1109/JIOT.2026.3712733}
}

📜 License

This repository is licensed under the Apache-2.0 License.

About

Real-time Industrial Anomaly Defect Inference Detection implemented by cpp(实时工业缺陷检测cpp)

Topics

Resources

Stars

7 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

🏭 FastVision: Fast Industrial Vision Platform (C++)

High-throughput, low-latency industrial vision platform based on TensorRT 10 and CUDA 11

C++17CUDATensorRTOpenCVImGuiWindows

English | 中文

📄 Paper🐍 Python Prototype


📸 System Demo

System Demo
Real-time three-screen display: Original Capture (Left) | Anomaly Heatmap (Middle) | Defect Overlay (Right)

✨ Core Highlights

🚀 Extreme Performance

Pure CUDA operator post-processing + OpenGL zero-copy rendering
Inference latency < 15ms (80+ FPS)

🎯 Anomalib Compatible

Fully compatible with Anomalib ecosystem
Seamless loading of standard models, supporting pixel-level defect segmentation and localization

🖥️ Interactive Panel

Modern console built with ImGui
Support for real-time threshold adjustment and model hot-switching

📦 Ready to Use

Provides Windows one-click installer (Setup.exe)
No need to configure Python/CUDA environment, just double-click to run


⚡ Quick Start

👥 I'm an End User

Don't want to code, just want to run the software?

  1. Download the latest release: 👉 Download (Release)
  2. Need help? Check the documentation: 📖 User Manual

👨‍💻 I'm a Developer

Want to modify the source code or develop secondary applications?

Please refer to the compilation and build guide: 🛠️ Developer Environment Setup


📌 Performance Benchmark

We conducted comprehensive tests on an RTX 3060 platform and achieved more than 2x throughput improvement compared to the original implementation.

Optimization StagePreprocessing (ms)Inference (ms)Postprocessing (ms)Total Latency (ms)Throughput (FPS)
Original C++ Implementation1117533~30
+ CUDA Preprocessing417526~38
+ Zero-copy Postprocessing417324~42
+ FP16 Quantization48315~67
+ Rendering Pipeline Optimization483< 15> 80

🏗️ System Architecture

The system adopts a modular design, achieving efficient interoperation between computation (CUDA) and display (OpenGL).

  • 🧠 Inference Core
    • Encapsulates TensorRT 10, supports FP32/FP16 dynamic precision.
    • Multi-threaded pipeline design, separating input IO from GPU computation.
  • ⚡ CUDA Acceleration Layer
    • Preprocessor: Color space conversion, normalization, Resize (NPP).
    • Postprocessor: Anomaly map generation, threshold segmentation, heatmap rendering (Custom Kernels).
  • 🎨 Visualization & Interaction
    • Dashboard: Control panel based on ImGui.
    • Renderer: Uses CUDA-OpenGL Interop to directly map VRAM textures, eliminating CPU-GPU bandwidth bottleneck.

📅 Changelog

v1.0.0 - 2026-01-28: Architecture Refactoring and Performance Optimization (Click to expand)
  • New Modular Architecture: Refactored configuration management, rendering, inference engine, and other independent modules.
  • Rendering Optimization: Implemented CUDA-OpenGL interoperability, eliminating Host-to-Device copy overhead.
  • Pipeline Enhancement: Multi-threaded Pipeline, support for FP16 acceleration.
  • UI Upgrade: Integrated ImGui for interactive parameter adjustment.

For complete records, please refer to CHANGELOG.md


🤝 Acknowledgements & Feedback

If this project helps your research or work, please give it a ⭐ Star on GitHub!

If you find any bugs or have improvement suggestions, please submit an Issue or Pull Request.

📚 Reference

If you find this project useful in your research or work, please cite this project:

@article{liao2026limr,
author = {Shaowei Liao and Wenyong Yu and Shaolin Liao},
title = {Lightweight Masked Reconstruction for Real-Time Sensor-Driven
Anomaly Detection in Industrial IoT},
journal = {IEEE Internet of Things Journal},
year = {2026},
doi = {10.1109/JIOT.2026.3712733}
}

📜 License

This repository is licensed under the Apache-2.0 License.

About

Real-time Industrial Anomaly Defect Inference Detection implemented by cpp(实时工业缺陷检测cpp)

Topics

Resources

Stars

7 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

🏭 FastVision: Fast Industrial Vision Platform (C++)

High-throughput, low-latency industrial vision platform based on TensorRT 10 and CUDA 11

C++17CUDATensorRTOpenCVImGuiWindows

English | 中文

📄 Paper🐍 Python Prototype


📸 System Demo

System Demo
Real-time three-screen display: Original Capture (Left) | Anomaly Heatmap (Middle) | Defect Overlay (Right)

✨ Core Highlights

🚀 Extreme Performance

Pure CUDA operator post-processing + OpenGL zero-copy rendering
Inference latency < 15ms (80+ FPS)

🎯 Anomalib Compatible

Fully compatible with Anomalib ecosystem
Seamless loading of standard models, supporting pixel-level defect segmentation and localization

🖥️ Interactive Panel

Modern console built with ImGui
Support for real-time threshold adjustment and model hot-switching

📦 Ready to Use

Provides Windows one-click installer (Setup.exe)
No need to configure Python/CUDA environment, just double-click to run


⚡ Quick Start

👥 I'm an End User

Don't want to code, just want to run the software?

  1. Download the latest release: 👉 Download (Release)
  2. Need help? Check the documentation: 📖 User Manual

👨‍💻 I'm a Developer

Want to modify the source code or develop secondary applications?

Please refer to the compilation and build guide: 🛠️ Developer Environment Setup


📌 Performance Benchmark

We conducted comprehensive tests on an RTX 3060 platform and achieved more than 2x throughput improvement compared to the original implementation.

Optimization StagePreprocessing (ms)Inference (ms)Postprocessing (ms)Total Latency (ms)Throughput (FPS)
Original C++ Implementation1117533~30
+ CUDA Preprocessing417526~38
+ Zero-copy Postprocessing417324~42
+ FP16 Quantization48315~67
+ Rendering Pipeline Optimization483< 15> 80

🏗️ System Architecture

The system adopts a modular design, achieving efficient interoperation between computation (CUDA) and display (OpenGL).

  • 🧠 Inference Core
    • Encapsulates TensorRT 10, supports FP32/FP16 dynamic precision.
    • Multi-threaded pipeline design, separating input IO from GPU computation.
  • ⚡ CUDA Acceleration Layer
    • Preprocessor: Color space conversion, normalization, Resize (NPP).
    • Postprocessor: Anomaly map generation, threshold segmentation, heatmap rendering (Custom Kernels).
  • 🎨 Visualization & Interaction
    • Dashboard: Control panel based on ImGui.
    • Renderer: Uses CUDA-OpenGL Interop to directly map VRAM textures, eliminating CPU-GPU bandwidth bottleneck.

📅 Changelog

v1.0.0 - 2026-01-28: Architecture Refactoring and Performance Optimization (Click to expand)
  • New Modular Architecture: Refactored configuration management, rendering, inference engine, and other independent modules.
  • Rendering Optimization: Implemented CUDA-OpenGL interoperability, eliminating Host-to-Device copy overhead.
  • Pipeline Enhancement: Multi-threaded Pipeline, support for FP16 acceleration.
  • UI Upgrade: Integrated ImGui for interactive parameter adjustment.

For complete records, please refer to CHANGELOG.md


🤝 Acknowledgements & Feedback

If this project helps your research or work, please give it a ⭐ Star on GitHub!

If you find any bugs or have improvement suggestions, please submit an Issue or Pull Request.

📚 Reference

If you find this project useful in your research or work, please cite this project:

@article{liao2026limr,
author = {Shaowei Liao and Wenyong Yu and Shaolin Liao},
title = {Lightweight Masked Reconstruction for Real-Time Sensor-Driven
Anomaly Detection in Industrial IoT},
journal = {IEEE Internet of Things Journal},
year = {2026},
doi = {10.1109/JIOT.2026.3712733}
}

📜 License

This repository is licensed under the Apache-2.0 License.

About

Real-time Industrial Anomaly Defect Inference Detection implemented by cpp(实时工业缺陷检测cpp)

Topics

Resources

Stars

7 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages