Latest commit

History

History

Folders and files

NameName
Last commit message
Last commit date

parent directory

..

README.md

English | 中文

Usage of FastDeploy model multi-thread or multi-process prediction

FastDeploy provides the following multi-thread or multi-process examples for python and cpp developers

Models that currently support multi-thread and multi-process predictions

task typeillustratemodel download link
Detectionsupport PaddleDetection series modelsPaddleDetection
Segmentationsupport PaddleSeg series modelsPaddleSeg
Classificationsupport PaddleClas series modelsPaddleClas
OCRsupport PaddleOCR series modelsPaddleOCR

Notice:

  • click the model download link above to download the model from the Download pre-training model module
  • OCR is a pipeline model. For multi-thread examples, please refer to the pipeline folder. Other single-model multi-thread examples are in the single_model folder.

Clone model when using multi-thread prediction

the inference process of vision model is consist of three stages

  • load the image, then the image is preprocessed, finally get the Tensor to be input to the model Runtime, that is the preprocess stage
  • the model Runtime receives Tensor, do the inference, and obtains the output tensor of Runtime, that is the infer stage
  • process the output tensor of Runtime to get the final structured information, such as DetectionResult, SegmentationResult, etc., that is the postprocess stage

For the above three stages: preprocess, inference, and postprocess, FastDeploy abstracted three corresponding classes, namely Preprocessor, Runtime, and PostProcessor

When using FastDeploy for multi-thread inference, several issues should be considered

  • Can the Preprocessor, Runtime, and Postprocessor support parallel processing respectively?
  • 在支持多线程并发的前提下,能否最大限度的减少内存或显存占用
  • Under the premise of supporting multi-thread concurrency, can the memory or video memory usage be minimized?

FastDeploy adopts the method of copying multiple objects separately for multi-thread inference, so each thread has an independent instance of Preprocessor, Runtime, and PostProcessor. In order to reduce the memory usage, the Runtime adopt sharing the model weights copy method. In this way, the memory usage caused by copying multiple objects is reduced.

FastDeploy provides the following interface to clone the model (take PaddleClas as an example)

  • Python: PaddleClasModel.clone()
  • C++: PaddleClasModel::Clone()

Python

import fastdeploy as fd
option = fd.RuntimeOption()
model = fd.vision.classification.PaddleClasModel(model_file,
params_file,
config_file,
runtime_option=option)
model2 = model.clone()
im = cv2.imread(image)
res = model.predict(im)

C++

auto model = fastdeploy::vision::classification::PaddleClasModel(model_file,
params_file,
config_file,
option);
auto model2 = model.Clone();
auto im = cv::imread(image_file);
fastdeploy::vision::ClassifyResult res;
model->Predict(im, &res)

Notice:Other models API refer to官方C++文档 and 官方Python文档

Python multi-thread and multi-process

Due to language limitations, Python has the existence of GIL lock. In computing-intensive scenarios, multithreading cannot make full use of hardware resources. Therefore, two examples of multi-process and multi-thread are provided on Python. The similarities and differences are as follows:

Comparison of multi-process and multi-thread inference in FastDeploy model

resource usagecomputationally intensiveI/O intensiveinter-process or inter-thread communication
multi-processlargefastfastslow
multi-threadlittleslowrelatively fastfast

注意: The above analysis is a theoretical analysis. In fact, Python has also made certain optimizations for different computing tasks. For example, the calculation of numpy can already be computed by multi-thread parallelly. In addition, the result aggregation between multiple processes involves time-consuming operation(inter-process communication), Besides, it is difficult to identify whether the task is computationally intensive or I/O intensive, so everything needs to be tested according to the task.

C++ multi-thread

The C++ multi-thread has the characteristics of occupying less resources and high speed.Therefore, multi-threaded inference is the best choice in C++

C++ comparition between multi-thread Clone and not Clone memory occupation

硬件:Intel(R) Xeon(R) Gold 6271C CPU @ 2.60GHz
模型:ResNet50_vd_infer
后端:CPU OPENVINO Backend

memory occupation of initializing multiple models in a single process

number of modelsafter model.Clone()after model->predict() with model.Clone()initializing model without model.Clone()after model->predict() without model.Clone()
1322M325M322M325M
2322M325M559M560M
3322M325M771M771M

memory occupation of multi-thread

thread numberafter model.Clone()after model->predict() with model.Clone()initialize model without model.Clone()after model->predict() without model.Clone()
1322M337M322M337M
2322M343M548M566M
3322M347M752M784M
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

History

Folders and files

NameName
Last commit message
Last commit date

parent directory

..

README.md

English | 中文

Usage of FastDeploy model multi-thread or multi-process prediction

FastDeploy provides the following multi-thread or multi-process examples for python and cpp developers

Models that currently support multi-thread and multi-process predictions

task typeillustratemodel download link
Detectionsupport PaddleDetection series modelsPaddleDetection
Segmentationsupport PaddleSeg series modelsPaddleSeg
Classificationsupport PaddleClas series modelsPaddleClas
OCRsupport PaddleOCR series modelsPaddleOCR

Notice:

  • click the model download link above to download the model from the Download pre-training model module
  • OCR is a pipeline model. For multi-thread examples, please refer to the pipeline folder. Other single-model multi-thread examples are in the single_model folder.

Clone model when using multi-thread prediction

the inference process of vision model is consist of three stages

  • load the image, then the image is preprocessed, finally get the Tensor to be input to the model Runtime, that is the preprocess stage
  • the model Runtime receives Tensor, do the inference, and obtains the output tensor of Runtime, that is the infer stage
  • process the output tensor of Runtime to get the final structured information, such as DetectionResult, SegmentationResult, etc., that is the postprocess stage

For the above three stages: preprocess, inference, and postprocess, FastDeploy abstracted three corresponding classes, namely Preprocessor, Runtime, and PostProcessor

When using FastDeploy for multi-thread inference, several issues should be considered

  • Can the Preprocessor, Runtime, and Postprocessor support parallel processing respectively?
  • 在支持多线程并发的前提下,能否最大限度的减少内存或显存占用
  • Under the premise of supporting multi-thread concurrency, can the memory or video memory usage be minimized?

FastDeploy adopts the method of copying multiple objects separately for multi-thread inference, so each thread has an independent instance of Preprocessor, Runtime, and PostProcessor. In order to reduce the memory usage, the Runtime adopt sharing the model weights copy method. In this way, the memory usage caused by copying multiple objects is reduced.

FastDeploy provides the following interface to clone the model (take PaddleClas as an example)

  • Python: PaddleClasModel.clone()
  • C++: PaddleClasModel::Clone()

Python

import fastdeploy as fd
option = fd.RuntimeOption()
model = fd.vision.classification.PaddleClasModel(model_file,
params_file,
config_file,
runtime_option=option)
model2 = model.clone()
im = cv2.imread(image)
res = model.predict(im)

C++

auto model = fastdeploy::vision::classification::PaddleClasModel(model_file,
params_file,
config_file,
option);
auto model2 = model.Clone();
auto im = cv::imread(image_file);
fastdeploy::vision::ClassifyResult res;
model->Predict(im, &res)

Notice:Other models API refer to官方C++文档 and 官方Python文档

Python multi-thread and multi-process

Due to language limitations, Python has the existence of GIL lock. In computing-intensive scenarios, multithreading cannot make full use of hardware resources. Therefore, two examples of multi-process and multi-thread are provided on Python. The similarities and differences are as follows:

Comparison of multi-process and multi-thread inference in FastDeploy model

resource usagecomputationally intensiveI/O intensiveinter-process or inter-thread communication
multi-processlargefastfastslow
multi-threadlittleslowrelatively fastfast

注意: The above analysis is a theoretical analysis. In fact, Python has also made certain optimizations for different computing tasks. For example, the calculation of numpy can already be computed by multi-thread parallelly. In addition, the result aggregation between multiple processes involves time-consuming operation(inter-process communication), Besides, it is difficult to identify whether the task is computationally intensive or I/O intensive, so everything needs to be tested according to the task.

C++ multi-thread

The C++ multi-thread has the characteristics of occupying less resources and high speed.Therefore, multi-threaded inference is the best choice in C++

C++ comparition between multi-thread Clone and not Clone memory occupation

硬件:Intel(R) Xeon(R) Gold 6271C CPU @ 2.60GHz
模型:ResNet50_vd_infer
后端:CPU OPENVINO Backend

memory occupation of initializing multiple models in a single process

number of modelsafter model.Clone()after model->predict() with model.Clone()initializing model without model.Clone()after model->predict() without model.Clone()
1322M325M322M325M
2322M325M559M560M
3322M325M771M771M

memory occupation of multi-thread

thread numberafter model.Clone()after model->predict() with model.Clone()initialize model without model.Clone()after model->predict() without model.Clone()
1322M337M322M337M
2322M343M548M566M
3322M347M752M784M
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

History

Folders and files

NameName
Last commit message
Last commit date

parent directory

..

README.md

English | 中文

Usage of FastDeploy model multi-thread or multi-process prediction

FastDeploy provides the following multi-thread or multi-process examples for python and cpp developers

Models that currently support multi-thread and multi-process predictions

task typeillustratemodel download link
Detectionsupport PaddleDetection series modelsPaddleDetection
Segmentationsupport PaddleSeg series modelsPaddleSeg
Classificationsupport PaddleClas series modelsPaddleClas
OCRsupport PaddleOCR series modelsPaddleOCR

Notice:

  • click the model download link above to download the model from the Download pre-training model module
  • OCR is a pipeline model. For multi-thread examples, please refer to the pipeline folder. Other single-model multi-thread examples are in the single_model folder.

Clone model when using multi-thread prediction

the inference process of vision model is consist of three stages

  • load the image, then the image is preprocessed, finally get the Tensor to be input to the model Runtime, that is the preprocess stage
  • the model Runtime receives Tensor, do the inference, and obtains the output tensor of Runtime, that is the infer stage
  • process the output tensor of Runtime to get the final structured information, such as DetectionResult, SegmentationResult, etc., that is the postprocess stage

For the above three stages: preprocess, inference, and postprocess, FastDeploy abstracted three corresponding classes, namely Preprocessor, Runtime, and PostProcessor

When using FastDeploy for multi-thread inference, several issues should be considered

  • Can the Preprocessor, Runtime, and Postprocessor support parallel processing respectively?
  • 在支持多线程并发的前提下,能否最大限度的减少内存或显存占用
  • Under the premise of supporting multi-thread concurrency, can the memory or video memory usage be minimized?

FastDeploy adopts the method of copying multiple objects separately for multi-thread inference, so each thread has an independent instance of Preprocessor, Runtime, and PostProcessor. In order to reduce the memory usage, the Runtime adopt sharing the model weights copy method. In this way, the memory usage caused by copying multiple objects is reduced.

FastDeploy provides the following interface to clone the model (take PaddleClas as an example)

  • Python: PaddleClasModel.clone()
  • C++: PaddleClasModel::Clone()

Python

import fastdeploy as fd
option = fd.RuntimeOption()
model = fd.vision.classification.PaddleClasModel(model_file,
params_file,
config_file,
runtime_option=option)
model2 = model.clone()
im = cv2.imread(image)
res = model.predict(im)

C++

auto model = fastdeploy::vision::classification::PaddleClasModel(model_file,
params_file,
config_file,
option);
auto model2 = model.Clone();
auto im = cv::imread(image_file);
fastdeploy::vision::ClassifyResult res;
model->Predict(im, &res)

Notice:Other models API refer to官方C++文档 and 官方Python文档

Python multi-thread and multi-process

Due to language limitations, Python has the existence of GIL lock. In computing-intensive scenarios, multithreading cannot make full use of hardware resources. Therefore, two examples of multi-process and multi-thread are provided on Python. The similarities and differences are as follows:

Comparison of multi-process and multi-thread inference in FastDeploy model

resource usagecomputationally intensiveI/O intensiveinter-process or inter-thread communication
multi-processlargefastfastslow
multi-threadlittleslowrelatively fastfast

注意: The above analysis is a theoretical analysis. In fact, Python has also made certain optimizations for different computing tasks. For example, the calculation of numpy can already be computed by multi-thread parallelly. In addition, the result aggregation between multiple processes involves time-consuming operation(inter-process communication), Besides, it is difficult to identify whether the task is computationally intensive or I/O intensive, so everything needs to be tested according to the task.

C++ multi-thread

The C++ multi-thread has the characteristics of occupying less resources and high speed.Therefore, multi-threaded inference is the best choice in C++

C++ comparition between multi-thread Clone and not Clone memory occupation

硬件:Intel(R) Xeon(R) Gold 6271C CPU @ 2.60GHz
模型:ResNet50_vd_infer
后端:CPU OPENVINO Backend

memory occupation of initializing multiple models in a single process

number of modelsafter model.Clone()after model->predict() with model.Clone()initializing model without model.Clone()after model->predict() without model.Clone()
1322M325M322M325M
2322M325M559M560M
3322M325M771M771M

memory occupation of multi-thread

thread numberafter model.Clone()after model->predict() with model.Clone()initialize model without model.Clone()after model->predict() without model.Clone()
1322M337M322M337M
2322M343M548M566M
3322M347M752M784M
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

History

Folders and files

NameName
Last commit message
Last commit date

parent directory

..

README.md

English | 中文

Usage of FastDeploy model multi-thread or multi-process prediction

FastDeploy provides the following multi-thread or multi-process examples for python and cpp developers

Models that currently support multi-thread and multi-process predictions

task typeillustratemodel download link
Detectionsupport PaddleDetection series modelsPaddleDetection
Segmentationsupport PaddleSeg series modelsPaddleSeg
Classificationsupport PaddleClas series modelsPaddleClas
OCRsupport PaddleOCR series modelsPaddleOCR

Notice:

  • click the model download link above to download the model from the Download pre-training model module
  • OCR is a pipeline model. For multi-thread examples, please refer to the pipeline folder. Other single-model multi-thread examples are in the single_model folder.

Clone model when using multi-thread prediction

the inference process of vision model is consist of three stages

  • load the image, then the image is preprocessed, finally get the Tensor to be input to the model Runtime, that is the preprocess stage
  • the model Runtime receives Tensor, do the inference, and obtains the output tensor of Runtime, that is the infer stage
  • process the output tensor of Runtime to get the final structured information, such as DetectionResult, SegmentationResult, etc., that is the postprocess stage

For the above three stages: preprocess, inference, and postprocess, FastDeploy abstracted three corresponding classes, namely Preprocessor, Runtime, and PostProcessor

When using FastDeploy for multi-thread inference, several issues should be considered

  • Can the Preprocessor, Runtime, and Postprocessor support parallel processing respectively?
  • 在支持多线程并发的前提下,能否最大限度的减少内存或显存占用
  • Under the premise of supporting multi-thread concurrency, can the memory or video memory usage be minimized?

FastDeploy adopts the method of copying multiple objects separately for multi-thread inference, so each thread has an independent instance of Preprocessor, Runtime, and PostProcessor. In order to reduce the memory usage, the Runtime adopt sharing the model weights copy method. In this way, the memory usage caused by copying multiple objects is reduced.

FastDeploy provides the following interface to clone the model (take PaddleClas as an example)

  • Python: PaddleClasModel.clone()
  • C++: PaddleClasModel::Clone()

Python

import fastdeploy as fd
option = fd.RuntimeOption()
model = fd.vision.classification.PaddleClasModel(model_file,
params_file,
config_file,
runtime_option=option)
model2 = model.clone()
im = cv2.imread(image)
res = model.predict(im)

C++

auto model = fastdeploy::vision::classification::PaddleClasModel(model_file,
params_file,
config_file,
option);
auto model2 = model.Clone();
auto im = cv::imread(image_file);
fastdeploy::vision::ClassifyResult res;
model->Predict(im, &res)

Notice:Other models API refer to官方C++文档 and 官方Python文档

Python multi-thread and multi-process

Due to language limitations, Python has the existence of GIL lock. In computing-intensive scenarios, multithreading cannot make full use of hardware resources. Therefore, two examples of multi-process and multi-thread are provided on Python. The similarities and differences are as follows:

Comparison of multi-process and multi-thread inference in FastDeploy model

resource usagecomputationally intensiveI/O intensiveinter-process or inter-thread communication
multi-processlargefastfastslow
multi-threadlittleslowrelatively fastfast

注意: The above analysis is a theoretical analysis. In fact, Python has also made certain optimizations for different computing tasks. For example, the calculation of numpy can already be computed by multi-thread parallelly. In addition, the result aggregation between multiple processes involves time-consuming operation(inter-process communication), Besides, it is difficult to identify whether the task is computationally intensive or I/O intensive, so everything needs to be tested according to the task.

C++ multi-thread

The C++ multi-thread has the characteristics of occupying less resources and high speed.Therefore, multi-threaded inference is the best choice in C++

C++ comparition between multi-thread Clone and not Clone memory occupation

硬件:Intel(R) Xeon(R) Gold 6271C CPU @ 2.60GHz
模型:ResNet50_vd_infer
后端:CPU OPENVINO Backend

memory occupation of initializing multiple models in a single process

number of modelsafter model.Clone()after model->predict() with model.Clone()initializing model without model.Clone()after model->predict() without model.Clone()
1322M325M322M325M
2322M325M559M560M
3322M325M771M771M

memory occupation of multi-thread

thread numberafter model.Clone()after model->predict() with model.Clone()initialize model without model.Clone()after model->predict() without model.Clone()
1322M337M322M337M
2322M343M548M566M
3322M347M752M784M
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

History

Folders and files

NameName
Last commit message
Last commit date

parent directory

..

README.md

English | 中文

Usage of FastDeploy model multi-thread or multi-process prediction

FastDeploy provides the following multi-thread or multi-process examples for python and cpp developers

Models that currently support multi-thread and multi-process predictions

task typeillustratemodel download link
Detectionsupport PaddleDetection series modelsPaddleDetection
Segmentationsupport PaddleSeg series modelsPaddleSeg
Classificationsupport PaddleClas series modelsPaddleClas
OCRsupport PaddleOCR series modelsPaddleOCR

Notice:

  • click the model download link above to download the model from the Download pre-training model module
  • OCR is a pipeline model. For multi-thread examples, please refer to the pipeline folder. Other single-model multi-thread examples are in the single_model folder.

Clone model when using multi-thread prediction

the inference process of vision model is consist of three stages

  • load the image, then the image is preprocessed, finally get the Tensor to be input to the model Runtime, that is the preprocess stage
  • the model Runtime receives Tensor, do the inference, and obtains the output tensor of Runtime, that is the infer stage
  • process the output tensor of Runtime to get the final structured information, such as DetectionResult, SegmentationResult, etc., that is the postprocess stage

For the above three stages: preprocess, inference, and postprocess, FastDeploy abstracted three corresponding classes, namely Preprocessor, Runtime, and PostProcessor

When using FastDeploy for multi-thread inference, several issues should be considered

  • Can the Preprocessor, Runtime, and Postprocessor support parallel processing respectively?
  • 在支持多线程并发的前提下,能否最大限度的减少内存或显存占用
  • Under the premise of supporting multi-thread concurrency, can the memory or video memory usage be minimized?

FastDeploy adopts the method of copying multiple objects separately for multi-thread inference, so each thread has an independent instance of Preprocessor, Runtime, and PostProcessor. In order to reduce the memory usage, the Runtime adopt sharing the model weights copy method. In this way, the memory usage caused by copying multiple objects is reduced.

FastDeploy provides the following interface to clone the model (take PaddleClas as an example)

  • Python: PaddleClasModel.clone()
  • C++: PaddleClasModel::Clone()

Python

import fastdeploy as fd
option = fd.RuntimeOption()
model = fd.vision.classification.PaddleClasModel(model_file,
params_file,
config_file,
runtime_option=option)
model2 = model.clone()
im = cv2.imread(image)
res = model.predict(im)

C++

auto model = fastdeploy::vision::classification::PaddleClasModel(model_file,
params_file,
config_file,
option);
auto model2 = model.Clone();
auto im = cv::imread(image_file);
fastdeploy::vision::ClassifyResult res;
model->Predict(im, &res)

Notice:Other models API refer to官方C++文档 and 官方Python文档

Python multi-thread and multi-process

Due to language limitations, Python has the existence of GIL lock. In computing-intensive scenarios, multithreading cannot make full use of hardware resources. Therefore, two examples of multi-process and multi-thread are provided on Python. The similarities and differences are as follows:

Comparison of multi-process and multi-thread inference in FastDeploy model

resource usagecomputationally intensiveI/O intensiveinter-process or inter-thread communication
multi-processlargefastfastslow
multi-threadlittleslowrelatively fastfast

注意: The above analysis is a theoretical analysis. In fact, Python has also made certain optimizations for different computing tasks. For example, the calculation of numpy can already be computed by multi-thread parallelly. In addition, the result aggregation between multiple processes involves time-consuming operation(inter-process communication), Besides, it is difficult to identify whether the task is computationally intensive or I/O intensive, so everything needs to be tested according to the task.

C++ multi-thread

The C++ multi-thread has the characteristics of occupying less resources and high speed.Therefore, multi-threaded inference is the best choice in C++

C++ comparition between multi-thread Clone and not Clone memory occupation

硬件:Intel(R) Xeon(R) Gold 6271C CPU @ 2.60GHz
模型:ResNet50_vd_infer
后端:CPU OPENVINO Backend

memory occupation of initializing multiple models in a single process

number of modelsafter model.Clone()after model->predict() with model.Clone()initializing model without model.Clone()after model->predict() without model.Clone()
1322M325M322M325M
2322M325M559M560M
3322M325M771M771M

memory occupation of multi-thread

thread numberafter model.Clone()after model->predict() with model.Clone()initialize model without model.Clone()after model->predict() without model.Clone()
1322M337M322M337M
2322M343M548M566M
3322M347M752M784M
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

History

Folders and files

NameName
Last commit message
Last commit date

parent directory

..

README.md

English | 中文

Usage of FastDeploy model multi-thread or multi-process prediction

FastDeploy provides the following multi-thread or multi-process examples for python and cpp developers

Models that currently support multi-thread and multi-process predictions

task typeillustratemodel download link
Detectionsupport PaddleDetection series modelsPaddleDetection
Segmentationsupport PaddleSeg series modelsPaddleSeg
Classificationsupport PaddleClas series modelsPaddleClas
OCRsupport PaddleOCR series modelsPaddleOCR

Notice:

  • click the model download link above to download the model from the Download pre-training model module
  • OCR is a pipeline model. For multi-thread examples, please refer to the pipeline folder. Other single-model multi-thread examples are in the single_model folder.

Clone model when using multi-thread prediction

the inference process of vision model is consist of three stages

  • load the image, then the image is preprocessed, finally get the Tensor to be input to the model Runtime, that is the preprocess stage
  • the model Runtime receives Tensor, do the inference, and obtains the output tensor of Runtime, that is the infer stage
  • process the output tensor of Runtime to get the final structured information, such as DetectionResult, SegmentationResult, etc., that is the postprocess stage

For the above three stages: preprocess, inference, and postprocess, FastDeploy abstracted three corresponding classes, namely Preprocessor, Runtime, and PostProcessor

When using FastDeploy for multi-thread inference, several issues should be considered

  • Can the Preprocessor, Runtime, and Postprocessor support parallel processing respectively?
  • 在支持多线程并发的前提下,能否最大限度的减少内存或显存占用
  • Under the premise of supporting multi-thread concurrency, can the memory or video memory usage be minimized?

FastDeploy adopts the method of copying multiple objects separately for multi-thread inference, so each thread has an independent instance of Preprocessor, Runtime, and PostProcessor. In order to reduce the memory usage, the Runtime adopt sharing the model weights copy method. In this way, the memory usage caused by copying multiple objects is reduced.

FastDeploy provides the following interface to clone the model (take PaddleClas as an example)

  • Python: PaddleClasModel.clone()
  • C++: PaddleClasModel::Clone()

Python

import fastdeploy as fd
option = fd.RuntimeOption()
model = fd.vision.classification.PaddleClasModel(model_file,
params_file,
config_file,
runtime_option=option)
model2 = model.clone()
im = cv2.imread(image)
res = model.predict(im)

C++

auto model = fastdeploy::vision::classification::PaddleClasModel(model_file,
params_file,
config_file,
option);
auto model2 = model.Clone();
auto im = cv::imread(image_file);
fastdeploy::vision::ClassifyResult res;
model->Predict(im, &res)

Notice:Other models API refer to官方C++文档 and 官方Python文档

Python multi-thread and multi-process

Due to language limitations, Python has the existence of GIL lock. In computing-intensive scenarios, multithreading cannot make full use of hardware resources. Therefore, two examples of multi-process and multi-thread are provided on Python. The similarities and differences are as follows:

Comparison of multi-process and multi-thread inference in FastDeploy model

resource usagecomputationally intensiveI/O intensiveinter-process or inter-thread communication
multi-processlargefastfastslow
multi-threadlittleslowrelatively fastfast

注意: The above analysis is a theoretical analysis. In fact, Python has also made certain optimizations for different computing tasks. For example, the calculation of numpy can already be computed by multi-thread parallelly. In addition, the result aggregation between multiple processes involves time-consuming operation(inter-process communication), Besides, it is difficult to identify whether the task is computationally intensive or I/O intensive, so everything needs to be tested according to the task.

C++ multi-thread

The C++ multi-thread has the characteristics of occupying less resources and high speed.Therefore, multi-threaded inference is the best choice in C++

C++ comparition between multi-thread Clone and not Clone memory occupation

硬件:Intel(R) Xeon(R) Gold 6271C CPU @ 2.60GHz
模型:ResNet50_vd_infer
后端:CPU OPENVINO Backend

memory occupation of initializing multiple models in a single process

number of modelsafter model.Clone()after model->predict() with model.Clone()initializing model without model.Clone()after model->predict() without model.Clone()
1322M325M322M325M
2322M325M559M560M
3322M325M771M771M

memory occupation of multi-thread

thread numberafter model.Clone()after model->predict() with model.Clone()initialize model without model.Clone()after model->predict() without model.Clone()
1322M337M322M337M
2322M343M548M566M
3322M347M752M784M
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

History

Folders and files

NameName
Last commit message
Last commit date

parent directory

..

README.md

English | 中文

Usage of FastDeploy model multi-thread or multi-process prediction

FastDeploy provides the following multi-thread or multi-process examples for python and cpp developers

Models that currently support multi-thread and multi-process predictions

task typeillustratemodel download link
Detectionsupport PaddleDetection series modelsPaddleDetection
Segmentationsupport PaddleSeg series modelsPaddleSeg
Classificationsupport PaddleClas series modelsPaddleClas
OCRsupport PaddleOCR series modelsPaddleOCR

Notice:

  • click the model download link above to download the model from the Download pre-training model module
  • OCR is a pipeline model. For multi-thread examples, please refer to the pipeline folder. Other single-model multi-thread examples are in the single_model folder.

Clone model when using multi-thread prediction

the inference process of vision model is consist of three stages

  • load the image, then the image is preprocessed, finally get the Tensor to be input to the model Runtime, that is the preprocess stage
  • the model Runtime receives Tensor, do the inference, and obtains the output tensor of Runtime, that is the infer stage
  • process the output tensor of Runtime to get the final structured information, such as DetectionResult, SegmentationResult, etc., that is the postprocess stage

For the above three stages: preprocess, inference, and postprocess, FastDeploy abstracted three corresponding classes, namely Preprocessor, Runtime, and PostProcessor

When using FastDeploy for multi-thread inference, several issues should be considered

  • Can the Preprocessor, Runtime, and Postprocessor support parallel processing respectively?
  • 在支持多线程并发的前提下,能否最大限度的减少内存或显存占用
  • Under the premise of supporting multi-thread concurrency, can the memory or video memory usage be minimized?

FastDeploy adopts the method of copying multiple objects separately for multi-thread inference, so each thread has an independent instance of Preprocessor, Runtime, and PostProcessor. In order to reduce the memory usage, the Runtime adopt sharing the model weights copy method. In this way, the memory usage caused by copying multiple objects is reduced.

FastDeploy provides the following interface to clone the model (take PaddleClas as an example)

  • Python: PaddleClasModel.clone()
  • C++: PaddleClasModel::Clone()

Python

import fastdeploy as fd
option = fd.RuntimeOption()
model = fd.vision.classification.PaddleClasModel(model_file,
params_file,
config_file,
runtime_option=option)
model2 = model.clone()
im = cv2.imread(image)
res = model.predict(im)

C++

auto model = fastdeploy::vision::classification::PaddleClasModel(model_file,
params_file,
config_file,
option);
auto model2 = model.Clone();
auto im = cv::imread(image_file);
fastdeploy::vision::ClassifyResult res;
model->Predict(im, &res)

Notice:Other models API refer to官方C++文档 and 官方Python文档

Python multi-thread and multi-process

Due to language limitations, Python has the existence of GIL lock. In computing-intensive scenarios, multithreading cannot make full use of hardware resources. Therefore, two examples of multi-process and multi-thread are provided on Python. The similarities and differences are as follows:

Comparison of multi-process and multi-thread inference in FastDeploy model

resource usagecomputationally intensiveI/O intensiveinter-process or inter-thread communication
multi-processlargefastfastslow
multi-threadlittleslowrelatively fastfast

注意: The above analysis is a theoretical analysis. In fact, Python has also made certain optimizations for different computing tasks. For example, the calculation of numpy can already be computed by multi-thread parallelly. In addition, the result aggregation between multiple processes involves time-consuming operation(inter-process communication), Besides, it is difficult to identify whether the task is computationally intensive or I/O intensive, so everything needs to be tested according to the task.

C++ multi-thread

The C++ multi-thread has the characteristics of occupying less resources and high speed.Therefore, multi-threaded inference is the best choice in C++

C++ comparition between multi-thread Clone and not Clone memory occupation

硬件:Intel(R) Xeon(R) Gold 6271C CPU @ 2.60GHz
模型:ResNet50_vd_infer
后端:CPU OPENVINO Backend

memory occupation of initializing multiple models in a single process

number of modelsafter model.Clone()after model->predict() with model.Clone()initializing model without model.Clone()after model->predict() without model.Clone()
1322M325M322M325M
2322M325M559M560M
3322M325M771M771M

memory occupation of multi-thread

thread numberafter model.Clone()after model->predict() with model.Clone()initialize model without model.Clone()after model->predict() without model.Clone()
1322M337M322M337M
2322M343M548M566M
3322M347M752M784M
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

History

Folders and files

NameName
Last commit message
Last commit date

parent directory

..

README.md

English | 中文

Usage of FastDeploy model multi-thread or multi-process prediction

FastDeploy provides the following multi-thread or multi-process examples for python and cpp developers

Models that currently support multi-thread and multi-process predictions

task typeillustratemodel download link
Detectionsupport PaddleDetection series modelsPaddleDetection
Segmentationsupport PaddleSeg series modelsPaddleSeg
Classificationsupport PaddleClas series modelsPaddleClas
OCRsupport PaddleOCR series modelsPaddleOCR

Notice:

  • click the model download link above to download the model from the Download pre-training model module
  • OCR is a pipeline model. For multi-thread examples, please refer to the pipeline folder. Other single-model multi-thread examples are in the single_model folder.

Clone model when using multi-thread prediction

the inference process of vision model is consist of three stages

  • load the image, then the image is preprocessed, finally get the Tensor to be input to the model Runtime, that is the preprocess stage
  • the model Runtime receives Tensor, do the inference, and obtains the output tensor of Runtime, that is the infer stage
  • process the output tensor of Runtime to get the final structured information, such as DetectionResult, SegmentationResult, etc., that is the postprocess stage

For the above three stages: preprocess, inference, and postprocess, FastDeploy abstracted three corresponding classes, namely Preprocessor, Runtime, and PostProcessor

When using FastDeploy for multi-thread inference, several issues should be considered

  • Can the Preprocessor, Runtime, and Postprocessor support parallel processing respectively?
  • 在支持多线程并发的前提下,能否最大限度的减少内存或显存占用
  • Under the premise of supporting multi-thread concurrency, can the memory or video memory usage be minimized?

FastDeploy adopts the method of copying multiple objects separately for multi-thread inference, so each thread has an independent instance of Preprocessor, Runtime, and PostProcessor. In order to reduce the memory usage, the Runtime adopt sharing the model weights copy method. In this way, the memory usage caused by copying multiple objects is reduced.

FastDeploy provides the following interface to clone the model (take PaddleClas as an example)

  • Python: PaddleClasModel.clone()
  • C++: PaddleClasModel::Clone()

Python

import fastdeploy as fd
option = fd.RuntimeOption()
model = fd.vision.classification.PaddleClasModel(model_file,
params_file,
config_file,
runtime_option=option)
model2 = model.clone()
im = cv2.imread(image)
res = model.predict(im)

C++

auto model = fastdeploy::vision::classification::PaddleClasModel(model_file,
params_file,
config_file,
option);
auto model2 = model.Clone();
auto im = cv::imread(image_file);
fastdeploy::vision::ClassifyResult res;
model->Predict(im, &res)

Notice:Other models API refer to官方C++文档 and 官方Python文档

Python multi-thread and multi-process

Due to language limitations, Python has the existence of GIL lock. In computing-intensive scenarios, multithreading cannot make full use of hardware resources. Therefore, two examples of multi-process and multi-thread are provided on Python. The similarities and differences are as follows:

Comparison of multi-process and multi-thread inference in FastDeploy model

resource usagecomputationally intensiveI/O intensiveinter-process or inter-thread communication
multi-processlargefastfastslow
multi-threadlittleslowrelatively fastfast

注意: The above analysis is a theoretical analysis. In fact, Python has also made certain optimizations for different computing tasks. For example, the calculation of numpy can already be computed by multi-thread parallelly. In addition, the result aggregation between multiple processes involves time-consuming operation(inter-process communication), Besides, it is difficult to identify whether the task is computationally intensive or I/O intensive, so everything needs to be tested according to the task.

C++ multi-thread

The C++ multi-thread has the characteristics of occupying less resources and high speed.Therefore, multi-threaded inference is the best choice in C++

C++ comparition between multi-thread Clone and not Clone memory occupation

硬件:Intel(R) Xeon(R) Gold 6271C CPU @ 2.60GHz
模型:ResNet50_vd_infer
后端:CPU OPENVINO Backend

memory occupation of initializing multiple models in a single process

number of modelsafter model.Clone()after model->predict() with model.Clone()initializing model without model.Clone()after model->predict() without model.Clone()
1322M325M322M325M
2322M325M559M560M
3322M325M771M771M

memory occupation of multi-thread

thread numberafter model.Clone()after model->predict() with model.Clone()initialize model without model.Clone()after model->predict() without model.Clone()
1322M337M322M337M
2322M343M548M566M
3322M347M752M784M