Latest commit

History

History

README.md

Examples

The series of demos featured in this directory exemplify a broad spectrum of workflows for deploying ML models on edge devices using ExecuTorch. These demos offer practical insights into key processes such as model exporting, quantization, backend delegation, module composition, memory planning, program saving and loading for inference on ExecuTorch runtime.

ExecuTorch's extensive support spans from simple modules like "Add" to comprehensive models like MobileNet V3, Wav2Letter, Llama 2, and more, showcasing its versatility in enabling the deployment of a wide spectrum of models across various edge AI applications.

Directory structure

examples
├── llm_manual # A storage place for the files that [LLM Maunal](https://pytorch.org/executorch/main/llm/getting-started) needs
├── models # Contains a set of popular and representative PyTorch models
├── portable # Contains end-to-end demos for ExecuTorch in portable mode
├── selective_build # Contains demos of selective build for optimizing the binary size of the ExecuTorch runtime
├── devtools # Contains demos of BundledProgram and ETDump
├── demo-apps # Contains demo apps for Android and iOS
├── xnnpack # Contains end-to-end ExecuTorch demos with first-party optimization using XNNPACK
├── apple
| └── coreml # Contains demos of Apple's Core ML backend
├── arm # Contains demos of the Arm TOSA and Ethos-U NPU flows
├── qualcomm # Contains demos of Qualcomm QNN backend
�├── samsung # Contains demos of Samsung Exynos backend
├── cadence # Contains demos of exporting and running a simple model on Xtensa DSPs
├── third-party # Third-party libraries required for working on the demos
└── README.md # This file

Using the examples

A user's journey may commence by exploring the demos located in the portable/ directory. Here, you will gain insights into the fundamental end-to-end workflow to generate a binary file from a ML model in portable mode and run it on the ExecuTorch runtime.

Demos Apps

Explore mobile apps with ExecuTorch models integrated and deployable on Android and iOS. This provides end-to-end instructions on how to export Llama models, load on device, build the app, and run it on device.

For specific details related to models and backend, you can explore the various subsections.

Llama Models

This page demonstrates how to run Llama 3.2 (1B, 3B), Llama 3.1 (8B), Llama 3 (8B), and Llama 2 7B models on mobile via ExecuTorch. We use XNNPACK, QNNPACK, and MediaTek to accelerate the performance and 4-bit groupwise PTQ quantization to fit the model on Android and iOS mobile phones.

Llava1.5 7B

This page demonstrates how to run Llava 1.5 7B model on mobile via ExecuTorch. We use XNNPACK to accelerate the performance and 4-bit groupwise PTQ quantization to fit the model on Android and iOS mobile phones.

Selective Build

To understand how to deploy the ExecuTorch runtime with optimization for binary size, explore the demos available in the selective_build directory. These demos are specifically designed to illustrate the Selective Build, offering insights into reducing the binary size while maintaining efficiency.

Developer Tools

You will find demos of ExecuTorch Developer Tools in the devtools directory. The examples focuses on exporting and executing BundledProgram for ExecuTorch model verification and ETDump for collecting profiling and debug data.

XNNPACK delegation

The demos in the xnnpack/ directory provide valuable insights into the process of lowering and executing an ExecuTorch model with built-in performance enhancements. These demos specifically showcase the workflow involving XNNPACK backend delegation and quantization.

Apple Backend

You will find demos of ExecuTorch Core ML Backend in the apple/coreml directory.

ARM Cortex-M55 + Ethos-U55 Backend

The arm directory contains scripts to help you run a PyTorch model on a ARM Corstone-300 platform via ExecuTorch.

QNN Backend

You will find demos of ExecuTorch QNN Backend in the qualcomm directory.

Exynos Backend

You will find demos of ExecuTorch Exynos Backend in the samsung directory.

Cadence HiFi4 DSP

The Cadence directory hosts a demo that showcases the process of exporting and executing a model on Xtensa Hifi4 DSP. You can utilize this tutorial to guide you in configuring the demo and running it.

Dependencies

Various models and workflows listed in this directory have dependencies on some other packages. You need to follow the setup guide in Setting up ExecuTorch from GitHub to have appropriate packages installed.

Disclaimer

The ExecuTorch Repository Content is provided without any guarantees about performance or compatibility. In particular, ExecuTorch makes available model architectures written in Python for PyTorch that may not perform in the same manner or meet the same standards as the original versions of those models. When using the ExecuTorch Repository Content, including any model architectures, you are solely responsible for determining the appropriateness of using or redistributing the ExecuTorch Repository Content and assume any risks associated with your use of the ExecuTorch Repository Content or any models, outputs, or results, both alone and in combination with any other technologies. Additionally, you may have other legal obligations that govern your use of other content, such as the terms of service for third-party models, weights, data, or other technologies, and you are solely responsible for complying with all such obligations.

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

History

README.md

Examples

The series of demos featured in this directory exemplify a broad spectrum of workflows for deploying ML models on edge devices using ExecuTorch. These demos offer practical insights into key processes such as model exporting, quantization, backend delegation, module composition, memory planning, program saving and loading for inference on ExecuTorch runtime.

ExecuTorch's extensive support spans from simple modules like "Add" to comprehensive models like MobileNet V3, Wav2Letter, Llama 2, and more, showcasing its versatility in enabling the deployment of a wide spectrum of models across various edge AI applications.

Directory structure

examples
├── llm_manual # A storage place for the files that [LLM Maunal](https://pytorch.org/executorch/main/llm/getting-started) needs
├── models # Contains a set of popular and representative PyTorch models
├── portable # Contains end-to-end demos for ExecuTorch in portable mode
├── selective_build # Contains demos of selective build for optimizing the binary size of the ExecuTorch runtime
├── devtools # Contains demos of BundledProgram and ETDump
├── demo-apps # Contains demo apps for Android and iOS
├── xnnpack # Contains end-to-end ExecuTorch demos with first-party optimization using XNNPACK
├── apple
| └── coreml # Contains demos of Apple's Core ML backend
├── arm # Contains demos of the Arm TOSA and Ethos-U NPU flows
├── qualcomm # Contains demos of Qualcomm QNN backend
�├── samsung # Contains demos of Samsung Exynos backend
├── cadence # Contains demos of exporting and running a simple model on Xtensa DSPs
├── third-party # Third-party libraries required for working on the demos
└── README.md # This file

Using the examples

A user's journey may commence by exploring the demos located in the portable/ directory. Here, you will gain insights into the fundamental end-to-end workflow to generate a binary file from a ML model in portable mode and run it on the ExecuTorch runtime.

Demos Apps

Explore mobile apps with ExecuTorch models integrated and deployable on Android and iOS. This provides end-to-end instructions on how to export Llama models, load on device, build the app, and run it on device.

For specific details related to models and backend, you can explore the various subsections.

Llama Models

This page demonstrates how to run Llama 3.2 (1B, 3B), Llama 3.1 (8B), Llama 3 (8B), and Llama 2 7B models on mobile via ExecuTorch. We use XNNPACK, QNNPACK, and MediaTek to accelerate the performance and 4-bit groupwise PTQ quantization to fit the model on Android and iOS mobile phones.

Llava1.5 7B

This page demonstrates how to run Llava 1.5 7B model on mobile via ExecuTorch. We use XNNPACK to accelerate the performance and 4-bit groupwise PTQ quantization to fit the model on Android and iOS mobile phones.

Selective Build

To understand how to deploy the ExecuTorch runtime with optimization for binary size, explore the demos available in the selective_build directory. These demos are specifically designed to illustrate the Selective Build, offering insights into reducing the binary size while maintaining efficiency.

Developer Tools

You will find demos of ExecuTorch Developer Tools in the devtools directory. The examples focuses on exporting and executing BundledProgram for ExecuTorch model verification and ETDump for collecting profiling and debug data.

XNNPACK delegation

The demos in the xnnpack/ directory provide valuable insights into the process of lowering and executing an ExecuTorch model with built-in performance enhancements. These demos specifically showcase the workflow involving XNNPACK backend delegation and quantization.

Apple Backend

You will find demos of ExecuTorch Core ML Backend in the apple/coreml directory.

ARM Cortex-M55 + Ethos-U55 Backend

The arm directory contains scripts to help you run a PyTorch model on a ARM Corstone-300 platform via ExecuTorch.

QNN Backend

You will find demos of ExecuTorch QNN Backend in the qualcomm directory.

Exynos Backend

You will find demos of ExecuTorch Exynos Backend in the samsung directory.

Cadence HiFi4 DSP

The Cadence directory hosts a demo that showcases the process of exporting and executing a model on Xtensa Hifi4 DSP. You can utilize this tutorial to guide you in configuring the demo and running it.

Dependencies

Various models and workflows listed in this directory have dependencies on some other packages. You need to follow the setup guide in Setting up ExecuTorch from GitHub to have appropriate packages installed.

Disclaimer

The ExecuTorch Repository Content is provided without any guarantees about performance or compatibility. In particular, ExecuTorch makes available model architectures written in Python for PyTorch that may not perform in the same manner or meet the same standards as the original versions of those models. When using the ExecuTorch Repository Content, including any model architectures, you are solely responsible for determining the appropriateness of using or redistributing the ExecuTorch Repository Content and assume any risks associated with your use of the ExecuTorch Repository Content or any models, outputs, or results, both alone and in combination with any other technologies. Additionally, you may have other legal obligations that govern your use of other content, such as the terms of service for third-party models, weights, data, or other technologies, and you are solely responsible for complying with all such obligations.

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

History

README.md

Examples

The series of demos featured in this directory exemplify a broad spectrum of workflows for deploying ML models on edge devices using ExecuTorch. These demos offer practical insights into key processes such as model exporting, quantization, backend delegation, module composition, memory planning, program saving and loading for inference on ExecuTorch runtime.

ExecuTorch's extensive support spans from simple modules like "Add" to comprehensive models like MobileNet V3, Wav2Letter, Llama 2, and more, showcasing its versatility in enabling the deployment of a wide spectrum of models across various edge AI applications.

Directory structure

examples
├── llm_manual # A storage place for the files that [LLM Maunal](https://pytorch.org/executorch/main/llm/getting-started) needs
├── models # Contains a set of popular and representative PyTorch models
├── portable # Contains end-to-end demos for ExecuTorch in portable mode
├── selective_build # Contains demos of selective build for optimizing the binary size of the ExecuTorch runtime
├── devtools # Contains demos of BundledProgram and ETDump
├── demo-apps # Contains demo apps for Android and iOS
├── xnnpack # Contains end-to-end ExecuTorch demos with first-party optimization using XNNPACK
├── apple
| └── coreml # Contains demos of Apple's Core ML backend
├── arm # Contains demos of the Arm TOSA and Ethos-U NPU flows
├── qualcomm # Contains demos of Qualcomm QNN backend
�├── samsung # Contains demos of Samsung Exynos backend
├── cadence # Contains demos of exporting and running a simple model on Xtensa DSPs
├── third-party # Third-party libraries required for working on the demos
└── README.md # This file

Using the examples

A user's journey may commence by exploring the demos located in the portable/ directory. Here, you will gain insights into the fundamental end-to-end workflow to generate a binary file from a ML model in portable mode and run it on the ExecuTorch runtime.

Demos Apps

Explore mobile apps with ExecuTorch models integrated and deployable on Android and iOS. This provides end-to-end instructions on how to export Llama models, load on device, build the app, and run it on device.

For specific details related to models and backend, you can explore the various subsections.

Llama Models

This page demonstrates how to run Llama 3.2 (1B, 3B), Llama 3.1 (8B), Llama 3 (8B), and Llama 2 7B models on mobile via ExecuTorch. We use XNNPACK, QNNPACK, and MediaTek to accelerate the performance and 4-bit groupwise PTQ quantization to fit the model on Android and iOS mobile phones.

Llava1.5 7B

This page demonstrates how to run Llava 1.5 7B model on mobile via ExecuTorch. We use XNNPACK to accelerate the performance and 4-bit groupwise PTQ quantization to fit the model on Android and iOS mobile phones.

Selective Build

To understand how to deploy the ExecuTorch runtime with optimization for binary size, explore the demos available in the selective_build directory. These demos are specifically designed to illustrate the Selective Build, offering insights into reducing the binary size while maintaining efficiency.

Developer Tools

You will find demos of ExecuTorch Developer Tools in the devtools directory. The examples focuses on exporting and executing BundledProgram for ExecuTorch model verification and ETDump for collecting profiling and debug data.

XNNPACK delegation

The demos in the xnnpack/ directory provide valuable insights into the process of lowering and executing an ExecuTorch model with built-in performance enhancements. These demos specifically showcase the workflow involving XNNPACK backend delegation and quantization.

Apple Backend

You will find demos of ExecuTorch Core ML Backend in the apple/coreml directory.

ARM Cortex-M55 + Ethos-U55 Backend

The arm directory contains scripts to help you run a PyTorch model on a ARM Corstone-300 platform via ExecuTorch.

QNN Backend

You will find demos of ExecuTorch QNN Backend in the qualcomm directory.

Exynos Backend

You will find demos of ExecuTorch Exynos Backend in the samsung directory.

Cadence HiFi4 DSP

The Cadence directory hosts a demo that showcases the process of exporting and executing a model on Xtensa Hifi4 DSP. You can utilize this tutorial to guide you in configuring the demo and running it.

Dependencies

Various models and workflows listed in this directory have dependencies on some other packages. You need to follow the setup guide in Setting up ExecuTorch from GitHub to have appropriate packages installed.

Disclaimer

The ExecuTorch Repository Content is provided without any guarantees about performance or compatibility. In particular, ExecuTorch makes available model architectures written in Python for PyTorch that may not perform in the same manner or meet the same standards as the original versions of those models. When using the ExecuTorch Repository Content, including any model architectures, you are solely responsible for determining the appropriateness of using or redistributing the ExecuTorch Repository Content and assume any risks associated with your use of the ExecuTorch Repository Content or any models, outputs, or results, both alone and in combination with any other technologies. Additionally, you may have other legal obligations that govern your use of other content, such as the terms of service for third-party models, weights, data, or other technologies, and you are solely responsible for complying with all such obligations.

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

History

README.md

Examples

The series of demos featured in this directory exemplify a broad spectrum of workflows for deploying ML models on edge devices using ExecuTorch. These demos offer practical insights into key processes such as model exporting, quantization, backend delegation, module composition, memory planning, program saving and loading for inference on ExecuTorch runtime.

ExecuTorch's extensive support spans from simple modules like "Add" to comprehensive models like MobileNet V3, Wav2Letter, Llama 2, and more, showcasing its versatility in enabling the deployment of a wide spectrum of models across various edge AI applications.

Directory structure

examples
├── llm_manual # A storage place for the files that [LLM Maunal](https://pytorch.org/executorch/main/llm/getting-started) needs
├── models # Contains a set of popular and representative PyTorch models
├── portable # Contains end-to-end demos for ExecuTorch in portable mode
├── selective_build # Contains demos of selective build for optimizing the binary size of the ExecuTorch runtime
├── devtools # Contains demos of BundledProgram and ETDump
├── demo-apps # Contains demo apps for Android and iOS
├── xnnpack # Contains end-to-end ExecuTorch demos with first-party optimization using XNNPACK
├── apple
| └── coreml # Contains demos of Apple's Core ML backend
├── arm # Contains demos of the Arm TOSA and Ethos-U NPU flows
├── qualcomm # Contains demos of Qualcomm QNN backend
�├── samsung # Contains demos of Samsung Exynos backend
├── cadence # Contains demos of exporting and running a simple model on Xtensa DSPs
├── third-party # Third-party libraries required for working on the demos
└── README.md # This file

Using the examples

A user's journey may commence by exploring the demos located in the portable/ directory. Here, you will gain insights into the fundamental end-to-end workflow to generate a binary file from a ML model in portable mode and run it on the ExecuTorch runtime.

Demos Apps

Explore mobile apps with ExecuTorch models integrated and deployable on Android and iOS. This provides end-to-end instructions on how to export Llama models, load on device, build the app, and run it on device.

For specific details related to models and backend, you can explore the various subsections.

Llama Models

This page demonstrates how to run Llama 3.2 (1B, 3B), Llama 3.1 (8B), Llama 3 (8B), and Llama 2 7B models on mobile via ExecuTorch. We use XNNPACK, QNNPACK, and MediaTek to accelerate the performance and 4-bit groupwise PTQ quantization to fit the model on Android and iOS mobile phones.

Llava1.5 7B

This page demonstrates how to run Llava 1.5 7B model on mobile via ExecuTorch. We use XNNPACK to accelerate the performance and 4-bit groupwise PTQ quantization to fit the model on Android and iOS mobile phones.

Selective Build

To understand how to deploy the ExecuTorch runtime with optimization for binary size, explore the demos available in the selective_build directory. These demos are specifically designed to illustrate the Selective Build, offering insights into reducing the binary size while maintaining efficiency.

Developer Tools

You will find demos of ExecuTorch Developer Tools in the devtools directory. The examples focuses on exporting and executing BundledProgram for ExecuTorch model verification and ETDump for collecting profiling and debug data.

XNNPACK delegation

The demos in the xnnpack/ directory provide valuable insights into the process of lowering and executing an ExecuTorch model with built-in performance enhancements. These demos specifically showcase the workflow involving XNNPACK backend delegation and quantization.

Apple Backend

You will find demos of ExecuTorch Core ML Backend in the apple/coreml directory.

ARM Cortex-M55 + Ethos-U55 Backend

The arm directory contains scripts to help you run a PyTorch model on a ARM Corstone-300 platform via ExecuTorch.

QNN Backend

You will find demos of ExecuTorch QNN Backend in the qualcomm directory.

Exynos Backend

You will find demos of ExecuTorch Exynos Backend in the samsung directory.

Cadence HiFi4 DSP

The Cadence directory hosts a demo that showcases the process of exporting and executing a model on Xtensa Hifi4 DSP. You can utilize this tutorial to guide you in configuring the demo and running it.

Dependencies

Various models and workflows listed in this directory have dependencies on some other packages. You need to follow the setup guide in Setting up ExecuTorch from GitHub to have appropriate packages installed.

Disclaimer

The ExecuTorch Repository Content is provided without any guarantees about performance or compatibility. In particular, ExecuTorch makes available model architectures written in Python for PyTorch that may not perform in the same manner or meet the same standards as the original versions of those models. When using the ExecuTorch Repository Content, including any model architectures, you are solely responsible for determining the appropriateness of using or redistributing the ExecuTorch Repository Content and assume any risks associated with your use of the ExecuTorch Repository Content or any models, outputs, or results, both alone and in combination with any other technologies. Additionally, you may have other legal obligations that govern your use of other content, such as the terms of service for third-party models, weights, data, or other technologies, and you are solely responsible for complying with all such obligations.

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

History

README.md

Examples

The series of demos featured in this directory exemplify a broad spectrum of workflows for deploying ML models on edge devices using ExecuTorch. These demos offer practical insights into key processes such as model exporting, quantization, backend delegation, module composition, memory planning, program saving and loading for inference on ExecuTorch runtime.

ExecuTorch's extensive support spans from simple modules like "Add" to comprehensive models like MobileNet V3, Wav2Letter, Llama 2, and more, showcasing its versatility in enabling the deployment of a wide spectrum of models across various edge AI applications.

Directory structure

examples
├── llm_manual # A storage place for the files that [LLM Maunal](https://pytorch.org/executorch/main/llm/getting-started) needs
├── models # Contains a set of popular and representative PyTorch models
├── portable # Contains end-to-end demos for ExecuTorch in portable mode
├── selective_build # Contains demos of selective build for optimizing the binary size of the ExecuTorch runtime
├── devtools # Contains demos of BundledProgram and ETDump
├── demo-apps # Contains demo apps for Android and iOS
├── xnnpack # Contains end-to-end ExecuTorch demos with first-party optimization using XNNPACK
├── apple
| └── coreml # Contains demos of Apple's Core ML backend
├── arm # Contains demos of the Arm TOSA and Ethos-U NPU flows
├── qualcomm # Contains demos of Qualcomm QNN backend
�├── samsung # Contains demos of Samsung Exynos backend
├── cadence # Contains demos of exporting and running a simple model on Xtensa DSPs
├── third-party # Third-party libraries required for working on the demos
└── README.md # This file

Using the examples

A user's journey may commence by exploring the demos located in the portable/ directory. Here, you will gain insights into the fundamental end-to-end workflow to generate a binary file from a ML model in portable mode and run it on the ExecuTorch runtime.

Demos Apps

Explore mobile apps with ExecuTorch models integrated and deployable on Android and iOS. This provides end-to-end instructions on how to export Llama models, load on device, build the app, and run it on device.

For specific details related to models and backend, you can explore the various subsections.

Llama Models

This page demonstrates how to run Llama 3.2 (1B, 3B), Llama 3.1 (8B), Llama 3 (8B), and Llama 2 7B models on mobile via ExecuTorch. We use XNNPACK, QNNPACK, and MediaTek to accelerate the performance and 4-bit groupwise PTQ quantization to fit the model on Android and iOS mobile phones.

Llava1.5 7B

This page demonstrates how to run Llava 1.5 7B model on mobile via ExecuTorch. We use XNNPACK to accelerate the performance and 4-bit groupwise PTQ quantization to fit the model on Android and iOS mobile phones.

Selective Build

To understand how to deploy the ExecuTorch runtime with optimization for binary size, explore the demos available in the selective_build directory. These demos are specifically designed to illustrate the Selective Build, offering insights into reducing the binary size while maintaining efficiency.

Developer Tools

You will find demos of ExecuTorch Developer Tools in the devtools directory. The examples focuses on exporting and executing BundledProgram for ExecuTorch model verification and ETDump for collecting profiling and debug data.

XNNPACK delegation

The demos in the xnnpack/ directory provide valuable insights into the process of lowering and executing an ExecuTorch model with built-in performance enhancements. These demos specifically showcase the workflow involving XNNPACK backend delegation and quantization.

Apple Backend

You will find demos of ExecuTorch Core ML Backend in the apple/coreml directory.

ARM Cortex-M55 + Ethos-U55 Backend

The arm directory contains scripts to help you run a PyTorch model on a ARM Corstone-300 platform via ExecuTorch.

QNN Backend

You will find demos of ExecuTorch QNN Backend in the qualcomm directory.

Exynos Backend

You will find demos of ExecuTorch Exynos Backend in the samsung directory.

Cadence HiFi4 DSP

The Cadence directory hosts a demo that showcases the process of exporting and executing a model on Xtensa Hifi4 DSP. You can utilize this tutorial to guide you in configuring the demo and running it.

Dependencies

Various models and workflows listed in this directory have dependencies on some other packages. You need to follow the setup guide in Setting up ExecuTorch from GitHub to have appropriate packages installed.

Disclaimer

The ExecuTorch Repository Content is provided without any guarantees about performance or compatibility. In particular, ExecuTorch makes available model architectures written in Python for PyTorch that may not perform in the same manner or meet the same standards as the original versions of those models. When using the ExecuTorch Repository Content, including any model architectures, you are solely responsible for determining the appropriateness of using or redistributing the ExecuTorch Repository Content and assume any risks associated with your use of the ExecuTorch Repository Content or any models, outputs, or results, both alone and in combination with any other technologies. Additionally, you may have other legal obligations that govern your use of other content, such as the terms of service for third-party models, weights, data, or other technologies, and you are solely responsible for complying with all such obligations.

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

History

README.md

Examples

The series of demos featured in this directory exemplify a broad spectrum of workflows for deploying ML models on edge devices using ExecuTorch. These demos offer practical insights into key processes such as model exporting, quantization, backend delegation, module composition, memory planning, program saving and loading for inference on ExecuTorch runtime.

ExecuTorch's extensive support spans from simple modules like "Add" to comprehensive models like MobileNet V3, Wav2Letter, Llama 2, and more, showcasing its versatility in enabling the deployment of a wide spectrum of models across various edge AI applications.

Directory structure

examples
├── llm_manual # A storage place for the files that [LLM Maunal](https://pytorch.org/executorch/main/llm/getting-started) needs
├── models # Contains a set of popular and representative PyTorch models
├── portable # Contains end-to-end demos for ExecuTorch in portable mode
├── selective_build # Contains demos of selective build for optimizing the binary size of the ExecuTorch runtime
├── devtools # Contains demos of BundledProgram and ETDump
├── demo-apps # Contains demo apps for Android and iOS
├── xnnpack # Contains end-to-end ExecuTorch demos with first-party optimization using XNNPACK
├── apple
| └── coreml # Contains demos of Apple's Core ML backend
├── arm # Contains demos of the Arm TOSA and Ethos-U NPU flows
├── qualcomm # Contains demos of Qualcomm QNN backend
�├── samsung # Contains demos of Samsung Exynos backend
├── cadence # Contains demos of exporting and running a simple model on Xtensa DSPs
├── third-party # Third-party libraries required for working on the demos
└── README.md # This file

Using the examples

A user's journey may commence by exploring the demos located in the portable/ directory. Here, you will gain insights into the fundamental end-to-end workflow to generate a binary file from a ML model in portable mode and run it on the ExecuTorch runtime.

Demos Apps

Explore mobile apps with ExecuTorch models integrated and deployable on Android and iOS. This provides end-to-end instructions on how to export Llama models, load on device, build the app, and run it on device.

For specific details related to models and backend, you can explore the various subsections.

Llama Models

This page demonstrates how to run Llama 3.2 (1B, 3B), Llama 3.1 (8B), Llama 3 (8B), and Llama 2 7B models on mobile via ExecuTorch. We use XNNPACK, QNNPACK, and MediaTek to accelerate the performance and 4-bit groupwise PTQ quantization to fit the model on Android and iOS mobile phones.

Llava1.5 7B

This page demonstrates how to run Llava 1.5 7B model on mobile via ExecuTorch. We use XNNPACK to accelerate the performance and 4-bit groupwise PTQ quantization to fit the model on Android and iOS mobile phones.

Selective Build

To understand how to deploy the ExecuTorch runtime with optimization for binary size, explore the demos available in the selective_build directory. These demos are specifically designed to illustrate the Selective Build, offering insights into reducing the binary size while maintaining efficiency.

Developer Tools

You will find demos of ExecuTorch Developer Tools in the devtools directory. The examples focuses on exporting and executing BundledProgram for ExecuTorch model verification and ETDump for collecting profiling and debug data.

XNNPACK delegation

The demos in the xnnpack/ directory provide valuable insights into the process of lowering and executing an ExecuTorch model with built-in performance enhancements. These demos specifically showcase the workflow involving XNNPACK backend delegation and quantization.

Apple Backend

You will find demos of ExecuTorch Core ML Backend in the apple/coreml directory.

ARM Cortex-M55 + Ethos-U55 Backend

The arm directory contains scripts to help you run a PyTorch model on a ARM Corstone-300 platform via ExecuTorch.

QNN Backend

You will find demos of ExecuTorch QNN Backend in the qualcomm directory.

Exynos Backend

You will find demos of ExecuTorch Exynos Backend in the samsung directory.

Cadence HiFi4 DSP

The Cadence directory hosts a demo that showcases the process of exporting and executing a model on Xtensa Hifi4 DSP. You can utilize this tutorial to guide you in configuring the demo and running it.

Dependencies

Various models and workflows listed in this directory have dependencies on some other packages. You need to follow the setup guide in Setting up ExecuTorch from GitHub to have appropriate packages installed.

Disclaimer

The ExecuTorch Repository Content is provided without any guarantees about performance or compatibility. In particular, ExecuTorch makes available model architectures written in Python for PyTorch that may not perform in the same manner or meet the same standards as the original versions of those models. When using the ExecuTorch Repository Content, including any model architectures, you are solely responsible for determining the appropriateness of using or redistributing the ExecuTorch Repository Content and assume any risks associated with your use of the ExecuTorch Repository Content or any models, outputs, or results, both alone and in combination with any other technologies. Additionally, you may have other legal obligations that govern your use of other content, such as the terms of service for third-party models, weights, data, or other technologies, and you are solely responsible for complying with all such obligations.

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

History

README.md

Examples

The series of demos featured in this directory exemplify a broad spectrum of workflows for deploying ML models on edge devices using ExecuTorch. These demos offer practical insights into key processes such as model exporting, quantization, backend delegation, module composition, memory planning, program saving and loading for inference on ExecuTorch runtime.

ExecuTorch's extensive support spans from simple modules like "Add" to comprehensive models like MobileNet V3, Wav2Letter, Llama 2, and more, showcasing its versatility in enabling the deployment of a wide spectrum of models across various edge AI applications.

Directory structure

examples
├── llm_manual # A storage place for the files that [LLM Maunal](https://pytorch.org/executorch/main/llm/getting-started) needs
├── models # Contains a set of popular and representative PyTorch models
├── portable # Contains end-to-end demos for ExecuTorch in portable mode
├── selective_build # Contains demos of selective build for optimizing the binary size of the ExecuTorch runtime
├── devtools # Contains demos of BundledProgram and ETDump
├── demo-apps # Contains demo apps for Android and iOS
├── xnnpack # Contains end-to-end ExecuTorch demos with first-party optimization using XNNPACK
├── apple
| └── coreml # Contains demos of Apple's Core ML backend
├── arm # Contains demos of the Arm TOSA and Ethos-U NPU flows
├── qualcomm # Contains demos of Qualcomm QNN backend
�├── samsung # Contains demos of Samsung Exynos backend
├── cadence # Contains demos of exporting and running a simple model on Xtensa DSPs
├── third-party # Third-party libraries required for working on the demos
└── README.md # This file

Using the examples

A user's journey may commence by exploring the demos located in the portable/ directory. Here, you will gain insights into the fundamental end-to-end workflow to generate a binary file from a ML model in portable mode and run it on the ExecuTorch runtime.

Demos Apps

Explore mobile apps with ExecuTorch models integrated and deployable on Android and iOS. This provides end-to-end instructions on how to export Llama models, load on device, build the app, and run it on device.

For specific details related to models and backend, you can explore the various subsections.

Llama Models

This page demonstrates how to run Llama 3.2 (1B, 3B), Llama 3.1 (8B), Llama 3 (8B), and Llama 2 7B models on mobile via ExecuTorch. We use XNNPACK, QNNPACK, and MediaTek to accelerate the performance and 4-bit groupwise PTQ quantization to fit the model on Android and iOS mobile phones.

Llava1.5 7B

This page demonstrates how to run Llava 1.5 7B model on mobile via ExecuTorch. We use XNNPACK to accelerate the performance and 4-bit groupwise PTQ quantization to fit the model on Android and iOS mobile phones.

Selective Build

To understand how to deploy the ExecuTorch runtime with optimization for binary size, explore the demos available in the selective_build directory. These demos are specifically designed to illustrate the Selective Build, offering insights into reducing the binary size while maintaining efficiency.

Developer Tools

You will find demos of ExecuTorch Developer Tools in the devtools directory. The examples focuses on exporting and executing BundledProgram for ExecuTorch model verification and ETDump for collecting profiling and debug data.

XNNPACK delegation

The demos in the xnnpack/ directory provide valuable insights into the process of lowering and executing an ExecuTorch model with built-in performance enhancements. These demos specifically showcase the workflow involving XNNPACK backend delegation and quantization.

Apple Backend

You will find demos of ExecuTorch Core ML Backend in the apple/coreml directory.

ARM Cortex-M55 + Ethos-U55 Backend

The arm directory contains scripts to help you run a PyTorch model on a ARM Corstone-300 platform via ExecuTorch.

QNN Backend

You will find demos of ExecuTorch QNN Backend in the qualcomm directory.

Exynos Backend

You will find demos of ExecuTorch Exynos Backend in the samsung directory.

Cadence HiFi4 DSP

The Cadence directory hosts a demo that showcases the process of exporting and executing a model on Xtensa Hifi4 DSP. You can utilize this tutorial to guide you in configuring the demo and running it.

Dependencies

Various models and workflows listed in this directory have dependencies on some other packages. You need to follow the setup guide in Setting up ExecuTorch from GitHub to have appropriate packages installed.

Disclaimer

The ExecuTorch Repository Content is provided without any guarantees about performance or compatibility. In particular, ExecuTorch makes available model architectures written in Python for PyTorch that may not perform in the same manner or meet the same standards as the original versions of those models. When using the ExecuTorch Repository Content, including any model architectures, you are solely responsible for determining the appropriateness of using or redistributing the ExecuTorch Repository Content and assume any risks associated with your use of the ExecuTorch Repository Content or any models, outputs, or results, both alone and in combination with any other technologies. Additionally, you may have other legal obligations that govern your use of other content, such as the terms of service for third-party models, weights, data, or other technologies, and you are solely responsible for complying with all such obligations.

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

History

README.md

Examples

The series of demos featured in this directory exemplify a broad spectrum of workflows for deploying ML models on edge devices using ExecuTorch. These demos offer practical insights into key processes such as model exporting, quantization, backend delegation, module composition, memory planning, program saving and loading for inference on ExecuTorch runtime.

ExecuTorch's extensive support spans from simple modules like "Add" to comprehensive models like MobileNet V3, Wav2Letter, Llama 2, and more, showcasing its versatility in enabling the deployment of a wide spectrum of models across various edge AI applications.

Directory structure

examples
├── llm_manual # A storage place for the files that [LLM Maunal](https://pytorch.org/executorch/main/llm/getting-started) needs
├── models # Contains a set of popular and representative PyTorch models
├── portable # Contains end-to-end demos for ExecuTorch in portable mode
├── selective_build # Contains demos of selective build for optimizing the binary size of the ExecuTorch runtime
├── devtools # Contains demos of BundledProgram and ETDump
├── demo-apps # Contains demo apps for Android and iOS
├── xnnpack # Contains end-to-end ExecuTorch demos with first-party optimization using XNNPACK
├── apple
| └── coreml # Contains demos of Apple's Core ML backend
├── arm # Contains demos of the Arm TOSA and Ethos-U NPU flows
├── qualcomm # Contains demos of Qualcomm QNN backend
�├── samsung # Contains demos of Samsung Exynos backend
├── cadence # Contains demos of exporting and running a simple model on Xtensa DSPs
├── third-party # Third-party libraries required for working on the demos
└── README.md # This file

Using the examples

A user's journey may commence by exploring the demos located in the portable/ directory. Here, you will gain insights into the fundamental end-to-end workflow to generate a binary file from a ML model in portable mode and run it on the ExecuTorch runtime.

Demos Apps

Explore mobile apps with ExecuTorch models integrated and deployable on Android and iOS. This provides end-to-end instructions on how to export Llama models, load on device, build the app, and run it on device.

For specific details related to models and backend, you can explore the various subsections.

Llama Models

This page demonstrates how to run Llama 3.2 (1B, 3B), Llama 3.1 (8B), Llama 3 (8B), and Llama 2 7B models on mobile via ExecuTorch. We use XNNPACK, QNNPACK, and MediaTek to accelerate the performance and 4-bit groupwise PTQ quantization to fit the model on Android and iOS mobile phones.

Llava1.5 7B

This page demonstrates how to run Llava 1.5 7B model on mobile via ExecuTorch. We use XNNPACK to accelerate the performance and 4-bit groupwise PTQ quantization to fit the model on Android and iOS mobile phones.

Selective Build

To understand how to deploy the ExecuTorch runtime with optimization for binary size, explore the demos available in the selective_build directory. These demos are specifically designed to illustrate the Selective Build, offering insights into reducing the binary size while maintaining efficiency.

Developer Tools

You will find demos of ExecuTorch Developer Tools in the devtools directory. The examples focuses on exporting and executing BundledProgram for ExecuTorch model verification and ETDump for collecting profiling and debug data.

XNNPACK delegation

The demos in the xnnpack/ directory provide valuable insights into the process of lowering and executing an ExecuTorch model with built-in performance enhancements. These demos specifically showcase the workflow involving XNNPACK backend delegation and quantization.

Apple Backend

You will find demos of ExecuTorch Core ML Backend in the apple/coreml directory.

ARM Cortex-M55 + Ethos-U55 Backend

The arm directory contains scripts to help you run a PyTorch model on a ARM Corstone-300 platform via ExecuTorch.

QNN Backend

You will find demos of ExecuTorch QNN Backend in the qualcomm directory.

Exynos Backend

You will find demos of ExecuTorch Exynos Backend in the samsung directory.

Cadence HiFi4 DSP

The Cadence directory hosts a demo that showcases the process of exporting and executing a model on Xtensa Hifi4 DSP. You can utilize this tutorial to guide you in configuring the demo and running it.

Dependencies

Various models and workflows listed in this directory have dependencies on some other packages. You need to follow the setup guide in Setting up ExecuTorch from GitHub to have appropriate packages installed.

Disclaimer

The ExecuTorch Repository Content is provided without any guarantees about performance or compatibility. In particular, ExecuTorch makes available model architectures written in Python for PyTorch that may not perform in the same manner or meet the same standards as the original versions of those models. When using the ExecuTorch Repository Content, including any model architectures, you are solely responsible for determining the appropriateness of using or redistributing the ExecuTorch Repository Content and assume any risks associated with your use of the ExecuTorch Repository Content or any models, outputs, or results, both alone and in combination with any other technologies. Additionally, you may have other legal obligations that govern your use of other content, such as the terms of service for third-party models, weights, data, or other technologies, and you are solely responsible for complying with all such obligations.