Skip to content

Latest commit

History

13,444 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

ExecuTorch logo mark

ExecuTorch

On-device AI inference powered by PyTorch

PyPI - VersionGitHub - ContributorsGitHub - StarsDiscord - Chat with UsDocumentation

ExecuTorch is PyTorch's unified solution for deploying AI models on-device—from smartphones to microcontrollers—built for privacy, performance, and portability. It powers Meta's on-device AI across Instagram, WhatsApp, Quest 3, Ray-Ban Meta Smart Glasses, and more.

Deploy LLMs, vision, speech, and multimodal models with the same PyTorch APIs you already know—accelerating research to production with seamless model export, optimization, and deployment. No manual C++ rewrites. No format conversions. No vendor lock-in.

📘 Table of Contents

Why ExecuTorch?

  • 🔒 Native PyTorch Export — Direct export from PyTorch. No .onnx, .tflite, or intermediate format conversions. Preserve model semantics.
  • ⚡ Production-Proven — Powers billions of users at Meta with real-time on-device inference.
  • 💾 Tiny Runtime — 50KB base footprint. Runs on microcontrollers to high-end smartphones.
  • 🚀 12+ Hardware Backends — Open-source acceleration for Apple, Samsung, Qualcomm, ARM, MediaTek, Vulkan, and more.
  • 🎯 One Export, Multiple Backends — Switch hardware targets with a single line change. Deploy the same model everywhere.

How It Works

ExecuTorch uses ahead-of-time (AOT) compilation to prepare PyTorch models for edge deployment:

  1. 🧩 Export — Capture your PyTorch model graph with torch.export()
  2. ⚙️ Compile — Quantize, optimize, and partition to hardware backends → .pte
  3. 🚀 Execute — Load .pte on-device via lightweight C++ runtime

Models use a standardized Core ATen operator set. Partitioners delegate subgraphs to specialized hardware (NPU/GPU) with CPU fallback.

Learn more: How ExecuTorch WorksArchitecture Guide

Quick Start

Installation

pip install executorch

For platform-specific setup (Android, iOS, embedded systems), see the Quick Start documentation for additional info.

Export and Deploy in 3 Steps

importtorchfromexecutorch.exirimportto_edge_transform_and_lowerfromexecutorch.backends.xnnpack.partition.xnnpack_partitionerimportXnnpackPartitioner# 1. Export your PyTorch modelmodel=MyModel().eval()
example_inputs= (torch.randn(1, 3, 224, 224),)
exported_program=torch.export.export(model, example_inputs)
# 2. Optimize for target hardware (switch backends with one line)program=to_edge_transform_and_lower(
exported_program,
partitioner=[XnnpackPartitioner()] # CPU | CoreMLPartitioner() for iOS | QnnPartitioner() for Qualcomm
).to_executorch()
# 3. Save for deploymentwithopen("model.pte", "wb") asf:
f.write(program.buffer)
# Test locally via ExecuTorch runtime's pybind API (optional)fromexecutorch.runtimeimportRuntimeruntime=Runtime.get()
method=runtime.load_program("model.pte").load_method("forward")
outputs=method.execute([torch.randn(1, 3, 224, 224)])

Run on Device

C++

#include<executorch/extension/module/module.h>
#include<executorch/extension/tensor/tensor.h>
Module module("model.pte");
auto tensor = make_tensor_ptr({2, 2}, {1.0f, 2.0f, 3.0f, 4.0f});
auto outputs = module.forward(tensor);

Swift (iOS)

import ExecuTorch
letmodule=Module(filePath:"model.pte")letinput=Tensor<Float>([1.0,2.0,3.0,4.0], shape:[2,2])letoutputs=try module.forward(input)

Kotlin (Android)

val module =Module.load("model.pte")
val inputTensor =Tensor.fromBlob(floatArrayOf(1.0f, 2.0f, 3.0f, 4.0f), longArrayOf(2, 2))
val outputs = module.forward(EValue.from(inputTensor))

LLM Example: Llama

Export Llama models using the export_llm script or Optimum-ExecuTorch:

# Using export_llm
python -m executorch.extension.llm.export.export_llm --model llama3_2 --output llama.pte
# Using Optimum-ExecuTorch
optimum-cli export executorch \
--model meta-llama/Llama-3.2-1B \
--task text-generation \
--recipe xnnpack \
--output_dir llama_model

Run on-device with the LLM runner API:

C++

#include<executorch/extension/llm/runner/text_llm_runner.h>auto runner = create_llama_runner("llama.pte", "tiktoken.bin");
executorch::extension::llm::GenerationConfig config{
.seq_len = 128, .temperature = 0.8f};
runner->generate("Hello, how are you?", config);

Swift (iOS)

import ExecuTorchLLM
letrunner=TextRunner(modelPath:"llama.pte", tokenizerPath:"tiktoken.bin")try runner.generate("Hello, how are you?",Config{
$0.sequenceLength =128}){ token inprint(token, terminator:"")}

Kotlin (Android)API DocsDemo App

val llmModule =LlmModule("llama.pte", "tiktoken.bin", 0.8f)
llmModule.load()
llmModule.generate("Hello, how are you?", 128, object:LlmCallback {
overridefunonResult(result:String) { print(result) }
overridefunonStats(stats:String) { }
})

For multimodal models (vision, audio), use the MultiModal runner API which extends the LLM runner to handle image and audio inputs alongside text. See Llava and Voxtral examples.

See examples/models/llama for complete workflow including quantization, mobile deployment, and advanced options.

Next Steps:

Platform & Hardware Support

PlatformSupported Backends
AndroidXNNPACK, Vulkan, Qualcomm, MediaTek, Samsung Exynos
iOSXNNPACK, CoreML (Neural Engine)
Linux / WindowsXNNPACK, OpenVINO, CUDA (experimental)
macOSXNNPACK, Metal (experimental), MLX (experimental)
Embedded / MCUXNNPACK, ARM Ethos-U, NXP, Cadence DSP

See Backend Documentation for detailed hardware requirements and optimization guides. For desktop/laptop GPU inference with CUDA and Metal, see the Desktop Guide. For Zephyr RTOS integration, see the Zephyr Guide.

Production Deployments

ExecuTorch powers on-device AI at scale across Meta's family of apps, VR/AR devices, and partner deployments. View success stories →

Examples & Models

LLMs:Llama 3.2/3.1/3, Qwen 3, Phi-4-mini, LiquidAI LFM2

Multimodal:Llava (vision-language), Voxtral (audio-language), Gemma (vision-language)

Vision/Speech:MobileNetV2, DeepLabV3, YOLO26, Whisper, Supertonic

Resources:examples/ directory • executorch-examples out-of-tree demos • Optimum-ExecuTorch for HuggingFace models • Unsloth for fine-tuned LLM deployment

Key Features

ExecuTorch provides advanced capabilities for production deployment:

  • Quantization — Built-in support via torchao for 8-bit, 4-bit, and dynamic quantization
  • Memory Planning — Optimize memory usage with ahead-of-time allocation strategies
  • Developer Tools — ETDump profiler, ETRecord inspector, and model debugger
  • Selective Build — Strip unused operators to minimize binary size
  • Custom Operators — Extend with domain-specific kernels
  • Dynamic Shapes — Support variable input sizes with bounded ranges

See Advanced Topics for quantization techniques, custom backends, and compiler passes.

Documentation

Community & Contributing

We welcome contributions from the community!

Citing ExecuTorch

If you found ExecuTorch helpful in your research and would like to acknowledge it, please cite us using the following BibTeX:

@article{executorch2026,
title={{ExecuTorch} - A Unified {PyTorch} Solution to Run {AI} Models On-Device},
author={Nachin, Mergen and Desai, Digant and Jia, Sicheng Stephen and Lai, Chen and Liu, Mengwei and Szwejbka, Jacob and Alvarez, Raziel and Ascani, RJ and Bort, Dave and Candales, Manuel and others},
journal={arXiv preprint arXiv:2605.08195},
url={https://github.com/pytorch/executorch},
year={2026}
}

License

ExecuTorch is BSD licensed, as found in the LICENSE file.




Part of the PyTorch ecosystem

GitHubDocumentation

About

End-to-end solution for enabling on-device AI across mobile and edge devices for PyTorch models

Resources

Code of conduct

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - oscarandersson8218/executorch: End-to-end solution for enabling on-device AI across mobile and edge devices for PyTorch models · GitHub
Skip to content

Latest commit

History

13,444 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

ExecuTorch logo mark

ExecuTorch

On-device AI inference powered by PyTorch

PyPI - VersionGitHub - ContributorsGitHub - StarsDiscord - Chat with UsDocumentation

ExecuTorch is PyTorch's unified solution for deploying AI models on-device—from smartphones to microcontrollers—built for privacy, performance, and portability. It powers Meta's on-device AI across Instagram, WhatsApp, Quest 3, Ray-Ban Meta Smart Glasses, and more.

Deploy LLMs, vision, speech, and multimodal models with the same PyTorch APIs you already know—accelerating research to production with seamless model export, optimization, and deployment. No manual C++ rewrites. No format conversions. No vendor lock-in.

📘 Table of Contents

Why ExecuTorch?

  • 🔒 Native PyTorch Export — Direct export from PyTorch. No .onnx, .tflite, or intermediate format conversions. Preserve model semantics.
  • ⚡ Production-Proven — Powers billions of users at Meta with real-time on-device inference.
  • 💾 Tiny Runtime — 50KB base footprint. Runs on microcontrollers to high-end smartphones.
  • 🚀 12+ Hardware Backends — Open-source acceleration for Apple, Samsung, Qualcomm, ARM, MediaTek, Vulkan, and more.
  • 🎯 One Export, Multiple Backends — Switch hardware targets with a single line change. Deploy the same model everywhere.

How It Works

ExecuTorch uses ahead-of-time (AOT) compilation to prepare PyTorch models for edge deployment:

  1. 🧩 Export — Capture your PyTorch model graph with torch.export()
  2. ⚙️ Compile — Quantize, optimize, and partition to hardware backends → .pte
  3. 🚀 Execute — Load .pte on-device via lightweight C++ runtime

Models use a standardized Core ATen operator set. Partitioners delegate subgraphs to specialized hardware (NPU/GPU) with CPU fallback.

Learn more: How ExecuTorch WorksArchitecture Guide

Quick Start

Installation

pip install executorch

For platform-specific setup (Android, iOS, embedded systems), see the Quick Start documentation for additional info.

Export and Deploy in 3 Steps

importtorchfromexecutorch.exirimportto_edge_transform_and_lowerfromexecutorch.backends.xnnpack.partition.xnnpack_partitionerimportXnnpackPartitioner# 1. Export your PyTorch modelmodel=MyModel().eval()
example_inputs= (torch.randn(1, 3, 224, 224),)
exported_program=torch.export.export(model, example_inputs)
# 2. Optimize for target hardware (switch backends with one line)program=to_edge_transform_and_lower(
exported_program,
partitioner=[XnnpackPartitioner()] # CPU | CoreMLPartitioner() for iOS | QnnPartitioner() for Qualcomm
).to_executorch()
# 3. Save for deploymentwithopen("model.pte", "wb") asf:
f.write(program.buffer)
# Test locally via ExecuTorch runtime's pybind API (optional)fromexecutorch.runtimeimportRuntimeruntime=Runtime.get()
method=runtime.load_program("model.pte").load_method("forward")
outputs=method.execute([torch.randn(1, 3, 224, 224)])

Run on Device

C++

#include<executorch/extension/module/module.h>
#include<executorch/extension/tensor/tensor.h>
Module module("model.pte");
auto tensor = make_tensor_ptr({2, 2}, {1.0f, 2.0f, 3.0f, 4.0f});
auto outputs = module.forward(tensor);

Swift (iOS)

import ExecuTorch
letmodule=Module(filePath:"model.pte")letinput=Tensor<Float>([1.0,2.0,3.0,4.0], shape:[2,2])letoutputs=try module.forward(input)

Kotlin (Android)

val module =Module.load("model.pte")
val inputTensor =Tensor.fromBlob(floatArrayOf(1.0f, 2.0f, 3.0f, 4.0f), longArrayOf(2, 2))
val outputs = module.forward(EValue.from(inputTensor))

LLM Example: Llama

Export Llama models using the export_llm script or Optimum-ExecuTorch:

# Using export_llm
python -m executorch.extension.llm.export.export_llm --model llama3_2 --output llama.pte
# Using Optimum-ExecuTorch
optimum-cli export executorch \
--model meta-llama/Llama-3.2-1B \
--task text-generation \
--recipe xnnpack \
--output_dir llama_model

Run on-device with the LLM runner API:

C++

#include<executorch/extension/llm/runner/text_llm_runner.h>auto runner = create_llama_runner("llama.pte", "tiktoken.bin");
executorch::extension::llm::GenerationConfig config{
.seq_len = 128, .temperature = 0.8f};
runner->generate("Hello, how are you?", config);

Swift (iOS)

import ExecuTorchLLM
letrunner=TextRunner(modelPath:"llama.pte", tokenizerPath:"tiktoken.bin")try runner.generate("Hello, how are you?",Config{
$0.sequenceLength =128}){ token inprint(token, terminator:"")}

Kotlin (Android)API DocsDemo App

val llmModule =LlmModule("llama.pte", "tiktoken.bin", 0.8f)
llmModule.load()
llmModule.generate("Hello, how are you?", 128, object:LlmCallback {
overridefunonResult(result:String) { print(result) }
overridefunonStats(stats:String) { }
})

For multimodal models (vision, audio), use the MultiModal runner API which extends the LLM runner to handle image and audio inputs alongside text. See Llava and Voxtral examples.

See examples/models/llama for complete workflow including quantization, mobile deployment, and advanced options.

Next Steps:

Platform & Hardware Support

PlatformSupported Backends
AndroidXNNPACK, Vulkan, Qualcomm, MediaTek, Samsung Exynos
iOSXNNPACK, CoreML (Neural Engine)
Linux / WindowsXNNPACK, OpenVINO, CUDA (experimental)
macOSXNNPACK, Metal (experimental), MLX (experimental)
Embedded / MCUXNNPACK, ARM Ethos-U, NXP, Cadence DSP

See Backend Documentation for detailed hardware requirements and optimization guides. For desktop/laptop GPU inference with CUDA and Metal, see the Desktop Guide. For Zephyr RTOS integration, see the Zephyr Guide.

Production Deployments

ExecuTorch powers on-device AI at scale across Meta's family of apps, VR/AR devices, and partner deployments. View success stories →

Examples & Models

LLMs:Llama 3.2/3.1/3, Qwen 3, Phi-4-mini, LiquidAI LFM2

Multimodal:Llava (vision-language), Voxtral (audio-language), Gemma (vision-language)

Vision/Speech:MobileNetV2, DeepLabV3, YOLO26, Whisper, Supertonic

Resources:examples/ directory • executorch-examples out-of-tree demos • Optimum-ExecuTorch for HuggingFace models • Unsloth for fine-tuned LLM deployment

Key Features

ExecuTorch provides advanced capabilities for production deployment:

  • Quantization — Built-in support via torchao for 8-bit, 4-bit, and dynamic quantization
  • Memory Planning — Optimize memory usage with ahead-of-time allocation strategies
  • Developer Tools — ETDump profiler, ETRecord inspector, and model debugger
  • Selective Build — Strip unused operators to minimize binary size
  • Custom Operators — Extend with domain-specific kernels
  • Dynamic Shapes — Support variable input sizes with bounded ranges

See Advanced Topics for quantization techniques, custom backends, and compiler passes.

Documentation

Community & Contributing

We welcome contributions from the community!

Citing ExecuTorch

If you found ExecuTorch helpful in your research and would like to acknowledge it, please cite us using the following BibTeX:

@article{executorch2026,
title={{ExecuTorch} - A Unified {PyTorch} Solution to Run {AI} Models On-Device},
author={Nachin, Mergen and Desai, Digant and Jia, Sicheng Stephen and Lai, Chen and Liu, Mengwei and Szwejbka, Jacob and Alvarez, Raziel and Ascani, RJ and Bort, Dave and Candales, Manuel and others},
journal={arXiv preprint arXiv:2605.08195},
url={https://github.com/pytorch/executorch},
year={2026}
}

License

ExecuTorch is BSD licensed, as found in the LICENSE file.




Part of the PyTorch ecosystem

GitHubDocumentation

About

End-to-end solution for enabling on-device AI across mobile and edge devices for PyTorch models

Resources

Code of conduct

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - oscarandersson8218/executorch: End-to-end solution for enabling on-device AI across mobile and edge devices for PyTorch models · GitHub
Skip to content

Latest commit

History

13,444 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

ExecuTorch logo mark

ExecuTorch

On-device AI inference powered by PyTorch

PyPI - VersionGitHub - ContributorsGitHub - StarsDiscord - Chat with UsDocumentation

ExecuTorch is PyTorch's unified solution for deploying AI models on-device—from smartphones to microcontrollers—built for privacy, performance, and portability. It powers Meta's on-device AI across Instagram, WhatsApp, Quest 3, Ray-Ban Meta Smart Glasses, and more.

Deploy LLMs, vision, speech, and multimodal models with the same PyTorch APIs you already know—accelerating research to production with seamless model export, optimization, and deployment. No manual C++ rewrites. No format conversions. No vendor lock-in.

📘 Table of Contents

Why ExecuTorch?

  • 🔒 Native PyTorch Export — Direct export from PyTorch. No .onnx, .tflite, or intermediate format conversions. Preserve model semantics.
  • ⚡ Production-Proven — Powers billions of users at Meta with real-time on-device inference.
  • 💾 Tiny Runtime — 50KB base footprint. Runs on microcontrollers to high-end smartphones.
  • 🚀 12+ Hardware Backends — Open-source acceleration for Apple, Samsung, Qualcomm, ARM, MediaTek, Vulkan, and more.
  • 🎯 One Export, Multiple Backends — Switch hardware targets with a single line change. Deploy the same model everywhere.

How It Works

ExecuTorch uses ahead-of-time (AOT) compilation to prepare PyTorch models for edge deployment:

  1. 🧩 Export — Capture your PyTorch model graph with torch.export()
  2. ⚙️ Compile — Quantize, optimize, and partition to hardware backends → .pte
  3. 🚀 Execute — Load .pte on-device via lightweight C++ runtime

Models use a standardized Core ATen operator set. Partitioners delegate subgraphs to specialized hardware (NPU/GPU) with CPU fallback.

Learn more: How ExecuTorch WorksArchitecture Guide

Quick Start

Installation

pip install executorch

For platform-specific setup (Android, iOS, embedded systems), see the Quick Start documentation for additional info.

Export and Deploy in 3 Steps

importtorchfromexecutorch.exirimportto_edge_transform_and_lowerfromexecutorch.backends.xnnpack.partition.xnnpack_partitionerimportXnnpackPartitioner# 1. Export your PyTorch modelmodel=MyModel().eval()
example_inputs= (torch.randn(1, 3, 224, 224),)
exported_program=torch.export.export(model, example_inputs)
# 2. Optimize for target hardware (switch backends with one line)program=to_edge_transform_and_lower(
exported_program,
partitioner=[XnnpackPartitioner()] # CPU | CoreMLPartitioner() for iOS | QnnPartitioner() for Qualcomm
).to_executorch()
# 3. Save for deploymentwithopen("model.pte", "wb") asf:
f.write(program.buffer)
# Test locally via ExecuTorch runtime's pybind API (optional)fromexecutorch.runtimeimportRuntimeruntime=Runtime.get()
method=runtime.load_program("model.pte").load_method("forward")
outputs=method.execute([torch.randn(1, 3, 224, 224)])

Run on Device

C++

#include<executorch/extension/module/module.h>
#include<executorch/extension/tensor/tensor.h>
Module module("model.pte");
auto tensor = make_tensor_ptr({2, 2}, {1.0f, 2.0f, 3.0f, 4.0f});
auto outputs = module.forward(tensor);

Swift (iOS)

import ExecuTorch
letmodule=Module(filePath:"model.pte")letinput=Tensor<Float>([1.0,2.0,3.0,4.0], shape:[2,2])letoutputs=try module.forward(input)

Kotlin (Android)

val module =Module.load("model.pte")
val inputTensor =Tensor.fromBlob(floatArrayOf(1.0f, 2.0f, 3.0f, 4.0f), longArrayOf(2, 2))
val outputs = module.forward(EValue.from(inputTensor))

LLM Example: Llama

Export Llama models using the export_llm script or Optimum-ExecuTorch:

# Using export_llm
python -m executorch.extension.llm.export.export_llm --model llama3_2 --output llama.pte
# Using Optimum-ExecuTorch
optimum-cli export executorch \
--model meta-llama/Llama-3.2-1B \
--task text-generation \
--recipe xnnpack \
--output_dir llama_model

Run on-device with the LLM runner API:

C++

#include<executorch/extension/llm/runner/text_llm_runner.h>auto runner = create_llama_runner("llama.pte", "tiktoken.bin");
executorch::extension::llm::GenerationConfig config{
.seq_len = 128, .temperature = 0.8f};
runner->generate("Hello, how are you?", config);

Swift (iOS)

import ExecuTorchLLM
letrunner=TextRunner(modelPath:"llama.pte", tokenizerPath:"tiktoken.bin")try runner.generate("Hello, how are you?",Config{
$0.sequenceLength =128}){ token inprint(token, terminator:"")}

Kotlin (Android)API DocsDemo App

val llmModule =LlmModule("llama.pte", "tiktoken.bin", 0.8f)
llmModule.load()
llmModule.generate("Hello, how are you?", 128, object:LlmCallback {
overridefunonResult(result:String) { print(result) }
overridefunonStats(stats:String) { }
})

For multimodal models (vision, audio), use the MultiModal runner API which extends the LLM runner to handle image and audio inputs alongside text. See Llava and Voxtral examples.

See examples/models/llama for complete workflow including quantization, mobile deployment, and advanced options.

Next Steps:

Platform & Hardware Support

PlatformSupported Backends
AndroidXNNPACK, Vulkan, Qualcomm, MediaTek, Samsung Exynos
iOSXNNPACK, CoreML (Neural Engine)
Linux / WindowsXNNPACK, OpenVINO, CUDA (experimental)
macOSXNNPACK, Metal (experimental), MLX (experimental)
Embedded / MCUXNNPACK, ARM Ethos-U, NXP, Cadence DSP

See Backend Documentation for detailed hardware requirements and optimization guides. For desktop/laptop GPU inference with CUDA and Metal, see the Desktop Guide. For Zephyr RTOS integration, see the Zephyr Guide.

Production Deployments

ExecuTorch powers on-device AI at scale across Meta's family of apps, VR/AR devices, and partner deployments. View success stories →

Examples & Models

LLMs:Llama 3.2/3.1/3, Qwen 3, Phi-4-mini, LiquidAI LFM2

Multimodal:Llava (vision-language), Voxtral (audio-language), Gemma (vision-language)

Vision/Speech:MobileNetV2, DeepLabV3, YOLO26, Whisper, Supertonic

Resources:examples/ directory • executorch-examples out-of-tree demos • Optimum-ExecuTorch for HuggingFace models • Unsloth for fine-tuned LLM deployment

Key Features

ExecuTorch provides advanced capabilities for production deployment:

  • Quantization — Built-in support via torchao for 8-bit, 4-bit, and dynamic quantization
  • Memory Planning — Optimize memory usage with ahead-of-time allocation strategies
  • Developer Tools — ETDump profiler, ETRecord inspector, and model debugger
  • Selective Build — Strip unused operators to minimize binary size
  • Custom Operators — Extend with domain-specific kernels
  • Dynamic Shapes — Support variable input sizes with bounded ranges

See Advanced Topics for quantization techniques, custom backends, and compiler passes.

Documentation

Community & Contributing

We welcome contributions from the community!

Citing ExecuTorch

If you found ExecuTorch helpful in your research and would like to acknowledge it, please cite us using the following BibTeX:

@article{executorch2026,
title={{ExecuTorch} - A Unified {PyTorch} Solution to Run {AI} Models On-Device},
author={Nachin, Mergen and Desai, Digant and Jia, Sicheng Stephen and Lai, Chen and Liu, Mengwei and Szwejbka, Jacob and Alvarez, Raziel and Ascani, RJ and Bort, Dave and Candales, Manuel and others},
journal={arXiv preprint arXiv:2605.08195},
url={https://github.com/pytorch/executorch},
year={2026}
}

License

ExecuTorch is BSD licensed, as found in the LICENSE file.




Part of the PyTorch ecosystem

GitHubDocumentation

About

End-to-end solution for enabling on-device AI across mobile and edge devices for PyTorch models

Resources

Code of conduct

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - oscarandersson8218/executorch: End-to-end solution for enabling on-device AI across mobile and edge devices for PyTorch models · GitHub
Skip to content

Latest commit

History

13,444 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

ExecuTorch logo mark

ExecuTorch

On-device AI inference powered by PyTorch

PyPI - VersionGitHub - ContributorsGitHub - StarsDiscord - Chat with UsDocumentation

ExecuTorch is PyTorch's unified solution for deploying AI models on-device—from smartphones to microcontrollers—built for privacy, performance, and portability. It powers Meta's on-device AI across Instagram, WhatsApp, Quest 3, Ray-Ban Meta Smart Glasses, and more.

Deploy LLMs, vision, speech, and multimodal models with the same PyTorch APIs you already know—accelerating research to production with seamless model export, optimization, and deployment. No manual C++ rewrites. No format conversions. No vendor lock-in.

📘 Table of Contents

Why ExecuTorch?

  • 🔒 Native PyTorch Export — Direct export from PyTorch. No .onnx, .tflite, or intermediate format conversions. Preserve model semantics.
  • ⚡ Production-Proven — Powers billions of users at Meta with real-time on-device inference.
  • 💾 Tiny Runtime — 50KB base footprint. Runs on microcontrollers to high-end smartphones.
  • 🚀 12+ Hardware Backends — Open-source acceleration for Apple, Samsung, Qualcomm, ARM, MediaTek, Vulkan, and more.
  • 🎯 One Export, Multiple Backends — Switch hardware targets with a single line change. Deploy the same model everywhere.

How It Works

ExecuTorch uses ahead-of-time (AOT) compilation to prepare PyTorch models for edge deployment:

  1. 🧩 Export — Capture your PyTorch model graph with torch.export()
  2. ⚙️ Compile — Quantize, optimize, and partition to hardware backends → .pte
  3. 🚀 Execute — Load .pte on-device via lightweight C++ runtime

Models use a standardized Core ATen operator set. Partitioners delegate subgraphs to specialized hardware (NPU/GPU) with CPU fallback.

Learn more: How ExecuTorch WorksArchitecture Guide

Quick Start

Installation

pip install executorch

For platform-specific setup (Android, iOS, embedded systems), see the Quick Start documentation for additional info.

Export and Deploy in 3 Steps

importtorchfromexecutorch.exirimportto_edge_transform_and_lowerfromexecutorch.backends.xnnpack.partition.xnnpack_partitionerimportXnnpackPartitioner# 1. Export your PyTorch modelmodel=MyModel().eval()
example_inputs= (torch.randn(1, 3, 224, 224),)
exported_program=torch.export.export(model, example_inputs)
# 2. Optimize for target hardware (switch backends with one line)program=to_edge_transform_and_lower(
exported_program,
partitioner=[XnnpackPartitioner()] # CPU | CoreMLPartitioner() for iOS | QnnPartitioner() for Qualcomm
).to_executorch()
# 3. Save for deploymentwithopen("model.pte", "wb") asf:
f.write(program.buffer)
# Test locally via ExecuTorch runtime's pybind API (optional)fromexecutorch.runtimeimportRuntimeruntime=Runtime.get()
method=runtime.load_program("model.pte").load_method("forward")
outputs=method.execute([torch.randn(1, 3, 224, 224)])

Run on Device

C++

#include<executorch/extension/module/module.h>
#include<executorch/extension/tensor/tensor.h>
Module module("model.pte");
auto tensor = make_tensor_ptr({2, 2}, {1.0f, 2.0f, 3.0f, 4.0f});
auto outputs = module.forward(tensor);

Swift (iOS)

import ExecuTorch
letmodule=Module(filePath:"model.pte")letinput=Tensor<Float>([1.0,2.0,3.0,4.0], shape:[2,2])letoutputs=try module.forward(input)

Kotlin (Android)

val module =Module.load("model.pte")
val inputTensor =Tensor.fromBlob(floatArrayOf(1.0f, 2.0f, 3.0f, 4.0f), longArrayOf(2, 2))
val outputs = module.forward(EValue.from(inputTensor))

LLM Example: Llama

Export Llama models using the export_llm script or Optimum-ExecuTorch:

# Using export_llm
python -m executorch.extension.llm.export.export_llm --model llama3_2 --output llama.pte
# Using Optimum-ExecuTorch
optimum-cli export executorch \
--model meta-llama/Llama-3.2-1B \
--task text-generation \
--recipe xnnpack \
--output_dir llama_model

Run on-device with the LLM runner API:

C++

#include<executorch/extension/llm/runner/text_llm_runner.h>auto runner = create_llama_runner("llama.pte", "tiktoken.bin");
executorch::extension::llm::GenerationConfig config{
.seq_len = 128, .temperature = 0.8f};
runner->generate("Hello, how are you?", config);

Swift (iOS)

import ExecuTorchLLM
letrunner=TextRunner(modelPath:"llama.pte", tokenizerPath:"tiktoken.bin")try runner.generate("Hello, how are you?",Config{
$0.sequenceLength =128}){ token inprint(token, terminator:"")}

Kotlin (Android)API DocsDemo App

val llmModule =LlmModule("llama.pte", "tiktoken.bin", 0.8f)
llmModule.load()
llmModule.generate("Hello, how are you?", 128, object:LlmCallback {
overridefunonResult(result:String) { print(result) }
overridefunonStats(stats:String) { }
})

For multimodal models (vision, audio), use the MultiModal runner API which extends the LLM runner to handle image and audio inputs alongside text. See Llava and Voxtral examples.

See examples/models/llama for complete workflow including quantization, mobile deployment, and advanced options.

Next Steps:

Platform & Hardware Support

PlatformSupported Backends
AndroidXNNPACK, Vulkan, Qualcomm, MediaTek, Samsung Exynos
iOSXNNPACK, CoreML (Neural Engine)
Linux / WindowsXNNPACK, OpenVINO, CUDA (experimental)
macOSXNNPACK, Metal (experimental), MLX (experimental)
Embedded / MCUXNNPACK, ARM Ethos-U, NXP, Cadence DSP

See Backend Documentation for detailed hardware requirements and optimization guides. For desktop/laptop GPU inference with CUDA and Metal, see the Desktop Guide. For Zephyr RTOS integration, see the Zephyr Guide.

Production Deployments

ExecuTorch powers on-device AI at scale across Meta's family of apps, VR/AR devices, and partner deployments. View success stories →

Examples & Models

LLMs:Llama 3.2/3.1/3, Qwen 3, Phi-4-mini, LiquidAI LFM2

Multimodal:Llava (vision-language), Voxtral (audio-language), Gemma (vision-language)

Vision/Speech:MobileNetV2, DeepLabV3, YOLO26, Whisper, Supertonic

Resources:examples/ directory • executorch-examples out-of-tree demos • Optimum-ExecuTorch for HuggingFace models • Unsloth for fine-tuned LLM deployment

Key Features

ExecuTorch provides advanced capabilities for production deployment:

  • Quantization — Built-in support via torchao for 8-bit, 4-bit, and dynamic quantization
  • Memory Planning — Optimize memory usage with ahead-of-time allocation strategies
  • Developer Tools — ETDump profiler, ETRecord inspector, and model debugger
  • Selective Build — Strip unused operators to minimize binary size
  • Custom Operators — Extend with domain-specific kernels
  • Dynamic Shapes — Support variable input sizes with bounded ranges

See Advanced Topics for quantization techniques, custom backends, and compiler passes.

Documentation

Community & Contributing

We welcome contributions from the community!

Citing ExecuTorch

If you found ExecuTorch helpful in your research and would like to acknowledge it, please cite us using the following BibTeX:

@article{executorch2026,
title={{ExecuTorch} - A Unified {PyTorch} Solution to Run {AI} Models On-Device},
author={Nachin, Mergen and Desai, Digant and Jia, Sicheng Stephen and Lai, Chen and Liu, Mengwei and Szwejbka, Jacob and Alvarez, Raziel and Ascani, RJ and Bort, Dave and Candales, Manuel and others},
journal={arXiv preprint arXiv:2605.08195},
url={https://github.com/pytorch/executorch},
year={2026}
}

License

ExecuTorch is BSD licensed, as found in the LICENSE file.




Part of the PyTorch ecosystem

GitHubDocumentation

About

End-to-end solution for enabling on-device AI across mobile and edge devices for PyTorch models

Resources

Code of conduct

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - oscarandersson8218/executorch: End-to-end solution for enabling on-device AI across mobile and edge devices for PyTorch models · GitHub
Skip to content

Latest commit

History

13,444 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

ExecuTorch logo mark

ExecuTorch

On-device AI inference powered by PyTorch

PyPI - VersionGitHub - ContributorsGitHub - StarsDiscord - Chat with UsDocumentation

ExecuTorch is PyTorch's unified solution for deploying AI models on-device—from smartphones to microcontrollers—built for privacy, performance, and portability. It powers Meta's on-device AI across Instagram, WhatsApp, Quest 3, Ray-Ban Meta Smart Glasses, and more.

Deploy LLMs, vision, speech, and multimodal models with the same PyTorch APIs you already know—accelerating research to production with seamless model export, optimization, and deployment. No manual C++ rewrites. No format conversions. No vendor lock-in.

📘 Table of Contents

Why ExecuTorch?

  • 🔒 Native PyTorch Export — Direct export from PyTorch. No .onnx, .tflite, or intermediate format conversions. Preserve model semantics.
  • ⚡ Production-Proven — Powers billions of users at Meta with real-time on-device inference.
  • 💾 Tiny Runtime — 50KB base footprint. Runs on microcontrollers to high-end smartphones.
  • 🚀 12+ Hardware Backends — Open-source acceleration for Apple, Samsung, Qualcomm, ARM, MediaTek, Vulkan, and more.
  • 🎯 One Export, Multiple Backends — Switch hardware targets with a single line change. Deploy the same model everywhere.

How It Works

ExecuTorch uses ahead-of-time (AOT) compilation to prepare PyTorch models for edge deployment:

  1. 🧩 Export — Capture your PyTorch model graph with torch.export()
  2. ⚙️ Compile — Quantize, optimize, and partition to hardware backends → .pte
  3. 🚀 Execute — Load .pte on-device via lightweight C++ runtime

Models use a standardized Core ATen operator set. Partitioners delegate subgraphs to specialized hardware (NPU/GPU) with CPU fallback.

Learn more: How ExecuTorch WorksArchitecture Guide

Quick Start

Installation

pip install executorch

For platform-specific setup (Android, iOS, embedded systems), see the Quick Start documentation for additional info.

Export and Deploy in 3 Steps

importtorchfromexecutorch.exirimportto_edge_transform_and_lowerfromexecutorch.backends.xnnpack.partition.xnnpack_partitionerimportXnnpackPartitioner# 1. Export your PyTorch modelmodel=MyModel().eval()
example_inputs= (torch.randn(1, 3, 224, 224),)
exported_program=torch.export.export(model, example_inputs)
# 2. Optimize for target hardware (switch backends with one line)program=to_edge_transform_and_lower(
exported_program,
partitioner=[XnnpackPartitioner()] # CPU | CoreMLPartitioner() for iOS | QnnPartitioner() for Qualcomm
).to_executorch()
# 3. Save for deploymentwithopen("model.pte", "wb") asf:
f.write(program.buffer)
# Test locally via ExecuTorch runtime's pybind API (optional)fromexecutorch.runtimeimportRuntimeruntime=Runtime.get()
method=runtime.load_program("model.pte").load_method("forward")
outputs=method.execute([torch.randn(1, 3, 224, 224)])

Run on Device

C++

#include<executorch/extension/module/module.h>
#include<executorch/extension/tensor/tensor.h>
Module module("model.pte");
auto tensor = make_tensor_ptr({2, 2}, {1.0f, 2.0f, 3.0f, 4.0f});
auto outputs = module.forward(tensor);

Swift (iOS)

import ExecuTorch
letmodule=Module(filePath:"model.pte")letinput=Tensor<Float>([1.0,2.0,3.0,4.0], shape:[2,2])letoutputs=try module.forward(input)

Kotlin (Android)

val module =Module.load("model.pte")
val inputTensor =Tensor.fromBlob(floatArrayOf(1.0f, 2.0f, 3.0f, 4.0f), longArrayOf(2, 2))
val outputs = module.forward(EValue.from(inputTensor))

LLM Example: Llama

Export Llama models using the export_llm script or Optimum-ExecuTorch:

# Using export_llm
python -m executorch.extension.llm.export.export_llm --model llama3_2 --output llama.pte
# Using Optimum-ExecuTorch
optimum-cli export executorch \
--model meta-llama/Llama-3.2-1B \
--task text-generation \
--recipe xnnpack \
--output_dir llama_model

Run on-device with the LLM runner API:

C++

#include<executorch/extension/llm/runner/text_llm_runner.h>auto runner = create_llama_runner("llama.pte", "tiktoken.bin");
executorch::extension::llm::GenerationConfig config{
.seq_len = 128, .temperature = 0.8f};
runner->generate("Hello, how are you?", config);

Swift (iOS)

import ExecuTorchLLM
letrunner=TextRunner(modelPath:"llama.pte", tokenizerPath:"tiktoken.bin")try runner.generate("Hello, how are you?",Config{
$0.sequenceLength =128}){ token inprint(token, terminator:"")}

Kotlin (Android)API DocsDemo App

val llmModule =LlmModule("llama.pte", "tiktoken.bin", 0.8f)
llmModule.load()
llmModule.generate("Hello, how are you?", 128, object:LlmCallback {
overridefunonResult(result:String) { print(result) }
overridefunonStats(stats:String) { }
})

For multimodal models (vision, audio), use the MultiModal runner API which extends the LLM runner to handle image and audio inputs alongside text. See Llava and Voxtral examples.

See examples/models/llama for complete workflow including quantization, mobile deployment, and advanced options.

Next Steps:

Platform & Hardware Support

PlatformSupported Backends
AndroidXNNPACK, Vulkan, Qualcomm, MediaTek, Samsung Exynos
iOSXNNPACK, CoreML (Neural Engine)
Linux / WindowsXNNPACK, OpenVINO, CUDA (experimental)
macOSXNNPACK, Metal (experimental), MLX (experimental)
Embedded / MCUXNNPACK, ARM Ethos-U, NXP, Cadence DSP

See Backend Documentation for detailed hardware requirements and optimization guides. For desktop/laptop GPU inference with CUDA and Metal, see the Desktop Guide. For Zephyr RTOS integration, see the Zephyr Guide.

Production Deployments

ExecuTorch powers on-device AI at scale across Meta's family of apps, VR/AR devices, and partner deployments. View success stories →

Examples & Models

LLMs:Llama 3.2/3.1/3, Qwen 3, Phi-4-mini, LiquidAI LFM2

Multimodal:Llava (vision-language), Voxtral (audio-language), Gemma (vision-language)

Vision/Speech:MobileNetV2, DeepLabV3, YOLO26, Whisper, Supertonic

Resources:examples/ directory • executorch-examples out-of-tree demos • Optimum-ExecuTorch for HuggingFace models • Unsloth for fine-tuned LLM deployment

Key Features

ExecuTorch provides advanced capabilities for production deployment:

  • Quantization — Built-in support via torchao for 8-bit, 4-bit, and dynamic quantization
  • Memory Planning — Optimize memory usage with ahead-of-time allocation strategies
  • Developer Tools — ETDump profiler, ETRecord inspector, and model debugger
  • Selective Build — Strip unused operators to minimize binary size
  • Custom Operators — Extend with domain-specific kernels
  • Dynamic Shapes — Support variable input sizes with bounded ranges

See Advanced Topics for quantization techniques, custom backends, and compiler passes.

Documentation

Community & Contributing

We welcome contributions from the community!

Citing ExecuTorch

If you found ExecuTorch helpful in your research and would like to acknowledge it, please cite us using the following BibTeX:

@article{executorch2026,
title={{ExecuTorch} - A Unified {PyTorch} Solution to Run {AI} Models On-Device},
author={Nachin, Mergen and Desai, Digant and Jia, Sicheng Stephen and Lai, Chen and Liu, Mengwei and Szwejbka, Jacob and Alvarez, Raziel and Ascani, RJ and Bort, Dave and Candales, Manuel and others},
journal={arXiv preprint arXiv:2605.08195},
url={https://github.com/pytorch/executorch},
year={2026}
}

License

ExecuTorch is BSD licensed, as found in the LICENSE file.




Part of the PyTorch ecosystem

GitHubDocumentation

About

End-to-end solution for enabling on-device AI across mobile and edge devices for PyTorch models

Resources

Code of conduct

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - oscarandersson8218/executorch: End-to-end solution for enabling on-device AI across mobile and edge devices for PyTorch models · GitHub
Skip to content

Latest commit

History

13,444 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

ExecuTorch logo mark

ExecuTorch

On-device AI inference powered by PyTorch

PyPI - VersionGitHub - ContributorsGitHub - StarsDiscord - Chat with UsDocumentation

ExecuTorch is PyTorch's unified solution for deploying AI models on-device—from smartphones to microcontrollers—built for privacy, performance, and portability. It powers Meta's on-device AI across Instagram, WhatsApp, Quest 3, Ray-Ban Meta Smart Glasses, and more.

Deploy LLMs, vision, speech, and multimodal models with the same PyTorch APIs you already know—accelerating research to production with seamless model export, optimization, and deployment. No manual C++ rewrites. No format conversions. No vendor lock-in.

📘 Table of Contents

Why ExecuTorch?

  • 🔒 Native PyTorch Export — Direct export from PyTorch. No .onnx, .tflite, or intermediate format conversions. Preserve model semantics.
  • ⚡ Production-Proven — Powers billions of users at Meta with real-time on-device inference.
  • 💾 Tiny Runtime — 50KB base footprint. Runs on microcontrollers to high-end smartphones.
  • 🚀 12+ Hardware Backends — Open-source acceleration for Apple, Samsung, Qualcomm, ARM, MediaTek, Vulkan, and more.
  • 🎯 One Export, Multiple Backends — Switch hardware targets with a single line change. Deploy the same model everywhere.

How It Works

ExecuTorch uses ahead-of-time (AOT) compilation to prepare PyTorch models for edge deployment:

  1. 🧩 Export — Capture your PyTorch model graph with torch.export()
  2. ⚙️ Compile — Quantize, optimize, and partition to hardware backends → .pte
  3. 🚀 Execute — Load .pte on-device via lightweight C++ runtime

Models use a standardized Core ATen operator set. Partitioners delegate subgraphs to specialized hardware (NPU/GPU) with CPU fallback.

Learn more: How ExecuTorch WorksArchitecture Guide

Quick Start

Installation

pip install executorch

For platform-specific setup (Android, iOS, embedded systems), see the Quick Start documentation for additional info.

Export and Deploy in 3 Steps

importtorchfromexecutorch.exirimportto_edge_transform_and_lowerfromexecutorch.backends.xnnpack.partition.xnnpack_partitionerimportXnnpackPartitioner# 1. Export your PyTorch modelmodel=MyModel().eval()
example_inputs= (torch.randn(1, 3, 224, 224),)
exported_program=torch.export.export(model, example_inputs)
# 2. Optimize for target hardware (switch backends with one line)program=to_edge_transform_and_lower(
exported_program,
partitioner=[XnnpackPartitioner()] # CPU | CoreMLPartitioner() for iOS | QnnPartitioner() for Qualcomm
).to_executorch()
# 3. Save for deploymentwithopen("model.pte", "wb") asf:
f.write(program.buffer)
# Test locally via ExecuTorch runtime's pybind API (optional)fromexecutorch.runtimeimportRuntimeruntime=Runtime.get()
method=runtime.load_program("model.pte").load_method("forward")
outputs=method.execute([torch.randn(1, 3, 224, 224)])

Run on Device

C++

#include<executorch/extension/module/module.h>
#include<executorch/extension/tensor/tensor.h>
Module module("model.pte");
auto tensor = make_tensor_ptr({2, 2}, {1.0f, 2.0f, 3.0f, 4.0f});
auto outputs = module.forward(tensor);

Swift (iOS)

import ExecuTorch
letmodule=Module(filePath:"model.pte")letinput=Tensor<Float>([1.0,2.0,3.0,4.0], shape:[2,2])letoutputs=try module.forward(input)

Kotlin (Android)

val module =Module.load("model.pte")
val inputTensor =Tensor.fromBlob(floatArrayOf(1.0f, 2.0f, 3.0f, 4.0f), longArrayOf(2, 2))
val outputs = module.forward(EValue.from(inputTensor))

LLM Example: Llama

Export Llama models using the export_llm script or Optimum-ExecuTorch:

# Using export_llm
python -m executorch.extension.llm.export.export_llm --model llama3_2 --output llama.pte
# Using Optimum-ExecuTorch
optimum-cli export executorch \
--model meta-llama/Llama-3.2-1B \
--task text-generation \
--recipe xnnpack \
--output_dir llama_model

Run on-device with the LLM runner API:

C++

#include<executorch/extension/llm/runner/text_llm_runner.h>auto runner = create_llama_runner("llama.pte", "tiktoken.bin");
executorch::extension::llm::GenerationConfig config{
.seq_len = 128, .temperature = 0.8f};
runner->generate("Hello, how are you?", config);

Swift (iOS)

import ExecuTorchLLM
letrunner=TextRunner(modelPath:"llama.pte", tokenizerPath:"tiktoken.bin")try runner.generate("Hello, how are you?",Config{
$0.sequenceLength =128}){ token inprint(token, terminator:"")}

Kotlin (Android)API DocsDemo App

val llmModule =LlmModule("llama.pte", "tiktoken.bin", 0.8f)
llmModule.load()
llmModule.generate("Hello, how are you?", 128, object:LlmCallback {
overridefunonResult(result:String) { print(result) }
overridefunonStats(stats:String) { }
})

For multimodal models (vision, audio), use the MultiModal runner API which extends the LLM runner to handle image and audio inputs alongside text. See Llava and Voxtral examples.

See examples/models/llama for complete workflow including quantization, mobile deployment, and advanced options.

Next Steps:

Platform & Hardware Support

PlatformSupported Backends
AndroidXNNPACK, Vulkan, Qualcomm, MediaTek, Samsung Exynos
iOSXNNPACK, CoreML (Neural Engine)
Linux / WindowsXNNPACK, OpenVINO, CUDA (experimental)
macOSXNNPACK, Metal (experimental), MLX (experimental)
Embedded / MCUXNNPACK, ARM Ethos-U, NXP, Cadence DSP

See Backend Documentation for detailed hardware requirements and optimization guides. For desktop/laptop GPU inference with CUDA and Metal, see the Desktop Guide. For Zephyr RTOS integration, see the Zephyr Guide.

Production Deployments

ExecuTorch powers on-device AI at scale across Meta's family of apps, VR/AR devices, and partner deployments. View success stories →

Examples & Models

LLMs:Llama 3.2/3.1/3, Qwen 3, Phi-4-mini, LiquidAI LFM2

Multimodal:Llava (vision-language), Voxtral (audio-language), Gemma (vision-language)

Vision/Speech:MobileNetV2, DeepLabV3, YOLO26, Whisper, Supertonic

Resources:examples/ directory • executorch-examples out-of-tree demos • Optimum-ExecuTorch for HuggingFace models • Unsloth for fine-tuned LLM deployment

Key Features

ExecuTorch provides advanced capabilities for production deployment:

  • Quantization — Built-in support via torchao for 8-bit, 4-bit, and dynamic quantization
  • Memory Planning — Optimize memory usage with ahead-of-time allocation strategies
  • Developer Tools — ETDump profiler, ETRecord inspector, and model debugger
  • Selective Build — Strip unused operators to minimize binary size
  • Custom Operators — Extend with domain-specific kernels
  • Dynamic Shapes — Support variable input sizes with bounded ranges

See Advanced Topics for quantization techniques, custom backends, and compiler passes.

Documentation

Community & Contributing

We welcome contributions from the community!

Citing ExecuTorch

If you found ExecuTorch helpful in your research and would like to acknowledge it, please cite us using the following BibTeX:

@article{executorch2026,
title={{ExecuTorch} - A Unified {PyTorch} Solution to Run {AI} Models On-Device},
author={Nachin, Mergen and Desai, Digant and Jia, Sicheng Stephen and Lai, Chen and Liu, Mengwei and Szwejbka, Jacob and Alvarez, Raziel and Ascani, RJ and Bort, Dave and Candales, Manuel and others},
journal={arXiv preprint arXiv:2605.08195},
url={https://github.com/pytorch/executorch},
year={2026}
}

License

ExecuTorch is BSD licensed, as found in the LICENSE file.




Part of the PyTorch ecosystem

GitHubDocumentation

About

End-to-end solution for enabling on-device AI across mobile and edge devices for PyTorch models

Resources

Code of conduct

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - oscarandersson8218/executorch: End-to-end solution for enabling on-device AI across mobile and edge devices for PyTorch models · GitHub
Skip to content

Latest commit

History

13,444 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

ExecuTorch logo mark

ExecuTorch

On-device AI inference powered by PyTorch

PyPI - VersionGitHub - ContributorsGitHub - StarsDiscord - Chat with UsDocumentation

ExecuTorch is PyTorch's unified solution for deploying AI models on-device—from smartphones to microcontrollers—built for privacy, performance, and portability. It powers Meta's on-device AI across Instagram, WhatsApp, Quest 3, Ray-Ban Meta Smart Glasses, and more.

Deploy LLMs, vision, speech, and multimodal models with the same PyTorch APIs you already know—accelerating research to production with seamless model export, optimization, and deployment. No manual C++ rewrites. No format conversions. No vendor lock-in.

📘 Table of Contents

Why ExecuTorch?

  • 🔒 Native PyTorch Export — Direct export from PyTorch. No .onnx, .tflite, or intermediate format conversions. Preserve model semantics.
  • ⚡ Production-Proven — Powers billions of users at Meta with real-time on-device inference.
  • 💾 Tiny Runtime — 50KB base footprint. Runs on microcontrollers to high-end smartphones.
  • 🚀 12+ Hardware Backends — Open-source acceleration for Apple, Samsung, Qualcomm, ARM, MediaTek, Vulkan, and more.
  • 🎯 One Export, Multiple Backends — Switch hardware targets with a single line change. Deploy the same model everywhere.

How It Works

ExecuTorch uses ahead-of-time (AOT) compilation to prepare PyTorch models for edge deployment:

  1. 🧩 Export — Capture your PyTorch model graph with torch.export()
  2. ⚙️ Compile — Quantize, optimize, and partition to hardware backends → .pte
  3. 🚀 Execute — Load .pte on-device via lightweight C++ runtime

Models use a standardized Core ATen operator set. Partitioners delegate subgraphs to specialized hardware (NPU/GPU) with CPU fallback.

Learn more: How ExecuTorch WorksArchitecture Guide

Quick Start

Installation

pip install executorch

For platform-specific setup (Android, iOS, embedded systems), see the Quick Start documentation for additional info.

Export and Deploy in 3 Steps

importtorchfromexecutorch.exirimportto_edge_transform_and_lowerfromexecutorch.backends.xnnpack.partition.xnnpack_partitionerimportXnnpackPartitioner# 1. Export your PyTorch modelmodel=MyModel().eval()
example_inputs= (torch.randn(1, 3, 224, 224),)
exported_program=torch.export.export(model, example_inputs)
# 2. Optimize for target hardware (switch backends with one line)program=to_edge_transform_and_lower(
exported_program,
partitioner=[XnnpackPartitioner()] # CPU | CoreMLPartitioner() for iOS | QnnPartitioner() for Qualcomm
).to_executorch()
# 3. Save for deploymentwithopen("model.pte", "wb") asf:
f.write(program.buffer)
# Test locally via ExecuTorch runtime's pybind API (optional)fromexecutorch.runtimeimportRuntimeruntime=Runtime.get()
method=runtime.load_program("model.pte").load_method("forward")
outputs=method.execute([torch.randn(1, 3, 224, 224)])

Run on Device

C++

#include<executorch/extension/module/module.h>
#include<executorch/extension/tensor/tensor.h>
Module module("model.pte");
auto tensor = make_tensor_ptr({2, 2}, {1.0f, 2.0f, 3.0f, 4.0f});
auto outputs = module.forward(tensor);

Swift (iOS)

import ExecuTorch
letmodule=Module(filePath:"model.pte")letinput=Tensor<Float>([1.0,2.0,3.0,4.0], shape:[2,2])letoutputs=try module.forward(input)

Kotlin (Android)

val module =Module.load("model.pte")
val inputTensor =Tensor.fromBlob(floatArrayOf(1.0f, 2.0f, 3.0f, 4.0f), longArrayOf(2, 2))
val outputs = module.forward(EValue.from(inputTensor))

LLM Example: Llama

Export Llama models using the export_llm script or Optimum-ExecuTorch:

# Using export_llm
python -m executorch.extension.llm.export.export_llm --model llama3_2 --output llama.pte
# Using Optimum-ExecuTorch
optimum-cli export executorch \
--model meta-llama/Llama-3.2-1B \
--task text-generation \
--recipe xnnpack \
--output_dir llama_model

Run on-device with the LLM runner API:

C++

#include<executorch/extension/llm/runner/text_llm_runner.h>auto runner = create_llama_runner("llama.pte", "tiktoken.bin");
executorch::extension::llm::GenerationConfig config{
.seq_len = 128, .temperature = 0.8f};
runner->generate("Hello, how are you?", config);

Swift (iOS)

import ExecuTorchLLM
letrunner=TextRunner(modelPath:"llama.pte", tokenizerPath:"tiktoken.bin")try runner.generate("Hello, how are you?",Config{
$0.sequenceLength =128}){ token inprint(token, terminator:"")}

Kotlin (Android)API DocsDemo App

val llmModule =LlmModule("llama.pte", "tiktoken.bin", 0.8f)
llmModule.load()
llmModule.generate("Hello, how are you?", 128, object:LlmCallback {
overridefunonResult(result:String) { print(result) }
overridefunonStats(stats:String) { }
})

For multimodal models (vision, audio), use the MultiModal runner API which extends the LLM runner to handle image and audio inputs alongside text. See Llava and Voxtral examples.

See examples/models/llama for complete workflow including quantization, mobile deployment, and advanced options.

Next Steps:

Platform & Hardware Support

PlatformSupported Backends
AndroidXNNPACK, Vulkan, Qualcomm, MediaTek, Samsung Exynos
iOSXNNPACK, CoreML (Neural Engine)
Linux / WindowsXNNPACK, OpenVINO, CUDA (experimental)
macOSXNNPACK, Metal (experimental), MLX (experimental)
Embedded / MCUXNNPACK, ARM Ethos-U, NXP, Cadence DSP

See Backend Documentation for detailed hardware requirements and optimization guides. For desktop/laptop GPU inference with CUDA and Metal, see the Desktop Guide. For Zephyr RTOS integration, see the Zephyr Guide.

Production Deployments

ExecuTorch powers on-device AI at scale across Meta's family of apps, VR/AR devices, and partner deployments. View success stories →

Examples & Models

LLMs:Llama 3.2/3.1/3, Qwen 3, Phi-4-mini, LiquidAI LFM2

Multimodal:Llava (vision-language), Voxtral (audio-language), Gemma (vision-language)

Vision/Speech:MobileNetV2, DeepLabV3, YOLO26, Whisper, Supertonic

Resources:examples/ directory • executorch-examples out-of-tree demos • Optimum-ExecuTorch for HuggingFace models • Unsloth for fine-tuned LLM deployment

Key Features

ExecuTorch provides advanced capabilities for production deployment:

  • Quantization — Built-in support via torchao for 8-bit, 4-bit, and dynamic quantization
  • Memory Planning — Optimize memory usage with ahead-of-time allocation strategies
  • Developer Tools — ETDump profiler, ETRecord inspector, and model debugger
  • Selective Build — Strip unused operators to minimize binary size
  • Custom Operators — Extend with domain-specific kernels
  • Dynamic Shapes — Support variable input sizes with bounded ranges

See Advanced Topics for quantization techniques, custom backends, and compiler passes.

Documentation

Community & Contributing

We welcome contributions from the community!

Citing ExecuTorch

If you found ExecuTorch helpful in your research and would like to acknowledge it, please cite us using the following BibTeX:

@article{executorch2026,
title={{ExecuTorch} - A Unified {PyTorch} Solution to Run {AI} Models On-Device},
author={Nachin, Mergen and Desai, Digant and Jia, Sicheng Stephen and Lai, Chen and Liu, Mengwei and Szwejbka, Jacob and Alvarez, Raziel and Ascani, RJ and Bort, Dave and Candales, Manuel and others},
journal={arXiv preprint arXiv:2605.08195},
url={https://github.com/pytorch/executorch},
year={2026}
}

License

ExecuTorch is BSD licensed, as found in the LICENSE file.




Part of the PyTorch ecosystem

GitHubDocumentation

About

End-to-end solution for enabling on-device AI across mobile and edge devices for PyTorch models

Resources

Code of conduct

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - oscarandersson8218/executorch: End-to-end solution for enabling on-device AI across mobile and edge devices for PyTorch models · GitHub
Skip to content

Latest commit

History

13,444 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

ExecuTorch logo mark

ExecuTorch

On-device AI inference powered by PyTorch

PyPI - VersionGitHub - ContributorsGitHub - StarsDiscord - Chat with UsDocumentation

ExecuTorch is PyTorch's unified solution for deploying AI models on-device—from smartphones to microcontrollers—built for privacy, performance, and portability. It powers Meta's on-device AI across Instagram, WhatsApp, Quest 3, Ray-Ban Meta Smart Glasses, and more.

Deploy LLMs, vision, speech, and multimodal models with the same PyTorch APIs you already know—accelerating research to production with seamless model export, optimization, and deployment. No manual C++ rewrites. No format conversions. No vendor lock-in.

📘 Table of Contents

Why ExecuTorch?

  • 🔒 Native PyTorch Export — Direct export from PyTorch. No .onnx, .tflite, or intermediate format conversions. Preserve model semantics.
  • ⚡ Production-Proven — Powers billions of users at Meta with real-time on-device inference.
  • 💾 Tiny Runtime — 50KB base footprint. Runs on microcontrollers to high-end smartphones.
  • 🚀 12+ Hardware Backends — Open-source acceleration for Apple, Samsung, Qualcomm, ARM, MediaTek, Vulkan, and more.
  • 🎯 One Export, Multiple Backends — Switch hardware targets with a single line change. Deploy the same model everywhere.

How It Works

ExecuTorch uses ahead-of-time (AOT) compilation to prepare PyTorch models for edge deployment:

  1. 🧩 Export — Capture your PyTorch model graph with torch.export()
  2. ⚙️ Compile — Quantize, optimize, and partition to hardware backends → .pte
  3. 🚀 Execute — Load .pte on-device via lightweight C++ runtime

Models use a standardized Core ATen operator set. Partitioners delegate subgraphs to specialized hardware (NPU/GPU) with CPU fallback.

Learn more: How ExecuTorch WorksArchitecture Guide

Quick Start

Installation

pip install executorch

For platform-specific setup (Android, iOS, embedded systems), see the Quick Start documentation for additional info.

Export and Deploy in 3 Steps

importtorchfromexecutorch.exirimportto_edge_transform_and_lowerfromexecutorch.backends.xnnpack.partition.xnnpack_partitionerimportXnnpackPartitioner# 1. Export your PyTorch modelmodel=MyModel().eval()
example_inputs= (torch.randn(1, 3, 224, 224),)
exported_program=torch.export.export(model, example_inputs)
# 2. Optimize for target hardware (switch backends with one line)program=to_edge_transform_and_lower(
exported_program,
partitioner=[XnnpackPartitioner()] # CPU | CoreMLPartitioner() for iOS | QnnPartitioner() for Qualcomm
).to_executorch()
# 3. Save for deploymentwithopen("model.pte", "wb") asf:
f.write(program.buffer)
# Test locally via ExecuTorch runtime's pybind API (optional)fromexecutorch.runtimeimportRuntimeruntime=Runtime.get()
method=runtime.load_program("model.pte").load_method("forward")
outputs=method.execute([torch.randn(1, 3, 224, 224)])

Run on Device

C++

#include<executorch/extension/module/module.h>
#include<executorch/extension/tensor/tensor.h>
Module module("model.pte");
auto tensor = make_tensor_ptr({2, 2}, {1.0f, 2.0f, 3.0f, 4.0f});
auto outputs = module.forward(tensor);

Swift (iOS)

import ExecuTorch
letmodule=Module(filePath:"model.pte")letinput=Tensor<Float>([1.0,2.0,3.0,4.0], shape:[2,2])letoutputs=try module.forward(input)

Kotlin (Android)

val module =Module.load("model.pte")
val inputTensor =Tensor.fromBlob(floatArrayOf(1.0f, 2.0f, 3.0f, 4.0f), longArrayOf(2, 2))
val outputs = module.forward(EValue.from(inputTensor))

LLM Example: Llama

Export Llama models using the export_llm script or Optimum-ExecuTorch:

# Using export_llm
python -m executorch.extension.llm.export.export_llm --model llama3_2 --output llama.pte
# Using Optimum-ExecuTorch
optimum-cli export executorch \
--model meta-llama/Llama-3.2-1B \
--task text-generation \
--recipe xnnpack \
--output_dir llama_model

Run on-device with the LLM runner API:

C++

#include<executorch/extension/llm/runner/text_llm_runner.h>auto runner = create_llama_runner("llama.pte", "tiktoken.bin");
executorch::extension::llm::GenerationConfig config{
.seq_len = 128, .temperature = 0.8f};
runner->generate("Hello, how are you?", config);

Swift (iOS)

import ExecuTorchLLM
letrunner=TextRunner(modelPath:"llama.pte", tokenizerPath:"tiktoken.bin")try runner.generate("Hello, how are you?",Config{
$0.sequenceLength =128}){ token inprint(token, terminator:"")}

Kotlin (Android)API DocsDemo App

val llmModule =LlmModule("llama.pte", "tiktoken.bin", 0.8f)
llmModule.load()
llmModule.generate("Hello, how are you?", 128, object:LlmCallback {
overridefunonResult(result:String) { print(result) }
overridefunonStats(stats:String) { }
})

For multimodal models (vision, audio), use the MultiModal runner API which extends the LLM runner to handle image and audio inputs alongside text. See Llava and Voxtral examples.

See examples/models/llama for complete workflow including quantization, mobile deployment, and advanced options.

Next Steps:

Platform & Hardware Support

PlatformSupported Backends
AndroidXNNPACK, Vulkan, Qualcomm, MediaTek, Samsung Exynos
iOSXNNPACK, CoreML (Neural Engine)
Linux / WindowsXNNPACK, OpenVINO, CUDA (experimental)
macOSXNNPACK, Metal (experimental), MLX (experimental)
Embedded / MCUXNNPACK, ARM Ethos-U, NXP, Cadence DSP

See Backend Documentation for detailed hardware requirements and optimization guides. For desktop/laptop GPU inference with CUDA and Metal, see the Desktop Guide. For Zephyr RTOS integration, see the Zephyr Guide.

Production Deployments

ExecuTorch powers on-device AI at scale across Meta's family of apps, VR/AR devices, and partner deployments. View success stories →

Examples & Models

LLMs:Llama 3.2/3.1/3, Qwen 3, Phi-4-mini, LiquidAI LFM2

Multimodal:Llava (vision-language), Voxtral (audio-language), Gemma (vision-language)

Vision/Speech:MobileNetV2, DeepLabV3, YOLO26, Whisper, Supertonic

Resources:examples/ directory • executorch-examples out-of-tree demos • Optimum-ExecuTorch for HuggingFace models • Unsloth for fine-tuned LLM deployment

Key Features

ExecuTorch provides advanced capabilities for production deployment:

  • Quantization — Built-in support via torchao for 8-bit, 4-bit, and dynamic quantization
  • Memory Planning — Optimize memory usage with ahead-of-time allocation strategies
  • Developer Tools — ETDump profiler, ETRecord inspector, and model debugger
  • Selective Build — Strip unused operators to minimize binary size
  • Custom Operators — Extend with domain-specific kernels
  • Dynamic Shapes — Support variable input sizes with bounded ranges

See Advanced Topics for quantization techniques, custom backends, and compiler passes.

Documentation

Community & Contributing

We welcome contributions from the community!

Citing ExecuTorch

If you found ExecuTorch helpful in your research and would like to acknowledge it, please cite us using the following BibTeX:

@article{executorch2026,
title={{ExecuTorch} - A Unified {PyTorch} Solution to Run {AI} Models On-Device},
author={Nachin, Mergen and Desai, Digant and Jia, Sicheng Stephen and Lai, Chen and Liu, Mengwei and Szwejbka, Jacob and Alvarez, Raziel and Ascani, RJ and Bort, Dave and Candales, Manuel and others},
journal={arXiv preprint arXiv:2605.08195},
url={https://github.com/pytorch/executorch},
year={2026}
}

License

ExecuTorch is BSD licensed, as found in the LICENSE file.




Part of the PyTorch ecosystem

GitHubDocumentation

About

End-to-end solution for enabling on-device AI across mobile and edge devices for PyTorch models

Resources

Code of conduct

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages