Pybind merge fix - #42

Merged
ynimmaga merged 41 commits into
ynimmaga:openvino_backendfrom
cavusmustafa:pybind_merge_fix
Mar 27, 2025
Merged

Pybind merge fix#42
ynimmaga merged 41 commits into
ynimmaga:openvino_backendfrom
cavusmustafa:pybind_merge_fix

Conversation

@cavusmustafa

@cavusmustafacavusmustafa commented Mar 27, 2025

Copy link
Copy Markdown
Collaborator
  • Merged latest executorch main into openvino_backend
  • Updated setup.py for latest pybind build updates

metascroyand others added 30 commits March 25, 2025 11:39
ETCoreML crashes in FB app. During debugging, we traced issue down to
this.
cc @kimishpatel@YifanShenSZ@cymbalrush
HF version bump. Ensure `optimum-executorch` can work on new
`transformers` models with `executorch==0.6.0`
### Test plan
CI to test HF models
Co-authored-by: Guang Yang <guangyang@fb.com>
Differential Revision: D71761219
Pull Request resolved: pytorch#9554
Summary: .
Reviewed By: bsoyluoglu
Differential Revision: D71752749
### Summary
Seeing this error in Linux wheel building jobs:
```
Collecting numpy (from torchvision==0.22.0.dev20250311)
Downloading numpy-2.2.4-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.metadata (62 kB)
INFO: pip is looking at multiple versions of torchvision to determine which version is compatible with other requirements. This could take a while.
The conflict is caused by:
The user requested torch==2.7.0.dev20250311
torchvision 0.22.0.dev20250311+cpu depends on torch==2.7.0.dev20250310
``` ### Test plan
CI
https://github.com/pytorch/executorch/actions/runs/14047575373/job/39331644423
There seems to be some CI issues with:
```
torch._dynamo.exc.FailOnRecompileLimitHit: recompile_limit reached with one_graph=True. Excessive recompilations can degrade performance due to the compilation overhead of each recompilation. To monitor recompilations, enable TORCH_LOGS=recompiles. If recompilations are expected, consider increasing
```
To help resolve this we reset dynamo at setup for all unittests. Let's
see if this helps
Differential Revision: D70329890
Pull Request resolved: pytorch#8772
### Summary
We seem to be using a combination of CMAKE_ARGS and environment
variables when creating wheels. Ultimately, CMake only uses the cmake
args, however we redefine some of these flags as env vars to help
`setup.py` determine if a certain feature is turned on. Specifically, it
looks for pybinding vars to bundle pybindings.
Let's remove this redundancy and just use the CMAKE_ARGS as the single
source of truth. For more details and other considerations, see
pytorch#9494 (abandoned).
Note that even in the wheel building jobs, we use cmake args instead of
environment variables to control features:
https://github.com/pytorch/executorch/blob/644b7ddf14180d97e348faa627f576e13d367d69/.ci/scripts/wheel/envvar_base.sh#L20https://github.com/pytorch/executorch/blob/644b7ddf14180d97e348faa627f576e13d367d69/.ci/scripts/wheel/envvar_macos.sh#L14-L15
### Test plan
build + check CMakeCache.txt to ensure flags are set
```bash
# Expected: EXECUTORCH_BUILD_PYBIND=OFF EXECUTORCH_BUILD_XNNPACK=OFF EXECUTORCH_BUILD_COREML=OFF
$ rm -rf pip-out dist && ./install_executorch.sh --pybind off
# Expected: EXECUTORCH_BUILD_PYBIND=ON EXECUTORCH_BUILD_XNNPACK=ON EXECUTORCH_BUILD_COREML=OFF
$ rm -rf pip-out dist && ./install_executorch.sh
# Expected: EXECUTORCH_BUILD_PYBIND=ON EXECUTORCH_BUILD_XNNPACK=OFF EXECUTORCH_BUILD_COREML=ON
$ rm -rf pip-out dist && ./install_executorch.sh --pybind coreml
# Expected: EXECUTORCH_BUILD_PYBIND=ON EXECUTORCH_BUILD_XNNPACK=ON EXECUTORCH_BUILD_COREML=ON
$ rm -rf pip-out dist && ./install_executorch.sh --pybind xnnpack coreml
# Throws an error
$ rm -rf pip-out dist && ./install_executorch.sh --pybind coreml off
```
cc @larryliu0820@lucylq
Differential Revision: D71752746
Pull Request resolved: pytorch#9597
Differential Revision: D71752743
Pull Request resolved: pytorch#9606
Differential Revision: D71752747
Pull Request resolved: pytorch#9608
Differential Revision: D71752748
Pull Request resolved: pytorch#9609
### Summary
We want `EXECUTORCH_BUILD_PYBIND` enabled if the user wants to build the
bindings — so let's just do it. Unless of course, they explicitly choose
not to by defining the arg themselves.
### Test plan
CI
cc @larryliu0820@lucylq
Summary: A few SoCs have been supported recently, updated the
documentation.
Differential Revision: D71827272
cc @mergennachin@byjlw
Number of delegates and tolerance has changed so update it.
This was (accidentally) removed at a refactoring. Also takes the chance
to use the new XfailIf.. decorator.
Signed-off-by: Erik Lundell <erik.lundell@arm.com>
…rch#9158)
### Summary
Add a proxy for an `export_llama` performance regression test by
comparing the ops in the graph before and after the PR. The export
happens without loading a checkpoint or params file, which means that
all of the base `ModelArgs` values for `llama_transformer` will be used.
### Test plan
N/A
Differential Revision: D71833608
Pull Request resolved: pytorch#9603
Add conv3d tests, though most are skipped since conv3d support is not
yet implemented.
Signed-off-by: Erik Lundell <erik.lundell@arm.com>
Seems to succeed in rare instances due to randomness.
Signed-off-by: Erik Lundell <erik.lundell@arm.com>
### Summary
This is the stage 1 of Mimi Enablement.
Stage 2 will consist of actual model enablement.
- Support OP:
- exp
- expm1
- elu
- transpose conv1d
- bitwise_and
- scalar_tensor
- stack
- unbind
Summary:
To support embedded system builds which threw an error on the warning
for left shift by 32 on a 32 bit dtype, the code was modified to:
```
memory_offset |= static_cast<size_t>(memory_offset_high)
<< (sizeof(size_t) - sizeof(uint32_t));
```
This fails for build of OSS qwen example however.
Instead, we modify to add a check for
```
sizeof(size_t) > sizeof(uint32_t)
```
in the conditional instead of changing the computation.
In our builds of interest, this compiles away the if branch
Reviewed By: digantdesai, dpalmasan
Differential Revision: D71488571
### Summary
Context: pytorch#9481
* Include the `executorchcoreml` pybinding in the builds
* Remove separate installation option
* Turn on CoreML by default for macOS builds
* Add a dependency on coremltools for macOS
### Test plan
CI
```
$ rm -rf cmake-out pip-out dist && ./install_executorch.sh
$ ./examples/models/llama/install_requirements.sh
$ .ci/scripts/test_llama.sh -model stories110M -build_tool cmake -dtype fp32 -mode coreml
$ .ci/scripts/test_llama.sh -model stories110M -build_tool cmake -dtype fp32 -mode xnnpack+custom+quantize_kv
```
cc @larryliu0820@lucylq
…ntization for Example Models (pytorch#9634)
### Summary
Changes:
1. When initializing Llama2 for aot_compiler, since checkpoints can only
e downloaded from hugging face, we initialize llama2 with uninitialized
weights. The problem with this is that when running quantization, we can
run into errors with the histogram if the unitialized values are nan. We
fix this by initializing the weights with zeros if no check point is
provided. This enforces that quantization step can still work.
2. Quant Type in AoT compiler. When looking at the model options
available to XNNPACK, everything is quantized with per-tensor static
quantization. This isn't the best option for all the models available.
For example transformer based models like Llama and MobileBert would
likely prefer dynamically quantized per channel weights, where has CNN
like MobileNet would prefer statically quantized per channel weights. We
add this type of Quant Type to the existing models options. This also
helps with Test Timeouts. per-tensor static quantization on a model like
llama can take a long time due to the introduction of MANY q/dq nodes,
and the complex partitions it creates. As a result, proposing partitions
can take a long time due to the constant BFS to find the largest
possible partition. By specifying the more apt quantization scheme like
dynamic per-channel quantization, we can avoid this complexity.
Overall this should help with flakey [nan, nan] errors in the
quantization histogram, and it should also help with CI timing out.
### Test plan
OSS XNNPACK CI for all model delegation
cc @digantdesai@cbilgin
### Summary
* After pytorch#9483, we should have
CoreML support out of the box for macOS
* Unfortunately, we still need
`backends/apple/coreml/scripts/install_requirements.sh` to use the
[coreml_executorch_runner](https://github.com/pytorch/executorch/tree/main/examples/apple/coreml/executor_runner)
(used for testing)
* I should have caught all the usage
### Test plan
Read
swolchokand others added 11 commits March 26, 2025 13:26
I have no idea what this file actually does, but it seems like we are
supposed to have this?
…ch#9509)
Disable one, fix the other.
Testing: built internally
pytorch#9511)
I planned to do this everywhere and forgot. Clean it all up, leave a
note, enforce the note with visibility. This makes sure everything in
buck-land gets ET_USE_THREADPOOL.
Test Plan: Profiled run on internal model, no longer seeing
parallel_for_no_threadpool
…AME_AS_COMPUTE (pytorch#9613)
As the title says, this is mostly a few related find-replaces, plus
marking SupportedTensorDtypes::SAME_AS_COMPUTE deprecated.
As the code comment says, these APIs are undergoing development (see
e.g. pytorch#9613) and it's pretty
inconvenient that they're incidentally committed-to externally. Mark
them deprecated so we have the option to drop that commitment in (IIUC)
0.7.
Previous attempt to bump HF transformers version to latest is reverted
due to llava model imcompatibility. This PR is to just ensure the CI are
able to test `optimum-executorch` with latest version of HF transformers
and upcoming `executorch==0.6.0`.
Note: This change is purely on CI and only for optimum-executorch,
should not affect other models like llava.
Co-authored-by: Guang Yang <guangyang@fb.com>
Differential Revision: D69994481
Pull Request resolved: pytorch#8703
@ynimmaga
ynimmaga merged commit 3f53cc2 into ynimmaga:openvino_backendMar 27, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

18 participants

@cavusmustafa@ynimmaga@metascroy@guangy10@shoumikhin@kirklandsign@jathu@mcr229@jackzhxng@larryliu0820@cccclai@mansnils@Erik-Lundell@winskuo-quic@JakeStevens@swolchok@bigfootjon@mergennachin
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Pybind merge fix - #42

Merged
ynimmaga merged 41 commits into
ynimmaga:openvino_backendfrom
cavusmustafa:pybind_merge_fix
Mar 27, 2025
Merged

Pybind merge fix#42
ynimmaga merged 41 commits into
ynimmaga:openvino_backendfrom
cavusmustafa:pybind_merge_fix

Conversation

@cavusmustafa

@cavusmustafacavusmustafa commented Mar 27, 2025

Copy link
Copy Markdown
Collaborator
  • Merged latest executorch main into openvino_backend
  • Updated setup.py for latest pybind build updates

metascroyand others added 30 commits March 25, 2025 11:39
ETCoreML crashes in FB app. During debugging, we traced issue down to
this.
cc @kimishpatel@YifanShenSZ@cymbalrush
HF version bump. Ensure `optimum-executorch` can work on new
`transformers` models with `executorch==0.6.0`
### Test plan
CI to test HF models
Co-authored-by: Guang Yang <guangyang@fb.com>
Differential Revision: D71761219
Pull Request resolved: pytorch#9554
Summary: .
Reviewed By: bsoyluoglu
Differential Revision: D71752749
### Summary
Seeing this error in Linux wheel building jobs:
```
Collecting numpy (from torchvision==0.22.0.dev20250311)
Downloading numpy-2.2.4-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.metadata (62 kB)
INFO: pip is looking at multiple versions of torchvision to determine which version is compatible with other requirements. This could take a while.
The conflict is caused by:
The user requested torch==2.7.0.dev20250311
torchvision 0.22.0.dev20250311+cpu depends on torch==2.7.0.dev20250310
``` ### Test plan
CI
https://github.com/pytorch/executorch/actions/runs/14047575373/job/39331644423
There seems to be some CI issues with:
```
torch._dynamo.exc.FailOnRecompileLimitHit: recompile_limit reached with one_graph=True. Excessive recompilations can degrade performance due to the compilation overhead of each recompilation. To monitor recompilations, enable TORCH_LOGS=recompiles. If recompilations are expected, consider increasing
```
To help resolve this we reset dynamo at setup for all unittests. Let's
see if this helps
Differential Revision: D70329890
Pull Request resolved: pytorch#8772
### Summary
We seem to be using a combination of CMAKE_ARGS and environment
variables when creating wheels. Ultimately, CMake only uses the cmake
args, however we redefine some of these flags as env vars to help
`setup.py` determine if a certain feature is turned on. Specifically, it
looks for pybinding vars to bundle pybindings.
Let's remove this redundancy and just use the CMAKE_ARGS as the single
source of truth. For more details and other considerations, see
pytorch#9494 (abandoned).
Note that even in the wheel building jobs, we use cmake args instead of
environment variables to control features:
https://github.com/pytorch/executorch/blob/644b7ddf14180d97e348faa627f576e13d367d69/.ci/scripts/wheel/envvar_base.sh#L20https://github.com/pytorch/executorch/blob/644b7ddf14180d97e348faa627f576e13d367d69/.ci/scripts/wheel/envvar_macos.sh#L14-L15
### Test plan
build + check CMakeCache.txt to ensure flags are set
```bash
# Expected: EXECUTORCH_BUILD_PYBIND=OFF EXECUTORCH_BUILD_XNNPACK=OFF EXECUTORCH_BUILD_COREML=OFF
$ rm -rf pip-out dist && ./install_executorch.sh --pybind off
# Expected: EXECUTORCH_BUILD_PYBIND=ON EXECUTORCH_BUILD_XNNPACK=ON EXECUTORCH_BUILD_COREML=OFF
$ rm -rf pip-out dist && ./install_executorch.sh
# Expected: EXECUTORCH_BUILD_PYBIND=ON EXECUTORCH_BUILD_XNNPACK=OFF EXECUTORCH_BUILD_COREML=ON
$ rm -rf pip-out dist && ./install_executorch.sh --pybind coreml
# Expected: EXECUTORCH_BUILD_PYBIND=ON EXECUTORCH_BUILD_XNNPACK=ON EXECUTORCH_BUILD_COREML=ON
$ rm -rf pip-out dist && ./install_executorch.sh --pybind xnnpack coreml
# Throws an error
$ rm -rf pip-out dist && ./install_executorch.sh --pybind coreml off
```
cc @larryliu0820@lucylq
Differential Revision: D71752746
Pull Request resolved: pytorch#9597
Differential Revision: D71752743
Pull Request resolved: pytorch#9606
Differential Revision: D71752747
Pull Request resolved: pytorch#9608
Differential Revision: D71752748
Pull Request resolved: pytorch#9609
### Summary
We want `EXECUTORCH_BUILD_PYBIND` enabled if the user wants to build the
bindings — so let's just do it. Unless of course, they explicitly choose
not to by defining the arg themselves.
### Test plan
CI
cc @larryliu0820@lucylq
Summary: A few SoCs have been supported recently, updated the
documentation.
Differential Revision: D71827272
cc @mergennachin@byjlw
Number of delegates and tolerance has changed so update it.
This was (accidentally) removed at a refactoring. Also takes the chance
to use the new XfailIf.. decorator.
Signed-off-by: Erik Lundell <erik.lundell@arm.com>
…rch#9158)
### Summary
Add a proxy for an `export_llama` performance regression test by
comparing the ops in the graph before and after the PR. The export
happens without loading a checkpoint or params file, which means that
all of the base `ModelArgs` values for `llama_transformer` will be used.
### Test plan
N/A
Differential Revision: D71833608
Pull Request resolved: pytorch#9603
Add conv3d tests, though most are skipped since conv3d support is not
yet implemented.
Signed-off-by: Erik Lundell <erik.lundell@arm.com>
Seems to succeed in rare instances due to randomness.
Signed-off-by: Erik Lundell <erik.lundell@arm.com>
### Summary
This is the stage 1 of Mimi Enablement.
Stage 2 will consist of actual model enablement.
- Support OP:
- exp
- expm1
- elu
- transpose conv1d
- bitwise_and
- scalar_tensor
- stack
- unbind
Summary:
To support embedded system builds which threw an error on the warning
for left shift by 32 on a 32 bit dtype, the code was modified to:
```
memory_offset |= static_cast<size_t>(memory_offset_high)
<< (sizeof(size_t) - sizeof(uint32_t));
```
This fails for build of OSS qwen example however.
Instead, we modify to add a check for
```
sizeof(size_t) > sizeof(uint32_t)
```
in the conditional instead of changing the computation.
In our builds of interest, this compiles away the if branch
Reviewed By: digantdesai, dpalmasan
Differential Revision: D71488571
### Summary
Context: pytorch#9481
* Include the `executorchcoreml` pybinding in the builds
* Remove separate installation option
* Turn on CoreML by default for macOS builds
* Add a dependency on coremltools for macOS
### Test plan
CI
```
$ rm -rf cmake-out pip-out dist && ./install_executorch.sh
$ ./examples/models/llama/install_requirements.sh
$ .ci/scripts/test_llama.sh -model stories110M -build_tool cmake -dtype fp32 -mode coreml
$ .ci/scripts/test_llama.sh -model stories110M -build_tool cmake -dtype fp32 -mode xnnpack+custom+quantize_kv
```
cc @larryliu0820@lucylq
…ntization for Example Models (pytorch#9634)
### Summary
Changes:
1. When initializing Llama2 for aot_compiler, since checkpoints can only
e downloaded from hugging face, we initialize llama2 with uninitialized
weights. The problem with this is that when running quantization, we can
run into errors with the histogram if the unitialized values are nan. We
fix this by initializing the weights with zeros if no check point is
provided. This enforces that quantization step can still work.
2. Quant Type in AoT compiler. When looking at the model options
available to XNNPACK, everything is quantized with per-tensor static
quantization. This isn't the best option for all the models available.
For example transformer based models like Llama and MobileBert would
likely prefer dynamically quantized per channel weights, where has CNN
like MobileNet would prefer statically quantized per channel weights. We
add this type of Quant Type to the existing models options. This also
helps with Test Timeouts. per-tensor static quantization on a model like
llama can take a long time due to the introduction of MANY q/dq nodes,
and the complex partitions it creates. As a result, proposing partitions
can take a long time due to the constant BFS to find the largest
possible partition. By specifying the more apt quantization scheme like
dynamic per-channel quantization, we can avoid this complexity.
Overall this should help with flakey [nan, nan] errors in the
quantization histogram, and it should also help with CI timing out.
### Test plan
OSS XNNPACK CI for all model delegation
cc @digantdesai@cbilgin
### Summary
* After pytorch#9483, we should have
CoreML support out of the box for macOS
* Unfortunately, we still need
`backends/apple/coreml/scripts/install_requirements.sh` to use the
[coreml_executorch_runner](https://github.com/pytorch/executorch/tree/main/examples/apple/coreml/executor_runner)
(used for testing)
* I should have caught all the usage
### Test plan
Read
swolchokand others added 11 commits March 26, 2025 13:26
I have no idea what this file actually does, but it seems like we are
supposed to have this?
…ch#9509)
Disable one, fix the other.
Testing: built internally
pytorch#9511)
I planned to do this everywhere and forgot. Clean it all up, leave a
note, enforce the note with visibility. This makes sure everything in
buck-land gets ET_USE_THREADPOOL.
Test Plan: Profiled run on internal model, no longer seeing
parallel_for_no_threadpool
…AME_AS_COMPUTE (pytorch#9613)
As the title says, this is mostly a few related find-replaces, plus
marking SupportedTensorDtypes::SAME_AS_COMPUTE deprecated.
As the code comment says, these APIs are undergoing development (see
e.g. pytorch#9613) and it's pretty
inconvenient that they're incidentally committed-to externally. Mark
them deprecated so we have the option to drop that commitment in (IIUC)
0.7.
Previous attempt to bump HF transformers version to latest is reverted
due to llava model imcompatibility. This PR is to just ensure the CI are
able to test `optimum-executorch` with latest version of HF transformers
and upcoming `executorch==0.6.0`.
Note: This change is purely on CI and only for optimum-executorch,
should not affect other models like llava.
Co-authored-by: Guang Yang <guangyang@fb.com>
Differential Revision: D69994481
Pull Request resolved: pytorch#8703
@ynimmaga
ynimmaga merged commit 3f53cc2 into ynimmaga:openvino_backendMar 27, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

18 participants

@cavusmustafa@ynimmaga@metascroy@guangy10@shoumikhin@kirklandsign@jathu@mcr229@jackzhxng@larryliu0820@cccclai@mansnils@Erik-Lundell@winskuo-quic@JakeStevens@swolchok@bigfootjon@mergennachin
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Pybind merge fix - #42

Merged
ynimmaga merged 41 commits into
ynimmaga:openvino_backendfrom
cavusmustafa:pybind_merge_fix
Mar 27, 2025
Merged

Pybind merge fix#42
ynimmaga merged 41 commits into
ynimmaga:openvino_backendfrom
cavusmustafa:pybind_merge_fix

Conversation

@cavusmustafa

@cavusmustafacavusmustafa commented Mar 27, 2025

Copy link
Copy Markdown
Collaborator
  • Merged latest executorch main into openvino_backend
  • Updated setup.py for latest pybind build updates

metascroyand others added 30 commits March 25, 2025 11:39
ETCoreML crashes in FB app. During debugging, we traced issue down to
this.
cc @kimishpatel@YifanShenSZ@cymbalrush
HF version bump. Ensure `optimum-executorch` can work on new
`transformers` models with `executorch==0.6.0`
### Test plan
CI to test HF models
Co-authored-by: Guang Yang <guangyang@fb.com>
Differential Revision: D71761219
Pull Request resolved: pytorch#9554
Summary: .
Reviewed By: bsoyluoglu
Differential Revision: D71752749
### Summary
Seeing this error in Linux wheel building jobs:
```
Collecting numpy (from torchvision==0.22.0.dev20250311)
Downloading numpy-2.2.4-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.metadata (62 kB)
INFO: pip is looking at multiple versions of torchvision to determine which version is compatible with other requirements. This could take a while.
The conflict is caused by:
The user requested torch==2.7.0.dev20250311
torchvision 0.22.0.dev20250311+cpu depends on torch==2.7.0.dev20250310
``` ### Test plan
CI
https://github.com/pytorch/executorch/actions/runs/14047575373/job/39331644423
There seems to be some CI issues with:
```
torch._dynamo.exc.FailOnRecompileLimitHit: recompile_limit reached with one_graph=True. Excessive recompilations can degrade performance due to the compilation overhead of each recompilation. To monitor recompilations, enable TORCH_LOGS=recompiles. If recompilations are expected, consider increasing
```
To help resolve this we reset dynamo at setup for all unittests. Let's
see if this helps
Differential Revision: D70329890
Pull Request resolved: pytorch#8772
### Summary
We seem to be using a combination of CMAKE_ARGS and environment
variables when creating wheels. Ultimately, CMake only uses the cmake
args, however we redefine some of these flags as env vars to help
`setup.py` determine if a certain feature is turned on. Specifically, it
looks for pybinding vars to bundle pybindings.
Let's remove this redundancy and just use the CMAKE_ARGS as the single
source of truth. For more details and other considerations, see
pytorch#9494 (abandoned).
Note that even in the wheel building jobs, we use cmake args instead of
environment variables to control features:
https://github.com/pytorch/executorch/blob/644b7ddf14180d97e348faa627f576e13d367d69/.ci/scripts/wheel/envvar_base.sh#L20https://github.com/pytorch/executorch/blob/644b7ddf14180d97e348faa627f576e13d367d69/.ci/scripts/wheel/envvar_macos.sh#L14-L15
### Test plan
build + check CMakeCache.txt to ensure flags are set
```bash
# Expected: EXECUTORCH_BUILD_PYBIND=OFF EXECUTORCH_BUILD_XNNPACK=OFF EXECUTORCH_BUILD_COREML=OFF
$ rm -rf pip-out dist && ./install_executorch.sh --pybind off
# Expected: EXECUTORCH_BUILD_PYBIND=ON EXECUTORCH_BUILD_XNNPACK=ON EXECUTORCH_BUILD_COREML=OFF
$ rm -rf pip-out dist && ./install_executorch.sh
# Expected: EXECUTORCH_BUILD_PYBIND=ON EXECUTORCH_BUILD_XNNPACK=OFF EXECUTORCH_BUILD_COREML=ON
$ rm -rf pip-out dist && ./install_executorch.sh --pybind coreml
# Expected: EXECUTORCH_BUILD_PYBIND=ON EXECUTORCH_BUILD_XNNPACK=ON EXECUTORCH_BUILD_COREML=ON
$ rm -rf pip-out dist && ./install_executorch.sh --pybind xnnpack coreml
# Throws an error
$ rm -rf pip-out dist && ./install_executorch.sh --pybind coreml off
```
cc @larryliu0820@lucylq
Differential Revision: D71752746
Pull Request resolved: pytorch#9597
Differential Revision: D71752743
Pull Request resolved: pytorch#9606
Differential Revision: D71752747
Pull Request resolved: pytorch#9608
Differential Revision: D71752748
Pull Request resolved: pytorch#9609
### Summary
We want `EXECUTORCH_BUILD_PYBIND` enabled if the user wants to build the
bindings — so let's just do it. Unless of course, they explicitly choose
not to by defining the arg themselves.
### Test plan
CI
cc @larryliu0820@lucylq
Summary: A few SoCs have been supported recently, updated the
documentation.
Differential Revision: D71827272
cc @mergennachin@byjlw
Number of delegates and tolerance has changed so update it.
This was (accidentally) removed at a refactoring. Also takes the chance
to use the new XfailIf.. decorator.
Signed-off-by: Erik Lundell <erik.lundell@arm.com>
…rch#9158)
### Summary
Add a proxy for an `export_llama` performance regression test by
comparing the ops in the graph before and after the PR. The export
happens without loading a checkpoint or params file, which means that
all of the base `ModelArgs` values for `llama_transformer` will be used.
### Test plan
N/A
Differential Revision: D71833608
Pull Request resolved: pytorch#9603
Add conv3d tests, though most are skipped since conv3d support is not
yet implemented.
Signed-off-by: Erik Lundell <erik.lundell@arm.com>
Seems to succeed in rare instances due to randomness.
Signed-off-by: Erik Lundell <erik.lundell@arm.com>
### Summary
This is the stage 1 of Mimi Enablement.
Stage 2 will consist of actual model enablement.
- Support OP:
- exp
- expm1
- elu
- transpose conv1d
- bitwise_and
- scalar_tensor
- stack
- unbind
Summary:
To support embedded system builds which threw an error on the warning
for left shift by 32 on a 32 bit dtype, the code was modified to:
```
memory_offset |= static_cast<size_t>(memory_offset_high)
<< (sizeof(size_t) - sizeof(uint32_t));
```
This fails for build of OSS qwen example however.
Instead, we modify to add a check for
```
sizeof(size_t) > sizeof(uint32_t)
```
in the conditional instead of changing the computation.
In our builds of interest, this compiles away the if branch
Reviewed By: digantdesai, dpalmasan
Differential Revision: D71488571
### Summary
Context: pytorch#9481
* Include the `executorchcoreml` pybinding in the builds
* Remove separate installation option
* Turn on CoreML by default for macOS builds
* Add a dependency on coremltools for macOS
### Test plan
CI
```
$ rm -rf cmake-out pip-out dist && ./install_executorch.sh
$ ./examples/models/llama/install_requirements.sh
$ .ci/scripts/test_llama.sh -model stories110M -build_tool cmake -dtype fp32 -mode coreml
$ .ci/scripts/test_llama.sh -model stories110M -build_tool cmake -dtype fp32 -mode xnnpack+custom+quantize_kv
```
cc @larryliu0820@lucylq
…ntization for Example Models (pytorch#9634)
### Summary
Changes:
1. When initializing Llama2 for aot_compiler, since checkpoints can only
e downloaded from hugging face, we initialize llama2 with uninitialized
weights. The problem with this is that when running quantization, we can
run into errors with the histogram if the unitialized values are nan. We
fix this by initializing the weights with zeros if no check point is
provided. This enforces that quantization step can still work.
2. Quant Type in AoT compiler. When looking at the model options
available to XNNPACK, everything is quantized with per-tensor static
quantization. This isn't the best option for all the models available.
For example transformer based models like Llama and MobileBert would
likely prefer dynamically quantized per channel weights, where has CNN
like MobileNet would prefer statically quantized per channel weights. We
add this type of Quant Type to the existing models options. This also
helps with Test Timeouts. per-tensor static quantization on a model like
llama can take a long time due to the introduction of MANY q/dq nodes,
and the complex partitions it creates. As a result, proposing partitions
can take a long time due to the constant BFS to find the largest
possible partition. By specifying the more apt quantization scheme like
dynamic per-channel quantization, we can avoid this complexity.
Overall this should help with flakey [nan, nan] errors in the
quantization histogram, and it should also help with CI timing out.
### Test plan
OSS XNNPACK CI for all model delegation
cc @digantdesai@cbilgin
### Summary
* After pytorch#9483, we should have
CoreML support out of the box for macOS
* Unfortunately, we still need
`backends/apple/coreml/scripts/install_requirements.sh` to use the
[coreml_executorch_runner](https://github.com/pytorch/executorch/tree/main/examples/apple/coreml/executor_runner)
(used for testing)
* I should have caught all the usage
### Test plan
Read
swolchokand others added 11 commits March 26, 2025 13:26
I have no idea what this file actually does, but it seems like we are
supposed to have this?
…ch#9509)
Disable one, fix the other.
Testing: built internally
pytorch#9511)
I planned to do this everywhere and forgot. Clean it all up, leave a
note, enforce the note with visibility. This makes sure everything in
buck-land gets ET_USE_THREADPOOL.
Test Plan: Profiled run on internal model, no longer seeing
parallel_for_no_threadpool
…AME_AS_COMPUTE (pytorch#9613)
As the title says, this is mostly a few related find-replaces, plus
marking SupportedTensorDtypes::SAME_AS_COMPUTE deprecated.
As the code comment says, these APIs are undergoing development (see
e.g. pytorch#9613) and it's pretty
inconvenient that they're incidentally committed-to externally. Mark
them deprecated so we have the option to drop that commitment in (IIUC)
0.7.
Previous attempt to bump HF transformers version to latest is reverted
due to llava model imcompatibility. This PR is to just ensure the CI are
able to test `optimum-executorch` with latest version of HF transformers
and upcoming `executorch==0.6.0`.
Note: This change is purely on CI and only for optimum-executorch,
should not affect other models like llava.
Co-authored-by: Guang Yang <guangyang@fb.com>
Differential Revision: D69994481
Pull Request resolved: pytorch#8703
@ynimmaga
ynimmaga merged commit 3f53cc2 into ynimmaga:openvino_backendMar 27, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

18 participants

@cavusmustafa@ynimmaga@metascroy@guangy10@shoumikhin@kirklandsign@jathu@mcr229@jackzhxng@larryliu0820@cccclai@mansnils@Erik-Lundell@winskuo-quic@JakeStevens@swolchok@bigfootjon@mergennachin
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Pybind merge fix - #42

Merged
ynimmaga merged 41 commits into
ynimmaga:openvino_backendfrom
cavusmustafa:pybind_merge_fix
Mar 27, 2025
Merged

Pybind merge fix#42
ynimmaga merged 41 commits into
ynimmaga:openvino_backendfrom
cavusmustafa:pybind_merge_fix

Conversation

@cavusmustafa

@cavusmustafacavusmustafa commented Mar 27, 2025

Copy link
Copy Markdown
Collaborator
  • Merged latest executorch main into openvino_backend
  • Updated setup.py for latest pybind build updates

metascroyand others added 30 commits March 25, 2025 11:39
ETCoreML crashes in FB app. During debugging, we traced issue down to
this.
cc @kimishpatel@YifanShenSZ@cymbalrush
HF version bump. Ensure `optimum-executorch` can work on new
`transformers` models with `executorch==0.6.0`
### Test plan
CI to test HF models
Co-authored-by: Guang Yang <guangyang@fb.com>
Differential Revision: D71761219
Pull Request resolved: pytorch#9554
Summary: .
Reviewed By: bsoyluoglu
Differential Revision: D71752749
### Summary
Seeing this error in Linux wheel building jobs:
```
Collecting numpy (from torchvision==0.22.0.dev20250311)
Downloading numpy-2.2.4-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.metadata (62 kB)
INFO: pip is looking at multiple versions of torchvision to determine which version is compatible with other requirements. This could take a while.
The conflict is caused by:
The user requested torch==2.7.0.dev20250311
torchvision 0.22.0.dev20250311+cpu depends on torch==2.7.0.dev20250310
``` ### Test plan
CI
https://github.com/pytorch/executorch/actions/runs/14047575373/job/39331644423
There seems to be some CI issues with:
```
torch._dynamo.exc.FailOnRecompileLimitHit: recompile_limit reached with one_graph=True. Excessive recompilations can degrade performance due to the compilation overhead of each recompilation. To monitor recompilations, enable TORCH_LOGS=recompiles. If recompilations are expected, consider increasing
```
To help resolve this we reset dynamo at setup for all unittests. Let's
see if this helps
Differential Revision: D70329890
Pull Request resolved: pytorch#8772
### Summary
We seem to be using a combination of CMAKE_ARGS and environment
variables when creating wheels. Ultimately, CMake only uses the cmake
args, however we redefine some of these flags as env vars to help
`setup.py` determine if a certain feature is turned on. Specifically, it
looks for pybinding vars to bundle pybindings.
Let's remove this redundancy and just use the CMAKE_ARGS as the single
source of truth. For more details and other considerations, see
pytorch#9494 (abandoned).
Note that even in the wheel building jobs, we use cmake args instead of
environment variables to control features:
https://github.com/pytorch/executorch/blob/644b7ddf14180d97e348faa627f576e13d367d69/.ci/scripts/wheel/envvar_base.sh#L20https://github.com/pytorch/executorch/blob/644b7ddf14180d97e348faa627f576e13d367d69/.ci/scripts/wheel/envvar_macos.sh#L14-L15
### Test plan
build + check CMakeCache.txt to ensure flags are set
```bash
# Expected: EXECUTORCH_BUILD_PYBIND=OFF EXECUTORCH_BUILD_XNNPACK=OFF EXECUTORCH_BUILD_COREML=OFF
$ rm -rf pip-out dist && ./install_executorch.sh --pybind off
# Expected: EXECUTORCH_BUILD_PYBIND=ON EXECUTORCH_BUILD_XNNPACK=ON EXECUTORCH_BUILD_COREML=OFF
$ rm -rf pip-out dist && ./install_executorch.sh
# Expected: EXECUTORCH_BUILD_PYBIND=ON EXECUTORCH_BUILD_XNNPACK=OFF EXECUTORCH_BUILD_COREML=ON
$ rm -rf pip-out dist && ./install_executorch.sh --pybind coreml
# Expected: EXECUTORCH_BUILD_PYBIND=ON EXECUTORCH_BUILD_XNNPACK=ON EXECUTORCH_BUILD_COREML=ON
$ rm -rf pip-out dist && ./install_executorch.sh --pybind xnnpack coreml
# Throws an error
$ rm -rf pip-out dist && ./install_executorch.sh --pybind coreml off
```
cc @larryliu0820@lucylq
Differential Revision: D71752746
Pull Request resolved: pytorch#9597
Differential Revision: D71752743
Pull Request resolved: pytorch#9606
Differential Revision: D71752747
Pull Request resolved: pytorch#9608
Differential Revision: D71752748
Pull Request resolved: pytorch#9609
### Summary
We want `EXECUTORCH_BUILD_PYBIND` enabled if the user wants to build the
bindings — so let's just do it. Unless of course, they explicitly choose
not to by defining the arg themselves.
### Test plan
CI
cc @larryliu0820@lucylq
Summary: A few SoCs have been supported recently, updated the
documentation.
Differential Revision: D71827272
cc @mergennachin@byjlw
Number of delegates and tolerance has changed so update it.
This was (accidentally) removed at a refactoring. Also takes the chance
to use the new XfailIf.. decorator.
Signed-off-by: Erik Lundell <erik.lundell@arm.com>
…rch#9158)
### Summary
Add a proxy for an `export_llama` performance regression test by
comparing the ops in the graph before and after the PR. The export
happens without loading a checkpoint or params file, which means that
all of the base `ModelArgs` values for `llama_transformer` will be used.
### Test plan
N/A
Differential Revision: D71833608
Pull Request resolved: pytorch#9603
Add conv3d tests, though most are skipped since conv3d support is not
yet implemented.
Signed-off-by: Erik Lundell <erik.lundell@arm.com>
Seems to succeed in rare instances due to randomness.
Signed-off-by: Erik Lundell <erik.lundell@arm.com>
### Summary
This is the stage 1 of Mimi Enablement.
Stage 2 will consist of actual model enablement.
- Support OP:
- exp
- expm1
- elu
- transpose conv1d
- bitwise_and
- scalar_tensor
- stack
- unbind
Summary:
To support embedded system builds which threw an error on the warning
for left shift by 32 on a 32 bit dtype, the code was modified to:
```
memory_offset |= static_cast<size_t>(memory_offset_high)
<< (sizeof(size_t) - sizeof(uint32_t));
```
This fails for build of OSS qwen example however.
Instead, we modify to add a check for
```
sizeof(size_t) > sizeof(uint32_t)
```
in the conditional instead of changing the computation.
In our builds of interest, this compiles away the if branch
Reviewed By: digantdesai, dpalmasan
Differential Revision: D71488571
### Summary
Context: pytorch#9481
* Include the `executorchcoreml` pybinding in the builds
* Remove separate installation option
* Turn on CoreML by default for macOS builds
* Add a dependency on coremltools for macOS
### Test plan
CI
```
$ rm -rf cmake-out pip-out dist && ./install_executorch.sh
$ ./examples/models/llama/install_requirements.sh
$ .ci/scripts/test_llama.sh -model stories110M -build_tool cmake -dtype fp32 -mode coreml
$ .ci/scripts/test_llama.sh -model stories110M -build_tool cmake -dtype fp32 -mode xnnpack+custom+quantize_kv
```
cc @larryliu0820@lucylq
…ntization for Example Models (pytorch#9634)
### Summary
Changes:
1. When initializing Llama2 for aot_compiler, since checkpoints can only
e downloaded from hugging face, we initialize llama2 with uninitialized
weights. The problem with this is that when running quantization, we can
run into errors with the histogram if the unitialized values are nan. We
fix this by initializing the weights with zeros if no check point is
provided. This enforces that quantization step can still work.
2. Quant Type in AoT compiler. When looking at the model options
available to XNNPACK, everything is quantized with per-tensor static
quantization. This isn't the best option for all the models available.
For example transformer based models like Llama and MobileBert would
likely prefer dynamically quantized per channel weights, where has CNN
like MobileNet would prefer statically quantized per channel weights. We
add this type of Quant Type to the existing models options. This also
helps with Test Timeouts. per-tensor static quantization on a model like
llama can take a long time due to the introduction of MANY q/dq nodes,
and the complex partitions it creates. As a result, proposing partitions
can take a long time due to the constant BFS to find the largest
possible partition. By specifying the more apt quantization scheme like
dynamic per-channel quantization, we can avoid this complexity.
Overall this should help with flakey [nan, nan] errors in the
quantization histogram, and it should also help with CI timing out.
### Test plan
OSS XNNPACK CI for all model delegation
cc @digantdesai@cbilgin
### Summary
* After pytorch#9483, we should have
CoreML support out of the box for macOS
* Unfortunately, we still need
`backends/apple/coreml/scripts/install_requirements.sh` to use the
[coreml_executorch_runner](https://github.com/pytorch/executorch/tree/main/examples/apple/coreml/executor_runner)
(used for testing)
* I should have caught all the usage
### Test plan
Read
swolchokand others added 11 commits March 26, 2025 13:26
I have no idea what this file actually does, but it seems like we are
supposed to have this?
…ch#9509)
Disable one, fix the other.
Testing: built internally
pytorch#9511)
I planned to do this everywhere and forgot. Clean it all up, leave a
note, enforce the note with visibility. This makes sure everything in
buck-land gets ET_USE_THREADPOOL.
Test Plan: Profiled run on internal model, no longer seeing
parallel_for_no_threadpool
…AME_AS_COMPUTE (pytorch#9613)
As the title says, this is mostly a few related find-replaces, plus
marking SupportedTensorDtypes::SAME_AS_COMPUTE deprecated.
As the code comment says, these APIs are undergoing development (see
e.g. pytorch#9613) and it's pretty
inconvenient that they're incidentally committed-to externally. Mark
them deprecated so we have the option to drop that commitment in (IIUC)
0.7.
Previous attempt to bump HF transformers version to latest is reverted
due to llava model imcompatibility. This PR is to just ensure the CI are
able to test `optimum-executorch` with latest version of HF transformers
and upcoming `executorch==0.6.0`.
Note: This change is purely on CI and only for optimum-executorch,
should not affect other models like llava.
Co-authored-by: Guang Yang <guangyang@fb.com>
Differential Revision: D69994481
Pull Request resolved: pytorch#8703
@ynimmaga
ynimmaga merged commit 3f53cc2 into ynimmaga:openvino_backendMar 27, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

18 participants

@cavusmustafa@ynimmaga@metascroy@guangy10@shoumikhin@kirklandsign@jathu@mcr229@jackzhxng@larryliu0820@cccclai@mansnils@Erik-Lundell@winskuo-quic@JakeStevens@swolchok@bigfootjon@mergennachin
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Pybind merge fix - #42

Merged
ynimmaga merged 41 commits into
ynimmaga:openvino_backendfrom
cavusmustafa:pybind_merge_fix
Mar 27, 2025
Merged

Pybind merge fix#42
ynimmaga merged 41 commits into
ynimmaga:openvino_backendfrom
cavusmustafa:pybind_merge_fix

Conversation

@cavusmustafa

@cavusmustafacavusmustafa commented Mar 27, 2025

Copy link
Copy Markdown
Collaborator
  • Merged latest executorch main into openvino_backend
  • Updated setup.py for latest pybind build updates

metascroyand others added 30 commits March 25, 2025 11:39
ETCoreML crashes in FB app. During debugging, we traced issue down to
this.
cc @kimishpatel@YifanShenSZ@cymbalrush
HF version bump. Ensure `optimum-executorch` can work on new
`transformers` models with `executorch==0.6.0`
### Test plan
CI to test HF models
Co-authored-by: Guang Yang <guangyang@fb.com>
Differential Revision: D71761219
Pull Request resolved: pytorch#9554
Summary: .
Reviewed By: bsoyluoglu
Differential Revision: D71752749
### Summary
Seeing this error in Linux wheel building jobs:
```
Collecting numpy (from torchvision==0.22.0.dev20250311)
Downloading numpy-2.2.4-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.metadata (62 kB)
INFO: pip is looking at multiple versions of torchvision to determine which version is compatible with other requirements. This could take a while.
The conflict is caused by:
The user requested torch==2.7.0.dev20250311
torchvision 0.22.0.dev20250311+cpu depends on torch==2.7.0.dev20250310
``` ### Test plan
CI
https://github.com/pytorch/executorch/actions/runs/14047575373/job/39331644423
There seems to be some CI issues with:
```
torch._dynamo.exc.FailOnRecompileLimitHit: recompile_limit reached with one_graph=True. Excessive recompilations can degrade performance due to the compilation overhead of each recompilation. To monitor recompilations, enable TORCH_LOGS=recompiles. If recompilations are expected, consider increasing
```
To help resolve this we reset dynamo at setup for all unittests. Let's
see if this helps
Differential Revision: D70329890
Pull Request resolved: pytorch#8772
### Summary
We seem to be using a combination of CMAKE_ARGS and environment
variables when creating wheels. Ultimately, CMake only uses the cmake
args, however we redefine some of these flags as env vars to help
`setup.py` determine if a certain feature is turned on. Specifically, it
looks for pybinding vars to bundle pybindings.
Let's remove this redundancy and just use the CMAKE_ARGS as the single
source of truth. For more details and other considerations, see
pytorch#9494 (abandoned).
Note that even in the wheel building jobs, we use cmake args instead of
environment variables to control features:
https://github.com/pytorch/executorch/blob/644b7ddf14180d97e348faa627f576e13d367d69/.ci/scripts/wheel/envvar_base.sh#L20https://github.com/pytorch/executorch/blob/644b7ddf14180d97e348faa627f576e13d367d69/.ci/scripts/wheel/envvar_macos.sh#L14-L15
### Test plan
build + check CMakeCache.txt to ensure flags are set
```bash
# Expected: EXECUTORCH_BUILD_PYBIND=OFF EXECUTORCH_BUILD_XNNPACK=OFF EXECUTORCH_BUILD_COREML=OFF
$ rm -rf pip-out dist && ./install_executorch.sh --pybind off
# Expected: EXECUTORCH_BUILD_PYBIND=ON EXECUTORCH_BUILD_XNNPACK=ON EXECUTORCH_BUILD_COREML=OFF
$ rm -rf pip-out dist && ./install_executorch.sh
# Expected: EXECUTORCH_BUILD_PYBIND=ON EXECUTORCH_BUILD_XNNPACK=OFF EXECUTORCH_BUILD_COREML=ON
$ rm -rf pip-out dist && ./install_executorch.sh --pybind coreml
# Expected: EXECUTORCH_BUILD_PYBIND=ON EXECUTORCH_BUILD_XNNPACK=ON EXECUTORCH_BUILD_COREML=ON
$ rm -rf pip-out dist && ./install_executorch.sh --pybind xnnpack coreml
# Throws an error
$ rm -rf pip-out dist && ./install_executorch.sh --pybind coreml off
```
cc @larryliu0820@lucylq
Differential Revision: D71752746
Pull Request resolved: pytorch#9597
Differential Revision: D71752743
Pull Request resolved: pytorch#9606
Differential Revision: D71752747
Pull Request resolved: pytorch#9608
Differential Revision: D71752748
Pull Request resolved: pytorch#9609
### Summary
We want `EXECUTORCH_BUILD_PYBIND` enabled if the user wants to build the
bindings — so let's just do it. Unless of course, they explicitly choose
not to by defining the arg themselves.
### Test plan
CI
cc @larryliu0820@lucylq
Summary: A few SoCs have been supported recently, updated the
documentation.
Differential Revision: D71827272
cc @mergennachin@byjlw
Number of delegates and tolerance has changed so update it.
This was (accidentally) removed at a refactoring. Also takes the chance
to use the new XfailIf.. decorator.
Signed-off-by: Erik Lundell <erik.lundell@arm.com>
…rch#9158)
### Summary
Add a proxy for an `export_llama` performance regression test by
comparing the ops in the graph before and after the PR. The export
happens without loading a checkpoint or params file, which means that
all of the base `ModelArgs` values for `llama_transformer` will be used.
### Test plan
N/A
Differential Revision: D71833608
Pull Request resolved: pytorch#9603
Add conv3d tests, though most are skipped since conv3d support is not
yet implemented.
Signed-off-by: Erik Lundell <erik.lundell@arm.com>
Seems to succeed in rare instances due to randomness.
Signed-off-by: Erik Lundell <erik.lundell@arm.com>
### Summary
This is the stage 1 of Mimi Enablement.
Stage 2 will consist of actual model enablement.
- Support OP:
- exp
- expm1
- elu
- transpose conv1d
- bitwise_and
- scalar_tensor
- stack
- unbind
Summary:
To support embedded system builds which threw an error on the warning
for left shift by 32 on a 32 bit dtype, the code was modified to:
```
memory_offset |= static_cast<size_t>(memory_offset_high)
<< (sizeof(size_t) - sizeof(uint32_t));
```
This fails for build of OSS qwen example however.
Instead, we modify to add a check for
```
sizeof(size_t) > sizeof(uint32_t)
```
in the conditional instead of changing the computation.
In our builds of interest, this compiles away the if branch
Reviewed By: digantdesai, dpalmasan
Differential Revision: D71488571
### Summary
Context: pytorch#9481
* Include the `executorchcoreml` pybinding in the builds
* Remove separate installation option
* Turn on CoreML by default for macOS builds
* Add a dependency on coremltools for macOS
### Test plan
CI
```
$ rm -rf cmake-out pip-out dist && ./install_executorch.sh
$ ./examples/models/llama/install_requirements.sh
$ .ci/scripts/test_llama.sh -model stories110M -build_tool cmake -dtype fp32 -mode coreml
$ .ci/scripts/test_llama.sh -model stories110M -build_tool cmake -dtype fp32 -mode xnnpack+custom+quantize_kv
```
cc @larryliu0820@lucylq
…ntization for Example Models (pytorch#9634)
### Summary
Changes:
1. When initializing Llama2 for aot_compiler, since checkpoints can only
e downloaded from hugging face, we initialize llama2 with uninitialized
weights. The problem with this is that when running quantization, we can
run into errors with the histogram if the unitialized values are nan. We
fix this by initializing the weights with zeros if no check point is
provided. This enforces that quantization step can still work.
2. Quant Type in AoT compiler. When looking at the model options
available to XNNPACK, everything is quantized with per-tensor static
quantization. This isn't the best option for all the models available.
For example transformer based models like Llama and MobileBert would
likely prefer dynamically quantized per channel weights, where has CNN
like MobileNet would prefer statically quantized per channel weights. We
add this type of Quant Type to the existing models options. This also
helps with Test Timeouts. per-tensor static quantization on a model like
llama can take a long time due to the introduction of MANY q/dq nodes,
and the complex partitions it creates. As a result, proposing partitions
can take a long time due to the constant BFS to find the largest
possible partition. By specifying the more apt quantization scheme like
dynamic per-channel quantization, we can avoid this complexity.
Overall this should help with flakey [nan, nan] errors in the
quantization histogram, and it should also help with CI timing out.
### Test plan
OSS XNNPACK CI for all model delegation
cc @digantdesai@cbilgin
### Summary
* After pytorch#9483, we should have
CoreML support out of the box for macOS
* Unfortunately, we still need
`backends/apple/coreml/scripts/install_requirements.sh` to use the
[coreml_executorch_runner](https://github.com/pytorch/executorch/tree/main/examples/apple/coreml/executor_runner)
(used for testing)
* I should have caught all the usage
### Test plan
Read
swolchokand others added 11 commits March 26, 2025 13:26
I have no idea what this file actually does, but it seems like we are
supposed to have this?
…ch#9509)
Disable one, fix the other.
Testing: built internally
pytorch#9511)
I planned to do this everywhere and forgot. Clean it all up, leave a
note, enforce the note with visibility. This makes sure everything in
buck-land gets ET_USE_THREADPOOL.
Test Plan: Profiled run on internal model, no longer seeing
parallel_for_no_threadpool
…AME_AS_COMPUTE (pytorch#9613)
As the title says, this is mostly a few related find-replaces, plus
marking SupportedTensorDtypes::SAME_AS_COMPUTE deprecated.
As the code comment says, these APIs are undergoing development (see
e.g. pytorch#9613) and it's pretty
inconvenient that they're incidentally committed-to externally. Mark
them deprecated so we have the option to drop that commitment in (IIUC)
0.7.
Previous attempt to bump HF transformers version to latest is reverted
due to llava model imcompatibility. This PR is to just ensure the CI are
able to test `optimum-executorch` with latest version of HF transformers
and upcoming `executorch==0.6.0`.
Note: This change is purely on CI and only for optimum-executorch,
should not affect other models like llava.
Co-authored-by: Guang Yang <guangyang@fb.com>
Differential Revision: D69994481
Pull Request resolved: pytorch#8703
@ynimmaga
ynimmaga merged commit 3f53cc2 into ynimmaga:openvino_backendMar 27, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

18 participants

@cavusmustafa@ynimmaga@metascroy@guangy10@shoumikhin@kirklandsign@jathu@mcr229@jackzhxng@larryliu0820@cccclai@mansnils@Erik-Lundell@winskuo-quic@JakeStevens@swolchok@bigfootjon@mergennachin
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Pybind merge fix - #42

Merged
ynimmaga merged 41 commits into
ynimmaga:openvino_backendfrom
cavusmustafa:pybind_merge_fix
Mar 27, 2025
Merged

Pybind merge fix#42
ynimmaga merged 41 commits into
ynimmaga:openvino_backendfrom
cavusmustafa:pybind_merge_fix

Conversation

@cavusmustafa

@cavusmustafacavusmustafa commented Mar 27, 2025

Copy link
Copy Markdown
Collaborator
  • Merged latest executorch main into openvino_backend
  • Updated setup.py for latest pybind build updates

metascroyand others added 30 commits March 25, 2025 11:39
ETCoreML crashes in FB app. During debugging, we traced issue down to
this.
cc @kimishpatel@YifanShenSZ@cymbalrush
HF version bump. Ensure `optimum-executorch` can work on new
`transformers` models with `executorch==0.6.0`
### Test plan
CI to test HF models
Co-authored-by: Guang Yang <guangyang@fb.com>
Differential Revision: D71761219
Pull Request resolved: pytorch#9554
Summary: .
Reviewed By: bsoyluoglu
Differential Revision: D71752749
### Summary
Seeing this error in Linux wheel building jobs:
```
Collecting numpy (from torchvision==0.22.0.dev20250311)
Downloading numpy-2.2.4-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.metadata (62 kB)
INFO: pip is looking at multiple versions of torchvision to determine which version is compatible with other requirements. This could take a while.
The conflict is caused by:
The user requested torch==2.7.0.dev20250311
torchvision 0.22.0.dev20250311+cpu depends on torch==2.7.0.dev20250310
``` ### Test plan
CI
https://github.com/pytorch/executorch/actions/runs/14047575373/job/39331644423
There seems to be some CI issues with:
```
torch._dynamo.exc.FailOnRecompileLimitHit: recompile_limit reached with one_graph=True. Excessive recompilations can degrade performance due to the compilation overhead of each recompilation. To monitor recompilations, enable TORCH_LOGS=recompiles. If recompilations are expected, consider increasing
```
To help resolve this we reset dynamo at setup for all unittests. Let's
see if this helps
Differential Revision: D70329890
Pull Request resolved: pytorch#8772
### Summary
We seem to be using a combination of CMAKE_ARGS and environment
variables when creating wheels. Ultimately, CMake only uses the cmake
args, however we redefine some of these flags as env vars to help
`setup.py` determine if a certain feature is turned on. Specifically, it
looks for pybinding vars to bundle pybindings.
Let's remove this redundancy and just use the CMAKE_ARGS as the single
source of truth. For more details and other considerations, see
pytorch#9494 (abandoned).
Note that even in the wheel building jobs, we use cmake args instead of
environment variables to control features:
https://github.com/pytorch/executorch/blob/644b7ddf14180d97e348faa627f576e13d367d69/.ci/scripts/wheel/envvar_base.sh#L20https://github.com/pytorch/executorch/blob/644b7ddf14180d97e348faa627f576e13d367d69/.ci/scripts/wheel/envvar_macos.sh#L14-L15
### Test plan
build + check CMakeCache.txt to ensure flags are set
```bash
# Expected: EXECUTORCH_BUILD_PYBIND=OFF EXECUTORCH_BUILD_XNNPACK=OFF EXECUTORCH_BUILD_COREML=OFF
$ rm -rf pip-out dist && ./install_executorch.sh --pybind off
# Expected: EXECUTORCH_BUILD_PYBIND=ON EXECUTORCH_BUILD_XNNPACK=ON EXECUTORCH_BUILD_COREML=OFF
$ rm -rf pip-out dist && ./install_executorch.sh
# Expected: EXECUTORCH_BUILD_PYBIND=ON EXECUTORCH_BUILD_XNNPACK=OFF EXECUTORCH_BUILD_COREML=ON
$ rm -rf pip-out dist && ./install_executorch.sh --pybind coreml
# Expected: EXECUTORCH_BUILD_PYBIND=ON EXECUTORCH_BUILD_XNNPACK=ON EXECUTORCH_BUILD_COREML=ON
$ rm -rf pip-out dist && ./install_executorch.sh --pybind xnnpack coreml
# Throws an error
$ rm -rf pip-out dist && ./install_executorch.sh --pybind coreml off
```
cc @larryliu0820@lucylq
Differential Revision: D71752746
Pull Request resolved: pytorch#9597
Differential Revision: D71752743
Pull Request resolved: pytorch#9606
Differential Revision: D71752747
Pull Request resolved: pytorch#9608
Differential Revision: D71752748
Pull Request resolved: pytorch#9609
### Summary
We want `EXECUTORCH_BUILD_PYBIND` enabled if the user wants to build the
bindings — so let's just do it. Unless of course, they explicitly choose
not to by defining the arg themselves.
### Test plan
CI
cc @larryliu0820@lucylq
Summary: A few SoCs have been supported recently, updated the
documentation.
Differential Revision: D71827272
cc @mergennachin@byjlw
Number of delegates and tolerance has changed so update it.
This was (accidentally) removed at a refactoring. Also takes the chance
to use the new XfailIf.. decorator.
Signed-off-by: Erik Lundell <erik.lundell@arm.com>
…rch#9158)
### Summary
Add a proxy for an `export_llama` performance regression test by
comparing the ops in the graph before and after the PR. The export
happens without loading a checkpoint or params file, which means that
all of the base `ModelArgs` values for `llama_transformer` will be used.
### Test plan
N/A
Differential Revision: D71833608
Pull Request resolved: pytorch#9603
Add conv3d tests, though most are skipped since conv3d support is not
yet implemented.
Signed-off-by: Erik Lundell <erik.lundell@arm.com>
Seems to succeed in rare instances due to randomness.
Signed-off-by: Erik Lundell <erik.lundell@arm.com>
### Summary
This is the stage 1 of Mimi Enablement.
Stage 2 will consist of actual model enablement.
- Support OP:
- exp
- expm1
- elu
- transpose conv1d
- bitwise_and
- scalar_tensor
- stack
- unbind
Summary:
To support embedded system builds which threw an error on the warning
for left shift by 32 on a 32 bit dtype, the code was modified to:
```
memory_offset |= static_cast<size_t>(memory_offset_high)
<< (sizeof(size_t) - sizeof(uint32_t));
```
This fails for build of OSS qwen example however.
Instead, we modify to add a check for
```
sizeof(size_t) > sizeof(uint32_t)
```
in the conditional instead of changing the computation.
In our builds of interest, this compiles away the if branch
Reviewed By: digantdesai, dpalmasan
Differential Revision: D71488571
### Summary
Context: pytorch#9481
* Include the `executorchcoreml` pybinding in the builds
* Remove separate installation option
* Turn on CoreML by default for macOS builds
* Add a dependency on coremltools for macOS
### Test plan
CI
```
$ rm -rf cmake-out pip-out dist && ./install_executorch.sh
$ ./examples/models/llama/install_requirements.sh
$ .ci/scripts/test_llama.sh -model stories110M -build_tool cmake -dtype fp32 -mode coreml
$ .ci/scripts/test_llama.sh -model stories110M -build_tool cmake -dtype fp32 -mode xnnpack+custom+quantize_kv
```
cc @larryliu0820@lucylq
…ntization for Example Models (pytorch#9634)
### Summary
Changes:
1. When initializing Llama2 for aot_compiler, since checkpoints can only
e downloaded from hugging face, we initialize llama2 with uninitialized
weights. The problem with this is that when running quantization, we can
run into errors with the histogram if the unitialized values are nan. We
fix this by initializing the weights with zeros if no check point is
provided. This enforces that quantization step can still work.
2. Quant Type in AoT compiler. When looking at the model options
available to XNNPACK, everything is quantized with per-tensor static
quantization. This isn't the best option for all the models available.
For example transformer based models like Llama and MobileBert would
likely prefer dynamically quantized per channel weights, where has CNN
like MobileNet would prefer statically quantized per channel weights. We
add this type of Quant Type to the existing models options. This also
helps with Test Timeouts. per-tensor static quantization on a model like
llama can take a long time due to the introduction of MANY q/dq nodes,
and the complex partitions it creates. As a result, proposing partitions
can take a long time due to the constant BFS to find the largest
possible partition. By specifying the more apt quantization scheme like
dynamic per-channel quantization, we can avoid this complexity.
Overall this should help with flakey [nan, nan] errors in the
quantization histogram, and it should also help with CI timing out.
### Test plan
OSS XNNPACK CI for all model delegation
cc @digantdesai@cbilgin
### Summary
* After pytorch#9483, we should have
CoreML support out of the box for macOS
* Unfortunately, we still need
`backends/apple/coreml/scripts/install_requirements.sh` to use the
[coreml_executorch_runner](https://github.com/pytorch/executorch/tree/main/examples/apple/coreml/executor_runner)
(used for testing)
* I should have caught all the usage
### Test plan
Read
swolchokand others added 11 commits March 26, 2025 13:26
I have no idea what this file actually does, but it seems like we are
supposed to have this?
…ch#9509)
Disable one, fix the other.
Testing: built internally
pytorch#9511)
I planned to do this everywhere and forgot. Clean it all up, leave a
note, enforce the note with visibility. This makes sure everything in
buck-land gets ET_USE_THREADPOOL.
Test Plan: Profiled run on internal model, no longer seeing
parallel_for_no_threadpool
…AME_AS_COMPUTE (pytorch#9613)
As the title says, this is mostly a few related find-replaces, plus
marking SupportedTensorDtypes::SAME_AS_COMPUTE deprecated.
As the code comment says, these APIs are undergoing development (see
e.g. pytorch#9613) and it's pretty
inconvenient that they're incidentally committed-to externally. Mark
them deprecated so we have the option to drop that commitment in (IIUC)
0.7.
Previous attempt to bump HF transformers version to latest is reverted
due to llava model imcompatibility. This PR is to just ensure the CI are
able to test `optimum-executorch` with latest version of HF transformers
and upcoming `executorch==0.6.0`.
Note: This change is purely on CI and only for optimum-executorch,
should not affect other models like llava.
Co-authored-by: Guang Yang <guangyang@fb.com>
Differential Revision: D69994481
Pull Request resolved: pytorch#8703
@ynimmaga
ynimmaga merged commit 3f53cc2 into ynimmaga:openvino_backendMar 27, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

18 participants

@cavusmustafa@ynimmaga@metascroy@guangy10@shoumikhin@kirklandsign@jathu@mcr229@jackzhxng@larryliu0820@cccclai@mansnils@Erik-Lundell@winskuo-quic@JakeStevens@swolchok@bigfootjon@mergennachin
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Pybind merge fix - #42

Merged
ynimmaga merged 41 commits into
ynimmaga:openvino_backendfrom
cavusmustafa:pybind_merge_fix
Mar 27, 2025
Merged

Pybind merge fix#42
ynimmaga merged 41 commits into
ynimmaga:openvino_backendfrom
cavusmustafa:pybind_merge_fix

Conversation

@cavusmustafa

@cavusmustafacavusmustafa commented Mar 27, 2025

Copy link
Copy Markdown
Collaborator
  • Merged latest executorch main into openvino_backend
  • Updated setup.py for latest pybind build updates

metascroyand others added 30 commits March 25, 2025 11:39
ETCoreML crashes in FB app. During debugging, we traced issue down to
this.
cc @kimishpatel@YifanShenSZ@cymbalrush
HF version bump. Ensure `optimum-executorch` can work on new
`transformers` models with `executorch==0.6.0`
### Test plan
CI to test HF models
Co-authored-by: Guang Yang <guangyang@fb.com>
Differential Revision: D71761219
Pull Request resolved: pytorch#9554
Summary: .
Reviewed By: bsoyluoglu
Differential Revision: D71752749
### Summary
Seeing this error in Linux wheel building jobs:
```
Collecting numpy (from torchvision==0.22.0.dev20250311)
Downloading numpy-2.2.4-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.metadata (62 kB)
INFO: pip is looking at multiple versions of torchvision to determine which version is compatible with other requirements. This could take a while.
The conflict is caused by:
The user requested torch==2.7.0.dev20250311
torchvision 0.22.0.dev20250311+cpu depends on torch==2.7.0.dev20250310
``` ### Test plan
CI
https://github.com/pytorch/executorch/actions/runs/14047575373/job/39331644423
There seems to be some CI issues with:
```
torch._dynamo.exc.FailOnRecompileLimitHit: recompile_limit reached with one_graph=True. Excessive recompilations can degrade performance due to the compilation overhead of each recompilation. To monitor recompilations, enable TORCH_LOGS=recompiles. If recompilations are expected, consider increasing
```
To help resolve this we reset dynamo at setup for all unittests. Let's
see if this helps
Differential Revision: D70329890
Pull Request resolved: pytorch#8772
### Summary
We seem to be using a combination of CMAKE_ARGS and environment
variables when creating wheels. Ultimately, CMake only uses the cmake
args, however we redefine some of these flags as env vars to help
`setup.py` determine if a certain feature is turned on. Specifically, it
looks for pybinding vars to bundle pybindings.
Let's remove this redundancy and just use the CMAKE_ARGS as the single
source of truth. For more details and other considerations, see
pytorch#9494 (abandoned).
Note that even in the wheel building jobs, we use cmake args instead of
environment variables to control features:
https://github.com/pytorch/executorch/blob/644b7ddf14180d97e348faa627f576e13d367d69/.ci/scripts/wheel/envvar_base.sh#L20https://github.com/pytorch/executorch/blob/644b7ddf14180d97e348faa627f576e13d367d69/.ci/scripts/wheel/envvar_macos.sh#L14-L15
### Test plan
build + check CMakeCache.txt to ensure flags are set
```bash
# Expected: EXECUTORCH_BUILD_PYBIND=OFF EXECUTORCH_BUILD_XNNPACK=OFF EXECUTORCH_BUILD_COREML=OFF
$ rm -rf pip-out dist && ./install_executorch.sh --pybind off
# Expected: EXECUTORCH_BUILD_PYBIND=ON EXECUTORCH_BUILD_XNNPACK=ON EXECUTORCH_BUILD_COREML=OFF
$ rm -rf pip-out dist && ./install_executorch.sh
# Expected: EXECUTORCH_BUILD_PYBIND=ON EXECUTORCH_BUILD_XNNPACK=OFF EXECUTORCH_BUILD_COREML=ON
$ rm -rf pip-out dist && ./install_executorch.sh --pybind coreml
# Expected: EXECUTORCH_BUILD_PYBIND=ON EXECUTORCH_BUILD_XNNPACK=ON EXECUTORCH_BUILD_COREML=ON
$ rm -rf pip-out dist && ./install_executorch.sh --pybind xnnpack coreml
# Throws an error
$ rm -rf pip-out dist && ./install_executorch.sh --pybind coreml off
```
cc @larryliu0820@lucylq
Differential Revision: D71752746
Pull Request resolved: pytorch#9597
Differential Revision: D71752743
Pull Request resolved: pytorch#9606
Differential Revision: D71752747
Pull Request resolved: pytorch#9608
Differential Revision: D71752748
Pull Request resolved: pytorch#9609
### Summary
We want `EXECUTORCH_BUILD_PYBIND` enabled if the user wants to build the
bindings — so let's just do it. Unless of course, they explicitly choose
not to by defining the arg themselves.
### Test plan
CI
cc @larryliu0820@lucylq
Summary: A few SoCs have been supported recently, updated the
documentation.
Differential Revision: D71827272
cc @mergennachin@byjlw
Number of delegates and tolerance has changed so update it.
This was (accidentally) removed at a refactoring. Also takes the chance
to use the new XfailIf.. decorator.
Signed-off-by: Erik Lundell <erik.lundell@arm.com>
…rch#9158)
### Summary
Add a proxy for an `export_llama` performance regression test by
comparing the ops in the graph before and after the PR. The export
happens without loading a checkpoint or params file, which means that
all of the base `ModelArgs` values for `llama_transformer` will be used.
### Test plan
N/A
Differential Revision: D71833608
Pull Request resolved: pytorch#9603
Add conv3d tests, though most are skipped since conv3d support is not
yet implemented.
Signed-off-by: Erik Lundell <erik.lundell@arm.com>
Seems to succeed in rare instances due to randomness.
Signed-off-by: Erik Lundell <erik.lundell@arm.com>
### Summary
This is the stage 1 of Mimi Enablement.
Stage 2 will consist of actual model enablement.
- Support OP:
- exp
- expm1
- elu
- transpose conv1d
- bitwise_and
- scalar_tensor
- stack
- unbind
Summary:
To support embedded system builds which threw an error on the warning
for left shift by 32 on a 32 bit dtype, the code was modified to:
```
memory_offset |= static_cast<size_t>(memory_offset_high)
<< (sizeof(size_t) - sizeof(uint32_t));
```
This fails for build of OSS qwen example however.
Instead, we modify to add a check for
```
sizeof(size_t) > sizeof(uint32_t)
```
in the conditional instead of changing the computation.
In our builds of interest, this compiles away the if branch
Reviewed By: digantdesai, dpalmasan
Differential Revision: D71488571
### Summary
Context: pytorch#9481
* Include the `executorchcoreml` pybinding in the builds
* Remove separate installation option
* Turn on CoreML by default for macOS builds
* Add a dependency on coremltools for macOS
### Test plan
CI
```
$ rm -rf cmake-out pip-out dist && ./install_executorch.sh
$ ./examples/models/llama/install_requirements.sh
$ .ci/scripts/test_llama.sh -model stories110M -build_tool cmake -dtype fp32 -mode coreml
$ .ci/scripts/test_llama.sh -model stories110M -build_tool cmake -dtype fp32 -mode xnnpack+custom+quantize_kv
```
cc @larryliu0820@lucylq
…ntization for Example Models (pytorch#9634)
### Summary
Changes:
1. When initializing Llama2 for aot_compiler, since checkpoints can only
e downloaded from hugging face, we initialize llama2 with uninitialized
weights. The problem with this is that when running quantization, we can
run into errors with the histogram if the unitialized values are nan. We
fix this by initializing the weights with zeros if no check point is
provided. This enforces that quantization step can still work.
2. Quant Type in AoT compiler. When looking at the model options
available to XNNPACK, everything is quantized with per-tensor static
quantization. This isn't the best option for all the models available.
For example transformer based models like Llama and MobileBert would
likely prefer dynamically quantized per channel weights, where has CNN
like MobileNet would prefer statically quantized per channel weights. We
add this type of Quant Type to the existing models options. This also
helps with Test Timeouts. per-tensor static quantization on a model like
llama can take a long time due to the introduction of MANY q/dq nodes,
and the complex partitions it creates. As a result, proposing partitions
can take a long time due to the constant BFS to find the largest
possible partition. By specifying the more apt quantization scheme like
dynamic per-channel quantization, we can avoid this complexity.
Overall this should help with flakey [nan, nan] errors in the
quantization histogram, and it should also help with CI timing out.
### Test plan
OSS XNNPACK CI for all model delegation
cc @digantdesai@cbilgin
### Summary
* After pytorch#9483, we should have
CoreML support out of the box for macOS
* Unfortunately, we still need
`backends/apple/coreml/scripts/install_requirements.sh` to use the
[coreml_executorch_runner](https://github.com/pytorch/executorch/tree/main/examples/apple/coreml/executor_runner)
(used for testing)
* I should have caught all the usage
### Test plan
Read
swolchokand others added 11 commits March 26, 2025 13:26
I have no idea what this file actually does, but it seems like we are
supposed to have this?
…ch#9509)
Disable one, fix the other.
Testing: built internally
pytorch#9511)
I planned to do this everywhere and forgot. Clean it all up, leave a
note, enforce the note with visibility. This makes sure everything in
buck-land gets ET_USE_THREADPOOL.
Test Plan: Profiled run on internal model, no longer seeing
parallel_for_no_threadpool
…AME_AS_COMPUTE (pytorch#9613)
As the title says, this is mostly a few related find-replaces, plus
marking SupportedTensorDtypes::SAME_AS_COMPUTE deprecated.
As the code comment says, these APIs are undergoing development (see
e.g. pytorch#9613) and it's pretty
inconvenient that they're incidentally committed-to externally. Mark
them deprecated so we have the option to drop that commitment in (IIUC)
0.7.
Previous attempt to bump HF transformers version to latest is reverted
due to llava model imcompatibility. This PR is to just ensure the CI are
able to test `optimum-executorch` with latest version of HF transformers
and upcoming `executorch==0.6.0`.
Note: This change is purely on CI and only for optimum-executorch,
should not affect other models like llava.
Co-authored-by: Guang Yang <guangyang@fb.com>
Differential Revision: D69994481
Pull Request resolved: pytorch#8703
@ynimmaga
ynimmaga merged commit 3f53cc2 into ynimmaga:openvino_backendMar 27, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

18 participants

@cavusmustafa@ynimmaga@metascroy@guangy10@shoumikhin@kirklandsign@jathu@mcr229@jackzhxng@larryliu0820@cccclai@mansnils@Erik-Lundell@winskuo-quic@JakeStevens@swolchok@bigfootjon@mergennachin
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Pybind merge fix - #42

Merged
ynimmaga merged 41 commits into
ynimmaga:openvino_backendfrom
cavusmustafa:pybind_merge_fix
Mar 27, 2025
Merged

Pybind merge fix#42
ynimmaga merged 41 commits into
ynimmaga:openvino_backendfrom
cavusmustafa:pybind_merge_fix

Conversation

@cavusmustafa

@cavusmustafacavusmustafa commented Mar 27, 2025

Copy link
Copy Markdown
Collaborator
  • Merged latest executorch main into openvino_backend
  • Updated setup.py for latest pybind build updates

metascroyand others added 30 commits March 25, 2025 11:39
ETCoreML crashes in FB app. During debugging, we traced issue down to
this.
cc @kimishpatel@YifanShenSZ@cymbalrush
HF version bump. Ensure `optimum-executorch` can work on new
`transformers` models with `executorch==0.6.0`
### Test plan
CI to test HF models
Co-authored-by: Guang Yang <guangyang@fb.com>
Differential Revision: D71761219
Pull Request resolved: pytorch#9554
Summary: .
Reviewed By: bsoyluoglu
Differential Revision: D71752749
### Summary
Seeing this error in Linux wheel building jobs:
```
Collecting numpy (from torchvision==0.22.0.dev20250311)
Downloading numpy-2.2.4-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.metadata (62 kB)
INFO: pip is looking at multiple versions of torchvision to determine which version is compatible with other requirements. This could take a while.
The conflict is caused by:
The user requested torch==2.7.0.dev20250311
torchvision 0.22.0.dev20250311+cpu depends on torch==2.7.0.dev20250310
``` ### Test plan
CI
https://github.com/pytorch/executorch/actions/runs/14047575373/job/39331644423
There seems to be some CI issues with:
```
torch._dynamo.exc.FailOnRecompileLimitHit: recompile_limit reached with one_graph=True. Excessive recompilations can degrade performance due to the compilation overhead of each recompilation. To monitor recompilations, enable TORCH_LOGS=recompiles. If recompilations are expected, consider increasing
```
To help resolve this we reset dynamo at setup for all unittests. Let's
see if this helps
Differential Revision: D70329890
Pull Request resolved: pytorch#8772
### Summary
We seem to be using a combination of CMAKE_ARGS and environment
variables when creating wheels. Ultimately, CMake only uses the cmake
args, however we redefine some of these flags as env vars to help
`setup.py` determine if a certain feature is turned on. Specifically, it
looks for pybinding vars to bundle pybindings.
Let's remove this redundancy and just use the CMAKE_ARGS as the single
source of truth. For more details and other considerations, see
pytorch#9494 (abandoned).
Note that even in the wheel building jobs, we use cmake args instead of
environment variables to control features:
https://github.com/pytorch/executorch/blob/644b7ddf14180d97e348faa627f576e13d367d69/.ci/scripts/wheel/envvar_base.sh#L20https://github.com/pytorch/executorch/blob/644b7ddf14180d97e348faa627f576e13d367d69/.ci/scripts/wheel/envvar_macos.sh#L14-L15
### Test plan
build + check CMakeCache.txt to ensure flags are set
```bash
# Expected: EXECUTORCH_BUILD_PYBIND=OFF EXECUTORCH_BUILD_XNNPACK=OFF EXECUTORCH_BUILD_COREML=OFF
$ rm -rf pip-out dist && ./install_executorch.sh --pybind off
# Expected: EXECUTORCH_BUILD_PYBIND=ON EXECUTORCH_BUILD_XNNPACK=ON EXECUTORCH_BUILD_COREML=OFF
$ rm -rf pip-out dist && ./install_executorch.sh
# Expected: EXECUTORCH_BUILD_PYBIND=ON EXECUTORCH_BUILD_XNNPACK=OFF EXECUTORCH_BUILD_COREML=ON
$ rm -rf pip-out dist && ./install_executorch.sh --pybind coreml
# Expected: EXECUTORCH_BUILD_PYBIND=ON EXECUTORCH_BUILD_XNNPACK=ON EXECUTORCH_BUILD_COREML=ON
$ rm -rf pip-out dist && ./install_executorch.sh --pybind xnnpack coreml
# Throws an error
$ rm -rf pip-out dist && ./install_executorch.sh --pybind coreml off
```
cc @larryliu0820@lucylq
Differential Revision: D71752746
Pull Request resolved: pytorch#9597
Differential Revision: D71752743
Pull Request resolved: pytorch#9606
Differential Revision: D71752747
Pull Request resolved: pytorch#9608
Differential Revision: D71752748
Pull Request resolved: pytorch#9609
### Summary
We want `EXECUTORCH_BUILD_PYBIND` enabled if the user wants to build the
bindings — so let's just do it. Unless of course, they explicitly choose
not to by defining the arg themselves.
### Test plan
CI
cc @larryliu0820@lucylq
Summary: A few SoCs have been supported recently, updated the
documentation.
Differential Revision: D71827272
cc @mergennachin@byjlw
Number of delegates and tolerance has changed so update it.
This was (accidentally) removed at a refactoring. Also takes the chance
to use the new XfailIf.. decorator.
Signed-off-by: Erik Lundell <erik.lundell@arm.com>
…rch#9158)
### Summary
Add a proxy for an `export_llama` performance regression test by
comparing the ops in the graph before and after the PR. The export
happens without loading a checkpoint or params file, which means that
all of the base `ModelArgs` values for `llama_transformer` will be used.
### Test plan
N/A
Differential Revision: D71833608
Pull Request resolved: pytorch#9603
Add conv3d tests, though most are skipped since conv3d support is not
yet implemented.
Signed-off-by: Erik Lundell <erik.lundell@arm.com>
Seems to succeed in rare instances due to randomness.
Signed-off-by: Erik Lundell <erik.lundell@arm.com>
### Summary
This is the stage 1 of Mimi Enablement.
Stage 2 will consist of actual model enablement.
- Support OP:
- exp
- expm1
- elu
- transpose conv1d
- bitwise_and
- scalar_tensor
- stack
- unbind
Summary:
To support embedded system builds which threw an error on the warning
for left shift by 32 on a 32 bit dtype, the code was modified to:
```
memory_offset |= static_cast<size_t>(memory_offset_high)
<< (sizeof(size_t) - sizeof(uint32_t));
```
This fails for build of OSS qwen example however.
Instead, we modify to add a check for
```
sizeof(size_t) > sizeof(uint32_t)
```
in the conditional instead of changing the computation.
In our builds of interest, this compiles away the if branch
Reviewed By: digantdesai, dpalmasan
Differential Revision: D71488571
### Summary
Context: pytorch#9481
* Include the `executorchcoreml` pybinding in the builds
* Remove separate installation option
* Turn on CoreML by default for macOS builds
* Add a dependency on coremltools for macOS
### Test plan
CI
```
$ rm -rf cmake-out pip-out dist && ./install_executorch.sh
$ ./examples/models/llama/install_requirements.sh
$ .ci/scripts/test_llama.sh -model stories110M -build_tool cmake -dtype fp32 -mode coreml
$ .ci/scripts/test_llama.sh -model stories110M -build_tool cmake -dtype fp32 -mode xnnpack+custom+quantize_kv
```
cc @larryliu0820@lucylq
…ntization for Example Models (pytorch#9634)
### Summary
Changes:
1. When initializing Llama2 for aot_compiler, since checkpoints can only
e downloaded from hugging face, we initialize llama2 with uninitialized
weights. The problem with this is that when running quantization, we can
run into errors with the histogram if the unitialized values are nan. We
fix this by initializing the weights with zeros if no check point is
provided. This enforces that quantization step can still work.
2. Quant Type in AoT compiler. When looking at the model options
available to XNNPACK, everything is quantized with per-tensor static
quantization. This isn't the best option for all the models available.
For example transformer based models like Llama and MobileBert would
likely prefer dynamically quantized per channel weights, where has CNN
like MobileNet would prefer statically quantized per channel weights. We
add this type of Quant Type to the existing models options. This also
helps with Test Timeouts. per-tensor static quantization on a model like
llama can take a long time due to the introduction of MANY q/dq nodes,
and the complex partitions it creates. As a result, proposing partitions
can take a long time due to the constant BFS to find the largest
possible partition. By specifying the more apt quantization scheme like
dynamic per-channel quantization, we can avoid this complexity.
Overall this should help with flakey [nan, nan] errors in the
quantization histogram, and it should also help with CI timing out.
### Test plan
OSS XNNPACK CI for all model delegation
cc @digantdesai@cbilgin
### Summary
* After pytorch#9483, we should have
CoreML support out of the box for macOS
* Unfortunately, we still need
`backends/apple/coreml/scripts/install_requirements.sh` to use the
[coreml_executorch_runner](https://github.com/pytorch/executorch/tree/main/examples/apple/coreml/executor_runner)
(used for testing)
* I should have caught all the usage
### Test plan
Read
swolchokand others added 11 commits March 26, 2025 13:26
I have no idea what this file actually does, but it seems like we are
supposed to have this?
…ch#9509)
Disable one, fix the other.
Testing: built internally
pytorch#9511)
I planned to do this everywhere and forgot. Clean it all up, leave a
note, enforce the note with visibility. This makes sure everything in
buck-land gets ET_USE_THREADPOOL.
Test Plan: Profiled run on internal model, no longer seeing
parallel_for_no_threadpool
…AME_AS_COMPUTE (pytorch#9613)
As the title says, this is mostly a few related find-replaces, plus
marking SupportedTensorDtypes::SAME_AS_COMPUTE deprecated.
As the code comment says, these APIs are undergoing development (see
e.g. pytorch#9613) and it's pretty
inconvenient that they're incidentally committed-to externally. Mark
them deprecated so we have the option to drop that commitment in (IIUC)
0.7.
Previous attempt to bump HF transformers version to latest is reverted
due to llava model imcompatibility. This PR is to just ensure the CI are
able to test `optimum-executorch` with latest version of HF transformers
and upcoming `executorch==0.6.0`.
Note: This change is purely on CI and only for optimum-executorch,
should not affect other models like llava.
Co-authored-by: Guang Yang <guangyang@fb.com>
Differential Revision: D69994481
Pull Request resolved: pytorch#8703
@ynimmaga
ynimmaga merged commit 3f53cc2 into ynimmaga:openvino_backendMar 27, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

18 participants

@cavusmustafa@ynimmaga@metascroy@guangy10@shoumikhin@kirklandsign@jathu@mcr229@jackzhxng@larryliu0820@cccclai@mansnils@Erik-Lundell@winskuo-quic@JakeStevens@swolchok@bigfootjon@mergennachin