Cortex-M backend: plan avg_pool2d scratch buffer AoT - #19825

Merged
rascani merged 5 commits into
pytorch:mainfrom
rascani:cortex-m-avgpool-aot-scratch
Jun 9, 2026
Merged

Cortex-M backend: plan avg_pool2d scratch buffer AoT#19825
rascani merged 5 commits into
pytorch:mainfrom
rascani:cortex-m-avgpool-aot-scratch

Conversation

@rascani

Copy link
Copy Markdown
Contributor

Summary

CMSIS-NN's DSP-variant arm_avgpool_s8 returns ARM_CMSIS_NN_ARG_ERROR when its ctx->buf is NULL and arm_avgpool_s8_get_buffer_size(...) is non-zero. The kernel hardcoded cmsis_ctx.buf = nullptr, which worked on Cortex-M55 because the MVE variant ignores ctx entirely, but failed on any DSP-class core (e.g. Cortex-M7). The AoT scratch-buffer planning system introduced for conv/depthwise-conv/transpose-conv/bmm missed quantized_avg_pool2d; this extends it to cover that op.

The quantized_avg_pool2d and .out schemas gain a Tensor scratch parameter. A new cmsis_nn_avgpool_buffer_size size function is registered, and avg_pool2d's lowering moves out of QuantizedOpFusionPass (which cannot create exir.memory.alloc nodes because it routes through ExportPass.call_operator) into ConvertToCortexMPass, alongside conv2d/bmm. The count_include_pad decomposition into an explicit cortex_m::pad node carries over to the new location. The kernel reads scratch.nbytes() and scratch.mutable_data_ptr<int8_t>() to wire cmsis_nn_context, with a CORTEX_M_ENABLE_RUNTIME_CHECKS-guarded assertion that the AoT and runtime buffer sizes agree — matching the conv2d pattern. The Python _NHWC_DIM_ORDER / to_physical_order helper that both passes need moves into passes_utils.py to avoid duplication.

Test plan

examples/arm/run.sh --model_name=mv2 --target=cortex-m7 --bundleio

A new dialect test exercises the ceil_mode=True fallback so that future refactors do not silently change which path it takes.

Authored with Claude.

@pytorch-bot

pytorch-botBot commented May 27, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/19825

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit 6d749a7 with merge base 189ffaa (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label May 27, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

CMSIS-NN's DSP-variant `arm_avgpool_s8` returns `ARM_CMSIS_NN_ARG_ERROR`
when its `ctx->buf` is NULL and `arm_avgpool_s8_get_buffer_size(...)` is
non-zero. The kernel hardcoded `cmsis_ctx.buf = nullptr`, which worked
on Cortex-M55 because the MVE variant ignores `ctx` entirely, but
failed on any DSP-class core (e.g. Cortex-M7). The AoT scratch-buffer
planning system introduced for conv/depthwise-conv/transpose-conv/bmm
missed `quantized_avg_pool2d`; this extends it to cover that op.
The `quantized_avg_pool2d` and `.out` schemas gain a `Tensor scratch`
parameter. A new `cmsis_nn_avgpool_buffer_size` size function is
registered, and avg_pool2d's lowering moves out of
`QuantizedOpFusionPass` (which cannot create `exir.memory.alloc` nodes
because it routes through `ExportPass.call_operator`) into
`ConvertToCortexMPass`, alongside conv2d/bmm. The `count_include_pad`
decomposition into an explicit `cortex_m::pad` node carries over to
the new location. The kernel reads `scratch.nbytes()` and
`scratch.mutable_data_ptr<int8_t>()` to wire `cmsis_nn_context`, with
a `CORTEX_M_ENABLE_RUNTIME_CHECKS`-guarded assertion that the AoT and
runtime buffer sizes agree — matching the conv2d pattern. The Python
`_NHWC_DIM_ORDER` / `to_physical_order` helper that both passes need
moves into `passes_utils.py` to avoid duplication.
A new dialect test exercises the `ceil_mode=True` fallback so that
future refactors do not silently change which path it takes.
Authored with Claude.
@rascani
rascaniforce-pushed the cortex-m-avgpool-aot-scratch branch from 4fd3c18 to 47e7fb8CompareJune 2, 2026 22:09
@rascani
rascani requested a review from digantdesaiJune 4, 2026 17:15
Comment threadbackends/cortex_m/test/ops/test_avg_pool2d.py
@digantdesai

Copy link
Copy Markdown
Contributor

@claude review this PR.

@claude

claudeBot commented Jun 9, 2026

Copy link
Copy Markdown

Claude finished @digantdesai's task in 3m 6s —— View job


PR Review: Cortex-M backend: plan avg_pool2d scratch buffer AoT

  • Read changed files
  • Analyze C++ kernel changes
  • Analyze Python pass changes
  • Analyze test coverage
  • Post review feedback

Overall: Clean, well-structured PR that follows the established scratch-buffer AoT planning pattern (conv2d, depthwise-conv, transpose-conv, bmm). The move from QuantizedOpFusionPass to ConvertToCortexMPass is well-motivated since the former routes through ExportPass.call_operator which cannot create exir.memory.alloc nodes. LGTM with minor observations below.


Findings

1. List aliasing in _get_avg_pool2d_replacement (convert_to_cortex_m_pass.py:514)

pre_pad=post_pad=to_physical_order([0, 0, pad_h, pad_w], input_tensor)

When is_channels_last(input_tensor) returns False, to_physical_order returns the same list object that was passed in (line 330 of passes_utils.py). Both pre_pad and post_pad then alias the same list. This is safe here since neither is mutated afterward, but the original code in QuantizedOpFusionPass explicitly created two independent lists. For symmetric padding the values are always equal, so this is correct but brittle if someone later mutates one of them.

Suggestion: pre_pad = post_pad = list(to_physical_order([0, 0, pad_h, pad_w], input_tensor)) — or just keep them as two separate calls for clarity. Very minor, not blocking.

2. No negative-return guard on arm_avgpool_s8_get_buffer_size (op_quantized_avg_pool2d.cpp:57-67)

The conv2d kernel's runtime check guards against negative return values from arm_convolve_wrapper_s8_get_buffer_size (which can return errors). The avgpool version skips this, which is correct — arm_avgpool_s8_get_buffer_size per the CMSIS-NN API always returns ≥ 0. Just noting this is intentional and consistent with the API contract.

3. Runtime check placement (op_quantized_avg_pool2d.cpp:56-68)

The runtime check validates the buffer size after the buffer pointer has already been assigned to cmsis_ctx.buf but beforearm_avgpool_s8 is called. This matches the conv2d pattern and is correct — the check is purely advisory/diagnostic and returns early before the kernel runs.

4. Reviewer question re: buffer size assertion in test (responding to @digantdesai's comment on line 108)

The test_dialect_avg_pool2d_fallback test verifies that ceil_mode=True correctly falls back to the aten op (leaving it unconverted). It doesn't exercise the scratch buffer path. The existingtest_dialect_avg_pool2d and test_implementation_avg_pool2d parametrized tests do exercise the scratch buffer path end-to-end (the _initialize_alloc_node_size call in ConvertToCortexMPass.call wires the buffer size). A unit test that directly asserts the buffer size value from cmsis_nn_avgpool_buffer_size would add confidence, but the existing e2e tests already catch mismatches (the kernel would fail at runtime with ARM_CMSIS_NN_ARG_ERROR).

5. Schema & YAML consistency — Verified: the .yaml, lib.define(...), @register_fake, @impl, and C++ signature all consistently include the new Tensor scratch parameter in the correct position.

6. Test coverage — Good. The PR adds:

  • fallback_test_cases + test_dialect_avg_pool2d_fallback: verifies ceil_mode=True stays as aten op
  • Existing parametrized tests (test_dialect_avg_pool2d, test_implementation_avg_pool2d) cover the happy path with the scratch buffer

Summary

This is a solid, well-scoped extension of the existing AoT scratch buffer system. The code correctly follows the established patterns, the kernel change is minimal and guarded, and the test coverage is reasonable. The only material suggestion is the minor list-aliasing nit in the pad decomposition path.

rascaniand others added 2 commits June 9, 2026 10:54
…t-scratch
# Conflicts:
#	backends/cortex_m/passes/convert_to_cortex_m_pass.py
Adapt _get_avg_pool2d_replacement from the old ConvertToCortexMPass class
to a @register_dialect_substitution function on AtenToCortexMPass,
matching the pattern established by pytorch#20070. Add casts to satisfy mypy's
node.args union type.
Add a structural assertion to test_dialect_avg_pool2d that verifies the
scratch buffer size matches cmsis_nn.avgpool_buffer_size() for the given
input/output dimensions.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@rascani
rascani merged commit 23b6ba0 into pytorch:mainJun 9, 2026
402 of 404 checks passed
@rascani
rascani deleted the cortex-m-avgpool-aot-scratch branch June 9, 2026 23:05
rascani added a commit that referenced this pull request Jun 9, 2026
)
Use `dialect_pass: AtenToDialectPass` parameter matching the
SubstitutionFn type updated in #19676. The `exported_program` parameter
caused a mypy arg-type error when both #19676 and #19825 landed on main.
Fixes the lintrunner-mypy failure on main.
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
psiddh added a commit to psiddh/executorch that referenced this pull request Aug 3, 2026
The library is generated from this repository, so an ExecuTorch change can
break it without touching examples/arduino at all. Nothing checked that.
The job builds the library, verifies the shipped models against it, compiles
every example for arduino:zephyr:unoq, and runs arduino-lint in Library
Manager submission mode. It runs on changes under examples/arduino,
backends/cortex_m, kernels/portable, runtime, and schema, which is the set
that can break any of the four.
verify_models.py is the part worth having. A .pte records how many values
each operator call puts on the stack, and the generated kernel wrappers
reject anything else; the two only agree when the model and the library came
from the same commit. Cortex-M schemas change - scratch was added to the conv
operators in pytorch#19636 and pytorch#19825 - and a mismatch is invisible until far too
late, because the program loads, every operator resolves, and then execute
returns InvalidProgram naming nothing useful. Tracking that down by hand took
the better part of a day. The check compares the two directly in about a
second, needs no board, and reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects 11
Compiling is deliberately not the whole job. Every failure worth finding in
this library compiled cleanly first, so the model check runs before the
sketches and arduino-lint guards the thing users actually install.
Authored with Claude Code (Opus 5).
psiddh added a commit to psiddh/executorch that referenced this pull request Aug 3, 2026
The library is generated from this repository, so an ExecuTorch change can
break it without touching examples/arduino at all. Nothing checked that.
The job builds the library, verifies the shipped models against it, compiles
every example for arduino:zephyr:unoq, and runs arduino-lint in Library
Manager submission mode. It runs on changes under examples/arduino,
backends/cortex_m, kernels/portable, runtime, and schema, which is the set
that can break any of the four.
verify_models.py is the part worth having. A .pte records how many values
each operator call puts on the stack, and the generated kernel wrappers
reject anything else; the two only agree when the model and the library came
from the same commit. Cortex-M schemas change - scratch was added to the conv
operators in pytorch#19636 and pytorch#19825 - and a mismatch is invisible until far too
late, because the program loads, every operator resolves, and then execute
returns InvalidProgram naming nothing useful. Tracking that down by hand took
the better part of a day. The check compares the two directly in about a
second, needs no board, and reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects 11
Compiling is deliberately not the whole job. Every failure worth finding in
this library compiled cleanly first, so the model check runs before the
sketches and arduino-lint guards the thing users actually install.
Authored with Claude Code (Opus 5).
psiddh added a commit to psiddh/executorch that referenced this pull request Aug 4, 2026
The library is generated from this repository, so an ExecuTorch change can
break it without touching examples/arduino at all. Nothing checked that.
The job builds the library, verifies the shipped models against it, compiles
every example for arduino:zephyr:unoq, and runs arduino-lint in Library
Manager submission mode. It runs on changes under examples/arduino,
backends/cortex_m, kernels/portable, runtime, and schema, which is the set
that can break any of the four.
verify_models.py is the part worth having. A .pte records how many values
each operator call puts on the stack, and the generated kernel wrappers
reject anything else; the two only agree when the model and the library came
from the same commit. Cortex-M schemas change - scratch was added to the conv
operators in pytorch#19636 and pytorch#19825 - and a mismatch is invisible until far too
late, because the program loads, every operator resolves, and then execute
returns InvalidProgram naming nothing useful. Tracking that down by hand took
the better part of a day. The check compares the two directly in about a
second, needs no board, and reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects 11
Compiling is deliberately not the whole job. Every failure worth finding in
this library compiled cleanly first, so the model check runs before the
sketches and arduino-lint guards the thing users actually install.
Authored with Claude Code (Opus 5).
psiddh added a commit that referenced this pull request Aug 4, 2026
…els (#21546)
## Summary
The Arduino library this directory generates could not compile, and had
it compiled it could not have run a model.
Kernels never reached the operator registry — ExecuTorch registers them
through codegen and the build script never ran it, so every
`Method::load` would have failed with `OperatorMissing`. CMSIS-NN was
never vendored, so the Cortex-M ops shipped without the library they
call. `schema/*.cpp` was never copied, the `*_aten.cpp` exclusion also
deleted the portable-mode `tensor_parser_exec_aten.cpp`, and `__errno`
was missing because the Zephyr core mixes picolibc with newlib's
`libm_nano`.
Both `minimal.cpp` and `zephyr.cpp` were vendored, so which `et_pal_*`
backend you got depended on link order, and every `ET_LOG` was discarded
either way. One backend now ships and its logs route to a hook the
examples implement against `Serial`.
The examples ship their models — previously none did, so every sketch
`#error`ed when opened from the IDE menu. The README pointed at the
Ethos-U `pte_to_header.py`, whose `network_model_sec` section no Arduino
core defines, so following the docs produced a model that fails
`Program::load`. Static link mode is mandatory and undocumented; the
Dynamic default yields a silently dead board. Renamed to `ExecuTorch`
because `arduino-lint` rejects "Arduino" in an Arduino library's name.
Registering every portable kernel costs 1.58 MB against 786 KB of flash,
so the op set is a curated default overridable via `ROOT_OPS`/`ALL_OPS`.
Models and libraries must come from the same ExecuTorch commit —
Cortex-M schemas change (`scratch` in #19636, #19825), and a mismatch
loads fine, resolves every operator, then fails at `Method::execute`.
The library now records and pins that commit. CI to enforce it follows
separately.
## Test plan
`arduino:zephyr:unoq`, board core 0.55.2, `link_mode=static`, flashed on
hardware:
```
HelloExecuTorch Model loaded OK!, 1 method 60% flash, 20% RAM
AddModel [1,2,3] + 1 = [2.00, 3.00, 4.00] 64% flash, 26% RAM
KeywordSpotting 10/10 keywords correct 70% flash, 53% RAM
```
All ten MFCC inputs in one sketch, exercising the CMSIS-NN conv /
depthwise / avgpool / linear kernels:
```
yes 8.95 no 4.78 up 4.63 down 9.72 left 7.87
right 7.41 on 12.03 off 8.02 stop 8.64 go 9.57
```
`arduino-lint --library-manager submit`: no errors, no warnings, under
both `specification` and `strict`. Clean regeneration is byte-identical
across all 626 generated files.
Only the Uno Q was tested; the other three boards remain marked Planned.
Authored with Claude Code (Opus 5).
psiddh added a commit that referenced this pull request Aug 11, 2026
The library is generated from this repository, so an ExecuTorch change
can break it without touching examples/arduino at all. Nothing checked
that.
The job builds the library, verifies the shipped models against it,
compiles every example for arduino:zephyr:unoq, and runs arduino-lint in
Library Manager submission mode. It runs on changes under
examples/arduino, backends/cortex_m, kernels/portable, runtime, and
schema, which is the set that can break any of the four.
verify_models.py is the part worth having. A .pte records how many
values each operator call puts on the stack, and the generated kernel
wrappers reject anything else; the two only agree when the model and the
library came from the same commit. Cortex-M schemas change - scratch was
added to the conv operators in #19636 and #19825 - and a mismatch is
invisible until far too late, because the program loads, every operator
resolves, and then execute returns InvalidProgram naming nothing useful.
Tracking that down by hand took the better part of a day. The check
compares the two directly in about a second, needs no board, and
reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects
11
Compiling is deliberately not the whole job. Every failure worth finding
in this library compiled cleanly first, so the model check runs before
the sketches and arduino-lint guards the thing users actually install.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@rascani@digantdesai
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all \u003cpre\u003e\u003ccode\u003e blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks"); } } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); } })(); (function(){ try { var __m = "github.com"; var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Cortex-M backend: plan avg_pool2d scratch buffer AoT - #19825

Merged
rascani merged 5 commits into
pytorch:mainfrom
rascani:cortex-m-avgpool-aot-scratch
Jun 9, 2026
Merged

Cortex-M backend: plan avg_pool2d scratch buffer AoT#19825
rascani merged 5 commits into
pytorch:mainfrom
rascani:cortex-m-avgpool-aot-scratch

Conversation

@rascani

Copy link
Copy Markdown
Contributor

Summary

CMSIS-NN's DSP-variant arm_avgpool_s8 returns ARM_CMSIS_NN_ARG_ERROR when its ctx->buf is NULL and arm_avgpool_s8_get_buffer_size(...) is non-zero. The kernel hardcoded cmsis_ctx.buf = nullptr, which worked on Cortex-M55 because the MVE variant ignores ctx entirely, but failed on any DSP-class core (e.g. Cortex-M7). The AoT scratch-buffer planning system introduced for conv/depthwise-conv/transpose-conv/bmm missed quantized_avg_pool2d; this extends it to cover that op.

The quantized_avg_pool2d and .out schemas gain a Tensor scratch parameter. A new cmsis_nn_avgpool_buffer_size size function is registered, and avg_pool2d's lowering moves out of QuantizedOpFusionPass (which cannot create exir.memory.alloc nodes because it routes through ExportPass.call_operator) into ConvertToCortexMPass, alongside conv2d/bmm. The count_include_pad decomposition into an explicit cortex_m::pad node carries over to the new location. The kernel reads scratch.nbytes() and scratch.mutable_data_ptr<int8_t>() to wire cmsis_nn_context, with a CORTEX_M_ENABLE_RUNTIME_CHECKS-guarded assertion that the AoT and runtime buffer sizes agree — matching the conv2d pattern. The Python _NHWC_DIM_ORDER / to_physical_order helper that both passes need moves into passes_utils.py to avoid duplication.

Test plan

examples/arm/run.sh --model_name=mv2 --target=cortex-m7 --bundleio

A new dialect test exercises the ceil_mode=True fallback so that future refactors do not silently change which path it takes.

Authored with Claude.

@pytorch-bot

pytorch-botBot commented May 27, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/19825

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit 6d749a7 with merge base 189ffaa (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label May 27, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

CMSIS-NN's DSP-variant `arm_avgpool_s8` returns `ARM_CMSIS_NN_ARG_ERROR`
when its `ctx->buf` is NULL and `arm_avgpool_s8_get_buffer_size(...)` is
non-zero. The kernel hardcoded `cmsis_ctx.buf = nullptr`, which worked
on Cortex-M55 because the MVE variant ignores `ctx` entirely, but
failed on any DSP-class core (e.g. Cortex-M7). The AoT scratch-buffer
planning system introduced for conv/depthwise-conv/transpose-conv/bmm
missed `quantized_avg_pool2d`; this extends it to cover that op.
The `quantized_avg_pool2d` and `.out` schemas gain a `Tensor scratch`
parameter. A new `cmsis_nn_avgpool_buffer_size` size function is
registered, and avg_pool2d's lowering moves out of
`QuantizedOpFusionPass` (which cannot create `exir.memory.alloc` nodes
because it routes through `ExportPass.call_operator`) into
`ConvertToCortexMPass`, alongside conv2d/bmm. The `count_include_pad`
decomposition into an explicit `cortex_m::pad` node carries over to
the new location. The kernel reads `scratch.nbytes()` and
`scratch.mutable_data_ptr<int8_t>()` to wire `cmsis_nn_context`, with
a `CORTEX_M_ENABLE_RUNTIME_CHECKS`-guarded assertion that the AoT and
runtime buffer sizes agree — matching the conv2d pattern. The Python
`_NHWC_DIM_ORDER` / `to_physical_order` helper that both passes need
moves into `passes_utils.py` to avoid duplication.
A new dialect test exercises the `ceil_mode=True` fallback so that
future refactors do not silently change which path it takes.
Authored with Claude.
@rascani
rascaniforce-pushed the cortex-m-avgpool-aot-scratch branch from 4fd3c18 to 47e7fb8CompareJune 2, 2026 22:09
@rascani
rascani requested a review from digantdesaiJune 4, 2026 17:15
Comment threadbackends/cortex_m/test/ops/test_avg_pool2d.py
@digantdesai

Copy link
Copy Markdown
Contributor

@claude review this PR.

@claude

claudeBot commented Jun 9, 2026

Copy link
Copy Markdown

Claude finished @digantdesai's task in 3m 6s —— View job


PR Review: Cortex-M backend: plan avg_pool2d scratch buffer AoT

  • Read changed files
  • Analyze C++ kernel changes
  • Analyze Python pass changes
  • Analyze test coverage
  • Post review feedback

Overall: Clean, well-structured PR that follows the established scratch-buffer AoT planning pattern (conv2d, depthwise-conv, transpose-conv, bmm). The move from QuantizedOpFusionPass to ConvertToCortexMPass is well-motivated since the former routes through ExportPass.call_operator which cannot create exir.memory.alloc nodes. LGTM with minor observations below.


Findings

1. List aliasing in _get_avg_pool2d_replacement (convert_to_cortex_m_pass.py:514)

pre_pad=post_pad=to_physical_order([0, 0, pad_h, pad_w], input_tensor)

When is_channels_last(input_tensor) returns False, to_physical_order returns the same list object that was passed in (line 330 of passes_utils.py). Both pre_pad and post_pad then alias the same list. This is safe here since neither is mutated afterward, but the original code in QuantizedOpFusionPass explicitly created two independent lists. For symmetric padding the values are always equal, so this is correct but brittle if someone later mutates one of them.

Suggestion: pre_pad = post_pad = list(to_physical_order([0, 0, pad_h, pad_w], input_tensor)) — or just keep them as two separate calls for clarity. Very minor, not blocking.

2. No negative-return guard on arm_avgpool_s8_get_buffer_size (op_quantized_avg_pool2d.cpp:57-67)

The conv2d kernel's runtime check guards against negative return values from arm_convolve_wrapper_s8_get_buffer_size (which can return errors). The avgpool version skips this, which is correct — arm_avgpool_s8_get_buffer_size per the CMSIS-NN API always returns ≥ 0. Just noting this is intentional and consistent with the API contract.

3. Runtime check placement (op_quantized_avg_pool2d.cpp:56-68)

The runtime check validates the buffer size after the buffer pointer has already been assigned to cmsis_ctx.buf but beforearm_avgpool_s8 is called. This matches the conv2d pattern and is correct — the check is purely advisory/diagnostic and returns early before the kernel runs.

4. Reviewer question re: buffer size assertion in test (responding to @digantdesai's comment on line 108)

The test_dialect_avg_pool2d_fallback test verifies that ceil_mode=True correctly falls back to the aten op (leaving it unconverted). It doesn't exercise the scratch buffer path. The existingtest_dialect_avg_pool2d and test_implementation_avg_pool2d parametrized tests do exercise the scratch buffer path end-to-end (the _initialize_alloc_node_size call in ConvertToCortexMPass.call wires the buffer size). A unit test that directly asserts the buffer size value from cmsis_nn_avgpool_buffer_size would add confidence, but the existing e2e tests already catch mismatches (the kernel would fail at runtime with ARM_CMSIS_NN_ARG_ERROR).

5. Schema & YAML consistency — Verified: the .yaml, lib.define(...), @register_fake, @impl, and C++ signature all consistently include the new Tensor scratch parameter in the correct position.

6. Test coverage — Good. The PR adds:

  • fallback_test_cases + test_dialect_avg_pool2d_fallback: verifies ceil_mode=True stays as aten op
  • Existing parametrized tests (test_dialect_avg_pool2d, test_implementation_avg_pool2d) cover the happy path with the scratch buffer

Summary

This is a solid, well-scoped extension of the existing AoT scratch buffer system. The code correctly follows the established patterns, the kernel change is minimal and guarded, and the test coverage is reasonable. The only material suggestion is the minor list-aliasing nit in the pad decomposition path.

rascaniand others added 2 commits June 9, 2026 10:54
…t-scratch
# Conflicts:
#	backends/cortex_m/passes/convert_to_cortex_m_pass.py
Adapt _get_avg_pool2d_replacement from the old ConvertToCortexMPass class
to a @register_dialect_substitution function on AtenToCortexMPass,
matching the pattern established by pytorch#20070. Add casts to satisfy mypy's
node.args union type.
Add a structural assertion to test_dialect_avg_pool2d that verifies the
scratch buffer size matches cmsis_nn.avgpool_buffer_size() for the given
input/output dimensions.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@rascani
rascani merged commit 23b6ba0 into pytorch:mainJun 9, 2026
402 of 404 checks passed
@rascani
rascani deleted the cortex-m-avgpool-aot-scratch branch June 9, 2026 23:05
rascani added a commit that referenced this pull request Jun 9, 2026
)
Use `dialect_pass: AtenToDialectPass` parameter matching the
SubstitutionFn type updated in #19676. The `exported_program` parameter
caused a mypy arg-type error when both #19676 and #19825 landed on main.
Fixes the lintrunner-mypy failure on main.
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
psiddh added a commit to psiddh/executorch that referenced this pull request Aug 3, 2026
The library is generated from this repository, so an ExecuTorch change can
break it without touching examples/arduino at all. Nothing checked that.
The job builds the library, verifies the shipped models against it, compiles
every example for arduino:zephyr:unoq, and runs arduino-lint in Library
Manager submission mode. It runs on changes under examples/arduino,
backends/cortex_m, kernels/portable, runtime, and schema, which is the set
that can break any of the four.
verify_models.py is the part worth having. A .pte records how many values
each operator call puts on the stack, and the generated kernel wrappers
reject anything else; the two only agree when the model and the library came
from the same commit. Cortex-M schemas change - scratch was added to the conv
operators in pytorch#19636 and pytorch#19825 - and a mismatch is invisible until far too
late, because the program loads, every operator resolves, and then execute
returns InvalidProgram naming nothing useful. Tracking that down by hand took
the better part of a day. The check compares the two directly in about a
second, needs no board, and reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects 11
Compiling is deliberately not the whole job. Every failure worth finding in
this library compiled cleanly first, so the model check runs before the
sketches and arduino-lint guards the thing users actually install.
Authored with Claude Code (Opus 5).
psiddh added a commit to psiddh/executorch that referenced this pull request Aug 3, 2026
The library is generated from this repository, so an ExecuTorch change can
break it without touching examples/arduino at all. Nothing checked that.
The job builds the library, verifies the shipped models against it, compiles
every example for arduino:zephyr:unoq, and runs arduino-lint in Library
Manager submission mode. It runs on changes under examples/arduino,
backends/cortex_m, kernels/portable, runtime, and schema, which is the set
that can break any of the four.
verify_models.py is the part worth having. A .pte records how many values
each operator call puts on the stack, and the generated kernel wrappers
reject anything else; the two only agree when the model and the library came
from the same commit. Cortex-M schemas change - scratch was added to the conv
operators in pytorch#19636 and pytorch#19825 - and a mismatch is invisible until far too
late, because the program loads, every operator resolves, and then execute
returns InvalidProgram naming nothing useful. Tracking that down by hand took
the better part of a day. The check compares the two directly in about a
second, needs no board, and reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects 11
Compiling is deliberately not the whole job. Every failure worth finding in
this library compiled cleanly first, so the model check runs before the
sketches and arduino-lint guards the thing users actually install.
Authored with Claude Code (Opus 5).
psiddh added a commit to psiddh/executorch that referenced this pull request Aug 4, 2026
The library is generated from this repository, so an ExecuTorch change can
break it without touching examples/arduino at all. Nothing checked that.
The job builds the library, verifies the shipped models against it, compiles
every example for arduino:zephyr:unoq, and runs arduino-lint in Library
Manager submission mode. It runs on changes under examples/arduino,
backends/cortex_m, kernels/portable, runtime, and schema, which is the set
that can break any of the four.
verify_models.py is the part worth having. A .pte records how many values
each operator call puts on the stack, and the generated kernel wrappers
reject anything else; the two only agree when the model and the library came
from the same commit. Cortex-M schemas change - scratch was added to the conv
operators in pytorch#19636 and pytorch#19825 - and a mismatch is invisible until far too
late, because the program loads, every operator resolves, and then execute
returns InvalidProgram naming nothing useful. Tracking that down by hand took
the better part of a day. The check compares the two directly in about a
second, needs no board, and reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects 11
Compiling is deliberately not the whole job. Every failure worth finding in
this library compiled cleanly first, so the model check runs before the
sketches and arduino-lint guards the thing users actually install.
Authored with Claude Code (Opus 5).
psiddh added a commit that referenced this pull request Aug 4, 2026
…els (#21546)
## Summary
The Arduino library this directory generates could not compile, and had
it compiled it could not have run a model.
Kernels never reached the operator registry — ExecuTorch registers them
through codegen and the build script never ran it, so every
`Method::load` would have failed with `OperatorMissing`. CMSIS-NN was
never vendored, so the Cortex-M ops shipped without the library they
call. `schema/*.cpp` was never copied, the `*_aten.cpp` exclusion also
deleted the portable-mode `tensor_parser_exec_aten.cpp`, and `__errno`
was missing because the Zephyr core mixes picolibc with newlib's
`libm_nano`.
Both `minimal.cpp` and `zephyr.cpp` were vendored, so which `et_pal_*`
backend you got depended on link order, and every `ET_LOG` was discarded
either way. One backend now ships and its logs route to a hook the
examples implement against `Serial`.
The examples ship their models — previously none did, so every sketch
`#error`ed when opened from the IDE menu. The README pointed at the
Ethos-U `pte_to_header.py`, whose `network_model_sec` section no Arduino
core defines, so following the docs produced a model that fails
`Program::load`. Static link mode is mandatory and undocumented; the
Dynamic default yields a silently dead board. Renamed to `ExecuTorch`
because `arduino-lint` rejects "Arduino" in an Arduino library's name.
Registering every portable kernel costs 1.58 MB against 786 KB of flash,
so the op set is a curated default overridable via `ROOT_OPS`/`ALL_OPS`.
Models and libraries must come from the same ExecuTorch commit —
Cortex-M schemas change (`scratch` in #19636, #19825), and a mismatch
loads fine, resolves every operator, then fails at `Method::execute`.
The library now records and pins that commit. CI to enforce it follows
separately.
## Test plan
`arduino:zephyr:unoq`, board core 0.55.2, `link_mode=static`, flashed on
hardware:
```
HelloExecuTorch Model loaded OK!, 1 method 60% flash, 20% RAM
AddModel [1,2,3] + 1 = [2.00, 3.00, 4.00] 64% flash, 26% RAM
KeywordSpotting 10/10 keywords correct 70% flash, 53% RAM
```
All ten MFCC inputs in one sketch, exercising the CMSIS-NN conv /
depthwise / avgpool / linear kernels:
```
yes 8.95 no 4.78 up 4.63 down 9.72 left 7.87
right 7.41 on 12.03 off 8.02 stop 8.64 go 9.57
```
`arduino-lint --library-manager submit`: no errors, no warnings, under
both `specification` and `strict`. Clean regeneration is byte-identical
across all 626 generated files.
Only the Uno Q was tested; the other three boards remain marked Planned.
Authored with Claude Code (Opus 5).
psiddh added a commit that referenced this pull request Aug 11, 2026
The library is generated from this repository, so an ExecuTorch change
can break it without touching examples/arduino at all. Nothing checked
that.
The job builds the library, verifies the shipped models against it,
compiles every example for arduino:zephyr:unoq, and runs arduino-lint in
Library Manager submission mode. It runs on changes under
examples/arduino, backends/cortex_m, kernels/portable, runtime, and
schema, which is the set that can break any of the four.
verify_models.py is the part worth having. A .pte records how many
values each operator call puts on the stack, and the generated kernel
wrappers reject anything else; the two only agree when the model and the
library came from the same commit. Cortex-M schemas change - scratch was
added to the conv operators in #19636 and #19825 - and a mismatch is
invisible until far too late, because the program loads, every operator
resolves, and then execute returns InvalidProgram naming nothing useful.
Tracking that down by hand took the better part of a day. The check
compares the two directly in about a second, needs no board, and
reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects
11
Compiling is deliberately not the whole job. Every failure worth finding
in this library compiled cleanly first, so the model check runs before
the sketches and arduino-lint guards the thing users actually install.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@rascani@digantdesai
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Cortex-M backend: plan avg_pool2d scratch buffer AoT - #19825

Merged
rascani merged 5 commits into
pytorch:mainfrom
rascani:cortex-m-avgpool-aot-scratch
Jun 9, 2026
Merged

Cortex-M backend: plan avg_pool2d scratch buffer AoT#19825
rascani merged 5 commits into
pytorch:mainfrom
rascani:cortex-m-avgpool-aot-scratch

Conversation

@rascani

Copy link
Copy Markdown
Contributor

Summary

CMSIS-NN's DSP-variant arm_avgpool_s8 returns ARM_CMSIS_NN_ARG_ERROR when its ctx->buf is NULL and arm_avgpool_s8_get_buffer_size(...) is non-zero. The kernel hardcoded cmsis_ctx.buf = nullptr, which worked on Cortex-M55 because the MVE variant ignores ctx entirely, but failed on any DSP-class core (e.g. Cortex-M7). The AoT scratch-buffer planning system introduced for conv/depthwise-conv/transpose-conv/bmm missed quantized_avg_pool2d; this extends it to cover that op.

The quantized_avg_pool2d and .out schemas gain a Tensor scratch parameter. A new cmsis_nn_avgpool_buffer_size size function is registered, and avg_pool2d's lowering moves out of QuantizedOpFusionPass (which cannot create exir.memory.alloc nodes because it routes through ExportPass.call_operator) into ConvertToCortexMPass, alongside conv2d/bmm. The count_include_pad decomposition into an explicit cortex_m::pad node carries over to the new location. The kernel reads scratch.nbytes() and scratch.mutable_data_ptr<int8_t>() to wire cmsis_nn_context, with a CORTEX_M_ENABLE_RUNTIME_CHECKS-guarded assertion that the AoT and runtime buffer sizes agree — matching the conv2d pattern. The Python _NHWC_DIM_ORDER / to_physical_order helper that both passes need moves into passes_utils.py to avoid duplication.

Test plan

examples/arm/run.sh --model_name=mv2 --target=cortex-m7 --bundleio

A new dialect test exercises the ceil_mode=True fallback so that future refactors do not silently change which path it takes.

Authored with Claude.

@pytorch-bot

pytorch-botBot commented May 27, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/19825

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit 6d749a7 with merge base 189ffaa (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label May 27, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

CMSIS-NN's DSP-variant `arm_avgpool_s8` returns `ARM_CMSIS_NN_ARG_ERROR`
when its `ctx->buf` is NULL and `arm_avgpool_s8_get_buffer_size(...)` is
non-zero. The kernel hardcoded `cmsis_ctx.buf = nullptr`, which worked
on Cortex-M55 because the MVE variant ignores `ctx` entirely, but
failed on any DSP-class core (e.g. Cortex-M7). The AoT scratch-buffer
planning system introduced for conv/depthwise-conv/transpose-conv/bmm
missed `quantized_avg_pool2d`; this extends it to cover that op.
The `quantized_avg_pool2d` and `.out` schemas gain a `Tensor scratch`
parameter. A new `cmsis_nn_avgpool_buffer_size` size function is
registered, and avg_pool2d's lowering moves out of
`QuantizedOpFusionPass` (which cannot create `exir.memory.alloc` nodes
because it routes through `ExportPass.call_operator`) into
`ConvertToCortexMPass`, alongside conv2d/bmm. The `count_include_pad`
decomposition into an explicit `cortex_m::pad` node carries over to
the new location. The kernel reads `scratch.nbytes()` and
`scratch.mutable_data_ptr<int8_t>()` to wire `cmsis_nn_context`, with
a `CORTEX_M_ENABLE_RUNTIME_CHECKS`-guarded assertion that the AoT and
runtime buffer sizes agree — matching the conv2d pattern. The Python
`_NHWC_DIM_ORDER` / `to_physical_order` helper that both passes need
moves into `passes_utils.py` to avoid duplication.
A new dialect test exercises the `ceil_mode=True` fallback so that
future refactors do not silently change which path it takes.
Authored with Claude.
@rascani
rascaniforce-pushed the cortex-m-avgpool-aot-scratch branch from 4fd3c18 to 47e7fb8CompareJune 2, 2026 22:09
@rascani
rascani requested a review from digantdesaiJune 4, 2026 17:15
Comment threadbackends/cortex_m/test/ops/test_avg_pool2d.py
@digantdesai

Copy link
Copy Markdown
Contributor

@claude review this PR.

@claude

claudeBot commented Jun 9, 2026

Copy link
Copy Markdown

Claude finished @digantdesai's task in 3m 6s —— View job


PR Review: Cortex-M backend: plan avg_pool2d scratch buffer AoT

  • Read changed files
  • Analyze C++ kernel changes
  • Analyze Python pass changes
  • Analyze test coverage
  • Post review feedback

Overall: Clean, well-structured PR that follows the established scratch-buffer AoT planning pattern (conv2d, depthwise-conv, transpose-conv, bmm). The move from QuantizedOpFusionPass to ConvertToCortexMPass is well-motivated since the former routes through ExportPass.call_operator which cannot create exir.memory.alloc nodes. LGTM with minor observations below.


Findings

1. List aliasing in _get_avg_pool2d_replacement (convert_to_cortex_m_pass.py:514)

pre_pad=post_pad=to_physical_order([0, 0, pad_h, pad_w], input_tensor)

When is_channels_last(input_tensor) returns False, to_physical_order returns the same list object that was passed in (line 330 of passes_utils.py). Both pre_pad and post_pad then alias the same list. This is safe here since neither is mutated afterward, but the original code in QuantizedOpFusionPass explicitly created two independent lists. For symmetric padding the values are always equal, so this is correct but brittle if someone later mutates one of them.

Suggestion: pre_pad = post_pad = list(to_physical_order([0, 0, pad_h, pad_w], input_tensor)) — or just keep them as two separate calls for clarity. Very minor, not blocking.

2. No negative-return guard on arm_avgpool_s8_get_buffer_size (op_quantized_avg_pool2d.cpp:57-67)

The conv2d kernel's runtime check guards against negative return values from arm_convolve_wrapper_s8_get_buffer_size (which can return errors). The avgpool version skips this, which is correct — arm_avgpool_s8_get_buffer_size per the CMSIS-NN API always returns ≥ 0. Just noting this is intentional and consistent with the API contract.

3. Runtime check placement (op_quantized_avg_pool2d.cpp:56-68)

The runtime check validates the buffer size after the buffer pointer has already been assigned to cmsis_ctx.buf but beforearm_avgpool_s8 is called. This matches the conv2d pattern and is correct — the check is purely advisory/diagnostic and returns early before the kernel runs.

4. Reviewer question re: buffer size assertion in test (responding to @digantdesai's comment on line 108)

The test_dialect_avg_pool2d_fallback test verifies that ceil_mode=True correctly falls back to the aten op (leaving it unconverted). It doesn't exercise the scratch buffer path. The existingtest_dialect_avg_pool2d and test_implementation_avg_pool2d parametrized tests do exercise the scratch buffer path end-to-end (the _initialize_alloc_node_size call in ConvertToCortexMPass.call wires the buffer size). A unit test that directly asserts the buffer size value from cmsis_nn_avgpool_buffer_size would add confidence, but the existing e2e tests already catch mismatches (the kernel would fail at runtime with ARM_CMSIS_NN_ARG_ERROR).

5. Schema & YAML consistency — Verified: the .yaml, lib.define(...), @register_fake, @impl, and C++ signature all consistently include the new Tensor scratch parameter in the correct position.

6. Test coverage — Good. The PR adds:

  • fallback_test_cases + test_dialect_avg_pool2d_fallback: verifies ceil_mode=True stays as aten op
  • Existing parametrized tests (test_dialect_avg_pool2d, test_implementation_avg_pool2d) cover the happy path with the scratch buffer

Summary

This is a solid, well-scoped extension of the existing AoT scratch buffer system. The code correctly follows the established patterns, the kernel change is minimal and guarded, and the test coverage is reasonable. The only material suggestion is the minor list-aliasing nit in the pad decomposition path.

rascaniand others added 2 commits June 9, 2026 10:54
…t-scratch
# Conflicts:
#	backends/cortex_m/passes/convert_to_cortex_m_pass.py
Adapt _get_avg_pool2d_replacement from the old ConvertToCortexMPass class
to a @register_dialect_substitution function on AtenToCortexMPass,
matching the pattern established by pytorch#20070. Add casts to satisfy mypy's
node.args union type.
Add a structural assertion to test_dialect_avg_pool2d that verifies the
scratch buffer size matches cmsis_nn.avgpool_buffer_size() for the given
input/output dimensions.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@rascani
rascani merged commit 23b6ba0 into pytorch:mainJun 9, 2026
402 of 404 checks passed
@rascani
rascani deleted the cortex-m-avgpool-aot-scratch branch June 9, 2026 23:05
rascani added a commit that referenced this pull request Jun 9, 2026
)
Use `dialect_pass: AtenToDialectPass` parameter matching the
SubstitutionFn type updated in #19676. The `exported_program` parameter
caused a mypy arg-type error when both #19676 and #19825 landed on main.
Fixes the lintrunner-mypy failure on main.
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
psiddh added a commit to psiddh/executorch that referenced this pull request Aug 3, 2026
The library is generated from this repository, so an ExecuTorch change can
break it without touching examples/arduino at all. Nothing checked that.
The job builds the library, verifies the shipped models against it, compiles
every example for arduino:zephyr:unoq, and runs arduino-lint in Library
Manager submission mode. It runs on changes under examples/arduino,
backends/cortex_m, kernels/portable, runtime, and schema, which is the set
that can break any of the four.
verify_models.py is the part worth having. A .pte records how many values
each operator call puts on the stack, and the generated kernel wrappers
reject anything else; the two only agree when the model and the library came
from the same commit. Cortex-M schemas change - scratch was added to the conv
operators in pytorch#19636 and pytorch#19825 - and a mismatch is invisible until far too
late, because the program loads, every operator resolves, and then execute
returns InvalidProgram naming nothing useful. Tracking that down by hand took
the better part of a day. The check compares the two directly in about a
second, needs no board, and reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects 11
Compiling is deliberately not the whole job. Every failure worth finding in
this library compiled cleanly first, so the model check runs before the
sketches and arduino-lint guards the thing users actually install.
Authored with Claude Code (Opus 5).
psiddh added a commit to psiddh/executorch that referenced this pull request Aug 3, 2026
The library is generated from this repository, so an ExecuTorch change can
break it without touching examples/arduino at all. Nothing checked that.
The job builds the library, verifies the shipped models against it, compiles
every example for arduino:zephyr:unoq, and runs arduino-lint in Library
Manager submission mode. It runs on changes under examples/arduino,
backends/cortex_m, kernels/portable, runtime, and schema, which is the set
that can break any of the four.
verify_models.py is the part worth having. A .pte records how many values
each operator call puts on the stack, and the generated kernel wrappers
reject anything else; the two only agree when the model and the library came
from the same commit. Cortex-M schemas change - scratch was added to the conv
operators in pytorch#19636 and pytorch#19825 - and a mismatch is invisible until far too
late, because the program loads, every operator resolves, and then execute
returns InvalidProgram naming nothing useful. Tracking that down by hand took
the better part of a day. The check compares the two directly in about a
second, needs no board, and reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects 11
Compiling is deliberately not the whole job. Every failure worth finding in
this library compiled cleanly first, so the model check runs before the
sketches and arduino-lint guards the thing users actually install.
Authored with Claude Code (Opus 5).
psiddh added a commit to psiddh/executorch that referenced this pull request Aug 4, 2026
The library is generated from this repository, so an ExecuTorch change can
break it without touching examples/arduino at all. Nothing checked that.
The job builds the library, verifies the shipped models against it, compiles
every example for arduino:zephyr:unoq, and runs arduino-lint in Library
Manager submission mode. It runs on changes under examples/arduino,
backends/cortex_m, kernels/portable, runtime, and schema, which is the set
that can break any of the four.
verify_models.py is the part worth having. A .pte records how many values
each operator call puts on the stack, and the generated kernel wrappers
reject anything else; the two only agree when the model and the library came
from the same commit. Cortex-M schemas change - scratch was added to the conv
operators in pytorch#19636 and pytorch#19825 - and a mismatch is invisible until far too
late, because the program loads, every operator resolves, and then execute
returns InvalidProgram naming nothing useful. Tracking that down by hand took
the better part of a day. The check compares the two directly in about a
second, needs no board, and reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects 11
Compiling is deliberately not the whole job. Every failure worth finding in
this library compiled cleanly first, so the model check runs before the
sketches and arduino-lint guards the thing users actually install.
Authored with Claude Code (Opus 5).
psiddh added a commit that referenced this pull request Aug 4, 2026
…els (#21546)
## Summary
The Arduino library this directory generates could not compile, and had
it compiled it could not have run a model.
Kernels never reached the operator registry — ExecuTorch registers them
through codegen and the build script never ran it, so every
`Method::load` would have failed with `OperatorMissing`. CMSIS-NN was
never vendored, so the Cortex-M ops shipped without the library they
call. `schema/*.cpp` was never copied, the `*_aten.cpp` exclusion also
deleted the portable-mode `tensor_parser_exec_aten.cpp`, and `__errno`
was missing because the Zephyr core mixes picolibc with newlib's
`libm_nano`.
Both `minimal.cpp` and `zephyr.cpp` were vendored, so which `et_pal_*`
backend you got depended on link order, and every `ET_LOG` was discarded
either way. One backend now ships and its logs route to a hook the
examples implement against `Serial`.
The examples ship their models — previously none did, so every sketch
`#error`ed when opened from the IDE menu. The README pointed at the
Ethos-U `pte_to_header.py`, whose `network_model_sec` section no Arduino
core defines, so following the docs produced a model that fails
`Program::load`. Static link mode is mandatory and undocumented; the
Dynamic default yields a silently dead board. Renamed to `ExecuTorch`
because `arduino-lint` rejects "Arduino" in an Arduino library's name.
Registering every portable kernel costs 1.58 MB against 786 KB of flash,
so the op set is a curated default overridable via `ROOT_OPS`/`ALL_OPS`.
Models and libraries must come from the same ExecuTorch commit —
Cortex-M schemas change (`scratch` in #19636, #19825), and a mismatch
loads fine, resolves every operator, then fails at `Method::execute`.
The library now records and pins that commit. CI to enforce it follows
separately.
## Test plan
`arduino:zephyr:unoq`, board core 0.55.2, `link_mode=static`, flashed on
hardware:
```
HelloExecuTorch Model loaded OK!, 1 method 60% flash, 20% RAM
AddModel [1,2,3] + 1 = [2.00, 3.00, 4.00] 64% flash, 26% RAM
KeywordSpotting 10/10 keywords correct 70% flash, 53% RAM
```
All ten MFCC inputs in one sketch, exercising the CMSIS-NN conv /
depthwise / avgpool / linear kernels:
```
yes 8.95 no 4.78 up 4.63 down 9.72 left 7.87
right 7.41 on 12.03 off 8.02 stop 8.64 go 9.57
```
`arduino-lint --library-manager submit`: no errors, no warnings, under
both `specification` and `strict`. Clean regeneration is byte-identical
across all 626 generated files.
Only the Uno Q was tested; the other three boards remain marked Planned.
Authored with Claude Code (Opus 5).
psiddh added a commit that referenced this pull request Aug 11, 2026
The library is generated from this repository, so an ExecuTorch change
can break it without touching examples/arduino at all. Nothing checked
that.
The job builds the library, verifies the shipped models against it,
compiles every example for arduino:zephyr:unoq, and runs arduino-lint in
Library Manager submission mode. It runs on changes under
examples/arduino, backends/cortex_m, kernels/portable, runtime, and
schema, which is the set that can break any of the four.
verify_models.py is the part worth having. A .pte records how many
values each operator call puts on the stack, and the generated kernel
wrappers reject anything else; the two only agree when the model and the
library came from the same commit. Cortex-M schemas change - scratch was
added to the conv operators in #19636 and #19825 - and a mismatch is
invisible until far too late, because the program loads, every operator
resolves, and then execute returns InvalidProgram naming nothing useful.
Tracking that down by hand took the better part of a day. The check
compares the two directly in about a second, needs no board, and
reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects
11
Compiling is deliberately not the whole job. Every failure worth finding
in this library compiled cleanly first, so the model check runs before
the sketches and arduino-lint guards the thing users actually install.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@rascani@digantdesai
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length \u003e 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Cortex-M backend: plan avg_pool2d scratch buffer AoT - #19825

Merged
rascani merged 5 commits into
pytorch:mainfrom
rascani:cortex-m-avgpool-aot-scratch
Jun 9, 2026
Merged

Cortex-M backend: plan avg_pool2d scratch buffer AoT#19825
rascani merged 5 commits into
pytorch:mainfrom
rascani:cortex-m-avgpool-aot-scratch

Conversation

@rascani

Copy link
Copy Markdown
Contributor

Summary

CMSIS-NN's DSP-variant arm_avgpool_s8 returns ARM_CMSIS_NN_ARG_ERROR when its ctx->buf is NULL and arm_avgpool_s8_get_buffer_size(...) is non-zero. The kernel hardcoded cmsis_ctx.buf = nullptr, which worked on Cortex-M55 because the MVE variant ignores ctx entirely, but failed on any DSP-class core (e.g. Cortex-M7). The AoT scratch-buffer planning system introduced for conv/depthwise-conv/transpose-conv/bmm missed quantized_avg_pool2d; this extends it to cover that op.

The quantized_avg_pool2d and .out schemas gain a Tensor scratch parameter. A new cmsis_nn_avgpool_buffer_size size function is registered, and avg_pool2d's lowering moves out of QuantizedOpFusionPass (which cannot create exir.memory.alloc nodes because it routes through ExportPass.call_operator) into ConvertToCortexMPass, alongside conv2d/bmm. The count_include_pad decomposition into an explicit cortex_m::pad node carries over to the new location. The kernel reads scratch.nbytes() and scratch.mutable_data_ptr<int8_t>() to wire cmsis_nn_context, with a CORTEX_M_ENABLE_RUNTIME_CHECKS-guarded assertion that the AoT and runtime buffer sizes agree — matching the conv2d pattern. The Python _NHWC_DIM_ORDER / to_physical_order helper that both passes need moves into passes_utils.py to avoid duplication.

Test plan

examples/arm/run.sh --model_name=mv2 --target=cortex-m7 --bundleio

A new dialect test exercises the ceil_mode=True fallback so that future refactors do not silently change which path it takes.

Authored with Claude.

@pytorch-bot

pytorch-botBot commented May 27, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/19825

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit 6d749a7 with merge base 189ffaa (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label May 27, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

CMSIS-NN's DSP-variant `arm_avgpool_s8` returns `ARM_CMSIS_NN_ARG_ERROR`
when its `ctx->buf` is NULL and `arm_avgpool_s8_get_buffer_size(...)` is
non-zero. The kernel hardcoded `cmsis_ctx.buf = nullptr`, which worked
on Cortex-M55 because the MVE variant ignores `ctx` entirely, but
failed on any DSP-class core (e.g. Cortex-M7). The AoT scratch-buffer
planning system introduced for conv/depthwise-conv/transpose-conv/bmm
missed `quantized_avg_pool2d`; this extends it to cover that op.
The `quantized_avg_pool2d` and `.out` schemas gain a `Tensor scratch`
parameter. A new `cmsis_nn_avgpool_buffer_size` size function is
registered, and avg_pool2d's lowering moves out of
`QuantizedOpFusionPass` (which cannot create `exir.memory.alloc` nodes
because it routes through `ExportPass.call_operator`) into
`ConvertToCortexMPass`, alongside conv2d/bmm. The `count_include_pad`
decomposition into an explicit `cortex_m::pad` node carries over to
the new location. The kernel reads `scratch.nbytes()` and
`scratch.mutable_data_ptr<int8_t>()` to wire `cmsis_nn_context`, with
a `CORTEX_M_ENABLE_RUNTIME_CHECKS`-guarded assertion that the AoT and
runtime buffer sizes agree — matching the conv2d pattern. The Python
`_NHWC_DIM_ORDER` / `to_physical_order` helper that both passes need
moves into `passes_utils.py` to avoid duplication.
A new dialect test exercises the `ceil_mode=True` fallback so that
future refactors do not silently change which path it takes.
Authored with Claude.
@rascani
rascaniforce-pushed the cortex-m-avgpool-aot-scratch branch from 4fd3c18 to 47e7fb8CompareJune 2, 2026 22:09
@rascani
rascani requested a review from digantdesaiJune 4, 2026 17:15
Comment threadbackends/cortex_m/test/ops/test_avg_pool2d.py
@digantdesai

Copy link
Copy Markdown
Contributor

@claude review this PR.

@claude

claudeBot commented Jun 9, 2026

Copy link
Copy Markdown

Claude finished @digantdesai's task in 3m 6s —— View job


PR Review: Cortex-M backend: plan avg_pool2d scratch buffer AoT

  • Read changed files
  • Analyze C++ kernel changes
  • Analyze Python pass changes
  • Analyze test coverage
  • Post review feedback

Overall: Clean, well-structured PR that follows the established scratch-buffer AoT planning pattern (conv2d, depthwise-conv, transpose-conv, bmm). The move from QuantizedOpFusionPass to ConvertToCortexMPass is well-motivated since the former routes through ExportPass.call_operator which cannot create exir.memory.alloc nodes. LGTM with minor observations below.


Findings

1. List aliasing in _get_avg_pool2d_replacement (convert_to_cortex_m_pass.py:514)

pre_pad=post_pad=to_physical_order([0, 0, pad_h, pad_w], input_tensor)

When is_channels_last(input_tensor) returns False, to_physical_order returns the same list object that was passed in (line 330 of passes_utils.py). Both pre_pad and post_pad then alias the same list. This is safe here since neither is mutated afterward, but the original code in QuantizedOpFusionPass explicitly created two independent lists. For symmetric padding the values are always equal, so this is correct but brittle if someone later mutates one of them.

Suggestion: pre_pad = post_pad = list(to_physical_order([0, 0, pad_h, pad_w], input_tensor)) — or just keep them as two separate calls for clarity. Very minor, not blocking.

2. No negative-return guard on arm_avgpool_s8_get_buffer_size (op_quantized_avg_pool2d.cpp:57-67)

The conv2d kernel's runtime check guards against negative return values from arm_convolve_wrapper_s8_get_buffer_size (which can return errors). The avgpool version skips this, which is correct — arm_avgpool_s8_get_buffer_size per the CMSIS-NN API always returns ≥ 0. Just noting this is intentional and consistent with the API contract.

3. Runtime check placement (op_quantized_avg_pool2d.cpp:56-68)

The runtime check validates the buffer size after the buffer pointer has already been assigned to cmsis_ctx.buf but beforearm_avgpool_s8 is called. This matches the conv2d pattern and is correct — the check is purely advisory/diagnostic and returns early before the kernel runs.

4. Reviewer question re: buffer size assertion in test (responding to @digantdesai's comment on line 108)

The test_dialect_avg_pool2d_fallback test verifies that ceil_mode=True correctly falls back to the aten op (leaving it unconverted). It doesn't exercise the scratch buffer path. The existingtest_dialect_avg_pool2d and test_implementation_avg_pool2d parametrized tests do exercise the scratch buffer path end-to-end (the _initialize_alloc_node_size call in ConvertToCortexMPass.call wires the buffer size). A unit test that directly asserts the buffer size value from cmsis_nn_avgpool_buffer_size would add confidence, but the existing e2e tests already catch mismatches (the kernel would fail at runtime with ARM_CMSIS_NN_ARG_ERROR).

5. Schema & YAML consistency — Verified: the .yaml, lib.define(...), @register_fake, @impl, and C++ signature all consistently include the new Tensor scratch parameter in the correct position.

6. Test coverage — Good. The PR adds:

  • fallback_test_cases + test_dialect_avg_pool2d_fallback: verifies ceil_mode=True stays as aten op
  • Existing parametrized tests (test_dialect_avg_pool2d, test_implementation_avg_pool2d) cover the happy path with the scratch buffer

Summary

This is a solid, well-scoped extension of the existing AoT scratch buffer system. The code correctly follows the established patterns, the kernel change is minimal and guarded, and the test coverage is reasonable. The only material suggestion is the minor list-aliasing nit in the pad decomposition path.

rascaniand others added 2 commits June 9, 2026 10:54
…t-scratch
# Conflicts:
#	backends/cortex_m/passes/convert_to_cortex_m_pass.py
Adapt _get_avg_pool2d_replacement from the old ConvertToCortexMPass class
to a @register_dialect_substitution function on AtenToCortexMPass,
matching the pattern established by pytorch#20070. Add casts to satisfy mypy's
node.args union type.
Add a structural assertion to test_dialect_avg_pool2d that verifies the
scratch buffer size matches cmsis_nn.avgpool_buffer_size() for the given
input/output dimensions.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@rascani
rascani merged commit 23b6ba0 into pytorch:mainJun 9, 2026
402 of 404 checks passed
@rascani
rascani deleted the cortex-m-avgpool-aot-scratch branch June 9, 2026 23:05
rascani added a commit that referenced this pull request Jun 9, 2026
)
Use `dialect_pass: AtenToDialectPass` parameter matching the
SubstitutionFn type updated in #19676. The `exported_program` parameter
caused a mypy arg-type error when both #19676 and #19825 landed on main.
Fixes the lintrunner-mypy failure on main.
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
psiddh added a commit to psiddh/executorch that referenced this pull request Aug 3, 2026
The library is generated from this repository, so an ExecuTorch change can
break it without touching examples/arduino at all. Nothing checked that.
The job builds the library, verifies the shipped models against it, compiles
every example for arduino:zephyr:unoq, and runs arduino-lint in Library
Manager submission mode. It runs on changes under examples/arduino,
backends/cortex_m, kernels/portable, runtime, and schema, which is the set
that can break any of the four.
verify_models.py is the part worth having. A .pte records how many values
each operator call puts on the stack, and the generated kernel wrappers
reject anything else; the two only agree when the model and the library came
from the same commit. Cortex-M schemas change - scratch was added to the conv
operators in pytorch#19636 and pytorch#19825 - and a mismatch is invisible until far too
late, because the program loads, every operator resolves, and then execute
returns InvalidProgram naming nothing useful. Tracking that down by hand took
the better part of a day. The check compares the two directly in about a
second, needs no board, and reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects 11
Compiling is deliberately not the whole job. Every failure worth finding in
this library compiled cleanly first, so the model check runs before the
sketches and arduino-lint guards the thing users actually install.
Authored with Claude Code (Opus 5).
psiddh added a commit to psiddh/executorch that referenced this pull request Aug 3, 2026
The library is generated from this repository, so an ExecuTorch change can
break it without touching examples/arduino at all. Nothing checked that.
The job builds the library, verifies the shipped models against it, compiles
every example for arduino:zephyr:unoq, and runs arduino-lint in Library
Manager submission mode. It runs on changes under examples/arduino,
backends/cortex_m, kernels/portable, runtime, and schema, which is the set
that can break any of the four.
verify_models.py is the part worth having. A .pte records how many values
each operator call puts on the stack, and the generated kernel wrappers
reject anything else; the two only agree when the model and the library came
from the same commit. Cortex-M schemas change - scratch was added to the conv
operators in pytorch#19636 and pytorch#19825 - and a mismatch is invisible until far too
late, because the program loads, every operator resolves, and then execute
returns InvalidProgram naming nothing useful. Tracking that down by hand took
the better part of a day. The check compares the two directly in about a
second, needs no board, and reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects 11
Compiling is deliberately not the whole job. Every failure worth finding in
this library compiled cleanly first, so the model check runs before the
sketches and arduino-lint guards the thing users actually install.
Authored with Claude Code (Opus 5).
psiddh added a commit to psiddh/executorch that referenced this pull request Aug 4, 2026
The library is generated from this repository, so an ExecuTorch change can
break it without touching examples/arduino at all. Nothing checked that.
The job builds the library, verifies the shipped models against it, compiles
every example for arduino:zephyr:unoq, and runs arduino-lint in Library
Manager submission mode. It runs on changes under examples/arduino,
backends/cortex_m, kernels/portable, runtime, and schema, which is the set
that can break any of the four.
verify_models.py is the part worth having. A .pte records how many values
each operator call puts on the stack, and the generated kernel wrappers
reject anything else; the two only agree when the model and the library came
from the same commit. Cortex-M schemas change - scratch was added to the conv
operators in pytorch#19636 and pytorch#19825 - and a mismatch is invisible until far too
late, because the program loads, every operator resolves, and then execute
returns InvalidProgram naming nothing useful. Tracking that down by hand took
the better part of a day. The check compares the two directly in about a
second, needs no board, and reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects 11
Compiling is deliberately not the whole job. Every failure worth finding in
this library compiled cleanly first, so the model check runs before the
sketches and arduino-lint guards the thing users actually install.
Authored with Claude Code (Opus 5).
psiddh added a commit that referenced this pull request Aug 4, 2026
…els (#21546)
## Summary
The Arduino library this directory generates could not compile, and had
it compiled it could not have run a model.
Kernels never reached the operator registry — ExecuTorch registers them
through codegen and the build script never ran it, so every
`Method::load` would have failed with `OperatorMissing`. CMSIS-NN was
never vendored, so the Cortex-M ops shipped without the library they
call. `schema/*.cpp` was never copied, the `*_aten.cpp` exclusion also
deleted the portable-mode `tensor_parser_exec_aten.cpp`, and `__errno`
was missing because the Zephyr core mixes picolibc with newlib's
`libm_nano`.
Both `minimal.cpp` and `zephyr.cpp` were vendored, so which `et_pal_*`
backend you got depended on link order, and every `ET_LOG` was discarded
either way. One backend now ships and its logs route to a hook the
examples implement against `Serial`.
The examples ship their models — previously none did, so every sketch
`#error`ed when opened from the IDE menu. The README pointed at the
Ethos-U `pte_to_header.py`, whose `network_model_sec` section no Arduino
core defines, so following the docs produced a model that fails
`Program::load`. Static link mode is mandatory and undocumented; the
Dynamic default yields a silently dead board. Renamed to `ExecuTorch`
because `arduino-lint` rejects "Arduino" in an Arduino library's name.
Registering every portable kernel costs 1.58 MB against 786 KB of flash,
so the op set is a curated default overridable via `ROOT_OPS`/`ALL_OPS`.
Models and libraries must come from the same ExecuTorch commit —
Cortex-M schemas change (`scratch` in #19636, #19825), and a mismatch
loads fine, resolves every operator, then fails at `Method::execute`.
The library now records and pins that commit. CI to enforce it follows
separately.
## Test plan
`arduino:zephyr:unoq`, board core 0.55.2, `link_mode=static`, flashed on
hardware:
```
HelloExecuTorch Model loaded OK!, 1 method 60% flash, 20% RAM
AddModel [1,2,3] + 1 = [2.00, 3.00, 4.00] 64% flash, 26% RAM
KeywordSpotting 10/10 keywords correct 70% flash, 53% RAM
```
All ten MFCC inputs in one sketch, exercising the CMSIS-NN conv /
depthwise / avgpool / linear kernels:
```
yes 8.95 no 4.78 up 4.63 down 9.72 left 7.87
right 7.41 on 12.03 off 8.02 stop 8.64 go 9.57
```
`arduino-lint --library-manager submit`: no errors, no warnings, under
both `specification` and `strict`. Clean regeneration is byte-identical
across all 626 generated files.
Only the Uno Q was tested; the other three boards remain marked Planned.
Authored with Claude Code (Opus 5).
psiddh added a commit that referenced this pull request Aug 11, 2026
The library is generated from this repository, so an ExecuTorch change
can break it without touching examples/arduino at all. Nothing checked
that.
The job builds the library, verifies the shipped models against it,
compiles every example for arduino:zephyr:unoq, and runs arduino-lint in
Library Manager submission mode. It runs on changes under
examples/arduino, backends/cortex_m, kernels/portable, runtime, and
schema, which is the set that can break any of the four.
verify_models.py is the part worth having. A .pte records how many
values each operator call puts on the stack, and the generated kernel
wrappers reject anything else; the two only agree when the model and the
library came from the same commit. Cortex-M schemas change - scratch was
added to the conv operators in #19636 and #19825 - and a mismatch is
invisible until far too late, because the program loads, every operator
resolves, and then execute returns InvalidProgram naming nothing useful.
Tracking that down by hand took the better part of a day. The check
compares the two directly in about a second, needs no board, and
reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects
11
Compiling is deliberately not the whole job. Every failure worth finding
in this library compiled cleanly first, so the model check runs before
the sketches and arduino-lint guards the thing users actually install.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@rascani@digantdesai
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Cortex-M backend: plan avg_pool2d scratch buffer AoT - #19825

Merged
rascani merged 5 commits into
pytorch:mainfrom
rascani:cortex-m-avgpool-aot-scratch
Jun 9, 2026
Merged

Cortex-M backend: plan avg_pool2d scratch buffer AoT#19825
rascani merged 5 commits into
pytorch:mainfrom
rascani:cortex-m-avgpool-aot-scratch

Conversation

@rascani

Copy link
Copy Markdown
Contributor

Summary

CMSIS-NN's DSP-variant arm_avgpool_s8 returns ARM_CMSIS_NN_ARG_ERROR when its ctx->buf is NULL and arm_avgpool_s8_get_buffer_size(...) is non-zero. The kernel hardcoded cmsis_ctx.buf = nullptr, which worked on Cortex-M55 because the MVE variant ignores ctx entirely, but failed on any DSP-class core (e.g. Cortex-M7). The AoT scratch-buffer planning system introduced for conv/depthwise-conv/transpose-conv/bmm missed quantized_avg_pool2d; this extends it to cover that op.

The quantized_avg_pool2d and .out schemas gain a Tensor scratch parameter. A new cmsis_nn_avgpool_buffer_size size function is registered, and avg_pool2d's lowering moves out of QuantizedOpFusionPass (which cannot create exir.memory.alloc nodes because it routes through ExportPass.call_operator) into ConvertToCortexMPass, alongside conv2d/bmm. The count_include_pad decomposition into an explicit cortex_m::pad node carries over to the new location. The kernel reads scratch.nbytes() and scratch.mutable_data_ptr<int8_t>() to wire cmsis_nn_context, with a CORTEX_M_ENABLE_RUNTIME_CHECKS-guarded assertion that the AoT and runtime buffer sizes agree — matching the conv2d pattern. The Python _NHWC_DIM_ORDER / to_physical_order helper that both passes need moves into passes_utils.py to avoid duplication.

Test plan

examples/arm/run.sh --model_name=mv2 --target=cortex-m7 --bundleio

A new dialect test exercises the ceil_mode=True fallback so that future refactors do not silently change which path it takes.

Authored with Claude.

@pytorch-bot

pytorch-botBot commented May 27, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/19825

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit 6d749a7 with merge base 189ffaa (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label May 27, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

CMSIS-NN's DSP-variant `arm_avgpool_s8` returns `ARM_CMSIS_NN_ARG_ERROR`
when its `ctx->buf` is NULL and `arm_avgpool_s8_get_buffer_size(...)` is
non-zero. The kernel hardcoded `cmsis_ctx.buf = nullptr`, which worked
on Cortex-M55 because the MVE variant ignores `ctx` entirely, but
failed on any DSP-class core (e.g. Cortex-M7). The AoT scratch-buffer
planning system introduced for conv/depthwise-conv/transpose-conv/bmm
missed `quantized_avg_pool2d`; this extends it to cover that op.
The `quantized_avg_pool2d` and `.out` schemas gain a `Tensor scratch`
parameter. A new `cmsis_nn_avgpool_buffer_size` size function is
registered, and avg_pool2d's lowering moves out of
`QuantizedOpFusionPass` (which cannot create `exir.memory.alloc` nodes
because it routes through `ExportPass.call_operator`) into
`ConvertToCortexMPass`, alongside conv2d/bmm. The `count_include_pad`
decomposition into an explicit `cortex_m::pad` node carries over to
the new location. The kernel reads `scratch.nbytes()` and
`scratch.mutable_data_ptr<int8_t>()` to wire `cmsis_nn_context`, with
a `CORTEX_M_ENABLE_RUNTIME_CHECKS`-guarded assertion that the AoT and
runtime buffer sizes agree — matching the conv2d pattern. The Python
`_NHWC_DIM_ORDER` / `to_physical_order` helper that both passes need
moves into `passes_utils.py` to avoid duplication.
A new dialect test exercises the `ceil_mode=True` fallback so that
future refactors do not silently change which path it takes.
Authored with Claude.
@rascani
rascaniforce-pushed the cortex-m-avgpool-aot-scratch branch from 4fd3c18 to 47e7fb8CompareJune 2, 2026 22:09
@rascani
rascani requested a review from digantdesaiJune 4, 2026 17:15
Comment threadbackends/cortex_m/test/ops/test_avg_pool2d.py
@digantdesai

Copy link
Copy Markdown
Contributor

@claude review this PR.

@claude

claudeBot commented Jun 9, 2026

Copy link
Copy Markdown

Claude finished @digantdesai's task in 3m 6s —— View job


PR Review: Cortex-M backend: plan avg_pool2d scratch buffer AoT

  • Read changed files
  • Analyze C++ kernel changes
  • Analyze Python pass changes
  • Analyze test coverage
  • Post review feedback

Overall: Clean, well-structured PR that follows the established scratch-buffer AoT planning pattern (conv2d, depthwise-conv, transpose-conv, bmm). The move from QuantizedOpFusionPass to ConvertToCortexMPass is well-motivated since the former routes through ExportPass.call_operator which cannot create exir.memory.alloc nodes. LGTM with minor observations below.


Findings

1. List aliasing in _get_avg_pool2d_replacement (convert_to_cortex_m_pass.py:514)

pre_pad=post_pad=to_physical_order([0, 0, pad_h, pad_w], input_tensor)

When is_channels_last(input_tensor) returns False, to_physical_order returns the same list object that was passed in (line 330 of passes_utils.py). Both pre_pad and post_pad then alias the same list. This is safe here since neither is mutated afterward, but the original code in QuantizedOpFusionPass explicitly created two independent lists. For symmetric padding the values are always equal, so this is correct but brittle if someone later mutates one of them.

Suggestion: pre_pad = post_pad = list(to_physical_order([0, 0, pad_h, pad_w], input_tensor)) — or just keep them as two separate calls for clarity. Very minor, not blocking.

2. No negative-return guard on arm_avgpool_s8_get_buffer_size (op_quantized_avg_pool2d.cpp:57-67)

The conv2d kernel's runtime check guards against negative return values from arm_convolve_wrapper_s8_get_buffer_size (which can return errors). The avgpool version skips this, which is correct — arm_avgpool_s8_get_buffer_size per the CMSIS-NN API always returns ≥ 0. Just noting this is intentional and consistent with the API contract.

3. Runtime check placement (op_quantized_avg_pool2d.cpp:56-68)

The runtime check validates the buffer size after the buffer pointer has already been assigned to cmsis_ctx.buf but beforearm_avgpool_s8 is called. This matches the conv2d pattern and is correct — the check is purely advisory/diagnostic and returns early before the kernel runs.

4. Reviewer question re: buffer size assertion in test (responding to @digantdesai's comment on line 108)

The test_dialect_avg_pool2d_fallback test verifies that ceil_mode=True correctly falls back to the aten op (leaving it unconverted). It doesn't exercise the scratch buffer path. The existingtest_dialect_avg_pool2d and test_implementation_avg_pool2d parametrized tests do exercise the scratch buffer path end-to-end (the _initialize_alloc_node_size call in ConvertToCortexMPass.call wires the buffer size). A unit test that directly asserts the buffer size value from cmsis_nn_avgpool_buffer_size would add confidence, but the existing e2e tests already catch mismatches (the kernel would fail at runtime with ARM_CMSIS_NN_ARG_ERROR).

5. Schema & YAML consistency — Verified: the .yaml, lib.define(...), @register_fake, @impl, and C++ signature all consistently include the new Tensor scratch parameter in the correct position.

6. Test coverage — Good. The PR adds:

  • fallback_test_cases + test_dialect_avg_pool2d_fallback: verifies ceil_mode=True stays as aten op
  • Existing parametrized tests (test_dialect_avg_pool2d, test_implementation_avg_pool2d) cover the happy path with the scratch buffer

Summary

This is a solid, well-scoped extension of the existing AoT scratch buffer system. The code correctly follows the established patterns, the kernel change is minimal and guarded, and the test coverage is reasonable. The only material suggestion is the minor list-aliasing nit in the pad decomposition path.

rascaniand others added 2 commits June 9, 2026 10:54
…t-scratch
# Conflicts:
#	backends/cortex_m/passes/convert_to_cortex_m_pass.py
Adapt _get_avg_pool2d_replacement from the old ConvertToCortexMPass class
to a @register_dialect_substitution function on AtenToCortexMPass,
matching the pattern established by pytorch#20070. Add casts to satisfy mypy's
node.args union type.
Add a structural assertion to test_dialect_avg_pool2d that verifies the
scratch buffer size matches cmsis_nn.avgpool_buffer_size() for the given
input/output dimensions.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@rascani
rascani merged commit 23b6ba0 into pytorch:mainJun 9, 2026
402 of 404 checks passed
@rascani
rascani deleted the cortex-m-avgpool-aot-scratch branch June 9, 2026 23:05
rascani added a commit that referenced this pull request Jun 9, 2026
)
Use `dialect_pass: AtenToDialectPass` parameter matching the
SubstitutionFn type updated in #19676. The `exported_program` parameter
caused a mypy arg-type error when both #19676 and #19825 landed on main.
Fixes the lintrunner-mypy failure on main.
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
psiddh added a commit to psiddh/executorch that referenced this pull request Aug 3, 2026
The library is generated from this repository, so an ExecuTorch change can
break it without touching examples/arduino at all. Nothing checked that.
The job builds the library, verifies the shipped models against it, compiles
every example for arduino:zephyr:unoq, and runs arduino-lint in Library
Manager submission mode. It runs on changes under examples/arduino,
backends/cortex_m, kernels/portable, runtime, and schema, which is the set
that can break any of the four.
verify_models.py is the part worth having. A .pte records how many values
each operator call puts on the stack, and the generated kernel wrappers
reject anything else; the two only agree when the model and the library came
from the same commit. Cortex-M schemas change - scratch was added to the conv
operators in pytorch#19636 and pytorch#19825 - and a mismatch is invisible until far too
late, because the program loads, every operator resolves, and then execute
returns InvalidProgram naming nothing useful. Tracking that down by hand took
the better part of a day. The check compares the two directly in about a
second, needs no board, and reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects 11
Compiling is deliberately not the whole job. Every failure worth finding in
this library compiled cleanly first, so the model check runs before the
sketches and arduino-lint guards the thing users actually install.
Authored with Claude Code (Opus 5).
psiddh added a commit to psiddh/executorch that referenced this pull request Aug 3, 2026
The library is generated from this repository, so an ExecuTorch change can
break it without touching examples/arduino at all. Nothing checked that.
The job builds the library, verifies the shipped models against it, compiles
every example for arduino:zephyr:unoq, and runs arduino-lint in Library
Manager submission mode. It runs on changes under examples/arduino,
backends/cortex_m, kernels/portable, runtime, and schema, which is the set
that can break any of the four.
verify_models.py is the part worth having. A .pte records how many values
each operator call puts on the stack, and the generated kernel wrappers
reject anything else; the two only agree when the model and the library came
from the same commit. Cortex-M schemas change - scratch was added to the conv
operators in pytorch#19636 and pytorch#19825 - and a mismatch is invisible until far too
late, because the program loads, every operator resolves, and then execute
returns InvalidProgram naming nothing useful. Tracking that down by hand took
the better part of a day. The check compares the two directly in about a
second, needs no board, and reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects 11
Compiling is deliberately not the whole job. Every failure worth finding in
this library compiled cleanly first, so the model check runs before the
sketches and arduino-lint guards the thing users actually install.
Authored with Claude Code (Opus 5).
psiddh added a commit to psiddh/executorch that referenced this pull request Aug 4, 2026
The library is generated from this repository, so an ExecuTorch change can
break it without touching examples/arduino at all. Nothing checked that.
The job builds the library, verifies the shipped models against it, compiles
every example for arduino:zephyr:unoq, and runs arduino-lint in Library
Manager submission mode. It runs on changes under examples/arduino,
backends/cortex_m, kernels/portable, runtime, and schema, which is the set
that can break any of the four.
verify_models.py is the part worth having. A .pte records how many values
each operator call puts on the stack, and the generated kernel wrappers
reject anything else; the two only agree when the model and the library came
from the same commit. Cortex-M schemas change - scratch was added to the conv
operators in pytorch#19636 and pytorch#19825 - and a mismatch is invisible until far too
late, because the program loads, every operator resolves, and then execute
returns InvalidProgram naming nothing useful. Tracking that down by hand took
the better part of a day. The check compares the two directly in about a
second, needs no board, and reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects 11
Compiling is deliberately not the whole job. Every failure worth finding in
this library compiled cleanly first, so the model check runs before the
sketches and arduino-lint guards the thing users actually install.
Authored with Claude Code (Opus 5).
psiddh added a commit that referenced this pull request Aug 4, 2026
…els (#21546)
## Summary
The Arduino library this directory generates could not compile, and had
it compiled it could not have run a model.
Kernels never reached the operator registry — ExecuTorch registers them
through codegen and the build script never ran it, so every
`Method::load` would have failed with `OperatorMissing`. CMSIS-NN was
never vendored, so the Cortex-M ops shipped without the library they
call. `schema/*.cpp` was never copied, the `*_aten.cpp` exclusion also
deleted the portable-mode `tensor_parser_exec_aten.cpp`, and `__errno`
was missing because the Zephyr core mixes picolibc with newlib's
`libm_nano`.
Both `minimal.cpp` and `zephyr.cpp` were vendored, so which `et_pal_*`
backend you got depended on link order, and every `ET_LOG` was discarded
either way. One backend now ships and its logs route to a hook the
examples implement against `Serial`.
The examples ship their models — previously none did, so every sketch
`#error`ed when opened from the IDE menu. The README pointed at the
Ethos-U `pte_to_header.py`, whose `network_model_sec` section no Arduino
core defines, so following the docs produced a model that fails
`Program::load`. Static link mode is mandatory and undocumented; the
Dynamic default yields a silently dead board. Renamed to `ExecuTorch`
because `arduino-lint` rejects "Arduino" in an Arduino library's name.
Registering every portable kernel costs 1.58 MB against 786 KB of flash,
so the op set is a curated default overridable via `ROOT_OPS`/`ALL_OPS`.
Models and libraries must come from the same ExecuTorch commit —
Cortex-M schemas change (`scratch` in #19636, #19825), and a mismatch
loads fine, resolves every operator, then fails at `Method::execute`.
The library now records and pins that commit. CI to enforce it follows
separately.
## Test plan
`arduino:zephyr:unoq`, board core 0.55.2, `link_mode=static`, flashed on
hardware:
```
HelloExecuTorch Model loaded OK!, 1 method 60% flash, 20% RAM
AddModel [1,2,3] + 1 = [2.00, 3.00, 4.00] 64% flash, 26% RAM
KeywordSpotting 10/10 keywords correct 70% flash, 53% RAM
```
All ten MFCC inputs in one sketch, exercising the CMSIS-NN conv /
depthwise / avgpool / linear kernels:
```
yes 8.95 no 4.78 up 4.63 down 9.72 left 7.87
right 7.41 on 12.03 off 8.02 stop 8.64 go 9.57
```
`arduino-lint --library-manager submit`: no errors, no warnings, under
both `specification` and `strict`. Clean regeneration is byte-identical
across all 626 generated files.
Only the Uno Q was tested; the other three boards remain marked Planned.
Authored with Claude Code (Opus 5).
psiddh added a commit that referenced this pull request Aug 11, 2026
The library is generated from this repository, so an ExecuTorch change
can break it without touching examples/arduino at all. Nothing checked
that.
The job builds the library, verifies the shipped models against it,
compiles every example for arduino:zephyr:unoq, and runs arduino-lint in
Library Manager submission mode. It runs on changes under
examples/arduino, backends/cortex_m, kernels/portable, runtime, and
schema, which is the set that can break any of the four.
verify_models.py is the part worth having. A .pte records how many
values each operator call puts on the stack, and the generated kernel
wrappers reject anything else; the two only agree when the model and the
library came from the same commit. Cortex-M schemas change - scratch was
added to the conv operators in #19636 and #19825 - and a mismatch is
invisible until far too late, because the program loads, every operator
resolves, and then execute returns InvalidProgram naming nothing useful.
Tracking that down by hand took the better part of a day. The check
compares the two directly in about a second, needs no board, and
reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects
11
Compiling is deliberately not the whole job. Every failure worth finding
in this library compiled cleanly first, so the model check runs before
the sketches and arduino-lint guards the thing users actually install.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@rascani@digantdesai
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Cortex-M backend: plan avg_pool2d scratch buffer AoT - #19825

Merged
rascani merged 5 commits into
pytorch:mainfrom
rascani:cortex-m-avgpool-aot-scratch
Jun 9, 2026
Merged

Cortex-M backend: plan avg_pool2d scratch buffer AoT#19825
rascani merged 5 commits into
pytorch:mainfrom
rascani:cortex-m-avgpool-aot-scratch

Conversation

@rascani

Copy link
Copy Markdown
Contributor

Summary

CMSIS-NN's DSP-variant arm_avgpool_s8 returns ARM_CMSIS_NN_ARG_ERROR when its ctx->buf is NULL and arm_avgpool_s8_get_buffer_size(...) is non-zero. The kernel hardcoded cmsis_ctx.buf = nullptr, which worked on Cortex-M55 because the MVE variant ignores ctx entirely, but failed on any DSP-class core (e.g. Cortex-M7). The AoT scratch-buffer planning system introduced for conv/depthwise-conv/transpose-conv/bmm missed quantized_avg_pool2d; this extends it to cover that op.

The quantized_avg_pool2d and .out schemas gain a Tensor scratch parameter. A new cmsis_nn_avgpool_buffer_size size function is registered, and avg_pool2d's lowering moves out of QuantizedOpFusionPass (which cannot create exir.memory.alloc nodes because it routes through ExportPass.call_operator) into ConvertToCortexMPass, alongside conv2d/bmm. The count_include_pad decomposition into an explicit cortex_m::pad node carries over to the new location. The kernel reads scratch.nbytes() and scratch.mutable_data_ptr<int8_t>() to wire cmsis_nn_context, with a CORTEX_M_ENABLE_RUNTIME_CHECKS-guarded assertion that the AoT and runtime buffer sizes agree — matching the conv2d pattern. The Python _NHWC_DIM_ORDER / to_physical_order helper that both passes need moves into passes_utils.py to avoid duplication.

Test plan

examples/arm/run.sh --model_name=mv2 --target=cortex-m7 --bundleio

A new dialect test exercises the ceil_mode=True fallback so that future refactors do not silently change which path it takes.

Authored with Claude.

@pytorch-bot

pytorch-botBot commented May 27, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/19825

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit 6d749a7 with merge base 189ffaa (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label May 27, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

CMSIS-NN's DSP-variant `arm_avgpool_s8` returns `ARM_CMSIS_NN_ARG_ERROR`
when its `ctx->buf` is NULL and `arm_avgpool_s8_get_buffer_size(...)` is
non-zero. The kernel hardcoded `cmsis_ctx.buf = nullptr`, which worked
on Cortex-M55 because the MVE variant ignores `ctx` entirely, but
failed on any DSP-class core (e.g. Cortex-M7). The AoT scratch-buffer
planning system introduced for conv/depthwise-conv/transpose-conv/bmm
missed `quantized_avg_pool2d`; this extends it to cover that op.
The `quantized_avg_pool2d` and `.out` schemas gain a `Tensor scratch`
parameter. A new `cmsis_nn_avgpool_buffer_size` size function is
registered, and avg_pool2d's lowering moves out of
`QuantizedOpFusionPass` (which cannot create `exir.memory.alloc` nodes
because it routes through `ExportPass.call_operator`) into
`ConvertToCortexMPass`, alongside conv2d/bmm. The `count_include_pad`
decomposition into an explicit `cortex_m::pad` node carries over to
the new location. The kernel reads `scratch.nbytes()` and
`scratch.mutable_data_ptr<int8_t>()` to wire `cmsis_nn_context`, with
a `CORTEX_M_ENABLE_RUNTIME_CHECKS`-guarded assertion that the AoT and
runtime buffer sizes agree — matching the conv2d pattern. The Python
`_NHWC_DIM_ORDER` / `to_physical_order` helper that both passes need
moves into `passes_utils.py` to avoid duplication.
A new dialect test exercises the `ceil_mode=True` fallback so that
future refactors do not silently change which path it takes.
Authored with Claude.
@rascani
rascaniforce-pushed the cortex-m-avgpool-aot-scratch branch from 4fd3c18 to 47e7fb8CompareJune 2, 2026 22:09
@rascani
rascani requested a review from digantdesaiJune 4, 2026 17:15
Comment threadbackends/cortex_m/test/ops/test_avg_pool2d.py
@digantdesai

Copy link
Copy Markdown
Contributor

@claude review this PR.

@claude

claudeBot commented Jun 9, 2026

Copy link
Copy Markdown

Claude finished @digantdesai's task in 3m 6s —— View job


PR Review: Cortex-M backend: plan avg_pool2d scratch buffer AoT

  • Read changed files
  • Analyze C++ kernel changes
  • Analyze Python pass changes
  • Analyze test coverage
  • Post review feedback

Overall: Clean, well-structured PR that follows the established scratch-buffer AoT planning pattern (conv2d, depthwise-conv, transpose-conv, bmm). The move from QuantizedOpFusionPass to ConvertToCortexMPass is well-motivated since the former routes through ExportPass.call_operator which cannot create exir.memory.alloc nodes. LGTM with minor observations below.


Findings

1. List aliasing in _get_avg_pool2d_replacement (convert_to_cortex_m_pass.py:514)

pre_pad=post_pad=to_physical_order([0, 0, pad_h, pad_w], input_tensor)

When is_channels_last(input_tensor) returns False, to_physical_order returns the same list object that was passed in (line 330 of passes_utils.py). Both pre_pad and post_pad then alias the same list. This is safe here since neither is mutated afterward, but the original code in QuantizedOpFusionPass explicitly created two independent lists. For symmetric padding the values are always equal, so this is correct but brittle if someone later mutates one of them.

Suggestion: pre_pad = post_pad = list(to_physical_order([0, 0, pad_h, pad_w], input_tensor)) — or just keep them as two separate calls for clarity. Very minor, not blocking.

2. No negative-return guard on arm_avgpool_s8_get_buffer_size (op_quantized_avg_pool2d.cpp:57-67)

The conv2d kernel's runtime check guards against negative return values from arm_convolve_wrapper_s8_get_buffer_size (which can return errors). The avgpool version skips this, which is correct — arm_avgpool_s8_get_buffer_size per the CMSIS-NN API always returns ≥ 0. Just noting this is intentional and consistent with the API contract.

3. Runtime check placement (op_quantized_avg_pool2d.cpp:56-68)

The runtime check validates the buffer size after the buffer pointer has already been assigned to cmsis_ctx.buf but beforearm_avgpool_s8 is called. This matches the conv2d pattern and is correct — the check is purely advisory/diagnostic and returns early before the kernel runs.

4. Reviewer question re: buffer size assertion in test (responding to @digantdesai's comment on line 108)

The test_dialect_avg_pool2d_fallback test verifies that ceil_mode=True correctly falls back to the aten op (leaving it unconverted). It doesn't exercise the scratch buffer path. The existingtest_dialect_avg_pool2d and test_implementation_avg_pool2d parametrized tests do exercise the scratch buffer path end-to-end (the _initialize_alloc_node_size call in ConvertToCortexMPass.call wires the buffer size). A unit test that directly asserts the buffer size value from cmsis_nn_avgpool_buffer_size would add confidence, but the existing e2e tests already catch mismatches (the kernel would fail at runtime with ARM_CMSIS_NN_ARG_ERROR).

5. Schema & YAML consistency — Verified: the .yaml, lib.define(...), @register_fake, @impl, and C++ signature all consistently include the new Tensor scratch parameter in the correct position.

6. Test coverage — Good. The PR adds:

  • fallback_test_cases + test_dialect_avg_pool2d_fallback: verifies ceil_mode=True stays as aten op
  • Existing parametrized tests (test_dialect_avg_pool2d, test_implementation_avg_pool2d) cover the happy path with the scratch buffer

Summary

This is a solid, well-scoped extension of the existing AoT scratch buffer system. The code correctly follows the established patterns, the kernel change is minimal and guarded, and the test coverage is reasonable. The only material suggestion is the minor list-aliasing nit in the pad decomposition path.

rascaniand others added 2 commits June 9, 2026 10:54
…t-scratch
# Conflicts:
#	backends/cortex_m/passes/convert_to_cortex_m_pass.py
Adapt _get_avg_pool2d_replacement from the old ConvertToCortexMPass class
to a @register_dialect_substitution function on AtenToCortexMPass,
matching the pattern established by pytorch#20070. Add casts to satisfy mypy's
node.args union type.
Add a structural assertion to test_dialect_avg_pool2d that verifies the
scratch buffer size matches cmsis_nn.avgpool_buffer_size() for the given
input/output dimensions.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@rascani
rascani merged commit 23b6ba0 into pytorch:mainJun 9, 2026
402 of 404 checks passed
@rascani
rascani deleted the cortex-m-avgpool-aot-scratch branch June 9, 2026 23:05
rascani added a commit that referenced this pull request Jun 9, 2026
)
Use `dialect_pass: AtenToDialectPass` parameter matching the
SubstitutionFn type updated in #19676. The `exported_program` parameter
caused a mypy arg-type error when both #19676 and #19825 landed on main.
Fixes the lintrunner-mypy failure on main.
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
psiddh added a commit to psiddh/executorch that referenced this pull request Aug 3, 2026
The library is generated from this repository, so an ExecuTorch change can
break it without touching examples/arduino at all. Nothing checked that.
The job builds the library, verifies the shipped models against it, compiles
every example for arduino:zephyr:unoq, and runs arduino-lint in Library
Manager submission mode. It runs on changes under examples/arduino,
backends/cortex_m, kernels/portable, runtime, and schema, which is the set
that can break any of the four.
verify_models.py is the part worth having. A .pte records how many values
each operator call puts on the stack, and the generated kernel wrappers
reject anything else; the two only agree when the model and the library came
from the same commit. Cortex-M schemas change - scratch was added to the conv
operators in pytorch#19636 and pytorch#19825 - and a mismatch is invisible until far too
late, because the program loads, every operator resolves, and then execute
returns InvalidProgram naming nothing useful. Tracking that down by hand took
the better part of a day. The check compares the two directly in about a
second, needs no board, and reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects 11
Compiling is deliberately not the whole job. Every failure worth finding in
this library compiled cleanly first, so the model check runs before the
sketches and arduino-lint guards the thing users actually install.
Authored with Claude Code (Opus 5).
psiddh added a commit to psiddh/executorch that referenced this pull request Aug 3, 2026
The library is generated from this repository, so an ExecuTorch change can
break it without touching examples/arduino at all. Nothing checked that.
The job builds the library, verifies the shipped models against it, compiles
every example for arduino:zephyr:unoq, and runs arduino-lint in Library
Manager submission mode. It runs on changes under examples/arduino,
backends/cortex_m, kernels/portable, runtime, and schema, which is the set
that can break any of the four.
verify_models.py is the part worth having. A .pte records how many values
each operator call puts on the stack, and the generated kernel wrappers
reject anything else; the two only agree when the model and the library came
from the same commit. Cortex-M schemas change - scratch was added to the conv
operators in pytorch#19636 and pytorch#19825 - and a mismatch is invisible until far too
late, because the program loads, every operator resolves, and then execute
returns InvalidProgram naming nothing useful. Tracking that down by hand took
the better part of a day. The check compares the two directly in about a
second, needs no board, and reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects 11
Compiling is deliberately not the whole job. Every failure worth finding in
this library compiled cleanly first, so the model check runs before the
sketches and arduino-lint guards the thing users actually install.
Authored with Claude Code (Opus 5).
psiddh added a commit to psiddh/executorch that referenced this pull request Aug 4, 2026
The library is generated from this repository, so an ExecuTorch change can
break it without touching examples/arduino at all. Nothing checked that.
The job builds the library, verifies the shipped models against it, compiles
every example for arduino:zephyr:unoq, and runs arduino-lint in Library
Manager submission mode. It runs on changes under examples/arduino,
backends/cortex_m, kernels/portable, runtime, and schema, which is the set
that can break any of the four.
verify_models.py is the part worth having. A .pte records how many values
each operator call puts on the stack, and the generated kernel wrappers
reject anything else; the two only agree when the model and the library came
from the same commit. Cortex-M schemas change - scratch was added to the conv
operators in pytorch#19636 and pytorch#19825 - and a mismatch is invisible until far too
late, because the program loads, every operator resolves, and then execute
returns InvalidProgram naming nothing useful. Tracking that down by hand took
the better part of a day. The check compares the two directly in about a
second, needs no board, and reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects 11
Compiling is deliberately not the whole job. Every failure worth finding in
this library compiled cleanly first, so the model check runs before the
sketches and arduino-lint guards the thing users actually install.
Authored with Claude Code (Opus 5).
psiddh added a commit that referenced this pull request Aug 4, 2026
…els (#21546)
## Summary
The Arduino library this directory generates could not compile, and had
it compiled it could not have run a model.
Kernels never reached the operator registry — ExecuTorch registers them
through codegen and the build script never ran it, so every
`Method::load` would have failed with `OperatorMissing`. CMSIS-NN was
never vendored, so the Cortex-M ops shipped without the library they
call. `schema/*.cpp` was never copied, the `*_aten.cpp` exclusion also
deleted the portable-mode `tensor_parser_exec_aten.cpp`, and `__errno`
was missing because the Zephyr core mixes picolibc with newlib's
`libm_nano`.
Both `minimal.cpp` and `zephyr.cpp` were vendored, so which `et_pal_*`
backend you got depended on link order, and every `ET_LOG` was discarded
either way. One backend now ships and its logs route to a hook the
examples implement against `Serial`.
The examples ship their models — previously none did, so every sketch
`#error`ed when opened from the IDE menu. The README pointed at the
Ethos-U `pte_to_header.py`, whose `network_model_sec` section no Arduino
core defines, so following the docs produced a model that fails
`Program::load`. Static link mode is mandatory and undocumented; the
Dynamic default yields a silently dead board. Renamed to `ExecuTorch`
because `arduino-lint` rejects "Arduino" in an Arduino library's name.
Registering every portable kernel costs 1.58 MB against 786 KB of flash,
so the op set is a curated default overridable via `ROOT_OPS`/`ALL_OPS`.
Models and libraries must come from the same ExecuTorch commit —
Cortex-M schemas change (`scratch` in #19636, #19825), and a mismatch
loads fine, resolves every operator, then fails at `Method::execute`.
The library now records and pins that commit. CI to enforce it follows
separately.
## Test plan
`arduino:zephyr:unoq`, board core 0.55.2, `link_mode=static`, flashed on
hardware:
```
HelloExecuTorch Model loaded OK!, 1 method 60% flash, 20% RAM
AddModel [1,2,3] + 1 = [2.00, 3.00, 4.00] 64% flash, 26% RAM
KeywordSpotting 10/10 keywords correct 70% flash, 53% RAM
```
All ten MFCC inputs in one sketch, exercising the CMSIS-NN conv /
depthwise / avgpool / linear kernels:
```
yes 8.95 no 4.78 up 4.63 down 9.72 left 7.87
right 7.41 on 12.03 off 8.02 stop 8.64 go 9.57
```
`arduino-lint --library-manager submit`: no errors, no warnings, under
both `specification` and `strict`. Clean regeneration is byte-identical
across all 626 generated files.
Only the Uno Q was tested; the other three boards remain marked Planned.
Authored with Claude Code (Opus 5).
psiddh added a commit that referenced this pull request Aug 11, 2026
The library is generated from this repository, so an ExecuTorch change
can break it without touching examples/arduino at all. Nothing checked
that.
The job builds the library, verifies the shipped models against it,
compiles every example for arduino:zephyr:unoq, and runs arduino-lint in
Library Manager submission mode. It runs on changes under
examples/arduino, backends/cortex_m, kernels/portable, runtime, and
schema, which is the set that can break any of the four.
verify_models.py is the part worth having. A .pte records how many
values each operator call puts on the stack, and the generated kernel
wrappers reject anything else; the two only agree when the model and the
library came from the same commit. Cortex-M schemas change - scratch was
added to the conv operators in #19636 and #19825 - and a mismatch is
invisible until far too late, because the program loads, every operator
resolves, and then execute returns InvalidProgram naming nothing useful.
Tracking that down by hand took the better part of a day. The check
compares the two directly in about a second, needs no board, and
reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects
11
Compiling is deliberately not the whole job. Every failure worth finding
in this library compiled cleanly first, so the model check runs before
the sketches and arduino-lint guards the thing users actually install.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@rascani@digantdesai
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Cortex-M backend: plan avg_pool2d scratch buffer AoT - #19825

Merged
rascani merged 5 commits into
pytorch:mainfrom
rascani:cortex-m-avgpool-aot-scratch
Jun 9, 2026
Merged

Cortex-M backend: plan avg_pool2d scratch buffer AoT#19825
rascani merged 5 commits into
pytorch:mainfrom
rascani:cortex-m-avgpool-aot-scratch

Conversation

@rascani

Copy link
Copy Markdown
Contributor

Summary

CMSIS-NN's DSP-variant arm_avgpool_s8 returns ARM_CMSIS_NN_ARG_ERROR when its ctx->buf is NULL and arm_avgpool_s8_get_buffer_size(...) is non-zero. The kernel hardcoded cmsis_ctx.buf = nullptr, which worked on Cortex-M55 because the MVE variant ignores ctx entirely, but failed on any DSP-class core (e.g. Cortex-M7). The AoT scratch-buffer planning system introduced for conv/depthwise-conv/transpose-conv/bmm missed quantized_avg_pool2d; this extends it to cover that op.

The quantized_avg_pool2d and .out schemas gain a Tensor scratch parameter. A new cmsis_nn_avgpool_buffer_size size function is registered, and avg_pool2d's lowering moves out of QuantizedOpFusionPass (which cannot create exir.memory.alloc nodes because it routes through ExportPass.call_operator) into ConvertToCortexMPass, alongside conv2d/bmm. The count_include_pad decomposition into an explicit cortex_m::pad node carries over to the new location. The kernel reads scratch.nbytes() and scratch.mutable_data_ptr<int8_t>() to wire cmsis_nn_context, with a CORTEX_M_ENABLE_RUNTIME_CHECKS-guarded assertion that the AoT and runtime buffer sizes agree — matching the conv2d pattern. The Python _NHWC_DIM_ORDER / to_physical_order helper that both passes need moves into passes_utils.py to avoid duplication.

Test plan

examples/arm/run.sh --model_name=mv2 --target=cortex-m7 --bundleio

A new dialect test exercises the ceil_mode=True fallback so that future refactors do not silently change which path it takes.

Authored with Claude.

@pytorch-bot

pytorch-botBot commented May 27, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/19825

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit 6d749a7 with merge base 189ffaa (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label May 27, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

CMSIS-NN's DSP-variant `arm_avgpool_s8` returns `ARM_CMSIS_NN_ARG_ERROR`
when its `ctx->buf` is NULL and `arm_avgpool_s8_get_buffer_size(...)` is
non-zero. The kernel hardcoded `cmsis_ctx.buf = nullptr`, which worked
on Cortex-M55 because the MVE variant ignores `ctx` entirely, but
failed on any DSP-class core (e.g. Cortex-M7). The AoT scratch-buffer
planning system introduced for conv/depthwise-conv/transpose-conv/bmm
missed `quantized_avg_pool2d`; this extends it to cover that op.
The `quantized_avg_pool2d` and `.out` schemas gain a `Tensor scratch`
parameter. A new `cmsis_nn_avgpool_buffer_size` size function is
registered, and avg_pool2d's lowering moves out of
`QuantizedOpFusionPass` (which cannot create `exir.memory.alloc` nodes
because it routes through `ExportPass.call_operator`) into
`ConvertToCortexMPass`, alongside conv2d/bmm. The `count_include_pad`
decomposition into an explicit `cortex_m::pad` node carries over to
the new location. The kernel reads `scratch.nbytes()` and
`scratch.mutable_data_ptr<int8_t>()` to wire `cmsis_nn_context`, with
a `CORTEX_M_ENABLE_RUNTIME_CHECKS`-guarded assertion that the AoT and
runtime buffer sizes agree — matching the conv2d pattern. The Python
`_NHWC_DIM_ORDER` / `to_physical_order` helper that both passes need
moves into `passes_utils.py` to avoid duplication.
A new dialect test exercises the `ceil_mode=True` fallback so that
future refactors do not silently change which path it takes.
Authored with Claude.
@rascani
rascaniforce-pushed the cortex-m-avgpool-aot-scratch branch from 4fd3c18 to 47e7fb8CompareJune 2, 2026 22:09
@rascani
rascani requested a review from digantdesaiJune 4, 2026 17:15
Comment threadbackends/cortex_m/test/ops/test_avg_pool2d.py
@digantdesai

Copy link
Copy Markdown
Contributor

@claude review this PR.

@claude

claudeBot commented Jun 9, 2026

Copy link
Copy Markdown

Claude finished @digantdesai's task in 3m 6s —— View job


PR Review: Cortex-M backend: plan avg_pool2d scratch buffer AoT

  • Read changed files
  • Analyze C++ kernel changes
  • Analyze Python pass changes
  • Analyze test coverage
  • Post review feedback

Overall: Clean, well-structured PR that follows the established scratch-buffer AoT planning pattern (conv2d, depthwise-conv, transpose-conv, bmm). The move from QuantizedOpFusionPass to ConvertToCortexMPass is well-motivated since the former routes through ExportPass.call_operator which cannot create exir.memory.alloc nodes. LGTM with minor observations below.


Findings

1. List aliasing in _get_avg_pool2d_replacement (convert_to_cortex_m_pass.py:514)

pre_pad=post_pad=to_physical_order([0, 0, pad_h, pad_w], input_tensor)

When is_channels_last(input_tensor) returns False, to_physical_order returns the same list object that was passed in (line 330 of passes_utils.py). Both pre_pad and post_pad then alias the same list. This is safe here since neither is mutated afterward, but the original code in QuantizedOpFusionPass explicitly created two independent lists. For symmetric padding the values are always equal, so this is correct but brittle if someone later mutates one of them.

Suggestion: pre_pad = post_pad = list(to_physical_order([0, 0, pad_h, pad_w], input_tensor)) — or just keep them as two separate calls for clarity. Very minor, not blocking.

2. No negative-return guard on arm_avgpool_s8_get_buffer_size (op_quantized_avg_pool2d.cpp:57-67)

The conv2d kernel's runtime check guards against negative return values from arm_convolve_wrapper_s8_get_buffer_size (which can return errors). The avgpool version skips this, which is correct — arm_avgpool_s8_get_buffer_size per the CMSIS-NN API always returns ≥ 0. Just noting this is intentional and consistent with the API contract.

3. Runtime check placement (op_quantized_avg_pool2d.cpp:56-68)

The runtime check validates the buffer size after the buffer pointer has already been assigned to cmsis_ctx.buf but beforearm_avgpool_s8 is called. This matches the conv2d pattern and is correct — the check is purely advisory/diagnostic and returns early before the kernel runs.

4. Reviewer question re: buffer size assertion in test (responding to @digantdesai's comment on line 108)

The test_dialect_avg_pool2d_fallback test verifies that ceil_mode=True correctly falls back to the aten op (leaving it unconverted). It doesn't exercise the scratch buffer path. The existingtest_dialect_avg_pool2d and test_implementation_avg_pool2d parametrized tests do exercise the scratch buffer path end-to-end (the _initialize_alloc_node_size call in ConvertToCortexMPass.call wires the buffer size). A unit test that directly asserts the buffer size value from cmsis_nn_avgpool_buffer_size would add confidence, but the existing e2e tests already catch mismatches (the kernel would fail at runtime with ARM_CMSIS_NN_ARG_ERROR).

5. Schema & YAML consistency — Verified: the .yaml, lib.define(...), @register_fake, @impl, and C++ signature all consistently include the new Tensor scratch parameter in the correct position.

6. Test coverage — Good. The PR adds:

  • fallback_test_cases + test_dialect_avg_pool2d_fallback: verifies ceil_mode=True stays as aten op
  • Existing parametrized tests (test_dialect_avg_pool2d, test_implementation_avg_pool2d) cover the happy path with the scratch buffer

Summary

This is a solid, well-scoped extension of the existing AoT scratch buffer system. The code correctly follows the established patterns, the kernel change is minimal and guarded, and the test coverage is reasonable. The only material suggestion is the minor list-aliasing nit in the pad decomposition path.

rascaniand others added 2 commits June 9, 2026 10:54
…t-scratch
# Conflicts:
#	backends/cortex_m/passes/convert_to_cortex_m_pass.py
Adapt _get_avg_pool2d_replacement from the old ConvertToCortexMPass class
to a @register_dialect_substitution function on AtenToCortexMPass,
matching the pattern established by pytorch#20070. Add casts to satisfy mypy's
node.args union type.
Add a structural assertion to test_dialect_avg_pool2d that verifies the
scratch buffer size matches cmsis_nn.avgpool_buffer_size() for the given
input/output dimensions.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@rascani
rascani merged commit 23b6ba0 into pytorch:mainJun 9, 2026
402 of 404 checks passed
@rascani
rascani deleted the cortex-m-avgpool-aot-scratch branch June 9, 2026 23:05
rascani added a commit that referenced this pull request Jun 9, 2026
)
Use `dialect_pass: AtenToDialectPass` parameter matching the
SubstitutionFn type updated in #19676. The `exported_program` parameter
caused a mypy arg-type error when both #19676 and #19825 landed on main.
Fixes the lintrunner-mypy failure on main.
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
psiddh added a commit to psiddh/executorch that referenced this pull request Aug 3, 2026
The library is generated from this repository, so an ExecuTorch change can
break it without touching examples/arduino at all. Nothing checked that.
The job builds the library, verifies the shipped models against it, compiles
every example for arduino:zephyr:unoq, and runs arduino-lint in Library
Manager submission mode. It runs on changes under examples/arduino,
backends/cortex_m, kernels/portable, runtime, and schema, which is the set
that can break any of the four.
verify_models.py is the part worth having. A .pte records how many values
each operator call puts on the stack, and the generated kernel wrappers
reject anything else; the two only agree when the model and the library came
from the same commit. Cortex-M schemas change - scratch was added to the conv
operators in pytorch#19636 and pytorch#19825 - and a mismatch is invisible until far too
late, because the program loads, every operator resolves, and then execute
returns InvalidProgram naming nothing useful. Tracking that down by hand took
the better part of a day. The check compares the two directly in about a
second, needs no board, and reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects 11
Compiling is deliberately not the whole job. Every failure worth finding in
this library compiled cleanly first, so the model check runs before the
sketches and arduino-lint guards the thing users actually install.
Authored with Claude Code (Opus 5).
psiddh added a commit to psiddh/executorch that referenced this pull request Aug 3, 2026
The library is generated from this repository, so an ExecuTorch change can
break it without touching examples/arduino at all. Nothing checked that.
The job builds the library, verifies the shipped models against it, compiles
every example for arduino:zephyr:unoq, and runs arduino-lint in Library
Manager submission mode. It runs on changes under examples/arduino,
backends/cortex_m, kernels/portable, runtime, and schema, which is the set
that can break any of the four.
verify_models.py is the part worth having. A .pte records how many values
each operator call puts on the stack, and the generated kernel wrappers
reject anything else; the two only agree when the model and the library came
from the same commit. Cortex-M schemas change - scratch was added to the conv
operators in pytorch#19636 and pytorch#19825 - and a mismatch is invisible until far too
late, because the program loads, every operator resolves, and then execute
returns InvalidProgram naming nothing useful. Tracking that down by hand took
the better part of a day. The check compares the two directly in about a
second, needs no board, and reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects 11
Compiling is deliberately not the whole job. Every failure worth finding in
this library compiled cleanly first, so the model check runs before the
sketches and arduino-lint guards the thing users actually install.
Authored with Claude Code (Opus 5).
psiddh added a commit to psiddh/executorch that referenced this pull request Aug 4, 2026
The library is generated from this repository, so an ExecuTorch change can
break it without touching examples/arduino at all. Nothing checked that.
The job builds the library, verifies the shipped models against it, compiles
every example for arduino:zephyr:unoq, and runs arduino-lint in Library
Manager submission mode. It runs on changes under examples/arduino,
backends/cortex_m, kernels/portable, runtime, and schema, which is the set
that can break any of the four.
verify_models.py is the part worth having. A .pte records how many values
each operator call puts on the stack, and the generated kernel wrappers
reject anything else; the two only agree when the model and the library came
from the same commit. Cortex-M schemas change - scratch was added to the conv
operators in pytorch#19636 and pytorch#19825 - and a mismatch is invisible until far too
late, because the program loads, every operator resolves, and then execute
returns InvalidProgram naming nothing useful. Tracking that down by hand took
the better part of a day. The check compares the two directly in about a
second, needs no board, and reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects 11
Compiling is deliberately not the whole job. Every failure worth finding in
this library compiled cleanly first, so the model check runs before the
sketches and arduino-lint guards the thing users actually install.
Authored with Claude Code (Opus 5).
psiddh added a commit that referenced this pull request Aug 4, 2026
…els (#21546)
## Summary
The Arduino library this directory generates could not compile, and had
it compiled it could not have run a model.
Kernels never reached the operator registry — ExecuTorch registers them
through codegen and the build script never ran it, so every
`Method::load` would have failed with `OperatorMissing`. CMSIS-NN was
never vendored, so the Cortex-M ops shipped without the library they
call. `schema/*.cpp` was never copied, the `*_aten.cpp` exclusion also
deleted the portable-mode `tensor_parser_exec_aten.cpp`, and `__errno`
was missing because the Zephyr core mixes picolibc with newlib's
`libm_nano`.
Both `minimal.cpp` and `zephyr.cpp` were vendored, so which `et_pal_*`
backend you got depended on link order, and every `ET_LOG` was discarded
either way. One backend now ships and its logs route to a hook the
examples implement against `Serial`.
The examples ship their models — previously none did, so every sketch
`#error`ed when opened from the IDE menu. The README pointed at the
Ethos-U `pte_to_header.py`, whose `network_model_sec` section no Arduino
core defines, so following the docs produced a model that fails
`Program::load`. Static link mode is mandatory and undocumented; the
Dynamic default yields a silently dead board. Renamed to `ExecuTorch`
because `arduino-lint` rejects "Arduino" in an Arduino library's name.
Registering every portable kernel costs 1.58 MB against 786 KB of flash,
so the op set is a curated default overridable via `ROOT_OPS`/`ALL_OPS`.
Models and libraries must come from the same ExecuTorch commit —
Cortex-M schemas change (`scratch` in #19636, #19825), and a mismatch
loads fine, resolves every operator, then fails at `Method::execute`.
The library now records and pins that commit. CI to enforce it follows
separately.
## Test plan
`arduino:zephyr:unoq`, board core 0.55.2, `link_mode=static`, flashed on
hardware:
```
HelloExecuTorch Model loaded OK!, 1 method 60% flash, 20% RAM
AddModel [1,2,3] + 1 = [2.00, 3.00, 4.00] 64% flash, 26% RAM
KeywordSpotting 10/10 keywords correct 70% flash, 53% RAM
```
All ten MFCC inputs in one sketch, exercising the CMSIS-NN conv /
depthwise / avgpool / linear kernels:
```
yes 8.95 no 4.78 up 4.63 down 9.72 left 7.87
right 7.41 on 12.03 off 8.02 stop 8.64 go 9.57
```
`arduino-lint --library-manager submit`: no errors, no warnings, under
both `specification` and `strict`. Clean regeneration is byte-identical
across all 626 generated files.
Only the Uno Q was tested; the other three boards remain marked Planned.
Authored with Claude Code (Opus 5).
psiddh added a commit that referenced this pull request Aug 11, 2026
The library is generated from this repository, so an ExecuTorch change
can break it without touching examples/arduino at all. Nothing checked
that.
The job builds the library, verifies the shipped models against it,
compiles every example for arduino:zephyr:unoq, and runs arduino-lint in
Library Manager submission mode. It runs on changes under
examples/arduino, backends/cortex_m, kernels/portable, runtime, and
schema, which is the set that can break any of the four.
verify_models.py is the part worth having. A .pte records how many
values each operator call puts on the stack, and the generated kernel
wrappers reject anything else; the two only agree when the model and the
library came from the same commit. Cortex-M schemas change - scratch was
added to the conv operators in #19636 and #19825 - and a mismatch is
invisible until far too late, because the program loads, every operator
resolves, and then execute returns InvalidProgram naming nothing useful.
Tracking that down by hand took the better part of a day. The check
compares the two directly in about a second, needs no board, and
reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects
11
Compiling is deliberately not the whole job. Every failure worth finding
in this library compiled cleanly first, so the model check runs before
the sketches and arduino-lint guards the thing users actually install.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@rascani@digantdesai
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Cortex-M backend: plan avg_pool2d scratch buffer AoT - #19825

Merged
rascani merged 5 commits into
pytorch:mainfrom
rascani:cortex-m-avgpool-aot-scratch
Jun 9, 2026
Merged

Cortex-M backend: plan avg_pool2d scratch buffer AoT#19825
rascani merged 5 commits into
pytorch:mainfrom
rascani:cortex-m-avgpool-aot-scratch

Conversation

@rascani

Copy link
Copy Markdown
Contributor

Summary

CMSIS-NN's DSP-variant arm_avgpool_s8 returns ARM_CMSIS_NN_ARG_ERROR when its ctx->buf is NULL and arm_avgpool_s8_get_buffer_size(...) is non-zero. The kernel hardcoded cmsis_ctx.buf = nullptr, which worked on Cortex-M55 because the MVE variant ignores ctx entirely, but failed on any DSP-class core (e.g. Cortex-M7). The AoT scratch-buffer planning system introduced for conv/depthwise-conv/transpose-conv/bmm missed quantized_avg_pool2d; this extends it to cover that op.

The quantized_avg_pool2d and .out schemas gain a Tensor scratch parameter. A new cmsis_nn_avgpool_buffer_size size function is registered, and avg_pool2d's lowering moves out of QuantizedOpFusionPass (which cannot create exir.memory.alloc nodes because it routes through ExportPass.call_operator) into ConvertToCortexMPass, alongside conv2d/bmm. The count_include_pad decomposition into an explicit cortex_m::pad node carries over to the new location. The kernel reads scratch.nbytes() and scratch.mutable_data_ptr<int8_t>() to wire cmsis_nn_context, with a CORTEX_M_ENABLE_RUNTIME_CHECKS-guarded assertion that the AoT and runtime buffer sizes agree — matching the conv2d pattern. The Python _NHWC_DIM_ORDER / to_physical_order helper that both passes need moves into passes_utils.py to avoid duplication.

Test plan

examples/arm/run.sh --model_name=mv2 --target=cortex-m7 --bundleio

A new dialect test exercises the ceil_mode=True fallback so that future refactors do not silently change which path it takes.

Authored with Claude.

@pytorch-bot

pytorch-botBot commented May 27, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/19825

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit 6d749a7 with merge base 189ffaa (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label May 27, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

CMSIS-NN's DSP-variant `arm_avgpool_s8` returns `ARM_CMSIS_NN_ARG_ERROR`
when its `ctx->buf` is NULL and `arm_avgpool_s8_get_buffer_size(...)` is
non-zero. The kernel hardcoded `cmsis_ctx.buf = nullptr`, which worked
on Cortex-M55 because the MVE variant ignores `ctx` entirely, but
failed on any DSP-class core (e.g. Cortex-M7). The AoT scratch-buffer
planning system introduced for conv/depthwise-conv/transpose-conv/bmm
missed `quantized_avg_pool2d`; this extends it to cover that op.
The `quantized_avg_pool2d` and `.out` schemas gain a `Tensor scratch`
parameter. A new `cmsis_nn_avgpool_buffer_size` size function is
registered, and avg_pool2d's lowering moves out of
`QuantizedOpFusionPass` (which cannot create `exir.memory.alloc` nodes
because it routes through `ExportPass.call_operator`) into
`ConvertToCortexMPass`, alongside conv2d/bmm. The `count_include_pad`
decomposition into an explicit `cortex_m::pad` node carries over to
the new location. The kernel reads `scratch.nbytes()` and
`scratch.mutable_data_ptr<int8_t>()` to wire `cmsis_nn_context`, with
a `CORTEX_M_ENABLE_RUNTIME_CHECKS`-guarded assertion that the AoT and
runtime buffer sizes agree — matching the conv2d pattern. The Python
`_NHWC_DIM_ORDER` / `to_physical_order` helper that both passes need
moves into `passes_utils.py` to avoid duplication.
A new dialect test exercises the `ceil_mode=True` fallback so that
future refactors do not silently change which path it takes.
Authored with Claude.
@rascani
rascaniforce-pushed the cortex-m-avgpool-aot-scratch branch from 4fd3c18 to 47e7fb8CompareJune 2, 2026 22:09
@rascani
rascani requested a review from digantdesaiJune 4, 2026 17:15
Comment threadbackends/cortex_m/test/ops/test_avg_pool2d.py
@digantdesai

Copy link
Copy Markdown
Contributor

@claude review this PR.

@claude

claudeBot commented Jun 9, 2026

Copy link
Copy Markdown

Claude finished @digantdesai's task in 3m 6s —— View job


PR Review: Cortex-M backend: plan avg_pool2d scratch buffer AoT

  • Read changed files
  • Analyze C++ kernel changes
  • Analyze Python pass changes
  • Analyze test coverage
  • Post review feedback

Overall: Clean, well-structured PR that follows the established scratch-buffer AoT planning pattern (conv2d, depthwise-conv, transpose-conv, bmm). The move from QuantizedOpFusionPass to ConvertToCortexMPass is well-motivated since the former routes through ExportPass.call_operator which cannot create exir.memory.alloc nodes. LGTM with minor observations below.


Findings

1. List aliasing in _get_avg_pool2d_replacement (convert_to_cortex_m_pass.py:514)

pre_pad=post_pad=to_physical_order([0, 0, pad_h, pad_w], input_tensor)

When is_channels_last(input_tensor) returns False, to_physical_order returns the same list object that was passed in (line 330 of passes_utils.py). Both pre_pad and post_pad then alias the same list. This is safe here since neither is mutated afterward, but the original code in QuantizedOpFusionPass explicitly created two independent lists. For symmetric padding the values are always equal, so this is correct but brittle if someone later mutates one of them.

Suggestion: pre_pad = post_pad = list(to_physical_order([0, 0, pad_h, pad_w], input_tensor)) — or just keep them as two separate calls for clarity. Very minor, not blocking.

2. No negative-return guard on arm_avgpool_s8_get_buffer_size (op_quantized_avg_pool2d.cpp:57-67)

The conv2d kernel's runtime check guards against negative return values from arm_convolve_wrapper_s8_get_buffer_size (which can return errors). The avgpool version skips this, which is correct — arm_avgpool_s8_get_buffer_size per the CMSIS-NN API always returns ≥ 0. Just noting this is intentional and consistent with the API contract.

3. Runtime check placement (op_quantized_avg_pool2d.cpp:56-68)

The runtime check validates the buffer size after the buffer pointer has already been assigned to cmsis_ctx.buf but beforearm_avgpool_s8 is called. This matches the conv2d pattern and is correct — the check is purely advisory/diagnostic and returns early before the kernel runs.

4. Reviewer question re: buffer size assertion in test (responding to @digantdesai's comment on line 108)

The test_dialect_avg_pool2d_fallback test verifies that ceil_mode=True correctly falls back to the aten op (leaving it unconverted). It doesn't exercise the scratch buffer path. The existingtest_dialect_avg_pool2d and test_implementation_avg_pool2d parametrized tests do exercise the scratch buffer path end-to-end (the _initialize_alloc_node_size call in ConvertToCortexMPass.call wires the buffer size). A unit test that directly asserts the buffer size value from cmsis_nn_avgpool_buffer_size would add confidence, but the existing e2e tests already catch mismatches (the kernel would fail at runtime with ARM_CMSIS_NN_ARG_ERROR).

5. Schema & YAML consistency — Verified: the .yaml, lib.define(...), @register_fake, @impl, and C++ signature all consistently include the new Tensor scratch parameter in the correct position.

6. Test coverage — Good. The PR adds:

  • fallback_test_cases + test_dialect_avg_pool2d_fallback: verifies ceil_mode=True stays as aten op
  • Existing parametrized tests (test_dialect_avg_pool2d, test_implementation_avg_pool2d) cover the happy path with the scratch buffer

Summary

This is a solid, well-scoped extension of the existing AoT scratch buffer system. The code correctly follows the established patterns, the kernel change is minimal and guarded, and the test coverage is reasonable. The only material suggestion is the minor list-aliasing nit in the pad decomposition path.

rascaniand others added 2 commits June 9, 2026 10:54
…t-scratch
# Conflicts:
#	backends/cortex_m/passes/convert_to_cortex_m_pass.py
Adapt _get_avg_pool2d_replacement from the old ConvertToCortexMPass class
to a @register_dialect_substitution function on AtenToCortexMPass,
matching the pattern established by pytorch#20070. Add casts to satisfy mypy's
node.args union type.
Add a structural assertion to test_dialect_avg_pool2d that verifies the
scratch buffer size matches cmsis_nn.avgpool_buffer_size() for the given
input/output dimensions.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@rascani
rascani merged commit 23b6ba0 into pytorch:mainJun 9, 2026
402 of 404 checks passed
@rascani
rascani deleted the cortex-m-avgpool-aot-scratch branch June 9, 2026 23:05
rascani added a commit that referenced this pull request Jun 9, 2026
)
Use `dialect_pass: AtenToDialectPass` parameter matching the
SubstitutionFn type updated in #19676. The `exported_program` parameter
caused a mypy arg-type error when both #19676 and #19825 landed on main.
Fixes the lintrunner-mypy failure on main.
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
psiddh added a commit to psiddh/executorch that referenced this pull request Aug 3, 2026
The library is generated from this repository, so an ExecuTorch change can
break it without touching examples/arduino at all. Nothing checked that.
The job builds the library, verifies the shipped models against it, compiles
every example for arduino:zephyr:unoq, and runs arduino-lint in Library
Manager submission mode. It runs on changes under examples/arduino,
backends/cortex_m, kernels/portable, runtime, and schema, which is the set
that can break any of the four.
verify_models.py is the part worth having. A .pte records how many values
each operator call puts on the stack, and the generated kernel wrappers
reject anything else; the two only agree when the model and the library came
from the same commit. Cortex-M schemas change - scratch was added to the conv
operators in pytorch#19636 and pytorch#19825 - and a mismatch is invisible until far too
late, because the program loads, every operator resolves, and then execute
returns InvalidProgram naming nothing useful. Tracking that down by hand took
the better part of a day. The check compares the two directly in about a
second, needs no board, and reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects 11
Compiling is deliberately not the whole job. Every failure worth finding in
this library compiled cleanly first, so the model check runs before the
sketches and arduino-lint guards the thing users actually install.
Authored with Claude Code (Opus 5).
psiddh added a commit to psiddh/executorch that referenced this pull request Aug 3, 2026
The library is generated from this repository, so an ExecuTorch change can
break it without touching examples/arduino at all. Nothing checked that.
The job builds the library, verifies the shipped models against it, compiles
every example for arduino:zephyr:unoq, and runs arduino-lint in Library
Manager submission mode. It runs on changes under examples/arduino,
backends/cortex_m, kernels/portable, runtime, and schema, which is the set
that can break any of the four.
verify_models.py is the part worth having. A .pte records how many values
each operator call puts on the stack, and the generated kernel wrappers
reject anything else; the two only agree when the model and the library came
from the same commit. Cortex-M schemas change - scratch was added to the conv
operators in pytorch#19636 and pytorch#19825 - and a mismatch is invisible until far too
late, because the program loads, every operator resolves, and then execute
returns InvalidProgram naming nothing useful. Tracking that down by hand took
the better part of a day. The check compares the two directly in about a
second, needs no board, and reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects 11
Compiling is deliberately not the whole job. Every failure worth finding in
this library compiled cleanly first, so the model check runs before the
sketches and arduino-lint guards the thing users actually install.
Authored with Claude Code (Opus 5).
psiddh added a commit to psiddh/executorch that referenced this pull request Aug 4, 2026
The library is generated from this repository, so an ExecuTorch change can
break it without touching examples/arduino at all. Nothing checked that.
The job builds the library, verifies the shipped models against it, compiles
every example for arduino:zephyr:unoq, and runs arduino-lint in Library
Manager submission mode. It runs on changes under examples/arduino,
backends/cortex_m, kernels/portable, runtime, and schema, which is the set
that can break any of the four.
verify_models.py is the part worth having. A .pte records how many values
each operator call puts on the stack, and the generated kernel wrappers
reject anything else; the two only agree when the model and the library came
from the same commit. Cortex-M schemas change - scratch was added to the conv
operators in pytorch#19636 and pytorch#19825 - and a mismatch is invisible until far too
late, because the program loads, every operator resolves, and then execute
returns InvalidProgram naming nothing useful. Tracking that down by hand took
the better part of a day. The check compares the two directly in about a
second, needs no board, and reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects 11
Compiling is deliberately not the whole job. Every failure worth finding in
this library compiled cleanly first, so the model check runs before the
sketches and arduino-lint guards the thing users actually install.
Authored with Claude Code (Opus 5).
psiddh added a commit that referenced this pull request Aug 4, 2026
…els (#21546)
## Summary
The Arduino library this directory generates could not compile, and had
it compiled it could not have run a model.
Kernels never reached the operator registry — ExecuTorch registers them
through codegen and the build script never ran it, so every
`Method::load` would have failed with `OperatorMissing`. CMSIS-NN was
never vendored, so the Cortex-M ops shipped without the library they
call. `schema/*.cpp` was never copied, the `*_aten.cpp` exclusion also
deleted the portable-mode `tensor_parser_exec_aten.cpp`, and `__errno`
was missing because the Zephyr core mixes picolibc with newlib's
`libm_nano`.
Both `minimal.cpp` and `zephyr.cpp` were vendored, so which `et_pal_*`
backend you got depended on link order, and every `ET_LOG` was discarded
either way. One backend now ships and its logs route to a hook the
examples implement against `Serial`.
The examples ship their models — previously none did, so every sketch
`#error`ed when opened from the IDE menu. The README pointed at the
Ethos-U `pte_to_header.py`, whose `network_model_sec` section no Arduino
core defines, so following the docs produced a model that fails
`Program::load`. Static link mode is mandatory and undocumented; the
Dynamic default yields a silently dead board. Renamed to `ExecuTorch`
because `arduino-lint` rejects "Arduino" in an Arduino library's name.
Registering every portable kernel costs 1.58 MB against 786 KB of flash,
so the op set is a curated default overridable via `ROOT_OPS`/`ALL_OPS`.
Models and libraries must come from the same ExecuTorch commit —
Cortex-M schemas change (`scratch` in #19636, #19825), and a mismatch
loads fine, resolves every operator, then fails at `Method::execute`.
The library now records and pins that commit. CI to enforce it follows
separately.
## Test plan
`arduino:zephyr:unoq`, board core 0.55.2, `link_mode=static`, flashed on
hardware:
```
HelloExecuTorch Model loaded OK!, 1 method 60% flash, 20% RAM
AddModel [1,2,3] + 1 = [2.00, 3.00, 4.00] 64% flash, 26% RAM
KeywordSpotting 10/10 keywords correct 70% flash, 53% RAM
```
All ten MFCC inputs in one sketch, exercising the CMSIS-NN conv /
depthwise / avgpool / linear kernels:
```
yes 8.95 no 4.78 up 4.63 down 9.72 left 7.87
right 7.41 on 12.03 off 8.02 stop 8.64 go 9.57
```
`arduino-lint --library-manager submit`: no errors, no warnings, under
both `specification` and `strict`. Clean regeneration is byte-identical
across all 626 generated files.
Only the Uno Q was tested; the other three boards remain marked Planned.
Authored with Claude Code (Opus 5).
psiddh added a commit that referenced this pull request Aug 11, 2026
The library is generated from this repository, so an ExecuTorch change
can break it without touching examples/arduino at all. Nothing checked
that.
The job builds the library, verifies the shipped models against it,
compiles every example for arduino:zephyr:unoq, and runs arduino-lint in
Library Manager submission mode. It runs on changes under
examples/arduino, backends/cortex_m, kernels/portable, runtime, and
schema, which is the set that can break any of the four.
verify_models.py is the part worth having. A .pte records how many
values each operator call puts on the stack, and the generated kernel
wrappers reject anything else; the two only agree when the model and the
library came from the same commit. Cortex-M schemas change - scratch was
added to the conv operators in #19636 and #19825 - and a mismatch is
invisible until far too late, because the program loads, every operator
resolves, and then execute returns InvalidProgram naming nothing useful.
Tracking that down by hand took the better part of a day. The check
compares the two directly in about a second, needs no board, and
reports:
FAIL KeywordSpotting
cortex_m::quantized_avg_pool2d.out: model supplies 10, library expects
11
Compiling is deliberately not the whole job. Every failure worth finding
in this library compiled cleanly first, so the model check runs before
the sketches and arduino-lint guards the thing users actually install.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@rascani@digantdesai