Delegate even-kernel 'same'-padding convs via a quantized static pad (#20553) - #20553

Merged
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
JakeStevens:export-D109871964
Jul 8, 2026
Merged

Delegate even-kernel 'same'-padding convs via a quantized static pad (#20553)#20553
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
JakeStevens:export-D109871964

Conversation

@JakeStevens

@JakeStevensJakeStevens commented Jun 26, 2026

Copy link
Copy Markdown
Contributor

Summary:

When using padding="same" for a conv with even kernels, torch.export creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in xnn_status_unsupported_parameter.

This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (ConvConfig._get_act_deps) so both delegate together: for fp32 the pad lowers directly as an XNNStaticConstantPad. For quantized graphs the pad is created by to_edge decomposition after the quantizer runs, so it arrives as dq -> pad -> conv with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). InsertPadQDQPass inserts an implicit quantize -> dequantize after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.

Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.

Differential Revision: D109871964

@pytorch-bot

pytorch-botBot commented Jun 26, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20553

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Pending, 2 Unrelated Failures

As of commit 0e81ab4 with merge base 3801496 (image):

NEW FAILURE - The following job has failed:

BROKEN TRUNK - The following jobs failed but were present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 26, 2026
@meta-codesync

Copy link
Copy Markdown
Contributor

@JakeStevens has exported this pull request. If you are a Meta employee, you can view the originating Diff in D109871964.

@meta-codesyncmeta-codesyncBot changed the title Fold constant_pad_nd into convolution input paddingDelegate even-kernel 'same'-padding convs via a quantized static pad (#20553)Jul 6, 2026
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 6, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
@JakeStevensJakeStevens added the release notes: xnnpack Changes to the XNNPack backend delegate label Jul 6, 2026
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 6, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 7, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 7, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964

An even-kernel 'same'-padding conv decomposes (after quantization) into
dequant -> constant_pad_nd -> convolution. Because the pad is introduced by
to_edge decomposition -- after the quantizer has run -- it is never annotated,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can we fuse it inside the preprocess for xnnpack or we have to run pad + conv in the XNNPACK graph? I am asking because if we can fuse it we might as well write that pass instead.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can fuse, I have some work in that direction, was going to publish as a quick follow up since there is an external contributor with the 1d work and the change is slightly larger so preferred to do it once these both land.

…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
@meta-codesync
meta-codesyncBot merged commit 420a8ab into pytorch:mainJul 8, 2026
181 of 184 checks passed
@JakeStevens
JakeStevens deleted the export-D109871964 branch July 9, 2026 13:41
SakshamKapoor2911 added a commit to SakshamKapoor2911/executorch that referenced this pull request Jul 9, 2026
 approach)
Instead of folding constant_pad_nd into asymmetric Conv2d padding,
keep it as an explicit XNNPACK pad op:
- Remove pad-folding logic from Conv1dUnsqueezePass
- Remove Conv1dFoldedPadMetaPass (no longer needed)
- Revert xnnpack_input_padding handling in op_conv2d.py
- Add InsertPadQDQPass: inserts QDQ after pad in quantized contexts
so it serializes as a quantized static pad
This matches the approach from pytorch#20553 and fixes the crash with
conv1d -> flatten (even kernel + same padding).
SakshamKapoor2911 added a commit to SakshamKapoor2911/executorch that referenced this pull request Jul 11, 2026
 approach)
Instead of folding constant_pad_nd into asymmetric Conv2d padding,
keep it as an explicit XNNPACK pad op:
- Remove pad-folding logic from Conv1dUnsqueezePass
- Remove Conv1dFoldedPadMetaPass (no longer needed)
- Revert xnnpack_input_padding handling in op_conv2d.py
- Add InsertPadQDQPass: inserts QDQ after pad in quantized contexts
so it serializes as a quantized static pad
This matches the approach from pytorch#20553 and fixes the crash with
conv1d -> flatten (even kernel + same padding).
JakeStevens added a commit that referenced this pull request Jul 16, 2026
Fixes#20558.
Related to #20553.
## Summary
Quantized `nn.Conv1d(..., padding="same")` with an even kernel exports
with
asymmetric padding (unequal left/right amounts), which cannot be folded
into
the convolution's symmetric padding field. The original pad-folding
approach
in this PR was replaced with the explicit-PAD approach from #20553:
- **InsertPadQDQPass** — inserts implicit quantize/dequantize pairs
after
constant_pad_nd nodes in quantized contexts so they serialize as
quantized
static pads. Refactored with additional guard checks (pad_value,
pad_amounts,
negative amounts) and correct idempotency.
- **ConvolutionConfig._get_act_deps** — pulls zero-valued
constant_pad_nd nodes
(and their QDQ chain if InsertPadQDQPass already ran) into the
convolution's
partition for both 1D and 2D convs. The merged method replaces separate
1D
and 2D implementations that existed on this branch.
- **Regression test** — quantized Conv1d even-kernel same-padding
covering both
symmetric and asymmetric pad cases, validated by numerical comparison.
- Removed unrelated __init__.py re-ordering and no-op
conv1d_unsqueeze_pass
changes.
With these changes, an even-kernel `padding="same"` conv1d graph:
```
dequant -> constant_pad_nd -> convolution
```
becomes (after XNNPACK preprocessing and partitioning):
```
[dequant -> pad -> q -> dq -> conv] (single XNNPACK delegate)
```
## Test plan
```
python -m pytest backends/xnnpack/test/ops/test_conv1d.py -q
# 7 passed (1 new regression test with 4 subTests)
python -m pytest backends/xnnpack/test/ops/test_conv2d.py -q
# 32 passed
python -m pytest backends/xnnpack/test/passes/test_insert_pad_qdq.py -q
# 3 passed
python -m pytest backends/xnnpack/test/ops/test_static_constant_pad.py -q
# 8 passed
lintrunner -a
# No lint issues
```
cc @GregoryComer@digantdesai@cbilgin@JakeStevens@freddan80@per@zingo@oscarandersson8218@mansnils@Sebastian-Larsson@robell@rascani
---------
Co-authored-by: Jacob Stevens <stevens.jacob1492@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exportedrelease notes: xnnpackChanges to the XNNPack backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Conv2d with padding 'same' fails (when quantised) at runtime on XNNPACK

3 participants

@JakeStevens@digantdesai@GregoryComer
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all \u003cpre\u003e\u003ccode\u003e blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks"); } } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); } })(); (function(){ try { var __m = "github.com"; var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Delegate even-kernel 'same'-padding convs via a quantized static pad (#20553) - #20553

Merged
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
JakeStevens:export-D109871964
Jul 8, 2026
Merged

Delegate even-kernel 'same'-padding convs via a quantized static pad (#20553)#20553
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
JakeStevens:export-D109871964

Conversation

@JakeStevens

@JakeStevensJakeStevens commented Jun 26, 2026

Copy link
Copy Markdown
Contributor

Summary:

When using padding="same" for a conv with even kernels, torch.export creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in xnn_status_unsupported_parameter.

This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (ConvConfig._get_act_deps) so both delegate together: for fp32 the pad lowers directly as an XNNStaticConstantPad. For quantized graphs the pad is created by to_edge decomposition after the quantizer runs, so it arrives as dq -> pad -> conv with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). InsertPadQDQPass inserts an implicit quantize -> dequantize after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.

Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.

Differential Revision: D109871964

@pytorch-bot

pytorch-botBot commented Jun 26, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20553

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Pending, 2 Unrelated Failures

As of commit 0e81ab4 with merge base 3801496 (image):

NEW FAILURE - The following job has failed:

BROKEN TRUNK - The following jobs failed but were present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 26, 2026
@meta-codesync

Copy link
Copy Markdown
Contributor

@JakeStevens has exported this pull request. If you are a Meta employee, you can view the originating Diff in D109871964.

@meta-codesyncmeta-codesyncBot changed the title Fold constant_pad_nd into convolution input paddingDelegate even-kernel 'same'-padding convs via a quantized static pad (#20553)Jul 6, 2026
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 6, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
@JakeStevensJakeStevens added the release notes: xnnpack Changes to the XNNPack backend delegate label Jul 6, 2026
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 6, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 7, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 7, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964

An even-kernel 'same'-padding conv decomposes (after quantization) into
dequant -> constant_pad_nd -> convolution. Because the pad is introduced by
to_edge decomposition -- after the quantizer has run -- it is never annotated,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can we fuse it inside the preprocess for xnnpack or we have to run pad + conv in the XNNPACK graph? I am asking because if we can fuse it we might as well write that pass instead.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can fuse, I have some work in that direction, was going to publish as a quick follow up since there is an external contributor with the 1d work and the change is slightly larger so preferred to do it once these both land.

…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
@meta-codesync
meta-codesyncBot merged commit 420a8ab into pytorch:mainJul 8, 2026
181 of 184 checks passed
@JakeStevens
JakeStevens deleted the export-D109871964 branch July 9, 2026 13:41
SakshamKapoor2911 added a commit to SakshamKapoor2911/executorch that referenced this pull request Jul 9, 2026
 approach)
Instead of folding constant_pad_nd into asymmetric Conv2d padding,
keep it as an explicit XNNPACK pad op:
- Remove pad-folding logic from Conv1dUnsqueezePass
- Remove Conv1dFoldedPadMetaPass (no longer needed)
- Revert xnnpack_input_padding handling in op_conv2d.py
- Add InsertPadQDQPass: inserts QDQ after pad in quantized contexts
so it serializes as a quantized static pad
This matches the approach from pytorch#20553 and fixes the crash with
conv1d -> flatten (even kernel + same padding).
SakshamKapoor2911 added a commit to SakshamKapoor2911/executorch that referenced this pull request Jul 11, 2026
 approach)
Instead of folding constant_pad_nd into asymmetric Conv2d padding,
keep it as an explicit XNNPACK pad op:
- Remove pad-folding logic from Conv1dUnsqueezePass
- Remove Conv1dFoldedPadMetaPass (no longer needed)
- Revert xnnpack_input_padding handling in op_conv2d.py
- Add InsertPadQDQPass: inserts QDQ after pad in quantized contexts
so it serializes as a quantized static pad
This matches the approach from pytorch#20553 and fixes the crash with
conv1d -> flatten (even kernel + same padding).
JakeStevens added a commit that referenced this pull request Jul 16, 2026
Fixes#20558.
Related to #20553.
## Summary
Quantized `nn.Conv1d(..., padding="same")` with an even kernel exports
with
asymmetric padding (unequal left/right amounts), which cannot be folded
into
the convolution's symmetric padding field. The original pad-folding
approach
in this PR was replaced with the explicit-PAD approach from #20553:
- **InsertPadQDQPass** — inserts implicit quantize/dequantize pairs
after
constant_pad_nd nodes in quantized contexts so they serialize as
quantized
static pads. Refactored with additional guard checks (pad_value,
pad_amounts,
negative amounts) and correct idempotency.
- **ConvolutionConfig._get_act_deps** — pulls zero-valued
constant_pad_nd nodes
(and their QDQ chain if InsertPadQDQPass already ran) into the
convolution's
partition for both 1D and 2D convs. The merged method replaces separate
1D
and 2D implementations that existed on this branch.
- **Regression test** — quantized Conv1d even-kernel same-padding
covering both
symmetric and asymmetric pad cases, validated by numerical comparison.
- Removed unrelated __init__.py re-ordering and no-op
conv1d_unsqueeze_pass
changes.
With these changes, an even-kernel `padding="same"` conv1d graph:
```
dequant -> constant_pad_nd -> convolution
```
becomes (after XNNPACK preprocessing and partitioning):
```
[dequant -> pad -> q -> dq -> conv] (single XNNPACK delegate)
```
## Test plan
```
python -m pytest backends/xnnpack/test/ops/test_conv1d.py -q
# 7 passed (1 new regression test with 4 subTests)
python -m pytest backends/xnnpack/test/ops/test_conv2d.py -q
# 32 passed
python -m pytest backends/xnnpack/test/passes/test_insert_pad_qdq.py -q
# 3 passed
python -m pytest backends/xnnpack/test/ops/test_static_constant_pad.py -q
# 8 passed
lintrunner -a
# No lint issues
```
cc @GregoryComer@digantdesai@cbilgin@JakeStevens@freddan80@per@zingo@oscarandersson8218@mansnils@Sebastian-Larsson@robell@rascani
---------
Co-authored-by: Jacob Stevens <stevens.jacob1492@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exportedrelease notes: xnnpackChanges to the XNNPack backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Conv2d with padding 'same' fails (when quantised) at runtime on XNNPACK

3 participants

@JakeStevens@digantdesai@GregoryComer
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Delegate even-kernel 'same'-padding convs via a quantized static pad (#20553) - #20553

Merged
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
JakeStevens:export-D109871964
Jul 8, 2026
Merged

Delegate even-kernel 'same'-padding convs via a quantized static pad (#20553)#20553
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
JakeStevens:export-D109871964

Conversation

@JakeStevens

@JakeStevensJakeStevens commented Jun 26, 2026

Copy link
Copy Markdown
Contributor

Summary:

When using padding="same" for a conv with even kernels, torch.export creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in xnn_status_unsupported_parameter.

This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (ConvConfig._get_act_deps) so both delegate together: for fp32 the pad lowers directly as an XNNStaticConstantPad. For quantized graphs the pad is created by to_edge decomposition after the quantizer runs, so it arrives as dq -> pad -> conv with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). InsertPadQDQPass inserts an implicit quantize -> dequantize after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.

Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.

Differential Revision: D109871964

@pytorch-bot

pytorch-botBot commented Jun 26, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20553

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Pending, 2 Unrelated Failures

As of commit 0e81ab4 with merge base 3801496 (image):

NEW FAILURE - The following job has failed:

BROKEN TRUNK - The following jobs failed but were present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 26, 2026
@meta-codesync

Copy link
Copy Markdown
Contributor

@JakeStevens has exported this pull request. If you are a Meta employee, you can view the originating Diff in D109871964.

@meta-codesyncmeta-codesyncBot changed the title Fold constant_pad_nd into convolution input paddingDelegate even-kernel 'same'-padding convs via a quantized static pad (#20553)Jul 6, 2026
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 6, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
@JakeStevensJakeStevens added the release notes: xnnpack Changes to the XNNPack backend delegate label Jul 6, 2026
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 6, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 7, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 7, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964

An even-kernel 'same'-padding conv decomposes (after quantization) into
dequant -> constant_pad_nd -> convolution. Because the pad is introduced by
to_edge decomposition -- after the quantizer has run -- it is never annotated,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can we fuse it inside the preprocess for xnnpack or we have to run pad + conv in the XNNPACK graph? I am asking because if we can fuse it we might as well write that pass instead.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can fuse, I have some work in that direction, was going to publish as a quick follow up since there is an external contributor with the 1d work and the change is slightly larger so preferred to do it once these both land.

…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
@meta-codesync
meta-codesyncBot merged commit 420a8ab into pytorch:mainJul 8, 2026
181 of 184 checks passed
@JakeStevens
JakeStevens deleted the export-D109871964 branch July 9, 2026 13:41
SakshamKapoor2911 added a commit to SakshamKapoor2911/executorch that referenced this pull request Jul 9, 2026
 approach)
Instead of folding constant_pad_nd into asymmetric Conv2d padding,
keep it as an explicit XNNPACK pad op:
- Remove pad-folding logic from Conv1dUnsqueezePass
- Remove Conv1dFoldedPadMetaPass (no longer needed)
- Revert xnnpack_input_padding handling in op_conv2d.py
- Add InsertPadQDQPass: inserts QDQ after pad in quantized contexts
so it serializes as a quantized static pad
This matches the approach from pytorch#20553 and fixes the crash with
conv1d -> flatten (even kernel + same padding).
SakshamKapoor2911 added a commit to SakshamKapoor2911/executorch that referenced this pull request Jul 11, 2026
 approach)
Instead of folding constant_pad_nd into asymmetric Conv2d padding,
keep it as an explicit XNNPACK pad op:
- Remove pad-folding logic from Conv1dUnsqueezePass
- Remove Conv1dFoldedPadMetaPass (no longer needed)
- Revert xnnpack_input_padding handling in op_conv2d.py
- Add InsertPadQDQPass: inserts QDQ after pad in quantized contexts
so it serializes as a quantized static pad
This matches the approach from pytorch#20553 and fixes the crash with
conv1d -> flatten (even kernel + same padding).
JakeStevens added a commit that referenced this pull request Jul 16, 2026
Fixes#20558.
Related to #20553.
## Summary
Quantized `nn.Conv1d(..., padding="same")` with an even kernel exports
with
asymmetric padding (unequal left/right amounts), which cannot be folded
into
the convolution's symmetric padding field. The original pad-folding
approach
in this PR was replaced with the explicit-PAD approach from #20553:
- **InsertPadQDQPass** — inserts implicit quantize/dequantize pairs
after
constant_pad_nd nodes in quantized contexts so they serialize as
quantized
static pads. Refactored with additional guard checks (pad_value,
pad_amounts,
negative amounts) and correct idempotency.
- **ConvolutionConfig._get_act_deps** — pulls zero-valued
constant_pad_nd nodes
(and their QDQ chain if InsertPadQDQPass already ran) into the
convolution's
partition for both 1D and 2D convs. The merged method replaces separate
1D
and 2D implementations that existed on this branch.
- **Regression test** — quantized Conv1d even-kernel same-padding
covering both
symmetric and asymmetric pad cases, validated by numerical comparison.
- Removed unrelated __init__.py re-ordering and no-op
conv1d_unsqueeze_pass
changes.
With these changes, an even-kernel `padding="same"` conv1d graph:
```
dequant -> constant_pad_nd -> convolution
```
becomes (after XNNPACK preprocessing and partitioning):
```
[dequant -> pad -> q -> dq -> conv] (single XNNPACK delegate)
```
## Test plan
```
python -m pytest backends/xnnpack/test/ops/test_conv1d.py -q
# 7 passed (1 new regression test with 4 subTests)
python -m pytest backends/xnnpack/test/ops/test_conv2d.py -q
# 32 passed
python -m pytest backends/xnnpack/test/passes/test_insert_pad_qdq.py -q
# 3 passed
python -m pytest backends/xnnpack/test/ops/test_static_constant_pad.py -q
# 8 passed
lintrunner -a
# No lint issues
```
cc @GregoryComer@digantdesai@cbilgin@JakeStevens@freddan80@per@zingo@oscarandersson8218@mansnils@Sebastian-Larsson@robell@rascani
---------
Co-authored-by: Jacob Stevens <stevens.jacob1492@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exportedrelease notes: xnnpackChanges to the XNNPack backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Conv2d with padding 'same' fails (when quantised) at runtime on XNNPACK

3 participants

@JakeStevens@digantdesai@GregoryComer
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length \u003e 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Delegate even-kernel 'same'-padding convs via a quantized static pad (#20553) - #20553

Merged
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
JakeStevens:export-D109871964
Jul 8, 2026
Merged

Delegate even-kernel 'same'-padding convs via a quantized static pad (#20553)#20553
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
JakeStevens:export-D109871964

Conversation

@JakeStevens

@JakeStevensJakeStevens commented Jun 26, 2026

Copy link
Copy Markdown
Contributor

Summary:

When using padding="same" for a conv with even kernels, torch.export creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in xnn_status_unsupported_parameter.

This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (ConvConfig._get_act_deps) so both delegate together: for fp32 the pad lowers directly as an XNNStaticConstantPad. For quantized graphs the pad is created by to_edge decomposition after the quantizer runs, so it arrives as dq -> pad -> conv with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). InsertPadQDQPass inserts an implicit quantize -> dequantize after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.

Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.

Differential Revision: D109871964

@pytorch-bot

pytorch-botBot commented Jun 26, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20553

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Pending, 2 Unrelated Failures

As of commit 0e81ab4 with merge base 3801496 (image):

NEW FAILURE - The following job has failed:

BROKEN TRUNK - The following jobs failed but were present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 26, 2026
@meta-codesync

Copy link
Copy Markdown
Contributor

@JakeStevens has exported this pull request. If you are a Meta employee, you can view the originating Diff in D109871964.

@meta-codesyncmeta-codesyncBot changed the title Fold constant_pad_nd into convolution input paddingDelegate even-kernel 'same'-padding convs via a quantized static pad (#20553)Jul 6, 2026
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 6, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
@JakeStevensJakeStevens added the release notes: xnnpack Changes to the XNNPack backend delegate label Jul 6, 2026
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 6, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 7, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 7, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964

An even-kernel 'same'-padding conv decomposes (after quantization) into
dequant -> constant_pad_nd -> convolution. Because the pad is introduced by
to_edge decomposition -- after the quantizer has run -- it is never annotated,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can we fuse it inside the preprocess for xnnpack or we have to run pad + conv in the XNNPACK graph? I am asking because if we can fuse it we might as well write that pass instead.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can fuse, I have some work in that direction, was going to publish as a quick follow up since there is an external contributor with the 1d work and the change is slightly larger so preferred to do it once these both land.

…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
@meta-codesync
meta-codesyncBot merged commit 420a8ab into pytorch:mainJul 8, 2026
181 of 184 checks passed
@JakeStevens
JakeStevens deleted the export-D109871964 branch July 9, 2026 13:41
SakshamKapoor2911 added a commit to SakshamKapoor2911/executorch that referenced this pull request Jul 9, 2026
 approach)
Instead of folding constant_pad_nd into asymmetric Conv2d padding,
keep it as an explicit XNNPACK pad op:
- Remove pad-folding logic from Conv1dUnsqueezePass
- Remove Conv1dFoldedPadMetaPass (no longer needed)
- Revert xnnpack_input_padding handling in op_conv2d.py
- Add InsertPadQDQPass: inserts QDQ after pad in quantized contexts
so it serializes as a quantized static pad
This matches the approach from pytorch#20553 and fixes the crash with
conv1d -> flatten (even kernel + same padding).
SakshamKapoor2911 added a commit to SakshamKapoor2911/executorch that referenced this pull request Jul 11, 2026
 approach)
Instead of folding constant_pad_nd into asymmetric Conv2d padding,
keep it as an explicit XNNPACK pad op:
- Remove pad-folding logic from Conv1dUnsqueezePass
- Remove Conv1dFoldedPadMetaPass (no longer needed)
- Revert xnnpack_input_padding handling in op_conv2d.py
- Add InsertPadQDQPass: inserts QDQ after pad in quantized contexts
so it serializes as a quantized static pad
This matches the approach from pytorch#20553 and fixes the crash with
conv1d -> flatten (even kernel + same padding).
JakeStevens added a commit that referenced this pull request Jul 16, 2026
Fixes#20558.
Related to #20553.
## Summary
Quantized `nn.Conv1d(..., padding="same")` with an even kernel exports
with
asymmetric padding (unequal left/right amounts), which cannot be folded
into
the convolution's symmetric padding field. The original pad-folding
approach
in this PR was replaced with the explicit-PAD approach from #20553:
- **InsertPadQDQPass** — inserts implicit quantize/dequantize pairs
after
constant_pad_nd nodes in quantized contexts so they serialize as
quantized
static pads. Refactored with additional guard checks (pad_value,
pad_amounts,
negative amounts) and correct idempotency.
- **ConvolutionConfig._get_act_deps** — pulls zero-valued
constant_pad_nd nodes
(and their QDQ chain if InsertPadQDQPass already ran) into the
convolution's
partition for both 1D and 2D convs. The merged method replaces separate
1D
and 2D implementations that existed on this branch.
- **Regression test** — quantized Conv1d even-kernel same-padding
covering both
symmetric and asymmetric pad cases, validated by numerical comparison.
- Removed unrelated __init__.py re-ordering and no-op
conv1d_unsqueeze_pass
changes.
With these changes, an even-kernel `padding="same"` conv1d graph:
```
dequant -> constant_pad_nd -> convolution
```
becomes (after XNNPACK preprocessing and partitioning):
```
[dequant -> pad -> q -> dq -> conv] (single XNNPACK delegate)
```
## Test plan
```
python -m pytest backends/xnnpack/test/ops/test_conv1d.py -q
# 7 passed (1 new regression test with 4 subTests)
python -m pytest backends/xnnpack/test/ops/test_conv2d.py -q
# 32 passed
python -m pytest backends/xnnpack/test/passes/test_insert_pad_qdq.py -q
# 3 passed
python -m pytest backends/xnnpack/test/ops/test_static_constant_pad.py -q
# 8 passed
lintrunner -a
# No lint issues
```
cc @GregoryComer@digantdesai@cbilgin@JakeStevens@freddan80@per@zingo@oscarandersson8218@mansnils@Sebastian-Larsson@robell@rascani
---------
Co-authored-by: Jacob Stevens <stevens.jacob1492@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exportedrelease notes: xnnpackChanges to the XNNPack backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Conv2d with padding 'same' fails (when quantised) at runtime on XNNPACK

3 participants

@JakeStevens@digantdesai@GregoryComer
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Delegate even-kernel 'same'-padding convs via a quantized static pad (#20553) - #20553

Merged
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
JakeStevens:export-D109871964
Jul 8, 2026
Merged

Delegate even-kernel 'same'-padding convs via a quantized static pad (#20553)#20553
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
JakeStevens:export-D109871964

Conversation

@JakeStevens

@JakeStevensJakeStevens commented Jun 26, 2026

Copy link
Copy Markdown
Contributor

Summary:

When using padding="same" for a conv with even kernels, torch.export creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in xnn_status_unsupported_parameter.

This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (ConvConfig._get_act_deps) so both delegate together: for fp32 the pad lowers directly as an XNNStaticConstantPad. For quantized graphs the pad is created by to_edge decomposition after the quantizer runs, so it arrives as dq -> pad -> conv with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). InsertPadQDQPass inserts an implicit quantize -> dequantize after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.

Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.

Differential Revision: D109871964

@pytorch-bot

pytorch-botBot commented Jun 26, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20553

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Pending, 2 Unrelated Failures

As of commit 0e81ab4 with merge base 3801496 (image):

NEW FAILURE - The following job has failed:

BROKEN TRUNK - The following jobs failed but were present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 26, 2026
@meta-codesync

Copy link
Copy Markdown
Contributor

@JakeStevens has exported this pull request. If you are a Meta employee, you can view the originating Diff in D109871964.

@meta-codesyncmeta-codesyncBot changed the title Fold constant_pad_nd into convolution input paddingDelegate even-kernel 'same'-padding convs via a quantized static pad (#20553)Jul 6, 2026
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 6, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
@JakeStevensJakeStevens added the release notes: xnnpack Changes to the XNNPack backend delegate label Jul 6, 2026
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 6, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 7, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 7, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964

An even-kernel 'same'-padding conv decomposes (after quantization) into
dequant -> constant_pad_nd -> convolution. Because the pad is introduced by
to_edge decomposition -- after the quantizer has run -- it is never annotated,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can we fuse it inside the preprocess for xnnpack or we have to run pad + conv in the XNNPACK graph? I am asking because if we can fuse it we might as well write that pass instead.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can fuse, I have some work in that direction, was going to publish as a quick follow up since there is an external contributor with the 1d work and the change is slightly larger so preferred to do it once these both land.

…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
@meta-codesync
meta-codesyncBot merged commit 420a8ab into pytorch:mainJul 8, 2026
181 of 184 checks passed
@JakeStevens
JakeStevens deleted the export-D109871964 branch July 9, 2026 13:41
SakshamKapoor2911 added a commit to SakshamKapoor2911/executorch that referenced this pull request Jul 9, 2026
 approach)
Instead of folding constant_pad_nd into asymmetric Conv2d padding,
keep it as an explicit XNNPACK pad op:
- Remove pad-folding logic from Conv1dUnsqueezePass
- Remove Conv1dFoldedPadMetaPass (no longer needed)
- Revert xnnpack_input_padding handling in op_conv2d.py
- Add InsertPadQDQPass: inserts QDQ after pad in quantized contexts
so it serializes as a quantized static pad
This matches the approach from pytorch#20553 and fixes the crash with
conv1d -> flatten (even kernel + same padding).
SakshamKapoor2911 added a commit to SakshamKapoor2911/executorch that referenced this pull request Jul 11, 2026
 approach)
Instead of folding constant_pad_nd into asymmetric Conv2d padding,
keep it as an explicit XNNPACK pad op:
- Remove pad-folding logic from Conv1dUnsqueezePass
- Remove Conv1dFoldedPadMetaPass (no longer needed)
- Revert xnnpack_input_padding handling in op_conv2d.py
- Add InsertPadQDQPass: inserts QDQ after pad in quantized contexts
so it serializes as a quantized static pad
This matches the approach from pytorch#20553 and fixes the crash with
conv1d -> flatten (even kernel + same padding).
JakeStevens added a commit that referenced this pull request Jul 16, 2026
Fixes#20558.
Related to #20553.
## Summary
Quantized `nn.Conv1d(..., padding="same")` with an even kernel exports
with
asymmetric padding (unequal left/right amounts), which cannot be folded
into
the convolution's symmetric padding field. The original pad-folding
approach
in this PR was replaced with the explicit-PAD approach from #20553:
- **InsertPadQDQPass** — inserts implicit quantize/dequantize pairs
after
constant_pad_nd nodes in quantized contexts so they serialize as
quantized
static pads. Refactored with additional guard checks (pad_value,
pad_amounts,
negative amounts) and correct idempotency.
- **ConvolutionConfig._get_act_deps** — pulls zero-valued
constant_pad_nd nodes
(and their QDQ chain if InsertPadQDQPass already ran) into the
convolution's
partition for both 1D and 2D convs. The merged method replaces separate
1D
and 2D implementations that existed on this branch.
- **Regression test** — quantized Conv1d even-kernel same-padding
covering both
symmetric and asymmetric pad cases, validated by numerical comparison.
- Removed unrelated __init__.py re-ordering and no-op
conv1d_unsqueeze_pass
changes.
With these changes, an even-kernel `padding="same"` conv1d graph:
```
dequant -> constant_pad_nd -> convolution
```
becomes (after XNNPACK preprocessing and partitioning):
```
[dequant -> pad -> q -> dq -> conv] (single XNNPACK delegate)
```
## Test plan
```
python -m pytest backends/xnnpack/test/ops/test_conv1d.py -q
# 7 passed (1 new regression test with 4 subTests)
python -m pytest backends/xnnpack/test/ops/test_conv2d.py -q
# 32 passed
python -m pytest backends/xnnpack/test/passes/test_insert_pad_qdq.py -q
# 3 passed
python -m pytest backends/xnnpack/test/ops/test_static_constant_pad.py -q
# 8 passed
lintrunner -a
# No lint issues
```
cc @GregoryComer@digantdesai@cbilgin@JakeStevens@freddan80@per@zingo@oscarandersson8218@mansnils@Sebastian-Larsson@robell@rascani
---------
Co-authored-by: Jacob Stevens <stevens.jacob1492@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exportedrelease notes: xnnpackChanges to the XNNPack backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Conv2d with padding 'same' fails (when quantised) at runtime on XNNPACK

3 participants

@JakeStevens@digantdesai@GregoryComer
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Delegate even-kernel 'same'-padding convs via a quantized static pad (#20553) - #20553

Merged
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
JakeStevens:export-D109871964
Jul 8, 2026
Merged

Delegate even-kernel 'same'-padding convs via a quantized static pad (#20553)#20553
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
JakeStevens:export-D109871964

Conversation

@JakeStevens

@JakeStevensJakeStevens commented Jun 26, 2026

Copy link
Copy Markdown
Contributor

Summary:

When using padding="same" for a conv with even kernels, torch.export creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in xnn_status_unsupported_parameter.

This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (ConvConfig._get_act_deps) so both delegate together: for fp32 the pad lowers directly as an XNNStaticConstantPad. For quantized graphs the pad is created by to_edge decomposition after the quantizer runs, so it arrives as dq -> pad -> conv with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). InsertPadQDQPass inserts an implicit quantize -> dequantize after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.

Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.

Differential Revision: D109871964

@pytorch-bot

pytorch-botBot commented Jun 26, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20553

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Pending, 2 Unrelated Failures

As of commit 0e81ab4 with merge base 3801496 (image):

NEW FAILURE - The following job has failed:

BROKEN TRUNK - The following jobs failed but were present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 26, 2026
@meta-codesync

Copy link
Copy Markdown
Contributor

@JakeStevens has exported this pull request. If you are a Meta employee, you can view the originating Diff in D109871964.

@meta-codesyncmeta-codesyncBot changed the title Fold constant_pad_nd into convolution input paddingDelegate even-kernel 'same'-padding convs via a quantized static pad (#20553)Jul 6, 2026
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 6, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
@JakeStevensJakeStevens added the release notes: xnnpack Changes to the XNNPack backend delegate label Jul 6, 2026
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 6, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 7, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 7, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964

An even-kernel 'same'-padding conv decomposes (after quantization) into
dequant -> constant_pad_nd -> convolution. Because the pad is introduced by
to_edge decomposition -- after the quantizer has run -- it is never annotated,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can we fuse it inside the preprocess for xnnpack or we have to run pad + conv in the XNNPACK graph? I am asking because if we can fuse it we might as well write that pass instead.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can fuse, I have some work in that direction, was going to publish as a quick follow up since there is an external contributor with the 1d work and the change is slightly larger so preferred to do it once these both land.

…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
@meta-codesync
meta-codesyncBot merged commit 420a8ab into pytorch:mainJul 8, 2026
181 of 184 checks passed
@JakeStevens
JakeStevens deleted the export-D109871964 branch July 9, 2026 13:41
SakshamKapoor2911 added a commit to SakshamKapoor2911/executorch that referenced this pull request Jul 9, 2026
 approach)
Instead of folding constant_pad_nd into asymmetric Conv2d padding,
keep it as an explicit XNNPACK pad op:
- Remove pad-folding logic from Conv1dUnsqueezePass
- Remove Conv1dFoldedPadMetaPass (no longer needed)
- Revert xnnpack_input_padding handling in op_conv2d.py
- Add InsertPadQDQPass: inserts QDQ after pad in quantized contexts
so it serializes as a quantized static pad
This matches the approach from pytorch#20553 and fixes the crash with
conv1d -> flatten (even kernel + same padding).
SakshamKapoor2911 added a commit to SakshamKapoor2911/executorch that referenced this pull request Jul 11, 2026
 approach)
Instead of folding constant_pad_nd into asymmetric Conv2d padding,
keep it as an explicit XNNPACK pad op:
- Remove pad-folding logic from Conv1dUnsqueezePass
- Remove Conv1dFoldedPadMetaPass (no longer needed)
- Revert xnnpack_input_padding handling in op_conv2d.py
- Add InsertPadQDQPass: inserts QDQ after pad in quantized contexts
so it serializes as a quantized static pad
This matches the approach from pytorch#20553 and fixes the crash with
conv1d -> flatten (even kernel + same padding).
JakeStevens added a commit that referenced this pull request Jul 16, 2026
Fixes#20558.
Related to #20553.
## Summary
Quantized `nn.Conv1d(..., padding="same")` with an even kernel exports
with
asymmetric padding (unequal left/right amounts), which cannot be folded
into
the convolution's symmetric padding field. The original pad-folding
approach
in this PR was replaced with the explicit-PAD approach from #20553:
- **InsertPadQDQPass** — inserts implicit quantize/dequantize pairs
after
constant_pad_nd nodes in quantized contexts so they serialize as
quantized
static pads. Refactored with additional guard checks (pad_value,
pad_amounts,
negative amounts) and correct idempotency.
- **ConvolutionConfig._get_act_deps** — pulls zero-valued
constant_pad_nd nodes
(and their QDQ chain if InsertPadQDQPass already ran) into the
convolution's
partition for both 1D and 2D convs. The merged method replaces separate
1D
and 2D implementations that existed on this branch.
- **Regression test** — quantized Conv1d even-kernel same-padding
covering both
symmetric and asymmetric pad cases, validated by numerical comparison.
- Removed unrelated __init__.py re-ordering and no-op
conv1d_unsqueeze_pass
changes.
With these changes, an even-kernel `padding="same"` conv1d graph:
```
dequant -> constant_pad_nd -> convolution
```
becomes (after XNNPACK preprocessing and partitioning):
```
[dequant -> pad -> q -> dq -> conv] (single XNNPACK delegate)
```
## Test plan
```
python -m pytest backends/xnnpack/test/ops/test_conv1d.py -q
# 7 passed (1 new regression test with 4 subTests)
python -m pytest backends/xnnpack/test/ops/test_conv2d.py -q
# 32 passed
python -m pytest backends/xnnpack/test/passes/test_insert_pad_qdq.py -q
# 3 passed
python -m pytest backends/xnnpack/test/ops/test_static_constant_pad.py -q
# 8 passed
lintrunner -a
# No lint issues
```
cc @GregoryComer@digantdesai@cbilgin@JakeStevens@freddan80@per@zingo@oscarandersson8218@mansnils@Sebastian-Larsson@robell@rascani
---------
Co-authored-by: Jacob Stevens <stevens.jacob1492@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exportedrelease notes: xnnpackChanges to the XNNPack backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Conv2d with padding 'same' fails (when quantised) at runtime on XNNPACK

3 participants

@JakeStevens@digantdesai@GregoryComer
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Delegate even-kernel 'same'-padding convs via a quantized static pad (#20553) - #20553

Merged
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
JakeStevens:export-D109871964
Jul 8, 2026
Merged

Delegate even-kernel 'same'-padding convs via a quantized static pad (#20553)#20553
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
JakeStevens:export-D109871964

Conversation

@JakeStevens

@JakeStevensJakeStevens commented Jun 26, 2026

Copy link
Copy Markdown
Contributor

Summary:

When using padding="same" for a conv with even kernels, torch.export creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in xnn_status_unsupported_parameter.

This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (ConvConfig._get_act_deps) so both delegate together: for fp32 the pad lowers directly as an XNNStaticConstantPad. For quantized graphs the pad is created by to_edge decomposition after the quantizer runs, so it arrives as dq -> pad -> conv with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). InsertPadQDQPass inserts an implicit quantize -> dequantize after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.

Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.

Differential Revision: D109871964

@pytorch-bot

pytorch-botBot commented Jun 26, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20553

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Pending, 2 Unrelated Failures

As of commit 0e81ab4 with merge base 3801496 (image):

NEW FAILURE - The following job has failed:

BROKEN TRUNK - The following jobs failed but were present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 26, 2026
@meta-codesync

Copy link
Copy Markdown
Contributor

@JakeStevens has exported this pull request. If you are a Meta employee, you can view the originating Diff in D109871964.

@meta-codesyncmeta-codesyncBot changed the title Fold constant_pad_nd into convolution input paddingDelegate even-kernel 'same'-padding convs via a quantized static pad (#20553)Jul 6, 2026
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 6, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
@JakeStevensJakeStevens added the release notes: xnnpack Changes to the XNNPack backend delegate label Jul 6, 2026
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 6, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 7, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 7, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964

An even-kernel 'same'-padding conv decomposes (after quantization) into
dequant -> constant_pad_nd -> convolution. Because the pad is introduced by
to_edge decomposition -- after the quantizer has run -- it is never annotated,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can we fuse it inside the preprocess for xnnpack or we have to run pad + conv in the XNNPACK graph? I am asking because if we can fuse it we might as well write that pass instead.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can fuse, I have some work in that direction, was going to publish as a quick follow up since there is an external contributor with the 1d work and the change is slightly larger so preferred to do it once these both land.

…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
@meta-codesync
meta-codesyncBot merged commit 420a8ab into pytorch:mainJul 8, 2026
181 of 184 checks passed
@JakeStevens
JakeStevens deleted the export-D109871964 branch July 9, 2026 13:41
SakshamKapoor2911 added a commit to SakshamKapoor2911/executorch that referenced this pull request Jul 9, 2026
 approach)
Instead of folding constant_pad_nd into asymmetric Conv2d padding,
keep it as an explicit XNNPACK pad op:
- Remove pad-folding logic from Conv1dUnsqueezePass
- Remove Conv1dFoldedPadMetaPass (no longer needed)
- Revert xnnpack_input_padding handling in op_conv2d.py
- Add InsertPadQDQPass: inserts QDQ after pad in quantized contexts
so it serializes as a quantized static pad
This matches the approach from pytorch#20553 and fixes the crash with
conv1d -> flatten (even kernel + same padding).
SakshamKapoor2911 added a commit to SakshamKapoor2911/executorch that referenced this pull request Jul 11, 2026
 approach)
Instead of folding constant_pad_nd into asymmetric Conv2d padding,
keep it as an explicit XNNPACK pad op:
- Remove pad-folding logic from Conv1dUnsqueezePass
- Remove Conv1dFoldedPadMetaPass (no longer needed)
- Revert xnnpack_input_padding handling in op_conv2d.py
- Add InsertPadQDQPass: inserts QDQ after pad in quantized contexts
so it serializes as a quantized static pad
This matches the approach from pytorch#20553 and fixes the crash with
conv1d -> flatten (even kernel + same padding).
JakeStevens added a commit that referenced this pull request Jul 16, 2026
Fixes#20558.
Related to #20553.
## Summary
Quantized `nn.Conv1d(..., padding="same")` with an even kernel exports
with
asymmetric padding (unequal left/right amounts), which cannot be folded
into
the convolution's symmetric padding field. The original pad-folding
approach
in this PR was replaced with the explicit-PAD approach from #20553:
- **InsertPadQDQPass** — inserts implicit quantize/dequantize pairs
after
constant_pad_nd nodes in quantized contexts so they serialize as
quantized
static pads. Refactored with additional guard checks (pad_value,
pad_amounts,
negative amounts) and correct idempotency.
- **ConvolutionConfig._get_act_deps** — pulls zero-valued
constant_pad_nd nodes
(and their QDQ chain if InsertPadQDQPass already ran) into the
convolution's
partition for both 1D and 2D convs. The merged method replaces separate
1D
and 2D implementations that existed on this branch.
- **Regression test** — quantized Conv1d even-kernel same-padding
covering both
symmetric and asymmetric pad cases, validated by numerical comparison.
- Removed unrelated __init__.py re-ordering and no-op
conv1d_unsqueeze_pass
changes.
With these changes, an even-kernel `padding="same"` conv1d graph:
```
dequant -> constant_pad_nd -> convolution
```
becomes (after XNNPACK preprocessing and partitioning):
```
[dequant -> pad -> q -> dq -> conv] (single XNNPACK delegate)
```
## Test plan
```
python -m pytest backends/xnnpack/test/ops/test_conv1d.py -q
# 7 passed (1 new regression test with 4 subTests)
python -m pytest backends/xnnpack/test/ops/test_conv2d.py -q
# 32 passed
python -m pytest backends/xnnpack/test/passes/test_insert_pad_qdq.py -q
# 3 passed
python -m pytest backends/xnnpack/test/ops/test_static_constant_pad.py -q
# 8 passed
lintrunner -a
# No lint issues
```
cc @GregoryComer@digantdesai@cbilgin@JakeStevens@freddan80@per@zingo@oscarandersson8218@mansnils@Sebastian-Larsson@robell@rascani
---------
Co-authored-by: Jacob Stevens <stevens.jacob1492@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exportedrelease notes: xnnpackChanges to the XNNPack backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Conv2d with padding 'same' fails (when quantised) at runtime on XNNPACK

3 participants

@JakeStevens@digantdesai@GregoryComer
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Delegate even-kernel 'same'-padding convs via a quantized static pad (#20553) - #20553

Merged
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
JakeStevens:export-D109871964
Jul 8, 2026
Merged

Delegate even-kernel 'same'-padding convs via a quantized static pad (#20553)#20553
meta-codesync[bot] merged 1 commit into
pytorch:mainfrom
JakeStevens:export-D109871964

Conversation

@JakeStevens

@JakeStevensJakeStevens commented Jun 26, 2026

Copy link
Copy Markdown
Contributor

Summary:

When using padding="same" for a conv with even kernels, torch.export creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in xnn_status_unsupported_parameter.

This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (ConvConfig._get_act_deps) so both delegate together: for fp32 the pad lowers directly as an XNNStaticConstantPad. For quantized graphs the pad is created by to_edge decomposition after the quantizer runs, so it arrives as dq -> pad -> conv with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). InsertPadQDQPass inserts an implicit quantize -> dequantize after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.

Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.

Differential Revision: D109871964

@pytorch-bot

pytorch-botBot commented Jun 26, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20553

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Pending, 2 Unrelated Failures

As of commit 0e81ab4 with merge base 3801496 (image):

NEW FAILURE - The following job has failed:

BROKEN TRUNK - The following jobs failed but were present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 26, 2026
@meta-codesync

Copy link
Copy Markdown
Contributor

@JakeStevens has exported this pull request. If you are a Meta employee, you can view the originating Diff in D109871964.

@meta-codesyncmeta-codesyncBot changed the title Fold constant_pad_nd into convolution input paddingDelegate even-kernel 'same'-padding convs via a quantized static pad (#20553)Jul 6, 2026
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 6, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
@JakeStevensJakeStevens added the release notes: xnnpack Changes to the XNNPack backend delegate label Jul 6, 2026
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 6, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 7, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
JakeStevens added a commit to JakeStevens/executorch that referenced this pull request Jul 7, 2026
…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964

An even-kernel 'same'-padding conv decomposes (after quantization) into
dequant -> constant_pad_nd -> convolution. Because the pad is introduced by
to_edge decomposition -- after the quantizer has run -- it is never annotated,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can we fuse it inside the preprocess for xnnpack or we have to run pad + conv in the XNNPACK graph? I am asking because if we can fuse it we might as well write that pass instead.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can fuse, I have some work in that direction, was going to publish as a quick follow up since there is an external contributor with the 1d work and the change is slightly larger so preferred to do it once these both land.

…ytorch#20553)
Summary:
When using padding="same" for a conv with even kernels, `torch.export` creates an explicit pad node, leading to a dq -> pad -> conv chain, which is not recognized, so conv ends up undelegated. The ReLU ends up in its own partition with a quantize. XNNPACK does not support this, resulting in `xnn_status_unsupported_parameter`.
This PR fixes this, allowing for full delegation to XNNPACK in this case, by pulling the zero-valued spatial pad into the conv's partition (`ConvConfig._get_act_deps`) so both delegate together: for fp32 the pad lowers directly as an `XNNStaticConstantPad`. For quantized graphs the pad is created by `to_edge` decomposition after the quantizer runs, so it arrives as `dq -> pad -> conv` with no quantize on its output and would serialize as fp32 (the conv would then reject its unquantized activation). `InsertPadQDQPass` inserts an implicit `quantize -> dequantize` after the pad, reusing the feeding dequant's params (a zero pad preserves quantization), so it lowers as a quantized static pad and the conv sees a proper dequantized activation.
Note: This results in full delegation, an improvement over the existing runtime error (or portable fallback as naive fix), but the pad is executed as a separate op and results in overhead.
Differential Revision: D109871964
@meta-codesync
meta-codesyncBot merged commit 420a8ab into pytorch:mainJul 8, 2026
181 of 184 checks passed
@JakeStevens
JakeStevens deleted the export-D109871964 branch July 9, 2026 13:41
SakshamKapoor2911 added a commit to SakshamKapoor2911/executorch that referenced this pull request Jul 9, 2026
 approach)
Instead of folding constant_pad_nd into asymmetric Conv2d padding,
keep it as an explicit XNNPACK pad op:
- Remove pad-folding logic from Conv1dUnsqueezePass
- Remove Conv1dFoldedPadMetaPass (no longer needed)
- Revert xnnpack_input_padding handling in op_conv2d.py
- Add InsertPadQDQPass: inserts QDQ after pad in quantized contexts
so it serializes as a quantized static pad
This matches the approach from pytorch#20553 and fixes the crash with
conv1d -> flatten (even kernel + same padding).
SakshamKapoor2911 added a commit to SakshamKapoor2911/executorch that referenced this pull request Jul 11, 2026
 approach)
Instead of folding constant_pad_nd into asymmetric Conv2d padding,
keep it as an explicit XNNPACK pad op:
- Remove pad-folding logic from Conv1dUnsqueezePass
- Remove Conv1dFoldedPadMetaPass (no longer needed)
- Revert xnnpack_input_padding handling in op_conv2d.py
- Add InsertPadQDQPass: inserts QDQ after pad in quantized contexts
so it serializes as a quantized static pad
This matches the approach from pytorch#20553 and fixes the crash with
conv1d -> flatten (even kernel + same padding).
JakeStevens added a commit that referenced this pull request Jul 16, 2026
Fixes#20558.
Related to #20553.
## Summary
Quantized `nn.Conv1d(..., padding="same")` with an even kernel exports
with
asymmetric padding (unequal left/right amounts), which cannot be folded
into
the convolution's symmetric padding field. The original pad-folding
approach
in this PR was replaced with the explicit-PAD approach from #20553:
- **InsertPadQDQPass** — inserts implicit quantize/dequantize pairs
after
constant_pad_nd nodes in quantized contexts so they serialize as
quantized
static pads. Refactored with additional guard checks (pad_value,
pad_amounts,
negative amounts) and correct idempotency.
- **ConvolutionConfig._get_act_deps** — pulls zero-valued
constant_pad_nd nodes
(and their QDQ chain if InsertPadQDQPass already ran) into the
convolution's
partition for both 1D and 2D convs. The merged method replaces separate
1D
and 2D implementations that existed on this branch.
- **Regression test** — quantized Conv1d even-kernel same-padding
covering both
symmetric and asymmetric pad cases, validated by numerical comparison.
- Removed unrelated __init__.py re-ordering and no-op
conv1d_unsqueeze_pass
changes.
With these changes, an even-kernel `padding="same"` conv1d graph:
```
dequant -> constant_pad_nd -> convolution
```
becomes (after XNNPACK preprocessing and partitioning):
```
[dequant -> pad -> q -> dq -> conv] (single XNNPACK delegate)
```
## Test plan
```
python -m pytest backends/xnnpack/test/ops/test_conv1d.py -q
# 7 passed (1 new regression test with 4 subTests)
python -m pytest backends/xnnpack/test/ops/test_conv2d.py -q
# 32 passed
python -m pytest backends/xnnpack/test/passes/test_insert_pad_qdq.py -q
# 3 passed
python -m pytest backends/xnnpack/test/ops/test_static_constant_pad.py -q
# 8 passed
lintrunner -a
# No lint issues
```
cc @GregoryComer@digantdesai@cbilgin@JakeStevens@freddan80@per@zingo@oscarandersson8218@mansnils@Sebastian-Larsson@robell@rascani
---------
Co-authored-by: Jacob Stevens <stevens.jacob1492@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exportedrelease notes: xnnpackChanges to the XNNPack backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Conv2d with padding 'same' fails (when quantised) at runtime on XNNPACK

3 participants

@JakeStevens@digantdesai@GregoryComer