[ExecuTorch][WebGPU] SymInt arithmetic ops (add/sub/mul/floordiv) for dynamic shapes - #20573

Merged
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/65/basefrom
gh/JulianCloudNTH/65/head
Jul 4, 2026
Merged

[ExecuTorch][WebGPU] SymInt arithmetic ops (add/sub/mul/floordiv) for dynamic shapes#20573
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/65/basefrom
gh/JulianCloudNTH/65/head

Conversation

@ghost

@ghostghost commented Jun 28, 2026

Copy link
Copy Markdown

Stack from ghstack (oldest at bottom):

Register scalar SymInt arithmetic so dynamic-shape graphs lower without an "unsupported op" failure.

Problem: A dynamic-shape exported program emits scalar SymInt arithmetic nodes (add/sub/mul/floordiv) to compute live sizes and positions (e.g. input_pos + S, seq_len // n). The WebGPU backend registered only et_vk.select_as_symint.default, so WebGPUGraph::build() threw unsupported op: add when loading any dynamic .pte.

Solution:

  • Before: only select_as_symint produced a live SymInt; any arithmetic on it was unsupported.
  • After: add/sub/mul/floordiv each recompute their output SymInt from the operands via a resize hook whenever a live operand changes.

Implementation:

  • register_sym_binary reads each operand (live SymInt via read_symint, else a static Int), seeds the build-time value, and registers a resize hook on any live operand (set_symint on recompute).
  • floordiv rounds toward negative infinity (Python semantics).
  • Mirrors Vulkan backends/vulkan/runtime/graph/ops/impl/SymIntOps.cpp (add/sub/floordiv/mul).
  • Registered under the bare targets add/sub/mul/floordiv — distinct registry keys from the tensor aten.add.Tensor/aten.mul.Tensor ops.

Constraints: An output that folded to a static Int is a no-op. No GPU kernel, no dispatch, no change to the static-shape path. sym_size.int is intentionally not added here (depends on the tensor-shape dim-source API in the following diff).

Co-authored-with: Claude Code.
@exported-using-ghexport

Differential Revision: D109906102

Differential Revision: D109906102

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jun 28, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20573

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit a918948 with merge base 124b677 (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@ghost
ghost temporarily deployed to cadence June 28, 2026 16:22 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 28, 2026 16:22 — with GitHub Actions Inactive
@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 28, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@ghost

Copy link
Copy Markdown
Author

@claude review and check for any areas or opportunities for modularization

@claude

claudeBot commented Jun 29, 2026

Copy link
Copy Markdown

Claude finished @JulianCloudNTH's task in 2m 6s —— View job


Review: WebGPU SymInt arithmetic ops

  • Read changed file SelectAsSymint.cpp
  • Compare against Vulkan reference SymIntOps.cpp
  • Check WebGPUGraph API (read_symint/set_symint/add_resize_hook/propagate_resize)
  • Analyze correctness & modularization opportunities

Overall this is a clean, well-scoped change. The resize-hook approach is a genuine improvement over the Vulkan reference it mirrors, and the diff does exactly what the description claims. Notes below, ordered by importance.


Modularization (the main ask)

1. register_sym_binary is the right call — and better than Vulkan. The Vulkan reference (backends/vulkan/runtime/graph/ops/impl/SymIntOps.cpp:53-166) copy-pastes the full _impl + resize_*_node + registration boilerplate four times. Collapsing all four ops into one helper parameterized by a std::function (SelectAsSymint.cpp:51-74) is the more maintainable design. Good instinct.

2. File/directory naming no longer matches contents. This is the clearest modularization opportunity. The file now contains general SymInt arithmetic (add/sub/mul/floordiv), but lives in ops/select_as_symint/SelectAsSymint.cpp — a directory named after one specific op. The repo convention is strictly one directory per op/concern (ops/add/, ops/mul/, ops/select/, …). A reader looking for where floordiv is registered won't find it under a select_as_symint path.

Consider extracting the arithmetic ops into a sibling ops/sym_int/SymIntOps.cpp (mirroring the Vulkan filename), leaving select_as_symint to own only select_as_symint. The shared helper read_scalar would move with them. This keeps the grep-to-op mapping the codebase relies on. Fix this →(if you'd rather keep the diff small and land the stack first, this is fine as a fast-follow — just flagging the convention.)


Correctness

3. Division by zero in floordiv is UB (SelectAsSymint.cpp:90-91). x / y and x % y with y == 0 are undefined behavior. The Vulkan reference has the same gap, so this isn't a regression, and a floordiv node with a zero divisor shouldn't arise from a valid export — but a one-line guard (throw, or return 0) would make it fail-loud like the rest of this file (read_symint, set_symint both throw on misuse) instead of crashing.

4. INT_MIN % y / INT_MIN / -1 overflow — theoretical only, since these values are tensor sizes/positions. Not worth handling; noting for completeness.

5. Self-referential operands register the hook twice. If a == b (e.g. x + x), both if branches at :68-72 fire and recompute is registered twice on the same SymInt. Harmless (recompute is idempotent, and set_symint no-ops when the value is unchanged), just slightly redundant. An if (a_is_symint) add_hook(a); else if (b_is_symint)… won't work since both may legitimately be distinct live operands — a b != a guard on the second branch would be the precise fix if you want it.


Minor / style

6. read_scalar silently assumes non-SymInt ⇒ Int (:43-48). If an operand were a Double (or any non-Int constant), get_int reads an unrelated flat array and returns a wrong value rather than erroring. Scalar arithmetic operands should only ever be Int/SymInt, so this is safe in practice; a comment or a type assert would document the assumption.

7. Chained arithmetic relies on registration order.propagate_resize (WebGPUGraph.cpp:148-158) iterates resize_hooks_ once in registration order; a hook that recomputes out marks out dirty so a later-registered hook keyed on out still fires in the same pass. This works only because ops are registered in topological order during build(). That's a reasonable invariant, but it's load-bearing and currently undocumented — worth a one-line comment on add_resize_hook or propagate_resize so a future single-pass refactor doesn't silently break multi-level SymInt chains (a + b, then that * c).


Verification

I could not build/run the WebGPU backend in this environment, so the above is static analysis only. Worth confirming with an actual dynamic-shape .pte that exercises a chained expression (e.g. (input_pos + S) // n) end-to-end, which would validate both the floor rounding (:88-95) and the chaining behavior in note 7.
· branch gh/JulianCloudNTH/65/head

[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:28 — with GitHub Actions Inactive
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 21:20 — with GitHub Actions Inactive
@meta-codesync
meta-codesyncBot merged commit ae4b83e into gh/JulianCloudNTH/65/baseJul 4, 2026
181 of 183 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JulianCloudNTH/65/head branch July 4, 2026 17:05
ghost pushed a commit that referenced this pull request Jul 4, 2026
… dynamic shapes (#20712)
This PR was created by the merge bot to help merge the original PR into
the main branch.
ghstack PR number: #20573 by
@JulianCloudNTH
^ Please use this as the source of truth for the PR details, comments,
and reviews
ghstack PR base:
https://github.com/pytorch/executorch/tree/gh/JulianCloudNTH/65/base
ghstack PR head:
https://github.com/pytorch/executorch/tree/gh/JulianCloudNTH/65/head
Merge bot PR base: https://github.com/pytorch/executorch/tree/main
Merge bot PR head:
https://github.com/pytorch/executorch/tree/gh/JulianCloudNTH/65/orig
@diff-train-skip-merge
---------
Co-authored-by: Julian Ng-Thow-Hing <juliannth@meta.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@psiddh@nil-is-all@JCNTH
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

[ExecuTorch][WebGPU] SymInt arithmetic ops (add/sub/mul/floordiv) for dynamic shapes - #20573

Merged
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/65/basefrom
gh/JulianCloudNTH/65/head
Jul 4, 2026
Merged

[ExecuTorch][WebGPU] SymInt arithmetic ops (add/sub/mul/floordiv) for dynamic shapes#20573
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/65/basefrom
gh/JulianCloudNTH/65/head

Conversation

@ghost

@ghostghost commented Jun 28, 2026

Copy link
Copy Markdown

Stack from ghstack (oldest at bottom):

Register scalar SymInt arithmetic so dynamic-shape graphs lower without an "unsupported op" failure.

Problem: A dynamic-shape exported program emits scalar SymInt arithmetic nodes (add/sub/mul/floordiv) to compute live sizes and positions (e.g. input_pos + S, seq_len // n). The WebGPU backend registered only et_vk.select_as_symint.default, so WebGPUGraph::build() threw unsupported op: add when loading any dynamic .pte.

Solution:

  • Before: only select_as_symint produced a live SymInt; any arithmetic on it was unsupported.
  • After: add/sub/mul/floordiv each recompute their output SymInt from the operands via a resize hook whenever a live operand changes.

Implementation:

  • register_sym_binary reads each operand (live SymInt via read_symint, else a static Int), seeds the build-time value, and registers a resize hook on any live operand (set_symint on recompute).
  • floordiv rounds toward negative infinity (Python semantics).
  • Mirrors Vulkan backends/vulkan/runtime/graph/ops/impl/SymIntOps.cpp (add/sub/floordiv/mul).
  • Registered under the bare targets add/sub/mul/floordiv — distinct registry keys from the tensor aten.add.Tensor/aten.mul.Tensor ops.

Constraints: An output that folded to a static Int is a no-op. No GPU kernel, no dispatch, no change to the static-shape path. sym_size.int is intentionally not added here (depends on the tensor-shape dim-source API in the following diff).

Co-authored-with: Claude Code.
@exported-using-ghexport

Differential Revision: D109906102

Differential Revision: D109906102

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jun 28, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20573

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit a918948 with merge base 124b677 (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@ghost
ghost temporarily deployed to cadence June 28, 2026 16:22 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 28, 2026 16:22 — with GitHub Actions Inactive
@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 28, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@ghost

Copy link
Copy Markdown
Author

@claude review and check for any areas or opportunities for modularization

@claude

claudeBot commented Jun 29, 2026

Copy link
Copy Markdown

Claude finished @JulianCloudNTH's task in 2m 6s —— View job


Review: WebGPU SymInt arithmetic ops

  • Read changed file SelectAsSymint.cpp
  • Compare against Vulkan reference SymIntOps.cpp
  • Check WebGPUGraph API (read_symint/set_symint/add_resize_hook/propagate_resize)
  • Analyze correctness & modularization opportunities

Overall this is a clean, well-scoped change. The resize-hook approach is a genuine improvement over the Vulkan reference it mirrors, and the diff does exactly what the description claims. Notes below, ordered by importance.


Modularization (the main ask)

1. register_sym_binary is the right call — and better than Vulkan. The Vulkan reference (backends/vulkan/runtime/graph/ops/impl/SymIntOps.cpp:53-166) copy-pastes the full _impl + resize_*_node + registration boilerplate four times. Collapsing all four ops into one helper parameterized by a std::function (SelectAsSymint.cpp:51-74) is the more maintainable design. Good instinct.

2. File/directory naming no longer matches contents. This is the clearest modularization opportunity. The file now contains general SymInt arithmetic (add/sub/mul/floordiv), but lives in ops/select_as_symint/SelectAsSymint.cpp — a directory named after one specific op. The repo convention is strictly one directory per op/concern (ops/add/, ops/mul/, ops/select/, …). A reader looking for where floordiv is registered won't find it under a select_as_symint path.

Consider extracting the arithmetic ops into a sibling ops/sym_int/SymIntOps.cpp (mirroring the Vulkan filename), leaving select_as_symint to own only select_as_symint. The shared helper read_scalar would move with them. This keeps the grep-to-op mapping the codebase relies on. Fix this →(if you'd rather keep the diff small and land the stack first, this is fine as a fast-follow — just flagging the convention.)


Correctness

3. Division by zero in floordiv is UB (SelectAsSymint.cpp:90-91). x / y and x % y with y == 0 are undefined behavior. The Vulkan reference has the same gap, so this isn't a regression, and a floordiv node with a zero divisor shouldn't arise from a valid export — but a one-line guard (throw, or return 0) would make it fail-loud like the rest of this file (read_symint, set_symint both throw on misuse) instead of crashing.

4. INT_MIN % y / INT_MIN / -1 overflow — theoretical only, since these values are tensor sizes/positions. Not worth handling; noting for completeness.

5. Self-referential operands register the hook twice. If a == b (e.g. x + x), both if branches at :68-72 fire and recompute is registered twice on the same SymInt. Harmless (recompute is idempotent, and set_symint no-ops when the value is unchanged), just slightly redundant. An if (a_is_symint) add_hook(a); else if (b_is_symint)… won't work since both may legitimately be distinct live operands — a b != a guard on the second branch would be the precise fix if you want it.


Minor / style

6. read_scalar silently assumes non-SymInt ⇒ Int (:43-48). If an operand were a Double (or any non-Int constant), get_int reads an unrelated flat array and returns a wrong value rather than erroring. Scalar arithmetic operands should only ever be Int/SymInt, so this is safe in practice; a comment or a type assert would document the assumption.

7. Chained arithmetic relies on registration order.propagate_resize (WebGPUGraph.cpp:148-158) iterates resize_hooks_ once in registration order; a hook that recomputes out marks out dirty so a later-registered hook keyed on out still fires in the same pass. This works only because ops are registered in topological order during build(). That's a reasonable invariant, but it's load-bearing and currently undocumented — worth a one-line comment on add_resize_hook or propagate_resize so a future single-pass refactor doesn't silently break multi-level SymInt chains (a + b, then that * c).


Verification

I could not build/run the WebGPU backend in this environment, so the above is static analysis only. Worth confirming with an actual dynamic-shape .pte that exercises a chained expression (e.g. (input_pos + S) // n) end-to-end, which would validate both the floor rounding (:88-95) and the chaining behavior in note 7.
· branch gh/JulianCloudNTH/65/head

[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:28 — with GitHub Actions Inactive
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 21:20 — with GitHub Actions Inactive
@meta-codesync
meta-codesyncBot merged commit ae4b83e into gh/JulianCloudNTH/65/baseJul 4, 2026
181 of 183 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JulianCloudNTH/65/head branch July 4, 2026 17:05
ghost pushed a commit that referenced this pull request Jul 4, 2026
… dynamic shapes (#20712)
This PR was created by the merge bot to help merge the original PR into
the main branch.
ghstack PR number: #20573 by
@JulianCloudNTH
^ Please use this as the source of truth for the PR details, comments,
and reviews
ghstack PR base:
https://github.com/pytorch/executorch/tree/gh/JulianCloudNTH/65/base
ghstack PR head:
https://github.com/pytorch/executorch/tree/gh/JulianCloudNTH/65/head
Merge bot PR base: https://github.com/pytorch/executorch/tree/main
Merge bot PR head:
https://github.com/pytorch/executorch/tree/gh/JulianCloudNTH/65/orig
@diff-train-skip-merge
---------
Co-authored-by: Julian Ng-Thow-Hing <juliannth@meta.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@psiddh@nil-is-all@JCNTH
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[ExecuTorch][WebGPU] SymInt arithmetic ops (add/sub/mul/floordiv) for dynamic shapes - #20573

Merged
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/65/basefrom
gh/JulianCloudNTH/65/head
Jul 4, 2026
Merged

[ExecuTorch][WebGPU] SymInt arithmetic ops (add/sub/mul/floordiv) for dynamic shapes#20573
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/65/basefrom
gh/JulianCloudNTH/65/head

Conversation

@ghost

@ghostghost commented Jun 28, 2026

Copy link
Copy Markdown

Stack from ghstack (oldest at bottom):

Register scalar SymInt arithmetic so dynamic-shape graphs lower without an "unsupported op" failure.

Problem: A dynamic-shape exported program emits scalar SymInt arithmetic nodes (add/sub/mul/floordiv) to compute live sizes and positions (e.g. input_pos + S, seq_len // n). The WebGPU backend registered only et_vk.select_as_symint.default, so WebGPUGraph::build() threw unsupported op: add when loading any dynamic .pte.

Solution:

  • Before: only select_as_symint produced a live SymInt; any arithmetic on it was unsupported.
  • After: add/sub/mul/floordiv each recompute their output SymInt from the operands via a resize hook whenever a live operand changes.

Implementation:

  • register_sym_binary reads each operand (live SymInt via read_symint, else a static Int), seeds the build-time value, and registers a resize hook on any live operand (set_symint on recompute).
  • floordiv rounds toward negative infinity (Python semantics).
  • Mirrors Vulkan backends/vulkan/runtime/graph/ops/impl/SymIntOps.cpp (add/sub/floordiv/mul).
  • Registered under the bare targets add/sub/mul/floordiv — distinct registry keys from the tensor aten.add.Tensor/aten.mul.Tensor ops.

Constraints: An output that folded to a static Int is a no-op. No GPU kernel, no dispatch, no change to the static-shape path. sym_size.int is intentionally not added here (depends on the tensor-shape dim-source API in the following diff).

Co-authored-with: Claude Code.
@exported-using-ghexport

Differential Revision: D109906102

Differential Revision: D109906102

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jun 28, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20573

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit a918948 with merge base 124b677 (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@ghost
ghost temporarily deployed to cadence June 28, 2026 16:22 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 28, 2026 16:22 — with GitHub Actions Inactive
@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 28, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@ghost

Copy link
Copy Markdown
Author

@claude review and check for any areas or opportunities for modularization

@claude

claudeBot commented Jun 29, 2026

Copy link
Copy Markdown

Claude finished @JulianCloudNTH's task in 2m 6s —— View job


Review: WebGPU SymInt arithmetic ops

  • Read changed file SelectAsSymint.cpp
  • Compare against Vulkan reference SymIntOps.cpp
  • Check WebGPUGraph API (read_symint/set_symint/add_resize_hook/propagate_resize)
  • Analyze correctness & modularization opportunities

Overall this is a clean, well-scoped change. The resize-hook approach is a genuine improvement over the Vulkan reference it mirrors, and the diff does exactly what the description claims. Notes below, ordered by importance.


Modularization (the main ask)

1. register_sym_binary is the right call — and better than Vulkan. The Vulkan reference (backends/vulkan/runtime/graph/ops/impl/SymIntOps.cpp:53-166) copy-pastes the full _impl + resize_*_node + registration boilerplate four times. Collapsing all four ops into one helper parameterized by a std::function (SelectAsSymint.cpp:51-74) is the more maintainable design. Good instinct.

2. File/directory naming no longer matches contents. This is the clearest modularization opportunity. The file now contains general SymInt arithmetic (add/sub/mul/floordiv), but lives in ops/select_as_symint/SelectAsSymint.cpp — a directory named after one specific op. The repo convention is strictly one directory per op/concern (ops/add/, ops/mul/, ops/select/, …). A reader looking for where floordiv is registered won't find it under a select_as_symint path.

Consider extracting the arithmetic ops into a sibling ops/sym_int/SymIntOps.cpp (mirroring the Vulkan filename), leaving select_as_symint to own only select_as_symint. The shared helper read_scalar would move with them. This keeps the grep-to-op mapping the codebase relies on. Fix this →(if you'd rather keep the diff small and land the stack first, this is fine as a fast-follow — just flagging the convention.)


Correctness

3. Division by zero in floordiv is UB (SelectAsSymint.cpp:90-91). x / y and x % y with y == 0 are undefined behavior. The Vulkan reference has the same gap, so this isn't a regression, and a floordiv node with a zero divisor shouldn't arise from a valid export — but a one-line guard (throw, or return 0) would make it fail-loud like the rest of this file (read_symint, set_symint both throw on misuse) instead of crashing.

4. INT_MIN % y / INT_MIN / -1 overflow — theoretical only, since these values are tensor sizes/positions. Not worth handling; noting for completeness.

5. Self-referential operands register the hook twice. If a == b (e.g. x + x), both if branches at :68-72 fire and recompute is registered twice on the same SymInt. Harmless (recompute is idempotent, and set_symint no-ops when the value is unchanged), just slightly redundant. An if (a_is_symint) add_hook(a); else if (b_is_symint)… won't work since both may legitimately be distinct live operands — a b != a guard on the second branch would be the precise fix if you want it.


Minor / style

6. read_scalar silently assumes non-SymInt ⇒ Int (:43-48). If an operand were a Double (or any non-Int constant), get_int reads an unrelated flat array and returns a wrong value rather than erroring. Scalar arithmetic operands should only ever be Int/SymInt, so this is safe in practice; a comment or a type assert would document the assumption.

7. Chained arithmetic relies on registration order.propagate_resize (WebGPUGraph.cpp:148-158) iterates resize_hooks_ once in registration order; a hook that recomputes out marks out dirty so a later-registered hook keyed on out still fires in the same pass. This works only because ops are registered in topological order during build(). That's a reasonable invariant, but it's load-bearing and currently undocumented — worth a one-line comment on add_resize_hook or propagate_resize so a future single-pass refactor doesn't silently break multi-level SymInt chains (a + b, then that * c).


Verification

I could not build/run the WebGPU backend in this environment, so the above is static analysis only. Worth confirming with an actual dynamic-shape .pte that exercises a chained expression (e.g. (input_pos + S) // n) end-to-end, which would validate both the floor rounding (:88-95) and the chaining behavior in note 7.
· branch gh/JulianCloudNTH/65/head

[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:28 — with GitHub Actions Inactive
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 21:20 — with GitHub Actions Inactive
@meta-codesync
meta-codesyncBot merged commit ae4b83e into gh/JulianCloudNTH/65/baseJul 4, 2026
181 of 183 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JulianCloudNTH/65/head branch July 4, 2026 17:05
ghost pushed a commit that referenced this pull request Jul 4, 2026
… dynamic shapes (#20712)
This PR was created by the merge bot to help merge the original PR into
the main branch.
ghstack PR number: #20573 by
@JulianCloudNTH
^ Please use this as the source of truth for the PR details, comments,
and reviews
ghstack PR base:
https://github.com/pytorch/executorch/tree/gh/JulianCloudNTH/65/base
ghstack PR head:
https://github.com/pytorch/executorch/tree/gh/JulianCloudNTH/65/head
Merge bot PR base: https://github.com/pytorch/executorch/tree/main
Merge bot PR head:
https://github.com/pytorch/executorch/tree/gh/JulianCloudNTH/65/orig
@diff-train-skip-merge
---------
Co-authored-by: Julian Ng-Thow-Hing <juliannth@meta.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@psiddh@nil-is-all@JCNTH
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[ExecuTorch][WebGPU] SymInt arithmetic ops (add/sub/mul/floordiv) for dynamic shapes - #20573

Merged
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/65/basefrom
gh/JulianCloudNTH/65/head
Jul 4, 2026
Merged

[ExecuTorch][WebGPU] SymInt arithmetic ops (add/sub/mul/floordiv) for dynamic shapes#20573
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/65/basefrom
gh/JulianCloudNTH/65/head

Conversation

@ghost

@ghostghost commented Jun 28, 2026

Copy link
Copy Markdown

Stack from ghstack (oldest at bottom):

Register scalar SymInt arithmetic so dynamic-shape graphs lower without an "unsupported op" failure.

Problem: A dynamic-shape exported program emits scalar SymInt arithmetic nodes (add/sub/mul/floordiv) to compute live sizes and positions (e.g. input_pos + S, seq_len // n). The WebGPU backend registered only et_vk.select_as_symint.default, so WebGPUGraph::build() threw unsupported op: add when loading any dynamic .pte.

Solution:

  • Before: only select_as_symint produced a live SymInt; any arithmetic on it was unsupported.
  • After: add/sub/mul/floordiv each recompute their output SymInt from the operands via a resize hook whenever a live operand changes.

Implementation:

  • register_sym_binary reads each operand (live SymInt via read_symint, else a static Int), seeds the build-time value, and registers a resize hook on any live operand (set_symint on recompute).
  • floordiv rounds toward negative infinity (Python semantics).
  • Mirrors Vulkan backends/vulkan/runtime/graph/ops/impl/SymIntOps.cpp (add/sub/floordiv/mul).
  • Registered under the bare targets add/sub/mul/floordiv — distinct registry keys from the tensor aten.add.Tensor/aten.mul.Tensor ops.

Constraints: An output that folded to a static Int is a no-op. No GPU kernel, no dispatch, no change to the static-shape path. sym_size.int is intentionally not added here (depends on the tensor-shape dim-source API in the following diff).

Co-authored-with: Claude Code.
@exported-using-ghexport

Differential Revision: D109906102

Differential Revision: D109906102

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jun 28, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20573

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit a918948 with merge base 124b677 (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@ghost
ghost temporarily deployed to cadence June 28, 2026 16:22 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 28, 2026 16:22 — with GitHub Actions Inactive
@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 28, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@ghost

Copy link
Copy Markdown
Author

@claude review and check for any areas or opportunities for modularization

@claude

claudeBot commented Jun 29, 2026

Copy link
Copy Markdown

Claude finished @JulianCloudNTH's task in 2m 6s —— View job


Review: WebGPU SymInt arithmetic ops

  • Read changed file SelectAsSymint.cpp
  • Compare against Vulkan reference SymIntOps.cpp
  • Check WebGPUGraph API (read_symint/set_symint/add_resize_hook/propagate_resize)
  • Analyze correctness & modularization opportunities

Overall this is a clean, well-scoped change. The resize-hook approach is a genuine improvement over the Vulkan reference it mirrors, and the diff does exactly what the description claims. Notes below, ordered by importance.


Modularization (the main ask)

1. register_sym_binary is the right call — and better than Vulkan. The Vulkan reference (backends/vulkan/runtime/graph/ops/impl/SymIntOps.cpp:53-166) copy-pastes the full _impl + resize_*_node + registration boilerplate four times. Collapsing all four ops into one helper parameterized by a std::function (SelectAsSymint.cpp:51-74) is the more maintainable design. Good instinct.

2. File/directory naming no longer matches contents. This is the clearest modularization opportunity. The file now contains general SymInt arithmetic (add/sub/mul/floordiv), but lives in ops/select_as_symint/SelectAsSymint.cpp — a directory named after one specific op. The repo convention is strictly one directory per op/concern (ops/add/, ops/mul/, ops/select/, …). A reader looking for where floordiv is registered won't find it under a select_as_symint path.

Consider extracting the arithmetic ops into a sibling ops/sym_int/SymIntOps.cpp (mirroring the Vulkan filename), leaving select_as_symint to own only select_as_symint. The shared helper read_scalar would move with them. This keeps the grep-to-op mapping the codebase relies on. Fix this →(if you'd rather keep the diff small and land the stack first, this is fine as a fast-follow — just flagging the convention.)


Correctness

3. Division by zero in floordiv is UB (SelectAsSymint.cpp:90-91). x / y and x % y with y == 0 are undefined behavior. The Vulkan reference has the same gap, so this isn't a regression, and a floordiv node with a zero divisor shouldn't arise from a valid export — but a one-line guard (throw, or return 0) would make it fail-loud like the rest of this file (read_symint, set_symint both throw on misuse) instead of crashing.

4. INT_MIN % y / INT_MIN / -1 overflow — theoretical only, since these values are tensor sizes/positions. Not worth handling; noting for completeness.

5. Self-referential operands register the hook twice. If a == b (e.g. x + x), both if branches at :68-72 fire and recompute is registered twice on the same SymInt. Harmless (recompute is idempotent, and set_symint no-ops when the value is unchanged), just slightly redundant. An if (a_is_symint) add_hook(a); else if (b_is_symint)… won't work since both may legitimately be distinct live operands — a b != a guard on the second branch would be the precise fix if you want it.


Minor / style

6. read_scalar silently assumes non-SymInt ⇒ Int (:43-48). If an operand were a Double (or any non-Int constant), get_int reads an unrelated flat array and returns a wrong value rather than erroring. Scalar arithmetic operands should only ever be Int/SymInt, so this is safe in practice; a comment or a type assert would document the assumption.

7. Chained arithmetic relies on registration order.propagate_resize (WebGPUGraph.cpp:148-158) iterates resize_hooks_ once in registration order; a hook that recomputes out marks out dirty so a later-registered hook keyed on out still fires in the same pass. This works only because ops are registered in topological order during build(). That's a reasonable invariant, but it's load-bearing and currently undocumented — worth a one-line comment on add_resize_hook or propagate_resize so a future single-pass refactor doesn't silently break multi-level SymInt chains (a + b, then that * c).


Verification

I could not build/run the WebGPU backend in this environment, so the above is static analysis only. Worth confirming with an actual dynamic-shape .pte that exercises a chained expression (e.g. (input_pos + S) // n) end-to-end, which would validate both the floor rounding (:88-95) and the chaining behavior in note 7.
· branch gh/JulianCloudNTH/65/head

[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:28 — with GitHub Actions Inactive
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 21:20 — with GitHub Actions Inactive
@meta-codesync
meta-codesyncBot merged commit ae4b83e into gh/JulianCloudNTH/65/baseJul 4, 2026
181 of 183 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JulianCloudNTH/65/head branch July 4, 2026 17:05
ghost pushed a commit that referenced this pull request Jul 4, 2026
… dynamic shapes (#20712)
This PR was created by the merge bot to help merge the original PR into
the main branch.
ghstack PR number: #20573 by
@JulianCloudNTH
^ Please use this as the source of truth for the PR details, comments,
and reviews
ghstack PR base:
https://github.com/pytorch/executorch/tree/gh/JulianCloudNTH/65/base
ghstack PR head:
https://github.com/pytorch/executorch/tree/gh/JulianCloudNTH/65/head
Merge bot PR base: https://github.com/pytorch/executorch/tree/main
Merge bot PR head:
https://github.com/pytorch/executorch/tree/gh/JulianCloudNTH/65/orig
@diff-train-skip-merge
---------
Co-authored-by: Julian Ng-Thow-Hing <juliannth@meta.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@psiddh@nil-is-all@JCNTH
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

[ExecuTorch][WebGPU] SymInt arithmetic ops (add/sub/mul/floordiv) for dynamic shapes - #20573

Merged
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/65/basefrom
gh/JulianCloudNTH/65/head
Jul 4, 2026
Merged

[ExecuTorch][WebGPU] SymInt arithmetic ops (add/sub/mul/floordiv) for dynamic shapes#20573
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/65/basefrom
gh/JulianCloudNTH/65/head

Conversation

@ghost

@ghostghost commented Jun 28, 2026

Copy link
Copy Markdown

Stack from ghstack (oldest at bottom):

Register scalar SymInt arithmetic so dynamic-shape graphs lower without an "unsupported op" failure.

Problem: A dynamic-shape exported program emits scalar SymInt arithmetic nodes (add/sub/mul/floordiv) to compute live sizes and positions (e.g. input_pos + S, seq_len // n). The WebGPU backend registered only et_vk.select_as_symint.default, so WebGPUGraph::build() threw unsupported op: add when loading any dynamic .pte.

Solution:

  • Before: only select_as_symint produced a live SymInt; any arithmetic on it was unsupported.
  • After: add/sub/mul/floordiv each recompute their output SymInt from the operands via a resize hook whenever a live operand changes.

Implementation:

  • register_sym_binary reads each operand (live SymInt via read_symint, else a static Int), seeds the build-time value, and registers a resize hook on any live operand (set_symint on recompute).
  • floordiv rounds toward negative infinity (Python semantics).
  • Mirrors Vulkan backends/vulkan/runtime/graph/ops/impl/SymIntOps.cpp (add/sub/floordiv/mul).
  • Registered under the bare targets add/sub/mul/floordiv — distinct registry keys from the tensor aten.add.Tensor/aten.mul.Tensor ops.

Constraints: An output that folded to a static Int is a no-op. No GPU kernel, no dispatch, no change to the static-shape path. sym_size.int is intentionally not added here (depends on the tensor-shape dim-source API in the following diff).

Co-authored-with: Claude Code.
@exported-using-ghexport

Differential Revision: D109906102

Differential Revision: D109906102

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jun 28, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20573

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit a918948 with merge base 124b677 (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@ghost
ghost temporarily deployed to cadence June 28, 2026 16:22 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 28, 2026 16:22 — with GitHub Actions Inactive
@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 28, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@ghost

Copy link
Copy Markdown
Author

@claude review and check for any areas or opportunities for modularization

@claude

claudeBot commented Jun 29, 2026

Copy link
Copy Markdown

Claude finished @JulianCloudNTH's task in 2m 6s —— View job


Review: WebGPU SymInt arithmetic ops

  • Read changed file SelectAsSymint.cpp
  • Compare against Vulkan reference SymIntOps.cpp
  • Check WebGPUGraph API (read_symint/set_symint/add_resize_hook/propagate_resize)
  • Analyze correctness & modularization opportunities

Overall this is a clean, well-scoped change. The resize-hook approach is a genuine improvement over the Vulkan reference it mirrors, and the diff does exactly what the description claims. Notes below, ordered by importance.


Modularization (the main ask)

1. register_sym_binary is the right call — and better than Vulkan. The Vulkan reference (backends/vulkan/runtime/graph/ops/impl/SymIntOps.cpp:53-166) copy-pastes the full _impl + resize_*_node + registration boilerplate four times. Collapsing all four ops into one helper parameterized by a std::function (SelectAsSymint.cpp:51-74) is the more maintainable design. Good instinct.

2. File/directory naming no longer matches contents. This is the clearest modularization opportunity. The file now contains general SymInt arithmetic (add/sub/mul/floordiv), but lives in ops/select_as_symint/SelectAsSymint.cpp — a directory named after one specific op. The repo convention is strictly one directory per op/concern (ops/add/, ops/mul/, ops/select/, …). A reader looking for where floordiv is registered won't find it under a select_as_symint path.

Consider extracting the arithmetic ops into a sibling ops/sym_int/SymIntOps.cpp (mirroring the Vulkan filename), leaving select_as_symint to own only select_as_symint. The shared helper read_scalar would move with them. This keeps the grep-to-op mapping the codebase relies on. Fix this →(if you'd rather keep the diff small and land the stack first, this is fine as a fast-follow — just flagging the convention.)


Correctness

3. Division by zero in floordiv is UB (SelectAsSymint.cpp:90-91). x / y and x % y with y == 0 are undefined behavior. The Vulkan reference has the same gap, so this isn't a regression, and a floordiv node with a zero divisor shouldn't arise from a valid export — but a one-line guard (throw, or return 0) would make it fail-loud like the rest of this file (read_symint, set_symint both throw on misuse) instead of crashing.

4. INT_MIN % y / INT_MIN / -1 overflow — theoretical only, since these values are tensor sizes/positions. Not worth handling; noting for completeness.

5. Self-referential operands register the hook twice. If a == b (e.g. x + x), both if branches at :68-72 fire and recompute is registered twice on the same SymInt. Harmless (recompute is idempotent, and set_symint no-ops when the value is unchanged), just slightly redundant. An if (a_is_symint) add_hook(a); else if (b_is_symint)… won't work since both may legitimately be distinct live operands — a b != a guard on the second branch would be the precise fix if you want it.


Minor / style

6. read_scalar silently assumes non-SymInt ⇒ Int (:43-48). If an operand were a Double (or any non-Int constant), get_int reads an unrelated flat array and returns a wrong value rather than erroring. Scalar arithmetic operands should only ever be Int/SymInt, so this is safe in practice; a comment or a type assert would document the assumption.

7. Chained arithmetic relies on registration order.propagate_resize (WebGPUGraph.cpp:148-158) iterates resize_hooks_ once in registration order; a hook that recomputes out marks out dirty so a later-registered hook keyed on out still fires in the same pass. This works only because ops are registered in topological order during build(). That's a reasonable invariant, but it's load-bearing and currently undocumented — worth a one-line comment on add_resize_hook or propagate_resize so a future single-pass refactor doesn't silently break multi-level SymInt chains (a + b, then that * c).


Verification

I could not build/run the WebGPU backend in this environment, so the above is static analysis only. Worth confirming with an actual dynamic-shape .pte that exercises a chained expression (e.g. (input_pos + S) // n) end-to-end, which would validate both the floor rounding (:88-95) and the chaining behavior in note 7.
· branch gh/JulianCloudNTH/65/head

[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:28 — with GitHub Actions Inactive
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 21:20 — with GitHub Actions Inactive
@meta-codesync
meta-codesyncBot merged commit ae4b83e into gh/JulianCloudNTH/65/baseJul 4, 2026
181 of 183 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JulianCloudNTH/65/head branch July 4, 2026 17:05
ghost pushed a commit that referenced this pull request Jul 4, 2026
… dynamic shapes (#20712)
This PR was created by the merge bot to help merge the original PR into
the main branch.
ghstack PR number: #20573 by
@JulianCloudNTH
^ Please use this as the source of truth for the PR details, comments,
and reviews
ghstack PR base:
https://github.com/pytorch/executorch/tree/gh/JulianCloudNTH/65/base
ghstack PR head:
https://github.com/pytorch/executorch/tree/gh/JulianCloudNTH/65/head
Merge bot PR base: https://github.com/pytorch/executorch/tree/main
Merge bot PR head:
https://github.com/pytorch/executorch/tree/gh/JulianCloudNTH/65/orig
@diff-train-skip-merge
---------
Co-authored-by: Julian Ng-Thow-Hing <juliannth@meta.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@psiddh@nil-is-all@JCNTH
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[ExecuTorch][WebGPU] SymInt arithmetic ops (add/sub/mul/floordiv) for dynamic shapes - #20573

Merged
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/65/basefrom
gh/JulianCloudNTH/65/head
Jul 4, 2026
Merged

[ExecuTorch][WebGPU] SymInt arithmetic ops (add/sub/mul/floordiv) for dynamic shapes#20573
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/65/basefrom
gh/JulianCloudNTH/65/head

Conversation

@ghost

@ghostghost commented Jun 28, 2026

Copy link
Copy Markdown

Stack from ghstack (oldest at bottom):

Register scalar SymInt arithmetic so dynamic-shape graphs lower without an "unsupported op" failure.

Problem: A dynamic-shape exported program emits scalar SymInt arithmetic nodes (add/sub/mul/floordiv) to compute live sizes and positions (e.g. input_pos + S, seq_len // n). The WebGPU backend registered only et_vk.select_as_symint.default, so WebGPUGraph::build() threw unsupported op: add when loading any dynamic .pte.

Solution:

  • Before: only select_as_symint produced a live SymInt; any arithmetic on it was unsupported.
  • After: add/sub/mul/floordiv each recompute their output SymInt from the operands via a resize hook whenever a live operand changes.

Implementation:

  • register_sym_binary reads each operand (live SymInt via read_symint, else a static Int), seeds the build-time value, and registers a resize hook on any live operand (set_symint on recompute).
  • floordiv rounds toward negative infinity (Python semantics).
  • Mirrors Vulkan backends/vulkan/runtime/graph/ops/impl/SymIntOps.cpp (add/sub/floordiv/mul).
  • Registered under the bare targets add/sub/mul/floordiv — distinct registry keys from the tensor aten.add.Tensor/aten.mul.Tensor ops.

Constraints: An output that folded to a static Int is a no-op. No GPU kernel, no dispatch, no change to the static-shape path. sym_size.int is intentionally not added here (depends on the tensor-shape dim-source API in the following diff).

Co-authored-with: Claude Code.
@exported-using-ghexport

Differential Revision: D109906102

Differential Revision: D109906102

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jun 28, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20573

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit a918948 with merge base 124b677 (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@ghost
ghost temporarily deployed to cadence June 28, 2026 16:22 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 28, 2026 16:22 — with GitHub Actions Inactive
@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 28, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@ghost

Copy link
Copy Markdown
Author

@claude review and check for any areas or opportunities for modularization

@claude

claudeBot commented Jun 29, 2026

Copy link
Copy Markdown

Claude finished @JulianCloudNTH's task in 2m 6s —— View job


Review: WebGPU SymInt arithmetic ops

  • Read changed file SelectAsSymint.cpp
  • Compare against Vulkan reference SymIntOps.cpp
  • Check WebGPUGraph API (read_symint/set_symint/add_resize_hook/propagate_resize)
  • Analyze correctness & modularization opportunities

Overall this is a clean, well-scoped change. The resize-hook approach is a genuine improvement over the Vulkan reference it mirrors, and the diff does exactly what the description claims. Notes below, ordered by importance.


Modularization (the main ask)

1. register_sym_binary is the right call — and better than Vulkan. The Vulkan reference (backends/vulkan/runtime/graph/ops/impl/SymIntOps.cpp:53-166) copy-pastes the full _impl + resize_*_node + registration boilerplate four times. Collapsing all four ops into one helper parameterized by a std::function (SelectAsSymint.cpp:51-74) is the more maintainable design. Good instinct.

2. File/directory naming no longer matches contents. This is the clearest modularization opportunity. The file now contains general SymInt arithmetic (add/sub/mul/floordiv), but lives in ops/select_as_symint/SelectAsSymint.cpp — a directory named after one specific op. The repo convention is strictly one directory per op/concern (ops/add/, ops/mul/, ops/select/, …). A reader looking for where floordiv is registered won't find it under a select_as_symint path.

Consider extracting the arithmetic ops into a sibling ops/sym_int/SymIntOps.cpp (mirroring the Vulkan filename), leaving select_as_symint to own only select_as_symint. The shared helper read_scalar would move with them. This keeps the grep-to-op mapping the codebase relies on. Fix this →(if you'd rather keep the diff small and land the stack first, this is fine as a fast-follow — just flagging the convention.)


Correctness

3. Division by zero in floordiv is UB (SelectAsSymint.cpp:90-91). x / y and x % y with y == 0 are undefined behavior. The Vulkan reference has the same gap, so this isn't a regression, and a floordiv node with a zero divisor shouldn't arise from a valid export — but a one-line guard (throw, or return 0) would make it fail-loud like the rest of this file (read_symint, set_symint both throw on misuse) instead of crashing.

4. INT_MIN % y / INT_MIN / -1 overflow — theoretical only, since these values are tensor sizes/positions. Not worth handling; noting for completeness.

5. Self-referential operands register the hook twice. If a == b (e.g. x + x), both if branches at :68-72 fire and recompute is registered twice on the same SymInt. Harmless (recompute is idempotent, and set_symint no-ops when the value is unchanged), just slightly redundant. An if (a_is_symint) add_hook(a); else if (b_is_symint)… won't work since both may legitimately be distinct live operands — a b != a guard on the second branch would be the precise fix if you want it.


Minor / style

6. read_scalar silently assumes non-SymInt ⇒ Int (:43-48). If an operand were a Double (or any non-Int constant), get_int reads an unrelated flat array and returns a wrong value rather than erroring. Scalar arithmetic operands should only ever be Int/SymInt, so this is safe in practice; a comment or a type assert would document the assumption.

7. Chained arithmetic relies on registration order.propagate_resize (WebGPUGraph.cpp:148-158) iterates resize_hooks_ once in registration order; a hook that recomputes out marks out dirty so a later-registered hook keyed on out still fires in the same pass. This works only because ops are registered in topological order during build(). That's a reasonable invariant, but it's load-bearing and currently undocumented — worth a one-line comment on add_resize_hook or propagate_resize so a future single-pass refactor doesn't silently break multi-level SymInt chains (a + b, then that * c).


Verification

I could not build/run the WebGPU backend in this environment, so the above is static analysis only. Worth confirming with an actual dynamic-shape .pte that exercises a chained expression (e.g. (input_pos + S) // n) end-to-end, which would validate both the floor rounding (:88-95) and the chaining behavior in note 7.
· branch gh/JulianCloudNTH/65/head

[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:28 — with GitHub Actions Inactive
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 21:20 — with GitHub Actions Inactive
@meta-codesync
meta-codesyncBot merged commit ae4b83e into gh/JulianCloudNTH/65/baseJul 4, 2026
181 of 183 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JulianCloudNTH/65/head branch July 4, 2026 17:05
ghost pushed a commit that referenced this pull request Jul 4, 2026
… dynamic shapes (#20712)
This PR was created by the merge bot to help merge the original PR into
the main branch.
ghstack PR number: #20573 by
@JulianCloudNTH
^ Please use this as the source of truth for the PR details, comments,
and reviews
ghstack PR base:
https://github.com/pytorch/executorch/tree/gh/JulianCloudNTH/65/base
ghstack PR head:
https://github.com/pytorch/executorch/tree/gh/JulianCloudNTH/65/head
Merge bot PR base: https://github.com/pytorch/executorch/tree/main
Merge bot PR head:
https://github.com/pytorch/executorch/tree/gh/JulianCloudNTH/65/orig
@diff-train-skip-merge
---------
Co-authored-by: Julian Ng-Thow-Hing <juliannth@meta.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@psiddh@nil-is-all@JCNTH
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[ExecuTorch][WebGPU] SymInt arithmetic ops (add/sub/mul/floordiv) for dynamic shapes - #20573

Merged
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/65/basefrom
gh/JulianCloudNTH/65/head
Jul 4, 2026
Merged

[ExecuTorch][WebGPU] SymInt arithmetic ops (add/sub/mul/floordiv) for dynamic shapes#20573
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/65/basefrom
gh/JulianCloudNTH/65/head

Conversation

@ghost

@ghostghost commented Jun 28, 2026

Copy link
Copy Markdown

Stack from ghstack (oldest at bottom):

Register scalar SymInt arithmetic so dynamic-shape graphs lower without an "unsupported op" failure.

Problem: A dynamic-shape exported program emits scalar SymInt arithmetic nodes (add/sub/mul/floordiv) to compute live sizes and positions (e.g. input_pos + S, seq_len // n). The WebGPU backend registered only et_vk.select_as_symint.default, so WebGPUGraph::build() threw unsupported op: add when loading any dynamic .pte.

Solution:

  • Before: only select_as_symint produced a live SymInt; any arithmetic on it was unsupported.
  • After: add/sub/mul/floordiv each recompute their output SymInt from the operands via a resize hook whenever a live operand changes.

Implementation:

  • register_sym_binary reads each operand (live SymInt via read_symint, else a static Int), seeds the build-time value, and registers a resize hook on any live operand (set_symint on recompute).
  • floordiv rounds toward negative infinity (Python semantics).
  • Mirrors Vulkan backends/vulkan/runtime/graph/ops/impl/SymIntOps.cpp (add/sub/floordiv/mul).
  • Registered under the bare targets add/sub/mul/floordiv — distinct registry keys from the tensor aten.add.Tensor/aten.mul.Tensor ops.

Constraints: An output that folded to a static Int is a no-op. No GPU kernel, no dispatch, no change to the static-shape path. sym_size.int is intentionally not added here (depends on the tensor-shape dim-source API in the following diff).

Co-authored-with: Claude Code.
@exported-using-ghexport

Differential Revision: D109906102

Differential Revision: D109906102

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jun 28, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20573

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit a918948 with merge base 124b677 (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@ghost
ghost temporarily deployed to cadence June 28, 2026 16:22 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 28, 2026 16:22 — with GitHub Actions Inactive
@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 28, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@ghost

Copy link
Copy Markdown
Author

@claude review and check for any areas or opportunities for modularization

@claude

claudeBot commented Jun 29, 2026

Copy link
Copy Markdown

Claude finished @JulianCloudNTH's task in 2m 6s —— View job


Review: WebGPU SymInt arithmetic ops

  • Read changed file SelectAsSymint.cpp
  • Compare against Vulkan reference SymIntOps.cpp
  • Check WebGPUGraph API (read_symint/set_symint/add_resize_hook/propagate_resize)
  • Analyze correctness & modularization opportunities

Overall this is a clean, well-scoped change. The resize-hook approach is a genuine improvement over the Vulkan reference it mirrors, and the diff does exactly what the description claims. Notes below, ordered by importance.


Modularization (the main ask)

1. register_sym_binary is the right call — and better than Vulkan. The Vulkan reference (backends/vulkan/runtime/graph/ops/impl/SymIntOps.cpp:53-166) copy-pastes the full _impl + resize_*_node + registration boilerplate four times. Collapsing all four ops into one helper parameterized by a std::function (SelectAsSymint.cpp:51-74) is the more maintainable design. Good instinct.

2. File/directory naming no longer matches contents. This is the clearest modularization opportunity. The file now contains general SymInt arithmetic (add/sub/mul/floordiv), but lives in ops/select_as_symint/SelectAsSymint.cpp — a directory named after one specific op. The repo convention is strictly one directory per op/concern (ops/add/, ops/mul/, ops/select/, …). A reader looking for where floordiv is registered won't find it under a select_as_symint path.

Consider extracting the arithmetic ops into a sibling ops/sym_int/SymIntOps.cpp (mirroring the Vulkan filename), leaving select_as_symint to own only select_as_symint. The shared helper read_scalar would move with them. This keeps the grep-to-op mapping the codebase relies on. Fix this →(if you'd rather keep the diff small and land the stack first, this is fine as a fast-follow — just flagging the convention.)


Correctness

3. Division by zero in floordiv is UB (SelectAsSymint.cpp:90-91). x / y and x % y with y == 0 are undefined behavior. The Vulkan reference has the same gap, so this isn't a regression, and a floordiv node with a zero divisor shouldn't arise from a valid export — but a one-line guard (throw, or return 0) would make it fail-loud like the rest of this file (read_symint, set_symint both throw on misuse) instead of crashing.

4. INT_MIN % y / INT_MIN / -1 overflow — theoretical only, since these values are tensor sizes/positions. Not worth handling; noting for completeness.

5. Self-referential operands register the hook twice. If a == b (e.g. x + x), both if branches at :68-72 fire and recompute is registered twice on the same SymInt. Harmless (recompute is idempotent, and set_symint no-ops when the value is unchanged), just slightly redundant. An if (a_is_symint) add_hook(a); else if (b_is_symint)… won't work since both may legitimately be distinct live operands — a b != a guard on the second branch would be the precise fix if you want it.


Minor / style

6. read_scalar silently assumes non-SymInt ⇒ Int (:43-48). If an operand were a Double (or any non-Int constant), get_int reads an unrelated flat array and returns a wrong value rather than erroring. Scalar arithmetic operands should only ever be Int/SymInt, so this is safe in practice; a comment or a type assert would document the assumption.

7. Chained arithmetic relies on registration order.propagate_resize (WebGPUGraph.cpp:148-158) iterates resize_hooks_ once in registration order; a hook that recomputes out marks out dirty so a later-registered hook keyed on out still fires in the same pass. This works only because ops are registered in topological order during build(). That's a reasonable invariant, but it's load-bearing and currently undocumented — worth a one-line comment on add_resize_hook or propagate_resize so a future single-pass refactor doesn't silently break multi-level SymInt chains (a + b, then that * c).


Verification

I could not build/run the WebGPU backend in this environment, so the above is static analysis only. Worth confirming with an actual dynamic-shape .pte that exercises a chained expression (e.g. (input_pos + S) // n) end-to-end, which would validate both the floor rounding (:88-95) and the chaining behavior in note 7.
· branch gh/JulianCloudNTH/65/head

[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:28 — with GitHub Actions Inactive
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 21:20 — with GitHub Actions Inactive
@meta-codesync
meta-codesyncBot merged commit ae4b83e into gh/JulianCloudNTH/65/baseJul 4, 2026
181 of 183 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JulianCloudNTH/65/head branch July 4, 2026 17:05
ghost pushed a commit that referenced this pull request Jul 4, 2026
… dynamic shapes (#20712)
This PR was created by the merge bot to help merge the original PR into
the main branch.
ghstack PR number: #20573 by
@JulianCloudNTH
^ Please use this as the source of truth for the PR details, comments,
and reviews
ghstack PR base:
https://github.com/pytorch/executorch/tree/gh/JulianCloudNTH/65/base
ghstack PR head:
https://github.com/pytorch/executorch/tree/gh/JulianCloudNTH/65/head
Merge bot PR base: https://github.com/pytorch/executorch/tree/main
Merge bot PR head:
https://github.com/pytorch/executorch/tree/gh/JulianCloudNTH/65/orig
@diff-train-skip-merge
---------
Co-authored-by: Julian Ng-Thow-Hing <juliannth@meta.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@psiddh@nil-is-all@JCNTH
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

[ExecuTorch][WebGPU] SymInt arithmetic ops (add/sub/mul/floordiv) for dynamic shapes - #20573

Merged
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/65/basefrom
gh/JulianCloudNTH/65/head
Jul 4, 2026
Merged

[ExecuTorch][WebGPU] SymInt arithmetic ops (add/sub/mul/floordiv) for dynamic shapes#20573
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/65/basefrom
gh/JulianCloudNTH/65/head

Conversation

@ghost

@ghostghost commented Jun 28, 2026

Copy link
Copy Markdown

Stack from ghstack (oldest at bottom):

Register scalar SymInt arithmetic so dynamic-shape graphs lower without an "unsupported op" failure.

Problem: A dynamic-shape exported program emits scalar SymInt arithmetic nodes (add/sub/mul/floordiv) to compute live sizes and positions (e.g. input_pos + S, seq_len // n). The WebGPU backend registered only et_vk.select_as_symint.default, so WebGPUGraph::build() threw unsupported op: add when loading any dynamic .pte.

Solution:

  • Before: only select_as_symint produced a live SymInt; any arithmetic on it was unsupported.
  • After: add/sub/mul/floordiv each recompute their output SymInt from the operands via a resize hook whenever a live operand changes.

Implementation:

  • register_sym_binary reads each operand (live SymInt via read_symint, else a static Int), seeds the build-time value, and registers a resize hook on any live operand (set_symint on recompute).
  • floordiv rounds toward negative infinity (Python semantics).
  • Mirrors Vulkan backends/vulkan/runtime/graph/ops/impl/SymIntOps.cpp (add/sub/floordiv/mul).
  • Registered under the bare targets add/sub/mul/floordiv — distinct registry keys from the tensor aten.add.Tensor/aten.mul.Tensor ops.

Constraints: An output that folded to a static Int is a no-op. No GPU kernel, no dispatch, no change to the static-shape path. sym_size.int is intentionally not added here (depends on the tensor-shape dim-source API in the following diff).

Co-authored-with: Claude Code.
@exported-using-ghexport

Differential Revision: D109906102

Differential Revision: D109906102

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jun 28, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20573

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit a918948 with merge base 124b677 (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@ghost
ghost temporarily deployed to cadence June 28, 2026 16:22 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 28, 2026 16:22 — with GitHub Actions Inactive
@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 28, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@ghost

Copy link
Copy Markdown
Author

@claude review and check for any areas or opportunities for modularization

@claude

claudeBot commented Jun 29, 2026

Copy link
Copy Markdown

Claude finished @JulianCloudNTH's task in 2m 6s —— View job


Review: WebGPU SymInt arithmetic ops

  • Read changed file SelectAsSymint.cpp
  • Compare against Vulkan reference SymIntOps.cpp
  • Check WebGPUGraph API (read_symint/set_symint/add_resize_hook/propagate_resize)
  • Analyze correctness & modularization opportunities

Overall this is a clean, well-scoped change. The resize-hook approach is a genuine improvement over the Vulkan reference it mirrors, and the diff does exactly what the description claims. Notes below, ordered by importance.


Modularization (the main ask)

1. register_sym_binary is the right call — and better than Vulkan. The Vulkan reference (backends/vulkan/runtime/graph/ops/impl/SymIntOps.cpp:53-166) copy-pastes the full _impl + resize_*_node + registration boilerplate four times. Collapsing all four ops into one helper parameterized by a std::function (SelectAsSymint.cpp:51-74) is the more maintainable design. Good instinct.

2. File/directory naming no longer matches contents. This is the clearest modularization opportunity. The file now contains general SymInt arithmetic (add/sub/mul/floordiv), but lives in ops/select_as_symint/SelectAsSymint.cpp — a directory named after one specific op. The repo convention is strictly one directory per op/concern (ops/add/, ops/mul/, ops/select/, …). A reader looking for where floordiv is registered won't find it under a select_as_symint path.

Consider extracting the arithmetic ops into a sibling ops/sym_int/SymIntOps.cpp (mirroring the Vulkan filename), leaving select_as_symint to own only select_as_symint. The shared helper read_scalar would move with them. This keeps the grep-to-op mapping the codebase relies on. Fix this →(if you'd rather keep the diff small and land the stack first, this is fine as a fast-follow — just flagging the convention.)


Correctness

3. Division by zero in floordiv is UB (SelectAsSymint.cpp:90-91). x / y and x % y with y == 0 are undefined behavior. The Vulkan reference has the same gap, so this isn't a regression, and a floordiv node with a zero divisor shouldn't arise from a valid export — but a one-line guard (throw, or return 0) would make it fail-loud like the rest of this file (read_symint, set_symint both throw on misuse) instead of crashing.

4. INT_MIN % y / INT_MIN / -1 overflow — theoretical only, since these values are tensor sizes/positions. Not worth handling; noting for completeness.

5. Self-referential operands register the hook twice. If a == b (e.g. x + x), both if branches at :68-72 fire and recompute is registered twice on the same SymInt. Harmless (recompute is idempotent, and set_symint no-ops when the value is unchanged), just slightly redundant. An if (a_is_symint) add_hook(a); else if (b_is_symint)… won't work since both may legitimately be distinct live operands — a b != a guard on the second branch would be the precise fix if you want it.


Minor / style

6. read_scalar silently assumes non-SymInt ⇒ Int (:43-48). If an operand were a Double (or any non-Int constant), get_int reads an unrelated flat array and returns a wrong value rather than erroring. Scalar arithmetic operands should only ever be Int/SymInt, so this is safe in practice; a comment or a type assert would document the assumption.

7. Chained arithmetic relies on registration order.propagate_resize (WebGPUGraph.cpp:148-158) iterates resize_hooks_ once in registration order; a hook that recomputes out marks out dirty so a later-registered hook keyed on out still fires in the same pass. This works only because ops are registered in topological order during build(). That's a reasonable invariant, but it's load-bearing and currently undocumented — worth a one-line comment on add_resize_hook or propagate_resize so a future single-pass refactor doesn't silently break multi-level SymInt chains (a + b, then that * c).


Verification

I could not build/run the WebGPU backend in this environment, so the above is static analysis only. Worth confirming with an actual dynamic-shape .pte that exercises a chained expression (e.g. (input_pos + S) // n) end-to-end, which would validate both the floor rounding (:88-95) and the chaining behavior in note 7.
· branch gh/JulianCloudNTH/65/head

[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:28 — with GitHub Actions Inactive
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 21:20 — with GitHub Actions Inactive
@meta-codesync
meta-codesyncBot merged commit ae4b83e into gh/JulianCloudNTH/65/baseJul 4, 2026
181 of 183 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JulianCloudNTH/65/head branch July 4, 2026 17:05
ghost pushed a commit that referenced this pull request Jul 4, 2026
… dynamic shapes (#20712)
This PR was created by the merge bot to help merge the original PR into
the main branch.
ghstack PR number: #20573 by
@JulianCloudNTH
^ Please use this as the source of truth for the PR details, comments,
and reviews
ghstack PR base:
https://github.com/pytorch/executorch/tree/gh/JulianCloudNTH/65/base
ghstack PR head:
https://github.com/pytorch/executorch/tree/gh/JulianCloudNTH/65/head
Merge bot PR base: https://github.com/pytorch/executorch/tree/main
Merge bot PR head:
https://github.com/pytorch/executorch/tree/gh/JulianCloudNTH/65/orig
@diff-train-skip-merge
---------
Co-authored-by: Julian Ng-Thow-Hing <juliannth@meta.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@psiddh@nil-is-all@JCNTH