[ExecuTorch][WebGPU] Dynamic resize hooks for sigmoid and select_copy - #20578

Merged
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/70/basefrom
gh/JulianCloudNTH/70/head
Jul 4, 2026
Merged

[ExecuTorch][WebGPU] Dynamic resize hooks for sigmoid and select_copy#20578
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/70/basefrom
gh/JulianCloudNTH/70/head

Conversation

@ghost

@ghostghost commented Jun 28, 2026

Copy link
Copy Markdown

Stack from ghstack (oldest at bottom):

Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.

Problem: Both ops baked their dispatch/params/output shape at build() for the max shape. select_copy was worse: a negative index (e.g. -1 for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.

Solution:

  • sigmoid (generic add_unary_op): a resize hook recomputes num_elements/dispatch and sets the output cur_dims (shape-preserving).
  • select_copy: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus dim), rebuilds the out/in TensorMeta UBOs and the dispatch.
  • Both keep their uniform buffer(s) alive via own_uniform_buffer.

Implementation:

  • The select out/in meta is rebuilt from synthetic WebGPUTensor{dims} via fill_tensor_meta (reads only .dims).
  • Mirrors Vulkan per-op resize_*_node.

Constraints: Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.

Co-authored-with: Claude Code.
@exported-using-ghexport

Differential Revision: D109906095

Differential Revision: D109906095

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jun 28, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20578

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Unrelated Failure

As of commit d658f5c with merge base 124b677 (image):

NEW FAILURE - The following job has failed:

FLAKY - The following job failed but was likely due to flakiness present on trunk:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@ghost

Copy link
Copy Markdown
Author

@claude review and check for any areas or opportunities for modularization

@claude

claudeBot commented Jun 29, 2026

Copy link
Copy Markdown

Claude finished @JulianCloudNTH's task in 2m 47s —— View job


Review: Dynamic resize hooks for sigmoid and select_copy

  • Read changed files & surrounding WebGPU runtime infrastructure
  • Analyze correctness of resize hooks
  • Identify modularization opportunities

The change is well-scoped and matches the established per-op resize-hook pattern (add, mul). The select_copy negative-index fix is correct and is a real bug fix: at build the index was normalized against the max dim, so -1 pointed past live data at smaller S; keeping raw_index and re-resolving against cur_dims(in_id)[dim] in the hook is the right fix. Both ops correctly switch to own_uniform_buffer so the hook can rewrite the UBOs.

Correctness — looks good, two minor notes

  1. Select.cpp — guards not re-checked on resize (acceptable). The build path validates nbytes == numel*4 (fp32) and the bind-group buffer size is pinned to the maxnbytes. The hook only rewrites meta/params/workgroup count, never buffer sizes — correct, since dtype and allocation are static and only live sizes shrink. Worth a one-line comment in the hook noting the bind-group sizes intentionally stay at max so it doesn't read as an omission.

  2. UnaryOp.cpp — build vs. hook compute numel from different tensors. Build uses out_tensor.nbytes/sizeof(float); the hook uses numel_of(cur_dims(in_id)). They're guaranteed equal (the op asserts in.nbytes == out.nbytes), so this is fine, but reading in in one place and out in the other is a small inconsistency. Computing both from the input would read more uniformly.

No issues with the lambda captures (all captured values are used; handles copied before own_uniform_buffer transfers ownership).

Modularization — the more substantive feedback

This PR is behavior-correct, but it lands on top of a growing amount of near-duplicate boilerplate across ops/. Three concrete, increasingly valuable extractions:

1. A meta-UBO rewrite helper (highest value, lowest risk). The "synthesize a WebGPUTensor from dims → fill_tensor_metawgpuQueueWriteBuffer" sequence now appears in select (Select.cpp:194-201) and mul (mul/BinaryOp.cpp:187-197), and is the kind of thing every future shape-changing op will copy. A small helper in TensorMeta.h would absorb it:

inlinevoidwrite_meta_from_dims(WebGPUGraph& g, WGPUBuffer buf,
const std::vector<int64_t>& dims) {
WebGPUTensor t; t.dims = dims;
TensorMeta m; fill_tensor_meta(t, &m);
wgpuQueueWriteBuffer(g.queue(), buf, 0, &m, sizeof(m));
}

Select's hook then drops to write_meta_from_dims(g, out_meta_buf, od); write_meta_from_dims(g, in_meta_buf, ind);. (A _broadcast overload covers mul.)

2. A set_workgroup_1d helper on the resize path. Every resize hook ends with the same g.dispatch_at(idx).workgroup_count_x = utils::compute_1d_workgroup_count(g.device(), numel, wg_size, name); (Select.cpp:206, UnaryOp.cpp:152, add/mul). A thin utils::set_dispatch_1d(g, dispatch_idx, numel, wg_size, name) removes the repeated dispatch_at(...).workgroup_count_x = plumbing and keeps the device/queue access in one place.

3. The 5-binding gather scaffold shared by slice and select.Select.cpp and Slice.cpp are ~110 lines of identical bind-group-layout / pipeline-layout / pipeline / bind-group setup for the canonical [in, out(rw), out_meta, in_meta, params] kernel — only the *Params struct and WGSL differ. This is the biggest duplication and the one most likely to drift. Consider a builder that returns the created objects:

structGatherPipeline { WGPUComputePipeline pipeline; WGPUBindGroup bind_group;
WGPUBuffer out_meta_buf, in_meta_buf, params_buf; };
GatherPipeline build_gather_op(graph, in_tensor, out_tensor, wgsl,
wg_size, constvoid* params, size_t params_size);

The get_or_create_bgl cache already exists in WebGPUGraph, so the layout could also be keyed/shared rather than rebuilt per op. This one is a larger refactor and out of scope for this PR — but it's the place where per-op boilerplate is accumulating fastest, and select/slice are the natural first consolidation. Start the refactor →

Given CLAUDE.md's "no trivial single-use helpers" rule, #1 and #2 clear that bar because they're now used by 2+ ops and will be reused by every shape-changing op; #3 is a real refactor worth a follow-up PR rather than expanding this one.

Recommendation

Ship the fix as-is (correct and minimal). Suggest doing #1 either here or as the immediate next stacked diff while the pattern is fresh, and tracking #3 as a follow-up.
· branch gh/JulianCloudNTH/70/head

[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
[ghstack-poisoned]

@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review automatically exported from Phabricator review in Meta.

[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:28 — with GitHub Actions Inactive
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 21:20 — with GitHub Actions Inactive
@meta-codesync
meta-codesyncBot merged commit 882b643 into gh/JulianCloudNTH/70/baseJul 4, 2026
180 of 183 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JulianCloudNTH/70/head branch July 4, 2026 17:05
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@SS-JIA@nil-is-all@JCNTH
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

[ExecuTorch][WebGPU] Dynamic resize hooks for sigmoid and select_copy - #20578

Merged
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/70/basefrom
gh/JulianCloudNTH/70/head
Jul 4, 2026
Merged

[ExecuTorch][WebGPU] Dynamic resize hooks for sigmoid and select_copy#20578
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/70/basefrom
gh/JulianCloudNTH/70/head

Conversation

@ghost

@ghostghost commented Jun 28, 2026

Copy link
Copy Markdown

Stack from ghstack (oldest at bottom):

Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.

Problem: Both ops baked their dispatch/params/output shape at build() for the max shape. select_copy was worse: a negative index (e.g. -1 for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.

Solution:

  • sigmoid (generic add_unary_op): a resize hook recomputes num_elements/dispatch and sets the output cur_dims (shape-preserving).
  • select_copy: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus dim), rebuilds the out/in TensorMeta UBOs and the dispatch.
  • Both keep their uniform buffer(s) alive via own_uniform_buffer.

Implementation:

  • The select out/in meta is rebuilt from synthetic WebGPUTensor{dims} via fill_tensor_meta (reads only .dims).
  • Mirrors Vulkan per-op resize_*_node.

Constraints: Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.

Co-authored-with: Claude Code.
@exported-using-ghexport

Differential Revision: D109906095

Differential Revision: D109906095

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jun 28, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20578

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Unrelated Failure

As of commit d658f5c with merge base 124b677 (image):

NEW FAILURE - The following job has failed:

FLAKY - The following job failed but was likely due to flakiness present on trunk:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@ghost

Copy link
Copy Markdown
Author

@claude review and check for any areas or opportunities for modularization

@claude

claudeBot commented Jun 29, 2026

Copy link
Copy Markdown

Claude finished @JulianCloudNTH's task in 2m 47s —— View job


Review: Dynamic resize hooks for sigmoid and select_copy

  • Read changed files & surrounding WebGPU runtime infrastructure
  • Analyze correctness of resize hooks
  • Identify modularization opportunities

The change is well-scoped and matches the established per-op resize-hook pattern (add, mul). The select_copy negative-index fix is correct and is a real bug fix: at build the index was normalized against the max dim, so -1 pointed past live data at smaller S; keeping raw_index and re-resolving against cur_dims(in_id)[dim] in the hook is the right fix. Both ops correctly switch to own_uniform_buffer so the hook can rewrite the UBOs.

Correctness — looks good, two minor notes

  1. Select.cpp — guards not re-checked on resize (acceptable). The build path validates nbytes == numel*4 (fp32) and the bind-group buffer size is pinned to the maxnbytes. The hook only rewrites meta/params/workgroup count, never buffer sizes — correct, since dtype and allocation are static and only live sizes shrink. Worth a one-line comment in the hook noting the bind-group sizes intentionally stay at max so it doesn't read as an omission.

  2. UnaryOp.cpp — build vs. hook compute numel from different tensors. Build uses out_tensor.nbytes/sizeof(float); the hook uses numel_of(cur_dims(in_id)). They're guaranteed equal (the op asserts in.nbytes == out.nbytes), so this is fine, but reading in in one place and out in the other is a small inconsistency. Computing both from the input would read more uniformly.

No issues with the lambda captures (all captured values are used; handles copied before own_uniform_buffer transfers ownership).

Modularization — the more substantive feedback

This PR is behavior-correct, but it lands on top of a growing amount of near-duplicate boilerplate across ops/. Three concrete, increasingly valuable extractions:

1. A meta-UBO rewrite helper (highest value, lowest risk). The "synthesize a WebGPUTensor from dims → fill_tensor_metawgpuQueueWriteBuffer" sequence now appears in select (Select.cpp:194-201) and mul (mul/BinaryOp.cpp:187-197), and is the kind of thing every future shape-changing op will copy. A small helper in TensorMeta.h would absorb it:

inlinevoidwrite_meta_from_dims(WebGPUGraph& g, WGPUBuffer buf,
const std::vector<int64_t>& dims) {
WebGPUTensor t; t.dims = dims;
TensorMeta m; fill_tensor_meta(t, &m);
wgpuQueueWriteBuffer(g.queue(), buf, 0, &m, sizeof(m));
}

Select's hook then drops to write_meta_from_dims(g, out_meta_buf, od); write_meta_from_dims(g, in_meta_buf, ind);. (A _broadcast overload covers mul.)

2. A set_workgroup_1d helper on the resize path. Every resize hook ends with the same g.dispatch_at(idx).workgroup_count_x = utils::compute_1d_workgroup_count(g.device(), numel, wg_size, name); (Select.cpp:206, UnaryOp.cpp:152, add/mul). A thin utils::set_dispatch_1d(g, dispatch_idx, numel, wg_size, name) removes the repeated dispatch_at(...).workgroup_count_x = plumbing and keeps the device/queue access in one place.

3. The 5-binding gather scaffold shared by slice and select.Select.cpp and Slice.cpp are ~110 lines of identical bind-group-layout / pipeline-layout / pipeline / bind-group setup for the canonical [in, out(rw), out_meta, in_meta, params] kernel — only the *Params struct and WGSL differ. This is the biggest duplication and the one most likely to drift. Consider a builder that returns the created objects:

structGatherPipeline { WGPUComputePipeline pipeline; WGPUBindGroup bind_group;
WGPUBuffer out_meta_buf, in_meta_buf, params_buf; };
GatherPipeline build_gather_op(graph, in_tensor, out_tensor, wgsl,
wg_size, constvoid* params, size_t params_size);

The get_or_create_bgl cache already exists in WebGPUGraph, so the layout could also be keyed/shared rather than rebuilt per op. This one is a larger refactor and out of scope for this PR — but it's the place where per-op boilerplate is accumulating fastest, and select/slice are the natural first consolidation. Start the refactor →

Given CLAUDE.md's "no trivial single-use helpers" rule, #1 and #2 clear that bar because they're now used by 2+ ops and will be reused by every shape-changing op; #3 is a real refactor worth a follow-up PR rather than expanding this one.

Recommendation

Ship the fix as-is (correct and minimal). Suggest doing #1 either here or as the immediate next stacked diff while the pattern is fresh, and tracking #3 as a follow-up.
· branch gh/JulianCloudNTH/70/head

[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
[ghstack-poisoned]

@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review automatically exported from Phabricator review in Meta.

[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:28 — with GitHub Actions Inactive
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 21:20 — with GitHub Actions Inactive
@meta-codesync
meta-codesyncBot merged commit 882b643 into gh/JulianCloudNTH/70/baseJul 4, 2026
180 of 183 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JulianCloudNTH/70/head branch July 4, 2026 17:05
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@SS-JIA@nil-is-all@JCNTH
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[ExecuTorch][WebGPU] Dynamic resize hooks for sigmoid and select_copy - #20578

Merged
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/70/basefrom
gh/JulianCloudNTH/70/head
Jul 4, 2026
Merged

[ExecuTorch][WebGPU] Dynamic resize hooks for sigmoid and select_copy#20578
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/70/basefrom
gh/JulianCloudNTH/70/head

Conversation

@ghost

@ghostghost commented Jun 28, 2026

Copy link
Copy Markdown

Stack from ghstack (oldest at bottom):

Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.

Problem: Both ops baked their dispatch/params/output shape at build() for the max shape. select_copy was worse: a negative index (e.g. -1 for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.

Solution:

  • sigmoid (generic add_unary_op): a resize hook recomputes num_elements/dispatch and sets the output cur_dims (shape-preserving).
  • select_copy: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus dim), rebuilds the out/in TensorMeta UBOs and the dispatch.
  • Both keep their uniform buffer(s) alive via own_uniform_buffer.

Implementation:

  • The select out/in meta is rebuilt from synthetic WebGPUTensor{dims} via fill_tensor_meta (reads only .dims).
  • Mirrors Vulkan per-op resize_*_node.

Constraints: Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.

Co-authored-with: Claude Code.
@exported-using-ghexport

Differential Revision: D109906095

Differential Revision: D109906095

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jun 28, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20578

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Unrelated Failure

As of commit d658f5c with merge base 124b677 (image):

NEW FAILURE - The following job has failed:

FLAKY - The following job failed but was likely due to flakiness present on trunk:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@ghost

Copy link
Copy Markdown
Author

@claude review and check for any areas or opportunities for modularization

@claude

claudeBot commented Jun 29, 2026

Copy link
Copy Markdown

Claude finished @JulianCloudNTH's task in 2m 47s —— View job


Review: Dynamic resize hooks for sigmoid and select_copy

  • Read changed files & surrounding WebGPU runtime infrastructure
  • Analyze correctness of resize hooks
  • Identify modularization opportunities

The change is well-scoped and matches the established per-op resize-hook pattern (add, mul). The select_copy negative-index fix is correct and is a real bug fix: at build the index was normalized against the max dim, so -1 pointed past live data at smaller S; keeping raw_index and re-resolving against cur_dims(in_id)[dim] in the hook is the right fix. Both ops correctly switch to own_uniform_buffer so the hook can rewrite the UBOs.

Correctness — looks good, two minor notes

  1. Select.cpp — guards not re-checked on resize (acceptable). The build path validates nbytes == numel*4 (fp32) and the bind-group buffer size is pinned to the maxnbytes. The hook only rewrites meta/params/workgroup count, never buffer sizes — correct, since dtype and allocation are static and only live sizes shrink. Worth a one-line comment in the hook noting the bind-group sizes intentionally stay at max so it doesn't read as an omission.

  2. UnaryOp.cpp — build vs. hook compute numel from different tensors. Build uses out_tensor.nbytes/sizeof(float); the hook uses numel_of(cur_dims(in_id)). They're guaranteed equal (the op asserts in.nbytes == out.nbytes), so this is fine, but reading in in one place and out in the other is a small inconsistency. Computing both from the input would read more uniformly.

No issues with the lambda captures (all captured values are used; handles copied before own_uniform_buffer transfers ownership).

Modularization — the more substantive feedback

This PR is behavior-correct, but it lands on top of a growing amount of near-duplicate boilerplate across ops/. Three concrete, increasingly valuable extractions:

1. A meta-UBO rewrite helper (highest value, lowest risk). The "synthesize a WebGPUTensor from dims → fill_tensor_metawgpuQueueWriteBuffer" sequence now appears in select (Select.cpp:194-201) and mul (mul/BinaryOp.cpp:187-197), and is the kind of thing every future shape-changing op will copy. A small helper in TensorMeta.h would absorb it:

inlinevoidwrite_meta_from_dims(WebGPUGraph& g, WGPUBuffer buf,
const std::vector<int64_t>& dims) {
WebGPUTensor t; t.dims = dims;
TensorMeta m; fill_tensor_meta(t, &m);
wgpuQueueWriteBuffer(g.queue(), buf, 0, &m, sizeof(m));
}

Select's hook then drops to write_meta_from_dims(g, out_meta_buf, od); write_meta_from_dims(g, in_meta_buf, ind);. (A _broadcast overload covers mul.)

2. A set_workgroup_1d helper on the resize path. Every resize hook ends with the same g.dispatch_at(idx).workgroup_count_x = utils::compute_1d_workgroup_count(g.device(), numel, wg_size, name); (Select.cpp:206, UnaryOp.cpp:152, add/mul). A thin utils::set_dispatch_1d(g, dispatch_idx, numel, wg_size, name) removes the repeated dispatch_at(...).workgroup_count_x = plumbing and keeps the device/queue access in one place.

3. The 5-binding gather scaffold shared by slice and select.Select.cpp and Slice.cpp are ~110 lines of identical bind-group-layout / pipeline-layout / pipeline / bind-group setup for the canonical [in, out(rw), out_meta, in_meta, params] kernel — only the *Params struct and WGSL differ. This is the biggest duplication and the one most likely to drift. Consider a builder that returns the created objects:

structGatherPipeline { WGPUComputePipeline pipeline; WGPUBindGroup bind_group;
WGPUBuffer out_meta_buf, in_meta_buf, params_buf; };
GatherPipeline build_gather_op(graph, in_tensor, out_tensor, wgsl,
wg_size, constvoid* params, size_t params_size);

The get_or_create_bgl cache already exists in WebGPUGraph, so the layout could also be keyed/shared rather than rebuilt per op. This one is a larger refactor and out of scope for this PR — but it's the place where per-op boilerplate is accumulating fastest, and select/slice are the natural first consolidation. Start the refactor →

Given CLAUDE.md's "no trivial single-use helpers" rule, #1 and #2 clear that bar because they're now used by 2+ ops and will be reused by every shape-changing op; #3 is a real refactor worth a follow-up PR rather than expanding this one.

Recommendation

Ship the fix as-is (correct and minimal). Suggest doing #1 either here or as the immediate next stacked diff while the pattern is fresh, and tracking #3 as a follow-up.
· branch gh/JulianCloudNTH/70/head

[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
[ghstack-poisoned]

@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review automatically exported from Phabricator review in Meta.

[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:28 — with GitHub Actions Inactive
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 21:20 — with GitHub Actions Inactive
@meta-codesync
meta-codesyncBot merged commit 882b643 into gh/JulianCloudNTH/70/baseJul 4, 2026
180 of 183 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JulianCloudNTH/70/head branch July 4, 2026 17:05
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@SS-JIA@nil-is-all@JCNTH
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[ExecuTorch][WebGPU] Dynamic resize hooks for sigmoid and select_copy - #20578

Merged
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/70/basefrom
gh/JulianCloudNTH/70/head
Jul 4, 2026
Merged

[ExecuTorch][WebGPU] Dynamic resize hooks for sigmoid and select_copy#20578
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/70/basefrom
gh/JulianCloudNTH/70/head

Conversation

@ghost

@ghostghost commented Jun 28, 2026

Copy link
Copy Markdown

Stack from ghstack (oldest at bottom):

Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.

Problem: Both ops baked their dispatch/params/output shape at build() for the max shape. select_copy was worse: a negative index (e.g. -1 for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.

Solution:

  • sigmoid (generic add_unary_op): a resize hook recomputes num_elements/dispatch and sets the output cur_dims (shape-preserving).
  • select_copy: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus dim), rebuilds the out/in TensorMeta UBOs and the dispatch.
  • Both keep their uniform buffer(s) alive via own_uniform_buffer.

Implementation:

  • The select out/in meta is rebuilt from synthetic WebGPUTensor{dims} via fill_tensor_meta (reads only .dims).
  • Mirrors Vulkan per-op resize_*_node.

Constraints: Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.

Co-authored-with: Claude Code.
@exported-using-ghexport

Differential Revision: D109906095

Differential Revision: D109906095

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jun 28, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20578

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Unrelated Failure

As of commit d658f5c with merge base 124b677 (image):

NEW FAILURE - The following job has failed:

FLAKY - The following job failed but was likely due to flakiness present on trunk:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@ghost

Copy link
Copy Markdown
Author

@claude review and check for any areas or opportunities for modularization

@claude

claudeBot commented Jun 29, 2026

Copy link
Copy Markdown

Claude finished @JulianCloudNTH's task in 2m 47s —— View job


Review: Dynamic resize hooks for sigmoid and select_copy

  • Read changed files & surrounding WebGPU runtime infrastructure
  • Analyze correctness of resize hooks
  • Identify modularization opportunities

The change is well-scoped and matches the established per-op resize-hook pattern (add, mul). The select_copy negative-index fix is correct and is a real bug fix: at build the index was normalized against the max dim, so -1 pointed past live data at smaller S; keeping raw_index and re-resolving against cur_dims(in_id)[dim] in the hook is the right fix. Both ops correctly switch to own_uniform_buffer so the hook can rewrite the UBOs.

Correctness — looks good, two minor notes

  1. Select.cpp — guards not re-checked on resize (acceptable). The build path validates nbytes == numel*4 (fp32) and the bind-group buffer size is pinned to the maxnbytes. The hook only rewrites meta/params/workgroup count, never buffer sizes — correct, since dtype and allocation are static and only live sizes shrink. Worth a one-line comment in the hook noting the bind-group sizes intentionally stay at max so it doesn't read as an omission.

  2. UnaryOp.cpp — build vs. hook compute numel from different tensors. Build uses out_tensor.nbytes/sizeof(float); the hook uses numel_of(cur_dims(in_id)). They're guaranteed equal (the op asserts in.nbytes == out.nbytes), so this is fine, but reading in in one place and out in the other is a small inconsistency. Computing both from the input would read more uniformly.

No issues with the lambda captures (all captured values are used; handles copied before own_uniform_buffer transfers ownership).

Modularization — the more substantive feedback

This PR is behavior-correct, but it lands on top of a growing amount of near-duplicate boilerplate across ops/. Three concrete, increasingly valuable extractions:

1. A meta-UBO rewrite helper (highest value, lowest risk). The "synthesize a WebGPUTensor from dims → fill_tensor_metawgpuQueueWriteBuffer" sequence now appears in select (Select.cpp:194-201) and mul (mul/BinaryOp.cpp:187-197), and is the kind of thing every future shape-changing op will copy. A small helper in TensorMeta.h would absorb it:

inlinevoidwrite_meta_from_dims(WebGPUGraph& g, WGPUBuffer buf,
const std::vector<int64_t>& dims) {
WebGPUTensor t; t.dims = dims;
TensorMeta m; fill_tensor_meta(t, &m);
wgpuQueueWriteBuffer(g.queue(), buf, 0, &m, sizeof(m));
}

Select's hook then drops to write_meta_from_dims(g, out_meta_buf, od); write_meta_from_dims(g, in_meta_buf, ind);. (A _broadcast overload covers mul.)

2. A set_workgroup_1d helper on the resize path. Every resize hook ends with the same g.dispatch_at(idx).workgroup_count_x = utils::compute_1d_workgroup_count(g.device(), numel, wg_size, name); (Select.cpp:206, UnaryOp.cpp:152, add/mul). A thin utils::set_dispatch_1d(g, dispatch_idx, numel, wg_size, name) removes the repeated dispatch_at(...).workgroup_count_x = plumbing and keeps the device/queue access in one place.

3. The 5-binding gather scaffold shared by slice and select.Select.cpp and Slice.cpp are ~110 lines of identical bind-group-layout / pipeline-layout / pipeline / bind-group setup for the canonical [in, out(rw), out_meta, in_meta, params] kernel — only the *Params struct and WGSL differ. This is the biggest duplication and the one most likely to drift. Consider a builder that returns the created objects:

structGatherPipeline { WGPUComputePipeline pipeline; WGPUBindGroup bind_group;
WGPUBuffer out_meta_buf, in_meta_buf, params_buf; };
GatherPipeline build_gather_op(graph, in_tensor, out_tensor, wgsl,
wg_size, constvoid* params, size_t params_size);

The get_or_create_bgl cache already exists in WebGPUGraph, so the layout could also be keyed/shared rather than rebuilt per op. This one is a larger refactor and out of scope for this PR — but it's the place where per-op boilerplate is accumulating fastest, and select/slice are the natural first consolidation. Start the refactor →

Given CLAUDE.md's "no trivial single-use helpers" rule, #1 and #2 clear that bar because they're now used by 2+ ops and will be reused by every shape-changing op; #3 is a real refactor worth a follow-up PR rather than expanding this one.

Recommendation

Ship the fix as-is (correct and minimal). Suggest doing #1 either here or as the immediate next stacked diff while the pattern is fresh, and tracking #3 as a follow-up.
· branch gh/JulianCloudNTH/70/head

[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
[ghstack-poisoned]

@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review automatically exported from Phabricator review in Meta.

[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:28 — with GitHub Actions Inactive
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 21:20 — with GitHub Actions Inactive
@meta-codesync
meta-codesyncBot merged commit 882b643 into gh/JulianCloudNTH/70/baseJul 4, 2026
180 of 183 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JulianCloudNTH/70/head branch July 4, 2026 17:05
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@SS-JIA@nil-is-all@JCNTH
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

[ExecuTorch][WebGPU] Dynamic resize hooks for sigmoid and select_copy - #20578

Merged
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/70/basefrom
gh/JulianCloudNTH/70/head
Jul 4, 2026
Merged

[ExecuTorch][WebGPU] Dynamic resize hooks for sigmoid and select_copy#20578
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/70/basefrom
gh/JulianCloudNTH/70/head

Conversation

@ghost

@ghostghost commented Jun 28, 2026

Copy link
Copy Markdown

Stack from ghstack (oldest at bottom):

Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.

Problem: Both ops baked their dispatch/params/output shape at build() for the max shape. select_copy was worse: a negative index (e.g. -1 for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.

Solution:

  • sigmoid (generic add_unary_op): a resize hook recomputes num_elements/dispatch and sets the output cur_dims (shape-preserving).
  • select_copy: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus dim), rebuilds the out/in TensorMeta UBOs and the dispatch.
  • Both keep their uniform buffer(s) alive via own_uniform_buffer.

Implementation:

  • The select out/in meta is rebuilt from synthetic WebGPUTensor{dims} via fill_tensor_meta (reads only .dims).
  • Mirrors Vulkan per-op resize_*_node.

Constraints: Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.

Co-authored-with: Claude Code.
@exported-using-ghexport

Differential Revision: D109906095

Differential Revision: D109906095

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jun 28, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20578

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Unrelated Failure

As of commit d658f5c with merge base 124b677 (image):

NEW FAILURE - The following job has failed:

FLAKY - The following job failed but was likely due to flakiness present on trunk:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@ghost

Copy link
Copy Markdown
Author

@claude review and check for any areas or opportunities for modularization

@claude

claudeBot commented Jun 29, 2026

Copy link
Copy Markdown

Claude finished @JulianCloudNTH's task in 2m 47s —— View job


Review: Dynamic resize hooks for sigmoid and select_copy

  • Read changed files & surrounding WebGPU runtime infrastructure
  • Analyze correctness of resize hooks
  • Identify modularization opportunities

The change is well-scoped and matches the established per-op resize-hook pattern (add, mul). The select_copy negative-index fix is correct and is a real bug fix: at build the index was normalized against the max dim, so -1 pointed past live data at smaller S; keeping raw_index and re-resolving against cur_dims(in_id)[dim] in the hook is the right fix. Both ops correctly switch to own_uniform_buffer so the hook can rewrite the UBOs.

Correctness — looks good, two minor notes

  1. Select.cpp — guards not re-checked on resize (acceptable). The build path validates nbytes == numel*4 (fp32) and the bind-group buffer size is pinned to the maxnbytes. The hook only rewrites meta/params/workgroup count, never buffer sizes — correct, since dtype and allocation are static and only live sizes shrink. Worth a one-line comment in the hook noting the bind-group sizes intentionally stay at max so it doesn't read as an omission.

  2. UnaryOp.cpp — build vs. hook compute numel from different tensors. Build uses out_tensor.nbytes/sizeof(float); the hook uses numel_of(cur_dims(in_id)). They're guaranteed equal (the op asserts in.nbytes == out.nbytes), so this is fine, but reading in in one place and out in the other is a small inconsistency. Computing both from the input would read more uniformly.

No issues with the lambda captures (all captured values are used; handles copied before own_uniform_buffer transfers ownership).

Modularization — the more substantive feedback

This PR is behavior-correct, but it lands on top of a growing amount of near-duplicate boilerplate across ops/. Three concrete, increasingly valuable extractions:

1. A meta-UBO rewrite helper (highest value, lowest risk). The "synthesize a WebGPUTensor from dims → fill_tensor_metawgpuQueueWriteBuffer" sequence now appears in select (Select.cpp:194-201) and mul (mul/BinaryOp.cpp:187-197), and is the kind of thing every future shape-changing op will copy. A small helper in TensorMeta.h would absorb it:

inlinevoidwrite_meta_from_dims(WebGPUGraph& g, WGPUBuffer buf,
const std::vector<int64_t>& dims) {
WebGPUTensor t; t.dims = dims;
TensorMeta m; fill_tensor_meta(t, &m);
wgpuQueueWriteBuffer(g.queue(), buf, 0, &m, sizeof(m));
}

Select's hook then drops to write_meta_from_dims(g, out_meta_buf, od); write_meta_from_dims(g, in_meta_buf, ind);. (A _broadcast overload covers mul.)

2. A set_workgroup_1d helper on the resize path. Every resize hook ends with the same g.dispatch_at(idx).workgroup_count_x = utils::compute_1d_workgroup_count(g.device(), numel, wg_size, name); (Select.cpp:206, UnaryOp.cpp:152, add/mul). A thin utils::set_dispatch_1d(g, dispatch_idx, numel, wg_size, name) removes the repeated dispatch_at(...).workgroup_count_x = plumbing and keeps the device/queue access in one place.

3. The 5-binding gather scaffold shared by slice and select.Select.cpp and Slice.cpp are ~110 lines of identical bind-group-layout / pipeline-layout / pipeline / bind-group setup for the canonical [in, out(rw), out_meta, in_meta, params] kernel — only the *Params struct and WGSL differ. This is the biggest duplication and the one most likely to drift. Consider a builder that returns the created objects:

structGatherPipeline { WGPUComputePipeline pipeline; WGPUBindGroup bind_group;
WGPUBuffer out_meta_buf, in_meta_buf, params_buf; };
GatherPipeline build_gather_op(graph, in_tensor, out_tensor, wgsl,
wg_size, constvoid* params, size_t params_size);

The get_or_create_bgl cache already exists in WebGPUGraph, so the layout could also be keyed/shared rather than rebuilt per op. This one is a larger refactor and out of scope for this PR — but it's the place where per-op boilerplate is accumulating fastest, and select/slice are the natural first consolidation. Start the refactor →

Given CLAUDE.md's "no trivial single-use helpers" rule, #1 and #2 clear that bar because they're now used by 2+ ops and will be reused by every shape-changing op; #3 is a real refactor worth a follow-up PR rather than expanding this one.

Recommendation

Ship the fix as-is (correct and minimal). Suggest doing #1 either here or as the immediate next stacked diff while the pattern is fresh, and tracking #3 as a follow-up.
· branch gh/JulianCloudNTH/70/head

[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
[ghstack-poisoned]

@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review automatically exported from Phabricator review in Meta.

[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:28 — with GitHub Actions Inactive
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 21:20 — with GitHub Actions Inactive
@meta-codesync
meta-codesyncBot merged commit 882b643 into gh/JulianCloudNTH/70/baseJul 4, 2026
180 of 183 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JulianCloudNTH/70/head branch July 4, 2026 17:05
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@SS-JIA@nil-is-all@JCNTH
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[ExecuTorch][WebGPU] Dynamic resize hooks for sigmoid and select_copy - #20578

Merged
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/70/basefrom
gh/JulianCloudNTH/70/head
Jul 4, 2026
Merged

[ExecuTorch][WebGPU] Dynamic resize hooks for sigmoid and select_copy#20578
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/70/basefrom
gh/JulianCloudNTH/70/head

Conversation

@ghost

@ghostghost commented Jun 28, 2026

Copy link
Copy Markdown

Stack from ghstack (oldest at bottom):

Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.

Problem: Both ops baked their dispatch/params/output shape at build() for the max shape. select_copy was worse: a negative index (e.g. -1 for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.

Solution:

  • sigmoid (generic add_unary_op): a resize hook recomputes num_elements/dispatch and sets the output cur_dims (shape-preserving).
  • select_copy: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus dim), rebuilds the out/in TensorMeta UBOs and the dispatch.
  • Both keep their uniform buffer(s) alive via own_uniform_buffer.

Implementation:

  • The select out/in meta is rebuilt from synthetic WebGPUTensor{dims} via fill_tensor_meta (reads only .dims).
  • Mirrors Vulkan per-op resize_*_node.

Constraints: Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.

Co-authored-with: Claude Code.
@exported-using-ghexport

Differential Revision: D109906095

Differential Revision: D109906095

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jun 28, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20578

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Unrelated Failure

As of commit d658f5c with merge base 124b677 (image):

NEW FAILURE - The following job has failed:

FLAKY - The following job failed but was likely due to flakiness present on trunk:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@ghost

Copy link
Copy Markdown
Author

@claude review and check for any areas or opportunities for modularization

@claude

claudeBot commented Jun 29, 2026

Copy link
Copy Markdown

Claude finished @JulianCloudNTH's task in 2m 47s —— View job


Review: Dynamic resize hooks for sigmoid and select_copy

  • Read changed files & surrounding WebGPU runtime infrastructure
  • Analyze correctness of resize hooks
  • Identify modularization opportunities

The change is well-scoped and matches the established per-op resize-hook pattern (add, mul). The select_copy negative-index fix is correct and is a real bug fix: at build the index was normalized against the max dim, so -1 pointed past live data at smaller S; keeping raw_index and re-resolving against cur_dims(in_id)[dim] in the hook is the right fix. Both ops correctly switch to own_uniform_buffer so the hook can rewrite the UBOs.

Correctness — looks good, two minor notes

  1. Select.cpp — guards not re-checked on resize (acceptable). The build path validates nbytes == numel*4 (fp32) and the bind-group buffer size is pinned to the maxnbytes. The hook only rewrites meta/params/workgroup count, never buffer sizes — correct, since dtype and allocation are static and only live sizes shrink. Worth a one-line comment in the hook noting the bind-group sizes intentionally stay at max so it doesn't read as an omission.

  2. UnaryOp.cpp — build vs. hook compute numel from different tensors. Build uses out_tensor.nbytes/sizeof(float); the hook uses numel_of(cur_dims(in_id)). They're guaranteed equal (the op asserts in.nbytes == out.nbytes), so this is fine, but reading in in one place and out in the other is a small inconsistency. Computing both from the input would read more uniformly.

No issues with the lambda captures (all captured values are used; handles copied before own_uniform_buffer transfers ownership).

Modularization — the more substantive feedback

This PR is behavior-correct, but it lands on top of a growing amount of near-duplicate boilerplate across ops/. Three concrete, increasingly valuable extractions:

1. A meta-UBO rewrite helper (highest value, lowest risk). The "synthesize a WebGPUTensor from dims → fill_tensor_metawgpuQueueWriteBuffer" sequence now appears in select (Select.cpp:194-201) and mul (mul/BinaryOp.cpp:187-197), and is the kind of thing every future shape-changing op will copy. A small helper in TensorMeta.h would absorb it:

inlinevoidwrite_meta_from_dims(WebGPUGraph& g, WGPUBuffer buf,
const std::vector<int64_t>& dims) {
WebGPUTensor t; t.dims = dims;
TensorMeta m; fill_tensor_meta(t, &m);
wgpuQueueWriteBuffer(g.queue(), buf, 0, &m, sizeof(m));
}

Select's hook then drops to write_meta_from_dims(g, out_meta_buf, od); write_meta_from_dims(g, in_meta_buf, ind);. (A _broadcast overload covers mul.)

2. A set_workgroup_1d helper on the resize path. Every resize hook ends with the same g.dispatch_at(idx).workgroup_count_x = utils::compute_1d_workgroup_count(g.device(), numel, wg_size, name); (Select.cpp:206, UnaryOp.cpp:152, add/mul). A thin utils::set_dispatch_1d(g, dispatch_idx, numel, wg_size, name) removes the repeated dispatch_at(...).workgroup_count_x = plumbing and keeps the device/queue access in one place.

3. The 5-binding gather scaffold shared by slice and select.Select.cpp and Slice.cpp are ~110 lines of identical bind-group-layout / pipeline-layout / pipeline / bind-group setup for the canonical [in, out(rw), out_meta, in_meta, params] kernel — only the *Params struct and WGSL differ. This is the biggest duplication and the one most likely to drift. Consider a builder that returns the created objects:

structGatherPipeline { WGPUComputePipeline pipeline; WGPUBindGroup bind_group;
WGPUBuffer out_meta_buf, in_meta_buf, params_buf; };
GatherPipeline build_gather_op(graph, in_tensor, out_tensor, wgsl,
wg_size, constvoid* params, size_t params_size);

The get_or_create_bgl cache already exists in WebGPUGraph, so the layout could also be keyed/shared rather than rebuilt per op. This one is a larger refactor and out of scope for this PR — but it's the place where per-op boilerplate is accumulating fastest, and select/slice are the natural first consolidation. Start the refactor →

Given CLAUDE.md's "no trivial single-use helpers" rule, #1 and #2 clear that bar because they're now used by 2+ ops and will be reused by every shape-changing op; #3 is a real refactor worth a follow-up PR rather than expanding this one.

Recommendation

Ship the fix as-is (correct and minimal). Suggest doing #1 either here or as the immediate next stacked diff while the pattern is fresh, and tracking #3 as a follow-up.
· branch gh/JulianCloudNTH/70/head

[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
[ghstack-poisoned]

@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review automatically exported from Phabricator review in Meta.

[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:28 — with GitHub Actions Inactive
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 21:20 — with GitHub Actions Inactive
@meta-codesync
meta-codesyncBot merged commit 882b643 into gh/JulianCloudNTH/70/baseJul 4, 2026
180 of 183 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JulianCloudNTH/70/head branch July 4, 2026 17:05
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@SS-JIA@nil-is-all@JCNTH
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[ExecuTorch][WebGPU] Dynamic resize hooks for sigmoid and select_copy - #20578

Merged
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/70/basefrom
gh/JulianCloudNTH/70/head
Jul 4, 2026
Merged

[ExecuTorch][WebGPU] Dynamic resize hooks for sigmoid and select_copy#20578
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/70/basefrom
gh/JulianCloudNTH/70/head

Conversation

@ghost

@ghostghost commented Jun 28, 2026

Copy link
Copy Markdown

Stack from ghstack (oldest at bottom):

Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.

Problem: Both ops baked their dispatch/params/output shape at build() for the max shape. select_copy was worse: a negative index (e.g. -1 for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.

Solution:

  • sigmoid (generic add_unary_op): a resize hook recomputes num_elements/dispatch and sets the output cur_dims (shape-preserving).
  • select_copy: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus dim), rebuilds the out/in TensorMeta UBOs and the dispatch.
  • Both keep their uniform buffer(s) alive via own_uniform_buffer.

Implementation:

  • The select out/in meta is rebuilt from synthetic WebGPUTensor{dims} via fill_tensor_meta (reads only .dims).
  • Mirrors Vulkan per-op resize_*_node.

Constraints: Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.

Co-authored-with: Claude Code.
@exported-using-ghexport

Differential Revision: D109906095

Differential Revision: D109906095

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jun 28, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20578

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Unrelated Failure

As of commit d658f5c with merge base 124b677 (image):

NEW FAILURE - The following job has failed:

FLAKY - The following job failed but was likely due to flakiness present on trunk:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@ghost

Copy link
Copy Markdown
Author

@claude review and check for any areas or opportunities for modularization

@claude

claudeBot commented Jun 29, 2026

Copy link
Copy Markdown

Claude finished @JulianCloudNTH's task in 2m 47s —— View job


Review: Dynamic resize hooks for sigmoid and select_copy

  • Read changed files & surrounding WebGPU runtime infrastructure
  • Analyze correctness of resize hooks
  • Identify modularization opportunities

The change is well-scoped and matches the established per-op resize-hook pattern (add, mul). The select_copy negative-index fix is correct and is a real bug fix: at build the index was normalized against the max dim, so -1 pointed past live data at smaller S; keeping raw_index and re-resolving against cur_dims(in_id)[dim] in the hook is the right fix. Both ops correctly switch to own_uniform_buffer so the hook can rewrite the UBOs.

Correctness — looks good, two minor notes

  1. Select.cpp — guards not re-checked on resize (acceptable). The build path validates nbytes == numel*4 (fp32) and the bind-group buffer size is pinned to the maxnbytes. The hook only rewrites meta/params/workgroup count, never buffer sizes — correct, since dtype and allocation are static and only live sizes shrink. Worth a one-line comment in the hook noting the bind-group sizes intentionally stay at max so it doesn't read as an omission.

  2. UnaryOp.cpp — build vs. hook compute numel from different tensors. Build uses out_tensor.nbytes/sizeof(float); the hook uses numel_of(cur_dims(in_id)). They're guaranteed equal (the op asserts in.nbytes == out.nbytes), so this is fine, but reading in in one place and out in the other is a small inconsistency. Computing both from the input would read more uniformly.

No issues with the lambda captures (all captured values are used; handles copied before own_uniform_buffer transfers ownership).

Modularization — the more substantive feedback

This PR is behavior-correct, but it lands on top of a growing amount of near-duplicate boilerplate across ops/. Three concrete, increasingly valuable extractions:

1. A meta-UBO rewrite helper (highest value, lowest risk). The "synthesize a WebGPUTensor from dims → fill_tensor_metawgpuQueueWriteBuffer" sequence now appears in select (Select.cpp:194-201) and mul (mul/BinaryOp.cpp:187-197), and is the kind of thing every future shape-changing op will copy. A small helper in TensorMeta.h would absorb it:

inlinevoidwrite_meta_from_dims(WebGPUGraph& g, WGPUBuffer buf,
const std::vector<int64_t>& dims) {
WebGPUTensor t; t.dims = dims;
TensorMeta m; fill_tensor_meta(t, &m);
wgpuQueueWriteBuffer(g.queue(), buf, 0, &m, sizeof(m));
}

Select's hook then drops to write_meta_from_dims(g, out_meta_buf, od); write_meta_from_dims(g, in_meta_buf, ind);. (A _broadcast overload covers mul.)

2. A set_workgroup_1d helper on the resize path. Every resize hook ends with the same g.dispatch_at(idx).workgroup_count_x = utils::compute_1d_workgroup_count(g.device(), numel, wg_size, name); (Select.cpp:206, UnaryOp.cpp:152, add/mul). A thin utils::set_dispatch_1d(g, dispatch_idx, numel, wg_size, name) removes the repeated dispatch_at(...).workgroup_count_x = plumbing and keeps the device/queue access in one place.

3. The 5-binding gather scaffold shared by slice and select.Select.cpp and Slice.cpp are ~110 lines of identical bind-group-layout / pipeline-layout / pipeline / bind-group setup for the canonical [in, out(rw), out_meta, in_meta, params] kernel — only the *Params struct and WGSL differ. This is the biggest duplication and the one most likely to drift. Consider a builder that returns the created objects:

structGatherPipeline { WGPUComputePipeline pipeline; WGPUBindGroup bind_group;
WGPUBuffer out_meta_buf, in_meta_buf, params_buf; };
GatherPipeline build_gather_op(graph, in_tensor, out_tensor, wgsl,
wg_size, constvoid* params, size_t params_size);

The get_or_create_bgl cache already exists in WebGPUGraph, so the layout could also be keyed/shared rather than rebuilt per op. This one is a larger refactor and out of scope for this PR — but it's the place where per-op boilerplate is accumulating fastest, and select/slice are the natural first consolidation. Start the refactor →

Given CLAUDE.md's "no trivial single-use helpers" rule, #1 and #2 clear that bar because they're now used by 2+ ops and will be reused by every shape-changing op; #3 is a real refactor worth a follow-up PR rather than expanding this one.

Recommendation

Ship the fix as-is (correct and minimal). Suggest doing #1 either here or as the immediate next stacked diff while the pattern is fresh, and tracking #3 as a follow-up.
· branch gh/JulianCloudNTH/70/head

[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
[ghstack-poisoned]

@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review automatically exported from Phabricator review in Meta.

[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:28 — with GitHub Actions Inactive
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 21:20 — with GitHub Actions Inactive
@meta-codesync
meta-codesyncBot merged commit 882b643 into gh/JulianCloudNTH/70/baseJul 4, 2026
180 of 183 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JulianCloudNTH/70/head branch July 4, 2026 17:05
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@SS-JIA@nil-is-all@JCNTH
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

[ExecuTorch][WebGPU] Dynamic resize hooks for sigmoid and select_copy - #20578

Merged
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/70/basefrom
gh/JulianCloudNTH/70/head
Jul 4, 2026
Merged

[ExecuTorch][WebGPU] Dynamic resize hooks for sigmoid and select_copy#20578
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/70/basefrom
gh/JulianCloudNTH/70/head

Conversation

@ghost

@ghostghost commented Jun 28, 2026

Copy link
Copy Markdown

Stack from ghstack (oldest at bottom):

Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.

Problem: Both ops baked their dispatch/params/output shape at build() for the max shape. select_copy was worse: a negative index (e.g. -1 for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.

Solution:

  • sigmoid (generic add_unary_op): a resize hook recomputes num_elements/dispatch and sets the output cur_dims (shape-preserving).
  • select_copy: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus dim), rebuilds the out/in TensorMeta UBOs and the dispatch.
  • Both keep their uniform buffer(s) alive via own_uniform_buffer.

Implementation:

  • The select out/in meta is rebuilt from synthetic WebGPUTensor{dims} via fill_tensor_meta (reads only .dims).
  • Mirrors Vulkan per-op resize_*_node.

Constraints: Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.

Co-authored-with: Claude Code.
@exported-using-ghexport

Differential Revision: D109906095

Differential Revision: D109906095

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jun 28, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20578

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Unrelated Failure

As of commit d658f5c with merge base 124b677 (image):

NEW FAILURE - The following job has failed:

FLAKY - The following job failed but was likely due to flakiness present on trunk:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@ghost

Copy link
Copy Markdown
Author

@claude review and check for any areas or opportunities for modularization

@claude

claudeBot commented Jun 29, 2026

Copy link
Copy Markdown

Claude finished @JulianCloudNTH's task in 2m 47s —— View job


Review: Dynamic resize hooks for sigmoid and select_copy

  • Read changed files & surrounding WebGPU runtime infrastructure
  • Analyze correctness of resize hooks
  • Identify modularization opportunities

The change is well-scoped and matches the established per-op resize-hook pattern (add, mul). The select_copy negative-index fix is correct and is a real bug fix: at build the index was normalized against the max dim, so -1 pointed past live data at smaller S; keeping raw_index and re-resolving against cur_dims(in_id)[dim] in the hook is the right fix. Both ops correctly switch to own_uniform_buffer so the hook can rewrite the UBOs.

Correctness — looks good, two minor notes

  1. Select.cpp — guards not re-checked on resize (acceptable). The build path validates nbytes == numel*4 (fp32) and the bind-group buffer size is pinned to the maxnbytes. The hook only rewrites meta/params/workgroup count, never buffer sizes — correct, since dtype and allocation are static and only live sizes shrink. Worth a one-line comment in the hook noting the bind-group sizes intentionally stay at max so it doesn't read as an omission.

  2. UnaryOp.cpp — build vs. hook compute numel from different tensors. Build uses out_tensor.nbytes/sizeof(float); the hook uses numel_of(cur_dims(in_id)). They're guaranteed equal (the op asserts in.nbytes == out.nbytes), so this is fine, but reading in in one place and out in the other is a small inconsistency. Computing both from the input would read more uniformly.

No issues with the lambda captures (all captured values are used; handles copied before own_uniform_buffer transfers ownership).

Modularization — the more substantive feedback

This PR is behavior-correct, but it lands on top of a growing amount of near-duplicate boilerplate across ops/. Three concrete, increasingly valuable extractions:

1. A meta-UBO rewrite helper (highest value, lowest risk). The "synthesize a WebGPUTensor from dims → fill_tensor_metawgpuQueueWriteBuffer" sequence now appears in select (Select.cpp:194-201) and mul (mul/BinaryOp.cpp:187-197), and is the kind of thing every future shape-changing op will copy. A small helper in TensorMeta.h would absorb it:

inlinevoidwrite_meta_from_dims(WebGPUGraph& g, WGPUBuffer buf,
const std::vector<int64_t>& dims) {
WebGPUTensor t; t.dims = dims;
TensorMeta m; fill_tensor_meta(t, &m);
wgpuQueueWriteBuffer(g.queue(), buf, 0, &m, sizeof(m));
}

Select's hook then drops to write_meta_from_dims(g, out_meta_buf, od); write_meta_from_dims(g, in_meta_buf, ind);. (A _broadcast overload covers mul.)

2. A set_workgroup_1d helper on the resize path. Every resize hook ends with the same g.dispatch_at(idx).workgroup_count_x = utils::compute_1d_workgroup_count(g.device(), numel, wg_size, name); (Select.cpp:206, UnaryOp.cpp:152, add/mul). A thin utils::set_dispatch_1d(g, dispatch_idx, numel, wg_size, name) removes the repeated dispatch_at(...).workgroup_count_x = plumbing and keeps the device/queue access in one place.

3. The 5-binding gather scaffold shared by slice and select.Select.cpp and Slice.cpp are ~110 lines of identical bind-group-layout / pipeline-layout / pipeline / bind-group setup for the canonical [in, out(rw), out_meta, in_meta, params] kernel — only the *Params struct and WGSL differ. This is the biggest duplication and the one most likely to drift. Consider a builder that returns the created objects:

structGatherPipeline { WGPUComputePipeline pipeline; WGPUBindGroup bind_group;
WGPUBuffer out_meta_buf, in_meta_buf, params_buf; };
GatherPipeline build_gather_op(graph, in_tensor, out_tensor, wgsl,
wg_size, constvoid* params, size_t params_size);

The get_or_create_bgl cache already exists in WebGPUGraph, so the layout could also be keyed/shared rather than rebuilt per op. This one is a larger refactor and out of scope for this PR — but it's the place where per-op boilerplate is accumulating fastest, and select/slice are the natural first consolidation. Start the refactor →

Given CLAUDE.md's "no trivial single-use helpers" rule, #1 and #2 clear that bar because they're now used by 2+ ops and will be reused by every shape-changing op; #3 is a real refactor worth a follow-up PR rather than expanding this one.

Recommendation

Ship the fix as-is (correct and minimal). Suggest doing #1 either here or as the immediate next stacked diff while the pattern is fresh, and tracking #3 as a follow-up.
· branch gh/JulianCloudNTH/70/head

[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
[ghstack-poisoned]

@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review automatically exported from Phabricator review in Meta.

[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:28 — with GitHub Actions Inactive
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 21:20 — with GitHub Actions Inactive
@meta-codesync
meta-codesyncBot merged commit 882b643 into gh/JulianCloudNTH/70/baseJul 4, 2026
180 of 183 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JulianCloudNTH/70/head branch July 4, 2026 17:05
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
Pull Request resolved: #20578
**Make sigmoid and select_copy serve any live shape from one graph; fix select's last-token index under dynamic shapes.**
**Problem:** Both ops baked their dispatch/params/output shape at `build()` for the max shape. `select_copy` was worse: a negative index (e.g. `-1` for the last token) was normalized against the build-time MAX dim, so at a smaller live S it selected a stale/zero position past the live data — producing wrong (often zero) output.
**Solution:**
- `sigmoid` (generic `add_unary_op`): a resize hook recomputes `num_elements`/dispatch and sets the output `cur_dims` (shape-preserving).
- `select_copy`: KEEP the raw (possibly negative) index at build; a resize hook re-resolves it against the LIVE dim, recomputes the output dims (= input minus `dim`), rebuilds the out/in `TensorMeta` UBOs and the dispatch.
- Both keep their uniform buffer(s) alive via `own_uniform_buffer`.
**Implementation:**
- The select out/in meta is rebuilt from synthetic `WebGPUTensor{dims}` via `fill_tensor_meta` (reads only `.dims`).
- Mirrors Vulkan per-op `resize_*_node`.
**Constraints:** Behavior-neutral on static graphs (hooks fire only when an input's live shape differs from the max). No kernel/WGSL/numerics change.
Co-authored-with: Claude Code.
ghstack-source-id: 399812832
@exported-using-ghexport
Differential Revision: [D109906095](https://our.internmc.facebook.com/intern/diff/D109906095/)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@SS-JIA@nil-is-all@JCNTH