[ExecuTorch][WebGPU] Dynamic resize hook for slice_copy (dynamic start/end) - #20581

Merged
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/73/basefrom
gh/JulianCloudNTH/73/head
Jul 4, 2026
Merged

[ExecuTorch][WebGPU] Dynamic resize hook for slice_copy (dynamic start/end)#20581
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/73/basefrom
gh/JulianCloudNTH/73/head

Conversation

@ghost

@ghostghost commented Jun 28, 2026

Copy link
Copy Markdown

Stack from ghstack (oldest at bottom):

Make slice_copy support a dynamic gather range so the RoPE-freqs slice [input_pos : input_pos + S] works under one dynamic graph.

Problem: the static slice handler read start via a scalar reader that throws on a SymInt and ignored end (output length baked AOT). The RoPE-freqs slice uses a SymInt input_pos for start and a live S for the range, so the static op could neither build nor resize for it.

Solution: read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.

  • Before: start is a static scalar (SymInt throws); end ignored; output length fixed at the serialized max.
  • After: start/end read via a SymInt-aware reader; a hook recomputes out[dim] = (end - start + step - 1) / step, rewrites out_meta/in_meta/params UBOs + the dispatch count, and sets the output cur_dims.

Implementation:

  • Hook registered on the start/end value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
  • Output/input TensorMeta rebuilt from live dims; dim/step stay static.
  • Keep the uniforms alive via own_uniform_buffer so the hook can rewrite them.
  • Mirrors Vulkan resize_slice_copy_node.

Constraints: fp32-only; dim/step static; numerics + layout unchanged; inert on a static graph.

NOTE (stacking): this diff sits on top of the in-review slice_copy op (D108793168); rebase onto it once that op lands on master.

Co-authored-with: Claude Code.
@exported-using-ghexport

Differential Revision: D109906092

Differential Revision: D109906092

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jun 28, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20581

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit 11a2c7b with merge base 124b677 (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 28, 2026
@ghost

Copy link
Copy Markdown
Author

@claude review and check for any areas or opportunities for modularization

@claude

claudeBot commented Jun 29, 2026

Copy link
Copy Markdown

Claude finished @JulianCloudNTH's task in 2m 32s —— View job


Code Review: Dynamic resize hook for slice_copy

  • Read PR diff and full Slice.cpp
  • Check WebGPUGraph resize API (read_symint, add_resize_hook, add_tensor_resize_hook, own_uniform_buffer, propagate_resize)
  • Compare against Vulkan resize_slice_copy_node
  • Identify correctness issues
  • Identify modularization opportunities

Overall this is clean and matches the established WebGPU dynamic-shape pattern well (SymInt hooks → tensor hooks, idempotent recompute, own_uniform_buffer to keep UBOs writable). The hook is correctly inert on a static graph. A few items below — one I'd consider a real (latent) correctness bug, the rest are quality/modularization.


1. Output non-slice dims should be derived from live input dims (divergence from Vulkan)

Slice.cpp:224-226

std::vector<int64_t> od = g.cur_dims(out_id); // <- seeds from output's *prior* dims
od[dim] = len;
g.set_cur_dims(out_id, od);

The Vulkan reference this claims to mirror derives the output shape from the input sizes each resize:

std::vector<int64_t> new_out_sizes = in_sizes; // backends/vulkan/.../Slice.cpp:63-64
new_out_sizes.at(dim) = (end - start + step - 1) / step;

For slice_copy, every non-dim axis of the output is by definition equal to the input's. By seeding od from the output's prior cur_dims and only rewriting od[dim], this hook leaves the other axes stale. The in_meta it writes does track in_dims, so if any non-slice axis of the input is also dynamic, out_meta and in_meta disagree on those axes — a silent shape mismatch.

For the RoPE-freqs case only dim is dynamic, so it happens to work today, but it's fragile and diverges from the reference. Suggest mirroring Vulkan:

std::vector<int64_t> od = in_dims; // already have `const auto& in_dims = g.cur_dims(in_id);`
od[dim] = len;

Fix this →


2. Redundant hook executions (minor; idempotent so not a bug)

The same recompute closure is registered up to three times — on start_id, end_id, and in_id. When more than one trigger is dirty in the same step (e.g. input_pos SymInt + an input resize), recompute runs 2–3× per propagate_resize. It's fully idempotent (recomputes from live state), so output is correct, just wasted wgpuQueueWriteBuffer calls. Worth a one-line comment noting the intentional redundancy, or skip if the duplicate writes are negligible.


3. Modularization opportunities (the explicit ask)

a. Extract shared slice geometry — highest value. The build path (Slice.cpp:100-116) and the hook (Slice.cpp:217-242) independently compute start/clamp and fill SliceParams (dim/start/step). A small helper would remove the duplication and make the build path and resize path provably consistent:

structSliceGeom { std::vector<int64_t> out_dims; SliceParams params; };
SliceGeom compute_slice_geom(const std::vector<int64_t>& in_dims,
int64_t dim, int64_t start, int64_t end, int64_t step);

Both sites call it; the hook additionally writes the UBOs/dispatch. This also naturally fixes finding #1, since the helper would build out_dims from in_dims.

b. read_index / is_symint / norm_clamp are reusable across transfer ops.slice, and soon select/narrow/index_select, all need "read a possibly-dynamic index" + "normalize+clamp into [0, size]". Vulkan already centralizes this as normalize_idx. Consider hoisting these three free functions into a shared WebGPU helper header (e.g. alongside TensorMeta.h or a new transfer helper) rather than re-declaring them per op. Even within this file, read_index and read_scalar (Slice.cpp:34-61) share the Int → INT64_MAX→default branch — read_scalar could be expressed as the strict variant of read_index.

c. Bind-group / pipeline boilerplate (out of scope here).Slice.cpp:122-200 is ~80 lines of BGL + pipeline + bind-group setup that is near-identical to other WebGPU ops. Not for this PR, but a make_compute_pipeline(...) / make_bind_group(...) helper would shrink every op file substantially. Flagging as a follow-up direction.


Minor

  • Slice.cpp:76 comment still reads end unread (out shape is AOT) — now stale, end_id is read in the hook. Update or drop.
  • len == 0 path → om.numel == 0compute_1d_workgroup_count(..., 0, ...). Confirm a 0-element dispatch is a safe no-op (Vulkan tolerates it); a quick assertion/comment would document the intent.

No blocking issues beyond #1, which I'd recommend addressing before this lands given it directly contradicts the "Mirrors Vulkan resize_slice_copy_node" claim in the description.
• Base: origin/gh/JulianCloudNTH/73/base

[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
[ghstack-poisoned]

@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review automatically exported from Phabricator review in Meta.

[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:28 — with GitHub Actions Inactive
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 21:20 — with GitHub Actions Inactive
@meta-codesync
meta-codesyncBot merged commit fb80edf into gh/JulianCloudNTH/73/baseJul 4, 2026
181 of 183 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JulianCloudNTH/73/head branch July 4, 2026 17:06
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@SS-JIA@nil-is-all@JCNTH
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

[ExecuTorch][WebGPU] Dynamic resize hook for slice_copy (dynamic start/end) - #20581

Merged
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/73/basefrom
gh/JulianCloudNTH/73/head
Jul 4, 2026
Merged

[ExecuTorch][WebGPU] Dynamic resize hook for slice_copy (dynamic start/end)#20581
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/73/basefrom
gh/JulianCloudNTH/73/head

Conversation

@ghost

@ghostghost commented Jun 28, 2026

Copy link
Copy Markdown

Stack from ghstack (oldest at bottom):

Make slice_copy support a dynamic gather range so the RoPE-freqs slice [input_pos : input_pos + S] works under one dynamic graph.

Problem: the static slice handler read start via a scalar reader that throws on a SymInt and ignored end (output length baked AOT). The RoPE-freqs slice uses a SymInt input_pos for start and a live S for the range, so the static op could neither build nor resize for it.

Solution: read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.

  • Before: start is a static scalar (SymInt throws); end ignored; output length fixed at the serialized max.
  • After: start/end read via a SymInt-aware reader; a hook recomputes out[dim] = (end - start + step - 1) / step, rewrites out_meta/in_meta/params UBOs + the dispatch count, and sets the output cur_dims.

Implementation:

  • Hook registered on the start/end value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
  • Output/input TensorMeta rebuilt from live dims; dim/step stay static.
  • Keep the uniforms alive via own_uniform_buffer so the hook can rewrite them.
  • Mirrors Vulkan resize_slice_copy_node.

Constraints: fp32-only; dim/step static; numerics + layout unchanged; inert on a static graph.

NOTE (stacking): this diff sits on top of the in-review slice_copy op (D108793168); rebase onto it once that op lands on master.

Co-authored-with: Claude Code.
@exported-using-ghexport

Differential Revision: D109906092

Differential Revision: D109906092

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jun 28, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20581

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit 11a2c7b with merge base 124b677 (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 28, 2026
@ghost

Copy link
Copy Markdown
Author

@claude review and check for any areas or opportunities for modularization

@claude

claudeBot commented Jun 29, 2026

Copy link
Copy Markdown

Claude finished @JulianCloudNTH's task in 2m 32s —— View job


Code Review: Dynamic resize hook for slice_copy

  • Read PR diff and full Slice.cpp
  • Check WebGPUGraph resize API (read_symint, add_resize_hook, add_tensor_resize_hook, own_uniform_buffer, propagate_resize)
  • Compare against Vulkan resize_slice_copy_node
  • Identify correctness issues
  • Identify modularization opportunities

Overall this is clean and matches the established WebGPU dynamic-shape pattern well (SymInt hooks → tensor hooks, idempotent recompute, own_uniform_buffer to keep UBOs writable). The hook is correctly inert on a static graph. A few items below — one I'd consider a real (latent) correctness bug, the rest are quality/modularization.


1. Output non-slice dims should be derived from live input dims (divergence from Vulkan)

Slice.cpp:224-226

std::vector<int64_t> od = g.cur_dims(out_id); // <- seeds from output's *prior* dims
od[dim] = len;
g.set_cur_dims(out_id, od);

The Vulkan reference this claims to mirror derives the output shape from the input sizes each resize:

std::vector<int64_t> new_out_sizes = in_sizes; // backends/vulkan/.../Slice.cpp:63-64
new_out_sizes.at(dim) = (end - start + step - 1) / step;

For slice_copy, every non-dim axis of the output is by definition equal to the input's. By seeding od from the output's prior cur_dims and only rewriting od[dim], this hook leaves the other axes stale. The in_meta it writes does track in_dims, so if any non-slice axis of the input is also dynamic, out_meta and in_meta disagree on those axes — a silent shape mismatch.

For the RoPE-freqs case only dim is dynamic, so it happens to work today, but it's fragile and diverges from the reference. Suggest mirroring Vulkan:

std::vector<int64_t> od = in_dims; // already have `const auto& in_dims = g.cur_dims(in_id);`
od[dim] = len;

Fix this →


2. Redundant hook executions (minor; idempotent so not a bug)

The same recompute closure is registered up to three times — on start_id, end_id, and in_id. When more than one trigger is dirty in the same step (e.g. input_pos SymInt + an input resize), recompute runs 2–3× per propagate_resize. It's fully idempotent (recomputes from live state), so output is correct, just wasted wgpuQueueWriteBuffer calls. Worth a one-line comment noting the intentional redundancy, or skip if the duplicate writes are negligible.


3. Modularization opportunities (the explicit ask)

a. Extract shared slice geometry — highest value. The build path (Slice.cpp:100-116) and the hook (Slice.cpp:217-242) independently compute start/clamp and fill SliceParams (dim/start/step). A small helper would remove the duplication and make the build path and resize path provably consistent:

structSliceGeom { std::vector<int64_t> out_dims; SliceParams params; };
SliceGeom compute_slice_geom(const std::vector<int64_t>& in_dims,
int64_t dim, int64_t start, int64_t end, int64_t step);

Both sites call it; the hook additionally writes the UBOs/dispatch. This also naturally fixes finding #1, since the helper would build out_dims from in_dims.

b. read_index / is_symint / norm_clamp are reusable across transfer ops.slice, and soon select/narrow/index_select, all need "read a possibly-dynamic index" + "normalize+clamp into [0, size]". Vulkan already centralizes this as normalize_idx. Consider hoisting these three free functions into a shared WebGPU helper header (e.g. alongside TensorMeta.h or a new transfer helper) rather than re-declaring them per op. Even within this file, read_index and read_scalar (Slice.cpp:34-61) share the Int → INT64_MAX→default branch — read_scalar could be expressed as the strict variant of read_index.

c. Bind-group / pipeline boilerplate (out of scope here).Slice.cpp:122-200 is ~80 lines of BGL + pipeline + bind-group setup that is near-identical to other WebGPU ops. Not for this PR, but a make_compute_pipeline(...) / make_bind_group(...) helper would shrink every op file substantially. Flagging as a follow-up direction.


Minor

  • Slice.cpp:76 comment still reads end unread (out shape is AOT) — now stale, end_id is read in the hook. Update or drop.
  • len == 0 path → om.numel == 0compute_1d_workgroup_count(..., 0, ...). Confirm a 0-element dispatch is a safe no-op (Vulkan tolerates it); a quick assertion/comment would document the intent.

No blocking issues beyond #1, which I'd recommend addressing before this lands given it directly contradicts the "Mirrors Vulkan resize_slice_copy_node" claim in the description.
• Base: origin/gh/JulianCloudNTH/73/base

[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
[ghstack-poisoned]

@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review automatically exported from Phabricator review in Meta.

[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:28 — with GitHub Actions Inactive
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 21:20 — with GitHub Actions Inactive
@meta-codesync
meta-codesyncBot merged commit fb80edf into gh/JulianCloudNTH/73/baseJul 4, 2026
181 of 183 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JulianCloudNTH/73/head branch July 4, 2026 17:06
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@SS-JIA@nil-is-all@JCNTH
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[ExecuTorch][WebGPU] Dynamic resize hook for slice_copy (dynamic start/end) - #20581

Merged
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/73/basefrom
gh/JulianCloudNTH/73/head
Jul 4, 2026
Merged

[ExecuTorch][WebGPU] Dynamic resize hook for slice_copy (dynamic start/end)#20581
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/73/basefrom
gh/JulianCloudNTH/73/head

Conversation

@ghost

@ghostghost commented Jun 28, 2026

Copy link
Copy Markdown

Stack from ghstack (oldest at bottom):

Make slice_copy support a dynamic gather range so the RoPE-freqs slice [input_pos : input_pos + S] works under one dynamic graph.

Problem: the static slice handler read start via a scalar reader that throws on a SymInt and ignored end (output length baked AOT). The RoPE-freqs slice uses a SymInt input_pos for start and a live S for the range, so the static op could neither build nor resize for it.

Solution: read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.

  • Before: start is a static scalar (SymInt throws); end ignored; output length fixed at the serialized max.
  • After: start/end read via a SymInt-aware reader; a hook recomputes out[dim] = (end - start + step - 1) / step, rewrites out_meta/in_meta/params UBOs + the dispatch count, and sets the output cur_dims.

Implementation:

  • Hook registered on the start/end value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
  • Output/input TensorMeta rebuilt from live dims; dim/step stay static.
  • Keep the uniforms alive via own_uniform_buffer so the hook can rewrite them.
  • Mirrors Vulkan resize_slice_copy_node.

Constraints: fp32-only; dim/step static; numerics + layout unchanged; inert on a static graph.

NOTE (stacking): this diff sits on top of the in-review slice_copy op (D108793168); rebase onto it once that op lands on master.

Co-authored-with: Claude Code.
@exported-using-ghexport

Differential Revision: D109906092

Differential Revision: D109906092

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jun 28, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20581

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit 11a2c7b with merge base 124b677 (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 28, 2026
@ghost

Copy link
Copy Markdown
Author

@claude review and check for any areas or opportunities for modularization

@claude

claudeBot commented Jun 29, 2026

Copy link
Copy Markdown

Claude finished @JulianCloudNTH's task in 2m 32s —— View job


Code Review: Dynamic resize hook for slice_copy

  • Read PR diff and full Slice.cpp
  • Check WebGPUGraph resize API (read_symint, add_resize_hook, add_tensor_resize_hook, own_uniform_buffer, propagate_resize)
  • Compare against Vulkan resize_slice_copy_node
  • Identify correctness issues
  • Identify modularization opportunities

Overall this is clean and matches the established WebGPU dynamic-shape pattern well (SymInt hooks → tensor hooks, idempotent recompute, own_uniform_buffer to keep UBOs writable). The hook is correctly inert on a static graph. A few items below — one I'd consider a real (latent) correctness bug, the rest are quality/modularization.


1. Output non-slice dims should be derived from live input dims (divergence from Vulkan)

Slice.cpp:224-226

std::vector<int64_t> od = g.cur_dims(out_id); // <- seeds from output's *prior* dims
od[dim] = len;
g.set_cur_dims(out_id, od);

The Vulkan reference this claims to mirror derives the output shape from the input sizes each resize:

std::vector<int64_t> new_out_sizes = in_sizes; // backends/vulkan/.../Slice.cpp:63-64
new_out_sizes.at(dim) = (end - start + step - 1) / step;

For slice_copy, every non-dim axis of the output is by definition equal to the input's. By seeding od from the output's prior cur_dims and only rewriting od[dim], this hook leaves the other axes stale. The in_meta it writes does track in_dims, so if any non-slice axis of the input is also dynamic, out_meta and in_meta disagree on those axes — a silent shape mismatch.

For the RoPE-freqs case only dim is dynamic, so it happens to work today, but it's fragile and diverges from the reference. Suggest mirroring Vulkan:

std::vector<int64_t> od = in_dims; // already have `const auto& in_dims = g.cur_dims(in_id);`
od[dim] = len;

Fix this →


2. Redundant hook executions (minor; idempotent so not a bug)

The same recompute closure is registered up to three times — on start_id, end_id, and in_id. When more than one trigger is dirty in the same step (e.g. input_pos SymInt + an input resize), recompute runs 2–3× per propagate_resize. It's fully idempotent (recomputes from live state), so output is correct, just wasted wgpuQueueWriteBuffer calls. Worth a one-line comment noting the intentional redundancy, or skip if the duplicate writes are negligible.


3. Modularization opportunities (the explicit ask)

a. Extract shared slice geometry — highest value. The build path (Slice.cpp:100-116) and the hook (Slice.cpp:217-242) independently compute start/clamp and fill SliceParams (dim/start/step). A small helper would remove the duplication and make the build path and resize path provably consistent:

structSliceGeom { std::vector<int64_t> out_dims; SliceParams params; };
SliceGeom compute_slice_geom(const std::vector<int64_t>& in_dims,
int64_t dim, int64_t start, int64_t end, int64_t step);

Both sites call it; the hook additionally writes the UBOs/dispatch. This also naturally fixes finding #1, since the helper would build out_dims from in_dims.

b. read_index / is_symint / norm_clamp are reusable across transfer ops.slice, and soon select/narrow/index_select, all need "read a possibly-dynamic index" + "normalize+clamp into [0, size]". Vulkan already centralizes this as normalize_idx. Consider hoisting these three free functions into a shared WebGPU helper header (e.g. alongside TensorMeta.h or a new transfer helper) rather than re-declaring them per op. Even within this file, read_index and read_scalar (Slice.cpp:34-61) share the Int → INT64_MAX→default branch — read_scalar could be expressed as the strict variant of read_index.

c. Bind-group / pipeline boilerplate (out of scope here).Slice.cpp:122-200 is ~80 lines of BGL + pipeline + bind-group setup that is near-identical to other WebGPU ops. Not for this PR, but a make_compute_pipeline(...) / make_bind_group(...) helper would shrink every op file substantially. Flagging as a follow-up direction.


Minor

  • Slice.cpp:76 comment still reads end unread (out shape is AOT) — now stale, end_id is read in the hook. Update or drop.
  • len == 0 path → om.numel == 0compute_1d_workgroup_count(..., 0, ...). Confirm a 0-element dispatch is a safe no-op (Vulkan tolerates it); a quick assertion/comment would document the intent.

No blocking issues beyond #1, which I'd recommend addressing before this lands given it directly contradicts the "Mirrors Vulkan resize_slice_copy_node" claim in the description.
• Base: origin/gh/JulianCloudNTH/73/base

[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
[ghstack-poisoned]

@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review automatically exported from Phabricator review in Meta.

[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:28 — with GitHub Actions Inactive
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 21:20 — with GitHub Actions Inactive
@meta-codesync
meta-codesyncBot merged commit fb80edf into gh/JulianCloudNTH/73/baseJul 4, 2026
181 of 183 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JulianCloudNTH/73/head branch July 4, 2026 17:06
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@SS-JIA@nil-is-all@JCNTH
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[ExecuTorch][WebGPU] Dynamic resize hook for slice_copy (dynamic start/end) - #20581

Merged
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/73/basefrom
gh/JulianCloudNTH/73/head
Jul 4, 2026
Merged

[ExecuTorch][WebGPU] Dynamic resize hook for slice_copy (dynamic start/end)#20581
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/73/basefrom
gh/JulianCloudNTH/73/head

Conversation

@ghost

@ghostghost commented Jun 28, 2026

Copy link
Copy Markdown

Stack from ghstack (oldest at bottom):

Make slice_copy support a dynamic gather range so the RoPE-freqs slice [input_pos : input_pos + S] works under one dynamic graph.

Problem: the static slice handler read start via a scalar reader that throws on a SymInt and ignored end (output length baked AOT). The RoPE-freqs slice uses a SymInt input_pos for start and a live S for the range, so the static op could neither build nor resize for it.

Solution: read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.

  • Before: start is a static scalar (SymInt throws); end ignored; output length fixed at the serialized max.
  • After: start/end read via a SymInt-aware reader; a hook recomputes out[dim] = (end - start + step - 1) / step, rewrites out_meta/in_meta/params UBOs + the dispatch count, and sets the output cur_dims.

Implementation:

  • Hook registered on the start/end value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
  • Output/input TensorMeta rebuilt from live dims; dim/step stay static.
  • Keep the uniforms alive via own_uniform_buffer so the hook can rewrite them.
  • Mirrors Vulkan resize_slice_copy_node.

Constraints: fp32-only; dim/step static; numerics + layout unchanged; inert on a static graph.

NOTE (stacking): this diff sits on top of the in-review slice_copy op (D108793168); rebase onto it once that op lands on master.

Co-authored-with: Claude Code.
@exported-using-ghexport

Differential Revision: D109906092

Differential Revision: D109906092

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jun 28, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20581

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit 11a2c7b with merge base 124b677 (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 28, 2026
@ghost

Copy link
Copy Markdown
Author

@claude review and check for any areas or opportunities for modularization

@claude

claudeBot commented Jun 29, 2026

Copy link
Copy Markdown

Claude finished @JulianCloudNTH's task in 2m 32s —— View job


Code Review: Dynamic resize hook for slice_copy

  • Read PR diff and full Slice.cpp
  • Check WebGPUGraph resize API (read_symint, add_resize_hook, add_tensor_resize_hook, own_uniform_buffer, propagate_resize)
  • Compare against Vulkan resize_slice_copy_node
  • Identify correctness issues
  • Identify modularization opportunities

Overall this is clean and matches the established WebGPU dynamic-shape pattern well (SymInt hooks → tensor hooks, idempotent recompute, own_uniform_buffer to keep UBOs writable). The hook is correctly inert on a static graph. A few items below — one I'd consider a real (latent) correctness bug, the rest are quality/modularization.


1. Output non-slice dims should be derived from live input dims (divergence from Vulkan)

Slice.cpp:224-226

std::vector<int64_t> od = g.cur_dims(out_id); // <- seeds from output's *prior* dims
od[dim] = len;
g.set_cur_dims(out_id, od);

The Vulkan reference this claims to mirror derives the output shape from the input sizes each resize:

std::vector<int64_t> new_out_sizes = in_sizes; // backends/vulkan/.../Slice.cpp:63-64
new_out_sizes.at(dim) = (end - start + step - 1) / step;

For slice_copy, every non-dim axis of the output is by definition equal to the input's. By seeding od from the output's prior cur_dims and only rewriting od[dim], this hook leaves the other axes stale. The in_meta it writes does track in_dims, so if any non-slice axis of the input is also dynamic, out_meta and in_meta disagree on those axes — a silent shape mismatch.

For the RoPE-freqs case only dim is dynamic, so it happens to work today, but it's fragile and diverges from the reference. Suggest mirroring Vulkan:

std::vector<int64_t> od = in_dims; // already have `const auto& in_dims = g.cur_dims(in_id);`
od[dim] = len;

Fix this →


2. Redundant hook executions (minor; idempotent so not a bug)

The same recompute closure is registered up to three times — on start_id, end_id, and in_id. When more than one trigger is dirty in the same step (e.g. input_pos SymInt + an input resize), recompute runs 2–3× per propagate_resize. It's fully idempotent (recomputes from live state), so output is correct, just wasted wgpuQueueWriteBuffer calls. Worth a one-line comment noting the intentional redundancy, or skip if the duplicate writes are negligible.


3. Modularization opportunities (the explicit ask)

a. Extract shared slice geometry — highest value. The build path (Slice.cpp:100-116) and the hook (Slice.cpp:217-242) independently compute start/clamp and fill SliceParams (dim/start/step). A small helper would remove the duplication and make the build path and resize path provably consistent:

structSliceGeom { std::vector<int64_t> out_dims; SliceParams params; };
SliceGeom compute_slice_geom(const std::vector<int64_t>& in_dims,
int64_t dim, int64_t start, int64_t end, int64_t step);

Both sites call it; the hook additionally writes the UBOs/dispatch. This also naturally fixes finding #1, since the helper would build out_dims from in_dims.

b. read_index / is_symint / norm_clamp are reusable across transfer ops.slice, and soon select/narrow/index_select, all need "read a possibly-dynamic index" + "normalize+clamp into [0, size]". Vulkan already centralizes this as normalize_idx. Consider hoisting these three free functions into a shared WebGPU helper header (e.g. alongside TensorMeta.h or a new transfer helper) rather than re-declaring them per op. Even within this file, read_index and read_scalar (Slice.cpp:34-61) share the Int → INT64_MAX→default branch — read_scalar could be expressed as the strict variant of read_index.

c. Bind-group / pipeline boilerplate (out of scope here).Slice.cpp:122-200 is ~80 lines of BGL + pipeline + bind-group setup that is near-identical to other WebGPU ops. Not for this PR, but a make_compute_pipeline(...) / make_bind_group(...) helper would shrink every op file substantially. Flagging as a follow-up direction.


Minor

  • Slice.cpp:76 comment still reads end unread (out shape is AOT) — now stale, end_id is read in the hook. Update or drop.
  • len == 0 path → om.numel == 0compute_1d_workgroup_count(..., 0, ...). Confirm a 0-element dispatch is a safe no-op (Vulkan tolerates it); a quick assertion/comment would document the intent.

No blocking issues beyond #1, which I'd recommend addressing before this lands given it directly contradicts the "Mirrors Vulkan resize_slice_copy_node" claim in the description.
• Base: origin/gh/JulianCloudNTH/73/base

[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
[ghstack-poisoned]

@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review automatically exported from Phabricator review in Meta.

[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:28 — with GitHub Actions Inactive
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 21:20 — with GitHub Actions Inactive
@meta-codesync
meta-codesyncBot merged commit fb80edf into gh/JulianCloudNTH/73/baseJul 4, 2026
181 of 183 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JulianCloudNTH/73/head branch July 4, 2026 17:06
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@SS-JIA@nil-is-all@JCNTH
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

[ExecuTorch][WebGPU] Dynamic resize hook for slice_copy (dynamic start/end) - #20581

Merged
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/73/basefrom
gh/JulianCloudNTH/73/head
Jul 4, 2026
Merged

[ExecuTorch][WebGPU] Dynamic resize hook for slice_copy (dynamic start/end)#20581
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/73/basefrom
gh/JulianCloudNTH/73/head

Conversation

@ghost

@ghostghost commented Jun 28, 2026

Copy link
Copy Markdown

Stack from ghstack (oldest at bottom):

Make slice_copy support a dynamic gather range so the RoPE-freqs slice [input_pos : input_pos + S] works under one dynamic graph.

Problem: the static slice handler read start via a scalar reader that throws on a SymInt and ignored end (output length baked AOT). The RoPE-freqs slice uses a SymInt input_pos for start and a live S for the range, so the static op could neither build nor resize for it.

Solution: read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.

  • Before: start is a static scalar (SymInt throws); end ignored; output length fixed at the serialized max.
  • After: start/end read via a SymInt-aware reader; a hook recomputes out[dim] = (end - start + step - 1) / step, rewrites out_meta/in_meta/params UBOs + the dispatch count, and sets the output cur_dims.

Implementation:

  • Hook registered on the start/end value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
  • Output/input TensorMeta rebuilt from live dims; dim/step stay static.
  • Keep the uniforms alive via own_uniform_buffer so the hook can rewrite them.
  • Mirrors Vulkan resize_slice_copy_node.

Constraints: fp32-only; dim/step static; numerics + layout unchanged; inert on a static graph.

NOTE (stacking): this diff sits on top of the in-review slice_copy op (D108793168); rebase onto it once that op lands on master.

Co-authored-with: Claude Code.
@exported-using-ghexport

Differential Revision: D109906092

Differential Revision: D109906092

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jun 28, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20581

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit 11a2c7b with merge base 124b677 (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 28, 2026
@ghost

Copy link
Copy Markdown
Author

@claude review and check for any areas or opportunities for modularization

@claude

claudeBot commented Jun 29, 2026

Copy link
Copy Markdown

Claude finished @JulianCloudNTH's task in 2m 32s —— View job


Code Review: Dynamic resize hook for slice_copy

  • Read PR diff and full Slice.cpp
  • Check WebGPUGraph resize API (read_symint, add_resize_hook, add_tensor_resize_hook, own_uniform_buffer, propagate_resize)
  • Compare against Vulkan resize_slice_copy_node
  • Identify correctness issues
  • Identify modularization opportunities

Overall this is clean and matches the established WebGPU dynamic-shape pattern well (SymInt hooks → tensor hooks, idempotent recompute, own_uniform_buffer to keep UBOs writable). The hook is correctly inert on a static graph. A few items below — one I'd consider a real (latent) correctness bug, the rest are quality/modularization.


1. Output non-slice dims should be derived from live input dims (divergence from Vulkan)

Slice.cpp:224-226

std::vector<int64_t> od = g.cur_dims(out_id); // <- seeds from output's *prior* dims
od[dim] = len;
g.set_cur_dims(out_id, od);

The Vulkan reference this claims to mirror derives the output shape from the input sizes each resize:

std::vector<int64_t> new_out_sizes = in_sizes; // backends/vulkan/.../Slice.cpp:63-64
new_out_sizes.at(dim) = (end - start + step - 1) / step;

For slice_copy, every non-dim axis of the output is by definition equal to the input's. By seeding od from the output's prior cur_dims and only rewriting od[dim], this hook leaves the other axes stale. The in_meta it writes does track in_dims, so if any non-slice axis of the input is also dynamic, out_meta and in_meta disagree on those axes — a silent shape mismatch.

For the RoPE-freqs case only dim is dynamic, so it happens to work today, but it's fragile and diverges from the reference. Suggest mirroring Vulkan:

std::vector<int64_t> od = in_dims; // already have `const auto& in_dims = g.cur_dims(in_id);`
od[dim] = len;

Fix this →


2. Redundant hook executions (minor; idempotent so not a bug)

The same recompute closure is registered up to three times — on start_id, end_id, and in_id. When more than one trigger is dirty in the same step (e.g. input_pos SymInt + an input resize), recompute runs 2–3× per propagate_resize. It's fully idempotent (recomputes from live state), so output is correct, just wasted wgpuQueueWriteBuffer calls. Worth a one-line comment noting the intentional redundancy, or skip if the duplicate writes are negligible.


3. Modularization opportunities (the explicit ask)

a. Extract shared slice geometry — highest value. The build path (Slice.cpp:100-116) and the hook (Slice.cpp:217-242) independently compute start/clamp and fill SliceParams (dim/start/step). A small helper would remove the duplication and make the build path and resize path provably consistent:

structSliceGeom { std::vector<int64_t> out_dims; SliceParams params; };
SliceGeom compute_slice_geom(const std::vector<int64_t>& in_dims,
int64_t dim, int64_t start, int64_t end, int64_t step);

Both sites call it; the hook additionally writes the UBOs/dispatch. This also naturally fixes finding #1, since the helper would build out_dims from in_dims.

b. read_index / is_symint / norm_clamp are reusable across transfer ops.slice, and soon select/narrow/index_select, all need "read a possibly-dynamic index" + "normalize+clamp into [0, size]". Vulkan already centralizes this as normalize_idx. Consider hoisting these three free functions into a shared WebGPU helper header (e.g. alongside TensorMeta.h or a new transfer helper) rather than re-declaring them per op. Even within this file, read_index and read_scalar (Slice.cpp:34-61) share the Int → INT64_MAX→default branch — read_scalar could be expressed as the strict variant of read_index.

c. Bind-group / pipeline boilerplate (out of scope here).Slice.cpp:122-200 is ~80 lines of BGL + pipeline + bind-group setup that is near-identical to other WebGPU ops. Not for this PR, but a make_compute_pipeline(...) / make_bind_group(...) helper would shrink every op file substantially. Flagging as a follow-up direction.


Minor

  • Slice.cpp:76 comment still reads end unread (out shape is AOT) — now stale, end_id is read in the hook. Update or drop.
  • len == 0 path → om.numel == 0compute_1d_workgroup_count(..., 0, ...). Confirm a 0-element dispatch is a safe no-op (Vulkan tolerates it); a quick assertion/comment would document the intent.

No blocking issues beyond #1, which I'd recommend addressing before this lands given it directly contradicts the "Mirrors Vulkan resize_slice_copy_node" claim in the description.
• Base: origin/gh/JulianCloudNTH/73/base

[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
[ghstack-poisoned]

@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review automatically exported from Phabricator review in Meta.

[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:28 — with GitHub Actions Inactive
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 21:20 — with GitHub Actions Inactive
@meta-codesync
meta-codesyncBot merged commit fb80edf into gh/JulianCloudNTH/73/baseJul 4, 2026
181 of 183 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JulianCloudNTH/73/head branch July 4, 2026 17:06
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@SS-JIA@nil-is-all@JCNTH
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[ExecuTorch][WebGPU] Dynamic resize hook for slice_copy (dynamic start/end) - #20581

Merged
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/73/basefrom
gh/JulianCloudNTH/73/head
Jul 4, 2026
Merged

[ExecuTorch][WebGPU] Dynamic resize hook for slice_copy (dynamic start/end)#20581
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/73/basefrom
gh/JulianCloudNTH/73/head

Conversation

@ghost

@ghostghost commented Jun 28, 2026

Copy link
Copy Markdown

Stack from ghstack (oldest at bottom):

Make slice_copy support a dynamic gather range so the RoPE-freqs slice [input_pos : input_pos + S] works under one dynamic graph.

Problem: the static slice handler read start via a scalar reader that throws on a SymInt and ignored end (output length baked AOT). The RoPE-freqs slice uses a SymInt input_pos for start and a live S for the range, so the static op could neither build nor resize for it.

Solution: read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.

  • Before: start is a static scalar (SymInt throws); end ignored; output length fixed at the serialized max.
  • After: start/end read via a SymInt-aware reader; a hook recomputes out[dim] = (end - start + step - 1) / step, rewrites out_meta/in_meta/params UBOs + the dispatch count, and sets the output cur_dims.

Implementation:

  • Hook registered on the start/end value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
  • Output/input TensorMeta rebuilt from live dims; dim/step stay static.
  • Keep the uniforms alive via own_uniform_buffer so the hook can rewrite them.
  • Mirrors Vulkan resize_slice_copy_node.

Constraints: fp32-only; dim/step static; numerics + layout unchanged; inert on a static graph.

NOTE (stacking): this diff sits on top of the in-review slice_copy op (D108793168); rebase onto it once that op lands on master.

Co-authored-with: Claude Code.
@exported-using-ghexport

Differential Revision: D109906092

Differential Revision: D109906092

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jun 28, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20581

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit 11a2c7b with merge base 124b677 (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 28, 2026
@ghost

Copy link
Copy Markdown
Author

@claude review and check for any areas or opportunities for modularization

@claude

claudeBot commented Jun 29, 2026

Copy link
Copy Markdown

Claude finished @JulianCloudNTH's task in 2m 32s —— View job


Code Review: Dynamic resize hook for slice_copy

  • Read PR diff and full Slice.cpp
  • Check WebGPUGraph resize API (read_symint, add_resize_hook, add_tensor_resize_hook, own_uniform_buffer, propagate_resize)
  • Compare against Vulkan resize_slice_copy_node
  • Identify correctness issues
  • Identify modularization opportunities

Overall this is clean and matches the established WebGPU dynamic-shape pattern well (SymInt hooks → tensor hooks, idempotent recompute, own_uniform_buffer to keep UBOs writable). The hook is correctly inert on a static graph. A few items below — one I'd consider a real (latent) correctness bug, the rest are quality/modularization.


1. Output non-slice dims should be derived from live input dims (divergence from Vulkan)

Slice.cpp:224-226

std::vector<int64_t> od = g.cur_dims(out_id); // <- seeds from output's *prior* dims
od[dim] = len;
g.set_cur_dims(out_id, od);

The Vulkan reference this claims to mirror derives the output shape from the input sizes each resize:

std::vector<int64_t> new_out_sizes = in_sizes; // backends/vulkan/.../Slice.cpp:63-64
new_out_sizes.at(dim) = (end - start + step - 1) / step;

For slice_copy, every non-dim axis of the output is by definition equal to the input's. By seeding od from the output's prior cur_dims and only rewriting od[dim], this hook leaves the other axes stale. The in_meta it writes does track in_dims, so if any non-slice axis of the input is also dynamic, out_meta and in_meta disagree on those axes — a silent shape mismatch.

For the RoPE-freqs case only dim is dynamic, so it happens to work today, but it's fragile and diverges from the reference. Suggest mirroring Vulkan:

std::vector<int64_t> od = in_dims; // already have `const auto& in_dims = g.cur_dims(in_id);`
od[dim] = len;

Fix this →


2. Redundant hook executions (minor; idempotent so not a bug)

The same recompute closure is registered up to three times — on start_id, end_id, and in_id. When more than one trigger is dirty in the same step (e.g. input_pos SymInt + an input resize), recompute runs 2–3× per propagate_resize. It's fully idempotent (recomputes from live state), so output is correct, just wasted wgpuQueueWriteBuffer calls. Worth a one-line comment noting the intentional redundancy, or skip if the duplicate writes are negligible.


3. Modularization opportunities (the explicit ask)

a. Extract shared slice geometry — highest value. The build path (Slice.cpp:100-116) and the hook (Slice.cpp:217-242) independently compute start/clamp and fill SliceParams (dim/start/step). A small helper would remove the duplication and make the build path and resize path provably consistent:

structSliceGeom { std::vector<int64_t> out_dims; SliceParams params; };
SliceGeom compute_slice_geom(const std::vector<int64_t>& in_dims,
int64_t dim, int64_t start, int64_t end, int64_t step);

Both sites call it; the hook additionally writes the UBOs/dispatch. This also naturally fixes finding #1, since the helper would build out_dims from in_dims.

b. read_index / is_symint / norm_clamp are reusable across transfer ops.slice, and soon select/narrow/index_select, all need "read a possibly-dynamic index" + "normalize+clamp into [0, size]". Vulkan already centralizes this as normalize_idx. Consider hoisting these three free functions into a shared WebGPU helper header (e.g. alongside TensorMeta.h or a new transfer helper) rather than re-declaring them per op. Even within this file, read_index and read_scalar (Slice.cpp:34-61) share the Int → INT64_MAX→default branch — read_scalar could be expressed as the strict variant of read_index.

c. Bind-group / pipeline boilerplate (out of scope here).Slice.cpp:122-200 is ~80 lines of BGL + pipeline + bind-group setup that is near-identical to other WebGPU ops. Not for this PR, but a make_compute_pipeline(...) / make_bind_group(...) helper would shrink every op file substantially. Flagging as a follow-up direction.


Minor

  • Slice.cpp:76 comment still reads end unread (out shape is AOT) — now stale, end_id is read in the hook. Update or drop.
  • len == 0 path → om.numel == 0compute_1d_workgroup_count(..., 0, ...). Confirm a 0-element dispatch is a safe no-op (Vulkan tolerates it); a quick assertion/comment would document the intent.

No blocking issues beyond #1, which I'd recommend addressing before this lands given it directly contradicts the "Mirrors Vulkan resize_slice_copy_node" claim in the description.
• Base: origin/gh/JulianCloudNTH/73/base

[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
[ghstack-poisoned]

@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review automatically exported from Phabricator review in Meta.

[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:28 — with GitHub Actions Inactive
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 21:20 — with GitHub Actions Inactive
@meta-codesync
meta-codesyncBot merged commit fb80edf into gh/JulianCloudNTH/73/baseJul 4, 2026
181 of 183 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JulianCloudNTH/73/head branch July 4, 2026 17:06
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@SS-JIA@nil-is-all@JCNTH
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[ExecuTorch][WebGPU] Dynamic resize hook for slice_copy (dynamic start/end) - #20581

Merged
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/73/basefrom
gh/JulianCloudNTH/73/head
Jul 4, 2026
Merged

[ExecuTorch][WebGPU] Dynamic resize hook for slice_copy (dynamic start/end)#20581
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/73/basefrom
gh/JulianCloudNTH/73/head

Conversation

@ghost

@ghostghost commented Jun 28, 2026

Copy link
Copy Markdown

Stack from ghstack (oldest at bottom):

Make slice_copy support a dynamic gather range so the RoPE-freqs slice [input_pos : input_pos + S] works under one dynamic graph.

Problem: the static slice handler read start via a scalar reader that throws on a SymInt and ignored end (output length baked AOT). The RoPE-freqs slice uses a SymInt input_pos for start and a live S for the range, so the static op could neither build nor resize for it.

Solution: read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.

  • Before: start is a static scalar (SymInt throws); end ignored; output length fixed at the serialized max.
  • After: start/end read via a SymInt-aware reader; a hook recomputes out[dim] = (end - start + step - 1) / step, rewrites out_meta/in_meta/params UBOs + the dispatch count, and sets the output cur_dims.

Implementation:

  • Hook registered on the start/end value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
  • Output/input TensorMeta rebuilt from live dims; dim/step stay static.
  • Keep the uniforms alive via own_uniform_buffer so the hook can rewrite them.
  • Mirrors Vulkan resize_slice_copy_node.

Constraints: fp32-only; dim/step static; numerics + layout unchanged; inert on a static graph.

NOTE (stacking): this diff sits on top of the in-review slice_copy op (D108793168); rebase onto it once that op lands on master.

Co-authored-with: Claude Code.
@exported-using-ghexport

Differential Revision: D109906092

Differential Revision: D109906092

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jun 28, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20581

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit 11a2c7b with merge base 124b677 (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 28, 2026
@ghost

Copy link
Copy Markdown
Author

@claude review and check for any areas or opportunities for modularization

@claude

claudeBot commented Jun 29, 2026

Copy link
Copy Markdown

Claude finished @JulianCloudNTH's task in 2m 32s —— View job


Code Review: Dynamic resize hook for slice_copy

  • Read PR diff and full Slice.cpp
  • Check WebGPUGraph resize API (read_symint, add_resize_hook, add_tensor_resize_hook, own_uniform_buffer, propagate_resize)
  • Compare against Vulkan resize_slice_copy_node
  • Identify correctness issues
  • Identify modularization opportunities

Overall this is clean and matches the established WebGPU dynamic-shape pattern well (SymInt hooks → tensor hooks, idempotent recompute, own_uniform_buffer to keep UBOs writable). The hook is correctly inert on a static graph. A few items below — one I'd consider a real (latent) correctness bug, the rest are quality/modularization.


1. Output non-slice dims should be derived from live input dims (divergence from Vulkan)

Slice.cpp:224-226

std::vector<int64_t> od = g.cur_dims(out_id); // <- seeds from output's *prior* dims
od[dim] = len;
g.set_cur_dims(out_id, od);

The Vulkan reference this claims to mirror derives the output shape from the input sizes each resize:

std::vector<int64_t> new_out_sizes = in_sizes; // backends/vulkan/.../Slice.cpp:63-64
new_out_sizes.at(dim) = (end - start + step - 1) / step;

For slice_copy, every non-dim axis of the output is by definition equal to the input's. By seeding od from the output's prior cur_dims and only rewriting od[dim], this hook leaves the other axes stale. The in_meta it writes does track in_dims, so if any non-slice axis of the input is also dynamic, out_meta and in_meta disagree on those axes — a silent shape mismatch.

For the RoPE-freqs case only dim is dynamic, so it happens to work today, but it's fragile and diverges from the reference. Suggest mirroring Vulkan:

std::vector<int64_t> od = in_dims; // already have `const auto& in_dims = g.cur_dims(in_id);`
od[dim] = len;

Fix this →


2. Redundant hook executions (minor; idempotent so not a bug)

The same recompute closure is registered up to three times — on start_id, end_id, and in_id. When more than one trigger is dirty in the same step (e.g. input_pos SymInt + an input resize), recompute runs 2–3× per propagate_resize. It's fully idempotent (recomputes from live state), so output is correct, just wasted wgpuQueueWriteBuffer calls. Worth a one-line comment noting the intentional redundancy, or skip if the duplicate writes are negligible.


3. Modularization opportunities (the explicit ask)

a. Extract shared slice geometry — highest value. The build path (Slice.cpp:100-116) and the hook (Slice.cpp:217-242) independently compute start/clamp and fill SliceParams (dim/start/step). A small helper would remove the duplication and make the build path and resize path provably consistent:

structSliceGeom { std::vector<int64_t> out_dims; SliceParams params; };
SliceGeom compute_slice_geom(const std::vector<int64_t>& in_dims,
int64_t dim, int64_t start, int64_t end, int64_t step);

Both sites call it; the hook additionally writes the UBOs/dispatch. This also naturally fixes finding #1, since the helper would build out_dims from in_dims.

b. read_index / is_symint / norm_clamp are reusable across transfer ops.slice, and soon select/narrow/index_select, all need "read a possibly-dynamic index" + "normalize+clamp into [0, size]". Vulkan already centralizes this as normalize_idx. Consider hoisting these three free functions into a shared WebGPU helper header (e.g. alongside TensorMeta.h or a new transfer helper) rather than re-declaring them per op. Even within this file, read_index and read_scalar (Slice.cpp:34-61) share the Int → INT64_MAX→default branch — read_scalar could be expressed as the strict variant of read_index.

c. Bind-group / pipeline boilerplate (out of scope here).Slice.cpp:122-200 is ~80 lines of BGL + pipeline + bind-group setup that is near-identical to other WebGPU ops. Not for this PR, but a make_compute_pipeline(...) / make_bind_group(...) helper would shrink every op file substantially. Flagging as a follow-up direction.


Minor

  • Slice.cpp:76 comment still reads end unread (out shape is AOT) — now stale, end_id is read in the hook. Update or drop.
  • len == 0 path → om.numel == 0compute_1d_workgroup_count(..., 0, ...). Confirm a 0-element dispatch is a safe no-op (Vulkan tolerates it); a quick assertion/comment would document the intent.

No blocking issues beyond #1, which I'd recommend addressing before this lands given it directly contradicts the "Mirrors Vulkan resize_slice_copy_node" claim in the description.
• Base: origin/gh/JulianCloudNTH/73/base

[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
[ghstack-poisoned]

@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review automatically exported from Phabricator review in Meta.

[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:28 — with GitHub Actions Inactive
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 21:20 — with GitHub Actions Inactive
@meta-codesync
meta-codesyncBot merged commit fb80edf into gh/JulianCloudNTH/73/baseJul 4, 2026
181 of 183 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JulianCloudNTH/73/head branch July 4, 2026 17:06
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@SS-JIA@nil-is-all@JCNTH
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

[ExecuTorch][WebGPU] Dynamic resize hook for slice_copy (dynamic start/end) - #20581

Merged
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/73/basefrom
gh/JulianCloudNTH/73/head
Jul 4, 2026
Merged

[ExecuTorch][WebGPU] Dynamic resize hook for slice_copy (dynamic start/end)#20581
meta-codesync[bot] merged 6 commits into
gh/JulianCloudNTH/73/basefrom
gh/JulianCloudNTH/73/head

Conversation

@ghost

@ghostghost commented Jun 28, 2026

Copy link
Copy Markdown

Stack from ghstack (oldest at bottom):

Make slice_copy support a dynamic gather range so the RoPE-freqs slice [input_pos : input_pos + S] works under one dynamic graph.

Problem: the static slice handler read start via a scalar reader that throws on a SymInt and ignored end (output length baked AOT). The RoPE-freqs slice uses a SymInt input_pos for start and a live S for the range, so the static op could neither build nor resize for it.

Solution: read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.

  • Before: start is a static scalar (SymInt throws); end ignored; output length fixed at the serialized max.
  • After: start/end read via a SymInt-aware reader; a hook recomputes out[dim] = (end - start + step - 1) / step, rewrites out_meta/in_meta/params UBOs + the dispatch count, and sets the output cur_dims.

Implementation:

  • Hook registered on the start/end value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
  • Output/input TensorMeta rebuilt from live dims; dim/step stay static.
  • Keep the uniforms alive via own_uniform_buffer so the hook can rewrite them.
  • Mirrors Vulkan resize_slice_copy_node.

Constraints: fp32-only; dim/step static; numerics + layout unchanged; inert on a static graph.

NOTE (stacking): this diff sits on top of the in-review slice_copy op (D108793168); rebase onto it once that op lands on master.

Co-authored-with: Claude Code.
@exported-using-ghexport

Differential Revision: D109906092

Differential Revision: D109906092

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jun 28, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20581

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit 11a2c7b with merge base 124b677 (image):

NEW FAILURE - The following job has failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 28, 2026
@ghost

Copy link
Copy Markdown
Author

@claude review and check for any areas or opportunities for modularization

@claude

claudeBot commented Jun 29, 2026

Copy link
Copy Markdown

Claude finished @JulianCloudNTH's task in 2m 32s —— View job


Code Review: Dynamic resize hook for slice_copy

  • Read PR diff and full Slice.cpp
  • Check WebGPUGraph resize API (read_symint, add_resize_hook, add_tensor_resize_hook, own_uniform_buffer, propagate_resize)
  • Compare against Vulkan resize_slice_copy_node
  • Identify correctness issues
  • Identify modularization opportunities

Overall this is clean and matches the established WebGPU dynamic-shape pattern well (SymInt hooks → tensor hooks, idempotent recompute, own_uniform_buffer to keep UBOs writable). The hook is correctly inert on a static graph. A few items below — one I'd consider a real (latent) correctness bug, the rest are quality/modularization.


1. Output non-slice dims should be derived from live input dims (divergence from Vulkan)

Slice.cpp:224-226

std::vector<int64_t> od = g.cur_dims(out_id); // <- seeds from output's *prior* dims
od[dim] = len;
g.set_cur_dims(out_id, od);

The Vulkan reference this claims to mirror derives the output shape from the input sizes each resize:

std::vector<int64_t> new_out_sizes = in_sizes; // backends/vulkan/.../Slice.cpp:63-64
new_out_sizes.at(dim) = (end - start + step - 1) / step;

For slice_copy, every non-dim axis of the output is by definition equal to the input's. By seeding od from the output's prior cur_dims and only rewriting od[dim], this hook leaves the other axes stale. The in_meta it writes does track in_dims, so if any non-slice axis of the input is also dynamic, out_meta and in_meta disagree on those axes — a silent shape mismatch.

For the RoPE-freqs case only dim is dynamic, so it happens to work today, but it's fragile and diverges from the reference. Suggest mirroring Vulkan:

std::vector<int64_t> od = in_dims; // already have `const auto& in_dims = g.cur_dims(in_id);`
od[dim] = len;

Fix this →


2. Redundant hook executions (minor; idempotent so not a bug)

The same recompute closure is registered up to three times — on start_id, end_id, and in_id. When more than one trigger is dirty in the same step (e.g. input_pos SymInt + an input resize), recompute runs 2–3× per propagate_resize. It's fully idempotent (recomputes from live state), so output is correct, just wasted wgpuQueueWriteBuffer calls. Worth a one-line comment noting the intentional redundancy, or skip if the duplicate writes are negligible.


3. Modularization opportunities (the explicit ask)

a. Extract shared slice geometry — highest value. The build path (Slice.cpp:100-116) and the hook (Slice.cpp:217-242) independently compute start/clamp and fill SliceParams (dim/start/step). A small helper would remove the duplication and make the build path and resize path provably consistent:

structSliceGeom { std::vector<int64_t> out_dims; SliceParams params; };
SliceGeom compute_slice_geom(const std::vector<int64_t>& in_dims,
int64_t dim, int64_t start, int64_t end, int64_t step);

Both sites call it; the hook additionally writes the UBOs/dispatch. This also naturally fixes finding #1, since the helper would build out_dims from in_dims.

b. read_index / is_symint / norm_clamp are reusable across transfer ops.slice, and soon select/narrow/index_select, all need "read a possibly-dynamic index" + "normalize+clamp into [0, size]". Vulkan already centralizes this as normalize_idx. Consider hoisting these three free functions into a shared WebGPU helper header (e.g. alongside TensorMeta.h or a new transfer helper) rather than re-declaring them per op. Even within this file, read_index and read_scalar (Slice.cpp:34-61) share the Int → INT64_MAX→default branch — read_scalar could be expressed as the strict variant of read_index.

c. Bind-group / pipeline boilerplate (out of scope here).Slice.cpp:122-200 is ~80 lines of BGL + pipeline + bind-group setup that is near-identical to other WebGPU ops. Not for this PR, but a make_compute_pipeline(...) / make_bind_group(...) helper would shrink every op file substantially. Flagging as a follow-up direction.


Minor

  • Slice.cpp:76 comment still reads end unread (out shape is AOT) — now stale, end_id is read in the hook. Update or drop.
  • len == 0 path → om.numel == 0compute_1d_workgroup_count(..., 0, ...). Confirm a 0-element dispatch is a safe no-op (Vulkan tolerates it); a quick assertion/comment would document the intent.

No blocking issues beyond #1, which I'd recommend addressing before this lands given it directly contradicts the "Mirrors Vulkan resize_slice_copy_node" claim in the description.
• Base: origin/gh/JulianCloudNTH/73/base

[ghstack-poisoned]
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence June 30, 2026 02:46 — with GitHub Actions Inactive
[ghstack-poisoned]

@SS-JIASS-JIA left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review automatically exported from Phabricator review in Meta.

[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:28 — with GitHub Actions Inactive
[ghstack-poisoned]
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 20:52 — with GitHub Actions Inactive
@ghost
ghost temporarily deployed to cadence July 3, 2026 21:20 — with GitHub Actions Inactive
@meta-codesync
meta-codesyncBot merged commit fb80edf into gh/JulianCloudNTH/73/baseJul 4, 2026
181 of 183 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JulianCloudNTH/73/head branch July 4, 2026 17:06
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
ghost pushed a commit that referenced this pull request Jul 4, 2026
…t/end)
Pull Request resolved: #20581
**Make `slice_copy` support a dynamic gather range so the RoPE-freqs slice `[input_pos : input_pos + S]` works under one dynamic graph.**
**Problem:** the static slice handler read `start` via a scalar reader that throws on a SymInt and ignored `end` (output length baked AOT). The RoPE-freqs slice uses a SymInt `input_pos` for start and a live S for the range, so the static op could neither build nor resize for it.
**Solution:** read start/end as possibly-dynamic SymInts and add a resize hook that recomputes the gather offset and live output length each step.
- Before: `start` is a static scalar (SymInt throws); `end` ignored; output length fixed at the serialized max.
- After: `start`/`end` read via a SymInt-aware reader; a hook recomputes `out[dim] = (end - start + step - 1) / step`, rewrites `out_meta`/`in_meta`/`params` UBOs + the dispatch count, and sets the output `cur_dims`.
**Implementation:**
- Hook registered on the `start`/`end` value-ids when they are SymInts and on the input tensor always (inert until resized, so a static slice is byte-identical).
- Output/input `TensorMeta` rebuilt from live dims; `dim`/`step` stay static.
- Keep the uniforms alive via `own_uniform_buffer` so the hook can rewrite them.
- Mirrors Vulkan `resize_slice_copy_node`.
**Constraints:** fp32-only; `dim`/`step` static; numerics + layout unchanged; inert on a static graph.
NOTE (stacking): this diff sits on top of the in-review `slice_copy` op (D108793168); rebase onto it once that op lands on master.
Co-authored-with: Claude Code.
ghstack-source-id: 399812835
@exported-using-ghexport
Differential Revision: [D109906092](https://our.internmc.facebook.com/intern/diff/D109906092/)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@SS-JIA@nil-is-all@JCNTH