Implement Portable bucketize - #20287

Open
Gallinator wants to merge 37 commits into
pytorch:mainfrom
Gallinator:bucketize
Open

Implement Portable bucketize#20287
Gallinator wants to merge 37 commits into
pytorch:mainfrom
Gallinator:bucketize

Conversation

@Gallinator

@GallinatorGallinator commented Jun 15, 2026

Copy link
Copy Markdown

Fixes#20270

Summary

Add portable scalar and tensor bucketize operator, based on the PyTorch implementation.
Key differences from the PyTorch implementation:

  • Implemented only bucketize, no searchsorted
  • Boundaries tensor size is not checked against output type, as SizesType is an int32_t and must be 1D
  • In the original implementation the input scalar is wrapped into a tensor, here two sets of functions are used instead
  • Supports both NHWC and NCHW as it is a pointwise kernel and there is no sorter argument. This avoids creating contiguous tensors
  • Stayed as close to the PyTorch version as possible, should a searchsorted implementation be added in the future

Release notes: ops & kernels

Test plan

C++ tests are extensive and include failure and edge cases tests.
Python tests compare the output of an exported module to the PyTorch implementation only for normal and edge cases, similar to #15893.

cc @larryliu0820@manuelcandales

@pytorch-bot

pytorch-botBot commented Jun 15, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20287

Note: Links to docs will display an error until the docs builds have been completed.

⚠️ 13 Awaiting Approval

As of commit ae0c4ae with merge base 9ffbe36 (image):

AWAITING APPROVAL - The following workflows need approval before CI can run:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla

meta-claBot commented Jun 15, 2026

Copy link
Copy Markdown

Hi @Gallinator!

Thank you for your pull request and welcome to our community.

Action Required

In order to merge any pull request (code, docs, etc.), we require contributors to sign our Contributor License Agreement, and we don't seem to have one on file for you.

Process

In order for us to review and merge your suggested changes, please sign at https://code.facebook.com/cla. If you are contributing on behalf of someone else (eg your employer), the individual CLA may not be sufficient and your employer may need to sign the corporate CLA.

Once the CLA is signed, our tooling will perform checks and validations. Afterwards, the pull request will be tagged with CLA signed. The tagging process may take up to 1 hour after signing. Please give it that time before contacting us about it.

If you have received this in error or have any questions, please contact us at cla@meta.com. Thanks!

@linux-foundation-easycla

linux-foundation-easyclaBot commented Jun 15, 2026

Copy link
Copy Markdown

CLA Signed
The committers listed above are authorized under a signed CLA.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 16, 2026
@meta-cla

meta-claBot commented Jun 16, 2026

Copy link
Copy Markdown

Thank you for signing our Contributor License Agreement. We can now accept your code for this (and any) Meta Open Source project. Thanks!

@Gallinator

Copy link
Copy Markdown
Author

@pytorchbot label "release notes: ops & kernels"

@pytorch-botpytorch-botBot added the release notes: ops & kernels Changes to the opset and any new / changed kernel implementations label Jun 16, 2026
@Gallinator
Gallinator marked this pull request as ready for review June 17, 2026 11:33
@Gasoonjia

Copy link
Copy Markdown
Contributor

@claude review this PR

@claude

claudeBot commented Jun 22, 2026

Copy link
Copy Markdown

Claude finished @Gasoonjia's task in 5m 32s —— View job


PR Review: Implement Portable bucketize

  • Read changed files (op_bucketize.cpp, tests, functions.yaml, build files)
  • Analyze correctness, conventions, and tests

Nice work — the kernel is clean, the C++ test coverage is genuinely thorough (edge cases, inf, empty, dtype matrix, failure paths), and the choice to support both NHWC/NCHW via the matching-dim-order check is correct: since self and out are required to have the same shape and the same dim order, iterating linear physical indices i for both loads and stores maps to the same logical element, so no contiguous copy is needed. Below are the issues I found.

Bugs

1. test_bucketize_scalar_out_int32 doesn't test int32kernels/test/test_bucketize.py:126-131
The "int32" test constructs BucketizeModule(out_int32=False, ...), so it's an exact duplicate of the int64 test and never exercises the int32 path. Set out_int32=True.

deftest_bucketize_scalar_out_int32(self):
model=BucketizeModule(out_int32=False, right=False) # should be True

Fix this →

Conventions

2. Missing BSD license headerop_bucketize.cpp, op_bucketize_test.cpp, test_bucketize.py
All three new files start directly with includes/code. Every other file under kernels/portable/cpu/ and kernels/test/ carries the standard Copyright (c) Meta Platforms, Inc. and affiliates. ... BSD-style license header. Please add it.

3. bucketize_tensor_out requires a pre-sized out instead of resizing itop_bucketize.cpp:177-180
The kernel checks tensors_have_same_shape(self, out) but never calls resize_tensor(out, self.sizes()). Most portable out-variant kernels (e.g. op_clamp.cpp:129) resize out to the input size so they work with dynamic / upper-bound shapes where the planner allocates out at a max size. With a hard same-shape check this op will fail under dynamic shapes. Consider:

ET_KERNEL_CHECK(
context, resize_tensor(out, self.sizes()) == Error::Ok, InvalidArgument, out);

(For the scalar variant, the out.dim() == 0 check is fine as-is.)

Minor / style

  • op_bucketize.cpp:69,108const bool& right passes a bool by const reference; pass by value (bool right).
  • op_bucketize.cpp:80,116boundaries.sizes().back() can be boundaries.numel() since boundaries is already validated as 1D; slightly clearer.
  • test_bucketize.py — every test docstring says "Test bucketize.Tensor_out", including the scalar tests (lines 120, 126, etc.). Worth correcting the scalar ones.

Things I checked that look correct

  • Compute-type promotion (promote_type_with_scalar for scalar, promoteTypes for tensor) and the REALHBF16 dtype switch.
  • out_int32 ↔ output dtype validation, complex-type and N-D-boundaries rejection paths.
  • Binary-search cus_lower_bound/cus_upper_bound for right/left semantics, empty boundaries (returns 0), and out-of-range / inf inputs — all covered by tests.
  • Build registration in functions.yaml, targets.bzl, and CMakeLists.txt.

Overall this is close to merge-ready; only item #1 is a real functional gap, and #2 is required for CI conventions.
· branch bucketize

@Gasoonjia

Copy link
Copy Markdown
Contributor

Thanks for contribution! I have a question left in the issue mind check that?

Comment threadkernels/test/test_bucketize.py Outdated
# LICENSE file in the root directory of this source tree.

"""
Test for bucketize operations in ExecuTorch.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

plz do not put export test script under kernels/test. Only the runtime test should be fine.

@GallinatorGallinatorJun 22, 2026

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@Gasoonjia So I have to remove test_bucketize.py ? I see that in #15893 they added an export test as well

@Gasoonjia

Copy link
Copy Markdown
Contributor

hi @Gallinator can you try to solve the ci issue?
also add test into https://github.com/pytorch/executorch/blob/main/kernels/test/targets.bzl plz

@Gallinator

Copy link
Copy Markdown
Author

@Gasoonjia CI should be fixed now

Signed-off-by: Gallinator <lcgallinator4@gmail.com>
 WRAP_TO_ATEN does not support Scalar inputs
@Gallinator

Gallinator commented Jul 29, 2026

Copy link
Copy Markdown
Author

@Gasoonjia I removed the scalar input AOT registration for now. If all CI tests pass I will open a new issue for WRAP_TO_ATEN not supporting Scalar inputs

@nil-is-all

Copy link
Copy Markdown
Contributor

@Gasoonjia I removed the scalar input AOT registration for now. If all CI tests pass I will open a new issue for WRAP_TO_ATEN not supporting Scalar inputs

Thanks, running CI again now

@Gallinator

Gallinator commented Aug 4, 2026

Copy link
Copy Markdown
Author

@nil-is-all@Gasoonjia The latest commit solved the 2 previously failing checks, the other failures seem related to the CI setup.
Is this enough to merge the PR?

@nil-is-all

Copy link
Copy Markdown
Contributor

@claude review and check CI failures

@claude

claudeBot commented Aug 5, 2026

Copy link
Copy Markdown

Claude finished @nil-is-all's task in 2m 18s —— View job


Review & CI check

  • Review PR code changes
  • Check CI failures
  • Summarize findings

CI failures — infrastructure, not this PR

I pulled the job logs. Every failure in the pull workflow (and the presets/docs jobs) fails at the same step: Checkout repository (pytorch/test-infra@), with:

API rate limit exceeded for installation. ... request ID DDDA:D8B26:... 2026-08-03 17:52:23 UTC
##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '.../pytorch/executorch'. No such file or directory

The bash/No such file or directory errors are the downstream cascade: the test-infra checkout never completed, so the working directory the later steps expect was never created. None of the failing jobs ever reached a build or test step. The Test QNN Windows MSVC build failure is flagged by Dr. CI as unclassified because that workflow didn't run on the merge base — it's not attributable to this PR either.

Meanwhile the jobs that did get a runner — Lint, Test XNNPACK, Test ARM, Test QNN, Test CoreML, Test Vulkan, Test WebGPU, periodic, test-mcu-cortex-m-backend, and the arm64 test-models-linux-basic variants — all passed. So this is a transient GitHub API rate-limit / infra outage on 2026-08-03, not a code problem. A re-run once the rate limit clears should turn these green — nothing to fix in the PR for CI.

Code review

The code is in good shape and the earlier review feedback has all been incorporated:

  • ✅ BSD license headers present on all new files (op_bucketize.cpp, op_bucketize_test.cpp).
  • bucketize_tensor_out now calls resize_tensor(out, self.sizes()) (op_bucketize.cpp:185-189) instead of a hard same-shape check — works under dynamic/upper-bound shapes.
  • right is now passed by value (bool) rather than const bool&.
  • ✅ The problematic Python export test was removed per @Gasoonjia's request; C++ coverage remains extensive.
  • ✅ Scalar runtime kernel is still registered in functions.yaml (bucketize.Scalar_out); only the AOT-test registration was dropped, matching the WRAP_TO_ATEN Scalar limitation described in the PR.

A couple of small, non-blocking observations:

  1. op_bucketize.cpp:14 still #includes .../thread_parallel_interface.h, but commit 89b9427 removed that target from the Buck deps in op_registration_util.bzl. parallel_for is transitively available via elementwise_util, so this likely compiles, but relying on a transitive include for a header you directly use is fragile under strict-deps. Consider either dropping the direct include (since elementwise_util.h provides parallel_for) or restoring the explicit dep so the include and dep agree. Other parallel_for users (op_amin.cpp) get it via reduce_util.

  2. Minor:boundaries.numel() is used directly for bd_end now (good, since boundaries is validated 1D) — consistent and clear.

Net: the code is essentially merge-ready; the red CI is a test-infra rate-limit outage and a re-run should clear it.
· branch bucketize

@Gallinator

Copy link
Copy Markdown
Author

@Gasoonjia@nil-is-all Is there anything left to improve? I see that #21869 lists other use cases for this kernel

@Gallinator

Copy link
Copy Markdown
Author

@Gasoonjia@nil-is-all CI should be fixed now, could you try to run it again?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.module: kernelsIssues related to kernel libraries and utilities, and code under kernels/release notes: ops & kernelsChanges to the opset and any new / changed kernel implementations

Projects

None yet

Development

Successfully merging this pull request may close these issues.

kernel 'aten::bucketize.Tensor_out' not found.

3 participants

@Gallinator@Gasoonjia@nil-is-all
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Implement Portable bucketize - #20287

Open
Gallinator wants to merge 37 commits into
pytorch:mainfrom
Gallinator:bucketize
Open

Implement Portable bucketize#20287
Gallinator wants to merge 37 commits into
pytorch:mainfrom
Gallinator:bucketize

Conversation

@Gallinator

@GallinatorGallinator commented Jun 15, 2026

Copy link
Copy Markdown

Fixes#20270

Summary

Add portable scalar and tensor bucketize operator, based on the PyTorch implementation.
Key differences from the PyTorch implementation:

  • Implemented only bucketize, no searchsorted
  • Boundaries tensor size is not checked against output type, as SizesType is an int32_t and must be 1D
  • In the original implementation the input scalar is wrapped into a tensor, here two sets of functions are used instead
  • Supports both NHWC and NCHW as it is a pointwise kernel and there is no sorter argument. This avoids creating contiguous tensors
  • Stayed as close to the PyTorch version as possible, should a searchsorted implementation be added in the future

Release notes: ops & kernels

Test plan

C++ tests are extensive and include failure and edge cases tests.
Python tests compare the output of an exported module to the PyTorch implementation only for normal and edge cases, similar to #15893.

cc @larryliu0820@manuelcandales

@pytorch-bot

pytorch-botBot commented Jun 15, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20287

Note: Links to docs will display an error until the docs builds have been completed.

⚠️ 13 Awaiting Approval

As of commit ae0c4ae with merge base 9ffbe36 (image):

AWAITING APPROVAL - The following workflows need approval before CI can run:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla

meta-claBot commented Jun 15, 2026

Copy link
Copy Markdown

Hi @Gallinator!

Thank you for your pull request and welcome to our community.

Action Required

In order to merge any pull request (code, docs, etc.), we require contributors to sign our Contributor License Agreement, and we don't seem to have one on file for you.

Process

In order for us to review and merge your suggested changes, please sign at https://code.facebook.com/cla. If you are contributing on behalf of someone else (eg your employer), the individual CLA may not be sufficient and your employer may need to sign the corporate CLA.

Once the CLA is signed, our tooling will perform checks and validations. Afterwards, the pull request will be tagged with CLA signed. The tagging process may take up to 1 hour after signing. Please give it that time before contacting us about it.

If you have received this in error or have any questions, please contact us at cla@meta.com. Thanks!

@linux-foundation-easycla

linux-foundation-easyclaBot commented Jun 15, 2026

Copy link
Copy Markdown

CLA Signed
The committers listed above are authorized under a signed CLA.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 16, 2026
@meta-cla

meta-claBot commented Jun 16, 2026

Copy link
Copy Markdown

Thank you for signing our Contributor License Agreement. We can now accept your code for this (and any) Meta Open Source project. Thanks!

@Gallinator

Copy link
Copy Markdown
Author

@pytorchbot label "release notes: ops & kernels"

@pytorch-botpytorch-botBot added the release notes: ops & kernels Changes to the opset and any new / changed kernel implementations label Jun 16, 2026
@Gallinator
Gallinator marked this pull request as ready for review June 17, 2026 11:33
@Gasoonjia

Copy link
Copy Markdown
Contributor

@claude review this PR

@claude

claudeBot commented Jun 22, 2026

Copy link
Copy Markdown

Claude finished @Gasoonjia's task in 5m 32s —— View job


PR Review: Implement Portable bucketize

  • Read changed files (op_bucketize.cpp, tests, functions.yaml, build files)
  • Analyze correctness, conventions, and tests

Nice work — the kernel is clean, the C++ test coverage is genuinely thorough (edge cases, inf, empty, dtype matrix, failure paths), and the choice to support both NHWC/NCHW via the matching-dim-order check is correct: since self and out are required to have the same shape and the same dim order, iterating linear physical indices i for both loads and stores maps to the same logical element, so no contiguous copy is needed. Below are the issues I found.

Bugs

1. test_bucketize_scalar_out_int32 doesn't test int32kernels/test/test_bucketize.py:126-131
The "int32" test constructs BucketizeModule(out_int32=False, ...), so it's an exact duplicate of the int64 test and never exercises the int32 path. Set out_int32=True.

deftest_bucketize_scalar_out_int32(self):
model=BucketizeModule(out_int32=False, right=False) # should be True

Fix this →

Conventions

2. Missing BSD license headerop_bucketize.cpp, op_bucketize_test.cpp, test_bucketize.py
All three new files start directly with includes/code. Every other file under kernels/portable/cpu/ and kernels/test/ carries the standard Copyright (c) Meta Platforms, Inc. and affiliates. ... BSD-style license header. Please add it.

3. bucketize_tensor_out requires a pre-sized out instead of resizing itop_bucketize.cpp:177-180
The kernel checks tensors_have_same_shape(self, out) but never calls resize_tensor(out, self.sizes()). Most portable out-variant kernels (e.g. op_clamp.cpp:129) resize out to the input size so they work with dynamic / upper-bound shapes where the planner allocates out at a max size. With a hard same-shape check this op will fail under dynamic shapes. Consider:

ET_KERNEL_CHECK(
context, resize_tensor(out, self.sizes()) == Error::Ok, InvalidArgument, out);

(For the scalar variant, the out.dim() == 0 check is fine as-is.)

Minor / style

  • op_bucketize.cpp:69,108const bool& right passes a bool by const reference; pass by value (bool right).
  • op_bucketize.cpp:80,116boundaries.sizes().back() can be boundaries.numel() since boundaries is already validated as 1D; slightly clearer.
  • test_bucketize.py — every test docstring says "Test bucketize.Tensor_out", including the scalar tests (lines 120, 126, etc.). Worth correcting the scalar ones.

Things I checked that look correct

  • Compute-type promotion (promote_type_with_scalar for scalar, promoteTypes for tensor) and the REALHBF16 dtype switch.
  • out_int32 ↔ output dtype validation, complex-type and N-D-boundaries rejection paths.
  • Binary-search cus_lower_bound/cus_upper_bound for right/left semantics, empty boundaries (returns 0), and out-of-range / inf inputs — all covered by tests.
  • Build registration in functions.yaml, targets.bzl, and CMakeLists.txt.

Overall this is close to merge-ready; only item #1 is a real functional gap, and #2 is required for CI conventions.
· branch bucketize

@Gasoonjia

Copy link
Copy Markdown
Contributor

Thanks for contribution! I have a question left in the issue mind check that?

Comment threadkernels/test/test_bucketize.py Outdated
# LICENSE file in the root directory of this source tree.

"""
Test for bucketize operations in ExecuTorch.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

plz do not put export test script under kernels/test. Only the runtime test should be fine.

@GallinatorGallinatorJun 22, 2026

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@Gasoonjia So I have to remove test_bucketize.py ? I see that in #15893 they added an export test as well

@Gasoonjia

Copy link
Copy Markdown
Contributor

hi @Gallinator can you try to solve the ci issue?
also add test into https://github.com/pytorch/executorch/blob/main/kernels/test/targets.bzl plz

@Gallinator

Copy link
Copy Markdown
Author

@Gasoonjia CI should be fixed now

Signed-off-by: Gallinator <lcgallinator4@gmail.com>
 WRAP_TO_ATEN does not support Scalar inputs
@Gallinator

Gallinator commented Jul 29, 2026

Copy link
Copy Markdown
Author

@Gasoonjia I removed the scalar input AOT registration for now. If all CI tests pass I will open a new issue for WRAP_TO_ATEN not supporting Scalar inputs

@nil-is-all

Copy link
Copy Markdown
Contributor

@Gasoonjia I removed the scalar input AOT registration for now. If all CI tests pass I will open a new issue for WRAP_TO_ATEN not supporting Scalar inputs

Thanks, running CI again now

@Gallinator

Gallinator commented Aug 4, 2026

Copy link
Copy Markdown
Author

@nil-is-all@Gasoonjia The latest commit solved the 2 previously failing checks, the other failures seem related to the CI setup.
Is this enough to merge the PR?

@nil-is-all

Copy link
Copy Markdown
Contributor

@claude review and check CI failures

@claude

claudeBot commented Aug 5, 2026

Copy link
Copy Markdown

Claude finished @nil-is-all's task in 2m 18s —— View job


Review & CI check

  • Review PR code changes
  • Check CI failures
  • Summarize findings

CI failures — infrastructure, not this PR

I pulled the job logs. Every failure in the pull workflow (and the presets/docs jobs) fails at the same step: Checkout repository (pytorch/test-infra@), with:

API rate limit exceeded for installation. ... request ID DDDA:D8B26:... 2026-08-03 17:52:23 UTC
##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '.../pytorch/executorch'. No such file or directory

The bash/No such file or directory errors are the downstream cascade: the test-infra checkout never completed, so the working directory the later steps expect was never created. None of the failing jobs ever reached a build or test step. The Test QNN Windows MSVC build failure is flagged by Dr. CI as unclassified because that workflow didn't run on the merge base — it's not attributable to this PR either.

Meanwhile the jobs that did get a runner — Lint, Test XNNPACK, Test ARM, Test QNN, Test CoreML, Test Vulkan, Test WebGPU, periodic, test-mcu-cortex-m-backend, and the arm64 test-models-linux-basic variants — all passed. So this is a transient GitHub API rate-limit / infra outage on 2026-08-03, not a code problem. A re-run once the rate limit clears should turn these green — nothing to fix in the PR for CI.

Code review

The code is in good shape and the earlier review feedback has all been incorporated:

  • ✅ BSD license headers present on all new files (op_bucketize.cpp, op_bucketize_test.cpp).
  • bucketize_tensor_out now calls resize_tensor(out, self.sizes()) (op_bucketize.cpp:185-189) instead of a hard same-shape check — works under dynamic/upper-bound shapes.
  • right is now passed by value (bool) rather than const bool&.
  • ✅ The problematic Python export test was removed per @Gasoonjia's request; C++ coverage remains extensive.
  • ✅ Scalar runtime kernel is still registered in functions.yaml (bucketize.Scalar_out); only the AOT-test registration was dropped, matching the WRAP_TO_ATEN Scalar limitation described in the PR.

A couple of small, non-blocking observations:

  1. op_bucketize.cpp:14 still #includes .../thread_parallel_interface.h, but commit 89b9427 removed that target from the Buck deps in op_registration_util.bzl. parallel_for is transitively available via elementwise_util, so this likely compiles, but relying on a transitive include for a header you directly use is fragile under strict-deps. Consider either dropping the direct include (since elementwise_util.h provides parallel_for) or restoring the explicit dep so the include and dep agree. Other parallel_for users (op_amin.cpp) get it via reduce_util.

  2. Minor:boundaries.numel() is used directly for bd_end now (good, since boundaries is validated 1D) — consistent and clear.

Net: the code is essentially merge-ready; the red CI is a test-infra rate-limit outage and a re-run should clear it.
· branch bucketize

@Gallinator

Copy link
Copy Markdown
Author

@Gasoonjia@nil-is-all Is there anything left to improve? I see that #21869 lists other use cases for this kernel

@Gallinator

Copy link
Copy Markdown
Author

@Gasoonjia@nil-is-all CI should be fixed now, could you try to run it again?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.module: kernelsIssues related to kernel libraries and utilities, and code under kernels/release notes: ops & kernelsChanges to the opset and any new / changed kernel implementations

Projects

None yet

Development

Successfully merging this pull request may close these issues.

kernel 'aten::bucketize.Tensor_out' not found.

3 participants

@Gallinator@Gasoonjia@nil-is-all
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Implement Portable bucketize - #20287

Open
Gallinator wants to merge 37 commits into
pytorch:mainfrom
Gallinator:bucketize
Open

Implement Portable bucketize#20287
Gallinator wants to merge 37 commits into
pytorch:mainfrom
Gallinator:bucketize

Conversation

@Gallinator

@GallinatorGallinator commented Jun 15, 2026

Copy link
Copy Markdown

Fixes#20270

Summary

Add portable scalar and tensor bucketize operator, based on the PyTorch implementation.
Key differences from the PyTorch implementation:

  • Implemented only bucketize, no searchsorted
  • Boundaries tensor size is not checked against output type, as SizesType is an int32_t and must be 1D
  • In the original implementation the input scalar is wrapped into a tensor, here two sets of functions are used instead
  • Supports both NHWC and NCHW as it is a pointwise kernel and there is no sorter argument. This avoids creating contiguous tensors
  • Stayed as close to the PyTorch version as possible, should a searchsorted implementation be added in the future

Release notes: ops & kernels

Test plan

C++ tests are extensive and include failure and edge cases tests.
Python tests compare the output of an exported module to the PyTorch implementation only for normal and edge cases, similar to #15893.

cc @larryliu0820@manuelcandales

@pytorch-bot

pytorch-botBot commented Jun 15, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20287

Note: Links to docs will display an error until the docs builds have been completed.

⚠️ 13 Awaiting Approval

As of commit ae0c4ae with merge base 9ffbe36 (image):

AWAITING APPROVAL - The following workflows need approval before CI can run:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla

meta-claBot commented Jun 15, 2026

Copy link
Copy Markdown

Hi @Gallinator!

Thank you for your pull request and welcome to our community.

Action Required

In order to merge any pull request (code, docs, etc.), we require contributors to sign our Contributor License Agreement, and we don't seem to have one on file for you.

Process

In order for us to review and merge your suggested changes, please sign at https://code.facebook.com/cla. If you are contributing on behalf of someone else (eg your employer), the individual CLA may not be sufficient and your employer may need to sign the corporate CLA.

Once the CLA is signed, our tooling will perform checks and validations. Afterwards, the pull request will be tagged with CLA signed. The tagging process may take up to 1 hour after signing. Please give it that time before contacting us about it.

If you have received this in error or have any questions, please contact us at cla@meta.com. Thanks!

@linux-foundation-easycla

linux-foundation-easyclaBot commented Jun 15, 2026

Copy link
Copy Markdown

CLA Signed
The committers listed above are authorized under a signed CLA.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 16, 2026
@meta-cla

meta-claBot commented Jun 16, 2026

Copy link
Copy Markdown

Thank you for signing our Contributor License Agreement. We can now accept your code for this (and any) Meta Open Source project. Thanks!

@Gallinator

Copy link
Copy Markdown
Author

@pytorchbot label "release notes: ops & kernels"

@pytorch-botpytorch-botBot added the release notes: ops & kernels Changes to the opset and any new / changed kernel implementations label Jun 16, 2026
@Gallinator
Gallinator marked this pull request as ready for review June 17, 2026 11:33
@Gasoonjia

Copy link
Copy Markdown
Contributor

@claude review this PR

@claude

claudeBot commented Jun 22, 2026

Copy link
Copy Markdown

Claude finished @Gasoonjia's task in 5m 32s —— View job


PR Review: Implement Portable bucketize

  • Read changed files (op_bucketize.cpp, tests, functions.yaml, build files)
  • Analyze correctness, conventions, and tests

Nice work — the kernel is clean, the C++ test coverage is genuinely thorough (edge cases, inf, empty, dtype matrix, failure paths), and the choice to support both NHWC/NCHW via the matching-dim-order check is correct: since self and out are required to have the same shape and the same dim order, iterating linear physical indices i for both loads and stores maps to the same logical element, so no contiguous copy is needed. Below are the issues I found.

Bugs

1. test_bucketize_scalar_out_int32 doesn't test int32kernels/test/test_bucketize.py:126-131
The "int32" test constructs BucketizeModule(out_int32=False, ...), so it's an exact duplicate of the int64 test and never exercises the int32 path. Set out_int32=True.

deftest_bucketize_scalar_out_int32(self):
model=BucketizeModule(out_int32=False, right=False) # should be True

Fix this →

Conventions

2. Missing BSD license headerop_bucketize.cpp, op_bucketize_test.cpp, test_bucketize.py
All three new files start directly with includes/code. Every other file under kernels/portable/cpu/ and kernels/test/ carries the standard Copyright (c) Meta Platforms, Inc. and affiliates. ... BSD-style license header. Please add it.

3. bucketize_tensor_out requires a pre-sized out instead of resizing itop_bucketize.cpp:177-180
The kernel checks tensors_have_same_shape(self, out) but never calls resize_tensor(out, self.sizes()). Most portable out-variant kernels (e.g. op_clamp.cpp:129) resize out to the input size so they work with dynamic / upper-bound shapes where the planner allocates out at a max size. With a hard same-shape check this op will fail under dynamic shapes. Consider:

ET_KERNEL_CHECK(
context, resize_tensor(out, self.sizes()) == Error::Ok, InvalidArgument, out);

(For the scalar variant, the out.dim() == 0 check is fine as-is.)

Minor / style

  • op_bucketize.cpp:69,108const bool& right passes a bool by const reference; pass by value (bool right).
  • op_bucketize.cpp:80,116boundaries.sizes().back() can be boundaries.numel() since boundaries is already validated as 1D; slightly clearer.
  • test_bucketize.py — every test docstring says "Test bucketize.Tensor_out", including the scalar tests (lines 120, 126, etc.). Worth correcting the scalar ones.

Things I checked that look correct

  • Compute-type promotion (promote_type_with_scalar for scalar, promoteTypes for tensor) and the REALHBF16 dtype switch.
  • out_int32 ↔ output dtype validation, complex-type and N-D-boundaries rejection paths.
  • Binary-search cus_lower_bound/cus_upper_bound for right/left semantics, empty boundaries (returns 0), and out-of-range / inf inputs — all covered by tests.
  • Build registration in functions.yaml, targets.bzl, and CMakeLists.txt.

Overall this is close to merge-ready; only item #1 is a real functional gap, and #2 is required for CI conventions.
· branch bucketize

@Gasoonjia

Copy link
Copy Markdown
Contributor

Thanks for contribution! I have a question left in the issue mind check that?

Comment threadkernels/test/test_bucketize.py Outdated
# LICENSE file in the root directory of this source tree.

"""
Test for bucketize operations in ExecuTorch.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

plz do not put export test script under kernels/test. Only the runtime test should be fine.

@GallinatorGallinatorJun 22, 2026

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@Gasoonjia So I have to remove test_bucketize.py ? I see that in #15893 they added an export test as well

@Gasoonjia

Copy link
Copy Markdown
Contributor

hi @Gallinator can you try to solve the ci issue?
also add test into https://github.com/pytorch/executorch/blob/main/kernels/test/targets.bzl plz

@Gallinator

Copy link
Copy Markdown
Author

@Gasoonjia CI should be fixed now

Signed-off-by: Gallinator <lcgallinator4@gmail.com>
 WRAP_TO_ATEN does not support Scalar inputs
@Gallinator

Gallinator commented Jul 29, 2026

Copy link
Copy Markdown
Author

@Gasoonjia I removed the scalar input AOT registration for now. If all CI tests pass I will open a new issue for WRAP_TO_ATEN not supporting Scalar inputs

@nil-is-all

Copy link
Copy Markdown
Contributor

@Gasoonjia I removed the scalar input AOT registration for now. If all CI tests pass I will open a new issue for WRAP_TO_ATEN not supporting Scalar inputs

Thanks, running CI again now

@Gallinator

Gallinator commented Aug 4, 2026

Copy link
Copy Markdown
Author

@nil-is-all@Gasoonjia The latest commit solved the 2 previously failing checks, the other failures seem related to the CI setup.
Is this enough to merge the PR?

@nil-is-all

Copy link
Copy Markdown
Contributor

@claude review and check CI failures

@claude

claudeBot commented Aug 5, 2026

Copy link
Copy Markdown

Claude finished @nil-is-all's task in 2m 18s —— View job


Review & CI check

  • Review PR code changes
  • Check CI failures
  • Summarize findings

CI failures — infrastructure, not this PR

I pulled the job logs. Every failure in the pull workflow (and the presets/docs jobs) fails at the same step: Checkout repository (pytorch/test-infra@), with:

API rate limit exceeded for installation. ... request ID DDDA:D8B26:... 2026-08-03 17:52:23 UTC
##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '.../pytorch/executorch'. No such file or directory

The bash/No such file or directory errors are the downstream cascade: the test-infra checkout never completed, so the working directory the later steps expect was never created. None of the failing jobs ever reached a build or test step. The Test QNN Windows MSVC build failure is flagged by Dr. CI as unclassified because that workflow didn't run on the merge base — it's not attributable to this PR either.

Meanwhile the jobs that did get a runner — Lint, Test XNNPACK, Test ARM, Test QNN, Test CoreML, Test Vulkan, Test WebGPU, periodic, test-mcu-cortex-m-backend, and the arm64 test-models-linux-basic variants — all passed. So this is a transient GitHub API rate-limit / infra outage on 2026-08-03, not a code problem. A re-run once the rate limit clears should turn these green — nothing to fix in the PR for CI.

Code review

The code is in good shape and the earlier review feedback has all been incorporated:

  • ✅ BSD license headers present on all new files (op_bucketize.cpp, op_bucketize_test.cpp).
  • bucketize_tensor_out now calls resize_tensor(out, self.sizes()) (op_bucketize.cpp:185-189) instead of a hard same-shape check — works under dynamic/upper-bound shapes.
  • right is now passed by value (bool) rather than const bool&.
  • ✅ The problematic Python export test was removed per @Gasoonjia's request; C++ coverage remains extensive.
  • ✅ Scalar runtime kernel is still registered in functions.yaml (bucketize.Scalar_out); only the AOT-test registration was dropped, matching the WRAP_TO_ATEN Scalar limitation described in the PR.

A couple of small, non-blocking observations:

  1. op_bucketize.cpp:14 still #includes .../thread_parallel_interface.h, but commit 89b9427 removed that target from the Buck deps in op_registration_util.bzl. parallel_for is transitively available via elementwise_util, so this likely compiles, but relying on a transitive include for a header you directly use is fragile under strict-deps. Consider either dropping the direct include (since elementwise_util.h provides parallel_for) or restoring the explicit dep so the include and dep agree. Other parallel_for users (op_amin.cpp) get it via reduce_util.

  2. Minor:boundaries.numel() is used directly for bd_end now (good, since boundaries is validated 1D) — consistent and clear.

Net: the code is essentially merge-ready; the red CI is a test-infra rate-limit outage and a re-run should clear it.
· branch bucketize

@Gallinator

Copy link
Copy Markdown
Author

@Gasoonjia@nil-is-all Is there anything left to improve? I see that #21869 lists other use cases for this kernel

@Gallinator

Copy link
Copy Markdown
Author

@Gasoonjia@nil-is-all CI should be fixed now, could you try to run it again?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.module: kernelsIssues related to kernel libraries and utilities, and code under kernels/release notes: ops & kernelsChanges to the opset and any new / changed kernel implementations

Projects

None yet

Development

Successfully merging this pull request may close these issues.

kernel 'aten::bucketize.Tensor_out' not found.

3 participants

@Gallinator@Gasoonjia@nil-is-all
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Implement Portable bucketize - #20287

Open
Gallinator wants to merge 37 commits into
pytorch:mainfrom
Gallinator:bucketize
Open

Implement Portable bucketize#20287
Gallinator wants to merge 37 commits into
pytorch:mainfrom
Gallinator:bucketize

Conversation

@Gallinator

@GallinatorGallinator commented Jun 15, 2026

Copy link
Copy Markdown

Fixes#20270

Summary

Add portable scalar and tensor bucketize operator, based on the PyTorch implementation.
Key differences from the PyTorch implementation:

  • Implemented only bucketize, no searchsorted
  • Boundaries tensor size is not checked against output type, as SizesType is an int32_t and must be 1D
  • In the original implementation the input scalar is wrapped into a tensor, here two sets of functions are used instead
  • Supports both NHWC and NCHW as it is a pointwise kernel and there is no sorter argument. This avoids creating contiguous tensors
  • Stayed as close to the PyTorch version as possible, should a searchsorted implementation be added in the future

Release notes: ops & kernels

Test plan

C++ tests are extensive and include failure and edge cases tests.
Python tests compare the output of an exported module to the PyTorch implementation only for normal and edge cases, similar to #15893.

cc @larryliu0820@manuelcandales

@pytorch-bot

pytorch-botBot commented Jun 15, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20287

Note: Links to docs will display an error until the docs builds have been completed.

⚠️ 13 Awaiting Approval

As of commit ae0c4ae with merge base 9ffbe36 (image):

AWAITING APPROVAL - The following workflows need approval before CI can run:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla

meta-claBot commented Jun 15, 2026

Copy link
Copy Markdown

Hi @Gallinator!

Thank you for your pull request and welcome to our community.

Action Required

In order to merge any pull request (code, docs, etc.), we require contributors to sign our Contributor License Agreement, and we don't seem to have one on file for you.

Process

In order for us to review and merge your suggested changes, please sign at https://code.facebook.com/cla. If you are contributing on behalf of someone else (eg your employer), the individual CLA may not be sufficient and your employer may need to sign the corporate CLA.

Once the CLA is signed, our tooling will perform checks and validations. Afterwards, the pull request will be tagged with CLA signed. The tagging process may take up to 1 hour after signing. Please give it that time before contacting us about it.

If you have received this in error or have any questions, please contact us at cla@meta.com. Thanks!

@linux-foundation-easycla

linux-foundation-easyclaBot commented Jun 15, 2026

Copy link
Copy Markdown

CLA Signed
The committers listed above are authorized under a signed CLA.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 16, 2026
@meta-cla

meta-claBot commented Jun 16, 2026

Copy link
Copy Markdown

Thank you for signing our Contributor License Agreement. We can now accept your code for this (and any) Meta Open Source project. Thanks!

@Gallinator

Copy link
Copy Markdown
Author

@pytorchbot label "release notes: ops & kernels"

@pytorch-botpytorch-botBot added the release notes: ops & kernels Changes to the opset and any new / changed kernel implementations label Jun 16, 2026
@Gallinator
Gallinator marked this pull request as ready for review June 17, 2026 11:33
@Gasoonjia

Copy link
Copy Markdown
Contributor

@claude review this PR

@claude

claudeBot commented Jun 22, 2026

Copy link
Copy Markdown

Claude finished @Gasoonjia's task in 5m 32s —— View job


PR Review: Implement Portable bucketize

  • Read changed files (op_bucketize.cpp, tests, functions.yaml, build files)
  • Analyze correctness, conventions, and tests

Nice work — the kernel is clean, the C++ test coverage is genuinely thorough (edge cases, inf, empty, dtype matrix, failure paths), and the choice to support both NHWC/NCHW via the matching-dim-order check is correct: since self and out are required to have the same shape and the same dim order, iterating linear physical indices i for both loads and stores maps to the same logical element, so no contiguous copy is needed. Below are the issues I found.

Bugs

1. test_bucketize_scalar_out_int32 doesn't test int32kernels/test/test_bucketize.py:126-131
The "int32" test constructs BucketizeModule(out_int32=False, ...), so it's an exact duplicate of the int64 test and never exercises the int32 path. Set out_int32=True.

deftest_bucketize_scalar_out_int32(self):
model=BucketizeModule(out_int32=False, right=False) # should be True

Fix this →

Conventions

2. Missing BSD license headerop_bucketize.cpp, op_bucketize_test.cpp, test_bucketize.py
All three new files start directly with includes/code. Every other file under kernels/portable/cpu/ and kernels/test/ carries the standard Copyright (c) Meta Platforms, Inc. and affiliates. ... BSD-style license header. Please add it.

3. bucketize_tensor_out requires a pre-sized out instead of resizing itop_bucketize.cpp:177-180
The kernel checks tensors_have_same_shape(self, out) but never calls resize_tensor(out, self.sizes()). Most portable out-variant kernels (e.g. op_clamp.cpp:129) resize out to the input size so they work with dynamic / upper-bound shapes where the planner allocates out at a max size. With a hard same-shape check this op will fail under dynamic shapes. Consider:

ET_KERNEL_CHECK(
context, resize_tensor(out, self.sizes()) == Error::Ok, InvalidArgument, out);

(For the scalar variant, the out.dim() == 0 check is fine as-is.)

Minor / style

  • op_bucketize.cpp:69,108const bool& right passes a bool by const reference; pass by value (bool right).
  • op_bucketize.cpp:80,116boundaries.sizes().back() can be boundaries.numel() since boundaries is already validated as 1D; slightly clearer.
  • test_bucketize.py — every test docstring says "Test bucketize.Tensor_out", including the scalar tests (lines 120, 126, etc.). Worth correcting the scalar ones.

Things I checked that look correct

  • Compute-type promotion (promote_type_with_scalar for scalar, promoteTypes for tensor) and the REALHBF16 dtype switch.
  • out_int32 ↔ output dtype validation, complex-type and N-D-boundaries rejection paths.
  • Binary-search cus_lower_bound/cus_upper_bound for right/left semantics, empty boundaries (returns 0), and out-of-range / inf inputs — all covered by tests.
  • Build registration in functions.yaml, targets.bzl, and CMakeLists.txt.

Overall this is close to merge-ready; only item #1 is a real functional gap, and #2 is required for CI conventions.
· branch bucketize

@Gasoonjia

Copy link
Copy Markdown
Contributor

Thanks for contribution! I have a question left in the issue mind check that?

Comment threadkernels/test/test_bucketize.py Outdated
# LICENSE file in the root directory of this source tree.

"""
Test for bucketize operations in ExecuTorch.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

plz do not put export test script under kernels/test. Only the runtime test should be fine.

@GallinatorGallinatorJun 22, 2026

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@Gasoonjia So I have to remove test_bucketize.py ? I see that in #15893 they added an export test as well

@Gasoonjia

Copy link
Copy Markdown
Contributor

hi @Gallinator can you try to solve the ci issue?
also add test into https://github.com/pytorch/executorch/blob/main/kernels/test/targets.bzl plz

@Gallinator

Copy link
Copy Markdown
Author

@Gasoonjia CI should be fixed now

Signed-off-by: Gallinator <lcgallinator4@gmail.com>
 WRAP_TO_ATEN does not support Scalar inputs
@Gallinator

Gallinator commented Jul 29, 2026

Copy link
Copy Markdown
Author

@Gasoonjia I removed the scalar input AOT registration for now. If all CI tests pass I will open a new issue for WRAP_TO_ATEN not supporting Scalar inputs

@nil-is-all

Copy link
Copy Markdown
Contributor

@Gasoonjia I removed the scalar input AOT registration for now. If all CI tests pass I will open a new issue for WRAP_TO_ATEN not supporting Scalar inputs

Thanks, running CI again now

@Gallinator

Gallinator commented Aug 4, 2026

Copy link
Copy Markdown
Author

@nil-is-all@Gasoonjia The latest commit solved the 2 previously failing checks, the other failures seem related to the CI setup.
Is this enough to merge the PR?

@nil-is-all

Copy link
Copy Markdown
Contributor

@claude review and check CI failures

@claude

claudeBot commented Aug 5, 2026

Copy link
Copy Markdown

Claude finished @nil-is-all's task in 2m 18s —— View job


Review & CI check

  • Review PR code changes
  • Check CI failures
  • Summarize findings

CI failures — infrastructure, not this PR

I pulled the job logs. Every failure in the pull workflow (and the presets/docs jobs) fails at the same step: Checkout repository (pytorch/test-infra@), with:

API rate limit exceeded for installation. ... request ID DDDA:D8B26:... 2026-08-03 17:52:23 UTC
##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '.../pytorch/executorch'. No such file or directory

The bash/No such file or directory errors are the downstream cascade: the test-infra checkout never completed, so the working directory the later steps expect was never created. None of the failing jobs ever reached a build or test step. The Test QNN Windows MSVC build failure is flagged by Dr. CI as unclassified because that workflow didn't run on the merge base — it's not attributable to this PR either.

Meanwhile the jobs that did get a runner — Lint, Test XNNPACK, Test ARM, Test QNN, Test CoreML, Test Vulkan, Test WebGPU, periodic, test-mcu-cortex-m-backend, and the arm64 test-models-linux-basic variants — all passed. So this is a transient GitHub API rate-limit / infra outage on 2026-08-03, not a code problem. A re-run once the rate limit clears should turn these green — nothing to fix in the PR for CI.

Code review

The code is in good shape and the earlier review feedback has all been incorporated:

  • ✅ BSD license headers present on all new files (op_bucketize.cpp, op_bucketize_test.cpp).
  • bucketize_tensor_out now calls resize_tensor(out, self.sizes()) (op_bucketize.cpp:185-189) instead of a hard same-shape check — works under dynamic/upper-bound shapes.
  • right is now passed by value (bool) rather than const bool&.
  • ✅ The problematic Python export test was removed per @Gasoonjia's request; C++ coverage remains extensive.
  • ✅ Scalar runtime kernel is still registered in functions.yaml (bucketize.Scalar_out); only the AOT-test registration was dropped, matching the WRAP_TO_ATEN Scalar limitation described in the PR.

A couple of small, non-blocking observations:

  1. op_bucketize.cpp:14 still #includes .../thread_parallel_interface.h, but commit 89b9427 removed that target from the Buck deps in op_registration_util.bzl. parallel_for is transitively available via elementwise_util, so this likely compiles, but relying on a transitive include for a header you directly use is fragile under strict-deps. Consider either dropping the direct include (since elementwise_util.h provides parallel_for) or restoring the explicit dep so the include and dep agree. Other parallel_for users (op_amin.cpp) get it via reduce_util.

  2. Minor:boundaries.numel() is used directly for bd_end now (good, since boundaries is validated 1D) — consistent and clear.

Net: the code is essentially merge-ready; the red CI is a test-infra rate-limit outage and a re-run should clear it.
· branch bucketize

@Gallinator

Copy link
Copy Markdown
Author

@Gasoonjia@nil-is-all Is there anything left to improve? I see that #21869 lists other use cases for this kernel

@Gallinator

Copy link
Copy Markdown
Author

@Gasoonjia@nil-is-all CI should be fixed now, could you try to run it again?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.module: kernelsIssues related to kernel libraries and utilities, and code under kernels/release notes: ops & kernelsChanges to the opset and any new / changed kernel implementations

Projects

None yet

Development

Successfully merging this pull request may close these issues.

kernel 'aten::bucketize.Tensor_out' not found.

3 participants

@Gallinator@Gasoonjia@nil-is-all
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Implement Portable bucketize - #20287

Open
Gallinator wants to merge 37 commits into
pytorch:mainfrom
Gallinator:bucketize
Open

Implement Portable bucketize#20287
Gallinator wants to merge 37 commits into
pytorch:mainfrom
Gallinator:bucketize

Conversation

@Gallinator

@GallinatorGallinator commented Jun 15, 2026

Copy link
Copy Markdown

Fixes#20270

Summary

Add portable scalar and tensor bucketize operator, based on the PyTorch implementation.
Key differences from the PyTorch implementation:

  • Implemented only bucketize, no searchsorted
  • Boundaries tensor size is not checked against output type, as SizesType is an int32_t and must be 1D
  • In the original implementation the input scalar is wrapped into a tensor, here two sets of functions are used instead
  • Supports both NHWC and NCHW as it is a pointwise kernel and there is no sorter argument. This avoids creating contiguous tensors
  • Stayed as close to the PyTorch version as possible, should a searchsorted implementation be added in the future

Release notes: ops & kernels

Test plan

C++ tests are extensive and include failure and edge cases tests.
Python tests compare the output of an exported module to the PyTorch implementation only for normal and edge cases, similar to #15893.

cc @larryliu0820@manuelcandales

@pytorch-bot

pytorch-botBot commented Jun 15, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20287

Note: Links to docs will display an error until the docs builds have been completed.

⚠️ 13 Awaiting Approval

As of commit ae0c4ae with merge base 9ffbe36 (image):

AWAITING APPROVAL - The following workflows need approval before CI can run:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla

meta-claBot commented Jun 15, 2026

Copy link
Copy Markdown

Hi @Gallinator!

Thank you for your pull request and welcome to our community.

Action Required

In order to merge any pull request (code, docs, etc.), we require contributors to sign our Contributor License Agreement, and we don't seem to have one on file for you.

Process

In order for us to review and merge your suggested changes, please sign at https://code.facebook.com/cla. If you are contributing on behalf of someone else (eg your employer), the individual CLA may not be sufficient and your employer may need to sign the corporate CLA.

Once the CLA is signed, our tooling will perform checks and validations. Afterwards, the pull request will be tagged with CLA signed. The tagging process may take up to 1 hour after signing. Please give it that time before contacting us about it.

If you have received this in error or have any questions, please contact us at cla@meta.com. Thanks!

@linux-foundation-easycla

linux-foundation-easyclaBot commented Jun 15, 2026

Copy link
Copy Markdown

CLA Signed
The committers listed above are authorized under a signed CLA.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 16, 2026
@meta-cla

meta-claBot commented Jun 16, 2026

Copy link
Copy Markdown

Thank you for signing our Contributor License Agreement. We can now accept your code for this (and any) Meta Open Source project. Thanks!

@Gallinator

Copy link
Copy Markdown
Author

@pytorchbot label "release notes: ops & kernels"

@pytorch-botpytorch-botBot added the release notes: ops & kernels Changes to the opset and any new / changed kernel implementations label Jun 16, 2026
@Gallinator
Gallinator marked this pull request as ready for review June 17, 2026 11:33
@Gasoonjia

Copy link
Copy Markdown
Contributor

@claude review this PR

@claude

claudeBot commented Jun 22, 2026

Copy link
Copy Markdown

Claude finished @Gasoonjia's task in 5m 32s —— View job


PR Review: Implement Portable bucketize

  • Read changed files (op_bucketize.cpp, tests, functions.yaml, build files)
  • Analyze correctness, conventions, and tests

Nice work — the kernel is clean, the C++ test coverage is genuinely thorough (edge cases, inf, empty, dtype matrix, failure paths), and the choice to support both NHWC/NCHW via the matching-dim-order check is correct: since self and out are required to have the same shape and the same dim order, iterating linear physical indices i for both loads and stores maps to the same logical element, so no contiguous copy is needed. Below are the issues I found.

Bugs

1. test_bucketize_scalar_out_int32 doesn't test int32kernels/test/test_bucketize.py:126-131
The "int32" test constructs BucketizeModule(out_int32=False, ...), so it's an exact duplicate of the int64 test and never exercises the int32 path. Set out_int32=True.

deftest_bucketize_scalar_out_int32(self):
model=BucketizeModule(out_int32=False, right=False) # should be True

Fix this →

Conventions

2. Missing BSD license headerop_bucketize.cpp, op_bucketize_test.cpp, test_bucketize.py
All three new files start directly with includes/code. Every other file under kernels/portable/cpu/ and kernels/test/ carries the standard Copyright (c) Meta Platforms, Inc. and affiliates. ... BSD-style license header. Please add it.

3. bucketize_tensor_out requires a pre-sized out instead of resizing itop_bucketize.cpp:177-180
The kernel checks tensors_have_same_shape(self, out) but never calls resize_tensor(out, self.sizes()). Most portable out-variant kernels (e.g. op_clamp.cpp:129) resize out to the input size so they work with dynamic / upper-bound shapes where the planner allocates out at a max size. With a hard same-shape check this op will fail under dynamic shapes. Consider:

ET_KERNEL_CHECK(
context, resize_tensor(out, self.sizes()) == Error::Ok, InvalidArgument, out);

(For the scalar variant, the out.dim() == 0 check is fine as-is.)

Minor / style

  • op_bucketize.cpp:69,108const bool& right passes a bool by const reference; pass by value (bool right).
  • op_bucketize.cpp:80,116boundaries.sizes().back() can be boundaries.numel() since boundaries is already validated as 1D; slightly clearer.
  • test_bucketize.py — every test docstring says "Test bucketize.Tensor_out", including the scalar tests (lines 120, 126, etc.). Worth correcting the scalar ones.

Things I checked that look correct

  • Compute-type promotion (promote_type_with_scalar for scalar, promoteTypes for tensor) and the REALHBF16 dtype switch.
  • out_int32 ↔ output dtype validation, complex-type and N-D-boundaries rejection paths.
  • Binary-search cus_lower_bound/cus_upper_bound for right/left semantics, empty boundaries (returns 0), and out-of-range / inf inputs — all covered by tests.
  • Build registration in functions.yaml, targets.bzl, and CMakeLists.txt.

Overall this is close to merge-ready; only item #1 is a real functional gap, and #2 is required for CI conventions.
· branch bucketize

@Gasoonjia

Copy link
Copy Markdown
Contributor

Thanks for contribution! I have a question left in the issue mind check that?

Comment threadkernels/test/test_bucketize.py Outdated
# LICENSE file in the root directory of this source tree.

"""
Test for bucketize operations in ExecuTorch.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

plz do not put export test script under kernels/test. Only the runtime test should be fine.

@GallinatorGallinatorJun 22, 2026

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@Gasoonjia So I have to remove test_bucketize.py ? I see that in #15893 they added an export test as well

@Gasoonjia

Copy link
Copy Markdown
Contributor

hi @Gallinator can you try to solve the ci issue?
also add test into https://github.com/pytorch/executorch/blob/main/kernels/test/targets.bzl plz

@Gallinator

Copy link
Copy Markdown
Author

@Gasoonjia CI should be fixed now

Signed-off-by: Gallinator <lcgallinator4@gmail.com>
 WRAP_TO_ATEN does not support Scalar inputs
@Gallinator

Gallinator commented Jul 29, 2026

Copy link
Copy Markdown
Author

@Gasoonjia I removed the scalar input AOT registration for now. If all CI tests pass I will open a new issue for WRAP_TO_ATEN not supporting Scalar inputs

@nil-is-all

Copy link
Copy Markdown
Contributor

@Gasoonjia I removed the scalar input AOT registration for now. If all CI tests pass I will open a new issue for WRAP_TO_ATEN not supporting Scalar inputs

Thanks, running CI again now

@Gallinator

Gallinator commented Aug 4, 2026

Copy link
Copy Markdown
Author

@nil-is-all@Gasoonjia The latest commit solved the 2 previously failing checks, the other failures seem related to the CI setup.
Is this enough to merge the PR?

@nil-is-all

Copy link
Copy Markdown
Contributor

@claude review and check CI failures

@claude

claudeBot commented Aug 5, 2026

Copy link
Copy Markdown

Claude finished @nil-is-all's task in 2m 18s —— View job


Review & CI check

  • Review PR code changes
  • Check CI failures
  • Summarize findings

CI failures — infrastructure, not this PR

I pulled the job logs. Every failure in the pull workflow (and the presets/docs jobs) fails at the same step: Checkout repository (pytorch/test-infra@), with:

API rate limit exceeded for installation. ... request ID DDDA:D8B26:... 2026-08-03 17:52:23 UTC
##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '.../pytorch/executorch'. No such file or directory

The bash/No such file or directory errors are the downstream cascade: the test-infra checkout never completed, so the working directory the later steps expect was never created. None of the failing jobs ever reached a build or test step. The Test QNN Windows MSVC build failure is flagged by Dr. CI as unclassified because that workflow didn't run on the merge base — it's not attributable to this PR either.

Meanwhile the jobs that did get a runner — Lint, Test XNNPACK, Test ARM, Test QNN, Test CoreML, Test Vulkan, Test WebGPU, periodic, test-mcu-cortex-m-backend, and the arm64 test-models-linux-basic variants — all passed. So this is a transient GitHub API rate-limit / infra outage on 2026-08-03, not a code problem. A re-run once the rate limit clears should turn these green — nothing to fix in the PR for CI.

Code review

The code is in good shape and the earlier review feedback has all been incorporated:

  • ✅ BSD license headers present on all new files (op_bucketize.cpp, op_bucketize_test.cpp).
  • bucketize_tensor_out now calls resize_tensor(out, self.sizes()) (op_bucketize.cpp:185-189) instead of a hard same-shape check — works under dynamic/upper-bound shapes.
  • right is now passed by value (bool) rather than const bool&.
  • ✅ The problematic Python export test was removed per @Gasoonjia's request; C++ coverage remains extensive.
  • ✅ Scalar runtime kernel is still registered in functions.yaml (bucketize.Scalar_out); only the AOT-test registration was dropped, matching the WRAP_TO_ATEN Scalar limitation described in the PR.

A couple of small, non-blocking observations:

  1. op_bucketize.cpp:14 still #includes .../thread_parallel_interface.h, but commit 89b9427 removed that target from the Buck deps in op_registration_util.bzl. parallel_for is transitively available via elementwise_util, so this likely compiles, but relying on a transitive include for a header you directly use is fragile under strict-deps. Consider either dropping the direct include (since elementwise_util.h provides parallel_for) or restoring the explicit dep so the include and dep agree. Other parallel_for users (op_amin.cpp) get it via reduce_util.

  2. Minor:boundaries.numel() is used directly for bd_end now (good, since boundaries is validated 1D) — consistent and clear.

Net: the code is essentially merge-ready; the red CI is a test-infra rate-limit outage and a re-run should clear it.
· branch bucketize

@Gallinator

Copy link
Copy Markdown
Author

@Gasoonjia@nil-is-all Is there anything left to improve? I see that #21869 lists other use cases for this kernel

@Gallinator

Copy link
Copy Markdown
Author

@Gasoonjia@nil-is-all CI should be fixed now, could you try to run it again?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.module: kernelsIssues related to kernel libraries and utilities, and code under kernels/release notes: ops & kernelsChanges to the opset and any new / changed kernel implementations

Projects

None yet

Development

Successfully merging this pull request may close these issues.

kernel 'aten::bucketize.Tensor_out' not found.

3 participants

@Gallinator@Gasoonjia@nil-is-all
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Implement Portable bucketize - #20287

Open
Gallinator wants to merge 37 commits into
pytorch:mainfrom
Gallinator:bucketize
Open

Implement Portable bucketize#20287
Gallinator wants to merge 37 commits into
pytorch:mainfrom
Gallinator:bucketize

Conversation

@Gallinator

@GallinatorGallinator commented Jun 15, 2026

Copy link
Copy Markdown

Fixes#20270

Summary

Add portable scalar and tensor bucketize operator, based on the PyTorch implementation.
Key differences from the PyTorch implementation:

  • Implemented only bucketize, no searchsorted
  • Boundaries tensor size is not checked against output type, as SizesType is an int32_t and must be 1D
  • In the original implementation the input scalar is wrapped into a tensor, here two sets of functions are used instead
  • Supports both NHWC and NCHW as it is a pointwise kernel and there is no sorter argument. This avoids creating contiguous tensors
  • Stayed as close to the PyTorch version as possible, should a searchsorted implementation be added in the future

Release notes: ops & kernels

Test plan

C++ tests are extensive and include failure and edge cases tests.
Python tests compare the output of an exported module to the PyTorch implementation only for normal and edge cases, similar to #15893.

cc @larryliu0820@manuelcandales

@pytorch-bot

pytorch-botBot commented Jun 15, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20287

Note: Links to docs will display an error until the docs builds have been completed.

⚠️ 13 Awaiting Approval

As of commit ae0c4ae with merge base 9ffbe36 (image):

AWAITING APPROVAL - The following workflows need approval before CI can run:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla

meta-claBot commented Jun 15, 2026

Copy link
Copy Markdown

Hi @Gallinator!

Thank you for your pull request and welcome to our community.

Action Required

In order to merge any pull request (code, docs, etc.), we require contributors to sign our Contributor License Agreement, and we don't seem to have one on file for you.

Process

In order for us to review and merge your suggested changes, please sign at https://code.facebook.com/cla. If you are contributing on behalf of someone else (eg your employer), the individual CLA may not be sufficient and your employer may need to sign the corporate CLA.

Once the CLA is signed, our tooling will perform checks and validations. Afterwards, the pull request will be tagged with CLA signed. The tagging process may take up to 1 hour after signing. Please give it that time before contacting us about it.

If you have received this in error or have any questions, please contact us at cla@meta.com. Thanks!

@linux-foundation-easycla

linux-foundation-easyclaBot commented Jun 15, 2026

Copy link
Copy Markdown

CLA Signed
The committers listed above are authorized under a signed CLA.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 16, 2026
@meta-cla

meta-claBot commented Jun 16, 2026

Copy link
Copy Markdown

Thank you for signing our Contributor License Agreement. We can now accept your code for this (and any) Meta Open Source project. Thanks!

@Gallinator

Copy link
Copy Markdown
Author

@pytorchbot label "release notes: ops & kernels"

@pytorch-botpytorch-botBot added the release notes: ops & kernels Changes to the opset and any new / changed kernel implementations label Jun 16, 2026
@Gallinator
Gallinator marked this pull request as ready for review June 17, 2026 11:33
@Gasoonjia

Copy link
Copy Markdown
Contributor

@claude review this PR

@claude

claudeBot commented Jun 22, 2026

Copy link
Copy Markdown

Claude finished @Gasoonjia's task in 5m 32s —— View job


PR Review: Implement Portable bucketize

  • Read changed files (op_bucketize.cpp, tests, functions.yaml, build files)
  • Analyze correctness, conventions, and tests

Nice work — the kernel is clean, the C++ test coverage is genuinely thorough (edge cases, inf, empty, dtype matrix, failure paths), and the choice to support both NHWC/NCHW via the matching-dim-order check is correct: since self and out are required to have the same shape and the same dim order, iterating linear physical indices i for both loads and stores maps to the same logical element, so no contiguous copy is needed. Below are the issues I found.

Bugs

1. test_bucketize_scalar_out_int32 doesn't test int32kernels/test/test_bucketize.py:126-131
The "int32" test constructs BucketizeModule(out_int32=False, ...), so it's an exact duplicate of the int64 test and never exercises the int32 path. Set out_int32=True.

deftest_bucketize_scalar_out_int32(self):
model=BucketizeModule(out_int32=False, right=False) # should be True

Fix this →

Conventions

2. Missing BSD license headerop_bucketize.cpp, op_bucketize_test.cpp, test_bucketize.py
All three new files start directly with includes/code. Every other file under kernels/portable/cpu/ and kernels/test/ carries the standard Copyright (c) Meta Platforms, Inc. and affiliates. ... BSD-style license header. Please add it.

3. bucketize_tensor_out requires a pre-sized out instead of resizing itop_bucketize.cpp:177-180
The kernel checks tensors_have_same_shape(self, out) but never calls resize_tensor(out, self.sizes()). Most portable out-variant kernels (e.g. op_clamp.cpp:129) resize out to the input size so they work with dynamic / upper-bound shapes where the planner allocates out at a max size. With a hard same-shape check this op will fail under dynamic shapes. Consider:

ET_KERNEL_CHECK(
context, resize_tensor(out, self.sizes()) == Error::Ok, InvalidArgument, out);

(For the scalar variant, the out.dim() == 0 check is fine as-is.)

Minor / style

  • op_bucketize.cpp:69,108const bool& right passes a bool by const reference; pass by value (bool right).
  • op_bucketize.cpp:80,116boundaries.sizes().back() can be boundaries.numel() since boundaries is already validated as 1D; slightly clearer.
  • test_bucketize.py — every test docstring says "Test bucketize.Tensor_out", including the scalar tests (lines 120, 126, etc.). Worth correcting the scalar ones.

Things I checked that look correct

  • Compute-type promotion (promote_type_with_scalar for scalar, promoteTypes for tensor) and the REALHBF16 dtype switch.
  • out_int32 ↔ output dtype validation, complex-type and N-D-boundaries rejection paths.
  • Binary-search cus_lower_bound/cus_upper_bound for right/left semantics, empty boundaries (returns 0), and out-of-range / inf inputs — all covered by tests.
  • Build registration in functions.yaml, targets.bzl, and CMakeLists.txt.

Overall this is close to merge-ready; only item #1 is a real functional gap, and #2 is required for CI conventions.
· branch bucketize

@Gasoonjia

Copy link
Copy Markdown
Contributor

Thanks for contribution! I have a question left in the issue mind check that?

Comment threadkernels/test/test_bucketize.py Outdated
# LICENSE file in the root directory of this source tree.

"""
Test for bucketize operations in ExecuTorch.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

plz do not put export test script under kernels/test. Only the runtime test should be fine.

@GallinatorGallinatorJun 22, 2026

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@Gasoonjia So I have to remove test_bucketize.py ? I see that in #15893 they added an export test as well

@Gasoonjia

Copy link
Copy Markdown
Contributor

hi @Gallinator can you try to solve the ci issue?
also add test into https://github.com/pytorch/executorch/blob/main/kernels/test/targets.bzl plz

@Gallinator

Copy link
Copy Markdown
Author

@Gasoonjia CI should be fixed now

Signed-off-by: Gallinator <lcgallinator4@gmail.com>
 WRAP_TO_ATEN does not support Scalar inputs
@Gallinator

Gallinator commented Jul 29, 2026

Copy link
Copy Markdown
Author

@Gasoonjia I removed the scalar input AOT registration for now. If all CI tests pass I will open a new issue for WRAP_TO_ATEN not supporting Scalar inputs

@nil-is-all

Copy link
Copy Markdown
Contributor

@Gasoonjia I removed the scalar input AOT registration for now. If all CI tests pass I will open a new issue for WRAP_TO_ATEN not supporting Scalar inputs

Thanks, running CI again now

@Gallinator

Gallinator commented Aug 4, 2026

Copy link
Copy Markdown
Author

@nil-is-all@Gasoonjia The latest commit solved the 2 previously failing checks, the other failures seem related to the CI setup.
Is this enough to merge the PR?

@nil-is-all

Copy link
Copy Markdown
Contributor

@claude review and check CI failures

@claude

claudeBot commented Aug 5, 2026

Copy link
Copy Markdown

Claude finished @nil-is-all's task in 2m 18s —— View job


Review & CI check

  • Review PR code changes
  • Check CI failures
  • Summarize findings

CI failures — infrastructure, not this PR

I pulled the job logs. Every failure in the pull workflow (and the presets/docs jobs) fails at the same step: Checkout repository (pytorch/test-infra@), with:

API rate limit exceeded for installation. ... request ID DDDA:D8B26:... 2026-08-03 17:52:23 UTC
##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '.../pytorch/executorch'. No such file or directory

The bash/No such file or directory errors are the downstream cascade: the test-infra checkout never completed, so the working directory the later steps expect was never created. None of the failing jobs ever reached a build or test step. The Test QNN Windows MSVC build failure is flagged by Dr. CI as unclassified because that workflow didn't run on the merge base — it's not attributable to this PR either.

Meanwhile the jobs that did get a runner — Lint, Test XNNPACK, Test ARM, Test QNN, Test CoreML, Test Vulkan, Test WebGPU, periodic, test-mcu-cortex-m-backend, and the arm64 test-models-linux-basic variants — all passed. So this is a transient GitHub API rate-limit / infra outage on 2026-08-03, not a code problem. A re-run once the rate limit clears should turn these green — nothing to fix in the PR for CI.

Code review

The code is in good shape and the earlier review feedback has all been incorporated:

  • ✅ BSD license headers present on all new files (op_bucketize.cpp, op_bucketize_test.cpp).
  • bucketize_tensor_out now calls resize_tensor(out, self.sizes()) (op_bucketize.cpp:185-189) instead of a hard same-shape check — works under dynamic/upper-bound shapes.
  • right is now passed by value (bool) rather than const bool&.
  • ✅ The problematic Python export test was removed per @Gasoonjia's request; C++ coverage remains extensive.
  • ✅ Scalar runtime kernel is still registered in functions.yaml (bucketize.Scalar_out); only the AOT-test registration was dropped, matching the WRAP_TO_ATEN Scalar limitation described in the PR.

A couple of small, non-blocking observations:

  1. op_bucketize.cpp:14 still #includes .../thread_parallel_interface.h, but commit 89b9427 removed that target from the Buck deps in op_registration_util.bzl. parallel_for is transitively available via elementwise_util, so this likely compiles, but relying on a transitive include for a header you directly use is fragile under strict-deps. Consider either dropping the direct include (since elementwise_util.h provides parallel_for) or restoring the explicit dep so the include and dep agree. Other parallel_for users (op_amin.cpp) get it via reduce_util.

  2. Minor:boundaries.numel() is used directly for bd_end now (good, since boundaries is validated 1D) — consistent and clear.

Net: the code is essentially merge-ready; the red CI is a test-infra rate-limit outage and a re-run should clear it.
· branch bucketize

@Gallinator

Copy link
Copy Markdown
Author

@Gasoonjia@nil-is-all Is there anything left to improve? I see that #21869 lists other use cases for this kernel

@Gallinator

Copy link
Copy Markdown
Author

@Gasoonjia@nil-is-all CI should be fixed now, could you try to run it again?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.module: kernelsIssues related to kernel libraries and utilities, and code under kernels/release notes: ops & kernelsChanges to the opset and any new / changed kernel implementations

Projects

None yet

Development

Successfully merging this pull request may close these issues.

kernel 'aten::bucketize.Tensor_out' not found.

3 participants

@Gallinator@Gasoonjia@nil-is-all
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Implement Portable bucketize - #20287

Open
Gallinator wants to merge 37 commits into
pytorch:mainfrom
Gallinator:bucketize
Open

Implement Portable bucketize#20287
Gallinator wants to merge 37 commits into
pytorch:mainfrom
Gallinator:bucketize

Conversation

@Gallinator

@GallinatorGallinator commented Jun 15, 2026

Copy link
Copy Markdown

Fixes#20270

Summary

Add portable scalar and tensor bucketize operator, based on the PyTorch implementation.
Key differences from the PyTorch implementation:

  • Implemented only bucketize, no searchsorted
  • Boundaries tensor size is not checked against output type, as SizesType is an int32_t and must be 1D
  • In the original implementation the input scalar is wrapped into a tensor, here two sets of functions are used instead
  • Supports both NHWC and NCHW as it is a pointwise kernel and there is no sorter argument. This avoids creating contiguous tensors
  • Stayed as close to the PyTorch version as possible, should a searchsorted implementation be added in the future

Release notes: ops & kernels

Test plan

C++ tests are extensive and include failure and edge cases tests.
Python tests compare the output of an exported module to the PyTorch implementation only for normal and edge cases, similar to #15893.

cc @larryliu0820@manuelcandales

@pytorch-bot

pytorch-botBot commented Jun 15, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20287

Note: Links to docs will display an error until the docs builds have been completed.

⚠️ 13 Awaiting Approval

As of commit ae0c4ae with merge base 9ffbe36 (image):

AWAITING APPROVAL - The following workflows need approval before CI can run:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla

meta-claBot commented Jun 15, 2026

Copy link
Copy Markdown

Hi @Gallinator!

Thank you for your pull request and welcome to our community.

Action Required

In order to merge any pull request (code, docs, etc.), we require contributors to sign our Contributor License Agreement, and we don't seem to have one on file for you.

Process

In order for us to review and merge your suggested changes, please sign at https://code.facebook.com/cla. If you are contributing on behalf of someone else (eg your employer), the individual CLA may not be sufficient and your employer may need to sign the corporate CLA.

Once the CLA is signed, our tooling will perform checks and validations. Afterwards, the pull request will be tagged with CLA signed. The tagging process may take up to 1 hour after signing. Please give it that time before contacting us about it.

If you have received this in error or have any questions, please contact us at cla@meta.com. Thanks!

@linux-foundation-easycla

linux-foundation-easyclaBot commented Jun 15, 2026

Copy link
Copy Markdown

CLA Signed
The committers listed above are authorized under a signed CLA.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 16, 2026
@meta-cla

meta-claBot commented Jun 16, 2026

Copy link
Copy Markdown

Thank you for signing our Contributor License Agreement. We can now accept your code for this (and any) Meta Open Source project. Thanks!

@Gallinator

Copy link
Copy Markdown
Author

@pytorchbot label "release notes: ops & kernels"

@pytorch-botpytorch-botBot added the release notes: ops & kernels Changes to the opset and any new / changed kernel implementations label Jun 16, 2026
@Gallinator
Gallinator marked this pull request as ready for review June 17, 2026 11:33
@Gasoonjia

Copy link
Copy Markdown
Contributor

@claude review this PR

@claude

claudeBot commented Jun 22, 2026

Copy link
Copy Markdown

Claude finished @Gasoonjia's task in 5m 32s —— View job


PR Review: Implement Portable bucketize

  • Read changed files (op_bucketize.cpp, tests, functions.yaml, build files)
  • Analyze correctness, conventions, and tests

Nice work — the kernel is clean, the C++ test coverage is genuinely thorough (edge cases, inf, empty, dtype matrix, failure paths), and the choice to support both NHWC/NCHW via the matching-dim-order check is correct: since self and out are required to have the same shape and the same dim order, iterating linear physical indices i for both loads and stores maps to the same logical element, so no contiguous copy is needed. Below are the issues I found.

Bugs

1. test_bucketize_scalar_out_int32 doesn't test int32kernels/test/test_bucketize.py:126-131
The "int32" test constructs BucketizeModule(out_int32=False, ...), so it's an exact duplicate of the int64 test and never exercises the int32 path. Set out_int32=True.

deftest_bucketize_scalar_out_int32(self):
model=BucketizeModule(out_int32=False, right=False) # should be True

Fix this →

Conventions

2. Missing BSD license headerop_bucketize.cpp, op_bucketize_test.cpp, test_bucketize.py
All three new files start directly with includes/code. Every other file under kernels/portable/cpu/ and kernels/test/ carries the standard Copyright (c) Meta Platforms, Inc. and affiliates. ... BSD-style license header. Please add it.

3. bucketize_tensor_out requires a pre-sized out instead of resizing itop_bucketize.cpp:177-180
The kernel checks tensors_have_same_shape(self, out) but never calls resize_tensor(out, self.sizes()). Most portable out-variant kernels (e.g. op_clamp.cpp:129) resize out to the input size so they work with dynamic / upper-bound shapes where the planner allocates out at a max size. With a hard same-shape check this op will fail under dynamic shapes. Consider:

ET_KERNEL_CHECK(
context, resize_tensor(out, self.sizes()) == Error::Ok, InvalidArgument, out);

(For the scalar variant, the out.dim() == 0 check is fine as-is.)

Minor / style

  • op_bucketize.cpp:69,108const bool& right passes a bool by const reference; pass by value (bool right).
  • op_bucketize.cpp:80,116boundaries.sizes().back() can be boundaries.numel() since boundaries is already validated as 1D; slightly clearer.
  • test_bucketize.py — every test docstring says "Test bucketize.Tensor_out", including the scalar tests (lines 120, 126, etc.). Worth correcting the scalar ones.

Things I checked that look correct

  • Compute-type promotion (promote_type_with_scalar for scalar, promoteTypes for tensor) and the REALHBF16 dtype switch.
  • out_int32 ↔ output dtype validation, complex-type and N-D-boundaries rejection paths.
  • Binary-search cus_lower_bound/cus_upper_bound for right/left semantics, empty boundaries (returns 0), and out-of-range / inf inputs — all covered by tests.
  • Build registration in functions.yaml, targets.bzl, and CMakeLists.txt.

Overall this is close to merge-ready; only item #1 is a real functional gap, and #2 is required for CI conventions.
· branch bucketize

@Gasoonjia

Copy link
Copy Markdown
Contributor

Thanks for contribution! I have a question left in the issue mind check that?

Comment threadkernels/test/test_bucketize.py Outdated
# LICENSE file in the root directory of this source tree.

"""
Test for bucketize operations in ExecuTorch.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

plz do not put export test script under kernels/test. Only the runtime test should be fine.

@GallinatorGallinatorJun 22, 2026

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@Gasoonjia So I have to remove test_bucketize.py ? I see that in #15893 they added an export test as well

@Gasoonjia

Copy link
Copy Markdown
Contributor

hi @Gallinator can you try to solve the ci issue?
also add test into https://github.com/pytorch/executorch/blob/main/kernels/test/targets.bzl plz

@Gallinator

Copy link
Copy Markdown
Author

@Gasoonjia CI should be fixed now

Signed-off-by: Gallinator <lcgallinator4@gmail.com>
 WRAP_TO_ATEN does not support Scalar inputs
@Gallinator

Gallinator commented Jul 29, 2026

Copy link
Copy Markdown
Author

@Gasoonjia I removed the scalar input AOT registration for now. If all CI tests pass I will open a new issue for WRAP_TO_ATEN not supporting Scalar inputs

@nil-is-all

Copy link
Copy Markdown
Contributor

@Gasoonjia I removed the scalar input AOT registration for now. If all CI tests pass I will open a new issue for WRAP_TO_ATEN not supporting Scalar inputs

Thanks, running CI again now

@Gallinator

Gallinator commented Aug 4, 2026

Copy link
Copy Markdown
Author

@nil-is-all@Gasoonjia The latest commit solved the 2 previously failing checks, the other failures seem related to the CI setup.
Is this enough to merge the PR?

@nil-is-all

Copy link
Copy Markdown
Contributor

@claude review and check CI failures

@claude

claudeBot commented Aug 5, 2026

Copy link
Copy Markdown

Claude finished @nil-is-all's task in 2m 18s —— View job


Review & CI check

  • Review PR code changes
  • Check CI failures
  • Summarize findings

CI failures — infrastructure, not this PR

I pulled the job logs. Every failure in the pull workflow (and the presets/docs jobs) fails at the same step: Checkout repository (pytorch/test-infra@), with:

API rate limit exceeded for installation. ... request ID DDDA:D8B26:... 2026-08-03 17:52:23 UTC
##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '.../pytorch/executorch'. No such file or directory

The bash/No such file or directory errors are the downstream cascade: the test-infra checkout never completed, so the working directory the later steps expect was never created. None of the failing jobs ever reached a build or test step. The Test QNN Windows MSVC build failure is flagged by Dr. CI as unclassified because that workflow didn't run on the merge base — it's not attributable to this PR either.

Meanwhile the jobs that did get a runner — Lint, Test XNNPACK, Test ARM, Test QNN, Test CoreML, Test Vulkan, Test WebGPU, periodic, test-mcu-cortex-m-backend, and the arm64 test-models-linux-basic variants — all passed. So this is a transient GitHub API rate-limit / infra outage on 2026-08-03, not a code problem. A re-run once the rate limit clears should turn these green — nothing to fix in the PR for CI.

Code review

The code is in good shape and the earlier review feedback has all been incorporated:

  • ✅ BSD license headers present on all new files (op_bucketize.cpp, op_bucketize_test.cpp).
  • bucketize_tensor_out now calls resize_tensor(out, self.sizes()) (op_bucketize.cpp:185-189) instead of a hard same-shape check — works under dynamic/upper-bound shapes.
  • right is now passed by value (bool) rather than const bool&.
  • ✅ The problematic Python export test was removed per @Gasoonjia's request; C++ coverage remains extensive.
  • ✅ Scalar runtime kernel is still registered in functions.yaml (bucketize.Scalar_out); only the AOT-test registration was dropped, matching the WRAP_TO_ATEN Scalar limitation described in the PR.

A couple of small, non-blocking observations:

  1. op_bucketize.cpp:14 still #includes .../thread_parallel_interface.h, but commit 89b9427 removed that target from the Buck deps in op_registration_util.bzl. parallel_for is transitively available via elementwise_util, so this likely compiles, but relying on a transitive include for a header you directly use is fragile under strict-deps. Consider either dropping the direct include (since elementwise_util.h provides parallel_for) or restoring the explicit dep so the include and dep agree. Other parallel_for users (op_amin.cpp) get it via reduce_util.

  2. Minor:boundaries.numel() is used directly for bd_end now (good, since boundaries is validated 1D) — consistent and clear.

Net: the code is essentially merge-ready; the red CI is a test-infra rate-limit outage and a re-run should clear it.
· branch bucketize

@Gallinator

Copy link
Copy Markdown
Author

@Gasoonjia@nil-is-all Is there anything left to improve? I see that #21869 lists other use cases for this kernel

@Gallinator

Copy link
Copy Markdown
Author

@Gasoonjia@nil-is-all CI should be fixed now, could you try to run it again?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.module: kernelsIssues related to kernel libraries and utilities, and code under kernels/release notes: ops & kernelsChanges to the opset and any new / changed kernel implementations

Projects

None yet

Development

Successfully merging this pull request may close these issues.

kernel 'aten::bucketize.Tensor_out' not found.

3 participants

@Gallinator@Gasoonjia@nil-is-all
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Implement Portable bucketize - #20287

Open
Gallinator wants to merge 37 commits into
pytorch:mainfrom
Gallinator:bucketize
Open

Implement Portable bucketize#20287
Gallinator wants to merge 37 commits into
pytorch:mainfrom
Gallinator:bucketize

Conversation

@Gallinator

@GallinatorGallinator commented Jun 15, 2026

Copy link
Copy Markdown

Fixes#20270

Summary

Add portable scalar and tensor bucketize operator, based on the PyTorch implementation.
Key differences from the PyTorch implementation:

  • Implemented only bucketize, no searchsorted
  • Boundaries tensor size is not checked against output type, as SizesType is an int32_t and must be 1D
  • In the original implementation the input scalar is wrapped into a tensor, here two sets of functions are used instead
  • Supports both NHWC and NCHW as it is a pointwise kernel and there is no sorter argument. This avoids creating contiguous tensors
  • Stayed as close to the PyTorch version as possible, should a searchsorted implementation be added in the future

Release notes: ops & kernels

Test plan

C++ tests are extensive and include failure and edge cases tests.
Python tests compare the output of an exported module to the PyTorch implementation only for normal and edge cases, similar to #15893.

cc @larryliu0820@manuelcandales

@pytorch-bot

pytorch-botBot commented Jun 15, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20287

Note: Links to docs will display an error until the docs builds have been completed.

⚠️ 13 Awaiting Approval

As of commit ae0c4ae with merge base 9ffbe36 (image):

AWAITING APPROVAL - The following workflows need approval before CI can run:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla

meta-claBot commented Jun 15, 2026

Copy link
Copy Markdown

Hi @Gallinator!

Thank you for your pull request and welcome to our community.

Action Required

In order to merge any pull request (code, docs, etc.), we require contributors to sign our Contributor License Agreement, and we don't seem to have one on file for you.

Process

In order for us to review and merge your suggested changes, please sign at https://code.facebook.com/cla. If you are contributing on behalf of someone else (eg your employer), the individual CLA may not be sufficient and your employer may need to sign the corporate CLA.

Once the CLA is signed, our tooling will perform checks and validations. Afterwards, the pull request will be tagged with CLA signed. The tagging process may take up to 1 hour after signing. Please give it that time before contacting us about it.

If you have received this in error or have any questions, please contact us at cla@meta.com. Thanks!

@linux-foundation-easycla

linux-foundation-easyclaBot commented Jun 15, 2026

Copy link
Copy Markdown

CLA Signed
The committers listed above are authorized under a signed CLA.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jun 16, 2026
@meta-cla

meta-claBot commented Jun 16, 2026

Copy link
Copy Markdown

Thank you for signing our Contributor License Agreement. We can now accept your code for this (and any) Meta Open Source project. Thanks!

@Gallinator

Copy link
Copy Markdown
Author

@pytorchbot label "release notes: ops & kernels"

@pytorch-botpytorch-botBot added the release notes: ops & kernels Changes to the opset and any new / changed kernel implementations label Jun 16, 2026
@Gallinator
Gallinator marked this pull request as ready for review June 17, 2026 11:33
@Gasoonjia

Copy link
Copy Markdown
Contributor

@claude review this PR

@claude

claudeBot commented Jun 22, 2026

Copy link
Copy Markdown

Claude finished @Gasoonjia's task in 5m 32s —— View job


PR Review: Implement Portable bucketize

  • Read changed files (op_bucketize.cpp, tests, functions.yaml, build files)
  • Analyze correctness, conventions, and tests

Nice work — the kernel is clean, the C++ test coverage is genuinely thorough (edge cases, inf, empty, dtype matrix, failure paths), and the choice to support both NHWC/NCHW via the matching-dim-order check is correct: since self and out are required to have the same shape and the same dim order, iterating linear physical indices i for both loads and stores maps to the same logical element, so no contiguous copy is needed. Below are the issues I found.

Bugs

1. test_bucketize_scalar_out_int32 doesn't test int32kernels/test/test_bucketize.py:126-131
The "int32" test constructs BucketizeModule(out_int32=False, ...), so it's an exact duplicate of the int64 test and never exercises the int32 path. Set out_int32=True.

deftest_bucketize_scalar_out_int32(self):
model=BucketizeModule(out_int32=False, right=False) # should be True

Fix this →

Conventions

2. Missing BSD license headerop_bucketize.cpp, op_bucketize_test.cpp, test_bucketize.py
All three new files start directly with includes/code. Every other file under kernels/portable/cpu/ and kernels/test/ carries the standard Copyright (c) Meta Platforms, Inc. and affiliates. ... BSD-style license header. Please add it.

3. bucketize_tensor_out requires a pre-sized out instead of resizing itop_bucketize.cpp:177-180
The kernel checks tensors_have_same_shape(self, out) but never calls resize_tensor(out, self.sizes()). Most portable out-variant kernels (e.g. op_clamp.cpp:129) resize out to the input size so they work with dynamic / upper-bound shapes where the planner allocates out at a max size. With a hard same-shape check this op will fail under dynamic shapes. Consider:

ET_KERNEL_CHECK(
context, resize_tensor(out, self.sizes()) == Error::Ok, InvalidArgument, out);

(For the scalar variant, the out.dim() == 0 check is fine as-is.)

Minor / style

  • op_bucketize.cpp:69,108const bool& right passes a bool by const reference; pass by value (bool right).
  • op_bucketize.cpp:80,116boundaries.sizes().back() can be boundaries.numel() since boundaries is already validated as 1D; slightly clearer.
  • test_bucketize.py — every test docstring says "Test bucketize.Tensor_out", including the scalar tests (lines 120, 126, etc.). Worth correcting the scalar ones.

Things I checked that look correct

  • Compute-type promotion (promote_type_with_scalar for scalar, promoteTypes for tensor) and the REALHBF16 dtype switch.
  • out_int32 ↔ output dtype validation, complex-type and N-D-boundaries rejection paths.
  • Binary-search cus_lower_bound/cus_upper_bound for right/left semantics, empty boundaries (returns 0), and out-of-range / inf inputs — all covered by tests.
  • Build registration in functions.yaml, targets.bzl, and CMakeLists.txt.

Overall this is close to merge-ready; only item #1 is a real functional gap, and #2 is required for CI conventions.
· branch bucketize

@Gasoonjia

Copy link
Copy Markdown
Contributor

Thanks for contribution! I have a question left in the issue mind check that?

Comment threadkernels/test/test_bucketize.py Outdated
# LICENSE file in the root directory of this source tree.

"""
Test for bucketize operations in ExecuTorch.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

plz do not put export test script under kernels/test. Only the runtime test should be fine.

@GallinatorGallinatorJun 22, 2026

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@Gasoonjia So I have to remove test_bucketize.py ? I see that in #15893 they added an export test as well

@Gasoonjia

Copy link
Copy Markdown
Contributor

hi @Gallinator can you try to solve the ci issue?
also add test into https://github.com/pytorch/executorch/blob/main/kernels/test/targets.bzl plz

@Gallinator

Copy link
Copy Markdown
Author

@Gasoonjia CI should be fixed now

Signed-off-by: Gallinator <lcgallinator4@gmail.com>
 WRAP_TO_ATEN does not support Scalar inputs
@Gallinator

Gallinator commented Jul 29, 2026

Copy link
Copy Markdown
Author

@Gasoonjia I removed the scalar input AOT registration for now. If all CI tests pass I will open a new issue for WRAP_TO_ATEN not supporting Scalar inputs

@nil-is-all

Copy link
Copy Markdown
Contributor

@Gasoonjia I removed the scalar input AOT registration for now. If all CI tests pass I will open a new issue for WRAP_TO_ATEN not supporting Scalar inputs

Thanks, running CI again now

@Gallinator

Gallinator commented Aug 4, 2026

Copy link
Copy Markdown
Author

@nil-is-all@Gasoonjia The latest commit solved the 2 previously failing checks, the other failures seem related to the CI setup.
Is this enough to merge the PR?

@nil-is-all

Copy link
Copy Markdown
Contributor

@claude review and check CI failures

@claude

claudeBot commented Aug 5, 2026

Copy link
Copy Markdown

Claude finished @nil-is-all's task in 2m 18s —— View job


Review & CI check

  • Review PR code changes
  • Check CI failures
  • Summarize findings

CI failures — infrastructure, not this PR

I pulled the job logs. Every failure in the pull workflow (and the presets/docs jobs) fails at the same step: Checkout repository (pytorch/test-infra@), with:

API rate limit exceeded for installation. ... request ID DDDA:D8B26:... 2026-08-03 17:52:23 UTC
##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '.../pytorch/executorch'. No such file or directory

The bash/No such file or directory errors are the downstream cascade: the test-infra checkout never completed, so the working directory the later steps expect was never created. None of the failing jobs ever reached a build or test step. The Test QNN Windows MSVC build failure is flagged by Dr. CI as unclassified because that workflow didn't run on the merge base — it's not attributable to this PR either.

Meanwhile the jobs that did get a runner — Lint, Test XNNPACK, Test ARM, Test QNN, Test CoreML, Test Vulkan, Test WebGPU, periodic, test-mcu-cortex-m-backend, and the arm64 test-models-linux-basic variants — all passed. So this is a transient GitHub API rate-limit / infra outage on 2026-08-03, not a code problem. A re-run once the rate limit clears should turn these green — nothing to fix in the PR for CI.

Code review

The code is in good shape and the earlier review feedback has all been incorporated:

  • ✅ BSD license headers present on all new files (op_bucketize.cpp, op_bucketize_test.cpp).
  • bucketize_tensor_out now calls resize_tensor(out, self.sizes()) (op_bucketize.cpp:185-189) instead of a hard same-shape check — works under dynamic/upper-bound shapes.
  • right is now passed by value (bool) rather than const bool&.
  • ✅ The problematic Python export test was removed per @Gasoonjia's request; C++ coverage remains extensive.
  • ✅ Scalar runtime kernel is still registered in functions.yaml (bucketize.Scalar_out); only the AOT-test registration was dropped, matching the WRAP_TO_ATEN Scalar limitation described in the PR.

A couple of small, non-blocking observations:

  1. op_bucketize.cpp:14 still #includes .../thread_parallel_interface.h, but commit 89b9427 removed that target from the Buck deps in op_registration_util.bzl. parallel_for is transitively available via elementwise_util, so this likely compiles, but relying on a transitive include for a header you directly use is fragile under strict-deps. Consider either dropping the direct include (since elementwise_util.h provides parallel_for) or restoring the explicit dep so the include and dep agree. Other parallel_for users (op_amin.cpp) get it via reduce_util.

  2. Minor:boundaries.numel() is used directly for bd_end now (good, since boundaries is validated 1D) — consistent and clear.

Net: the code is essentially merge-ready; the red CI is a test-infra rate-limit outage and a re-run should clear it.
· branch bucketize

@Gallinator

Copy link
Copy Markdown
Author

@Gasoonjia@nil-is-all Is there anything left to improve? I see that #21869 lists other use cases for this kernel

@Gallinator

Copy link
Copy Markdown
Author

@Gasoonjia@nil-is-all CI should be fixed now, could you try to run it again?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.module: kernelsIssues related to kernel libraries and utilities, and code under kernels/release notes: ops & kernelsChanges to the opset and any new / changed kernel implementations

Projects

None yet

Development

Successfully merging this pull request may close these issues.

kernel 'aten::bucketize.Tensor_out' not found.

3 participants

@Gallinator@Gasoonjia@nil-is-all