Dev1/winskuo/gh /aihub remove code - #2

Closed
winskuo-quic wants to merge 21 commits into
mainfrom
dev1/winskuo/gh_/aihub_remove_code
Closed

Dev1/winskuo/gh /aihub remove code#2
winskuo-quic wants to merge 21 commits into
mainfrom
dev1/winskuo/gh_/aihub_remove_code

Conversation

@winskuo-quic

Copy link
Copy Markdown
Collaborator

Summary

[PLEASE REMOVE] See CONTRIBUTING.md's Pull Requests for ExecuTorch PR guidelines.

[PLEASE REMOVE] If this PR closes an issue, please add a Fixes #<issue-id> line.

[PLEASE REMOVE] If this PR introduces a fix or feature that should be the upcoming release notes, please add a "Release notes: " label. For a list of available release notes labels, check out CONTRIBUTING.md's Pull Requests.

Test plan

[PLEASE REMOVE] How did you test this PR? Please write down any manual commands you used and note down tests that you have written if applicable.

JacobSzwejbkaand others added 21 commits April 24, 2026 15:25
Summary:
Attempt 3 to check numel and nbytes overflow. This time we defer
checking dynamic sized inputs until their size is realized.
Reviewed By: lucylq
Differential Revision: D98148157
…inear (pytorch#19117)
## Summary
The bilinear grid_sampler_2d portable kernel computes interpolation
weights via subtractions like `(ix_se - ix)` where both operands are
close integer-valued coordinates in pixel space. In fp16 (10 bits of
mantissa) that's classic catastrophic cancellation — the result has only
a handful of significant bits. The downstream weighted-sum accumulation
then loses further precision.
Measured on a unit test exercising interior grid points with fp16
inputs, the kernel drifts by ~0.1 absolute from an fp32 reference.
That's visible as incorrect depth / flow output near non-integer sample
points, which is most of them.
## Fix
An `AccType<CTYPE>` trait mapping `Half` and `BFloat16` to `float`,
leaving every other dtype unchanged. Used for intermediate coordinate,
weight computation, and `out_val` accumulation. Loads cast `CTYPE ->
ACC`; the final store casts `ACC -> CTYPE` once. Only internal math is
promoted; memory layout / public API / tensor dtypes are unchanged.
```cpp
template <typename CTYPE>
using AccType = std::conditional_t<
std::is_same_v<CTYPE, executorch::aten::Half> ||
std::is_same_v<CTYPE, executorch::aten::BFloat16>,
float,
CTYPE>;
```
## Effects
- **fp32 / Int / any non-half dtype**: `AccType<T>` is `T`, so the
generated code is byte-identical. No behavior change.
- **Half / BFloat16**: `max_abs` vs an fp32 reference drops from **~0.1
to 0** on the shapes I tested (N=1..2, C=7..64, H/W up to 96, both
`align_corners` values).
- **Perf**: a handful of fp16↔fp32 conversions per output element. Not
measurable at op level; well within the portable kernel's scalar cost
envelope.
## Scope
Only touches the bilinear interpolation path. The nearest-mode path
doesn't do weighted-sum accumulation and doesn't have the cancellation
issue — left alone in this change.
## Test plan
- [x] Builds clean for Android arm64 and host (Apple Clang 21).
- [x] Verified numerically via a standalone harness that runs the kernel
with matched fp32 / fp16 inputs and compares against an
fp32-then-downcast reference. All shapes pass within a single fp16 ULP
(or are bit-exact). fp32 tests remain bit-identical to the pre-change
kernel.
- [x] Existing `kernels/test/op_grid_sampler_2d_test.cpp` unit tests
continue to pass (both fp32 shapes that were previously tested, and the
fp16 path I'm specifically fixing).
Happy to add an fp16-specific test case to `op_grid_sampler_2d_test.cpp`
if useful for CI coverage here — just let me know the preferred
approach.
cc @larryliu0820@manuelcandales
Differential Revision: D101728720
Pull Request resolved: pytorch#19052
### Summary
Fixespytorch#18924
Extends `aten.bitwise_not` support in the MLX delegate to handle integer
tensors, not just boolean tensors.
Previously the handler only dispatched to `LogicalNotNode` for `bool`
and raised `NotImplementedError` for all other dtypes. This adds a
dedicated `BitwiseInvertNode` backed by `mlx::core::bitwise_invert`, and
updates the handler to dispatch based on dtype:
- `bool` → `LogicalNotNode` (unchanged)
- `int32`, `int64` → `BitwiseInvertNode`
#### Changes:
- `serialization/schema.fbs`: add `BitwiseInvertNode` table and append
to `OpNode` union
- `runtime/MLXInterpreter.h`: add `exec_bitwise_invert()` and dispatch
case
- `ops.py`: update `_bitwise_not_handler` to dispatch to
`BitwiseInvertNode` for integers
- `test/test_ops.py`: add `bitwise_not_int` test for `int32` and `int64`
### Test plan
All tests were ran on a machine with an Apple M1 Pro CPU, macOS 26.4.1.
- `python3 -m py_compile backends/mlx/ops.py
backends/mlx/test/test_ops.py`
- `python3 backends/mlx/serialization/generate.py`
- `python3 -m executorch.backends.mlx.test.run_all_tests
bitwise_not_int`
### Test output
```
============================================================
TEST SUMMARY
============================================================
Passed: 6
Failed: 0
============================================================
```
cc @metascroy
---------
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Scott Roy <161522778+metascroy@users.noreply.github.com>
Differential Revision: D102070375
Pull Request resolved: pytorch#19074
Add Closeable interface so Module can be used with try-with-resources.
close() delegates to destroy(). Also make destroy() idempotent by
checking mHybridData.isValid() before calling resetNative(), satisfying
the Closeable contract.
This commit was authored with the help of Claude.
TrainingModule: implement Closeable, replace Log.e + silent empty
returns with IllegalStateException throws. Add checkNotDestroyed() guard
on all public methods.
SGD: throw IllegalStateException instead of bare RuntimeException when
optimizer is destroyed.
AsrModule: throw ExecutorchRuntimeException instead of bare
RuntimeException on transcription failure.
ExecuTorchRuntime.validateFilePath: throw IllegalArgumentException
instead of bare RuntimeException, with descriptive message.
JNI constructors: wrap ExecuTorchJni and ExecuTorchLlmJni constructor
bodies in try-catch so C++ exceptions become ExecutorchRuntimeException
instead of generic RuntimeException.
This commit was authored with the help of Claude.
---------
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Differential Revision: D102425537
Pull Request resolved: pytorch#19128
…_nhwc operator + tests
Differential Revision: D96507563
Pull Request resolved: pytorch#18479
…ch#19028 (pytorch#19133)
## Summary
Reverts the following Android PRs:
- pytorch#19099 — Android: consistent error types across all modules
- pytorch#19124 — Android: Module implements Closeable
- pytorch#19092 — Android: improve error diagnostics for LlmModule and
exceptions
- pytorch#19028 — Ignored Module tests: provide required input tensor
Authored with Claude.
Differential Revision: D102488314
Pull Request resolved: pytorch#19134
Differential Revision: D102385104
Pull Request resolved: pytorch#19122
Differential Revision: D102493794
Pull Request resolved: pytorch#19136
@winskuo-quic
winskuo-quic changed the base branch from dev1/winskuo/gh_aihub_remove_readme to mainApril 27, 2026 02:47
@winskuo-quic
winskuo-quic deleted the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:47
@winskuo-quic
winskuo-quic restored the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:48
@winskuo-quic
winskuo-quic deleted the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:49
@winskuo-quic
winskuo-quic restored the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:49
@winskuo-quic
winskuo-quic deleted the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:50
@winskuo-quic
winskuo-quic restored the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:50
quic-boyuc added a commit that referenced this pull request May 5, 2026
Add observe_pass, a decorator that wraps any PassBase subclass or callable
pass to automatically collect input/output graphs via Observatory. Names
are derived from the class/function name and deduplicated by collect()
itself (#2, pytorch#3, ...) so repeated calls never silently overwrite records.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request May 18, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request May 22, 2026
In compare mode each archive column already shows the label in its sticky
header (commit 5ae2a72), so the `<label>/` that compare_archives prepends
to every record and session name renders as redundant noise inside the
column and squeezes the natural-width layout from c9e7bda.
Strip is display-only — `rec.name` / `session.name` in `state.data` stay
prefixed so graph_ref resolution and the global record map keep working.
A `displayPrefix` arg threads through `renderRecordItem`,
`_renderSessionDashboardLink`, `renderIndexFlat`, and `renderIndexTree`;
`_renderArchiveColumn` passes `archive.label`. Single-archive paths pass
no prefix and are unchanged. Collision suffixes (`A/foo #2`) strip cleanly
to `foo #2` since only the head matches.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
Add observe_pass, a decorator that wraps any PassBase subclass or callable
pass to automatically collect input/output graphs via Observatory. Names
are derived from the class/function name and deduplicated by collect()
itself (#2, pytorch#3, ...) so repeated calls never silently overwrite records.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
In compare mode each archive column already shows the label in its sticky
header (commit 5ae2a72), so the `<label>/` that compare_archives prepends
to every record and session name renders as redundant noise inside the
column and squeezes the natural-width layout from c9e7bda.
Strip is display-only — `rec.name` / `session.name` in `state.data` stay
prefixed so graph_ref resolution and the global record map keep working.
A `displayPrefix` arg threads through `renderRecordItem`,
`_renderSessionDashboardLink`, `renderIndexFlat`, and `renderIndexTree`;
`_renderArchiveColumn` passes `archive.label`. Single-archive paths pass
no prefix and are unchanged. Collision suffixes (`A/foo #2`) strip cleanly
to `foo #2` since only the head matches.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
qti-horodnic pushed a commit that referenced this pull request Jul 28, 2026
### Summary
- fill the test gap for newly added op
- utils test
### Test plan
pytest backends/qualcomm/tests/rework/utils/test.py
quic-boyuc pushed a commit that referenced this pull request Aug 25, 2026
…2095)
### Summary
ETDump records an intermediate tensor by handing the tensor's data
pointer to a
data sink, and every sink reads those bytes with a plain host read:
```cpp
memcpy(cur_data_begin, ptr, length);
```
When the tensor lives on an accelerator that pointer is not host memory,
so the
read segfaults. A program placed on CUDA crashes as soon as tracing is
turned
on, which is exactly when someone is trying to debug it.
Returning an error instead would not help. All four callers wrap the
result in
`ET_CHECK_MSG`, so an error aborts the process rather than skipping the
tensor.
This change brings the data back to host memory first. When the tensor
is not on
CPU, ETDump looks up the allocator registered for that device type,
stages the
bytes into a temporary host buffer with `copy_device_to_host`, writes
that buffer
to the sink and frees it. A tensor on CPU keeps the old path and copies
nothing
extra.
If no allocator is registered for the device, ETDump now reports
`NotFound` and
logs the device type instead of reading the pointer anyway.
### Test plan
Two new test files, each with a CMake target and a Buck target.
`devtools/etdump/tests/etdump_device_test.cpp` registers the existing
`MockCudaAllocator`, which backs its device memory with host memory, and
has two
tests. One logs a tensor tagged as CUDA and checks both that ETDump went
through
the allocator and that the bytes reached the debug buffer. The other
logs a
tensor on CPU and checks that the allocator was not used at all, so the
CPU path
is unchanged.
`devtools/etdump/tests/etdump_device_no_allocator_test.cpp` covers the
case where
nothing is registered for the device. The registry is a process wide
static with
no way to remove an entry, so that case needs a binary that never
registers
anything, which is why it is a second file.
`devtools/etdump/tests/CMakeLists.txt` was not referenced by any parent
`CMakeLists.txt`, so nothing in that directory was built by CMake. This
adds
`add_subdirectory(tests)` to `devtools/etdump/CMakeLists.txt` under
`BUILD_TESTING`, so both new tests are picked up by `ctest`, which is
how the C++
tests run.
The pre-existing `sdk_etdump_tests` target stays out of the CMake build.
It
compiles `etdump_test.cpp`, which includes `etdump_filter.h`, which
needs re2,
and the devtools build does not pull re2 in. It is now guarded on re2
being
available rather than being silently unreachable.
With this change both new binaries pass under `ctest`:
```
1/2 Test #1: etdump_device_test ................ Passed
2/2 Test #2: etdump_device_no_allocator_test ... Passed
```
With `etdump_flatcc.cpp` reverted to the old code and everything
rebuilt, both
fail:
```
Expected equality of these values:
g_mock_cuda.d2h_count_
Which is: 0
1
Death test: etdump_gen.log_evalue(EValue(tensor))
Result: failed to die.
```
Also reproduced the real crash on one NVIDIA H100, with a small program
that
allocates through `cudaMalloc`, tags a tensor as CUDA and logs it:
```
before: Segmentation fault (core dumped)
after: debug buffer holds 1.5 2.5 3.5 4.5
```
Checked that `etdump_flatcc.cpp` still compiles with `-DUSE_ATEN_LIB`.
`clang-format` reports no changes needed on the four touched C++ files.
### Landing order
pytorch#22058 adds device-planned arenas to the Python bindings that ETDump can
be handed
pointers into, and without this change `BufferDataSink::write` would
`memcpy` device
memory from the host. This should land before or together with it.
### Not covered
The ATen mode branch of the device type conversion only compiles. There
is no
ATen mode CMake build to run it in, so nothing here executes it. It is
reachable
in principle: `runtime/executor/tensor_parser_aten.cpp` reads the
serialized
device type and index, including CUDA, and builds the ATen tensor on
that device.
An earlier version of this description said ATen tensors carry no device
metadata,
which is wrong.
`LogTensorOnCpuDoesNotStageThroughTheAllocator` passes with the
production change
reverted, since CPU tensors already went straight to the data sink. It
documents
the CPU path rather than locking the fix; the other two tests are the
ones that
require the new branch.
---------
Co-authored-by: r <r@e>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

10 participants

@winskuo-quic@JacobSzwejbka@jgibson2@navsud@AlessandroVacca@lucylq@psiddh@DrJessop@hsharma35@kirklandsign
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Dev1/winskuo/gh /aihub remove code - #2

Closed
winskuo-quic wants to merge 21 commits into
mainfrom
dev1/winskuo/gh_/aihub_remove_code
Closed

Dev1/winskuo/gh /aihub remove code#2
winskuo-quic wants to merge 21 commits into
mainfrom
dev1/winskuo/gh_/aihub_remove_code

Conversation

@winskuo-quic

Copy link
Copy Markdown
Collaborator

Summary

[PLEASE REMOVE] See CONTRIBUTING.md's Pull Requests for ExecuTorch PR guidelines.

[PLEASE REMOVE] If this PR closes an issue, please add a Fixes #<issue-id> line.

[PLEASE REMOVE] If this PR introduces a fix or feature that should be the upcoming release notes, please add a "Release notes: " label. For a list of available release notes labels, check out CONTRIBUTING.md's Pull Requests.

Test plan

[PLEASE REMOVE] How did you test this PR? Please write down any manual commands you used and note down tests that you have written if applicable.

JacobSzwejbkaand others added 21 commits April 24, 2026 15:25
Summary:
Attempt 3 to check numel and nbytes overflow. This time we defer
checking dynamic sized inputs until their size is realized.
Reviewed By: lucylq
Differential Revision: D98148157
…inear (pytorch#19117)
## Summary
The bilinear grid_sampler_2d portable kernel computes interpolation
weights via subtractions like `(ix_se - ix)` where both operands are
close integer-valued coordinates in pixel space. In fp16 (10 bits of
mantissa) that's classic catastrophic cancellation — the result has only
a handful of significant bits. The downstream weighted-sum accumulation
then loses further precision.
Measured on a unit test exercising interior grid points with fp16
inputs, the kernel drifts by ~0.1 absolute from an fp32 reference.
That's visible as incorrect depth / flow output near non-integer sample
points, which is most of them.
## Fix
An `AccType<CTYPE>` trait mapping `Half` and `BFloat16` to `float`,
leaving every other dtype unchanged. Used for intermediate coordinate,
weight computation, and `out_val` accumulation. Loads cast `CTYPE ->
ACC`; the final store casts `ACC -> CTYPE` once. Only internal math is
promoted; memory layout / public API / tensor dtypes are unchanged.
```cpp
template <typename CTYPE>
using AccType = std::conditional_t<
std::is_same_v<CTYPE, executorch::aten::Half> ||
std::is_same_v<CTYPE, executorch::aten::BFloat16>,
float,
CTYPE>;
```
## Effects
- **fp32 / Int / any non-half dtype**: `AccType<T>` is `T`, so the
generated code is byte-identical. No behavior change.
- **Half / BFloat16**: `max_abs` vs an fp32 reference drops from **~0.1
to 0** on the shapes I tested (N=1..2, C=7..64, H/W up to 96, both
`align_corners` values).
- **Perf**: a handful of fp16↔fp32 conversions per output element. Not
measurable at op level; well within the portable kernel's scalar cost
envelope.
## Scope
Only touches the bilinear interpolation path. The nearest-mode path
doesn't do weighted-sum accumulation and doesn't have the cancellation
issue — left alone in this change.
## Test plan
- [x] Builds clean for Android arm64 and host (Apple Clang 21).
- [x] Verified numerically via a standalone harness that runs the kernel
with matched fp32 / fp16 inputs and compares against an
fp32-then-downcast reference. All shapes pass within a single fp16 ULP
(or are bit-exact). fp32 tests remain bit-identical to the pre-change
kernel.
- [x] Existing `kernels/test/op_grid_sampler_2d_test.cpp` unit tests
continue to pass (both fp32 shapes that were previously tested, and the
fp16 path I'm specifically fixing).
Happy to add an fp16-specific test case to `op_grid_sampler_2d_test.cpp`
if useful for CI coverage here — just let me know the preferred
approach.
cc @larryliu0820@manuelcandales
Differential Revision: D101728720
Pull Request resolved: pytorch#19052
### Summary
Fixespytorch#18924
Extends `aten.bitwise_not` support in the MLX delegate to handle integer
tensors, not just boolean tensors.
Previously the handler only dispatched to `LogicalNotNode` for `bool`
and raised `NotImplementedError` for all other dtypes. This adds a
dedicated `BitwiseInvertNode` backed by `mlx::core::bitwise_invert`, and
updates the handler to dispatch based on dtype:
- `bool` → `LogicalNotNode` (unchanged)
- `int32`, `int64` → `BitwiseInvertNode`
#### Changes:
- `serialization/schema.fbs`: add `BitwiseInvertNode` table and append
to `OpNode` union
- `runtime/MLXInterpreter.h`: add `exec_bitwise_invert()` and dispatch
case
- `ops.py`: update `_bitwise_not_handler` to dispatch to
`BitwiseInvertNode` for integers
- `test/test_ops.py`: add `bitwise_not_int` test for `int32` and `int64`
### Test plan
All tests were ran on a machine with an Apple M1 Pro CPU, macOS 26.4.1.
- `python3 -m py_compile backends/mlx/ops.py
backends/mlx/test/test_ops.py`
- `python3 backends/mlx/serialization/generate.py`
- `python3 -m executorch.backends.mlx.test.run_all_tests
bitwise_not_int`
### Test output
```
============================================================
TEST SUMMARY
============================================================
Passed: 6
Failed: 0
============================================================
```
cc @metascroy
---------
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Scott Roy <161522778+metascroy@users.noreply.github.com>
Differential Revision: D102070375
Pull Request resolved: pytorch#19074
Add Closeable interface so Module can be used with try-with-resources.
close() delegates to destroy(). Also make destroy() idempotent by
checking mHybridData.isValid() before calling resetNative(), satisfying
the Closeable contract.
This commit was authored with the help of Claude.
TrainingModule: implement Closeable, replace Log.e + silent empty
returns with IllegalStateException throws. Add checkNotDestroyed() guard
on all public methods.
SGD: throw IllegalStateException instead of bare RuntimeException when
optimizer is destroyed.
AsrModule: throw ExecutorchRuntimeException instead of bare
RuntimeException on transcription failure.
ExecuTorchRuntime.validateFilePath: throw IllegalArgumentException
instead of bare RuntimeException, with descriptive message.
JNI constructors: wrap ExecuTorchJni and ExecuTorchLlmJni constructor
bodies in try-catch so C++ exceptions become ExecutorchRuntimeException
instead of generic RuntimeException.
This commit was authored with the help of Claude.
---------
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Differential Revision: D102425537
Pull Request resolved: pytorch#19128
…_nhwc operator + tests
Differential Revision: D96507563
Pull Request resolved: pytorch#18479
…ch#19028 (pytorch#19133)
## Summary
Reverts the following Android PRs:
- pytorch#19099 — Android: consistent error types across all modules
- pytorch#19124 — Android: Module implements Closeable
- pytorch#19092 — Android: improve error diagnostics for LlmModule and
exceptions
- pytorch#19028 — Ignored Module tests: provide required input tensor
Authored with Claude.
Differential Revision: D102488314
Pull Request resolved: pytorch#19134
Differential Revision: D102385104
Pull Request resolved: pytorch#19122
Differential Revision: D102493794
Pull Request resolved: pytorch#19136
@winskuo-quic
winskuo-quic changed the base branch from dev1/winskuo/gh_aihub_remove_readme to mainApril 27, 2026 02:47
@winskuo-quic
winskuo-quic deleted the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:47
@winskuo-quic
winskuo-quic restored the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:48
@winskuo-quic
winskuo-quic deleted the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:49
@winskuo-quic
winskuo-quic restored the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:49
@winskuo-quic
winskuo-quic deleted the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:50
@winskuo-quic
winskuo-quic restored the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:50
quic-boyuc added a commit that referenced this pull request May 5, 2026
Add observe_pass, a decorator that wraps any PassBase subclass or callable
pass to automatically collect input/output graphs via Observatory. Names
are derived from the class/function name and deduplicated by collect()
itself (#2, pytorch#3, ...) so repeated calls never silently overwrite records.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request May 18, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request May 22, 2026
In compare mode each archive column already shows the label in its sticky
header (commit 5ae2a72), so the `<label>/` that compare_archives prepends
to every record and session name renders as redundant noise inside the
column and squeezes the natural-width layout from c9e7bda.
Strip is display-only — `rec.name` / `session.name` in `state.data` stay
prefixed so graph_ref resolution and the global record map keep working.
A `displayPrefix` arg threads through `renderRecordItem`,
`_renderSessionDashboardLink`, `renderIndexFlat`, and `renderIndexTree`;
`_renderArchiveColumn` passes `archive.label`. Single-archive paths pass
no prefix and are unchanged. Collision suffixes (`A/foo #2`) strip cleanly
to `foo #2` since only the head matches.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
Add observe_pass, a decorator that wraps any PassBase subclass or callable
pass to automatically collect input/output graphs via Observatory. Names
are derived from the class/function name and deduplicated by collect()
itself (#2, pytorch#3, ...) so repeated calls never silently overwrite records.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
In compare mode each archive column already shows the label in its sticky
header (commit 5ae2a72), so the `<label>/` that compare_archives prepends
to every record and session name renders as redundant noise inside the
column and squeezes the natural-width layout from c9e7bda.
Strip is display-only — `rec.name` / `session.name` in `state.data` stay
prefixed so graph_ref resolution and the global record map keep working.
A `displayPrefix` arg threads through `renderRecordItem`,
`_renderSessionDashboardLink`, `renderIndexFlat`, and `renderIndexTree`;
`_renderArchiveColumn` passes `archive.label`. Single-archive paths pass
no prefix and are unchanged. Collision suffixes (`A/foo #2`) strip cleanly
to `foo #2` since only the head matches.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
qti-horodnic pushed a commit that referenced this pull request Jul 28, 2026
### Summary
- fill the test gap for newly added op
- utils test
### Test plan
pytest backends/qualcomm/tests/rework/utils/test.py
quic-boyuc pushed a commit that referenced this pull request Aug 25, 2026
…2095)
### Summary
ETDump records an intermediate tensor by handing the tensor's data
pointer to a
data sink, and every sink reads those bytes with a plain host read:
```cpp
memcpy(cur_data_begin, ptr, length);
```
When the tensor lives on an accelerator that pointer is not host memory,
so the
read segfaults. A program placed on CUDA crashes as soon as tracing is
turned
on, which is exactly when someone is trying to debug it.
Returning an error instead would not help. All four callers wrap the
result in
`ET_CHECK_MSG`, so an error aborts the process rather than skipping the
tensor.
This change brings the data back to host memory first. When the tensor
is not on
CPU, ETDump looks up the allocator registered for that device type,
stages the
bytes into a temporary host buffer with `copy_device_to_host`, writes
that buffer
to the sink and frees it. A tensor on CPU keeps the old path and copies
nothing
extra.
If no allocator is registered for the device, ETDump now reports
`NotFound` and
logs the device type instead of reading the pointer anyway.
### Test plan
Two new test files, each with a CMake target and a Buck target.
`devtools/etdump/tests/etdump_device_test.cpp` registers the existing
`MockCudaAllocator`, which backs its device memory with host memory, and
has two
tests. One logs a tensor tagged as CUDA and checks both that ETDump went
through
the allocator and that the bytes reached the debug buffer. The other
logs a
tensor on CPU and checks that the allocator was not used at all, so the
CPU path
is unchanged.
`devtools/etdump/tests/etdump_device_no_allocator_test.cpp` covers the
case where
nothing is registered for the device. The registry is a process wide
static with
no way to remove an entry, so that case needs a binary that never
registers
anything, which is why it is a second file.
`devtools/etdump/tests/CMakeLists.txt` was not referenced by any parent
`CMakeLists.txt`, so nothing in that directory was built by CMake. This
adds
`add_subdirectory(tests)` to `devtools/etdump/CMakeLists.txt` under
`BUILD_TESTING`, so both new tests are picked up by `ctest`, which is
how the C++
tests run.
The pre-existing `sdk_etdump_tests` target stays out of the CMake build.
It
compiles `etdump_test.cpp`, which includes `etdump_filter.h`, which
needs re2,
and the devtools build does not pull re2 in. It is now guarded on re2
being
available rather than being silently unreachable.
With this change both new binaries pass under `ctest`:
```
1/2 Test #1: etdump_device_test ................ Passed
2/2 Test #2: etdump_device_no_allocator_test ... Passed
```
With `etdump_flatcc.cpp` reverted to the old code and everything
rebuilt, both
fail:
```
Expected equality of these values:
g_mock_cuda.d2h_count_
Which is: 0
1
Death test: etdump_gen.log_evalue(EValue(tensor))
Result: failed to die.
```
Also reproduced the real crash on one NVIDIA H100, with a small program
that
allocates through `cudaMalloc`, tags a tensor as CUDA and logs it:
```
before: Segmentation fault (core dumped)
after: debug buffer holds 1.5 2.5 3.5 4.5
```
Checked that `etdump_flatcc.cpp` still compiles with `-DUSE_ATEN_LIB`.
`clang-format` reports no changes needed on the four touched C++ files.
### Landing order
pytorch#22058 adds device-planned arenas to the Python bindings that ETDump can
be handed
pointers into, and without this change `BufferDataSink::write` would
`memcpy` device
memory from the host. This should land before or together with it.
### Not covered
The ATen mode branch of the device type conversion only compiles. There
is no
ATen mode CMake build to run it in, so nothing here executes it. It is
reachable
in principle: `runtime/executor/tensor_parser_aten.cpp` reads the
serialized
device type and index, including CUDA, and builds the ATen tensor on
that device.
An earlier version of this description said ATen tensors carry no device
metadata,
which is wrong.
`LogTensorOnCpuDoesNotStageThroughTheAllocator` passes with the
production change
reverted, since CPU tensors already went straight to the data sink. It
documents
the CPU path rather than locking the fix; the other two tests are the
ones that
require the new branch.
---------
Co-authored-by: r <r@e>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

10 participants

@winskuo-quic@JacobSzwejbka@jgibson2@navsud@AlessandroVacca@lucylq@psiddh@DrJessop@hsharma35@kirklandsign
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Dev1/winskuo/gh /aihub remove code - #2

Closed
winskuo-quic wants to merge 21 commits into
mainfrom
dev1/winskuo/gh_/aihub_remove_code
Closed

Dev1/winskuo/gh /aihub remove code#2
winskuo-quic wants to merge 21 commits into
mainfrom
dev1/winskuo/gh_/aihub_remove_code

Conversation

@winskuo-quic

Copy link
Copy Markdown
Collaborator

Summary

[PLEASE REMOVE] See CONTRIBUTING.md's Pull Requests for ExecuTorch PR guidelines.

[PLEASE REMOVE] If this PR closes an issue, please add a Fixes #<issue-id> line.

[PLEASE REMOVE] If this PR introduces a fix or feature that should be the upcoming release notes, please add a "Release notes: " label. For a list of available release notes labels, check out CONTRIBUTING.md's Pull Requests.

Test plan

[PLEASE REMOVE] How did you test this PR? Please write down any manual commands you used and note down tests that you have written if applicable.

JacobSzwejbkaand others added 21 commits April 24, 2026 15:25
Summary:
Attempt 3 to check numel and nbytes overflow. This time we defer
checking dynamic sized inputs until their size is realized.
Reviewed By: lucylq
Differential Revision: D98148157
…inear (pytorch#19117)
## Summary
The bilinear grid_sampler_2d portable kernel computes interpolation
weights via subtractions like `(ix_se - ix)` where both operands are
close integer-valued coordinates in pixel space. In fp16 (10 bits of
mantissa) that's classic catastrophic cancellation — the result has only
a handful of significant bits. The downstream weighted-sum accumulation
then loses further precision.
Measured on a unit test exercising interior grid points with fp16
inputs, the kernel drifts by ~0.1 absolute from an fp32 reference.
That's visible as incorrect depth / flow output near non-integer sample
points, which is most of them.
## Fix
An `AccType<CTYPE>` trait mapping `Half` and `BFloat16` to `float`,
leaving every other dtype unchanged. Used for intermediate coordinate,
weight computation, and `out_val` accumulation. Loads cast `CTYPE ->
ACC`; the final store casts `ACC -> CTYPE` once. Only internal math is
promoted; memory layout / public API / tensor dtypes are unchanged.
```cpp
template <typename CTYPE>
using AccType = std::conditional_t<
std::is_same_v<CTYPE, executorch::aten::Half> ||
std::is_same_v<CTYPE, executorch::aten::BFloat16>,
float,
CTYPE>;
```
## Effects
- **fp32 / Int / any non-half dtype**: `AccType<T>` is `T`, so the
generated code is byte-identical. No behavior change.
- **Half / BFloat16**: `max_abs` vs an fp32 reference drops from **~0.1
to 0** on the shapes I tested (N=1..2, C=7..64, H/W up to 96, both
`align_corners` values).
- **Perf**: a handful of fp16↔fp32 conversions per output element. Not
measurable at op level; well within the portable kernel's scalar cost
envelope.
## Scope
Only touches the bilinear interpolation path. The nearest-mode path
doesn't do weighted-sum accumulation and doesn't have the cancellation
issue — left alone in this change.
## Test plan
- [x] Builds clean for Android arm64 and host (Apple Clang 21).
- [x] Verified numerically via a standalone harness that runs the kernel
with matched fp32 / fp16 inputs and compares against an
fp32-then-downcast reference. All shapes pass within a single fp16 ULP
(or are bit-exact). fp32 tests remain bit-identical to the pre-change
kernel.
- [x] Existing `kernels/test/op_grid_sampler_2d_test.cpp` unit tests
continue to pass (both fp32 shapes that were previously tested, and the
fp16 path I'm specifically fixing).
Happy to add an fp16-specific test case to `op_grid_sampler_2d_test.cpp`
if useful for CI coverage here — just let me know the preferred
approach.
cc @larryliu0820@manuelcandales
Differential Revision: D101728720
Pull Request resolved: pytorch#19052
### Summary
Fixespytorch#18924
Extends `aten.bitwise_not` support in the MLX delegate to handle integer
tensors, not just boolean tensors.
Previously the handler only dispatched to `LogicalNotNode` for `bool`
and raised `NotImplementedError` for all other dtypes. This adds a
dedicated `BitwiseInvertNode` backed by `mlx::core::bitwise_invert`, and
updates the handler to dispatch based on dtype:
- `bool` → `LogicalNotNode` (unchanged)
- `int32`, `int64` → `BitwiseInvertNode`
#### Changes:
- `serialization/schema.fbs`: add `BitwiseInvertNode` table and append
to `OpNode` union
- `runtime/MLXInterpreter.h`: add `exec_bitwise_invert()` and dispatch
case
- `ops.py`: update `_bitwise_not_handler` to dispatch to
`BitwiseInvertNode` for integers
- `test/test_ops.py`: add `bitwise_not_int` test for `int32` and `int64`
### Test plan
All tests were ran on a machine with an Apple M1 Pro CPU, macOS 26.4.1.
- `python3 -m py_compile backends/mlx/ops.py
backends/mlx/test/test_ops.py`
- `python3 backends/mlx/serialization/generate.py`
- `python3 -m executorch.backends.mlx.test.run_all_tests
bitwise_not_int`
### Test output
```
============================================================
TEST SUMMARY
============================================================
Passed: 6
Failed: 0
============================================================
```
cc @metascroy
---------
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Scott Roy <161522778+metascroy@users.noreply.github.com>
Differential Revision: D102070375
Pull Request resolved: pytorch#19074
Add Closeable interface so Module can be used with try-with-resources.
close() delegates to destroy(). Also make destroy() idempotent by
checking mHybridData.isValid() before calling resetNative(), satisfying
the Closeable contract.
This commit was authored with the help of Claude.
TrainingModule: implement Closeable, replace Log.e + silent empty
returns with IllegalStateException throws. Add checkNotDestroyed() guard
on all public methods.
SGD: throw IllegalStateException instead of bare RuntimeException when
optimizer is destroyed.
AsrModule: throw ExecutorchRuntimeException instead of bare
RuntimeException on transcription failure.
ExecuTorchRuntime.validateFilePath: throw IllegalArgumentException
instead of bare RuntimeException, with descriptive message.
JNI constructors: wrap ExecuTorchJni and ExecuTorchLlmJni constructor
bodies in try-catch so C++ exceptions become ExecutorchRuntimeException
instead of generic RuntimeException.
This commit was authored with the help of Claude.
---------
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Differential Revision: D102425537
Pull Request resolved: pytorch#19128
…_nhwc operator + tests
Differential Revision: D96507563
Pull Request resolved: pytorch#18479
…ch#19028 (pytorch#19133)
## Summary
Reverts the following Android PRs:
- pytorch#19099 — Android: consistent error types across all modules
- pytorch#19124 — Android: Module implements Closeable
- pytorch#19092 — Android: improve error diagnostics for LlmModule and
exceptions
- pytorch#19028 — Ignored Module tests: provide required input tensor
Authored with Claude.
Differential Revision: D102488314
Pull Request resolved: pytorch#19134
Differential Revision: D102385104
Pull Request resolved: pytorch#19122
Differential Revision: D102493794
Pull Request resolved: pytorch#19136
@winskuo-quic
winskuo-quic changed the base branch from dev1/winskuo/gh_aihub_remove_readme to mainApril 27, 2026 02:47
@winskuo-quic
winskuo-quic deleted the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:47
@winskuo-quic
winskuo-quic restored the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:48
@winskuo-quic
winskuo-quic deleted the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:49
@winskuo-quic
winskuo-quic restored the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:49
@winskuo-quic
winskuo-quic deleted the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:50
@winskuo-quic
winskuo-quic restored the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:50
quic-boyuc added a commit that referenced this pull request May 5, 2026
Add observe_pass, a decorator that wraps any PassBase subclass or callable
pass to automatically collect input/output graphs via Observatory. Names
are derived from the class/function name and deduplicated by collect()
itself (#2, pytorch#3, ...) so repeated calls never silently overwrite records.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request May 18, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request May 22, 2026
In compare mode each archive column already shows the label in its sticky
header (commit 5ae2a72), so the `<label>/` that compare_archives prepends
to every record and session name renders as redundant noise inside the
column and squeezes the natural-width layout from c9e7bda.
Strip is display-only — `rec.name` / `session.name` in `state.data` stay
prefixed so graph_ref resolution and the global record map keep working.
A `displayPrefix` arg threads through `renderRecordItem`,
`_renderSessionDashboardLink`, `renderIndexFlat`, and `renderIndexTree`;
`_renderArchiveColumn` passes `archive.label`. Single-archive paths pass
no prefix and are unchanged. Collision suffixes (`A/foo #2`) strip cleanly
to `foo #2` since only the head matches.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
Add observe_pass, a decorator that wraps any PassBase subclass or callable
pass to automatically collect input/output graphs via Observatory. Names
are derived from the class/function name and deduplicated by collect()
itself (#2, pytorch#3, ...) so repeated calls never silently overwrite records.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
In compare mode each archive column already shows the label in its sticky
header (commit 5ae2a72), so the `<label>/` that compare_archives prepends
to every record and session name renders as redundant noise inside the
column and squeezes the natural-width layout from c9e7bda.
Strip is display-only — `rec.name` / `session.name` in `state.data` stay
prefixed so graph_ref resolution and the global record map keep working.
A `displayPrefix` arg threads through `renderRecordItem`,
`_renderSessionDashboardLink`, `renderIndexFlat`, and `renderIndexTree`;
`_renderArchiveColumn` passes `archive.label`. Single-archive paths pass
no prefix and are unchanged. Collision suffixes (`A/foo #2`) strip cleanly
to `foo #2` since only the head matches.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
qti-horodnic pushed a commit that referenced this pull request Jul 28, 2026
### Summary
- fill the test gap for newly added op
- utils test
### Test plan
pytest backends/qualcomm/tests/rework/utils/test.py
quic-boyuc pushed a commit that referenced this pull request Aug 25, 2026
…2095)
### Summary
ETDump records an intermediate tensor by handing the tensor's data
pointer to a
data sink, and every sink reads those bytes with a plain host read:
```cpp
memcpy(cur_data_begin, ptr, length);
```
When the tensor lives on an accelerator that pointer is not host memory,
so the
read segfaults. A program placed on CUDA crashes as soon as tracing is
turned
on, which is exactly when someone is trying to debug it.
Returning an error instead would not help. All four callers wrap the
result in
`ET_CHECK_MSG`, so an error aborts the process rather than skipping the
tensor.
This change brings the data back to host memory first. When the tensor
is not on
CPU, ETDump looks up the allocator registered for that device type,
stages the
bytes into a temporary host buffer with `copy_device_to_host`, writes
that buffer
to the sink and frees it. A tensor on CPU keeps the old path and copies
nothing
extra.
If no allocator is registered for the device, ETDump now reports
`NotFound` and
logs the device type instead of reading the pointer anyway.
### Test plan
Two new test files, each with a CMake target and a Buck target.
`devtools/etdump/tests/etdump_device_test.cpp` registers the existing
`MockCudaAllocator`, which backs its device memory with host memory, and
has two
tests. One logs a tensor tagged as CUDA and checks both that ETDump went
through
the allocator and that the bytes reached the debug buffer. The other
logs a
tensor on CPU and checks that the allocator was not used at all, so the
CPU path
is unchanged.
`devtools/etdump/tests/etdump_device_no_allocator_test.cpp` covers the
case where
nothing is registered for the device. The registry is a process wide
static with
no way to remove an entry, so that case needs a binary that never
registers
anything, which is why it is a second file.
`devtools/etdump/tests/CMakeLists.txt` was not referenced by any parent
`CMakeLists.txt`, so nothing in that directory was built by CMake. This
adds
`add_subdirectory(tests)` to `devtools/etdump/CMakeLists.txt` under
`BUILD_TESTING`, so both new tests are picked up by `ctest`, which is
how the C++
tests run.
The pre-existing `sdk_etdump_tests` target stays out of the CMake build.
It
compiles `etdump_test.cpp`, which includes `etdump_filter.h`, which
needs re2,
and the devtools build does not pull re2 in. It is now guarded on re2
being
available rather than being silently unreachable.
With this change both new binaries pass under `ctest`:
```
1/2 Test #1: etdump_device_test ................ Passed
2/2 Test #2: etdump_device_no_allocator_test ... Passed
```
With `etdump_flatcc.cpp` reverted to the old code and everything
rebuilt, both
fail:
```
Expected equality of these values:
g_mock_cuda.d2h_count_
Which is: 0
1
Death test: etdump_gen.log_evalue(EValue(tensor))
Result: failed to die.
```
Also reproduced the real crash on one NVIDIA H100, with a small program
that
allocates through `cudaMalloc`, tags a tensor as CUDA and logs it:
```
before: Segmentation fault (core dumped)
after: debug buffer holds 1.5 2.5 3.5 4.5
```
Checked that `etdump_flatcc.cpp` still compiles with `-DUSE_ATEN_LIB`.
`clang-format` reports no changes needed on the four touched C++ files.
### Landing order
pytorch#22058 adds device-planned arenas to the Python bindings that ETDump can
be handed
pointers into, and without this change `BufferDataSink::write` would
`memcpy` device
memory from the host. This should land before or together with it.
### Not covered
The ATen mode branch of the device type conversion only compiles. There
is no
ATen mode CMake build to run it in, so nothing here executes it. It is
reachable
in principle: `runtime/executor/tensor_parser_aten.cpp` reads the
serialized
device type and index, including CUDA, and builds the ATen tensor on
that device.
An earlier version of this description said ATen tensors carry no device
metadata,
which is wrong.
`LogTensorOnCpuDoesNotStageThroughTheAllocator` passes with the
production change
reverted, since CPU tensors already went straight to the data sink. It
documents
the CPU path rather than locking the fix; the other two tests are the
ones that
require the new branch.
---------
Co-authored-by: r <r@e>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

10 participants

@winskuo-quic@JacobSzwejbka@jgibson2@navsud@AlessandroVacca@lucylq@psiddh@DrJessop@hsharma35@kirklandsign
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Dev1/winskuo/gh /aihub remove code - #2

Closed
winskuo-quic wants to merge 21 commits into
mainfrom
dev1/winskuo/gh_/aihub_remove_code
Closed

Dev1/winskuo/gh /aihub remove code#2
winskuo-quic wants to merge 21 commits into
mainfrom
dev1/winskuo/gh_/aihub_remove_code

Conversation

@winskuo-quic

Copy link
Copy Markdown
Collaborator

Summary

[PLEASE REMOVE] See CONTRIBUTING.md's Pull Requests for ExecuTorch PR guidelines.

[PLEASE REMOVE] If this PR closes an issue, please add a Fixes #<issue-id> line.

[PLEASE REMOVE] If this PR introduces a fix or feature that should be the upcoming release notes, please add a "Release notes: " label. For a list of available release notes labels, check out CONTRIBUTING.md's Pull Requests.

Test plan

[PLEASE REMOVE] How did you test this PR? Please write down any manual commands you used and note down tests that you have written if applicable.

JacobSzwejbkaand others added 21 commits April 24, 2026 15:25
Summary:
Attempt 3 to check numel and nbytes overflow. This time we defer
checking dynamic sized inputs until their size is realized.
Reviewed By: lucylq
Differential Revision: D98148157
…inear (pytorch#19117)
## Summary
The bilinear grid_sampler_2d portable kernel computes interpolation
weights via subtractions like `(ix_se - ix)` where both operands are
close integer-valued coordinates in pixel space. In fp16 (10 bits of
mantissa) that's classic catastrophic cancellation — the result has only
a handful of significant bits. The downstream weighted-sum accumulation
then loses further precision.
Measured on a unit test exercising interior grid points with fp16
inputs, the kernel drifts by ~0.1 absolute from an fp32 reference.
That's visible as incorrect depth / flow output near non-integer sample
points, which is most of them.
## Fix
An `AccType<CTYPE>` trait mapping `Half` and `BFloat16` to `float`,
leaving every other dtype unchanged. Used for intermediate coordinate,
weight computation, and `out_val` accumulation. Loads cast `CTYPE ->
ACC`; the final store casts `ACC -> CTYPE` once. Only internal math is
promoted; memory layout / public API / tensor dtypes are unchanged.
```cpp
template <typename CTYPE>
using AccType = std::conditional_t<
std::is_same_v<CTYPE, executorch::aten::Half> ||
std::is_same_v<CTYPE, executorch::aten::BFloat16>,
float,
CTYPE>;
```
## Effects
- **fp32 / Int / any non-half dtype**: `AccType<T>` is `T`, so the
generated code is byte-identical. No behavior change.
- **Half / BFloat16**: `max_abs` vs an fp32 reference drops from **~0.1
to 0** on the shapes I tested (N=1..2, C=7..64, H/W up to 96, both
`align_corners` values).
- **Perf**: a handful of fp16↔fp32 conversions per output element. Not
measurable at op level; well within the portable kernel's scalar cost
envelope.
## Scope
Only touches the bilinear interpolation path. The nearest-mode path
doesn't do weighted-sum accumulation and doesn't have the cancellation
issue — left alone in this change.
## Test plan
- [x] Builds clean for Android arm64 and host (Apple Clang 21).
- [x] Verified numerically via a standalone harness that runs the kernel
with matched fp32 / fp16 inputs and compares against an
fp32-then-downcast reference. All shapes pass within a single fp16 ULP
(or are bit-exact). fp32 tests remain bit-identical to the pre-change
kernel.
- [x] Existing `kernels/test/op_grid_sampler_2d_test.cpp` unit tests
continue to pass (both fp32 shapes that were previously tested, and the
fp16 path I'm specifically fixing).
Happy to add an fp16-specific test case to `op_grid_sampler_2d_test.cpp`
if useful for CI coverage here — just let me know the preferred
approach.
cc @larryliu0820@manuelcandales
Differential Revision: D101728720
Pull Request resolved: pytorch#19052
### Summary
Fixespytorch#18924
Extends `aten.bitwise_not` support in the MLX delegate to handle integer
tensors, not just boolean tensors.
Previously the handler only dispatched to `LogicalNotNode` for `bool`
and raised `NotImplementedError` for all other dtypes. This adds a
dedicated `BitwiseInvertNode` backed by `mlx::core::bitwise_invert`, and
updates the handler to dispatch based on dtype:
- `bool` → `LogicalNotNode` (unchanged)
- `int32`, `int64` → `BitwiseInvertNode`
#### Changes:
- `serialization/schema.fbs`: add `BitwiseInvertNode` table and append
to `OpNode` union
- `runtime/MLXInterpreter.h`: add `exec_bitwise_invert()` and dispatch
case
- `ops.py`: update `_bitwise_not_handler` to dispatch to
`BitwiseInvertNode` for integers
- `test/test_ops.py`: add `bitwise_not_int` test for `int32` and `int64`
### Test plan
All tests were ran on a machine with an Apple M1 Pro CPU, macOS 26.4.1.
- `python3 -m py_compile backends/mlx/ops.py
backends/mlx/test/test_ops.py`
- `python3 backends/mlx/serialization/generate.py`
- `python3 -m executorch.backends.mlx.test.run_all_tests
bitwise_not_int`
### Test output
```
============================================================
TEST SUMMARY
============================================================
Passed: 6
Failed: 0
============================================================
```
cc @metascroy
---------
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Scott Roy <161522778+metascroy@users.noreply.github.com>
Differential Revision: D102070375
Pull Request resolved: pytorch#19074
Add Closeable interface so Module can be used with try-with-resources.
close() delegates to destroy(). Also make destroy() idempotent by
checking mHybridData.isValid() before calling resetNative(), satisfying
the Closeable contract.
This commit was authored with the help of Claude.
TrainingModule: implement Closeable, replace Log.e + silent empty
returns with IllegalStateException throws. Add checkNotDestroyed() guard
on all public methods.
SGD: throw IllegalStateException instead of bare RuntimeException when
optimizer is destroyed.
AsrModule: throw ExecutorchRuntimeException instead of bare
RuntimeException on transcription failure.
ExecuTorchRuntime.validateFilePath: throw IllegalArgumentException
instead of bare RuntimeException, with descriptive message.
JNI constructors: wrap ExecuTorchJni and ExecuTorchLlmJni constructor
bodies in try-catch so C++ exceptions become ExecutorchRuntimeException
instead of generic RuntimeException.
This commit was authored with the help of Claude.
---------
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Differential Revision: D102425537
Pull Request resolved: pytorch#19128
…_nhwc operator + tests
Differential Revision: D96507563
Pull Request resolved: pytorch#18479
…ch#19028 (pytorch#19133)
## Summary
Reverts the following Android PRs:
- pytorch#19099 — Android: consistent error types across all modules
- pytorch#19124 — Android: Module implements Closeable
- pytorch#19092 — Android: improve error diagnostics for LlmModule and
exceptions
- pytorch#19028 — Ignored Module tests: provide required input tensor
Authored with Claude.
Differential Revision: D102488314
Pull Request resolved: pytorch#19134
Differential Revision: D102385104
Pull Request resolved: pytorch#19122
Differential Revision: D102493794
Pull Request resolved: pytorch#19136
@winskuo-quic
winskuo-quic changed the base branch from dev1/winskuo/gh_aihub_remove_readme to mainApril 27, 2026 02:47
@winskuo-quic
winskuo-quic deleted the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:47
@winskuo-quic
winskuo-quic restored the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:48
@winskuo-quic
winskuo-quic deleted the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:49
@winskuo-quic
winskuo-quic restored the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:49
@winskuo-quic
winskuo-quic deleted the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:50
@winskuo-quic
winskuo-quic restored the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:50
quic-boyuc added a commit that referenced this pull request May 5, 2026
Add observe_pass, a decorator that wraps any PassBase subclass or callable
pass to automatically collect input/output graphs via Observatory. Names
are derived from the class/function name and deduplicated by collect()
itself (#2, pytorch#3, ...) so repeated calls never silently overwrite records.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request May 18, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request May 22, 2026
In compare mode each archive column already shows the label in its sticky
header (commit 5ae2a72), so the `<label>/` that compare_archives prepends
to every record and session name renders as redundant noise inside the
column and squeezes the natural-width layout from c9e7bda.
Strip is display-only — `rec.name` / `session.name` in `state.data` stay
prefixed so graph_ref resolution and the global record map keep working.
A `displayPrefix` arg threads through `renderRecordItem`,
`_renderSessionDashboardLink`, `renderIndexFlat`, and `renderIndexTree`;
`_renderArchiveColumn` passes `archive.label`. Single-archive paths pass
no prefix and are unchanged. Collision suffixes (`A/foo #2`) strip cleanly
to `foo #2` since only the head matches.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
Add observe_pass, a decorator that wraps any PassBase subclass or callable
pass to automatically collect input/output graphs via Observatory. Names
are derived from the class/function name and deduplicated by collect()
itself (#2, pytorch#3, ...) so repeated calls never silently overwrite records.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
In compare mode each archive column already shows the label in its sticky
header (commit 5ae2a72), so the `<label>/` that compare_archives prepends
to every record and session name renders as redundant noise inside the
column and squeezes the natural-width layout from c9e7bda.
Strip is display-only — `rec.name` / `session.name` in `state.data` stay
prefixed so graph_ref resolution and the global record map keep working.
A `displayPrefix` arg threads through `renderRecordItem`,
`_renderSessionDashboardLink`, `renderIndexFlat`, and `renderIndexTree`;
`_renderArchiveColumn` passes `archive.label`. Single-archive paths pass
no prefix and are unchanged. Collision suffixes (`A/foo #2`) strip cleanly
to `foo #2` since only the head matches.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
qti-horodnic pushed a commit that referenced this pull request Jul 28, 2026
### Summary
- fill the test gap for newly added op
- utils test
### Test plan
pytest backends/qualcomm/tests/rework/utils/test.py
quic-boyuc pushed a commit that referenced this pull request Aug 25, 2026
…2095)
### Summary
ETDump records an intermediate tensor by handing the tensor's data
pointer to a
data sink, and every sink reads those bytes with a plain host read:
```cpp
memcpy(cur_data_begin, ptr, length);
```
When the tensor lives on an accelerator that pointer is not host memory,
so the
read segfaults. A program placed on CUDA crashes as soon as tracing is
turned
on, which is exactly when someone is trying to debug it.
Returning an error instead would not help. All four callers wrap the
result in
`ET_CHECK_MSG`, so an error aborts the process rather than skipping the
tensor.
This change brings the data back to host memory first. When the tensor
is not on
CPU, ETDump looks up the allocator registered for that device type,
stages the
bytes into a temporary host buffer with `copy_device_to_host`, writes
that buffer
to the sink and frees it. A tensor on CPU keeps the old path and copies
nothing
extra.
If no allocator is registered for the device, ETDump now reports
`NotFound` and
logs the device type instead of reading the pointer anyway.
### Test plan
Two new test files, each with a CMake target and a Buck target.
`devtools/etdump/tests/etdump_device_test.cpp` registers the existing
`MockCudaAllocator`, which backs its device memory with host memory, and
has two
tests. One logs a tensor tagged as CUDA and checks both that ETDump went
through
the allocator and that the bytes reached the debug buffer. The other
logs a
tensor on CPU and checks that the allocator was not used at all, so the
CPU path
is unchanged.
`devtools/etdump/tests/etdump_device_no_allocator_test.cpp` covers the
case where
nothing is registered for the device. The registry is a process wide
static with
no way to remove an entry, so that case needs a binary that never
registers
anything, which is why it is a second file.
`devtools/etdump/tests/CMakeLists.txt` was not referenced by any parent
`CMakeLists.txt`, so nothing in that directory was built by CMake. This
adds
`add_subdirectory(tests)` to `devtools/etdump/CMakeLists.txt` under
`BUILD_TESTING`, so both new tests are picked up by `ctest`, which is
how the C++
tests run.
The pre-existing `sdk_etdump_tests` target stays out of the CMake build.
It
compiles `etdump_test.cpp`, which includes `etdump_filter.h`, which
needs re2,
and the devtools build does not pull re2 in. It is now guarded on re2
being
available rather than being silently unreachable.
With this change both new binaries pass under `ctest`:
```
1/2 Test #1: etdump_device_test ................ Passed
2/2 Test #2: etdump_device_no_allocator_test ... Passed
```
With `etdump_flatcc.cpp` reverted to the old code and everything
rebuilt, both
fail:
```
Expected equality of these values:
g_mock_cuda.d2h_count_
Which is: 0
1
Death test: etdump_gen.log_evalue(EValue(tensor))
Result: failed to die.
```
Also reproduced the real crash on one NVIDIA H100, with a small program
that
allocates through `cudaMalloc`, tags a tensor as CUDA and logs it:
```
before: Segmentation fault (core dumped)
after: debug buffer holds 1.5 2.5 3.5 4.5
```
Checked that `etdump_flatcc.cpp` still compiles with `-DUSE_ATEN_LIB`.
`clang-format` reports no changes needed on the four touched C++ files.
### Landing order
pytorch#22058 adds device-planned arenas to the Python bindings that ETDump can
be handed
pointers into, and without this change `BufferDataSink::write` would
`memcpy` device
memory from the host. This should land before or together with it.
### Not covered
The ATen mode branch of the device type conversion only compiles. There
is no
ATen mode CMake build to run it in, so nothing here executes it. It is
reachable
in principle: `runtime/executor/tensor_parser_aten.cpp` reads the
serialized
device type and index, including CUDA, and builds the ATen tensor on
that device.
An earlier version of this description said ATen tensors carry no device
metadata,
which is wrong.
`LogTensorOnCpuDoesNotStageThroughTheAllocator` passes with the
production change
reverted, since CPU tensors already went straight to the data sink. It
documents
the CPU path rather than locking the fix; the other two tests are the
ones that
require the new branch.
---------
Co-authored-by: r <r@e>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

10 participants

@winskuo-quic@JacobSzwejbka@jgibson2@navsud@AlessandroVacca@lucylq@psiddh@DrJessop@hsharma35@kirklandsign
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Dev1/winskuo/gh /aihub remove code - #2

Closed
winskuo-quic wants to merge 21 commits into
mainfrom
dev1/winskuo/gh_/aihub_remove_code
Closed

Dev1/winskuo/gh /aihub remove code#2
winskuo-quic wants to merge 21 commits into
mainfrom
dev1/winskuo/gh_/aihub_remove_code

Conversation

@winskuo-quic

Copy link
Copy Markdown
Collaborator

Summary

[PLEASE REMOVE] See CONTRIBUTING.md's Pull Requests for ExecuTorch PR guidelines.

[PLEASE REMOVE] If this PR closes an issue, please add a Fixes #<issue-id> line.

[PLEASE REMOVE] If this PR introduces a fix or feature that should be the upcoming release notes, please add a "Release notes: " label. For a list of available release notes labels, check out CONTRIBUTING.md's Pull Requests.

Test plan

[PLEASE REMOVE] How did you test this PR? Please write down any manual commands you used and note down tests that you have written if applicable.

JacobSzwejbkaand others added 21 commits April 24, 2026 15:25
Summary:
Attempt 3 to check numel and nbytes overflow. This time we defer
checking dynamic sized inputs until their size is realized.
Reviewed By: lucylq
Differential Revision: D98148157
…inear (pytorch#19117)
## Summary
The bilinear grid_sampler_2d portable kernel computes interpolation
weights via subtractions like `(ix_se - ix)` where both operands are
close integer-valued coordinates in pixel space. In fp16 (10 bits of
mantissa) that's classic catastrophic cancellation — the result has only
a handful of significant bits. The downstream weighted-sum accumulation
then loses further precision.
Measured on a unit test exercising interior grid points with fp16
inputs, the kernel drifts by ~0.1 absolute from an fp32 reference.
That's visible as incorrect depth / flow output near non-integer sample
points, which is most of them.
## Fix
An `AccType<CTYPE>` trait mapping `Half` and `BFloat16` to `float`,
leaving every other dtype unchanged. Used for intermediate coordinate,
weight computation, and `out_val` accumulation. Loads cast `CTYPE ->
ACC`; the final store casts `ACC -> CTYPE` once. Only internal math is
promoted; memory layout / public API / tensor dtypes are unchanged.
```cpp
template <typename CTYPE>
using AccType = std::conditional_t<
std::is_same_v<CTYPE, executorch::aten::Half> ||
std::is_same_v<CTYPE, executorch::aten::BFloat16>,
float,
CTYPE>;
```
## Effects
- **fp32 / Int / any non-half dtype**: `AccType<T>` is `T`, so the
generated code is byte-identical. No behavior change.
- **Half / BFloat16**: `max_abs` vs an fp32 reference drops from **~0.1
to 0** on the shapes I tested (N=1..2, C=7..64, H/W up to 96, both
`align_corners` values).
- **Perf**: a handful of fp16↔fp32 conversions per output element. Not
measurable at op level; well within the portable kernel's scalar cost
envelope.
## Scope
Only touches the bilinear interpolation path. The nearest-mode path
doesn't do weighted-sum accumulation and doesn't have the cancellation
issue — left alone in this change.
## Test plan
- [x] Builds clean for Android arm64 and host (Apple Clang 21).
- [x] Verified numerically via a standalone harness that runs the kernel
with matched fp32 / fp16 inputs and compares against an
fp32-then-downcast reference. All shapes pass within a single fp16 ULP
(or are bit-exact). fp32 tests remain bit-identical to the pre-change
kernel.
- [x] Existing `kernels/test/op_grid_sampler_2d_test.cpp` unit tests
continue to pass (both fp32 shapes that were previously tested, and the
fp16 path I'm specifically fixing).
Happy to add an fp16-specific test case to `op_grid_sampler_2d_test.cpp`
if useful for CI coverage here — just let me know the preferred
approach.
cc @larryliu0820@manuelcandales
Differential Revision: D101728720
Pull Request resolved: pytorch#19052
### Summary
Fixespytorch#18924
Extends `aten.bitwise_not` support in the MLX delegate to handle integer
tensors, not just boolean tensors.
Previously the handler only dispatched to `LogicalNotNode` for `bool`
and raised `NotImplementedError` for all other dtypes. This adds a
dedicated `BitwiseInvertNode` backed by `mlx::core::bitwise_invert`, and
updates the handler to dispatch based on dtype:
- `bool` → `LogicalNotNode` (unchanged)
- `int32`, `int64` → `BitwiseInvertNode`
#### Changes:
- `serialization/schema.fbs`: add `BitwiseInvertNode` table and append
to `OpNode` union
- `runtime/MLXInterpreter.h`: add `exec_bitwise_invert()` and dispatch
case
- `ops.py`: update `_bitwise_not_handler` to dispatch to
`BitwiseInvertNode` for integers
- `test/test_ops.py`: add `bitwise_not_int` test for `int32` and `int64`
### Test plan
All tests were ran on a machine with an Apple M1 Pro CPU, macOS 26.4.1.
- `python3 -m py_compile backends/mlx/ops.py
backends/mlx/test/test_ops.py`
- `python3 backends/mlx/serialization/generate.py`
- `python3 -m executorch.backends.mlx.test.run_all_tests
bitwise_not_int`
### Test output
```
============================================================
TEST SUMMARY
============================================================
Passed: 6
Failed: 0
============================================================
```
cc @metascroy
---------
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Scott Roy <161522778+metascroy@users.noreply.github.com>
Differential Revision: D102070375
Pull Request resolved: pytorch#19074
Add Closeable interface so Module can be used with try-with-resources.
close() delegates to destroy(). Also make destroy() idempotent by
checking mHybridData.isValid() before calling resetNative(), satisfying
the Closeable contract.
This commit was authored with the help of Claude.
TrainingModule: implement Closeable, replace Log.e + silent empty
returns with IllegalStateException throws. Add checkNotDestroyed() guard
on all public methods.
SGD: throw IllegalStateException instead of bare RuntimeException when
optimizer is destroyed.
AsrModule: throw ExecutorchRuntimeException instead of bare
RuntimeException on transcription failure.
ExecuTorchRuntime.validateFilePath: throw IllegalArgumentException
instead of bare RuntimeException, with descriptive message.
JNI constructors: wrap ExecuTorchJni and ExecuTorchLlmJni constructor
bodies in try-catch so C++ exceptions become ExecutorchRuntimeException
instead of generic RuntimeException.
This commit was authored with the help of Claude.
---------
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Differential Revision: D102425537
Pull Request resolved: pytorch#19128
…_nhwc operator + tests
Differential Revision: D96507563
Pull Request resolved: pytorch#18479
…ch#19028 (pytorch#19133)
## Summary
Reverts the following Android PRs:
- pytorch#19099 — Android: consistent error types across all modules
- pytorch#19124 — Android: Module implements Closeable
- pytorch#19092 — Android: improve error diagnostics for LlmModule and
exceptions
- pytorch#19028 — Ignored Module tests: provide required input tensor
Authored with Claude.
Differential Revision: D102488314
Pull Request resolved: pytorch#19134
Differential Revision: D102385104
Pull Request resolved: pytorch#19122
Differential Revision: D102493794
Pull Request resolved: pytorch#19136
@winskuo-quic
winskuo-quic changed the base branch from dev1/winskuo/gh_aihub_remove_readme to mainApril 27, 2026 02:47
@winskuo-quic
winskuo-quic deleted the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:47
@winskuo-quic
winskuo-quic restored the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:48
@winskuo-quic
winskuo-quic deleted the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:49
@winskuo-quic
winskuo-quic restored the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:49
@winskuo-quic
winskuo-quic deleted the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:50
@winskuo-quic
winskuo-quic restored the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:50
quic-boyuc added a commit that referenced this pull request May 5, 2026
Add observe_pass, a decorator that wraps any PassBase subclass or callable
pass to automatically collect input/output graphs via Observatory. Names
are derived from the class/function name and deduplicated by collect()
itself (#2, pytorch#3, ...) so repeated calls never silently overwrite records.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request May 18, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request May 22, 2026
In compare mode each archive column already shows the label in its sticky
header (commit 5ae2a72), so the `<label>/` that compare_archives prepends
to every record and session name renders as redundant noise inside the
column and squeezes the natural-width layout from c9e7bda.
Strip is display-only — `rec.name` / `session.name` in `state.data` stay
prefixed so graph_ref resolution and the global record map keep working.
A `displayPrefix` arg threads through `renderRecordItem`,
`_renderSessionDashboardLink`, `renderIndexFlat`, and `renderIndexTree`;
`_renderArchiveColumn` passes `archive.label`. Single-archive paths pass
no prefix and are unchanged. Collision suffixes (`A/foo #2`) strip cleanly
to `foo #2` since only the head matches.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
Add observe_pass, a decorator that wraps any PassBase subclass or callable
pass to automatically collect input/output graphs via Observatory. Names
are derived from the class/function name and deduplicated by collect()
itself (#2, pytorch#3, ...) so repeated calls never silently overwrite records.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
In compare mode each archive column already shows the label in its sticky
header (commit 5ae2a72), so the `<label>/` that compare_archives prepends
to every record and session name renders as redundant noise inside the
column and squeezes the natural-width layout from c9e7bda.
Strip is display-only — `rec.name` / `session.name` in `state.data` stay
prefixed so graph_ref resolution and the global record map keep working.
A `displayPrefix` arg threads through `renderRecordItem`,
`_renderSessionDashboardLink`, `renderIndexFlat`, and `renderIndexTree`;
`_renderArchiveColumn` passes `archive.label`. Single-archive paths pass
no prefix and are unchanged. Collision suffixes (`A/foo #2`) strip cleanly
to `foo #2` since only the head matches.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
qti-horodnic pushed a commit that referenced this pull request Jul 28, 2026
### Summary
- fill the test gap for newly added op
- utils test
### Test plan
pytest backends/qualcomm/tests/rework/utils/test.py
quic-boyuc pushed a commit that referenced this pull request Aug 25, 2026
…2095)
### Summary
ETDump records an intermediate tensor by handing the tensor's data
pointer to a
data sink, and every sink reads those bytes with a plain host read:
```cpp
memcpy(cur_data_begin, ptr, length);
```
When the tensor lives on an accelerator that pointer is not host memory,
so the
read segfaults. A program placed on CUDA crashes as soon as tracing is
turned
on, which is exactly when someone is trying to debug it.
Returning an error instead would not help. All four callers wrap the
result in
`ET_CHECK_MSG`, so an error aborts the process rather than skipping the
tensor.
This change brings the data back to host memory first. When the tensor
is not on
CPU, ETDump looks up the allocator registered for that device type,
stages the
bytes into a temporary host buffer with `copy_device_to_host`, writes
that buffer
to the sink and frees it. A tensor on CPU keeps the old path and copies
nothing
extra.
If no allocator is registered for the device, ETDump now reports
`NotFound` and
logs the device type instead of reading the pointer anyway.
### Test plan
Two new test files, each with a CMake target and a Buck target.
`devtools/etdump/tests/etdump_device_test.cpp` registers the existing
`MockCudaAllocator`, which backs its device memory with host memory, and
has two
tests. One logs a tensor tagged as CUDA and checks both that ETDump went
through
the allocator and that the bytes reached the debug buffer. The other
logs a
tensor on CPU and checks that the allocator was not used at all, so the
CPU path
is unchanged.
`devtools/etdump/tests/etdump_device_no_allocator_test.cpp` covers the
case where
nothing is registered for the device. The registry is a process wide
static with
no way to remove an entry, so that case needs a binary that never
registers
anything, which is why it is a second file.
`devtools/etdump/tests/CMakeLists.txt` was not referenced by any parent
`CMakeLists.txt`, so nothing in that directory was built by CMake. This
adds
`add_subdirectory(tests)` to `devtools/etdump/CMakeLists.txt` under
`BUILD_TESTING`, so both new tests are picked up by `ctest`, which is
how the C++
tests run.
The pre-existing `sdk_etdump_tests` target stays out of the CMake build.
It
compiles `etdump_test.cpp`, which includes `etdump_filter.h`, which
needs re2,
and the devtools build does not pull re2 in. It is now guarded on re2
being
available rather than being silently unreachable.
With this change both new binaries pass under `ctest`:
```
1/2 Test #1: etdump_device_test ................ Passed
2/2 Test #2: etdump_device_no_allocator_test ... Passed
```
With `etdump_flatcc.cpp` reverted to the old code and everything
rebuilt, both
fail:
```
Expected equality of these values:
g_mock_cuda.d2h_count_
Which is: 0
1
Death test: etdump_gen.log_evalue(EValue(tensor))
Result: failed to die.
```
Also reproduced the real crash on one NVIDIA H100, with a small program
that
allocates through `cudaMalloc`, tags a tensor as CUDA and logs it:
```
before: Segmentation fault (core dumped)
after: debug buffer holds 1.5 2.5 3.5 4.5
```
Checked that `etdump_flatcc.cpp` still compiles with `-DUSE_ATEN_LIB`.
`clang-format` reports no changes needed on the four touched C++ files.
### Landing order
pytorch#22058 adds device-planned arenas to the Python bindings that ETDump can
be handed
pointers into, and without this change `BufferDataSink::write` would
`memcpy` device
memory from the host. This should land before or together with it.
### Not covered
The ATen mode branch of the device type conversion only compiles. There
is no
ATen mode CMake build to run it in, so nothing here executes it. It is
reachable
in principle: `runtime/executor/tensor_parser_aten.cpp` reads the
serialized
device type and index, including CUDA, and builds the ATen tensor on
that device.
An earlier version of this description said ATen tensors carry no device
metadata,
which is wrong.
`LogTensorOnCpuDoesNotStageThroughTheAllocator` passes with the
production change
reverted, since CPU tensors already went straight to the data sink. It
documents
the CPU path rather than locking the fix; the other two tests are the
ones that
require the new branch.
---------
Co-authored-by: r <r@e>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

10 participants

@winskuo-quic@JacobSzwejbka@jgibson2@navsud@AlessandroVacca@lucylq@psiddh@DrJessop@hsharma35@kirklandsign
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Dev1/winskuo/gh /aihub remove code - #2

Closed
winskuo-quic wants to merge 21 commits into
mainfrom
dev1/winskuo/gh_/aihub_remove_code
Closed

Dev1/winskuo/gh /aihub remove code#2
winskuo-quic wants to merge 21 commits into
mainfrom
dev1/winskuo/gh_/aihub_remove_code

Conversation

@winskuo-quic

Copy link
Copy Markdown
Collaborator

Summary

[PLEASE REMOVE] See CONTRIBUTING.md's Pull Requests for ExecuTorch PR guidelines.

[PLEASE REMOVE] If this PR closes an issue, please add a Fixes #<issue-id> line.

[PLEASE REMOVE] If this PR introduces a fix or feature that should be the upcoming release notes, please add a "Release notes: " label. For a list of available release notes labels, check out CONTRIBUTING.md's Pull Requests.

Test plan

[PLEASE REMOVE] How did you test this PR? Please write down any manual commands you used and note down tests that you have written if applicable.

JacobSzwejbkaand others added 21 commits April 24, 2026 15:25
Summary:
Attempt 3 to check numel and nbytes overflow. This time we defer
checking dynamic sized inputs until their size is realized.
Reviewed By: lucylq
Differential Revision: D98148157
…inear (pytorch#19117)
## Summary
The bilinear grid_sampler_2d portable kernel computes interpolation
weights via subtractions like `(ix_se - ix)` where both operands are
close integer-valued coordinates in pixel space. In fp16 (10 bits of
mantissa) that's classic catastrophic cancellation — the result has only
a handful of significant bits. The downstream weighted-sum accumulation
then loses further precision.
Measured on a unit test exercising interior grid points with fp16
inputs, the kernel drifts by ~0.1 absolute from an fp32 reference.
That's visible as incorrect depth / flow output near non-integer sample
points, which is most of them.
## Fix
An `AccType<CTYPE>` trait mapping `Half` and `BFloat16` to `float`,
leaving every other dtype unchanged. Used for intermediate coordinate,
weight computation, and `out_val` accumulation. Loads cast `CTYPE ->
ACC`; the final store casts `ACC -> CTYPE` once. Only internal math is
promoted; memory layout / public API / tensor dtypes are unchanged.
```cpp
template <typename CTYPE>
using AccType = std::conditional_t<
std::is_same_v<CTYPE, executorch::aten::Half> ||
std::is_same_v<CTYPE, executorch::aten::BFloat16>,
float,
CTYPE>;
```
## Effects
- **fp32 / Int / any non-half dtype**: `AccType<T>` is `T`, so the
generated code is byte-identical. No behavior change.
- **Half / BFloat16**: `max_abs` vs an fp32 reference drops from **~0.1
to 0** on the shapes I tested (N=1..2, C=7..64, H/W up to 96, both
`align_corners` values).
- **Perf**: a handful of fp16↔fp32 conversions per output element. Not
measurable at op level; well within the portable kernel's scalar cost
envelope.
## Scope
Only touches the bilinear interpolation path. The nearest-mode path
doesn't do weighted-sum accumulation and doesn't have the cancellation
issue — left alone in this change.
## Test plan
- [x] Builds clean for Android arm64 and host (Apple Clang 21).
- [x] Verified numerically via a standalone harness that runs the kernel
with matched fp32 / fp16 inputs and compares against an
fp32-then-downcast reference. All shapes pass within a single fp16 ULP
(or are bit-exact). fp32 tests remain bit-identical to the pre-change
kernel.
- [x] Existing `kernels/test/op_grid_sampler_2d_test.cpp` unit tests
continue to pass (both fp32 shapes that were previously tested, and the
fp16 path I'm specifically fixing).
Happy to add an fp16-specific test case to `op_grid_sampler_2d_test.cpp`
if useful for CI coverage here — just let me know the preferred
approach.
cc @larryliu0820@manuelcandales
Differential Revision: D101728720
Pull Request resolved: pytorch#19052
### Summary
Fixespytorch#18924
Extends `aten.bitwise_not` support in the MLX delegate to handle integer
tensors, not just boolean tensors.
Previously the handler only dispatched to `LogicalNotNode` for `bool`
and raised `NotImplementedError` for all other dtypes. This adds a
dedicated `BitwiseInvertNode` backed by `mlx::core::bitwise_invert`, and
updates the handler to dispatch based on dtype:
- `bool` → `LogicalNotNode` (unchanged)
- `int32`, `int64` → `BitwiseInvertNode`
#### Changes:
- `serialization/schema.fbs`: add `BitwiseInvertNode` table and append
to `OpNode` union
- `runtime/MLXInterpreter.h`: add `exec_bitwise_invert()` and dispatch
case
- `ops.py`: update `_bitwise_not_handler` to dispatch to
`BitwiseInvertNode` for integers
- `test/test_ops.py`: add `bitwise_not_int` test for `int32` and `int64`
### Test plan
All tests were ran on a machine with an Apple M1 Pro CPU, macOS 26.4.1.
- `python3 -m py_compile backends/mlx/ops.py
backends/mlx/test/test_ops.py`
- `python3 backends/mlx/serialization/generate.py`
- `python3 -m executorch.backends.mlx.test.run_all_tests
bitwise_not_int`
### Test output
```
============================================================
TEST SUMMARY
============================================================
Passed: 6
Failed: 0
============================================================
```
cc @metascroy
---------
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Scott Roy <161522778+metascroy@users.noreply.github.com>
Differential Revision: D102070375
Pull Request resolved: pytorch#19074
Add Closeable interface so Module can be used with try-with-resources.
close() delegates to destroy(). Also make destroy() idempotent by
checking mHybridData.isValid() before calling resetNative(), satisfying
the Closeable contract.
This commit was authored with the help of Claude.
TrainingModule: implement Closeable, replace Log.e + silent empty
returns with IllegalStateException throws. Add checkNotDestroyed() guard
on all public methods.
SGD: throw IllegalStateException instead of bare RuntimeException when
optimizer is destroyed.
AsrModule: throw ExecutorchRuntimeException instead of bare
RuntimeException on transcription failure.
ExecuTorchRuntime.validateFilePath: throw IllegalArgumentException
instead of bare RuntimeException, with descriptive message.
JNI constructors: wrap ExecuTorchJni and ExecuTorchLlmJni constructor
bodies in try-catch so C++ exceptions become ExecutorchRuntimeException
instead of generic RuntimeException.
This commit was authored with the help of Claude.
---------
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Differential Revision: D102425537
Pull Request resolved: pytorch#19128
…_nhwc operator + tests
Differential Revision: D96507563
Pull Request resolved: pytorch#18479
…ch#19028 (pytorch#19133)
## Summary
Reverts the following Android PRs:
- pytorch#19099 — Android: consistent error types across all modules
- pytorch#19124 — Android: Module implements Closeable
- pytorch#19092 — Android: improve error diagnostics for LlmModule and
exceptions
- pytorch#19028 — Ignored Module tests: provide required input tensor
Authored with Claude.
Differential Revision: D102488314
Pull Request resolved: pytorch#19134
Differential Revision: D102385104
Pull Request resolved: pytorch#19122
Differential Revision: D102493794
Pull Request resolved: pytorch#19136
@winskuo-quic
winskuo-quic changed the base branch from dev1/winskuo/gh_aihub_remove_readme to mainApril 27, 2026 02:47
@winskuo-quic
winskuo-quic deleted the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:47
@winskuo-quic
winskuo-quic restored the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:48
@winskuo-quic
winskuo-quic deleted the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:49
@winskuo-quic
winskuo-quic restored the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:49
@winskuo-quic
winskuo-quic deleted the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:50
@winskuo-quic
winskuo-quic restored the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:50
quic-boyuc added a commit that referenced this pull request May 5, 2026
Add observe_pass, a decorator that wraps any PassBase subclass or callable
pass to automatically collect input/output graphs via Observatory. Names
are derived from the class/function name and deduplicated by collect()
itself (#2, pytorch#3, ...) so repeated calls never silently overwrite records.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request May 18, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request May 22, 2026
In compare mode each archive column already shows the label in its sticky
header (commit 5ae2a72), so the `<label>/` that compare_archives prepends
to every record and session name renders as redundant noise inside the
column and squeezes the natural-width layout from c9e7bda.
Strip is display-only — `rec.name` / `session.name` in `state.data` stay
prefixed so graph_ref resolution and the global record map keep working.
A `displayPrefix` arg threads through `renderRecordItem`,
`_renderSessionDashboardLink`, `renderIndexFlat`, and `renderIndexTree`;
`_renderArchiveColumn` passes `archive.label`. Single-archive paths pass
no prefix and are unchanged. Collision suffixes (`A/foo #2`) strip cleanly
to `foo #2` since only the head matches.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
Add observe_pass, a decorator that wraps any PassBase subclass or callable
pass to automatically collect input/output graphs via Observatory. Names
are derived from the class/function name and deduplicated by collect()
itself (#2, pytorch#3, ...) so repeated calls never silently overwrite records.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
In compare mode each archive column already shows the label in its sticky
header (commit 5ae2a72), so the `<label>/` that compare_archives prepends
to every record and session name renders as redundant noise inside the
column and squeezes the natural-width layout from c9e7bda.
Strip is display-only — `rec.name` / `session.name` in `state.data` stay
prefixed so graph_ref resolution and the global record map keep working.
A `displayPrefix` arg threads through `renderRecordItem`,
`_renderSessionDashboardLink`, `renderIndexFlat`, and `renderIndexTree`;
`_renderArchiveColumn` passes `archive.label`. Single-archive paths pass
no prefix and are unchanged. Collision suffixes (`A/foo #2`) strip cleanly
to `foo #2` since only the head matches.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
qti-horodnic pushed a commit that referenced this pull request Jul 28, 2026
### Summary
- fill the test gap for newly added op
- utils test
### Test plan
pytest backends/qualcomm/tests/rework/utils/test.py
quic-boyuc pushed a commit that referenced this pull request Aug 25, 2026
…2095)
### Summary
ETDump records an intermediate tensor by handing the tensor's data
pointer to a
data sink, and every sink reads those bytes with a plain host read:
```cpp
memcpy(cur_data_begin, ptr, length);
```
When the tensor lives on an accelerator that pointer is not host memory,
so the
read segfaults. A program placed on CUDA crashes as soon as tracing is
turned
on, which is exactly when someone is trying to debug it.
Returning an error instead would not help. All four callers wrap the
result in
`ET_CHECK_MSG`, so an error aborts the process rather than skipping the
tensor.
This change brings the data back to host memory first. When the tensor
is not on
CPU, ETDump looks up the allocator registered for that device type,
stages the
bytes into a temporary host buffer with `copy_device_to_host`, writes
that buffer
to the sink and frees it. A tensor on CPU keeps the old path and copies
nothing
extra.
If no allocator is registered for the device, ETDump now reports
`NotFound` and
logs the device type instead of reading the pointer anyway.
### Test plan
Two new test files, each with a CMake target and a Buck target.
`devtools/etdump/tests/etdump_device_test.cpp` registers the existing
`MockCudaAllocator`, which backs its device memory with host memory, and
has two
tests. One logs a tensor tagged as CUDA and checks both that ETDump went
through
the allocator and that the bytes reached the debug buffer. The other
logs a
tensor on CPU and checks that the allocator was not used at all, so the
CPU path
is unchanged.
`devtools/etdump/tests/etdump_device_no_allocator_test.cpp` covers the
case where
nothing is registered for the device. The registry is a process wide
static with
no way to remove an entry, so that case needs a binary that never
registers
anything, which is why it is a second file.
`devtools/etdump/tests/CMakeLists.txt` was not referenced by any parent
`CMakeLists.txt`, so nothing in that directory was built by CMake. This
adds
`add_subdirectory(tests)` to `devtools/etdump/CMakeLists.txt` under
`BUILD_TESTING`, so both new tests are picked up by `ctest`, which is
how the C++
tests run.
The pre-existing `sdk_etdump_tests` target stays out of the CMake build.
It
compiles `etdump_test.cpp`, which includes `etdump_filter.h`, which
needs re2,
and the devtools build does not pull re2 in. It is now guarded on re2
being
available rather than being silently unreachable.
With this change both new binaries pass under `ctest`:
```
1/2 Test #1: etdump_device_test ................ Passed
2/2 Test #2: etdump_device_no_allocator_test ... Passed
```
With `etdump_flatcc.cpp` reverted to the old code and everything
rebuilt, both
fail:
```
Expected equality of these values:
g_mock_cuda.d2h_count_
Which is: 0
1
Death test: etdump_gen.log_evalue(EValue(tensor))
Result: failed to die.
```
Also reproduced the real crash on one NVIDIA H100, with a small program
that
allocates through `cudaMalloc`, tags a tensor as CUDA and logs it:
```
before: Segmentation fault (core dumped)
after: debug buffer holds 1.5 2.5 3.5 4.5
```
Checked that `etdump_flatcc.cpp` still compiles with `-DUSE_ATEN_LIB`.
`clang-format` reports no changes needed on the four touched C++ files.
### Landing order
pytorch#22058 adds device-planned arenas to the Python bindings that ETDump can
be handed
pointers into, and without this change `BufferDataSink::write` would
`memcpy` device
memory from the host. This should land before or together with it.
### Not covered
The ATen mode branch of the device type conversion only compiles. There
is no
ATen mode CMake build to run it in, so nothing here executes it. It is
reachable
in principle: `runtime/executor/tensor_parser_aten.cpp` reads the
serialized
device type and index, including CUDA, and builds the ATen tensor on
that device.
An earlier version of this description said ATen tensors carry no device
metadata,
which is wrong.
`LogTensorOnCpuDoesNotStageThroughTheAllocator` passes with the
production change
reverted, since CPU tensors already went straight to the data sink. It
documents
the CPU path rather than locking the fix; the other two tests are the
ones that
require the new branch.
---------
Co-authored-by: r <r@e>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

10 participants

@winskuo-quic@JacobSzwejbka@jgibson2@navsud@AlessandroVacca@lucylq@psiddh@DrJessop@hsharma35@kirklandsign
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Dev1/winskuo/gh /aihub remove code - #2

Closed
winskuo-quic wants to merge 21 commits into
mainfrom
dev1/winskuo/gh_/aihub_remove_code
Closed

Dev1/winskuo/gh /aihub remove code#2
winskuo-quic wants to merge 21 commits into
mainfrom
dev1/winskuo/gh_/aihub_remove_code

Conversation

@winskuo-quic

Copy link
Copy Markdown
Collaborator

Summary

[PLEASE REMOVE] See CONTRIBUTING.md's Pull Requests for ExecuTorch PR guidelines.

[PLEASE REMOVE] If this PR closes an issue, please add a Fixes #<issue-id> line.

[PLEASE REMOVE] If this PR introduces a fix or feature that should be the upcoming release notes, please add a "Release notes: " label. For a list of available release notes labels, check out CONTRIBUTING.md's Pull Requests.

Test plan

[PLEASE REMOVE] How did you test this PR? Please write down any manual commands you used and note down tests that you have written if applicable.

JacobSzwejbkaand others added 21 commits April 24, 2026 15:25
Summary:
Attempt 3 to check numel and nbytes overflow. This time we defer
checking dynamic sized inputs until their size is realized.
Reviewed By: lucylq
Differential Revision: D98148157
…inear (pytorch#19117)
## Summary
The bilinear grid_sampler_2d portable kernel computes interpolation
weights via subtractions like `(ix_se - ix)` where both operands are
close integer-valued coordinates in pixel space. In fp16 (10 bits of
mantissa) that's classic catastrophic cancellation — the result has only
a handful of significant bits. The downstream weighted-sum accumulation
then loses further precision.
Measured on a unit test exercising interior grid points with fp16
inputs, the kernel drifts by ~0.1 absolute from an fp32 reference.
That's visible as incorrect depth / flow output near non-integer sample
points, which is most of them.
## Fix
An `AccType<CTYPE>` trait mapping `Half` and `BFloat16` to `float`,
leaving every other dtype unchanged. Used for intermediate coordinate,
weight computation, and `out_val` accumulation. Loads cast `CTYPE ->
ACC`; the final store casts `ACC -> CTYPE` once. Only internal math is
promoted; memory layout / public API / tensor dtypes are unchanged.
```cpp
template <typename CTYPE>
using AccType = std::conditional_t<
std::is_same_v<CTYPE, executorch::aten::Half> ||
std::is_same_v<CTYPE, executorch::aten::BFloat16>,
float,
CTYPE>;
```
## Effects
- **fp32 / Int / any non-half dtype**: `AccType<T>` is `T`, so the
generated code is byte-identical. No behavior change.
- **Half / BFloat16**: `max_abs` vs an fp32 reference drops from **~0.1
to 0** on the shapes I tested (N=1..2, C=7..64, H/W up to 96, both
`align_corners` values).
- **Perf**: a handful of fp16↔fp32 conversions per output element. Not
measurable at op level; well within the portable kernel's scalar cost
envelope.
## Scope
Only touches the bilinear interpolation path. The nearest-mode path
doesn't do weighted-sum accumulation and doesn't have the cancellation
issue — left alone in this change.
## Test plan
- [x] Builds clean for Android arm64 and host (Apple Clang 21).
- [x] Verified numerically via a standalone harness that runs the kernel
with matched fp32 / fp16 inputs and compares against an
fp32-then-downcast reference. All shapes pass within a single fp16 ULP
(or are bit-exact). fp32 tests remain bit-identical to the pre-change
kernel.
- [x] Existing `kernels/test/op_grid_sampler_2d_test.cpp` unit tests
continue to pass (both fp32 shapes that were previously tested, and the
fp16 path I'm specifically fixing).
Happy to add an fp16-specific test case to `op_grid_sampler_2d_test.cpp`
if useful for CI coverage here — just let me know the preferred
approach.
cc @larryliu0820@manuelcandales
Differential Revision: D101728720
Pull Request resolved: pytorch#19052
### Summary
Fixespytorch#18924
Extends `aten.bitwise_not` support in the MLX delegate to handle integer
tensors, not just boolean tensors.
Previously the handler only dispatched to `LogicalNotNode` for `bool`
and raised `NotImplementedError` for all other dtypes. This adds a
dedicated `BitwiseInvertNode` backed by `mlx::core::bitwise_invert`, and
updates the handler to dispatch based on dtype:
- `bool` → `LogicalNotNode` (unchanged)
- `int32`, `int64` → `BitwiseInvertNode`
#### Changes:
- `serialization/schema.fbs`: add `BitwiseInvertNode` table and append
to `OpNode` union
- `runtime/MLXInterpreter.h`: add `exec_bitwise_invert()` and dispatch
case
- `ops.py`: update `_bitwise_not_handler` to dispatch to
`BitwiseInvertNode` for integers
- `test/test_ops.py`: add `bitwise_not_int` test for `int32` and `int64`
### Test plan
All tests were ran on a machine with an Apple M1 Pro CPU, macOS 26.4.1.
- `python3 -m py_compile backends/mlx/ops.py
backends/mlx/test/test_ops.py`
- `python3 backends/mlx/serialization/generate.py`
- `python3 -m executorch.backends.mlx.test.run_all_tests
bitwise_not_int`
### Test output
```
============================================================
TEST SUMMARY
============================================================
Passed: 6
Failed: 0
============================================================
```
cc @metascroy
---------
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Scott Roy <161522778+metascroy@users.noreply.github.com>
Differential Revision: D102070375
Pull Request resolved: pytorch#19074
Add Closeable interface so Module can be used with try-with-resources.
close() delegates to destroy(). Also make destroy() idempotent by
checking mHybridData.isValid() before calling resetNative(), satisfying
the Closeable contract.
This commit was authored with the help of Claude.
TrainingModule: implement Closeable, replace Log.e + silent empty
returns with IllegalStateException throws. Add checkNotDestroyed() guard
on all public methods.
SGD: throw IllegalStateException instead of bare RuntimeException when
optimizer is destroyed.
AsrModule: throw ExecutorchRuntimeException instead of bare
RuntimeException on transcription failure.
ExecuTorchRuntime.validateFilePath: throw IllegalArgumentException
instead of bare RuntimeException, with descriptive message.
JNI constructors: wrap ExecuTorchJni and ExecuTorchLlmJni constructor
bodies in try-catch so C++ exceptions become ExecutorchRuntimeException
instead of generic RuntimeException.
This commit was authored with the help of Claude.
---------
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Differential Revision: D102425537
Pull Request resolved: pytorch#19128
…_nhwc operator + tests
Differential Revision: D96507563
Pull Request resolved: pytorch#18479
…ch#19028 (pytorch#19133)
## Summary
Reverts the following Android PRs:
- pytorch#19099 — Android: consistent error types across all modules
- pytorch#19124 — Android: Module implements Closeable
- pytorch#19092 — Android: improve error diagnostics for LlmModule and
exceptions
- pytorch#19028 — Ignored Module tests: provide required input tensor
Authored with Claude.
Differential Revision: D102488314
Pull Request resolved: pytorch#19134
Differential Revision: D102385104
Pull Request resolved: pytorch#19122
Differential Revision: D102493794
Pull Request resolved: pytorch#19136
@winskuo-quic
winskuo-quic changed the base branch from dev1/winskuo/gh_aihub_remove_readme to mainApril 27, 2026 02:47
@winskuo-quic
winskuo-quic deleted the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:47
@winskuo-quic
winskuo-quic restored the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:48
@winskuo-quic
winskuo-quic deleted the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:49
@winskuo-quic
winskuo-quic restored the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:49
@winskuo-quic
winskuo-quic deleted the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:50
@winskuo-quic
winskuo-quic restored the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:50
quic-boyuc added a commit that referenced this pull request May 5, 2026
Add observe_pass, a decorator that wraps any PassBase subclass or callable
pass to automatically collect input/output graphs via Observatory. Names
are derived from the class/function name and deduplicated by collect()
itself (#2, pytorch#3, ...) so repeated calls never silently overwrite records.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request May 18, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request May 22, 2026
In compare mode each archive column already shows the label in its sticky
header (commit 5ae2a72), so the `<label>/` that compare_archives prepends
to every record and session name renders as redundant noise inside the
column and squeezes the natural-width layout from c9e7bda.
Strip is display-only — `rec.name` / `session.name` in `state.data` stay
prefixed so graph_ref resolution and the global record map keep working.
A `displayPrefix` arg threads through `renderRecordItem`,
`_renderSessionDashboardLink`, `renderIndexFlat`, and `renderIndexTree`;
`_renderArchiveColumn` passes `archive.label`. Single-archive paths pass
no prefix and are unchanged. Collision suffixes (`A/foo #2`) strip cleanly
to `foo #2` since only the head matches.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
Add observe_pass, a decorator that wraps any PassBase subclass or callable
pass to automatically collect input/output graphs via Observatory. Names
are derived from the class/function name and deduplicated by collect()
itself (#2, pytorch#3, ...) so repeated calls never silently overwrite records.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
In compare mode each archive column already shows the label in its sticky
header (commit 5ae2a72), so the `<label>/` that compare_archives prepends
to every record and session name renders as redundant noise inside the
column and squeezes the natural-width layout from c9e7bda.
Strip is display-only — `rec.name` / `session.name` in `state.data` stay
prefixed so graph_ref resolution and the global record map keep working.
A `displayPrefix` arg threads through `renderRecordItem`,
`_renderSessionDashboardLink`, `renderIndexFlat`, and `renderIndexTree`;
`_renderArchiveColumn` passes `archive.label`. Single-archive paths pass
no prefix and are unchanged. Collision suffixes (`A/foo #2`) strip cleanly
to `foo #2` since only the head matches.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
qti-horodnic pushed a commit that referenced this pull request Jul 28, 2026
### Summary
- fill the test gap for newly added op
- utils test
### Test plan
pytest backends/qualcomm/tests/rework/utils/test.py
quic-boyuc pushed a commit that referenced this pull request Aug 25, 2026
…2095)
### Summary
ETDump records an intermediate tensor by handing the tensor's data
pointer to a
data sink, and every sink reads those bytes with a plain host read:
```cpp
memcpy(cur_data_begin, ptr, length);
```
When the tensor lives on an accelerator that pointer is not host memory,
so the
read segfaults. A program placed on CUDA crashes as soon as tracing is
turned
on, which is exactly when someone is trying to debug it.
Returning an error instead would not help. All four callers wrap the
result in
`ET_CHECK_MSG`, so an error aborts the process rather than skipping the
tensor.
This change brings the data back to host memory first. When the tensor
is not on
CPU, ETDump looks up the allocator registered for that device type,
stages the
bytes into a temporary host buffer with `copy_device_to_host`, writes
that buffer
to the sink and frees it. A tensor on CPU keeps the old path and copies
nothing
extra.
If no allocator is registered for the device, ETDump now reports
`NotFound` and
logs the device type instead of reading the pointer anyway.
### Test plan
Two new test files, each with a CMake target and a Buck target.
`devtools/etdump/tests/etdump_device_test.cpp` registers the existing
`MockCudaAllocator`, which backs its device memory with host memory, and
has two
tests. One logs a tensor tagged as CUDA and checks both that ETDump went
through
the allocator and that the bytes reached the debug buffer. The other
logs a
tensor on CPU and checks that the allocator was not used at all, so the
CPU path
is unchanged.
`devtools/etdump/tests/etdump_device_no_allocator_test.cpp` covers the
case where
nothing is registered for the device. The registry is a process wide
static with
no way to remove an entry, so that case needs a binary that never
registers
anything, which is why it is a second file.
`devtools/etdump/tests/CMakeLists.txt` was not referenced by any parent
`CMakeLists.txt`, so nothing in that directory was built by CMake. This
adds
`add_subdirectory(tests)` to `devtools/etdump/CMakeLists.txt` under
`BUILD_TESTING`, so both new tests are picked up by `ctest`, which is
how the C++
tests run.
The pre-existing `sdk_etdump_tests` target stays out of the CMake build.
It
compiles `etdump_test.cpp`, which includes `etdump_filter.h`, which
needs re2,
and the devtools build does not pull re2 in. It is now guarded on re2
being
available rather than being silently unreachable.
With this change both new binaries pass under `ctest`:
```
1/2 Test #1: etdump_device_test ................ Passed
2/2 Test #2: etdump_device_no_allocator_test ... Passed
```
With `etdump_flatcc.cpp` reverted to the old code and everything
rebuilt, both
fail:
```
Expected equality of these values:
g_mock_cuda.d2h_count_
Which is: 0
1
Death test: etdump_gen.log_evalue(EValue(tensor))
Result: failed to die.
```
Also reproduced the real crash on one NVIDIA H100, with a small program
that
allocates through `cudaMalloc`, tags a tensor as CUDA and logs it:
```
before: Segmentation fault (core dumped)
after: debug buffer holds 1.5 2.5 3.5 4.5
```
Checked that `etdump_flatcc.cpp` still compiles with `-DUSE_ATEN_LIB`.
`clang-format` reports no changes needed on the four touched C++ files.
### Landing order
pytorch#22058 adds device-planned arenas to the Python bindings that ETDump can
be handed
pointers into, and without this change `BufferDataSink::write` would
`memcpy` device
memory from the host. This should land before or together with it.
### Not covered
The ATen mode branch of the device type conversion only compiles. There
is no
ATen mode CMake build to run it in, so nothing here executes it. It is
reachable
in principle: `runtime/executor/tensor_parser_aten.cpp` reads the
serialized
device type and index, including CUDA, and builds the ATen tensor on
that device.
An earlier version of this description said ATen tensors carry no device
metadata,
which is wrong.
`LogTensorOnCpuDoesNotStageThroughTheAllocator` passes with the
production change
reverted, since CPU tensors already went straight to the data sink. It
documents
the CPU path rather than locking the fix; the other two tests are the
ones that
require the new branch.
---------
Co-authored-by: r <r@e>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

10 participants

@winskuo-quic@JacobSzwejbka@jgibson2@navsud@AlessandroVacca@lucylq@psiddh@DrJessop@hsharma35@kirklandsign
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Dev1/winskuo/gh /aihub remove code - #2

Closed
winskuo-quic wants to merge 21 commits into
mainfrom
dev1/winskuo/gh_/aihub_remove_code
Closed

Dev1/winskuo/gh /aihub remove code#2
winskuo-quic wants to merge 21 commits into
mainfrom
dev1/winskuo/gh_/aihub_remove_code

Conversation

@winskuo-quic

Copy link
Copy Markdown
Collaborator

Summary

[PLEASE REMOVE] See CONTRIBUTING.md's Pull Requests for ExecuTorch PR guidelines.

[PLEASE REMOVE] If this PR closes an issue, please add a Fixes #<issue-id> line.

[PLEASE REMOVE] If this PR introduces a fix or feature that should be the upcoming release notes, please add a "Release notes: " label. For a list of available release notes labels, check out CONTRIBUTING.md's Pull Requests.

Test plan

[PLEASE REMOVE] How did you test this PR? Please write down any manual commands you used and note down tests that you have written if applicable.

JacobSzwejbkaand others added 21 commits April 24, 2026 15:25
Summary:
Attempt 3 to check numel and nbytes overflow. This time we defer
checking dynamic sized inputs until their size is realized.
Reviewed By: lucylq
Differential Revision: D98148157
…inear (pytorch#19117)
## Summary
The bilinear grid_sampler_2d portable kernel computes interpolation
weights via subtractions like `(ix_se - ix)` where both operands are
close integer-valued coordinates in pixel space. In fp16 (10 bits of
mantissa) that's classic catastrophic cancellation — the result has only
a handful of significant bits. The downstream weighted-sum accumulation
then loses further precision.
Measured on a unit test exercising interior grid points with fp16
inputs, the kernel drifts by ~0.1 absolute from an fp32 reference.
That's visible as incorrect depth / flow output near non-integer sample
points, which is most of them.
## Fix
An `AccType<CTYPE>` trait mapping `Half` and `BFloat16` to `float`,
leaving every other dtype unchanged. Used for intermediate coordinate,
weight computation, and `out_val` accumulation. Loads cast `CTYPE ->
ACC`; the final store casts `ACC -> CTYPE` once. Only internal math is
promoted; memory layout / public API / tensor dtypes are unchanged.
```cpp
template <typename CTYPE>
using AccType = std::conditional_t<
std::is_same_v<CTYPE, executorch::aten::Half> ||
std::is_same_v<CTYPE, executorch::aten::BFloat16>,
float,
CTYPE>;
```
## Effects
- **fp32 / Int / any non-half dtype**: `AccType<T>` is `T`, so the
generated code is byte-identical. No behavior change.
- **Half / BFloat16**: `max_abs` vs an fp32 reference drops from **~0.1
to 0** on the shapes I tested (N=1..2, C=7..64, H/W up to 96, both
`align_corners` values).
- **Perf**: a handful of fp16↔fp32 conversions per output element. Not
measurable at op level; well within the portable kernel's scalar cost
envelope.
## Scope
Only touches the bilinear interpolation path. The nearest-mode path
doesn't do weighted-sum accumulation and doesn't have the cancellation
issue — left alone in this change.
## Test plan
- [x] Builds clean for Android arm64 and host (Apple Clang 21).
- [x] Verified numerically via a standalone harness that runs the kernel
with matched fp32 / fp16 inputs and compares against an
fp32-then-downcast reference. All shapes pass within a single fp16 ULP
(or are bit-exact). fp32 tests remain bit-identical to the pre-change
kernel.
- [x] Existing `kernels/test/op_grid_sampler_2d_test.cpp` unit tests
continue to pass (both fp32 shapes that were previously tested, and the
fp16 path I'm specifically fixing).
Happy to add an fp16-specific test case to `op_grid_sampler_2d_test.cpp`
if useful for CI coverage here — just let me know the preferred
approach.
cc @larryliu0820@manuelcandales
Differential Revision: D101728720
Pull Request resolved: pytorch#19052
### Summary
Fixespytorch#18924
Extends `aten.bitwise_not` support in the MLX delegate to handle integer
tensors, not just boolean tensors.
Previously the handler only dispatched to `LogicalNotNode` for `bool`
and raised `NotImplementedError` for all other dtypes. This adds a
dedicated `BitwiseInvertNode` backed by `mlx::core::bitwise_invert`, and
updates the handler to dispatch based on dtype:
- `bool` → `LogicalNotNode` (unchanged)
- `int32`, `int64` → `BitwiseInvertNode`
#### Changes:
- `serialization/schema.fbs`: add `BitwiseInvertNode` table and append
to `OpNode` union
- `runtime/MLXInterpreter.h`: add `exec_bitwise_invert()` and dispatch
case
- `ops.py`: update `_bitwise_not_handler` to dispatch to
`BitwiseInvertNode` for integers
- `test/test_ops.py`: add `bitwise_not_int` test for `int32` and `int64`
### Test plan
All tests were ran on a machine with an Apple M1 Pro CPU, macOS 26.4.1.
- `python3 -m py_compile backends/mlx/ops.py
backends/mlx/test/test_ops.py`
- `python3 backends/mlx/serialization/generate.py`
- `python3 -m executorch.backends.mlx.test.run_all_tests
bitwise_not_int`
### Test output
```
============================================================
TEST SUMMARY
============================================================
Passed: 6
Failed: 0
============================================================
```
cc @metascroy
---------
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Scott Roy <161522778+metascroy@users.noreply.github.com>
Differential Revision: D102070375
Pull Request resolved: pytorch#19074
Add Closeable interface so Module can be used with try-with-resources.
close() delegates to destroy(). Also make destroy() idempotent by
checking mHybridData.isValid() before calling resetNative(), satisfying
the Closeable contract.
This commit was authored with the help of Claude.
TrainingModule: implement Closeable, replace Log.e + silent empty
returns with IllegalStateException throws. Add checkNotDestroyed() guard
on all public methods.
SGD: throw IllegalStateException instead of bare RuntimeException when
optimizer is destroyed.
AsrModule: throw ExecutorchRuntimeException instead of bare
RuntimeException on transcription failure.
ExecuTorchRuntime.validateFilePath: throw IllegalArgumentException
instead of bare RuntimeException, with descriptive message.
JNI constructors: wrap ExecuTorchJni and ExecuTorchLlmJni constructor
bodies in try-catch so C++ exceptions become ExecutorchRuntimeException
instead of generic RuntimeException.
This commit was authored with the help of Claude.
---------
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Differential Revision: D102425537
Pull Request resolved: pytorch#19128
…_nhwc operator + tests
Differential Revision: D96507563
Pull Request resolved: pytorch#18479
…ch#19028 (pytorch#19133)
## Summary
Reverts the following Android PRs:
- pytorch#19099 — Android: consistent error types across all modules
- pytorch#19124 — Android: Module implements Closeable
- pytorch#19092 — Android: improve error diagnostics for LlmModule and
exceptions
- pytorch#19028 — Ignored Module tests: provide required input tensor
Authored with Claude.
Differential Revision: D102488314
Pull Request resolved: pytorch#19134
Differential Revision: D102385104
Pull Request resolved: pytorch#19122
Differential Revision: D102493794
Pull Request resolved: pytorch#19136
@winskuo-quic
winskuo-quic changed the base branch from dev1/winskuo/gh_aihub_remove_readme to mainApril 27, 2026 02:47
@winskuo-quic
winskuo-quic deleted the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:47
@winskuo-quic
winskuo-quic restored the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:48
@winskuo-quic
winskuo-quic deleted the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:49
@winskuo-quic
winskuo-quic restored the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:49
@winskuo-quic
winskuo-quic deleted the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:50
@winskuo-quic
winskuo-quic restored the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:50
quic-boyuc added a commit that referenced this pull request May 5, 2026
Add observe_pass, a decorator that wraps any PassBase subclass or callable
pass to automatically collect input/output graphs via Observatory. Names
are derived from the class/function name and deduplicated by collect()
itself (#2, pytorch#3, ...) so repeated calls never silently overwrite records.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request May 18, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request May 22, 2026
In compare mode each archive column already shows the label in its sticky
header (commit 5ae2a72), so the `<label>/` that compare_archives prepends
to every record and session name renders as redundant noise inside the
column and squeezes the natural-width layout from c9e7bda.
Strip is display-only — `rec.name` / `session.name` in `state.data` stay
prefixed so graph_ref resolution and the global record map keep working.
A `displayPrefix` arg threads through `renderRecordItem`,
`_renderSessionDashboardLink`, `renderIndexFlat`, and `renderIndexTree`;
`_renderArchiveColumn` passes `archive.label`. Single-archive paths pass
no prefix and are unchanged. Collision suffixes (`A/foo #2`) strip cleanly
to `foo #2` since only the head matches.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
Add observe_pass, a decorator that wraps any PassBase subclass or callable
pass to automatically collect input/output graphs via Observatory. Names
are derived from the class/function name and deduplicated by collect()
itself (#2, pytorch#3, ...) so repeated calls never silently overwrite records.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
In compare mode each archive column already shows the label in its sticky
header (commit 5ae2a72), so the `<label>/` that compare_archives prepends
to every record and session name renders as redundant noise inside the
column and squeezes the natural-width layout from c9e7bda.
Strip is display-only — `rec.name` / `session.name` in `state.data` stay
prefixed so graph_ref resolution and the global record map keep working.
A `displayPrefix` arg threads through `renderRecordItem`,
`_renderSessionDashboardLink`, `renderIndexFlat`, and `renderIndexTree`;
`_renderArchiveColumn` passes `archive.label`. Single-archive paths pass
no prefix and are unchanged. Collision suffixes (`A/foo #2`) strip cleanly
to `foo #2` since only the head matches.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
qti-horodnic pushed a commit that referenced this pull request Jul 28, 2026
### Summary
- fill the test gap for newly added op
- utils test
### Test plan
pytest backends/qualcomm/tests/rework/utils/test.py
quic-boyuc pushed a commit that referenced this pull request Aug 25, 2026
…2095)
### Summary
ETDump records an intermediate tensor by handing the tensor's data
pointer to a
data sink, and every sink reads those bytes with a plain host read:
```cpp
memcpy(cur_data_begin, ptr, length);
```
When the tensor lives on an accelerator that pointer is not host memory,
so the
read segfaults. A program placed on CUDA crashes as soon as tracing is
turned
on, which is exactly when someone is trying to debug it.
Returning an error instead would not help. All four callers wrap the
result in
`ET_CHECK_MSG`, so an error aborts the process rather than skipping the
tensor.
This change brings the data back to host memory first. When the tensor
is not on
CPU, ETDump looks up the allocator registered for that device type,
stages the
bytes into a temporary host buffer with `copy_device_to_host`, writes
that buffer
to the sink and frees it. A tensor on CPU keeps the old path and copies
nothing
extra.
If no allocator is registered for the device, ETDump now reports
`NotFound` and
logs the device type instead of reading the pointer anyway.
### Test plan
Two new test files, each with a CMake target and a Buck target.
`devtools/etdump/tests/etdump_device_test.cpp` registers the existing
`MockCudaAllocator`, which backs its device memory with host memory, and
has two
tests. One logs a tensor tagged as CUDA and checks both that ETDump went
through
the allocator and that the bytes reached the debug buffer. The other
logs a
tensor on CPU and checks that the allocator was not used at all, so the
CPU path
is unchanged.
`devtools/etdump/tests/etdump_device_no_allocator_test.cpp` covers the
case where
nothing is registered for the device. The registry is a process wide
static with
no way to remove an entry, so that case needs a binary that never
registers
anything, which is why it is a second file.
`devtools/etdump/tests/CMakeLists.txt` was not referenced by any parent
`CMakeLists.txt`, so nothing in that directory was built by CMake. This
adds
`add_subdirectory(tests)` to `devtools/etdump/CMakeLists.txt` under
`BUILD_TESTING`, so both new tests are picked up by `ctest`, which is
how the C++
tests run.
The pre-existing `sdk_etdump_tests` target stays out of the CMake build.
It
compiles `etdump_test.cpp`, which includes `etdump_filter.h`, which
needs re2,
and the devtools build does not pull re2 in. It is now guarded on re2
being
available rather than being silently unreachable.
With this change both new binaries pass under `ctest`:
```
1/2 Test #1: etdump_device_test ................ Passed
2/2 Test #2: etdump_device_no_allocator_test ... Passed
```
With `etdump_flatcc.cpp` reverted to the old code and everything
rebuilt, both
fail:
```
Expected equality of these values:
g_mock_cuda.d2h_count_
Which is: 0
1
Death test: etdump_gen.log_evalue(EValue(tensor))
Result: failed to die.
```
Also reproduced the real crash on one NVIDIA H100, with a small program
that
allocates through `cudaMalloc`, tags a tensor as CUDA and logs it:
```
before: Segmentation fault (core dumped)
after: debug buffer holds 1.5 2.5 3.5 4.5
```
Checked that `etdump_flatcc.cpp` still compiles with `-DUSE_ATEN_LIB`.
`clang-format` reports no changes needed on the four touched C++ files.
### Landing order
pytorch#22058 adds device-planned arenas to the Python bindings that ETDump can
be handed
pointers into, and without this change `BufferDataSink::write` would
`memcpy` device
memory from the host. This should land before or together with it.
### Not covered
The ATen mode branch of the device type conversion only compiles. There
is no
ATen mode CMake build to run it in, so nothing here executes it. It is
reachable
in principle: `runtime/executor/tensor_parser_aten.cpp` reads the
serialized
device type and index, including CUDA, and builds the ATen tensor on
that device.
An earlier version of this description said ATen tensors carry no device
metadata,
which is wrong.
`LogTensorOnCpuDoesNotStageThroughTheAllocator` passes with the
production change
reverted, since CPU tensors already went straight to the data sink. It
documents
the CPU path rather than locking the fix; the other two tests are the
ones that
require the new branch.
---------
Co-authored-by: r <r@e>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

10 participants

@winskuo-quic@JacobSzwejbka@jgibson2@navsud@AlessandroVacca@lucylq@psiddh@DrJessop@hsharma35@kirklandsign