[DO NOT MERGE]Qualcomm AI Engine Direct - Mimi Decoder low SQNR reproduce - #1

Draft
winskuo-quic wants to merge 1 commit into
dev1/winskuo/mimi_stage2from
dev1/winskuo/mimi_streaming_low_sqnr_reproduce
Draft

[DO NOT MERGE]Qualcomm AI Engine Direct - Mimi Decoder low SQNR reproduce#1
winskuo-quic wants to merge 1 commit into
dev1/winskuo/mimi_stage2from
dev1/winskuo/mimi_streaming_low_sqnr_reproduce

Conversation

@winskuo-quic

Copy link
Copy Markdown
Collaborator

Summary

Please DO NOT MERGE this PR.
This PR is to reproduce the issue where nn.Module inference is working fine, however, once after torch.export.export, sqnr went from 120 -> 8.

Command:
python examples/qualcomm/oss_scripts/moshi/mimi.py -b build-android -s $device -m SM8650 --chunks_per_batch 125
To run nn.Module where it is working fine, please uncomment line 264-267 and comment out line 269-273

@winskuo-quic
winskuo-quic changed the base branch from dev1/winskuo/mimi_stage2 to mainApril 15, 2025 03:18
@winskuo-quic
winskuo-quic changed the base branch from main to dev1/winskuo/mimi_stage2April 15, 2025 03:19
chenweng-quic pushed a commit that referenced this pull request Jun 9, 2025
Differential Revision: D75911655
Pull Request resolved: pytorch#11344
shewu-quic pushed a commit that referenced this pull request Aug 1, 2025
BNNS copy crashes the process when the dtypes differ
(pytorch#11714).
With the example in this PR
(pytorch#11714), we crash the
process on main. Here is the stack trace from LLDB:
```
Process 19234 stopped
* thread #1, queue = 'com.apple.main-thread', stop reason = signal SIGABRT
frame #0: 0x0000000190ac9388 libsystem_kernel.dylib`__pthread_kill + 8
libsystem_kernel.dylib`__pthread_kill:
-> 0x190ac9388 <+8>: b.lo 0x190ac93a8 ; <+40>
0x190ac938c <+12>: pacibsp 0x190ac9390 <+16>: stp x29, x30, [sp, #-0x10]!
0x190ac9394 <+20>: mov x29, sp
(lldb) bt
* thread #1, queue = 'com.apple.main-thread', stop reason = signal SIGABRT
* frame #0: 0x0000000190ac9388 libsystem_kernel.dylib`__pthread_kill + 8
frame #1: 0x0000000190b0288c libsystem_pthread.dylib`pthread_kill + 296
frame #2: 0x0000000190a0bc60 libsystem_c.dylib`abort + 124
frame pytorch#3: 0x0000000190910174 libsystem_malloc.dylib`malloc_vreport + 892
frame pytorch#4: 0x0000000190913c90 libsystem_malloc.dylib`malloc_report + 64
frame pytorch#5: 0x000000019091821c libsystem_malloc.dylib`___BUG_IN_CLIENT_OF_LIBMALLOC_POINTER_BEING_FREED_WAS_NOT_ALLOCATED + 32
frame pytorch#6: 0x000000019d2f4084 libBNNS.dylib`___lldb_unnamed_symbol1620 + 564
frame pytorch#7: 0x000000019d2f5bac libBNNS.dylib`___lldb_unnamed_symbol1628 + 680
frame pytorch#8: 0x000000019d69ce48 libBNNS.dylib`BNNSCopy + 616
frame pytorch#9: 0x000000030c74d950 _portable_lib.cpython-310-darwin.so`(anonymous namespace)::copy_using_bnns(executorchcoreml::MultiArray const&, executorchcoreml::MultiArray&) + 188
frame pytorch#10: 0x000000030c74cfdc _portable_lib.cpython-310-darwin.so`(anonymous namespace)::copy(executorchcoreml::MultiArray const&, executorchcoreml::MultiArray&, executorchcoreml::MultiArray::CopyOptions) + 72
frame pytorch#11: 0x000000030c74ceec _portable_lib.cpython-310-darwin.so`executorchcoreml::MultiArray::copy(executorchcoreml::MultiArray&, executorchcoreml::MultiArray::CopyOptions) const + 148
frame pytorch#12: 0x000000030c7488d4 _portable_lib.cpython-310-darwin.so`invocation function for block in (anonymous namespace)::copy(MLMultiArray*, executorchcoreml::MultiArray&) + 376
frame pytorch#13: 0x000000030c748ac8 _portable_lib.cpython-310-darwin.so`invocation function for block in (anonymous namespace)::copy(MLMultiArray*, executorchcoreml::MultiArray&) + 52
frame pytorch#14: 0x000000019ad33f4c CoreML`CoreML::MultiArrayBuffer::getBytesWithHandler(void (void const*, unsigned long) block_pointer) const + 340
frame pytorch#15: 0x000000019ad34138 CoreML`-[MLMultiArray(ScopedBufferAccess) getBytesWithHandler:] + 152
frame pytorch#16: 0x000000030c7485ec _portable_lib.cpython-310-darwin.so`(anonymous namespace)::copy(MLMultiArray*, executorchcoreml::MultiArray&) + 296
frame pytorch#17: 0x000000030c744f68 _portable_lib.cpython-310-darwin.so`(anonymous namespace)::set_outputs(std::__1::vector<executorchcoreml::MultiArray, std::__1::allocator<executorchcoreml::MultiArray>>&, NSArray<MLMultiArray*>*) + 180
```
With this PR, the process succeeds.
quic-boyuc added a commit that referenced this pull request May 18, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc pushed a commit that referenced this pull request Aug 25, 2026
…2095)
### Summary
ETDump records an intermediate tensor by handing the tensor's data
pointer to a
data sink, and every sink reads those bytes with a plain host read:
```cpp
memcpy(cur_data_begin, ptr, length);
```
When the tensor lives on an accelerator that pointer is not host memory,
so the
read segfaults. A program placed on CUDA crashes as soon as tracing is
turned
on, which is exactly when someone is trying to debug it.
Returning an error instead would not help. All four callers wrap the
result in
`ET_CHECK_MSG`, so an error aborts the process rather than skipping the
tensor.
This change brings the data back to host memory first. When the tensor
is not on
CPU, ETDump looks up the allocator registered for that device type,
stages the
bytes into a temporary host buffer with `copy_device_to_host`, writes
that buffer
to the sink and frees it. A tensor on CPU keeps the old path and copies
nothing
extra.
If no allocator is registered for the device, ETDump now reports
`NotFound` and
logs the device type instead of reading the pointer anyway.
### Test plan
Two new test files, each with a CMake target and a Buck target.
`devtools/etdump/tests/etdump_device_test.cpp` registers the existing
`MockCudaAllocator`, which backs its device memory with host memory, and
has two
tests. One logs a tensor tagged as CUDA and checks both that ETDump went
through
the allocator and that the bytes reached the debug buffer. The other
logs a
tensor on CPU and checks that the allocator was not used at all, so the
CPU path
is unchanged.
`devtools/etdump/tests/etdump_device_no_allocator_test.cpp` covers the
case where
nothing is registered for the device. The registry is a process wide
static with
no way to remove an entry, so that case needs a binary that never
registers
anything, which is why it is a second file.
`devtools/etdump/tests/CMakeLists.txt` was not referenced by any parent
`CMakeLists.txt`, so nothing in that directory was built by CMake. This
adds
`add_subdirectory(tests)` to `devtools/etdump/CMakeLists.txt` under
`BUILD_TESTING`, so both new tests are picked up by `ctest`, which is
how the C++
tests run.
The pre-existing `sdk_etdump_tests` target stays out of the CMake build.
It
compiles `etdump_test.cpp`, which includes `etdump_filter.h`, which
needs re2,
and the devtools build does not pull re2 in. It is now guarded on re2
being
available rather than being silently unreachable.
With this change both new binaries pass under `ctest`:
```
1/2 Test #1: etdump_device_test ................ Passed
2/2 Test #2: etdump_device_no_allocator_test ... Passed
```
With `etdump_flatcc.cpp` reverted to the old code and everything
rebuilt, both
fail:
```
Expected equality of these values:
g_mock_cuda.d2h_count_
Which is: 0
1
Death test: etdump_gen.log_evalue(EValue(tensor))
Result: failed to die.
```
Also reproduced the real crash on one NVIDIA H100, with a small program
that
allocates through `cudaMalloc`, tags a tensor as CUDA and logs it:
```
before: Segmentation fault (core dumped)
after: debug buffer holds 1.5 2.5 3.5 4.5
```
Checked that `etdump_flatcc.cpp` still compiles with `-DUSE_ATEN_LIB`.
`clang-format` reports no changes needed on the four touched C++ files.
### Landing order
pytorch#22058 adds device-planned arenas to the Python bindings that ETDump can
be handed
pointers into, and without this change `BufferDataSink::write` would
`memcpy` device
memory from the host. This should land before or together with it.
### Not covered
The ATen mode branch of the device type conversion only compiles. There
is no
ATen mode CMake build to run it in, so nothing here executes it. It is
reachable
in principle: `runtime/executor/tensor_parser_aten.cpp` reads the
serialized
device type and index, including CUDA, and builds the ATen tensor on
that device.
An earlier version of this description said ATen tensors carry no device
metadata,
which is wrong.
`LogTensorOnCpuDoesNotStageThroughTheAllocator` passes with the
production change
reverted, since CPU tensors already went straight to the data sink. It
documents
the CPU path rather than locking the fix; the other two tests are the
ones that
require the new branch.
---------
Co-authored-by: r <r@e>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@winskuo-quic
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

[DO NOT MERGE]Qualcomm AI Engine Direct - Mimi Decoder low SQNR reproduce - #1

Draft
winskuo-quic wants to merge 1 commit into
dev1/winskuo/mimi_stage2from
dev1/winskuo/mimi_streaming_low_sqnr_reproduce
Draft

[DO NOT MERGE]Qualcomm AI Engine Direct - Mimi Decoder low SQNR reproduce#1
winskuo-quic wants to merge 1 commit into
dev1/winskuo/mimi_stage2from
dev1/winskuo/mimi_streaming_low_sqnr_reproduce

Conversation

@winskuo-quic

Copy link
Copy Markdown
Collaborator

Summary

Please DO NOT MERGE this PR.
This PR is to reproduce the issue where nn.Module inference is working fine, however, once after torch.export.export, sqnr went from 120 -> 8.

Command:
python examples/qualcomm/oss_scripts/moshi/mimi.py -b build-android -s $device -m SM8650 --chunks_per_batch 125
To run nn.Module where it is working fine, please uncomment line 264-267 and comment out line 269-273

@winskuo-quic
winskuo-quic changed the base branch from dev1/winskuo/mimi_stage2 to mainApril 15, 2025 03:18
@winskuo-quic
winskuo-quic changed the base branch from main to dev1/winskuo/mimi_stage2April 15, 2025 03:19
chenweng-quic pushed a commit that referenced this pull request Jun 9, 2025
Differential Revision: D75911655
Pull Request resolved: pytorch#11344
shewu-quic pushed a commit that referenced this pull request Aug 1, 2025
BNNS copy crashes the process when the dtypes differ
(pytorch#11714).
With the example in this PR
(pytorch#11714), we crash the
process on main. Here is the stack trace from LLDB:
```
Process 19234 stopped
* thread #1, queue = 'com.apple.main-thread', stop reason = signal SIGABRT
frame #0: 0x0000000190ac9388 libsystem_kernel.dylib`__pthread_kill + 8
libsystem_kernel.dylib`__pthread_kill:
-> 0x190ac9388 <+8>: b.lo 0x190ac93a8 ; <+40>
0x190ac938c <+12>: pacibsp 0x190ac9390 <+16>: stp x29, x30, [sp, #-0x10]!
0x190ac9394 <+20>: mov x29, sp
(lldb) bt
* thread #1, queue = 'com.apple.main-thread', stop reason = signal SIGABRT
* frame #0: 0x0000000190ac9388 libsystem_kernel.dylib`__pthread_kill + 8
frame #1: 0x0000000190b0288c libsystem_pthread.dylib`pthread_kill + 296
frame #2: 0x0000000190a0bc60 libsystem_c.dylib`abort + 124
frame pytorch#3: 0x0000000190910174 libsystem_malloc.dylib`malloc_vreport + 892
frame pytorch#4: 0x0000000190913c90 libsystem_malloc.dylib`malloc_report + 64
frame pytorch#5: 0x000000019091821c libsystem_malloc.dylib`___BUG_IN_CLIENT_OF_LIBMALLOC_POINTER_BEING_FREED_WAS_NOT_ALLOCATED + 32
frame pytorch#6: 0x000000019d2f4084 libBNNS.dylib`___lldb_unnamed_symbol1620 + 564
frame pytorch#7: 0x000000019d2f5bac libBNNS.dylib`___lldb_unnamed_symbol1628 + 680
frame pytorch#8: 0x000000019d69ce48 libBNNS.dylib`BNNSCopy + 616
frame pytorch#9: 0x000000030c74d950 _portable_lib.cpython-310-darwin.so`(anonymous namespace)::copy_using_bnns(executorchcoreml::MultiArray const&, executorchcoreml::MultiArray&) + 188
frame pytorch#10: 0x000000030c74cfdc _portable_lib.cpython-310-darwin.so`(anonymous namespace)::copy(executorchcoreml::MultiArray const&, executorchcoreml::MultiArray&, executorchcoreml::MultiArray::CopyOptions) + 72
frame pytorch#11: 0x000000030c74ceec _portable_lib.cpython-310-darwin.so`executorchcoreml::MultiArray::copy(executorchcoreml::MultiArray&, executorchcoreml::MultiArray::CopyOptions) const + 148
frame pytorch#12: 0x000000030c7488d4 _portable_lib.cpython-310-darwin.so`invocation function for block in (anonymous namespace)::copy(MLMultiArray*, executorchcoreml::MultiArray&) + 376
frame pytorch#13: 0x000000030c748ac8 _portable_lib.cpython-310-darwin.so`invocation function for block in (anonymous namespace)::copy(MLMultiArray*, executorchcoreml::MultiArray&) + 52
frame pytorch#14: 0x000000019ad33f4c CoreML`CoreML::MultiArrayBuffer::getBytesWithHandler(void (void const*, unsigned long) block_pointer) const + 340
frame pytorch#15: 0x000000019ad34138 CoreML`-[MLMultiArray(ScopedBufferAccess) getBytesWithHandler:] + 152
frame pytorch#16: 0x000000030c7485ec _portable_lib.cpython-310-darwin.so`(anonymous namespace)::copy(MLMultiArray*, executorchcoreml::MultiArray&) + 296
frame pytorch#17: 0x000000030c744f68 _portable_lib.cpython-310-darwin.so`(anonymous namespace)::set_outputs(std::__1::vector<executorchcoreml::MultiArray, std::__1::allocator<executorchcoreml::MultiArray>>&, NSArray<MLMultiArray*>*) + 180
```
With this PR, the process succeeds.
quic-boyuc added a commit that referenced this pull request May 18, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc pushed a commit that referenced this pull request Aug 25, 2026
…2095)
### Summary
ETDump records an intermediate tensor by handing the tensor's data
pointer to a
data sink, and every sink reads those bytes with a plain host read:
```cpp
memcpy(cur_data_begin, ptr, length);
```
When the tensor lives on an accelerator that pointer is not host memory,
so the
read segfaults. A program placed on CUDA crashes as soon as tracing is
turned
on, which is exactly when someone is trying to debug it.
Returning an error instead would not help. All four callers wrap the
result in
`ET_CHECK_MSG`, so an error aborts the process rather than skipping the
tensor.
This change brings the data back to host memory first. When the tensor
is not on
CPU, ETDump looks up the allocator registered for that device type,
stages the
bytes into a temporary host buffer with `copy_device_to_host`, writes
that buffer
to the sink and frees it. A tensor on CPU keeps the old path and copies
nothing
extra.
If no allocator is registered for the device, ETDump now reports
`NotFound` and
logs the device type instead of reading the pointer anyway.
### Test plan
Two new test files, each with a CMake target and a Buck target.
`devtools/etdump/tests/etdump_device_test.cpp` registers the existing
`MockCudaAllocator`, which backs its device memory with host memory, and
has two
tests. One logs a tensor tagged as CUDA and checks both that ETDump went
through
the allocator and that the bytes reached the debug buffer. The other
logs a
tensor on CPU and checks that the allocator was not used at all, so the
CPU path
is unchanged.
`devtools/etdump/tests/etdump_device_no_allocator_test.cpp` covers the
case where
nothing is registered for the device. The registry is a process wide
static with
no way to remove an entry, so that case needs a binary that never
registers
anything, which is why it is a second file.
`devtools/etdump/tests/CMakeLists.txt` was not referenced by any parent
`CMakeLists.txt`, so nothing in that directory was built by CMake. This
adds
`add_subdirectory(tests)` to `devtools/etdump/CMakeLists.txt` under
`BUILD_TESTING`, so both new tests are picked up by `ctest`, which is
how the C++
tests run.
The pre-existing `sdk_etdump_tests` target stays out of the CMake build.
It
compiles `etdump_test.cpp`, which includes `etdump_filter.h`, which
needs re2,
and the devtools build does not pull re2 in. It is now guarded on re2
being
available rather than being silently unreachable.
With this change both new binaries pass under `ctest`:
```
1/2 Test #1: etdump_device_test ................ Passed
2/2 Test #2: etdump_device_no_allocator_test ... Passed
```
With `etdump_flatcc.cpp` reverted to the old code and everything
rebuilt, both
fail:
```
Expected equality of these values:
g_mock_cuda.d2h_count_
Which is: 0
1
Death test: etdump_gen.log_evalue(EValue(tensor))
Result: failed to die.
```
Also reproduced the real crash on one NVIDIA H100, with a small program
that
allocates through `cudaMalloc`, tags a tensor as CUDA and logs it:
```
before: Segmentation fault (core dumped)
after: debug buffer holds 1.5 2.5 3.5 4.5
```
Checked that `etdump_flatcc.cpp` still compiles with `-DUSE_ATEN_LIB`.
`clang-format` reports no changes needed on the four touched C++ files.
### Landing order
pytorch#22058 adds device-planned arenas to the Python bindings that ETDump can
be handed
pointers into, and without this change `BufferDataSink::write` would
`memcpy` device
memory from the host. This should land before or together with it.
### Not covered
The ATen mode branch of the device type conversion only compiles. There
is no
ATen mode CMake build to run it in, so nothing here executes it. It is
reachable
in principle: `runtime/executor/tensor_parser_aten.cpp` reads the
serialized
device type and index, including CUDA, and builds the ATen tensor on
that device.
An earlier version of this description said ATen tensors carry no device
metadata,
which is wrong.
`LogTensorOnCpuDoesNotStageThroughTheAllocator` passes with the
production change
reverted, since CPU tensors already went straight to the data sink. It
documents
the CPU path rather than locking the fix; the other two tests are the
ones that
require the new branch.
---------
Co-authored-by: r <r@e>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@winskuo-quic
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[DO NOT MERGE]Qualcomm AI Engine Direct - Mimi Decoder low SQNR reproduce - #1

Draft
winskuo-quic wants to merge 1 commit into
dev1/winskuo/mimi_stage2from
dev1/winskuo/mimi_streaming_low_sqnr_reproduce
Draft

[DO NOT MERGE]Qualcomm AI Engine Direct - Mimi Decoder low SQNR reproduce#1
winskuo-quic wants to merge 1 commit into
dev1/winskuo/mimi_stage2from
dev1/winskuo/mimi_streaming_low_sqnr_reproduce

Conversation

@winskuo-quic

Copy link
Copy Markdown
Collaborator

Summary

Please DO NOT MERGE this PR.
This PR is to reproduce the issue where nn.Module inference is working fine, however, once after torch.export.export, sqnr went from 120 -> 8.

Command:
python examples/qualcomm/oss_scripts/moshi/mimi.py -b build-android -s $device -m SM8650 --chunks_per_batch 125
To run nn.Module where it is working fine, please uncomment line 264-267 and comment out line 269-273

@winskuo-quic
winskuo-quic changed the base branch from dev1/winskuo/mimi_stage2 to mainApril 15, 2025 03:18
@winskuo-quic
winskuo-quic changed the base branch from main to dev1/winskuo/mimi_stage2April 15, 2025 03:19
chenweng-quic pushed a commit that referenced this pull request Jun 9, 2025
Differential Revision: D75911655
Pull Request resolved: pytorch#11344
shewu-quic pushed a commit that referenced this pull request Aug 1, 2025
BNNS copy crashes the process when the dtypes differ
(pytorch#11714).
With the example in this PR
(pytorch#11714), we crash the
process on main. Here is the stack trace from LLDB:
```
Process 19234 stopped
* thread #1, queue = 'com.apple.main-thread', stop reason = signal SIGABRT
frame #0: 0x0000000190ac9388 libsystem_kernel.dylib`__pthread_kill + 8
libsystem_kernel.dylib`__pthread_kill:
-> 0x190ac9388 <+8>: b.lo 0x190ac93a8 ; <+40>
0x190ac938c <+12>: pacibsp 0x190ac9390 <+16>: stp x29, x30, [sp, #-0x10]!
0x190ac9394 <+20>: mov x29, sp
(lldb) bt
* thread #1, queue = 'com.apple.main-thread', stop reason = signal SIGABRT
* frame #0: 0x0000000190ac9388 libsystem_kernel.dylib`__pthread_kill + 8
frame #1: 0x0000000190b0288c libsystem_pthread.dylib`pthread_kill + 296
frame #2: 0x0000000190a0bc60 libsystem_c.dylib`abort + 124
frame pytorch#3: 0x0000000190910174 libsystem_malloc.dylib`malloc_vreport + 892
frame pytorch#4: 0x0000000190913c90 libsystem_malloc.dylib`malloc_report + 64
frame pytorch#5: 0x000000019091821c libsystem_malloc.dylib`___BUG_IN_CLIENT_OF_LIBMALLOC_POINTER_BEING_FREED_WAS_NOT_ALLOCATED + 32
frame pytorch#6: 0x000000019d2f4084 libBNNS.dylib`___lldb_unnamed_symbol1620 + 564
frame pytorch#7: 0x000000019d2f5bac libBNNS.dylib`___lldb_unnamed_symbol1628 + 680
frame pytorch#8: 0x000000019d69ce48 libBNNS.dylib`BNNSCopy + 616
frame pytorch#9: 0x000000030c74d950 _portable_lib.cpython-310-darwin.so`(anonymous namespace)::copy_using_bnns(executorchcoreml::MultiArray const&, executorchcoreml::MultiArray&) + 188
frame pytorch#10: 0x000000030c74cfdc _portable_lib.cpython-310-darwin.so`(anonymous namespace)::copy(executorchcoreml::MultiArray const&, executorchcoreml::MultiArray&, executorchcoreml::MultiArray::CopyOptions) + 72
frame pytorch#11: 0x000000030c74ceec _portable_lib.cpython-310-darwin.so`executorchcoreml::MultiArray::copy(executorchcoreml::MultiArray&, executorchcoreml::MultiArray::CopyOptions) const + 148
frame pytorch#12: 0x000000030c7488d4 _portable_lib.cpython-310-darwin.so`invocation function for block in (anonymous namespace)::copy(MLMultiArray*, executorchcoreml::MultiArray&) + 376
frame pytorch#13: 0x000000030c748ac8 _portable_lib.cpython-310-darwin.so`invocation function for block in (anonymous namespace)::copy(MLMultiArray*, executorchcoreml::MultiArray&) + 52
frame pytorch#14: 0x000000019ad33f4c CoreML`CoreML::MultiArrayBuffer::getBytesWithHandler(void (void const*, unsigned long) block_pointer) const + 340
frame pytorch#15: 0x000000019ad34138 CoreML`-[MLMultiArray(ScopedBufferAccess) getBytesWithHandler:] + 152
frame pytorch#16: 0x000000030c7485ec _portable_lib.cpython-310-darwin.so`(anonymous namespace)::copy(MLMultiArray*, executorchcoreml::MultiArray&) + 296
frame pytorch#17: 0x000000030c744f68 _portable_lib.cpython-310-darwin.so`(anonymous namespace)::set_outputs(std::__1::vector<executorchcoreml::MultiArray, std::__1::allocator<executorchcoreml::MultiArray>>&, NSArray<MLMultiArray*>*) + 180
```
With this PR, the process succeeds.
quic-boyuc added a commit that referenced this pull request May 18, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc pushed a commit that referenced this pull request Aug 25, 2026
…2095)
### Summary
ETDump records an intermediate tensor by handing the tensor's data
pointer to a
data sink, and every sink reads those bytes with a plain host read:
```cpp
memcpy(cur_data_begin, ptr, length);
```
When the tensor lives on an accelerator that pointer is not host memory,
so the
read segfaults. A program placed on CUDA crashes as soon as tracing is
turned
on, which is exactly when someone is trying to debug it.
Returning an error instead would not help. All four callers wrap the
result in
`ET_CHECK_MSG`, so an error aborts the process rather than skipping the
tensor.
This change brings the data back to host memory first. When the tensor
is not on
CPU, ETDump looks up the allocator registered for that device type,
stages the
bytes into a temporary host buffer with `copy_device_to_host`, writes
that buffer
to the sink and frees it. A tensor on CPU keeps the old path and copies
nothing
extra.
If no allocator is registered for the device, ETDump now reports
`NotFound` and
logs the device type instead of reading the pointer anyway.
### Test plan
Two new test files, each with a CMake target and a Buck target.
`devtools/etdump/tests/etdump_device_test.cpp` registers the existing
`MockCudaAllocator`, which backs its device memory with host memory, and
has two
tests. One logs a tensor tagged as CUDA and checks both that ETDump went
through
the allocator and that the bytes reached the debug buffer. The other
logs a
tensor on CPU and checks that the allocator was not used at all, so the
CPU path
is unchanged.
`devtools/etdump/tests/etdump_device_no_allocator_test.cpp` covers the
case where
nothing is registered for the device. The registry is a process wide
static with
no way to remove an entry, so that case needs a binary that never
registers
anything, which is why it is a second file.
`devtools/etdump/tests/CMakeLists.txt` was not referenced by any parent
`CMakeLists.txt`, so nothing in that directory was built by CMake. This
adds
`add_subdirectory(tests)` to `devtools/etdump/CMakeLists.txt` under
`BUILD_TESTING`, so both new tests are picked up by `ctest`, which is
how the C++
tests run.
The pre-existing `sdk_etdump_tests` target stays out of the CMake build.
It
compiles `etdump_test.cpp`, which includes `etdump_filter.h`, which
needs re2,
and the devtools build does not pull re2 in. It is now guarded on re2
being
available rather than being silently unreachable.
With this change both new binaries pass under `ctest`:
```
1/2 Test #1: etdump_device_test ................ Passed
2/2 Test #2: etdump_device_no_allocator_test ... Passed
```
With `etdump_flatcc.cpp` reverted to the old code and everything
rebuilt, both
fail:
```
Expected equality of these values:
g_mock_cuda.d2h_count_
Which is: 0
1
Death test: etdump_gen.log_evalue(EValue(tensor))
Result: failed to die.
```
Also reproduced the real crash on one NVIDIA H100, with a small program
that
allocates through `cudaMalloc`, tags a tensor as CUDA and logs it:
```
before: Segmentation fault (core dumped)
after: debug buffer holds 1.5 2.5 3.5 4.5
```
Checked that `etdump_flatcc.cpp` still compiles with `-DUSE_ATEN_LIB`.
`clang-format` reports no changes needed on the four touched C++ files.
### Landing order
pytorch#22058 adds device-planned arenas to the Python bindings that ETDump can
be handed
pointers into, and without this change `BufferDataSink::write` would
`memcpy` device
memory from the host. This should land before or together with it.
### Not covered
The ATen mode branch of the device type conversion only compiles. There
is no
ATen mode CMake build to run it in, so nothing here executes it. It is
reachable
in principle: `runtime/executor/tensor_parser_aten.cpp` reads the
serialized
device type and index, including CUDA, and builds the ATen tensor on
that device.
An earlier version of this description said ATen tensors carry no device
metadata,
which is wrong.
`LogTensorOnCpuDoesNotStageThroughTheAllocator` passes with the
production change
reverted, since CPU tensors already went straight to the data sink. It
documents
the CPU path rather than locking the fix; the other two tests are the
ones that
require the new branch.
---------
Co-authored-by: r <r@e>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@winskuo-quic
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[DO NOT MERGE]Qualcomm AI Engine Direct - Mimi Decoder low SQNR reproduce - #1

Draft
winskuo-quic wants to merge 1 commit into
dev1/winskuo/mimi_stage2from
dev1/winskuo/mimi_streaming_low_sqnr_reproduce
Draft

[DO NOT MERGE]Qualcomm AI Engine Direct - Mimi Decoder low SQNR reproduce#1
winskuo-quic wants to merge 1 commit into
dev1/winskuo/mimi_stage2from
dev1/winskuo/mimi_streaming_low_sqnr_reproduce

Conversation

@winskuo-quic

Copy link
Copy Markdown
Collaborator

Summary

Please DO NOT MERGE this PR.
This PR is to reproduce the issue where nn.Module inference is working fine, however, once after torch.export.export, sqnr went from 120 -> 8.

Command:
python examples/qualcomm/oss_scripts/moshi/mimi.py -b build-android -s $device -m SM8650 --chunks_per_batch 125
To run nn.Module where it is working fine, please uncomment line 264-267 and comment out line 269-273

@winskuo-quic
winskuo-quic changed the base branch from dev1/winskuo/mimi_stage2 to mainApril 15, 2025 03:18
@winskuo-quic
winskuo-quic changed the base branch from main to dev1/winskuo/mimi_stage2April 15, 2025 03:19
chenweng-quic pushed a commit that referenced this pull request Jun 9, 2025
Differential Revision: D75911655
Pull Request resolved: pytorch#11344
shewu-quic pushed a commit that referenced this pull request Aug 1, 2025
BNNS copy crashes the process when the dtypes differ
(pytorch#11714).
With the example in this PR
(pytorch#11714), we crash the
process on main. Here is the stack trace from LLDB:
```
Process 19234 stopped
* thread #1, queue = 'com.apple.main-thread', stop reason = signal SIGABRT
frame #0: 0x0000000190ac9388 libsystem_kernel.dylib`__pthread_kill + 8
libsystem_kernel.dylib`__pthread_kill:
-> 0x190ac9388 <+8>: b.lo 0x190ac93a8 ; <+40>
0x190ac938c <+12>: pacibsp 0x190ac9390 <+16>: stp x29, x30, [sp, #-0x10]!
0x190ac9394 <+20>: mov x29, sp
(lldb) bt
* thread #1, queue = 'com.apple.main-thread', stop reason = signal SIGABRT
* frame #0: 0x0000000190ac9388 libsystem_kernel.dylib`__pthread_kill + 8
frame #1: 0x0000000190b0288c libsystem_pthread.dylib`pthread_kill + 296
frame #2: 0x0000000190a0bc60 libsystem_c.dylib`abort + 124
frame pytorch#3: 0x0000000190910174 libsystem_malloc.dylib`malloc_vreport + 892
frame pytorch#4: 0x0000000190913c90 libsystem_malloc.dylib`malloc_report + 64
frame pytorch#5: 0x000000019091821c libsystem_malloc.dylib`___BUG_IN_CLIENT_OF_LIBMALLOC_POINTER_BEING_FREED_WAS_NOT_ALLOCATED + 32
frame pytorch#6: 0x000000019d2f4084 libBNNS.dylib`___lldb_unnamed_symbol1620 + 564
frame pytorch#7: 0x000000019d2f5bac libBNNS.dylib`___lldb_unnamed_symbol1628 + 680
frame pytorch#8: 0x000000019d69ce48 libBNNS.dylib`BNNSCopy + 616
frame pytorch#9: 0x000000030c74d950 _portable_lib.cpython-310-darwin.so`(anonymous namespace)::copy_using_bnns(executorchcoreml::MultiArray const&, executorchcoreml::MultiArray&) + 188
frame pytorch#10: 0x000000030c74cfdc _portable_lib.cpython-310-darwin.so`(anonymous namespace)::copy(executorchcoreml::MultiArray const&, executorchcoreml::MultiArray&, executorchcoreml::MultiArray::CopyOptions) + 72
frame pytorch#11: 0x000000030c74ceec _portable_lib.cpython-310-darwin.so`executorchcoreml::MultiArray::copy(executorchcoreml::MultiArray&, executorchcoreml::MultiArray::CopyOptions) const + 148
frame pytorch#12: 0x000000030c7488d4 _portable_lib.cpython-310-darwin.so`invocation function for block in (anonymous namespace)::copy(MLMultiArray*, executorchcoreml::MultiArray&) + 376
frame pytorch#13: 0x000000030c748ac8 _portable_lib.cpython-310-darwin.so`invocation function for block in (anonymous namespace)::copy(MLMultiArray*, executorchcoreml::MultiArray&) + 52
frame pytorch#14: 0x000000019ad33f4c CoreML`CoreML::MultiArrayBuffer::getBytesWithHandler(void (void const*, unsigned long) block_pointer) const + 340
frame pytorch#15: 0x000000019ad34138 CoreML`-[MLMultiArray(ScopedBufferAccess) getBytesWithHandler:] + 152
frame pytorch#16: 0x000000030c7485ec _portable_lib.cpython-310-darwin.so`(anonymous namespace)::copy(MLMultiArray*, executorchcoreml::MultiArray&) + 296
frame pytorch#17: 0x000000030c744f68 _portable_lib.cpython-310-darwin.so`(anonymous namespace)::set_outputs(std::__1::vector<executorchcoreml::MultiArray, std::__1::allocator<executorchcoreml::MultiArray>>&, NSArray<MLMultiArray*>*) + 180
```
With this PR, the process succeeds.
quic-boyuc added a commit that referenced this pull request May 18, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc pushed a commit that referenced this pull request Aug 25, 2026
…2095)
### Summary
ETDump records an intermediate tensor by handing the tensor's data
pointer to a
data sink, and every sink reads those bytes with a plain host read:
```cpp
memcpy(cur_data_begin, ptr, length);
```
When the tensor lives on an accelerator that pointer is not host memory,
so the
read segfaults. A program placed on CUDA crashes as soon as tracing is
turned
on, which is exactly when someone is trying to debug it.
Returning an error instead would not help. All four callers wrap the
result in
`ET_CHECK_MSG`, so an error aborts the process rather than skipping the
tensor.
This change brings the data back to host memory first. When the tensor
is not on
CPU, ETDump looks up the allocator registered for that device type,
stages the
bytes into a temporary host buffer with `copy_device_to_host`, writes
that buffer
to the sink and frees it. A tensor on CPU keeps the old path and copies
nothing
extra.
If no allocator is registered for the device, ETDump now reports
`NotFound` and
logs the device type instead of reading the pointer anyway.
### Test plan
Two new test files, each with a CMake target and a Buck target.
`devtools/etdump/tests/etdump_device_test.cpp` registers the existing
`MockCudaAllocator`, which backs its device memory with host memory, and
has two
tests. One logs a tensor tagged as CUDA and checks both that ETDump went
through
the allocator and that the bytes reached the debug buffer. The other
logs a
tensor on CPU and checks that the allocator was not used at all, so the
CPU path
is unchanged.
`devtools/etdump/tests/etdump_device_no_allocator_test.cpp` covers the
case where
nothing is registered for the device. The registry is a process wide
static with
no way to remove an entry, so that case needs a binary that never
registers
anything, which is why it is a second file.
`devtools/etdump/tests/CMakeLists.txt` was not referenced by any parent
`CMakeLists.txt`, so nothing in that directory was built by CMake. This
adds
`add_subdirectory(tests)` to `devtools/etdump/CMakeLists.txt` under
`BUILD_TESTING`, so both new tests are picked up by `ctest`, which is
how the C++
tests run.
The pre-existing `sdk_etdump_tests` target stays out of the CMake build.
It
compiles `etdump_test.cpp`, which includes `etdump_filter.h`, which
needs re2,
and the devtools build does not pull re2 in. It is now guarded on re2
being
available rather than being silently unreachable.
With this change both new binaries pass under `ctest`:
```
1/2 Test #1: etdump_device_test ................ Passed
2/2 Test #2: etdump_device_no_allocator_test ... Passed
```
With `etdump_flatcc.cpp` reverted to the old code and everything
rebuilt, both
fail:
```
Expected equality of these values:
g_mock_cuda.d2h_count_
Which is: 0
1
Death test: etdump_gen.log_evalue(EValue(tensor))
Result: failed to die.
```
Also reproduced the real crash on one NVIDIA H100, with a small program
that
allocates through `cudaMalloc`, tags a tensor as CUDA and logs it:
```
before: Segmentation fault (core dumped)
after: debug buffer holds 1.5 2.5 3.5 4.5
```
Checked that `etdump_flatcc.cpp` still compiles with `-DUSE_ATEN_LIB`.
`clang-format` reports no changes needed on the four touched C++ files.
### Landing order
pytorch#22058 adds device-planned arenas to the Python bindings that ETDump can
be handed
pointers into, and without this change `BufferDataSink::write` would
`memcpy` device
memory from the host. This should land before or together with it.
### Not covered
The ATen mode branch of the device type conversion only compiles. There
is no
ATen mode CMake build to run it in, so nothing here executes it. It is
reachable
in principle: `runtime/executor/tensor_parser_aten.cpp` reads the
serialized
device type and index, including CUDA, and builds the ATen tensor on
that device.
An earlier version of this description said ATen tensors carry no device
metadata,
which is wrong.
`LogTensorOnCpuDoesNotStageThroughTheAllocator` passes with the
production change
reverted, since CPU tensors already went straight to the data sink. It
documents
the CPU path rather than locking the fix; the other two tests are the
ones that
require the new branch.
---------
Co-authored-by: r <r@e>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@winskuo-quic
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

[DO NOT MERGE]Qualcomm AI Engine Direct - Mimi Decoder low SQNR reproduce - #1

Draft
winskuo-quic wants to merge 1 commit into
dev1/winskuo/mimi_stage2from
dev1/winskuo/mimi_streaming_low_sqnr_reproduce
Draft

[DO NOT MERGE]Qualcomm AI Engine Direct - Mimi Decoder low SQNR reproduce#1
winskuo-quic wants to merge 1 commit into
dev1/winskuo/mimi_stage2from
dev1/winskuo/mimi_streaming_low_sqnr_reproduce

Conversation

@winskuo-quic

Copy link
Copy Markdown
Collaborator

Summary

Please DO NOT MERGE this PR.
This PR is to reproduce the issue where nn.Module inference is working fine, however, once after torch.export.export, sqnr went from 120 -> 8.

Command:
python examples/qualcomm/oss_scripts/moshi/mimi.py -b build-android -s $device -m SM8650 --chunks_per_batch 125
To run nn.Module where it is working fine, please uncomment line 264-267 and comment out line 269-273

@winskuo-quic
winskuo-quic changed the base branch from dev1/winskuo/mimi_stage2 to mainApril 15, 2025 03:18
@winskuo-quic
winskuo-quic changed the base branch from main to dev1/winskuo/mimi_stage2April 15, 2025 03:19
chenweng-quic pushed a commit that referenced this pull request Jun 9, 2025
Differential Revision: D75911655
Pull Request resolved: pytorch#11344
shewu-quic pushed a commit that referenced this pull request Aug 1, 2025
BNNS copy crashes the process when the dtypes differ
(pytorch#11714).
With the example in this PR
(pytorch#11714), we crash the
process on main. Here is the stack trace from LLDB:
```
Process 19234 stopped
* thread #1, queue = 'com.apple.main-thread', stop reason = signal SIGABRT
frame #0: 0x0000000190ac9388 libsystem_kernel.dylib`__pthread_kill + 8
libsystem_kernel.dylib`__pthread_kill:
-> 0x190ac9388 <+8>: b.lo 0x190ac93a8 ; <+40>
0x190ac938c <+12>: pacibsp 0x190ac9390 <+16>: stp x29, x30, [sp, #-0x10]!
0x190ac9394 <+20>: mov x29, sp
(lldb) bt
* thread #1, queue = 'com.apple.main-thread', stop reason = signal SIGABRT
* frame #0: 0x0000000190ac9388 libsystem_kernel.dylib`__pthread_kill + 8
frame #1: 0x0000000190b0288c libsystem_pthread.dylib`pthread_kill + 296
frame #2: 0x0000000190a0bc60 libsystem_c.dylib`abort + 124
frame pytorch#3: 0x0000000190910174 libsystem_malloc.dylib`malloc_vreport + 892
frame pytorch#4: 0x0000000190913c90 libsystem_malloc.dylib`malloc_report + 64
frame pytorch#5: 0x000000019091821c libsystem_malloc.dylib`___BUG_IN_CLIENT_OF_LIBMALLOC_POINTER_BEING_FREED_WAS_NOT_ALLOCATED + 32
frame pytorch#6: 0x000000019d2f4084 libBNNS.dylib`___lldb_unnamed_symbol1620 + 564
frame pytorch#7: 0x000000019d2f5bac libBNNS.dylib`___lldb_unnamed_symbol1628 + 680
frame pytorch#8: 0x000000019d69ce48 libBNNS.dylib`BNNSCopy + 616
frame pytorch#9: 0x000000030c74d950 _portable_lib.cpython-310-darwin.so`(anonymous namespace)::copy_using_bnns(executorchcoreml::MultiArray const&, executorchcoreml::MultiArray&) + 188
frame pytorch#10: 0x000000030c74cfdc _portable_lib.cpython-310-darwin.so`(anonymous namespace)::copy(executorchcoreml::MultiArray const&, executorchcoreml::MultiArray&, executorchcoreml::MultiArray::CopyOptions) + 72
frame pytorch#11: 0x000000030c74ceec _portable_lib.cpython-310-darwin.so`executorchcoreml::MultiArray::copy(executorchcoreml::MultiArray&, executorchcoreml::MultiArray::CopyOptions) const + 148
frame pytorch#12: 0x000000030c7488d4 _portable_lib.cpython-310-darwin.so`invocation function for block in (anonymous namespace)::copy(MLMultiArray*, executorchcoreml::MultiArray&) + 376
frame pytorch#13: 0x000000030c748ac8 _portable_lib.cpython-310-darwin.so`invocation function for block in (anonymous namespace)::copy(MLMultiArray*, executorchcoreml::MultiArray&) + 52
frame pytorch#14: 0x000000019ad33f4c CoreML`CoreML::MultiArrayBuffer::getBytesWithHandler(void (void const*, unsigned long) block_pointer) const + 340
frame pytorch#15: 0x000000019ad34138 CoreML`-[MLMultiArray(ScopedBufferAccess) getBytesWithHandler:] + 152
frame pytorch#16: 0x000000030c7485ec _portable_lib.cpython-310-darwin.so`(anonymous namespace)::copy(MLMultiArray*, executorchcoreml::MultiArray&) + 296
frame pytorch#17: 0x000000030c744f68 _portable_lib.cpython-310-darwin.so`(anonymous namespace)::set_outputs(std::__1::vector<executorchcoreml::MultiArray, std::__1::allocator<executorchcoreml::MultiArray>>&, NSArray<MLMultiArray*>*) + 180
```
With this PR, the process succeeds.
quic-boyuc added a commit that referenced this pull request May 18, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc pushed a commit that referenced this pull request Aug 25, 2026
…2095)
### Summary
ETDump records an intermediate tensor by handing the tensor's data
pointer to a
data sink, and every sink reads those bytes with a plain host read:
```cpp
memcpy(cur_data_begin, ptr, length);
```
When the tensor lives on an accelerator that pointer is not host memory,
so the
read segfaults. A program placed on CUDA crashes as soon as tracing is
turned
on, which is exactly when someone is trying to debug it.
Returning an error instead would not help. All four callers wrap the
result in
`ET_CHECK_MSG`, so an error aborts the process rather than skipping the
tensor.
This change brings the data back to host memory first. When the tensor
is not on
CPU, ETDump looks up the allocator registered for that device type,
stages the
bytes into a temporary host buffer with `copy_device_to_host`, writes
that buffer
to the sink and frees it. A tensor on CPU keeps the old path and copies
nothing
extra.
If no allocator is registered for the device, ETDump now reports
`NotFound` and
logs the device type instead of reading the pointer anyway.
### Test plan
Two new test files, each with a CMake target and a Buck target.
`devtools/etdump/tests/etdump_device_test.cpp` registers the existing
`MockCudaAllocator`, which backs its device memory with host memory, and
has two
tests. One logs a tensor tagged as CUDA and checks both that ETDump went
through
the allocator and that the bytes reached the debug buffer. The other
logs a
tensor on CPU and checks that the allocator was not used at all, so the
CPU path
is unchanged.
`devtools/etdump/tests/etdump_device_no_allocator_test.cpp` covers the
case where
nothing is registered for the device. The registry is a process wide
static with
no way to remove an entry, so that case needs a binary that never
registers
anything, which is why it is a second file.
`devtools/etdump/tests/CMakeLists.txt` was not referenced by any parent
`CMakeLists.txt`, so nothing in that directory was built by CMake. This
adds
`add_subdirectory(tests)` to `devtools/etdump/CMakeLists.txt` under
`BUILD_TESTING`, so both new tests are picked up by `ctest`, which is
how the C++
tests run.
The pre-existing `sdk_etdump_tests` target stays out of the CMake build.
It
compiles `etdump_test.cpp`, which includes `etdump_filter.h`, which
needs re2,
and the devtools build does not pull re2 in. It is now guarded on re2
being
available rather than being silently unreachable.
With this change both new binaries pass under `ctest`:
```
1/2 Test #1: etdump_device_test ................ Passed
2/2 Test #2: etdump_device_no_allocator_test ... Passed
```
With `etdump_flatcc.cpp` reverted to the old code and everything
rebuilt, both
fail:
```
Expected equality of these values:
g_mock_cuda.d2h_count_
Which is: 0
1
Death test: etdump_gen.log_evalue(EValue(tensor))
Result: failed to die.
```
Also reproduced the real crash on one NVIDIA H100, with a small program
that
allocates through `cudaMalloc`, tags a tensor as CUDA and logs it:
```
before: Segmentation fault (core dumped)
after: debug buffer holds 1.5 2.5 3.5 4.5
```
Checked that `etdump_flatcc.cpp` still compiles with `-DUSE_ATEN_LIB`.
`clang-format` reports no changes needed on the four touched C++ files.
### Landing order
pytorch#22058 adds device-planned arenas to the Python bindings that ETDump can
be handed
pointers into, and without this change `BufferDataSink::write` would
`memcpy` device
memory from the host. This should land before or together with it.
### Not covered
The ATen mode branch of the device type conversion only compiles. There
is no
ATen mode CMake build to run it in, so nothing here executes it. It is
reachable
in principle: `runtime/executor/tensor_parser_aten.cpp` reads the
serialized
device type and index, including CUDA, and builds the ATen tensor on
that device.
An earlier version of this description said ATen tensors carry no device
metadata,
which is wrong.
`LogTensorOnCpuDoesNotStageThroughTheAllocator` passes with the
production change
reverted, since CPU tensors already went straight to the data sink. It
documents
the CPU path rather than locking the fix; the other two tests are the
ones that
require the new branch.
---------
Co-authored-by: r <r@e>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@winskuo-quic
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[DO NOT MERGE]Qualcomm AI Engine Direct - Mimi Decoder low SQNR reproduce - #1

Draft
winskuo-quic wants to merge 1 commit into
dev1/winskuo/mimi_stage2from
dev1/winskuo/mimi_streaming_low_sqnr_reproduce
Draft

[DO NOT MERGE]Qualcomm AI Engine Direct - Mimi Decoder low SQNR reproduce#1
winskuo-quic wants to merge 1 commit into
dev1/winskuo/mimi_stage2from
dev1/winskuo/mimi_streaming_low_sqnr_reproduce

Conversation

@winskuo-quic

Copy link
Copy Markdown
Collaborator

Summary

Please DO NOT MERGE this PR.
This PR is to reproduce the issue where nn.Module inference is working fine, however, once after torch.export.export, sqnr went from 120 -> 8.

Command:
python examples/qualcomm/oss_scripts/moshi/mimi.py -b build-android -s $device -m SM8650 --chunks_per_batch 125
To run nn.Module where it is working fine, please uncomment line 264-267 and comment out line 269-273

@winskuo-quic
winskuo-quic changed the base branch from dev1/winskuo/mimi_stage2 to mainApril 15, 2025 03:18
@winskuo-quic
winskuo-quic changed the base branch from main to dev1/winskuo/mimi_stage2April 15, 2025 03:19
chenweng-quic pushed a commit that referenced this pull request Jun 9, 2025
Differential Revision: D75911655
Pull Request resolved: pytorch#11344
shewu-quic pushed a commit that referenced this pull request Aug 1, 2025
BNNS copy crashes the process when the dtypes differ
(pytorch#11714).
With the example in this PR
(pytorch#11714), we crash the
process on main. Here is the stack trace from LLDB:
```
Process 19234 stopped
* thread #1, queue = 'com.apple.main-thread', stop reason = signal SIGABRT
frame #0: 0x0000000190ac9388 libsystem_kernel.dylib`__pthread_kill + 8
libsystem_kernel.dylib`__pthread_kill:
-> 0x190ac9388 <+8>: b.lo 0x190ac93a8 ; <+40>
0x190ac938c <+12>: pacibsp 0x190ac9390 <+16>: stp x29, x30, [sp, #-0x10]!
0x190ac9394 <+20>: mov x29, sp
(lldb) bt
* thread #1, queue = 'com.apple.main-thread', stop reason = signal SIGABRT
* frame #0: 0x0000000190ac9388 libsystem_kernel.dylib`__pthread_kill + 8
frame #1: 0x0000000190b0288c libsystem_pthread.dylib`pthread_kill + 296
frame #2: 0x0000000190a0bc60 libsystem_c.dylib`abort + 124
frame pytorch#3: 0x0000000190910174 libsystem_malloc.dylib`malloc_vreport + 892
frame pytorch#4: 0x0000000190913c90 libsystem_malloc.dylib`malloc_report + 64
frame pytorch#5: 0x000000019091821c libsystem_malloc.dylib`___BUG_IN_CLIENT_OF_LIBMALLOC_POINTER_BEING_FREED_WAS_NOT_ALLOCATED + 32
frame pytorch#6: 0x000000019d2f4084 libBNNS.dylib`___lldb_unnamed_symbol1620 + 564
frame pytorch#7: 0x000000019d2f5bac libBNNS.dylib`___lldb_unnamed_symbol1628 + 680
frame pytorch#8: 0x000000019d69ce48 libBNNS.dylib`BNNSCopy + 616
frame pytorch#9: 0x000000030c74d950 _portable_lib.cpython-310-darwin.so`(anonymous namespace)::copy_using_bnns(executorchcoreml::MultiArray const&, executorchcoreml::MultiArray&) + 188
frame pytorch#10: 0x000000030c74cfdc _portable_lib.cpython-310-darwin.so`(anonymous namespace)::copy(executorchcoreml::MultiArray const&, executorchcoreml::MultiArray&, executorchcoreml::MultiArray::CopyOptions) + 72
frame pytorch#11: 0x000000030c74ceec _portable_lib.cpython-310-darwin.so`executorchcoreml::MultiArray::copy(executorchcoreml::MultiArray&, executorchcoreml::MultiArray::CopyOptions) const + 148
frame pytorch#12: 0x000000030c7488d4 _portable_lib.cpython-310-darwin.so`invocation function for block in (anonymous namespace)::copy(MLMultiArray*, executorchcoreml::MultiArray&) + 376
frame pytorch#13: 0x000000030c748ac8 _portable_lib.cpython-310-darwin.so`invocation function for block in (anonymous namespace)::copy(MLMultiArray*, executorchcoreml::MultiArray&) + 52
frame pytorch#14: 0x000000019ad33f4c CoreML`CoreML::MultiArrayBuffer::getBytesWithHandler(void (void const*, unsigned long) block_pointer) const + 340
frame pytorch#15: 0x000000019ad34138 CoreML`-[MLMultiArray(ScopedBufferAccess) getBytesWithHandler:] + 152
frame pytorch#16: 0x000000030c7485ec _portable_lib.cpython-310-darwin.so`(anonymous namespace)::copy(MLMultiArray*, executorchcoreml::MultiArray&) + 296
frame pytorch#17: 0x000000030c744f68 _portable_lib.cpython-310-darwin.so`(anonymous namespace)::set_outputs(std::__1::vector<executorchcoreml::MultiArray, std::__1::allocator<executorchcoreml::MultiArray>>&, NSArray<MLMultiArray*>*) + 180
```
With this PR, the process succeeds.
quic-boyuc added a commit that referenced this pull request May 18, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc pushed a commit that referenced this pull request Aug 25, 2026
…2095)
### Summary
ETDump records an intermediate tensor by handing the tensor's data
pointer to a
data sink, and every sink reads those bytes with a plain host read:
```cpp
memcpy(cur_data_begin, ptr, length);
```
When the tensor lives on an accelerator that pointer is not host memory,
so the
read segfaults. A program placed on CUDA crashes as soon as tracing is
turned
on, which is exactly when someone is trying to debug it.
Returning an error instead would not help. All four callers wrap the
result in
`ET_CHECK_MSG`, so an error aborts the process rather than skipping the
tensor.
This change brings the data back to host memory first. When the tensor
is not on
CPU, ETDump looks up the allocator registered for that device type,
stages the
bytes into a temporary host buffer with `copy_device_to_host`, writes
that buffer
to the sink and frees it. A tensor on CPU keeps the old path and copies
nothing
extra.
If no allocator is registered for the device, ETDump now reports
`NotFound` and
logs the device type instead of reading the pointer anyway.
### Test plan
Two new test files, each with a CMake target and a Buck target.
`devtools/etdump/tests/etdump_device_test.cpp` registers the existing
`MockCudaAllocator`, which backs its device memory with host memory, and
has two
tests. One logs a tensor tagged as CUDA and checks both that ETDump went
through
the allocator and that the bytes reached the debug buffer. The other
logs a
tensor on CPU and checks that the allocator was not used at all, so the
CPU path
is unchanged.
`devtools/etdump/tests/etdump_device_no_allocator_test.cpp` covers the
case where
nothing is registered for the device. The registry is a process wide
static with
no way to remove an entry, so that case needs a binary that never
registers
anything, which is why it is a second file.
`devtools/etdump/tests/CMakeLists.txt` was not referenced by any parent
`CMakeLists.txt`, so nothing in that directory was built by CMake. This
adds
`add_subdirectory(tests)` to `devtools/etdump/CMakeLists.txt` under
`BUILD_TESTING`, so both new tests are picked up by `ctest`, which is
how the C++
tests run.
The pre-existing `sdk_etdump_tests` target stays out of the CMake build.
It
compiles `etdump_test.cpp`, which includes `etdump_filter.h`, which
needs re2,
and the devtools build does not pull re2 in. It is now guarded on re2
being
available rather than being silently unreachable.
With this change both new binaries pass under `ctest`:
```
1/2 Test #1: etdump_device_test ................ Passed
2/2 Test #2: etdump_device_no_allocator_test ... Passed
```
With `etdump_flatcc.cpp` reverted to the old code and everything
rebuilt, both
fail:
```
Expected equality of these values:
g_mock_cuda.d2h_count_
Which is: 0
1
Death test: etdump_gen.log_evalue(EValue(tensor))
Result: failed to die.
```
Also reproduced the real crash on one NVIDIA H100, with a small program
that
allocates through `cudaMalloc`, tags a tensor as CUDA and logs it:
```
before: Segmentation fault (core dumped)
after: debug buffer holds 1.5 2.5 3.5 4.5
```
Checked that `etdump_flatcc.cpp` still compiles with `-DUSE_ATEN_LIB`.
`clang-format` reports no changes needed on the four touched C++ files.
### Landing order
pytorch#22058 adds device-planned arenas to the Python bindings that ETDump can
be handed
pointers into, and without this change `BufferDataSink::write` would
`memcpy` device
memory from the host. This should land before or together with it.
### Not covered
The ATen mode branch of the device type conversion only compiles. There
is no
ATen mode CMake build to run it in, so nothing here executes it. It is
reachable
in principle: `runtime/executor/tensor_parser_aten.cpp` reads the
serialized
device type and index, including CUDA, and builds the ATen tensor on
that device.
An earlier version of this description said ATen tensors carry no device
metadata,
which is wrong.
`LogTensorOnCpuDoesNotStageThroughTheAllocator` passes with the
production change
reverted, since CPU tensors already went straight to the data sink. It
documents
the CPU path rather than locking the fix; the other two tests are the
ones that
require the new branch.
---------
Co-authored-by: r <r@e>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@winskuo-quic
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[DO NOT MERGE]Qualcomm AI Engine Direct - Mimi Decoder low SQNR reproduce - #1

Draft
winskuo-quic wants to merge 1 commit into
dev1/winskuo/mimi_stage2from
dev1/winskuo/mimi_streaming_low_sqnr_reproduce
Draft

[DO NOT MERGE]Qualcomm AI Engine Direct - Mimi Decoder low SQNR reproduce#1
winskuo-quic wants to merge 1 commit into
dev1/winskuo/mimi_stage2from
dev1/winskuo/mimi_streaming_low_sqnr_reproduce

Conversation

@winskuo-quic

Copy link
Copy Markdown
Collaborator

Summary

Please DO NOT MERGE this PR.
This PR is to reproduce the issue where nn.Module inference is working fine, however, once after torch.export.export, sqnr went from 120 -> 8.

Command:
python examples/qualcomm/oss_scripts/moshi/mimi.py -b build-android -s $device -m SM8650 --chunks_per_batch 125
To run nn.Module where it is working fine, please uncomment line 264-267 and comment out line 269-273

@winskuo-quic
winskuo-quic changed the base branch from dev1/winskuo/mimi_stage2 to mainApril 15, 2025 03:18
@winskuo-quic
winskuo-quic changed the base branch from main to dev1/winskuo/mimi_stage2April 15, 2025 03:19
chenweng-quic pushed a commit that referenced this pull request Jun 9, 2025
Differential Revision: D75911655
Pull Request resolved: pytorch#11344
shewu-quic pushed a commit that referenced this pull request Aug 1, 2025
BNNS copy crashes the process when the dtypes differ
(pytorch#11714).
With the example in this PR
(pytorch#11714), we crash the
process on main. Here is the stack trace from LLDB:
```
Process 19234 stopped
* thread #1, queue = 'com.apple.main-thread', stop reason = signal SIGABRT
frame #0: 0x0000000190ac9388 libsystem_kernel.dylib`__pthread_kill + 8
libsystem_kernel.dylib`__pthread_kill:
-> 0x190ac9388 <+8>: b.lo 0x190ac93a8 ; <+40>
0x190ac938c <+12>: pacibsp 0x190ac9390 <+16>: stp x29, x30, [sp, #-0x10]!
0x190ac9394 <+20>: mov x29, sp
(lldb) bt
* thread #1, queue = 'com.apple.main-thread', stop reason = signal SIGABRT
* frame #0: 0x0000000190ac9388 libsystem_kernel.dylib`__pthread_kill + 8
frame #1: 0x0000000190b0288c libsystem_pthread.dylib`pthread_kill + 296
frame #2: 0x0000000190a0bc60 libsystem_c.dylib`abort + 124
frame pytorch#3: 0x0000000190910174 libsystem_malloc.dylib`malloc_vreport + 892
frame pytorch#4: 0x0000000190913c90 libsystem_malloc.dylib`malloc_report + 64
frame pytorch#5: 0x000000019091821c libsystem_malloc.dylib`___BUG_IN_CLIENT_OF_LIBMALLOC_POINTER_BEING_FREED_WAS_NOT_ALLOCATED + 32
frame pytorch#6: 0x000000019d2f4084 libBNNS.dylib`___lldb_unnamed_symbol1620 + 564
frame pytorch#7: 0x000000019d2f5bac libBNNS.dylib`___lldb_unnamed_symbol1628 + 680
frame pytorch#8: 0x000000019d69ce48 libBNNS.dylib`BNNSCopy + 616
frame pytorch#9: 0x000000030c74d950 _portable_lib.cpython-310-darwin.so`(anonymous namespace)::copy_using_bnns(executorchcoreml::MultiArray const&, executorchcoreml::MultiArray&) + 188
frame pytorch#10: 0x000000030c74cfdc _portable_lib.cpython-310-darwin.so`(anonymous namespace)::copy(executorchcoreml::MultiArray const&, executorchcoreml::MultiArray&, executorchcoreml::MultiArray::CopyOptions) + 72
frame pytorch#11: 0x000000030c74ceec _portable_lib.cpython-310-darwin.so`executorchcoreml::MultiArray::copy(executorchcoreml::MultiArray&, executorchcoreml::MultiArray::CopyOptions) const + 148
frame pytorch#12: 0x000000030c7488d4 _portable_lib.cpython-310-darwin.so`invocation function for block in (anonymous namespace)::copy(MLMultiArray*, executorchcoreml::MultiArray&) + 376
frame pytorch#13: 0x000000030c748ac8 _portable_lib.cpython-310-darwin.so`invocation function for block in (anonymous namespace)::copy(MLMultiArray*, executorchcoreml::MultiArray&) + 52
frame pytorch#14: 0x000000019ad33f4c CoreML`CoreML::MultiArrayBuffer::getBytesWithHandler(void (void const*, unsigned long) block_pointer) const + 340
frame pytorch#15: 0x000000019ad34138 CoreML`-[MLMultiArray(ScopedBufferAccess) getBytesWithHandler:] + 152
frame pytorch#16: 0x000000030c7485ec _portable_lib.cpython-310-darwin.so`(anonymous namespace)::copy(MLMultiArray*, executorchcoreml::MultiArray&) + 296
frame pytorch#17: 0x000000030c744f68 _portable_lib.cpython-310-darwin.so`(anonymous namespace)::set_outputs(std::__1::vector<executorchcoreml::MultiArray, std::__1::allocator<executorchcoreml::MultiArray>>&, NSArray<MLMultiArray*>*) + 180
```
With this PR, the process succeeds.
quic-boyuc added a commit that referenced this pull request May 18, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc pushed a commit that referenced this pull request Aug 25, 2026
…2095)
### Summary
ETDump records an intermediate tensor by handing the tensor's data
pointer to a
data sink, and every sink reads those bytes with a plain host read:
```cpp
memcpy(cur_data_begin, ptr, length);
```
When the tensor lives on an accelerator that pointer is not host memory,
so the
read segfaults. A program placed on CUDA crashes as soon as tracing is
turned
on, which is exactly when someone is trying to debug it.
Returning an error instead would not help. All four callers wrap the
result in
`ET_CHECK_MSG`, so an error aborts the process rather than skipping the
tensor.
This change brings the data back to host memory first. When the tensor
is not on
CPU, ETDump looks up the allocator registered for that device type,
stages the
bytes into a temporary host buffer with `copy_device_to_host`, writes
that buffer
to the sink and frees it. A tensor on CPU keeps the old path and copies
nothing
extra.
If no allocator is registered for the device, ETDump now reports
`NotFound` and
logs the device type instead of reading the pointer anyway.
### Test plan
Two new test files, each with a CMake target and a Buck target.
`devtools/etdump/tests/etdump_device_test.cpp` registers the existing
`MockCudaAllocator`, which backs its device memory with host memory, and
has two
tests. One logs a tensor tagged as CUDA and checks both that ETDump went
through
the allocator and that the bytes reached the debug buffer. The other
logs a
tensor on CPU and checks that the allocator was not used at all, so the
CPU path
is unchanged.
`devtools/etdump/tests/etdump_device_no_allocator_test.cpp` covers the
case where
nothing is registered for the device. The registry is a process wide
static with
no way to remove an entry, so that case needs a binary that never
registers
anything, which is why it is a second file.
`devtools/etdump/tests/CMakeLists.txt` was not referenced by any parent
`CMakeLists.txt`, so nothing in that directory was built by CMake. This
adds
`add_subdirectory(tests)` to `devtools/etdump/CMakeLists.txt` under
`BUILD_TESTING`, so both new tests are picked up by `ctest`, which is
how the C++
tests run.
The pre-existing `sdk_etdump_tests` target stays out of the CMake build.
It
compiles `etdump_test.cpp`, which includes `etdump_filter.h`, which
needs re2,
and the devtools build does not pull re2 in. It is now guarded on re2
being
available rather than being silently unreachable.
With this change both new binaries pass under `ctest`:
```
1/2 Test #1: etdump_device_test ................ Passed
2/2 Test #2: etdump_device_no_allocator_test ... Passed
```
With `etdump_flatcc.cpp` reverted to the old code and everything
rebuilt, both
fail:
```
Expected equality of these values:
g_mock_cuda.d2h_count_
Which is: 0
1
Death test: etdump_gen.log_evalue(EValue(tensor))
Result: failed to die.
```
Also reproduced the real crash on one NVIDIA H100, with a small program
that
allocates through `cudaMalloc`, tags a tensor as CUDA and logs it:
```
before: Segmentation fault (core dumped)
after: debug buffer holds 1.5 2.5 3.5 4.5
```
Checked that `etdump_flatcc.cpp` still compiles with `-DUSE_ATEN_LIB`.
`clang-format` reports no changes needed on the four touched C++ files.
### Landing order
pytorch#22058 adds device-planned arenas to the Python bindings that ETDump can
be handed
pointers into, and without this change `BufferDataSink::write` would
`memcpy` device
memory from the host. This should land before or together with it.
### Not covered
The ATen mode branch of the device type conversion only compiles. There
is no
ATen mode CMake build to run it in, so nothing here executes it. It is
reachable
in principle: `runtime/executor/tensor_parser_aten.cpp` reads the
serialized
device type and index, including CUDA, and builds the ATen tensor on
that device.
An earlier version of this description said ATen tensors carry no device
metadata,
which is wrong.
`LogTensorOnCpuDoesNotStageThroughTheAllocator` passes with the
production change
reverted, since CPU tensors already went straight to the data sink. It
documents
the CPU path rather than locking the fix; the other two tests are the
ones that
require the new branch.
---------
Co-authored-by: r <r@e>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@winskuo-quic
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

[DO NOT MERGE]Qualcomm AI Engine Direct - Mimi Decoder low SQNR reproduce - #1

Draft
winskuo-quic wants to merge 1 commit into
dev1/winskuo/mimi_stage2from
dev1/winskuo/mimi_streaming_low_sqnr_reproduce
Draft

[DO NOT MERGE]Qualcomm AI Engine Direct - Mimi Decoder low SQNR reproduce#1
winskuo-quic wants to merge 1 commit into
dev1/winskuo/mimi_stage2from
dev1/winskuo/mimi_streaming_low_sqnr_reproduce

Conversation

@winskuo-quic

Copy link
Copy Markdown
Collaborator

Summary

Please DO NOT MERGE this PR.
This PR is to reproduce the issue where nn.Module inference is working fine, however, once after torch.export.export, sqnr went from 120 -> 8.

Command:
python examples/qualcomm/oss_scripts/moshi/mimi.py -b build-android -s $device -m SM8650 --chunks_per_batch 125
To run nn.Module where it is working fine, please uncomment line 264-267 and comment out line 269-273

@winskuo-quic
winskuo-quic changed the base branch from dev1/winskuo/mimi_stage2 to mainApril 15, 2025 03:18
@winskuo-quic
winskuo-quic changed the base branch from main to dev1/winskuo/mimi_stage2April 15, 2025 03:19
chenweng-quic pushed a commit that referenced this pull request Jun 9, 2025
Differential Revision: D75911655
Pull Request resolved: pytorch#11344
shewu-quic pushed a commit that referenced this pull request Aug 1, 2025
BNNS copy crashes the process when the dtypes differ
(pytorch#11714).
With the example in this PR
(pytorch#11714), we crash the
process on main. Here is the stack trace from LLDB:
```
Process 19234 stopped
* thread #1, queue = 'com.apple.main-thread', stop reason = signal SIGABRT
frame #0: 0x0000000190ac9388 libsystem_kernel.dylib`__pthread_kill + 8
libsystem_kernel.dylib`__pthread_kill:
-> 0x190ac9388 <+8>: b.lo 0x190ac93a8 ; <+40>
0x190ac938c <+12>: pacibsp 0x190ac9390 <+16>: stp x29, x30, [sp, #-0x10]!
0x190ac9394 <+20>: mov x29, sp
(lldb) bt
* thread #1, queue = 'com.apple.main-thread', stop reason = signal SIGABRT
* frame #0: 0x0000000190ac9388 libsystem_kernel.dylib`__pthread_kill + 8
frame #1: 0x0000000190b0288c libsystem_pthread.dylib`pthread_kill + 296
frame #2: 0x0000000190a0bc60 libsystem_c.dylib`abort + 124
frame pytorch#3: 0x0000000190910174 libsystem_malloc.dylib`malloc_vreport + 892
frame pytorch#4: 0x0000000190913c90 libsystem_malloc.dylib`malloc_report + 64
frame pytorch#5: 0x000000019091821c libsystem_malloc.dylib`___BUG_IN_CLIENT_OF_LIBMALLOC_POINTER_BEING_FREED_WAS_NOT_ALLOCATED + 32
frame pytorch#6: 0x000000019d2f4084 libBNNS.dylib`___lldb_unnamed_symbol1620 + 564
frame pytorch#7: 0x000000019d2f5bac libBNNS.dylib`___lldb_unnamed_symbol1628 + 680
frame pytorch#8: 0x000000019d69ce48 libBNNS.dylib`BNNSCopy + 616
frame pytorch#9: 0x000000030c74d950 _portable_lib.cpython-310-darwin.so`(anonymous namespace)::copy_using_bnns(executorchcoreml::MultiArray const&, executorchcoreml::MultiArray&) + 188
frame pytorch#10: 0x000000030c74cfdc _portable_lib.cpython-310-darwin.so`(anonymous namespace)::copy(executorchcoreml::MultiArray const&, executorchcoreml::MultiArray&, executorchcoreml::MultiArray::CopyOptions) + 72
frame pytorch#11: 0x000000030c74ceec _portable_lib.cpython-310-darwin.so`executorchcoreml::MultiArray::copy(executorchcoreml::MultiArray&, executorchcoreml::MultiArray::CopyOptions) const + 148
frame pytorch#12: 0x000000030c7488d4 _portable_lib.cpython-310-darwin.so`invocation function for block in (anonymous namespace)::copy(MLMultiArray*, executorchcoreml::MultiArray&) + 376
frame pytorch#13: 0x000000030c748ac8 _portable_lib.cpython-310-darwin.so`invocation function for block in (anonymous namespace)::copy(MLMultiArray*, executorchcoreml::MultiArray&) + 52
frame pytorch#14: 0x000000019ad33f4c CoreML`CoreML::MultiArrayBuffer::getBytesWithHandler(void (void const*, unsigned long) block_pointer) const + 340
frame pytorch#15: 0x000000019ad34138 CoreML`-[MLMultiArray(ScopedBufferAccess) getBytesWithHandler:] + 152
frame pytorch#16: 0x000000030c7485ec _portable_lib.cpython-310-darwin.so`(anonymous namespace)::copy(MLMultiArray*, executorchcoreml::MultiArray&) + 296
frame pytorch#17: 0x000000030c744f68 _portable_lib.cpython-310-darwin.so`(anonymous namespace)::set_outputs(std::__1::vector<executorchcoreml::MultiArray, std::__1::allocator<executorchcoreml::MultiArray>>&, NSArray<MLMultiArray*>*) + 180
```
With this PR, the process succeeds.
quic-boyuc added a commit that referenced this pull request May 18, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc pushed a commit that referenced this pull request Aug 25, 2026
…2095)
### Summary
ETDump records an intermediate tensor by handing the tensor's data
pointer to a
data sink, and every sink reads those bytes with a plain host read:
```cpp
memcpy(cur_data_begin, ptr, length);
```
When the tensor lives on an accelerator that pointer is not host memory,
so the
read segfaults. A program placed on CUDA crashes as soon as tracing is
turned
on, which is exactly when someone is trying to debug it.
Returning an error instead would not help. All four callers wrap the
result in
`ET_CHECK_MSG`, so an error aborts the process rather than skipping the
tensor.
This change brings the data back to host memory first. When the tensor
is not on
CPU, ETDump looks up the allocator registered for that device type,
stages the
bytes into a temporary host buffer with `copy_device_to_host`, writes
that buffer
to the sink and frees it. A tensor on CPU keeps the old path and copies
nothing
extra.
If no allocator is registered for the device, ETDump now reports
`NotFound` and
logs the device type instead of reading the pointer anyway.
### Test plan
Two new test files, each with a CMake target and a Buck target.
`devtools/etdump/tests/etdump_device_test.cpp` registers the existing
`MockCudaAllocator`, which backs its device memory with host memory, and
has two
tests. One logs a tensor tagged as CUDA and checks both that ETDump went
through
the allocator and that the bytes reached the debug buffer. The other
logs a
tensor on CPU and checks that the allocator was not used at all, so the
CPU path
is unchanged.
`devtools/etdump/tests/etdump_device_no_allocator_test.cpp` covers the
case where
nothing is registered for the device. The registry is a process wide
static with
no way to remove an entry, so that case needs a binary that never
registers
anything, which is why it is a second file.
`devtools/etdump/tests/CMakeLists.txt` was not referenced by any parent
`CMakeLists.txt`, so nothing in that directory was built by CMake. This
adds
`add_subdirectory(tests)` to `devtools/etdump/CMakeLists.txt` under
`BUILD_TESTING`, so both new tests are picked up by `ctest`, which is
how the C++
tests run.
The pre-existing `sdk_etdump_tests` target stays out of the CMake build.
It
compiles `etdump_test.cpp`, which includes `etdump_filter.h`, which
needs re2,
and the devtools build does not pull re2 in. It is now guarded on re2
being
available rather than being silently unreachable.
With this change both new binaries pass under `ctest`:
```
1/2 Test #1: etdump_device_test ................ Passed
2/2 Test #2: etdump_device_no_allocator_test ... Passed
```
With `etdump_flatcc.cpp` reverted to the old code and everything
rebuilt, both
fail:
```
Expected equality of these values:
g_mock_cuda.d2h_count_
Which is: 0
1
Death test: etdump_gen.log_evalue(EValue(tensor))
Result: failed to die.
```
Also reproduced the real crash on one NVIDIA H100, with a small program
that
allocates through `cudaMalloc`, tags a tensor as CUDA and logs it:
```
before: Segmentation fault (core dumped)
after: debug buffer holds 1.5 2.5 3.5 4.5
```
Checked that `etdump_flatcc.cpp` still compiles with `-DUSE_ATEN_LIB`.
`clang-format` reports no changes needed on the four touched C++ files.
### Landing order
pytorch#22058 adds device-planned arenas to the Python bindings that ETDump can
be handed
pointers into, and without this change `BufferDataSink::write` would
`memcpy` device
memory from the host. This should land before or together with it.
### Not covered
The ATen mode branch of the device type conversion only compiles. There
is no
ATen mode CMake build to run it in, so nothing here executes it. It is
reachable
in principle: `runtime/executor/tensor_parser_aten.cpp` reads the
serialized
device type and index, including CUDA, and builds the ATen tensor on
that device.
An earlier version of this description said ATen tensors carry no device
metadata,
which is wrong.
`LogTensorOnCpuDoesNotStageThroughTheAllocator` passes with the
production change
reverted, since CPU tensors already went straight to the data sink. It
documents
the CPU path rather than locking the fix; the other two tests are the
ones that
require the new branch.
---------
Co-authored-by: r <r@e>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@winskuo-quic