Qualcomm AI Engine Direct - oss model enablement (EfficientSAM) - #9266

Merged
cccclai merged 1 commit into
pytorch:mainfrom
CodeLinaro:dev1/danny/EfficientSAM_enablement
Apr 10, 2025
Merged

Qualcomm AI Engine Direct - oss model enablement (EfficientSAM)#9266
cccclai merged 1 commit into
pytorch:mainfrom
CodeLinaro:dev1/danny/EfficientSAM_enablement

Conversation

@DannyYuyang-quic

@DannyYuyang-quicDannyYuyang-quic commented Mar 14, 2025

Copy link
Copy Markdown
Contributor

Summary

Test plan

python ./examples/qualcomm/oss_scripts/efficientSAM/efficientSAM.py -m ${soc} -b build-android -H ${host_id} -s ${device_id} --oss_repo ${Path_to_oss_repo} --pretrained_weight ${Path_to_pretrained_weight} -d ${Path_to_dataset_dir}

@pytorch-bot

pytorch-botBot commented Mar 14, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/9266

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit 42412d7 with merge base c9c5481 (image):

NEW FAILURE - The following job has failed:

  • pull / unittest-arm / linux-job (gh)
    RuntimeError: Command docker exec -t c50233835032a0388f502a805413875acb5526b953584d9a5d02ffbe166713f5 /exec failed with exit code 1

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 14, 2025
@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

@pytorchbot label "release notes: qualcomm"

@pytorch-botpytorch-botBot added the release notes: qualcomm Changes to the Qualcomm backend delegate label Mar 14, 2025
@DannyYuyang-quic
DannyYuyang-quicforce-pushed the dev1/danny/EfficientSAM_enablement branch from f36af71 to 1f614eaCompareMarch 14, 2025 08:56
@DannyYuyang-quic

DannyYuyang-quic commented Mar 14, 2025

Copy link
Copy Markdown
ContributorAuthor

@cccclai
This PR is one of the model enablement requests. Please have a look.
BTW, I can't trigger CI. Could you please help with this?

Thanks!

@cccclaicccclai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for enabling EfficientSAM! As we enable more models, let's add it to part of the CI, similar to #8616

@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

Thank you for enabling EfficientSAM! As we enable more models, let's add it to part of the CI, similar to #8616

Regarding the CI, would you prefer I add it in this PR or add it in the next one?

@cccclai

Copy link
Copy Markdown
Contributor

Thank you for enabling EfficientSAM! As we enable more models, let's add it to part of the CI, similar to #8616

Next PR is fine

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

def test_qnn_backend_cumsum(self):
module = CumSum() # noqa: F405
sample_input = (torch.randn(4),)
self.lower_module_and_test_output(module, sample_input)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems like the fp test is failing, can you double check? The quantized one is passing.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could you please share the log~?
Both the fp and quantized tests are passing from my side.

@cccclai

Copy link
Copy Markdown
Contributor

The error message is

[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Initializing HtpProvider
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend parameters
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend
[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
/data/sandcastle/boxes/eden-trunk-hg-full-fbsource/buck-out/v2/gen/fbcode/ec7059d5161b31ff/executorch/backends/qualcomm/tests/fb/__test_qnn_delegate_simulator__/test_qnn_delegate_simulator#link-tree/executorch/backends/qualcomm/qnn_preprocess.py:69: Visiting: aten_cumsum_default, aten.cumsum.default
[ERROR] [Qnn ExecuTorch]: tcm_migration.cc:1863:ERROR:no properties registered for q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:210:ERROR:could not create op: q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:1403:ERROR:Op 0x10 preparation failed with err:-1
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> "aten_cumsum_default" generated: could not create op
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> RouterX86 graph prepare failed 12
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> Failed to finalize graph (id: 1) with err 1002
[ERROR] [Qnn ExecuTorch]: Failed to finalize Qnn Graph with error: 1002
[ERROR] [Qnn ExecuTorch]: Fail to compile QNN graph
FAIL
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend

@DannyYuyang-quicDannyYuyang-quic left a comment

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The error message is

[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Initializing HtpProvider
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend parameters
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend
[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
/data/sandcastle/boxes/eden-trunk-hg-full-fbsource/buck-out/v2/gen/fbcode/ec7059d5161b31ff/executorch/backends/qualcomm/tests/fb/__test_qnn_delegate_simulator__/test_qnn_delegate_simulator#link-tree/executorch/backends/qualcomm/qnn_preprocess.py:69: Visiting: aten_cumsum_default, aten.cumsum.default
[ERROR] [Qnn ExecuTorch]: tcm_migration.cc:1863:ERROR:no properties registered for q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:210:ERROR:could not create op: q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:1403:ERROR:Op 0x10 preparation failed with err:-1
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> "aten_cumsum_default" generated: could not create op
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> RouterX86 graph prepare failed 12
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> Failed to finalize graph (id: 1) with err 1002
[ERROR] [Qnn ExecuTorch]: Failed to finalize Qnn Graph with error: 1002
[ERROR] [Qnn ExecuTorch]: Fail to compile QNN graph
FAIL
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend

I tried QNN versions 2.27, 2.28, 2.31, and 2.32 with android-ndk-r26c in this PR but still can't reproduce this error. I think I need more details to tackle this issue. Could you please check which QNN version you used? Also, check the QNN version in $LD_LIBRARY_PATH, thanks!

@cccclai

cccclai commented Mar 25, 2025

Copy link
Copy Markdown
Contributor

I'm currently using qnn 2.28, but not sure the android ndk version. It seems like there are quite a few fp tests continuously failing while most are passing. We can merge this PR for now, but will need some help to identify the root cause. See following for the failing fp tests, the one under fp but not marked are the passing tests.

 # Overwrite because unit test is failing, once passing, remove this function
def test_qnn_backend_interpolate_nearest_2d(self):
pass
# Overwrite because unit test is failing, once passing, remove this function
def test_qnn_backend_interpolate_bilinear_2d(self):
pass
def test_qnn_backend_element_wise_ceil(self):
pass
def test_qnn_backend_embedding(self):
pass
def test_qnn_backend_gelu(self):
pass
def test_qnn_backend_group_norm(self):
pass
def test_qnn_backend_hardsigmoid(self):
pass
def test_qnn_backend_index(self):
pass
def test_qnn_backend_index_put(self):
pass
def test_qnn_backend_instance_norm_2d(self):
pass
def test_qnn_backend_rms_norm(self):
pass
def test_qnn_backend_stack(self):
pass
def test_qnn_backend_where(self):
pass
def test_qnn_backend_conv_transpose2d(self):
pass

We target SM8650 for the unit test btw.

@DannyYuyang-quic

DannyYuyang-quic commented Mar 25, 2025

Copy link
Copy Markdown
ContributorAuthor

I think you can check backend_options = generate_htp_compiler_spec(use_fp16=True) setting.
The use_fp16 option should be set to True in all fp tests, as shown in the figure below.
I'm not sure if you have modified this setting or not~

image

@cccclai

Copy link
Copy Markdown
Contributor

I think you can check backend_options = generate_htp_compiler_spec(use_fp16=True) setting. The use_fp16 option should be set to True in all fp tests, as shown in the figure below. I'm not sure if you have modified this setting or not~

image

Ah good catch. I inherit the base class TestQNNFloatingPointOperator and accidently use TestQNNQuantizedOperator to setup instead of TestQNNFloatingPointOperator. Let me update.

@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

I think you can check backend_options = generate_htp_compiler_spec(use_fp16=True) setting. The use_fp16 option should be set to True in all fp tests, as shown in the figure below. I'm not sure if you have modified this setting or not~
image

Ah good catch. I inherit the base class TestQNNFloatingPointOperator and accidently use TestQNNQuantizedOperator to setup instead of TestQNNFloatingPointOperator. Let me update.

Cool! Let me know if there is any issue.

@cccclai

Copy link
Copy Markdown
Contributor

It should be good now. Mind rebasing?

@DannyYuyang-quic

DannyYuyang-quic commented Apr 2, 2025

Copy link
Copy Markdown
ContributorAuthor

I've rebased the branch. Thanks!

I was on the older base in the last commit, but now it's on the newest one.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclaicccclai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you!

@cccclai

Copy link
Copy Markdown
Contributor

Everything is green, looks like it's still need rebase...mind rebasing again?

 - e2e script for https://github.com/yformer/EfficientSAM
- Fastvit breakage fix
- Add support for cum_sum
- Add bicubic interpolate transform pass
- Fix stack op
@DannyYuyang-quic
DannyYuyang-quicforce-pushed the dev1/danny/EfficientSAM_enablement branch from 4bb1800 to 42412d7CompareApril 10, 2025 06:30
@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

Everything is green, looks like it's still need rebase...mind rebasing again?

Done!
Please have a look.
Thanks!

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai
cccclai merged commit 5e4f045 into pytorch:mainApr 10, 2025
kirklandsign pushed a commit that referenced this pull request Apr 11, 2025
### Summary
- e2e script for https://github.com/yformer/EfficientSAM
- Fastvit breakage fix
- Add support for cum_sum
- Add bicubic interpolate transform pass
- Fix stack op
### Test plan
``` bash
python ./examples/qualcomm/oss_scripts/efficientSAM/efficientSAM.py -m ${soc} -b build-android -H ${host_id} -s ${device_id} --oss_repo ${Path_to_oss_repo} --pretrained_weight ${Path_to_pretrained_weight} -d ${Path_to_dataset_dir}
```
keyprocedure pushed a commit to keyprocedure/executorch that referenced this pull request Apr 21, 2025
…rch#9266)
### Summary
- e2e script for https://github.com/yformer/EfficientSAM
- Fastvit breakage fix
- Add support for cum_sum
- Add bicubic interpolate transform pass
- Fix stack op
### Test plan
``` bash
python ./examples/qualcomm/oss_scripts/efficientSAM/efficientSAM.py -m ${soc} -b build-android -H ${host_id} -s ${device_id} --oss_repo ${Path_to_oss_repo} --pretrained_weight ${Path_to_pretrained_weight} -d ${Path_to_dataset_dir}
```
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.release notes: qualcommChanges to the Qualcomm backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@DannyYuyang-quic@cccclai@facebook-github-bot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Qualcomm AI Engine Direct - oss model enablement (EfficientSAM) - #9266

Merged
cccclai merged 1 commit into
pytorch:mainfrom
CodeLinaro:dev1/danny/EfficientSAM_enablement
Apr 10, 2025
Merged

Qualcomm AI Engine Direct - oss model enablement (EfficientSAM)#9266
cccclai merged 1 commit into
pytorch:mainfrom
CodeLinaro:dev1/danny/EfficientSAM_enablement

Conversation

@DannyYuyang-quic

@DannyYuyang-quicDannyYuyang-quic commented Mar 14, 2025

Copy link
Copy Markdown
Contributor

Summary

Test plan

python ./examples/qualcomm/oss_scripts/efficientSAM/efficientSAM.py -m ${soc} -b build-android -H ${host_id} -s ${device_id} --oss_repo ${Path_to_oss_repo} --pretrained_weight ${Path_to_pretrained_weight} -d ${Path_to_dataset_dir}

@pytorch-bot

pytorch-botBot commented Mar 14, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/9266

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit 42412d7 with merge base c9c5481 (image):

NEW FAILURE - The following job has failed:

  • pull / unittest-arm / linux-job (gh)
    RuntimeError: Command docker exec -t c50233835032a0388f502a805413875acb5526b953584d9a5d02ffbe166713f5 /exec failed with exit code 1

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 14, 2025
@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

@pytorchbot label "release notes: qualcomm"

@pytorch-botpytorch-botBot added the release notes: qualcomm Changes to the Qualcomm backend delegate label Mar 14, 2025
@DannyYuyang-quic
DannyYuyang-quicforce-pushed the dev1/danny/EfficientSAM_enablement branch from f36af71 to 1f614eaCompareMarch 14, 2025 08:56
@DannyYuyang-quic

DannyYuyang-quic commented Mar 14, 2025

Copy link
Copy Markdown
ContributorAuthor

@cccclai
This PR is one of the model enablement requests. Please have a look.
BTW, I can't trigger CI. Could you please help with this?

Thanks!

@cccclaicccclai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for enabling EfficientSAM! As we enable more models, let's add it to part of the CI, similar to #8616

@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

Thank you for enabling EfficientSAM! As we enable more models, let's add it to part of the CI, similar to #8616

Regarding the CI, would you prefer I add it in this PR or add it in the next one?

@cccclai

Copy link
Copy Markdown
Contributor

Thank you for enabling EfficientSAM! As we enable more models, let's add it to part of the CI, similar to #8616

Next PR is fine

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

def test_qnn_backend_cumsum(self):
module = CumSum() # noqa: F405
sample_input = (torch.randn(4),)
self.lower_module_and_test_output(module, sample_input)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems like the fp test is failing, can you double check? The quantized one is passing.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could you please share the log~?
Both the fp and quantized tests are passing from my side.

@cccclai

Copy link
Copy Markdown
Contributor

The error message is

[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Initializing HtpProvider
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend parameters
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend
[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
/data/sandcastle/boxes/eden-trunk-hg-full-fbsource/buck-out/v2/gen/fbcode/ec7059d5161b31ff/executorch/backends/qualcomm/tests/fb/__test_qnn_delegate_simulator__/test_qnn_delegate_simulator#link-tree/executorch/backends/qualcomm/qnn_preprocess.py:69: Visiting: aten_cumsum_default, aten.cumsum.default
[ERROR] [Qnn ExecuTorch]: tcm_migration.cc:1863:ERROR:no properties registered for q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:210:ERROR:could not create op: q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:1403:ERROR:Op 0x10 preparation failed with err:-1
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> "aten_cumsum_default" generated: could not create op
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> RouterX86 graph prepare failed 12
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> Failed to finalize graph (id: 1) with err 1002
[ERROR] [Qnn ExecuTorch]: Failed to finalize Qnn Graph with error: 1002
[ERROR] [Qnn ExecuTorch]: Fail to compile QNN graph
FAIL
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend

@DannyYuyang-quicDannyYuyang-quic left a comment

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The error message is

[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Initializing HtpProvider
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend parameters
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend
[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
/data/sandcastle/boxes/eden-trunk-hg-full-fbsource/buck-out/v2/gen/fbcode/ec7059d5161b31ff/executorch/backends/qualcomm/tests/fb/__test_qnn_delegate_simulator__/test_qnn_delegate_simulator#link-tree/executorch/backends/qualcomm/qnn_preprocess.py:69: Visiting: aten_cumsum_default, aten.cumsum.default
[ERROR] [Qnn ExecuTorch]: tcm_migration.cc:1863:ERROR:no properties registered for q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:210:ERROR:could not create op: q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:1403:ERROR:Op 0x10 preparation failed with err:-1
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> "aten_cumsum_default" generated: could not create op
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> RouterX86 graph prepare failed 12
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> Failed to finalize graph (id: 1) with err 1002
[ERROR] [Qnn ExecuTorch]: Failed to finalize Qnn Graph with error: 1002
[ERROR] [Qnn ExecuTorch]: Fail to compile QNN graph
FAIL
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend

I tried QNN versions 2.27, 2.28, 2.31, and 2.32 with android-ndk-r26c in this PR but still can't reproduce this error. I think I need more details to tackle this issue. Could you please check which QNN version you used? Also, check the QNN version in $LD_LIBRARY_PATH, thanks!

@cccclai

cccclai commented Mar 25, 2025

Copy link
Copy Markdown
Contributor

I'm currently using qnn 2.28, but not sure the android ndk version. It seems like there are quite a few fp tests continuously failing while most are passing. We can merge this PR for now, but will need some help to identify the root cause. See following for the failing fp tests, the one under fp but not marked are the passing tests.

 # Overwrite because unit test is failing, once passing, remove this function
def test_qnn_backend_interpolate_nearest_2d(self):
pass
# Overwrite because unit test is failing, once passing, remove this function
def test_qnn_backend_interpolate_bilinear_2d(self):
pass
def test_qnn_backend_element_wise_ceil(self):
pass
def test_qnn_backend_embedding(self):
pass
def test_qnn_backend_gelu(self):
pass
def test_qnn_backend_group_norm(self):
pass
def test_qnn_backend_hardsigmoid(self):
pass
def test_qnn_backend_index(self):
pass
def test_qnn_backend_index_put(self):
pass
def test_qnn_backend_instance_norm_2d(self):
pass
def test_qnn_backend_rms_norm(self):
pass
def test_qnn_backend_stack(self):
pass
def test_qnn_backend_where(self):
pass
def test_qnn_backend_conv_transpose2d(self):
pass

We target SM8650 for the unit test btw.

@DannyYuyang-quic

DannyYuyang-quic commented Mar 25, 2025

Copy link
Copy Markdown
ContributorAuthor

I think you can check backend_options = generate_htp_compiler_spec(use_fp16=True) setting.
The use_fp16 option should be set to True in all fp tests, as shown in the figure below.
I'm not sure if you have modified this setting or not~

image

@cccclai

Copy link
Copy Markdown
Contributor

I think you can check backend_options = generate_htp_compiler_spec(use_fp16=True) setting. The use_fp16 option should be set to True in all fp tests, as shown in the figure below. I'm not sure if you have modified this setting or not~

image

Ah good catch. I inherit the base class TestQNNFloatingPointOperator and accidently use TestQNNQuantizedOperator to setup instead of TestQNNFloatingPointOperator. Let me update.

@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

I think you can check backend_options = generate_htp_compiler_spec(use_fp16=True) setting. The use_fp16 option should be set to True in all fp tests, as shown in the figure below. I'm not sure if you have modified this setting or not~
image

Ah good catch. I inherit the base class TestQNNFloatingPointOperator and accidently use TestQNNQuantizedOperator to setup instead of TestQNNFloatingPointOperator. Let me update.

Cool! Let me know if there is any issue.

@cccclai

Copy link
Copy Markdown
Contributor

It should be good now. Mind rebasing?

@DannyYuyang-quic

DannyYuyang-quic commented Apr 2, 2025

Copy link
Copy Markdown
ContributorAuthor

I've rebased the branch. Thanks!

I was on the older base in the last commit, but now it's on the newest one.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclaicccclai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you!

@cccclai

Copy link
Copy Markdown
Contributor

Everything is green, looks like it's still need rebase...mind rebasing again?

 - e2e script for https://github.com/yformer/EfficientSAM
- Fastvit breakage fix
- Add support for cum_sum
- Add bicubic interpolate transform pass
- Fix stack op
@DannyYuyang-quic
DannyYuyang-quicforce-pushed the dev1/danny/EfficientSAM_enablement branch from 4bb1800 to 42412d7CompareApril 10, 2025 06:30
@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

Everything is green, looks like it's still need rebase...mind rebasing again?

Done!
Please have a look.
Thanks!

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai
cccclai merged commit 5e4f045 into pytorch:mainApr 10, 2025
kirklandsign pushed a commit that referenced this pull request Apr 11, 2025
### Summary
- e2e script for https://github.com/yformer/EfficientSAM
- Fastvit breakage fix
- Add support for cum_sum
- Add bicubic interpolate transform pass
- Fix stack op
### Test plan
``` bash
python ./examples/qualcomm/oss_scripts/efficientSAM/efficientSAM.py -m ${soc} -b build-android -H ${host_id} -s ${device_id} --oss_repo ${Path_to_oss_repo} --pretrained_weight ${Path_to_pretrained_weight} -d ${Path_to_dataset_dir}
```
keyprocedure pushed a commit to keyprocedure/executorch that referenced this pull request Apr 21, 2025
…rch#9266)
### Summary
- e2e script for https://github.com/yformer/EfficientSAM
- Fastvit breakage fix
- Add support for cum_sum
- Add bicubic interpolate transform pass
- Fix stack op
### Test plan
``` bash
python ./examples/qualcomm/oss_scripts/efficientSAM/efficientSAM.py -m ${soc} -b build-android -H ${host_id} -s ${device_id} --oss_repo ${Path_to_oss_repo} --pretrained_weight ${Path_to_pretrained_weight} -d ${Path_to_dataset_dir}
```
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.release notes: qualcommChanges to the Qualcomm backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@DannyYuyang-quic@cccclai@facebook-github-bot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Qualcomm AI Engine Direct - oss model enablement (EfficientSAM) - #9266

Merged
cccclai merged 1 commit into
pytorch:mainfrom
CodeLinaro:dev1/danny/EfficientSAM_enablement
Apr 10, 2025
Merged

Qualcomm AI Engine Direct - oss model enablement (EfficientSAM)#9266
cccclai merged 1 commit into
pytorch:mainfrom
CodeLinaro:dev1/danny/EfficientSAM_enablement

Conversation

@DannyYuyang-quic

@DannyYuyang-quicDannyYuyang-quic commented Mar 14, 2025

Copy link
Copy Markdown
Contributor

Summary

Test plan

python ./examples/qualcomm/oss_scripts/efficientSAM/efficientSAM.py -m ${soc} -b build-android -H ${host_id} -s ${device_id} --oss_repo ${Path_to_oss_repo} --pretrained_weight ${Path_to_pretrained_weight} -d ${Path_to_dataset_dir}

@pytorch-bot

pytorch-botBot commented Mar 14, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/9266

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit 42412d7 with merge base c9c5481 (image):

NEW FAILURE - The following job has failed:

  • pull / unittest-arm / linux-job (gh)
    RuntimeError: Command docker exec -t c50233835032a0388f502a805413875acb5526b953584d9a5d02ffbe166713f5 /exec failed with exit code 1

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 14, 2025
@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

@pytorchbot label "release notes: qualcomm"

@pytorch-botpytorch-botBot added the release notes: qualcomm Changes to the Qualcomm backend delegate label Mar 14, 2025
@DannyYuyang-quic
DannyYuyang-quicforce-pushed the dev1/danny/EfficientSAM_enablement branch from f36af71 to 1f614eaCompareMarch 14, 2025 08:56
@DannyYuyang-quic

DannyYuyang-quic commented Mar 14, 2025

Copy link
Copy Markdown
ContributorAuthor

@cccclai
This PR is one of the model enablement requests. Please have a look.
BTW, I can't trigger CI. Could you please help with this?

Thanks!

@cccclaicccclai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for enabling EfficientSAM! As we enable more models, let's add it to part of the CI, similar to #8616

@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

Thank you for enabling EfficientSAM! As we enable more models, let's add it to part of the CI, similar to #8616

Regarding the CI, would you prefer I add it in this PR or add it in the next one?

@cccclai

Copy link
Copy Markdown
Contributor

Thank you for enabling EfficientSAM! As we enable more models, let's add it to part of the CI, similar to #8616

Next PR is fine

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

def test_qnn_backend_cumsum(self):
module = CumSum() # noqa: F405
sample_input = (torch.randn(4),)
self.lower_module_and_test_output(module, sample_input)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems like the fp test is failing, can you double check? The quantized one is passing.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could you please share the log~?
Both the fp and quantized tests are passing from my side.

@cccclai

Copy link
Copy Markdown
Contributor

The error message is

[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Initializing HtpProvider
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend parameters
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend
[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
/data/sandcastle/boxes/eden-trunk-hg-full-fbsource/buck-out/v2/gen/fbcode/ec7059d5161b31ff/executorch/backends/qualcomm/tests/fb/__test_qnn_delegate_simulator__/test_qnn_delegate_simulator#link-tree/executorch/backends/qualcomm/qnn_preprocess.py:69: Visiting: aten_cumsum_default, aten.cumsum.default
[ERROR] [Qnn ExecuTorch]: tcm_migration.cc:1863:ERROR:no properties registered for q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:210:ERROR:could not create op: q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:1403:ERROR:Op 0x10 preparation failed with err:-1
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> "aten_cumsum_default" generated: could not create op
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> RouterX86 graph prepare failed 12
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> Failed to finalize graph (id: 1) with err 1002
[ERROR] [Qnn ExecuTorch]: Failed to finalize Qnn Graph with error: 1002
[ERROR] [Qnn ExecuTorch]: Fail to compile QNN graph
FAIL
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend

@DannyYuyang-quicDannyYuyang-quic left a comment

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The error message is

[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Initializing HtpProvider
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend parameters
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend
[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
/data/sandcastle/boxes/eden-trunk-hg-full-fbsource/buck-out/v2/gen/fbcode/ec7059d5161b31ff/executorch/backends/qualcomm/tests/fb/__test_qnn_delegate_simulator__/test_qnn_delegate_simulator#link-tree/executorch/backends/qualcomm/qnn_preprocess.py:69: Visiting: aten_cumsum_default, aten.cumsum.default
[ERROR] [Qnn ExecuTorch]: tcm_migration.cc:1863:ERROR:no properties registered for q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:210:ERROR:could not create op: q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:1403:ERROR:Op 0x10 preparation failed with err:-1
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> "aten_cumsum_default" generated: could not create op
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> RouterX86 graph prepare failed 12
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> Failed to finalize graph (id: 1) with err 1002
[ERROR] [Qnn ExecuTorch]: Failed to finalize Qnn Graph with error: 1002
[ERROR] [Qnn ExecuTorch]: Fail to compile QNN graph
FAIL
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend

I tried QNN versions 2.27, 2.28, 2.31, and 2.32 with android-ndk-r26c in this PR but still can't reproduce this error. I think I need more details to tackle this issue. Could you please check which QNN version you used? Also, check the QNN version in $LD_LIBRARY_PATH, thanks!

@cccclai

cccclai commented Mar 25, 2025

Copy link
Copy Markdown
Contributor

I'm currently using qnn 2.28, but not sure the android ndk version. It seems like there are quite a few fp tests continuously failing while most are passing. We can merge this PR for now, but will need some help to identify the root cause. See following for the failing fp tests, the one under fp but not marked are the passing tests.

 # Overwrite because unit test is failing, once passing, remove this function
def test_qnn_backend_interpolate_nearest_2d(self):
pass
# Overwrite because unit test is failing, once passing, remove this function
def test_qnn_backend_interpolate_bilinear_2d(self):
pass
def test_qnn_backend_element_wise_ceil(self):
pass
def test_qnn_backend_embedding(self):
pass
def test_qnn_backend_gelu(self):
pass
def test_qnn_backend_group_norm(self):
pass
def test_qnn_backend_hardsigmoid(self):
pass
def test_qnn_backend_index(self):
pass
def test_qnn_backend_index_put(self):
pass
def test_qnn_backend_instance_norm_2d(self):
pass
def test_qnn_backend_rms_norm(self):
pass
def test_qnn_backend_stack(self):
pass
def test_qnn_backend_where(self):
pass
def test_qnn_backend_conv_transpose2d(self):
pass

We target SM8650 for the unit test btw.

@DannyYuyang-quic

DannyYuyang-quic commented Mar 25, 2025

Copy link
Copy Markdown
ContributorAuthor

I think you can check backend_options = generate_htp_compiler_spec(use_fp16=True) setting.
The use_fp16 option should be set to True in all fp tests, as shown in the figure below.
I'm not sure if you have modified this setting or not~

image

@cccclai

Copy link
Copy Markdown
Contributor

I think you can check backend_options = generate_htp_compiler_spec(use_fp16=True) setting. The use_fp16 option should be set to True in all fp tests, as shown in the figure below. I'm not sure if you have modified this setting or not~

image

Ah good catch. I inherit the base class TestQNNFloatingPointOperator and accidently use TestQNNQuantizedOperator to setup instead of TestQNNFloatingPointOperator. Let me update.

@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

I think you can check backend_options = generate_htp_compiler_spec(use_fp16=True) setting. The use_fp16 option should be set to True in all fp tests, as shown in the figure below. I'm not sure if you have modified this setting or not~
image

Ah good catch. I inherit the base class TestQNNFloatingPointOperator and accidently use TestQNNQuantizedOperator to setup instead of TestQNNFloatingPointOperator. Let me update.

Cool! Let me know if there is any issue.

@cccclai

Copy link
Copy Markdown
Contributor

It should be good now. Mind rebasing?

@DannyYuyang-quic

DannyYuyang-quic commented Apr 2, 2025

Copy link
Copy Markdown
ContributorAuthor

I've rebased the branch. Thanks!

I was on the older base in the last commit, but now it's on the newest one.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclaicccclai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you!

@cccclai

Copy link
Copy Markdown
Contributor

Everything is green, looks like it's still need rebase...mind rebasing again?

 - e2e script for https://github.com/yformer/EfficientSAM
- Fastvit breakage fix
- Add support for cum_sum
- Add bicubic interpolate transform pass
- Fix stack op
@DannyYuyang-quic
DannyYuyang-quicforce-pushed the dev1/danny/EfficientSAM_enablement branch from 4bb1800 to 42412d7CompareApril 10, 2025 06:30
@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

Everything is green, looks like it's still need rebase...mind rebasing again?

Done!
Please have a look.
Thanks!

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai
cccclai merged commit 5e4f045 into pytorch:mainApr 10, 2025
kirklandsign pushed a commit that referenced this pull request Apr 11, 2025
### Summary
- e2e script for https://github.com/yformer/EfficientSAM
- Fastvit breakage fix
- Add support for cum_sum
- Add bicubic interpolate transform pass
- Fix stack op
### Test plan
``` bash
python ./examples/qualcomm/oss_scripts/efficientSAM/efficientSAM.py -m ${soc} -b build-android -H ${host_id} -s ${device_id} --oss_repo ${Path_to_oss_repo} --pretrained_weight ${Path_to_pretrained_weight} -d ${Path_to_dataset_dir}
```
keyprocedure pushed a commit to keyprocedure/executorch that referenced this pull request Apr 21, 2025
…rch#9266)
### Summary
- e2e script for https://github.com/yformer/EfficientSAM
- Fastvit breakage fix
- Add support for cum_sum
- Add bicubic interpolate transform pass
- Fix stack op
### Test plan
``` bash
python ./examples/qualcomm/oss_scripts/efficientSAM/efficientSAM.py -m ${soc} -b build-android -H ${host_id} -s ${device_id} --oss_repo ${Path_to_oss_repo} --pretrained_weight ${Path_to_pretrained_weight} -d ${Path_to_dataset_dir}
```
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.release notes: qualcommChanges to the Qualcomm backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@DannyYuyang-quic@cccclai@facebook-github-bot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Qualcomm AI Engine Direct - oss model enablement (EfficientSAM) - #9266

Merged
cccclai merged 1 commit into
pytorch:mainfrom
CodeLinaro:dev1/danny/EfficientSAM_enablement
Apr 10, 2025
Merged

Qualcomm AI Engine Direct - oss model enablement (EfficientSAM)#9266
cccclai merged 1 commit into
pytorch:mainfrom
CodeLinaro:dev1/danny/EfficientSAM_enablement

Conversation

@DannyYuyang-quic

@DannyYuyang-quicDannyYuyang-quic commented Mar 14, 2025

Copy link
Copy Markdown
Contributor

Summary

Test plan

python ./examples/qualcomm/oss_scripts/efficientSAM/efficientSAM.py -m ${soc} -b build-android -H ${host_id} -s ${device_id} --oss_repo ${Path_to_oss_repo} --pretrained_weight ${Path_to_pretrained_weight} -d ${Path_to_dataset_dir}

@pytorch-bot

pytorch-botBot commented Mar 14, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/9266

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit 42412d7 with merge base c9c5481 (image):

NEW FAILURE - The following job has failed:

  • pull / unittest-arm / linux-job (gh)
    RuntimeError: Command docker exec -t c50233835032a0388f502a805413875acb5526b953584d9a5d02ffbe166713f5 /exec failed with exit code 1

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 14, 2025
@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

@pytorchbot label "release notes: qualcomm"

@pytorch-botpytorch-botBot added the release notes: qualcomm Changes to the Qualcomm backend delegate label Mar 14, 2025
@DannyYuyang-quic
DannyYuyang-quicforce-pushed the dev1/danny/EfficientSAM_enablement branch from f36af71 to 1f614eaCompareMarch 14, 2025 08:56
@DannyYuyang-quic

DannyYuyang-quic commented Mar 14, 2025

Copy link
Copy Markdown
ContributorAuthor

@cccclai
This PR is one of the model enablement requests. Please have a look.
BTW, I can't trigger CI. Could you please help with this?

Thanks!

@cccclaicccclai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for enabling EfficientSAM! As we enable more models, let's add it to part of the CI, similar to #8616

@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

Thank you for enabling EfficientSAM! As we enable more models, let's add it to part of the CI, similar to #8616

Regarding the CI, would you prefer I add it in this PR or add it in the next one?

@cccclai

Copy link
Copy Markdown
Contributor

Thank you for enabling EfficientSAM! As we enable more models, let's add it to part of the CI, similar to #8616

Next PR is fine

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

def test_qnn_backend_cumsum(self):
module = CumSum() # noqa: F405
sample_input = (torch.randn(4),)
self.lower_module_and_test_output(module, sample_input)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems like the fp test is failing, can you double check? The quantized one is passing.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could you please share the log~?
Both the fp and quantized tests are passing from my side.

@cccclai

Copy link
Copy Markdown
Contributor

The error message is

[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Initializing HtpProvider
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend parameters
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend
[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
/data/sandcastle/boxes/eden-trunk-hg-full-fbsource/buck-out/v2/gen/fbcode/ec7059d5161b31ff/executorch/backends/qualcomm/tests/fb/__test_qnn_delegate_simulator__/test_qnn_delegate_simulator#link-tree/executorch/backends/qualcomm/qnn_preprocess.py:69: Visiting: aten_cumsum_default, aten.cumsum.default
[ERROR] [Qnn ExecuTorch]: tcm_migration.cc:1863:ERROR:no properties registered for q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:210:ERROR:could not create op: q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:1403:ERROR:Op 0x10 preparation failed with err:-1
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> "aten_cumsum_default" generated: could not create op
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> RouterX86 graph prepare failed 12
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> Failed to finalize graph (id: 1) with err 1002
[ERROR] [Qnn ExecuTorch]: Failed to finalize Qnn Graph with error: 1002
[ERROR] [Qnn ExecuTorch]: Fail to compile QNN graph
FAIL
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend

@DannyYuyang-quicDannyYuyang-quic left a comment

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The error message is

[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Initializing HtpProvider
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend parameters
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend
[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
/data/sandcastle/boxes/eden-trunk-hg-full-fbsource/buck-out/v2/gen/fbcode/ec7059d5161b31ff/executorch/backends/qualcomm/tests/fb/__test_qnn_delegate_simulator__/test_qnn_delegate_simulator#link-tree/executorch/backends/qualcomm/qnn_preprocess.py:69: Visiting: aten_cumsum_default, aten.cumsum.default
[ERROR] [Qnn ExecuTorch]: tcm_migration.cc:1863:ERROR:no properties registered for q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:210:ERROR:could not create op: q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:1403:ERROR:Op 0x10 preparation failed with err:-1
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> "aten_cumsum_default" generated: could not create op
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> RouterX86 graph prepare failed 12
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> Failed to finalize graph (id: 1) with err 1002
[ERROR] [Qnn ExecuTorch]: Failed to finalize Qnn Graph with error: 1002
[ERROR] [Qnn ExecuTorch]: Fail to compile QNN graph
FAIL
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend

I tried QNN versions 2.27, 2.28, 2.31, and 2.32 with android-ndk-r26c in this PR but still can't reproduce this error. I think I need more details to tackle this issue. Could you please check which QNN version you used? Also, check the QNN version in $LD_LIBRARY_PATH, thanks!

@cccclai

cccclai commented Mar 25, 2025

Copy link
Copy Markdown
Contributor

I'm currently using qnn 2.28, but not sure the android ndk version. It seems like there are quite a few fp tests continuously failing while most are passing. We can merge this PR for now, but will need some help to identify the root cause. See following for the failing fp tests, the one under fp but not marked are the passing tests.

 # Overwrite because unit test is failing, once passing, remove this function
def test_qnn_backend_interpolate_nearest_2d(self):
pass
# Overwrite because unit test is failing, once passing, remove this function
def test_qnn_backend_interpolate_bilinear_2d(self):
pass
def test_qnn_backend_element_wise_ceil(self):
pass
def test_qnn_backend_embedding(self):
pass
def test_qnn_backend_gelu(self):
pass
def test_qnn_backend_group_norm(self):
pass
def test_qnn_backend_hardsigmoid(self):
pass
def test_qnn_backend_index(self):
pass
def test_qnn_backend_index_put(self):
pass
def test_qnn_backend_instance_norm_2d(self):
pass
def test_qnn_backend_rms_norm(self):
pass
def test_qnn_backend_stack(self):
pass
def test_qnn_backend_where(self):
pass
def test_qnn_backend_conv_transpose2d(self):
pass

We target SM8650 for the unit test btw.

@DannyYuyang-quic

DannyYuyang-quic commented Mar 25, 2025

Copy link
Copy Markdown
ContributorAuthor

I think you can check backend_options = generate_htp_compiler_spec(use_fp16=True) setting.
The use_fp16 option should be set to True in all fp tests, as shown in the figure below.
I'm not sure if you have modified this setting or not~

image

@cccclai

Copy link
Copy Markdown
Contributor

I think you can check backend_options = generate_htp_compiler_spec(use_fp16=True) setting. The use_fp16 option should be set to True in all fp tests, as shown in the figure below. I'm not sure if you have modified this setting or not~

image

Ah good catch. I inherit the base class TestQNNFloatingPointOperator and accidently use TestQNNQuantizedOperator to setup instead of TestQNNFloatingPointOperator. Let me update.

@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

I think you can check backend_options = generate_htp_compiler_spec(use_fp16=True) setting. The use_fp16 option should be set to True in all fp tests, as shown in the figure below. I'm not sure if you have modified this setting or not~
image

Ah good catch. I inherit the base class TestQNNFloatingPointOperator and accidently use TestQNNQuantizedOperator to setup instead of TestQNNFloatingPointOperator. Let me update.

Cool! Let me know if there is any issue.

@cccclai

Copy link
Copy Markdown
Contributor

It should be good now. Mind rebasing?

@DannyYuyang-quic

DannyYuyang-quic commented Apr 2, 2025

Copy link
Copy Markdown
ContributorAuthor

I've rebased the branch. Thanks!

I was on the older base in the last commit, but now it's on the newest one.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclaicccclai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you!

@cccclai

Copy link
Copy Markdown
Contributor

Everything is green, looks like it's still need rebase...mind rebasing again?

 - e2e script for https://github.com/yformer/EfficientSAM
- Fastvit breakage fix
- Add support for cum_sum
- Add bicubic interpolate transform pass
- Fix stack op
@DannyYuyang-quic
DannyYuyang-quicforce-pushed the dev1/danny/EfficientSAM_enablement branch from 4bb1800 to 42412d7CompareApril 10, 2025 06:30
@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

Everything is green, looks like it's still need rebase...mind rebasing again?

Done!
Please have a look.
Thanks!

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai
cccclai merged commit 5e4f045 into pytorch:mainApr 10, 2025
kirklandsign pushed a commit that referenced this pull request Apr 11, 2025
### Summary
- e2e script for https://github.com/yformer/EfficientSAM
- Fastvit breakage fix
- Add support for cum_sum
- Add bicubic interpolate transform pass
- Fix stack op
### Test plan
``` bash
python ./examples/qualcomm/oss_scripts/efficientSAM/efficientSAM.py -m ${soc} -b build-android -H ${host_id} -s ${device_id} --oss_repo ${Path_to_oss_repo} --pretrained_weight ${Path_to_pretrained_weight} -d ${Path_to_dataset_dir}
```
keyprocedure pushed a commit to keyprocedure/executorch that referenced this pull request Apr 21, 2025
…rch#9266)
### Summary
- e2e script for https://github.com/yformer/EfficientSAM
- Fastvit breakage fix
- Add support for cum_sum
- Add bicubic interpolate transform pass
- Fix stack op
### Test plan
``` bash
python ./examples/qualcomm/oss_scripts/efficientSAM/efficientSAM.py -m ${soc} -b build-android -H ${host_id} -s ${device_id} --oss_repo ${Path_to_oss_repo} --pretrained_weight ${Path_to_pretrained_weight} -d ${Path_to_dataset_dir}
```
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.release notes: qualcommChanges to the Qualcomm backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@DannyYuyang-quic@cccclai@facebook-github-bot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Qualcomm AI Engine Direct - oss model enablement (EfficientSAM) - #9266

Merged
cccclai merged 1 commit into
pytorch:mainfrom
CodeLinaro:dev1/danny/EfficientSAM_enablement
Apr 10, 2025
Merged

Qualcomm AI Engine Direct - oss model enablement (EfficientSAM)#9266
cccclai merged 1 commit into
pytorch:mainfrom
CodeLinaro:dev1/danny/EfficientSAM_enablement

Conversation

@DannyYuyang-quic

@DannyYuyang-quicDannyYuyang-quic commented Mar 14, 2025

Copy link
Copy Markdown
Contributor

Summary

Test plan

python ./examples/qualcomm/oss_scripts/efficientSAM/efficientSAM.py -m ${soc} -b build-android -H ${host_id} -s ${device_id} --oss_repo ${Path_to_oss_repo} --pretrained_weight ${Path_to_pretrained_weight} -d ${Path_to_dataset_dir}

@pytorch-bot

pytorch-botBot commented Mar 14, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/9266

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit 42412d7 with merge base c9c5481 (image):

NEW FAILURE - The following job has failed:

  • pull / unittest-arm / linux-job (gh)
    RuntimeError: Command docker exec -t c50233835032a0388f502a805413875acb5526b953584d9a5d02ffbe166713f5 /exec failed with exit code 1

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 14, 2025
@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

@pytorchbot label "release notes: qualcomm"

@pytorch-botpytorch-botBot added the release notes: qualcomm Changes to the Qualcomm backend delegate label Mar 14, 2025
@DannyYuyang-quic
DannyYuyang-quicforce-pushed the dev1/danny/EfficientSAM_enablement branch from f36af71 to 1f614eaCompareMarch 14, 2025 08:56
@DannyYuyang-quic

DannyYuyang-quic commented Mar 14, 2025

Copy link
Copy Markdown
ContributorAuthor

@cccclai
This PR is one of the model enablement requests. Please have a look.
BTW, I can't trigger CI. Could you please help with this?

Thanks!

@cccclaicccclai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for enabling EfficientSAM! As we enable more models, let's add it to part of the CI, similar to #8616

@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

Thank you for enabling EfficientSAM! As we enable more models, let's add it to part of the CI, similar to #8616

Regarding the CI, would you prefer I add it in this PR or add it in the next one?

@cccclai

Copy link
Copy Markdown
Contributor

Thank you for enabling EfficientSAM! As we enable more models, let's add it to part of the CI, similar to #8616

Next PR is fine

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

def test_qnn_backend_cumsum(self):
module = CumSum() # noqa: F405
sample_input = (torch.randn(4),)
self.lower_module_and_test_output(module, sample_input)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems like the fp test is failing, can you double check? The quantized one is passing.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could you please share the log~?
Both the fp and quantized tests are passing from my side.

@cccclai

Copy link
Copy Markdown
Contributor

The error message is

[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Initializing HtpProvider
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend parameters
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend
[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
/data/sandcastle/boxes/eden-trunk-hg-full-fbsource/buck-out/v2/gen/fbcode/ec7059d5161b31ff/executorch/backends/qualcomm/tests/fb/__test_qnn_delegate_simulator__/test_qnn_delegate_simulator#link-tree/executorch/backends/qualcomm/qnn_preprocess.py:69: Visiting: aten_cumsum_default, aten.cumsum.default
[ERROR] [Qnn ExecuTorch]: tcm_migration.cc:1863:ERROR:no properties registered for q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:210:ERROR:could not create op: q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:1403:ERROR:Op 0x10 preparation failed with err:-1
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> "aten_cumsum_default" generated: could not create op
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> RouterX86 graph prepare failed 12
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> Failed to finalize graph (id: 1) with err 1002
[ERROR] [Qnn ExecuTorch]: Failed to finalize Qnn Graph with error: 1002
[ERROR] [Qnn ExecuTorch]: Fail to compile QNN graph
FAIL
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend

@DannyYuyang-quicDannyYuyang-quic left a comment

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The error message is

[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Initializing HtpProvider
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend parameters
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend
[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
/data/sandcastle/boxes/eden-trunk-hg-full-fbsource/buck-out/v2/gen/fbcode/ec7059d5161b31ff/executorch/backends/qualcomm/tests/fb/__test_qnn_delegate_simulator__/test_qnn_delegate_simulator#link-tree/executorch/backends/qualcomm/qnn_preprocess.py:69: Visiting: aten_cumsum_default, aten.cumsum.default
[ERROR] [Qnn ExecuTorch]: tcm_migration.cc:1863:ERROR:no properties registered for q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:210:ERROR:could not create op: q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:1403:ERROR:Op 0x10 preparation failed with err:-1
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> "aten_cumsum_default" generated: could not create op
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> RouterX86 graph prepare failed 12
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> Failed to finalize graph (id: 1) with err 1002
[ERROR] [Qnn ExecuTorch]: Failed to finalize Qnn Graph with error: 1002
[ERROR] [Qnn ExecuTorch]: Fail to compile QNN graph
FAIL
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend

I tried QNN versions 2.27, 2.28, 2.31, and 2.32 with android-ndk-r26c in this PR but still can't reproduce this error. I think I need more details to tackle this issue. Could you please check which QNN version you used? Also, check the QNN version in $LD_LIBRARY_PATH, thanks!

@cccclai

cccclai commented Mar 25, 2025

Copy link
Copy Markdown
Contributor

I'm currently using qnn 2.28, but not sure the android ndk version. It seems like there are quite a few fp tests continuously failing while most are passing. We can merge this PR for now, but will need some help to identify the root cause. See following for the failing fp tests, the one under fp but not marked are the passing tests.

 # Overwrite because unit test is failing, once passing, remove this function
def test_qnn_backend_interpolate_nearest_2d(self):
pass
# Overwrite because unit test is failing, once passing, remove this function
def test_qnn_backend_interpolate_bilinear_2d(self):
pass
def test_qnn_backend_element_wise_ceil(self):
pass
def test_qnn_backend_embedding(self):
pass
def test_qnn_backend_gelu(self):
pass
def test_qnn_backend_group_norm(self):
pass
def test_qnn_backend_hardsigmoid(self):
pass
def test_qnn_backend_index(self):
pass
def test_qnn_backend_index_put(self):
pass
def test_qnn_backend_instance_norm_2d(self):
pass
def test_qnn_backend_rms_norm(self):
pass
def test_qnn_backend_stack(self):
pass
def test_qnn_backend_where(self):
pass
def test_qnn_backend_conv_transpose2d(self):
pass

We target SM8650 for the unit test btw.

@DannyYuyang-quic

DannyYuyang-quic commented Mar 25, 2025

Copy link
Copy Markdown
ContributorAuthor

I think you can check backend_options = generate_htp_compiler_spec(use_fp16=True) setting.
The use_fp16 option should be set to True in all fp tests, as shown in the figure below.
I'm not sure if you have modified this setting or not~

image

@cccclai

Copy link
Copy Markdown
Contributor

I think you can check backend_options = generate_htp_compiler_spec(use_fp16=True) setting. The use_fp16 option should be set to True in all fp tests, as shown in the figure below. I'm not sure if you have modified this setting or not~

image

Ah good catch. I inherit the base class TestQNNFloatingPointOperator and accidently use TestQNNQuantizedOperator to setup instead of TestQNNFloatingPointOperator. Let me update.

@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

I think you can check backend_options = generate_htp_compiler_spec(use_fp16=True) setting. The use_fp16 option should be set to True in all fp tests, as shown in the figure below. I'm not sure if you have modified this setting or not~
image

Ah good catch. I inherit the base class TestQNNFloatingPointOperator and accidently use TestQNNQuantizedOperator to setup instead of TestQNNFloatingPointOperator. Let me update.

Cool! Let me know if there is any issue.

@cccclai

Copy link
Copy Markdown
Contributor

It should be good now. Mind rebasing?

@DannyYuyang-quic

DannyYuyang-quic commented Apr 2, 2025

Copy link
Copy Markdown
ContributorAuthor

I've rebased the branch. Thanks!

I was on the older base in the last commit, but now it's on the newest one.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclaicccclai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you!

@cccclai

Copy link
Copy Markdown
Contributor

Everything is green, looks like it's still need rebase...mind rebasing again?

 - e2e script for https://github.com/yformer/EfficientSAM
- Fastvit breakage fix
- Add support for cum_sum
- Add bicubic interpolate transform pass
- Fix stack op
@DannyYuyang-quic
DannyYuyang-quicforce-pushed the dev1/danny/EfficientSAM_enablement branch from 4bb1800 to 42412d7CompareApril 10, 2025 06:30
@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

Everything is green, looks like it's still need rebase...mind rebasing again?

Done!
Please have a look.
Thanks!

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai
cccclai merged commit 5e4f045 into pytorch:mainApr 10, 2025
kirklandsign pushed a commit that referenced this pull request Apr 11, 2025
### Summary
- e2e script for https://github.com/yformer/EfficientSAM
- Fastvit breakage fix
- Add support for cum_sum
- Add bicubic interpolate transform pass
- Fix stack op
### Test plan
``` bash
python ./examples/qualcomm/oss_scripts/efficientSAM/efficientSAM.py -m ${soc} -b build-android -H ${host_id} -s ${device_id} --oss_repo ${Path_to_oss_repo} --pretrained_weight ${Path_to_pretrained_weight} -d ${Path_to_dataset_dir}
```
keyprocedure pushed a commit to keyprocedure/executorch that referenced this pull request Apr 21, 2025
…rch#9266)
### Summary
- e2e script for https://github.com/yformer/EfficientSAM
- Fastvit breakage fix
- Add support for cum_sum
- Add bicubic interpolate transform pass
- Fix stack op
### Test plan
``` bash
python ./examples/qualcomm/oss_scripts/efficientSAM/efficientSAM.py -m ${soc} -b build-android -H ${host_id} -s ${device_id} --oss_repo ${Path_to_oss_repo} --pretrained_weight ${Path_to_pretrained_weight} -d ${Path_to_dataset_dir}
```
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.release notes: qualcommChanges to the Qualcomm backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@DannyYuyang-quic@cccclai@facebook-github-bot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Qualcomm AI Engine Direct - oss model enablement (EfficientSAM) - #9266

Merged
cccclai merged 1 commit into
pytorch:mainfrom
CodeLinaro:dev1/danny/EfficientSAM_enablement
Apr 10, 2025
Merged

Qualcomm AI Engine Direct - oss model enablement (EfficientSAM)#9266
cccclai merged 1 commit into
pytorch:mainfrom
CodeLinaro:dev1/danny/EfficientSAM_enablement

Conversation

@DannyYuyang-quic

@DannyYuyang-quicDannyYuyang-quic commented Mar 14, 2025

Copy link
Copy Markdown
Contributor

Summary

Test plan

python ./examples/qualcomm/oss_scripts/efficientSAM/efficientSAM.py -m ${soc} -b build-android -H ${host_id} -s ${device_id} --oss_repo ${Path_to_oss_repo} --pretrained_weight ${Path_to_pretrained_weight} -d ${Path_to_dataset_dir}

@pytorch-bot

pytorch-botBot commented Mar 14, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/9266

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit 42412d7 with merge base c9c5481 (image):

NEW FAILURE - The following job has failed:

  • pull / unittest-arm / linux-job (gh)
    RuntimeError: Command docker exec -t c50233835032a0388f502a805413875acb5526b953584d9a5d02ffbe166713f5 /exec failed with exit code 1

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 14, 2025
@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

@pytorchbot label "release notes: qualcomm"

@pytorch-botpytorch-botBot added the release notes: qualcomm Changes to the Qualcomm backend delegate label Mar 14, 2025
@DannyYuyang-quic
DannyYuyang-quicforce-pushed the dev1/danny/EfficientSAM_enablement branch from f36af71 to 1f614eaCompareMarch 14, 2025 08:56
@DannyYuyang-quic

DannyYuyang-quic commented Mar 14, 2025

Copy link
Copy Markdown
ContributorAuthor

@cccclai
This PR is one of the model enablement requests. Please have a look.
BTW, I can't trigger CI. Could you please help with this?

Thanks!

@cccclaicccclai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for enabling EfficientSAM! As we enable more models, let's add it to part of the CI, similar to #8616

@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

Thank you for enabling EfficientSAM! As we enable more models, let's add it to part of the CI, similar to #8616

Regarding the CI, would you prefer I add it in this PR or add it in the next one?

@cccclai

Copy link
Copy Markdown
Contributor

Thank you for enabling EfficientSAM! As we enable more models, let's add it to part of the CI, similar to #8616

Next PR is fine

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

def test_qnn_backend_cumsum(self):
module = CumSum() # noqa: F405
sample_input = (torch.randn(4),)
self.lower_module_and_test_output(module, sample_input)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems like the fp test is failing, can you double check? The quantized one is passing.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could you please share the log~?
Both the fp and quantized tests are passing from my side.

@cccclai

Copy link
Copy Markdown
Contributor

The error message is

[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Initializing HtpProvider
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend parameters
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend
[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
/data/sandcastle/boxes/eden-trunk-hg-full-fbsource/buck-out/v2/gen/fbcode/ec7059d5161b31ff/executorch/backends/qualcomm/tests/fb/__test_qnn_delegate_simulator__/test_qnn_delegate_simulator#link-tree/executorch/backends/qualcomm/qnn_preprocess.py:69: Visiting: aten_cumsum_default, aten.cumsum.default
[ERROR] [Qnn ExecuTorch]: tcm_migration.cc:1863:ERROR:no properties registered for q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:210:ERROR:could not create op: q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:1403:ERROR:Op 0x10 preparation failed with err:-1
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> "aten_cumsum_default" generated: could not create op
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> RouterX86 graph prepare failed 12
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> Failed to finalize graph (id: 1) with err 1002
[ERROR] [Qnn ExecuTorch]: Failed to finalize Qnn Graph with error: 1002
[ERROR] [Qnn ExecuTorch]: Fail to compile QNN graph
FAIL
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend

@DannyYuyang-quicDannyYuyang-quic left a comment

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The error message is

[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Initializing HtpProvider
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend parameters
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend
[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
/data/sandcastle/boxes/eden-trunk-hg-full-fbsource/buck-out/v2/gen/fbcode/ec7059d5161b31ff/executorch/backends/qualcomm/tests/fb/__test_qnn_delegate_simulator__/test_qnn_delegate_simulator#link-tree/executorch/backends/qualcomm/qnn_preprocess.py:69: Visiting: aten_cumsum_default, aten.cumsum.default
[ERROR] [Qnn ExecuTorch]: tcm_migration.cc:1863:ERROR:no properties registered for q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:210:ERROR:could not create op: q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:1403:ERROR:Op 0x10 preparation failed with err:-1
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> "aten_cumsum_default" generated: could not create op
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> RouterX86 graph prepare failed 12
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> Failed to finalize graph (id: 1) with err 1002
[ERROR] [Qnn ExecuTorch]: Failed to finalize Qnn Graph with error: 1002
[ERROR] [Qnn ExecuTorch]: Fail to compile QNN graph
FAIL
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend

I tried QNN versions 2.27, 2.28, 2.31, and 2.32 with android-ndk-r26c in this PR but still can't reproduce this error. I think I need more details to tackle this issue. Could you please check which QNN version you used? Also, check the QNN version in $LD_LIBRARY_PATH, thanks!

@cccclai

cccclai commented Mar 25, 2025

Copy link
Copy Markdown
Contributor

I'm currently using qnn 2.28, but not sure the android ndk version. It seems like there are quite a few fp tests continuously failing while most are passing. We can merge this PR for now, but will need some help to identify the root cause. See following for the failing fp tests, the one under fp but not marked are the passing tests.

 # Overwrite because unit test is failing, once passing, remove this function
def test_qnn_backend_interpolate_nearest_2d(self):
pass
# Overwrite because unit test is failing, once passing, remove this function
def test_qnn_backend_interpolate_bilinear_2d(self):
pass
def test_qnn_backend_element_wise_ceil(self):
pass
def test_qnn_backend_embedding(self):
pass
def test_qnn_backend_gelu(self):
pass
def test_qnn_backend_group_norm(self):
pass
def test_qnn_backend_hardsigmoid(self):
pass
def test_qnn_backend_index(self):
pass
def test_qnn_backend_index_put(self):
pass
def test_qnn_backend_instance_norm_2d(self):
pass
def test_qnn_backend_rms_norm(self):
pass
def test_qnn_backend_stack(self):
pass
def test_qnn_backend_where(self):
pass
def test_qnn_backend_conv_transpose2d(self):
pass

We target SM8650 for the unit test btw.

@DannyYuyang-quic

DannyYuyang-quic commented Mar 25, 2025

Copy link
Copy Markdown
ContributorAuthor

I think you can check backend_options = generate_htp_compiler_spec(use_fp16=True) setting.
The use_fp16 option should be set to True in all fp tests, as shown in the figure below.
I'm not sure if you have modified this setting or not~

image

@cccclai

Copy link
Copy Markdown
Contributor

I think you can check backend_options = generate_htp_compiler_spec(use_fp16=True) setting. The use_fp16 option should be set to True in all fp tests, as shown in the figure below. I'm not sure if you have modified this setting or not~

image

Ah good catch. I inherit the base class TestQNNFloatingPointOperator and accidently use TestQNNQuantizedOperator to setup instead of TestQNNFloatingPointOperator. Let me update.

@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

I think you can check backend_options = generate_htp_compiler_spec(use_fp16=True) setting. The use_fp16 option should be set to True in all fp tests, as shown in the figure below. I'm not sure if you have modified this setting or not~
image

Ah good catch. I inherit the base class TestQNNFloatingPointOperator and accidently use TestQNNQuantizedOperator to setup instead of TestQNNFloatingPointOperator. Let me update.

Cool! Let me know if there is any issue.

@cccclai

Copy link
Copy Markdown
Contributor

It should be good now. Mind rebasing?

@DannyYuyang-quic

DannyYuyang-quic commented Apr 2, 2025

Copy link
Copy Markdown
ContributorAuthor

I've rebased the branch. Thanks!

I was on the older base in the last commit, but now it's on the newest one.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclaicccclai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you!

@cccclai

Copy link
Copy Markdown
Contributor

Everything is green, looks like it's still need rebase...mind rebasing again?

 - e2e script for https://github.com/yformer/EfficientSAM
- Fastvit breakage fix
- Add support for cum_sum
- Add bicubic interpolate transform pass
- Fix stack op
@DannyYuyang-quic
DannyYuyang-quicforce-pushed the dev1/danny/EfficientSAM_enablement branch from 4bb1800 to 42412d7CompareApril 10, 2025 06:30
@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

Everything is green, looks like it's still need rebase...mind rebasing again?

Done!
Please have a look.
Thanks!

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai
cccclai merged commit 5e4f045 into pytorch:mainApr 10, 2025
kirklandsign pushed a commit that referenced this pull request Apr 11, 2025
### Summary
- e2e script for https://github.com/yformer/EfficientSAM
- Fastvit breakage fix
- Add support for cum_sum
- Add bicubic interpolate transform pass
- Fix stack op
### Test plan
``` bash
python ./examples/qualcomm/oss_scripts/efficientSAM/efficientSAM.py -m ${soc} -b build-android -H ${host_id} -s ${device_id} --oss_repo ${Path_to_oss_repo} --pretrained_weight ${Path_to_pretrained_weight} -d ${Path_to_dataset_dir}
```
keyprocedure pushed a commit to keyprocedure/executorch that referenced this pull request Apr 21, 2025
…rch#9266)
### Summary
- e2e script for https://github.com/yformer/EfficientSAM
- Fastvit breakage fix
- Add support for cum_sum
- Add bicubic interpolate transform pass
- Fix stack op
### Test plan
``` bash
python ./examples/qualcomm/oss_scripts/efficientSAM/efficientSAM.py -m ${soc} -b build-android -H ${host_id} -s ${device_id} --oss_repo ${Path_to_oss_repo} --pretrained_weight ${Path_to_pretrained_weight} -d ${Path_to_dataset_dir}
```
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.release notes: qualcommChanges to the Qualcomm backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@DannyYuyang-quic@cccclai@facebook-github-bot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Qualcomm AI Engine Direct - oss model enablement (EfficientSAM) - #9266

Merged
cccclai merged 1 commit into
pytorch:mainfrom
CodeLinaro:dev1/danny/EfficientSAM_enablement
Apr 10, 2025
Merged

Qualcomm AI Engine Direct - oss model enablement (EfficientSAM)#9266
cccclai merged 1 commit into
pytorch:mainfrom
CodeLinaro:dev1/danny/EfficientSAM_enablement

Conversation

@DannyYuyang-quic

@DannyYuyang-quicDannyYuyang-quic commented Mar 14, 2025

Copy link
Copy Markdown
Contributor

Summary

Test plan

python ./examples/qualcomm/oss_scripts/efficientSAM/efficientSAM.py -m ${soc} -b build-android -H ${host_id} -s ${device_id} --oss_repo ${Path_to_oss_repo} --pretrained_weight ${Path_to_pretrained_weight} -d ${Path_to_dataset_dir}

@pytorch-bot

pytorch-botBot commented Mar 14, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/9266

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit 42412d7 with merge base c9c5481 (image):

NEW FAILURE - The following job has failed:

  • pull / unittest-arm / linux-job (gh)
    RuntimeError: Command docker exec -t c50233835032a0388f502a805413875acb5526b953584d9a5d02ffbe166713f5 /exec failed with exit code 1

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 14, 2025
@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

@pytorchbot label "release notes: qualcomm"

@pytorch-botpytorch-botBot added the release notes: qualcomm Changes to the Qualcomm backend delegate label Mar 14, 2025
@DannyYuyang-quic
DannyYuyang-quicforce-pushed the dev1/danny/EfficientSAM_enablement branch from f36af71 to 1f614eaCompareMarch 14, 2025 08:56
@DannyYuyang-quic

DannyYuyang-quic commented Mar 14, 2025

Copy link
Copy Markdown
ContributorAuthor

@cccclai
This PR is one of the model enablement requests. Please have a look.
BTW, I can't trigger CI. Could you please help with this?

Thanks!

@cccclaicccclai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for enabling EfficientSAM! As we enable more models, let's add it to part of the CI, similar to #8616

@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

Thank you for enabling EfficientSAM! As we enable more models, let's add it to part of the CI, similar to #8616

Regarding the CI, would you prefer I add it in this PR or add it in the next one?

@cccclai

Copy link
Copy Markdown
Contributor

Thank you for enabling EfficientSAM! As we enable more models, let's add it to part of the CI, similar to #8616

Next PR is fine

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

def test_qnn_backend_cumsum(self):
module = CumSum() # noqa: F405
sample_input = (torch.randn(4),)
self.lower_module_and_test_output(module, sample_input)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems like the fp test is failing, can you double check? The quantized one is passing.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could you please share the log~?
Both the fp and quantized tests are passing from my side.

@cccclai

Copy link
Copy Markdown
Contributor

The error message is

[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Initializing HtpProvider
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend parameters
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend
[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
/data/sandcastle/boxes/eden-trunk-hg-full-fbsource/buck-out/v2/gen/fbcode/ec7059d5161b31ff/executorch/backends/qualcomm/tests/fb/__test_qnn_delegate_simulator__/test_qnn_delegate_simulator#link-tree/executorch/backends/qualcomm/qnn_preprocess.py:69: Visiting: aten_cumsum_default, aten.cumsum.default
[ERROR] [Qnn ExecuTorch]: tcm_migration.cc:1863:ERROR:no properties registered for q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:210:ERROR:could not create op: q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:1403:ERROR:Op 0x10 preparation failed with err:-1
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> "aten_cumsum_default" generated: could not create op
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> RouterX86 graph prepare failed 12
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> Failed to finalize graph (id: 1) with err 1002
[ERROR] [Qnn ExecuTorch]: Failed to finalize Qnn Graph with error: 1002
[ERROR] [Qnn ExecuTorch]: Fail to compile QNN graph
FAIL
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend

@DannyYuyang-quicDannyYuyang-quic left a comment

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The error message is

[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Initializing HtpProvider
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend parameters
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend
[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
/data/sandcastle/boxes/eden-trunk-hg-full-fbsource/buck-out/v2/gen/fbcode/ec7059d5161b31ff/executorch/backends/qualcomm/tests/fb/__test_qnn_delegate_simulator__/test_qnn_delegate_simulator#link-tree/executorch/backends/qualcomm/qnn_preprocess.py:69: Visiting: aten_cumsum_default, aten.cumsum.default
[ERROR] [Qnn ExecuTorch]: tcm_migration.cc:1863:ERROR:no properties registered for q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:210:ERROR:could not create op: q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:1403:ERROR:Op 0x10 preparation failed with err:-1
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> "aten_cumsum_default" generated: could not create op
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> RouterX86 graph prepare failed 12
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> Failed to finalize graph (id: 1) with err 1002
[ERROR] [Qnn ExecuTorch]: Failed to finalize Qnn Graph with error: 1002
[ERROR] [Qnn ExecuTorch]: Fail to compile QNN graph
FAIL
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend

I tried QNN versions 2.27, 2.28, 2.31, and 2.32 with android-ndk-r26c in this PR but still can't reproduce this error. I think I need more details to tackle this issue. Could you please check which QNN version you used? Also, check the QNN version in $LD_LIBRARY_PATH, thanks!

@cccclai

cccclai commented Mar 25, 2025

Copy link
Copy Markdown
Contributor

I'm currently using qnn 2.28, but not sure the android ndk version. It seems like there are quite a few fp tests continuously failing while most are passing. We can merge this PR for now, but will need some help to identify the root cause. See following for the failing fp tests, the one under fp but not marked are the passing tests.

 # Overwrite because unit test is failing, once passing, remove this function
def test_qnn_backend_interpolate_nearest_2d(self):
pass
# Overwrite because unit test is failing, once passing, remove this function
def test_qnn_backend_interpolate_bilinear_2d(self):
pass
def test_qnn_backend_element_wise_ceil(self):
pass
def test_qnn_backend_embedding(self):
pass
def test_qnn_backend_gelu(self):
pass
def test_qnn_backend_group_norm(self):
pass
def test_qnn_backend_hardsigmoid(self):
pass
def test_qnn_backend_index(self):
pass
def test_qnn_backend_index_put(self):
pass
def test_qnn_backend_instance_norm_2d(self):
pass
def test_qnn_backend_rms_norm(self):
pass
def test_qnn_backend_stack(self):
pass
def test_qnn_backend_where(self):
pass
def test_qnn_backend_conv_transpose2d(self):
pass

We target SM8650 for the unit test btw.

@DannyYuyang-quic

DannyYuyang-quic commented Mar 25, 2025

Copy link
Copy Markdown
ContributorAuthor

I think you can check backend_options = generate_htp_compiler_spec(use_fp16=True) setting.
The use_fp16 option should be set to True in all fp tests, as shown in the figure below.
I'm not sure if you have modified this setting or not~

image

@cccclai

Copy link
Copy Markdown
Contributor

I think you can check backend_options = generate_htp_compiler_spec(use_fp16=True) setting. The use_fp16 option should be set to True in all fp tests, as shown in the figure below. I'm not sure if you have modified this setting or not~

image

Ah good catch. I inherit the base class TestQNNFloatingPointOperator and accidently use TestQNNQuantizedOperator to setup instead of TestQNNFloatingPointOperator. Let me update.

@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

I think you can check backend_options = generate_htp_compiler_spec(use_fp16=True) setting. The use_fp16 option should be set to True in all fp tests, as shown in the figure below. I'm not sure if you have modified this setting or not~
image

Ah good catch. I inherit the base class TestQNNFloatingPointOperator and accidently use TestQNNQuantizedOperator to setup instead of TestQNNFloatingPointOperator. Let me update.

Cool! Let me know if there is any issue.

@cccclai

Copy link
Copy Markdown
Contributor

It should be good now. Mind rebasing?

@DannyYuyang-quic

DannyYuyang-quic commented Apr 2, 2025

Copy link
Copy Markdown
ContributorAuthor

I've rebased the branch. Thanks!

I was on the older base in the last commit, but now it's on the newest one.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclaicccclai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you!

@cccclai

Copy link
Copy Markdown
Contributor

Everything is green, looks like it's still need rebase...mind rebasing again?

 - e2e script for https://github.com/yformer/EfficientSAM
- Fastvit breakage fix
- Add support for cum_sum
- Add bicubic interpolate transform pass
- Fix stack op
@DannyYuyang-quic
DannyYuyang-quicforce-pushed the dev1/danny/EfficientSAM_enablement branch from 4bb1800 to 42412d7CompareApril 10, 2025 06:30
@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

Everything is green, looks like it's still need rebase...mind rebasing again?

Done!
Please have a look.
Thanks!

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai
cccclai merged commit 5e4f045 into pytorch:mainApr 10, 2025
kirklandsign pushed a commit that referenced this pull request Apr 11, 2025
### Summary
- e2e script for https://github.com/yformer/EfficientSAM
- Fastvit breakage fix
- Add support for cum_sum
- Add bicubic interpolate transform pass
- Fix stack op
### Test plan
``` bash
python ./examples/qualcomm/oss_scripts/efficientSAM/efficientSAM.py -m ${soc} -b build-android -H ${host_id} -s ${device_id} --oss_repo ${Path_to_oss_repo} --pretrained_weight ${Path_to_pretrained_weight} -d ${Path_to_dataset_dir}
```
keyprocedure pushed a commit to keyprocedure/executorch that referenced this pull request Apr 21, 2025
…rch#9266)
### Summary
- e2e script for https://github.com/yformer/EfficientSAM
- Fastvit breakage fix
- Add support for cum_sum
- Add bicubic interpolate transform pass
- Fix stack op
### Test plan
``` bash
python ./examples/qualcomm/oss_scripts/efficientSAM/efficientSAM.py -m ${soc} -b build-android -H ${host_id} -s ${device_id} --oss_repo ${Path_to_oss_repo} --pretrained_weight ${Path_to_pretrained_weight} -d ${Path_to_dataset_dir}
```
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.release notes: qualcommChanges to the Qualcomm backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@DannyYuyang-quic@cccclai@facebook-github-bot
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Qualcomm AI Engine Direct - oss model enablement (EfficientSAM) - #9266

Merged
cccclai merged 1 commit into
pytorch:mainfrom
CodeLinaro:dev1/danny/EfficientSAM_enablement
Apr 10, 2025
Merged

Qualcomm AI Engine Direct - oss model enablement (EfficientSAM)#9266
cccclai merged 1 commit into
pytorch:mainfrom
CodeLinaro:dev1/danny/EfficientSAM_enablement

Conversation

@DannyYuyang-quic

@DannyYuyang-quicDannyYuyang-quic commented Mar 14, 2025

Copy link
Copy Markdown
Contributor

Summary

Test plan

python ./examples/qualcomm/oss_scripts/efficientSAM/efficientSAM.py -m ${soc} -b build-android -H ${host_id} -s ${device_id} --oss_repo ${Path_to_oss_repo} --pretrained_weight ${Path_to_pretrained_weight} -d ${Path_to_dataset_dir}

@pytorch-bot

pytorch-botBot commented Mar 14, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/9266

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit 42412d7 with merge base c9c5481 (image):

NEW FAILURE - The following job has failed:

  • pull / unittest-arm / linux-job (gh)
    RuntimeError: Command docker exec -t c50233835032a0388f502a805413875acb5526b953584d9a5d02ffbe166713f5 /exec failed with exit code 1

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 14, 2025
@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

@pytorchbot label "release notes: qualcomm"

@pytorch-botpytorch-botBot added the release notes: qualcomm Changes to the Qualcomm backend delegate label Mar 14, 2025
@DannyYuyang-quic
DannyYuyang-quicforce-pushed the dev1/danny/EfficientSAM_enablement branch from f36af71 to 1f614eaCompareMarch 14, 2025 08:56
@DannyYuyang-quic

DannyYuyang-quic commented Mar 14, 2025

Copy link
Copy Markdown
ContributorAuthor

@cccclai
This PR is one of the model enablement requests. Please have a look.
BTW, I can't trigger CI. Could you please help with this?

Thanks!

@cccclaicccclai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for enabling EfficientSAM! As we enable more models, let's add it to part of the CI, similar to #8616

@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

Thank you for enabling EfficientSAM! As we enable more models, let's add it to part of the CI, similar to #8616

Regarding the CI, would you prefer I add it in this PR or add it in the next one?

@cccclai

Copy link
Copy Markdown
Contributor

Thank you for enabling EfficientSAM! As we enable more models, let's add it to part of the CI, similar to #8616

Next PR is fine

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

def test_qnn_backend_cumsum(self):
module = CumSum() # noqa: F405
sample_input = (torch.randn(4),)
self.lower_module_and_test_output(module, sample_input)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems like the fp test is failing, can you double check? The quantized one is passing.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could you please share the log~?
Both the fp and quantized tests are passing from my side.

@cccclai

Copy link
Copy Markdown
Contributor

The error message is

[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Initializing HtpProvider
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend parameters
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend
[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
/data/sandcastle/boxes/eden-trunk-hg-full-fbsource/buck-out/v2/gen/fbcode/ec7059d5161b31ff/executorch/backends/qualcomm/tests/fb/__test_qnn_delegate_simulator__/test_qnn_delegate_simulator#link-tree/executorch/backends/qualcomm/qnn_preprocess.py:69: Visiting: aten_cumsum_default, aten.cumsum.default
[ERROR] [Qnn ExecuTorch]: tcm_migration.cc:1863:ERROR:no properties registered for q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:210:ERROR:could not create op: q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:1403:ERROR:Op 0x10 preparation failed with err:-1
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> "aten_cumsum_default" generated: could not create op
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> RouterX86 graph prepare failed 12
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> Failed to finalize graph (id: 1) with err 1002
[ERROR] [Qnn ExecuTorch]: Failed to finalize Qnn Graph with error: 1002
[ERROR] [Qnn ExecuTorch]: Fail to compile QNN graph
FAIL
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend

@DannyYuyang-quicDannyYuyang-quic left a comment

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The error message is

[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Initializing HtpProvider
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend parameters
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend
[INFO] [Qnn ExecuTorch]: create QNN Logger with log_level 2
[INFO] [Qnn ExecuTorch]: Initialize Qnn backend parameters for Qnn executorch backend type 2
[INFO] [Qnn ExecuTorch]: Caching: Caching is in SAVE MODE.
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Performance Estimates unsupported
[WARNING] [Qnn ExecuTorch]: QnnDsp <W> Arch 68 set by custom config is different from arch associated with SoC 57, will overwrite it to 75
[INFO] [Qnn ExecuTorch]: Running level=3 optimization.
/data/sandcastle/boxes/eden-trunk-hg-full-fbsource/buck-out/v2/gen/fbcode/ec7059d5161b31ff/executorch/backends/qualcomm/tests/fb/__test_qnn_delegate_simulator__/test_qnn_delegate_simulator#link-tree/executorch/backends/qualcomm/qnn_preprocess.py:69: Visiting: aten_cumsum_default, aten.cumsum.default
[ERROR] [Qnn ExecuTorch]: tcm_migration.cc:1863:ERROR:no properties registered for q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:210:ERROR:could not create op: q::QNN_CumulativeSum
[ERROR] [Qnn ExecuTorch]: graph_prepare.cc:1403:ERROR:Op 0x10 preparation failed with err:-1
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> "aten_cumsum_default" generated: could not create op
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> RouterX86 graph prepare failed 12
[ERROR] [Qnn ExecuTorch]: QnnDsp <E> Failed to finalize graph (id: 1) with err 1002
[ERROR] [Qnn ExecuTorch]: Failed to finalize Qnn Graph with error: 1002
[ERROR] [Qnn ExecuTorch]: Fail to compile QNN graph
FAIL
[INFO] [Qnn ExecuTorch]: Destroy Qnn context
[INFO] [Qnn ExecuTorch]: Destroy Qnn device
[INFO] [Qnn ExecuTorch]: Destroy Qnn backend

I tried QNN versions 2.27, 2.28, 2.31, and 2.32 with android-ndk-r26c in this PR but still can't reproduce this error. I think I need more details to tackle this issue. Could you please check which QNN version you used? Also, check the QNN version in $LD_LIBRARY_PATH, thanks!

@cccclai

cccclai commented Mar 25, 2025

Copy link
Copy Markdown
Contributor

I'm currently using qnn 2.28, but not sure the android ndk version. It seems like there are quite a few fp tests continuously failing while most are passing. We can merge this PR for now, but will need some help to identify the root cause. See following for the failing fp tests, the one under fp but not marked are the passing tests.

 # Overwrite because unit test is failing, once passing, remove this function
def test_qnn_backend_interpolate_nearest_2d(self):
pass
# Overwrite because unit test is failing, once passing, remove this function
def test_qnn_backend_interpolate_bilinear_2d(self):
pass
def test_qnn_backend_element_wise_ceil(self):
pass
def test_qnn_backend_embedding(self):
pass
def test_qnn_backend_gelu(self):
pass
def test_qnn_backend_group_norm(self):
pass
def test_qnn_backend_hardsigmoid(self):
pass
def test_qnn_backend_index(self):
pass
def test_qnn_backend_index_put(self):
pass
def test_qnn_backend_instance_norm_2d(self):
pass
def test_qnn_backend_rms_norm(self):
pass
def test_qnn_backend_stack(self):
pass
def test_qnn_backend_where(self):
pass
def test_qnn_backend_conv_transpose2d(self):
pass

We target SM8650 for the unit test btw.

@DannyYuyang-quic

DannyYuyang-quic commented Mar 25, 2025

Copy link
Copy Markdown
ContributorAuthor

I think you can check backend_options = generate_htp_compiler_spec(use_fp16=True) setting.
The use_fp16 option should be set to True in all fp tests, as shown in the figure below.
I'm not sure if you have modified this setting or not~

image

@cccclai

Copy link
Copy Markdown
Contributor

I think you can check backend_options = generate_htp_compiler_spec(use_fp16=True) setting. The use_fp16 option should be set to True in all fp tests, as shown in the figure below. I'm not sure if you have modified this setting or not~

image

Ah good catch. I inherit the base class TestQNNFloatingPointOperator and accidently use TestQNNQuantizedOperator to setup instead of TestQNNFloatingPointOperator. Let me update.

@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

I think you can check backend_options = generate_htp_compiler_spec(use_fp16=True) setting. The use_fp16 option should be set to True in all fp tests, as shown in the figure below. I'm not sure if you have modified this setting or not~
image

Ah good catch. I inherit the base class TestQNNFloatingPointOperator and accidently use TestQNNQuantizedOperator to setup instead of TestQNNFloatingPointOperator. Let me update.

Cool! Let me know if there is any issue.

@cccclai

Copy link
Copy Markdown
Contributor

It should be good now. Mind rebasing?

@DannyYuyang-quic

DannyYuyang-quic commented Apr 2, 2025

Copy link
Copy Markdown
ContributorAuthor

I've rebased the branch. Thanks!

I was on the older base in the last commit, but now it's on the newest one.

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclaicccclai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you!

@cccclai

Copy link
Copy Markdown
Contributor

Everything is green, looks like it's still need rebase...mind rebasing again?

 - e2e script for https://github.com/yformer/EfficientSAM
- Fastvit breakage fix
- Add support for cum_sum
- Add bicubic interpolate transform pass
- Fix stack op
@DannyYuyang-quic
DannyYuyang-quicforce-pushed the dev1/danny/EfficientSAM_enablement branch from 4bb1800 to 42412d7CompareApril 10, 2025 06:30
@DannyYuyang-quic

Copy link
Copy Markdown
ContributorAuthor

Everything is green, looks like it's still need rebase...mind rebasing again?

Done!
Please have a look.
Thanks!

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai
cccclai merged commit 5e4f045 into pytorch:mainApr 10, 2025
kirklandsign pushed a commit that referenced this pull request Apr 11, 2025
### Summary
- e2e script for https://github.com/yformer/EfficientSAM
- Fastvit breakage fix
- Add support for cum_sum
- Add bicubic interpolate transform pass
- Fix stack op
### Test plan
``` bash
python ./examples/qualcomm/oss_scripts/efficientSAM/efficientSAM.py -m ${soc} -b build-android -H ${host_id} -s ${device_id} --oss_repo ${Path_to_oss_repo} --pretrained_weight ${Path_to_pretrained_weight} -d ${Path_to_dataset_dir}
```
keyprocedure pushed a commit to keyprocedure/executorch that referenced this pull request Apr 21, 2025
…rch#9266)
### Summary
- e2e script for https://github.com/yformer/EfficientSAM
- Fastvit breakage fix
- Add support for cum_sum
- Add bicubic interpolate transform pass
- Fix stack op
### Test plan
``` bash
python ./examples/qualcomm/oss_scripts/efficientSAM/efficientSAM.py -m ${soc} -b build-android -H ${host_id} -s ${device_id} --oss_repo ${Path_to_oss_repo} --pretrained_weight ${Path_to_pretrained_weight} -d ${Path_to_dataset_dir}
```
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.release notes: qualcommChanges to the Qualcomm backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@DannyYuyang-quic@cccclai@facebook-github-bot