Qualcomm AI Engine Direct - Optimize performance of pcq embedding - #20686

Merged
psiddh merged 3 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_pcq_embedding
Jul 7, 2026
Merged

Qualcomm AI Engine Direct - Optimize performance of pcq embedding#20686
psiddh merged 3 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_pcq_embedding

Conversation

@shewu-quic

Copy link
Copy Markdown
Collaborator

Summary:

  • Change pcq embedding pattern for backend optimization
    • Note that it is supported after QNN 2.48

Test plan

python3 backends/qualcomm/tests/test_qnn_delegate.py TestQNNQuantizedOperator.test_qnn_backend_embedding_per_channel --build_folder build-android --host {HOST} --device {DEVICE} --soc_model SM8850 -a {ARTIFACTS} 

@pytorch-bot

pytorch-botBot commented Jul 2, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20686

Note: Links to docs will display an error until the docs builds have been completed.

❌ 2 New Failures

As of commit a551ca3 with merge base f41d66d (image):

NEW FAILURES - The following jobs have failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jul 2, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_pcq_embedding branch from 231b11e to 11bd40cCompareJuly 2, 2026 04:05
@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Hi @psiddh,
This PR is to optimize performance for PCQ embedding on HTP.
Could you please have a look?
Thanks,
Hutton

@psiddh

Copy link
Copy Markdown
Contributor

@claude Can you review this PR ?

@claude

claudeBot commented Jul 6, 2026

Copy link
Copy Markdown

Claude finished @psiddh's task in 3m 23s —— View job


PR Review: Optimize performance of PCQ embedding

  • Read the PR diff and changed files
  • Analyze op_embedding.py restructure
  • Analyze test and llm_wrappers.py changes
  • Post review

Nice cleanup — the new pattern is clearly better than the old one. Here's my review.

Summary of the change

The core change in op_embedding.py inverts the op order for per-channel-quantized (PCQ) embeddings:

  • Before:Gather(per-channel weight) → intermediate output where the gather output was forced to a per-tensor approximation (scale = max(scales), zp = max(zero_points)), then a Convert requantized to the activation encoding. That max-scale approximation is exactly what caused the accuracy problem the test was skipped for ("Bad accuracy for HTP").
  • After:Convert(per-channel weight) → per-tensor activation-encoded weight table → Gather → output. The per-channel→per-tensor requant now happens on the weight table itself (properly, per-channel), and gather just indexes it. This removes the lossy max(scales) heuristic entirely, which is why the test can now be enabled.

This is a solid structural improvement, and dropping the node.meta[QCOM_QUANT_ATTRS] mutate-and-restore dance (node_quant_attrs/intermediate_quant_attrs) removes a fragile side effect. Removing the now-unused QCOM_QUANT_MAX/MIN/SCALE(S)/ZERO_POINT(S) imports is good hygiene.

Correctness notes (looks right)

  • The final output wrapper is now registered under node.name (op_embedding.py:96-103), so downstream consumers resolve the correct tensor. The Convert output is an intermediate _convert native tensor feeding into Gather. ✅
  • Gather input (converted weight) and Gather output share the same activation encoding via get_quant_encoding_conf(node, node) — correct, since gather only indexes and must not change encoding. ✅
  • op_wrapper_list ordering is correct: convert_op is appended before gather_op, so the convert executes first. ✅

Questions / things worth confirming

1. Memory/perf of converting the full weight table (op_embedding.py:71-89)
The Convert now runs over the entire embedding table (dims=weight_tensor.size(), i.e. vocab × hidden) instead of over the small gather output. For LLM-sized vocabularies this is a large tensor. Since the convert input is a QNN_TENSOR_TYPE_STATIC weight, I assume the intent is that HTP constant-folds this convert at compile time so only the gather runs at inference — which would make this a genuine perf win. Can you confirm the convert-of-static is folded (and doesn't materialize a full-size runtime NATIVE tensor each inference)? A one-line note in the code or PR would help future readers understand why this is faster.

2. llm_wrappers.py getattr/setattr asymmetry (llm_wrappers.py:785-797)
The new _get_attr correctly traverses dotted targets (e.g. static_llama.tok_embedding.weight) for the read side. But the write side still uses a flat setattr(unquantized_model, unquantized_node.target, ...). If unquantized_node.target can also be dotted, setattr will create a single literal attribute named "a.b.c" on the model rather than setting model.a.b.c. Is unquantized_node.target guaranteed to be flat here, while only quantized_node.target is nested? If not, the write path needs the same traversal.

3. Version gate
is_qnn_sdk_version_less_than("2.48") only compares within the same major version (current_major == target_major and current_minor < target_minor). So on a hypothetical QNN 3.x the check returns False and the test runs — which is the desired behavior here, just flagging that the helper is major-scoped by design.

Minor

  • Test now covers both use_16a8w and use_16a4w with subTest — good added coverage. Consider adding a short comment noting the QNN 2.48 requirement mirrors the backend change (the skip message "UT pass after QNN 2.48" covers it, so optional).

Overall this looks like a clean, well-motivated change. My main ask is confirming point #1 (that the static convert is folded so this is actually faster, not just more accurate) and clarifying the setattr asymmetry in #2.

@psiddh

Copy link
Copy Markdown
Contributor

One qq: The test is skipped below QNN 2.48. does that mean the new pcq-embedding pattern only works on 2.48+,
or does it still work on 2.37 and 2.48 is just where the accuracy test passes?

attr = getattr(attr, target_atom)
return attr

setattr(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should setattr match the getattr walk pattern here?

@psiddh

Copy link
Copy Markdown
Contributor

Also just for my understanding : The convert now runs over the whole weight table instead of just the gathered rows. Since the weight is static, is this convert constant-folded during QNN compile (so it's computed once and the runtime
graph is just the gather)? Just want to confirm it's a one-time AOT cost and not a per-inference convert of the full table.

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

One qq: The test is skipped below QNN 2.48. does that mean the new pcq-embedding pattern only works on 2.48+, or does it still work on 2.37 and 2.48 is just where the accuracy test passes?

Yes, this new optimize pattern is only supported after QNN 2.48. Otherwise, it will failed to compile

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Also just for my understanding : The convert now runs over the whole weight table instead of just the gathered rows. Since the weight is static, is this convert constant-folded during QNN compile (so it's computed once and the runtime graph is just the gather)? Just want to confirm it's a one-time AOT cost and not a per-inference convert of the full table.

Yes, it will be fused by QNN HTP during AOT.

@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_pcq_embedding branch from f749d49 to f105816CompareJuly 6, 2026 14:28
@psiddh

Copy link
Copy Markdown
Contributor

One qq: The test is skipped below QNN 2.48. does that mean the new pcq-embedding pattern only works on 2.48+, or does it still work on 2.37 and 2.48 is just where the accuracy test passes?

Yes, this new optimize pattern is only supported after QNN 2.48. Otherwise, it will failed to compile

Thanks for confirming. Since it fails to compile below 2.48 and the builder emits this pattern unconditionally, won't pcq embedding now break for anyone on our pinned 2.37? Should we either gate the builder on is_qnn_sdk_version_less_than("2.48") (fall back to the old pattern below 2.48), or bump the minimum QNN version to 2.48 and move CI accordingly? Otherwise the test skip hides a compile failure for 2.37 users

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Thanks for confirming. Since it fails to compile below 2.48 and the builder emits this pattern unconditionally, won't pcq embedding now break for anyone on our pinned 2.37? Should we either gate the builder on is_qnn_sdk_version_less_than("2.48") (fall back to the old pattern below 2.48), or bump the minimum QNN version to 2.48 and move CI accordingly? Otherwise the test skip hides a compile failure for 2.37 users

Thanks for flagging that. I agree — I keep the legacy pattern for pcq embedding to maintain backward compatibility.

@psiddh

Copy link
Copy Markdown
Contributor

Thanks for all the follow up changes.. lgtm

@psiddh
psiddh merged commit 7c3f2ed into pytorch:mainJul 7, 2026
191 of 193 checks passed
psiddh added a commit that referenced this pull request Jul 21, 2026
Summary:
D110960687 (#20686) rewrote `Embedding.define_node` to
select between the optimized and legacy pcq-embedding lowerings based on
`is_qnn_sdk_version_less_than("2.48")`. That helper resolves the SDK
version via
`get_sdk_build_id`, which builds a path from
`os.environ["QNN_SDK_ROOT"]`.
`define_node` runs during AOT partitioning (`is_node_supported`), and
AOT-only
environments do not necessarily have `QNN_SDK_ROOT` set. In that case
`os.path.join(os.environ.get("QNN_SDK_ROOT", None), ...)` raised
`TypeError: expected str, bytes or os.PathLike object, not NoneType`,
breaking
every QNN lowering that contains an embedding — including the internal
`test_dummy_llama_qnn_16a4w_aot_and_runtime`.
Fall back to the legacy embedding lowering (valid on all QNN versions)
when
`QNN_SDK_ROOT` is unavailable, so the version-gated optimization is only
taken
when the SDK version can actually be determined. Also make
`get_sdk_build_id`
raise a clear `EnvironmentError` instead of a cryptic `TypeError` when
the
environment variable is missing.
This diff was authored with Claude Code.
Differential Revision: D112944232
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@shewu-quic@psiddh
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Qualcomm AI Engine Direct - Optimize performance of pcq embedding - #20686

Merged
psiddh merged 3 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_pcq_embedding
Jul 7, 2026
Merged

Qualcomm AI Engine Direct - Optimize performance of pcq embedding#20686
psiddh merged 3 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_pcq_embedding

Conversation

@shewu-quic

Copy link
Copy Markdown
Collaborator

Summary:

  • Change pcq embedding pattern for backend optimization
    • Note that it is supported after QNN 2.48

Test plan

python3 backends/qualcomm/tests/test_qnn_delegate.py TestQNNQuantizedOperator.test_qnn_backend_embedding_per_channel --build_folder build-android --host {HOST} --device {DEVICE} --soc_model SM8850 -a {ARTIFACTS} 

@pytorch-bot

pytorch-botBot commented Jul 2, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20686

Note: Links to docs will display an error until the docs builds have been completed.

❌ 2 New Failures

As of commit a551ca3 with merge base f41d66d (image):

NEW FAILURES - The following jobs have failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jul 2, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_pcq_embedding branch from 231b11e to 11bd40cCompareJuly 2, 2026 04:05
@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Hi @psiddh,
This PR is to optimize performance for PCQ embedding on HTP.
Could you please have a look?
Thanks,
Hutton

@psiddh

Copy link
Copy Markdown
Contributor

@claude Can you review this PR ?

@claude

claudeBot commented Jul 6, 2026

Copy link
Copy Markdown

Claude finished @psiddh's task in 3m 23s —— View job


PR Review: Optimize performance of PCQ embedding

  • Read the PR diff and changed files
  • Analyze op_embedding.py restructure
  • Analyze test and llm_wrappers.py changes
  • Post review

Nice cleanup — the new pattern is clearly better than the old one. Here's my review.

Summary of the change

The core change in op_embedding.py inverts the op order for per-channel-quantized (PCQ) embeddings:

  • Before:Gather(per-channel weight) → intermediate output where the gather output was forced to a per-tensor approximation (scale = max(scales), zp = max(zero_points)), then a Convert requantized to the activation encoding. That max-scale approximation is exactly what caused the accuracy problem the test was skipped for ("Bad accuracy for HTP").
  • After:Convert(per-channel weight) → per-tensor activation-encoded weight table → Gather → output. The per-channel→per-tensor requant now happens on the weight table itself (properly, per-channel), and gather just indexes it. This removes the lossy max(scales) heuristic entirely, which is why the test can now be enabled.

This is a solid structural improvement, and dropping the node.meta[QCOM_QUANT_ATTRS] mutate-and-restore dance (node_quant_attrs/intermediate_quant_attrs) removes a fragile side effect. Removing the now-unused QCOM_QUANT_MAX/MIN/SCALE(S)/ZERO_POINT(S) imports is good hygiene.

Correctness notes (looks right)

  • The final output wrapper is now registered under node.name (op_embedding.py:96-103), so downstream consumers resolve the correct tensor. The Convert output is an intermediate _convert native tensor feeding into Gather. ✅
  • Gather input (converted weight) and Gather output share the same activation encoding via get_quant_encoding_conf(node, node) — correct, since gather only indexes and must not change encoding. ✅
  • op_wrapper_list ordering is correct: convert_op is appended before gather_op, so the convert executes first. ✅

Questions / things worth confirming

1. Memory/perf of converting the full weight table (op_embedding.py:71-89)
The Convert now runs over the entire embedding table (dims=weight_tensor.size(), i.e. vocab × hidden) instead of over the small gather output. For LLM-sized vocabularies this is a large tensor. Since the convert input is a QNN_TENSOR_TYPE_STATIC weight, I assume the intent is that HTP constant-folds this convert at compile time so only the gather runs at inference — which would make this a genuine perf win. Can you confirm the convert-of-static is folded (and doesn't materialize a full-size runtime NATIVE tensor each inference)? A one-line note in the code or PR would help future readers understand why this is faster.

2. llm_wrappers.py getattr/setattr asymmetry (llm_wrappers.py:785-797)
The new _get_attr correctly traverses dotted targets (e.g. static_llama.tok_embedding.weight) for the read side. But the write side still uses a flat setattr(unquantized_model, unquantized_node.target, ...). If unquantized_node.target can also be dotted, setattr will create a single literal attribute named "a.b.c" on the model rather than setting model.a.b.c. Is unquantized_node.target guaranteed to be flat here, while only quantized_node.target is nested? If not, the write path needs the same traversal.

3. Version gate
is_qnn_sdk_version_less_than("2.48") only compares within the same major version (current_major == target_major and current_minor < target_minor). So on a hypothetical QNN 3.x the check returns False and the test runs — which is the desired behavior here, just flagging that the helper is major-scoped by design.

Minor

  • Test now covers both use_16a8w and use_16a4w with subTest — good added coverage. Consider adding a short comment noting the QNN 2.48 requirement mirrors the backend change (the skip message "UT pass after QNN 2.48" covers it, so optional).

Overall this looks like a clean, well-motivated change. My main ask is confirming point #1 (that the static convert is folded so this is actually faster, not just more accurate) and clarifying the setattr asymmetry in #2.

@psiddh

Copy link
Copy Markdown
Contributor

One qq: The test is skipped below QNN 2.48. does that mean the new pcq-embedding pattern only works on 2.48+,
or does it still work on 2.37 and 2.48 is just where the accuracy test passes?

attr = getattr(attr, target_atom)
return attr

setattr(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should setattr match the getattr walk pattern here?

@psiddh

Copy link
Copy Markdown
Contributor

Also just for my understanding : The convert now runs over the whole weight table instead of just the gathered rows. Since the weight is static, is this convert constant-folded during QNN compile (so it's computed once and the runtime
graph is just the gather)? Just want to confirm it's a one-time AOT cost and not a per-inference convert of the full table.

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

One qq: The test is skipped below QNN 2.48. does that mean the new pcq-embedding pattern only works on 2.48+, or does it still work on 2.37 and 2.48 is just where the accuracy test passes?

Yes, this new optimize pattern is only supported after QNN 2.48. Otherwise, it will failed to compile

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Also just for my understanding : The convert now runs over the whole weight table instead of just the gathered rows. Since the weight is static, is this convert constant-folded during QNN compile (so it's computed once and the runtime graph is just the gather)? Just want to confirm it's a one-time AOT cost and not a per-inference convert of the full table.

Yes, it will be fused by QNN HTP during AOT.

@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_pcq_embedding branch from f749d49 to f105816CompareJuly 6, 2026 14:28
@psiddh

Copy link
Copy Markdown
Contributor

One qq: The test is skipped below QNN 2.48. does that mean the new pcq-embedding pattern only works on 2.48+, or does it still work on 2.37 and 2.48 is just where the accuracy test passes?

Yes, this new optimize pattern is only supported after QNN 2.48. Otherwise, it will failed to compile

Thanks for confirming. Since it fails to compile below 2.48 and the builder emits this pattern unconditionally, won't pcq embedding now break for anyone on our pinned 2.37? Should we either gate the builder on is_qnn_sdk_version_less_than("2.48") (fall back to the old pattern below 2.48), or bump the minimum QNN version to 2.48 and move CI accordingly? Otherwise the test skip hides a compile failure for 2.37 users

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Thanks for confirming. Since it fails to compile below 2.48 and the builder emits this pattern unconditionally, won't pcq embedding now break for anyone on our pinned 2.37? Should we either gate the builder on is_qnn_sdk_version_less_than("2.48") (fall back to the old pattern below 2.48), or bump the minimum QNN version to 2.48 and move CI accordingly? Otherwise the test skip hides a compile failure for 2.37 users

Thanks for flagging that. I agree — I keep the legacy pattern for pcq embedding to maintain backward compatibility.

@psiddh

Copy link
Copy Markdown
Contributor

Thanks for all the follow up changes.. lgtm

@psiddh
psiddh merged commit 7c3f2ed into pytorch:mainJul 7, 2026
191 of 193 checks passed
psiddh added a commit that referenced this pull request Jul 21, 2026
Summary:
D110960687 (#20686) rewrote `Embedding.define_node` to
select between the optimized and legacy pcq-embedding lowerings based on
`is_qnn_sdk_version_less_than("2.48")`. That helper resolves the SDK
version via
`get_sdk_build_id`, which builds a path from
`os.environ["QNN_SDK_ROOT"]`.
`define_node` runs during AOT partitioning (`is_node_supported`), and
AOT-only
environments do not necessarily have `QNN_SDK_ROOT` set. In that case
`os.path.join(os.environ.get("QNN_SDK_ROOT", None), ...)` raised
`TypeError: expected str, bytes or os.PathLike object, not NoneType`,
breaking
every QNN lowering that contains an embedding — including the internal
`test_dummy_llama_qnn_16a4w_aot_and_runtime`.
Fall back to the legacy embedding lowering (valid on all QNN versions)
when
`QNN_SDK_ROOT` is unavailable, so the version-gated optimization is only
taken
when the SDK version can actually be determined. Also make
`get_sdk_build_id`
raise a clear `EnvironmentError` instead of a cryptic `TypeError` when
the
environment variable is missing.
This diff was authored with Claude Code.
Differential Revision: D112944232
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@shewu-quic@psiddh
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Qualcomm AI Engine Direct - Optimize performance of pcq embedding - #20686

Merged
psiddh merged 3 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_pcq_embedding
Jul 7, 2026
Merged

Qualcomm AI Engine Direct - Optimize performance of pcq embedding#20686
psiddh merged 3 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_pcq_embedding

Conversation

@shewu-quic

Copy link
Copy Markdown
Collaborator

Summary:

  • Change pcq embedding pattern for backend optimization
    • Note that it is supported after QNN 2.48

Test plan

python3 backends/qualcomm/tests/test_qnn_delegate.py TestQNNQuantizedOperator.test_qnn_backend_embedding_per_channel --build_folder build-android --host {HOST} --device {DEVICE} --soc_model SM8850 -a {ARTIFACTS} 

@pytorch-bot

pytorch-botBot commented Jul 2, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20686

Note: Links to docs will display an error until the docs builds have been completed.

❌ 2 New Failures

As of commit a551ca3 with merge base f41d66d (image):

NEW FAILURES - The following jobs have failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jul 2, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_pcq_embedding branch from 231b11e to 11bd40cCompareJuly 2, 2026 04:05
@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Hi @psiddh,
This PR is to optimize performance for PCQ embedding on HTP.
Could you please have a look?
Thanks,
Hutton

@psiddh

Copy link
Copy Markdown
Contributor

@claude Can you review this PR ?

@claude

claudeBot commented Jul 6, 2026

Copy link
Copy Markdown

Claude finished @psiddh's task in 3m 23s —— View job


PR Review: Optimize performance of PCQ embedding

  • Read the PR diff and changed files
  • Analyze op_embedding.py restructure
  • Analyze test and llm_wrappers.py changes
  • Post review

Nice cleanup — the new pattern is clearly better than the old one. Here's my review.

Summary of the change

The core change in op_embedding.py inverts the op order for per-channel-quantized (PCQ) embeddings:

  • Before:Gather(per-channel weight) → intermediate output where the gather output was forced to a per-tensor approximation (scale = max(scales), zp = max(zero_points)), then a Convert requantized to the activation encoding. That max-scale approximation is exactly what caused the accuracy problem the test was skipped for ("Bad accuracy for HTP").
  • After:Convert(per-channel weight) → per-tensor activation-encoded weight table → Gather → output. The per-channel→per-tensor requant now happens on the weight table itself (properly, per-channel), and gather just indexes it. This removes the lossy max(scales) heuristic entirely, which is why the test can now be enabled.

This is a solid structural improvement, and dropping the node.meta[QCOM_QUANT_ATTRS] mutate-and-restore dance (node_quant_attrs/intermediate_quant_attrs) removes a fragile side effect. Removing the now-unused QCOM_QUANT_MAX/MIN/SCALE(S)/ZERO_POINT(S) imports is good hygiene.

Correctness notes (looks right)

  • The final output wrapper is now registered under node.name (op_embedding.py:96-103), so downstream consumers resolve the correct tensor. The Convert output is an intermediate _convert native tensor feeding into Gather. ✅
  • Gather input (converted weight) and Gather output share the same activation encoding via get_quant_encoding_conf(node, node) — correct, since gather only indexes and must not change encoding. ✅
  • op_wrapper_list ordering is correct: convert_op is appended before gather_op, so the convert executes first. ✅

Questions / things worth confirming

1. Memory/perf of converting the full weight table (op_embedding.py:71-89)
The Convert now runs over the entire embedding table (dims=weight_tensor.size(), i.e. vocab × hidden) instead of over the small gather output. For LLM-sized vocabularies this is a large tensor. Since the convert input is a QNN_TENSOR_TYPE_STATIC weight, I assume the intent is that HTP constant-folds this convert at compile time so only the gather runs at inference — which would make this a genuine perf win. Can you confirm the convert-of-static is folded (and doesn't materialize a full-size runtime NATIVE tensor each inference)? A one-line note in the code or PR would help future readers understand why this is faster.

2. llm_wrappers.py getattr/setattr asymmetry (llm_wrappers.py:785-797)
The new _get_attr correctly traverses dotted targets (e.g. static_llama.tok_embedding.weight) for the read side. But the write side still uses a flat setattr(unquantized_model, unquantized_node.target, ...). If unquantized_node.target can also be dotted, setattr will create a single literal attribute named "a.b.c" on the model rather than setting model.a.b.c. Is unquantized_node.target guaranteed to be flat here, while only quantized_node.target is nested? If not, the write path needs the same traversal.

3. Version gate
is_qnn_sdk_version_less_than("2.48") only compares within the same major version (current_major == target_major and current_minor < target_minor). So on a hypothetical QNN 3.x the check returns False and the test runs — which is the desired behavior here, just flagging that the helper is major-scoped by design.

Minor

  • Test now covers both use_16a8w and use_16a4w with subTest — good added coverage. Consider adding a short comment noting the QNN 2.48 requirement mirrors the backend change (the skip message "UT pass after QNN 2.48" covers it, so optional).

Overall this looks like a clean, well-motivated change. My main ask is confirming point #1 (that the static convert is folded so this is actually faster, not just more accurate) and clarifying the setattr asymmetry in #2.

@psiddh

Copy link
Copy Markdown
Contributor

One qq: The test is skipped below QNN 2.48. does that mean the new pcq-embedding pattern only works on 2.48+,
or does it still work on 2.37 and 2.48 is just where the accuracy test passes?

attr = getattr(attr, target_atom)
return attr

setattr(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should setattr match the getattr walk pattern here?

@psiddh

Copy link
Copy Markdown
Contributor

Also just for my understanding : The convert now runs over the whole weight table instead of just the gathered rows. Since the weight is static, is this convert constant-folded during QNN compile (so it's computed once and the runtime
graph is just the gather)? Just want to confirm it's a one-time AOT cost and not a per-inference convert of the full table.

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

One qq: The test is skipped below QNN 2.48. does that mean the new pcq-embedding pattern only works on 2.48+, or does it still work on 2.37 and 2.48 is just where the accuracy test passes?

Yes, this new optimize pattern is only supported after QNN 2.48. Otherwise, it will failed to compile

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Also just for my understanding : The convert now runs over the whole weight table instead of just the gathered rows. Since the weight is static, is this convert constant-folded during QNN compile (so it's computed once and the runtime graph is just the gather)? Just want to confirm it's a one-time AOT cost and not a per-inference convert of the full table.

Yes, it will be fused by QNN HTP during AOT.

@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_pcq_embedding branch from f749d49 to f105816CompareJuly 6, 2026 14:28
@psiddh

Copy link
Copy Markdown
Contributor

One qq: The test is skipped below QNN 2.48. does that mean the new pcq-embedding pattern only works on 2.48+, or does it still work on 2.37 and 2.48 is just where the accuracy test passes?

Yes, this new optimize pattern is only supported after QNN 2.48. Otherwise, it will failed to compile

Thanks for confirming. Since it fails to compile below 2.48 and the builder emits this pattern unconditionally, won't pcq embedding now break for anyone on our pinned 2.37? Should we either gate the builder on is_qnn_sdk_version_less_than("2.48") (fall back to the old pattern below 2.48), or bump the minimum QNN version to 2.48 and move CI accordingly? Otherwise the test skip hides a compile failure for 2.37 users

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Thanks for confirming. Since it fails to compile below 2.48 and the builder emits this pattern unconditionally, won't pcq embedding now break for anyone on our pinned 2.37? Should we either gate the builder on is_qnn_sdk_version_less_than("2.48") (fall back to the old pattern below 2.48), or bump the minimum QNN version to 2.48 and move CI accordingly? Otherwise the test skip hides a compile failure for 2.37 users

Thanks for flagging that. I agree — I keep the legacy pattern for pcq embedding to maintain backward compatibility.

@psiddh

Copy link
Copy Markdown
Contributor

Thanks for all the follow up changes.. lgtm

@psiddh
psiddh merged commit 7c3f2ed into pytorch:mainJul 7, 2026
191 of 193 checks passed
psiddh added a commit that referenced this pull request Jul 21, 2026
Summary:
D110960687 (#20686) rewrote `Embedding.define_node` to
select between the optimized and legacy pcq-embedding lowerings based on
`is_qnn_sdk_version_less_than("2.48")`. That helper resolves the SDK
version via
`get_sdk_build_id`, which builds a path from
`os.environ["QNN_SDK_ROOT"]`.
`define_node` runs during AOT partitioning (`is_node_supported`), and
AOT-only
environments do not necessarily have `QNN_SDK_ROOT` set. In that case
`os.path.join(os.environ.get("QNN_SDK_ROOT", None), ...)` raised
`TypeError: expected str, bytes or os.PathLike object, not NoneType`,
breaking
every QNN lowering that contains an embedding — including the internal
`test_dummy_llama_qnn_16a4w_aot_and_runtime`.
Fall back to the legacy embedding lowering (valid on all QNN versions)
when
`QNN_SDK_ROOT` is unavailable, so the version-gated optimization is only
taken
when the SDK version can actually be determined. Also make
`get_sdk_build_id`
raise a clear `EnvironmentError` instead of a cryptic `TypeError` when
the
environment variable is missing.
This diff was authored with Claude Code.
Differential Revision: D112944232
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@shewu-quic@psiddh
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Qualcomm AI Engine Direct - Optimize performance of pcq embedding - #20686

Merged
psiddh merged 3 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_pcq_embedding
Jul 7, 2026
Merged

Qualcomm AI Engine Direct - Optimize performance of pcq embedding#20686
psiddh merged 3 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_pcq_embedding

Conversation

@shewu-quic

Copy link
Copy Markdown
Collaborator

Summary:

  • Change pcq embedding pattern for backend optimization
    • Note that it is supported after QNN 2.48

Test plan

python3 backends/qualcomm/tests/test_qnn_delegate.py TestQNNQuantizedOperator.test_qnn_backend_embedding_per_channel --build_folder build-android --host {HOST} --device {DEVICE} --soc_model SM8850 -a {ARTIFACTS} 

@pytorch-bot

pytorch-botBot commented Jul 2, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20686

Note: Links to docs will display an error until the docs builds have been completed.

❌ 2 New Failures

As of commit a551ca3 with merge base f41d66d (image):

NEW FAILURES - The following jobs have failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jul 2, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_pcq_embedding branch from 231b11e to 11bd40cCompareJuly 2, 2026 04:05
@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Hi @psiddh,
This PR is to optimize performance for PCQ embedding on HTP.
Could you please have a look?
Thanks,
Hutton

@psiddh

Copy link
Copy Markdown
Contributor

@claude Can you review this PR ?

@claude

claudeBot commented Jul 6, 2026

Copy link
Copy Markdown

Claude finished @psiddh's task in 3m 23s —— View job


PR Review: Optimize performance of PCQ embedding

  • Read the PR diff and changed files
  • Analyze op_embedding.py restructure
  • Analyze test and llm_wrappers.py changes
  • Post review

Nice cleanup — the new pattern is clearly better than the old one. Here's my review.

Summary of the change

The core change in op_embedding.py inverts the op order for per-channel-quantized (PCQ) embeddings:

  • Before:Gather(per-channel weight) → intermediate output where the gather output was forced to a per-tensor approximation (scale = max(scales), zp = max(zero_points)), then a Convert requantized to the activation encoding. That max-scale approximation is exactly what caused the accuracy problem the test was skipped for ("Bad accuracy for HTP").
  • After:Convert(per-channel weight) → per-tensor activation-encoded weight table → Gather → output. The per-channel→per-tensor requant now happens on the weight table itself (properly, per-channel), and gather just indexes it. This removes the lossy max(scales) heuristic entirely, which is why the test can now be enabled.

This is a solid structural improvement, and dropping the node.meta[QCOM_QUANT_ATTRS] mutate-and-restore dance (node_quant_attrs/intermediate_quant_attrs) removes a fragile side effect. Removing the now-unused QCOM_QUANT_MAX/MIN/SCALE(S)/ZERO_POINT(S) imports is good hygiene.

Correctness notes (looks right)

  • The final output wrapper is now registered under node.name (op_embedding.py:96-103), so downstream consumers resolve the correct tensor. The Convert output is an intermediate _convert native tensor feeding into Gather. ✅
  • Gather input (converted weight) and Gather output share the same activation encoding via get_quant_encoding_conf(node, node) — correct, since gather only indexes and must not change encoding. ✅
  • op_wrapper_list ordering is correct: convert_op is appended before gather_op, so the convert executes first. ✅

Questions / things worth confirming

1. Memory/perf of converting the full weight table (op_embedding.py:71-89)
The Convert now runs over the entire embedding table (dims=weight_tensor.size(), i.e. vocab × hidden) instead of over the small gather output. For LLM-sized vocabularies this is a large tensor. Since the convert input is a QNN_TENSOR_TYPE_STATIC weight, I assume the intent is that HTP constant-folds this convert at compile time so only the gather runs at inference — which would make this a genuine perf win. Can you confirm the convert-of-static is folded (and doesn't materialize a full-size runtime NATIVE tensor each inference)? A one-line note in the code or PR would help future readers understand why this is faster.

2. llm_wrappers.py getattr/setattr asymmetry (llm_wrappers.py:785-797)
The new _get_attr correctly traverses dotted targets (e.g. static_llama.tok_embedding.weight) for the read side. But the write side still uses a flat setattr(unquantized_model, unquantized_node.target, ...). If unquantized_node.target can also be dotted, setattr will create a single literal attribute named "a.b.c" on the model rather than setting model.a.b.c. Is unquantized_node.target guaranteed to be flat here, while only quantized_node.target is nested? If not, the write path needs the same traversal.

3. Version gate
is_qnn_sdk_version_less_than("2.48") only compares within the same major version (current_major == target_major and current_minor < target_minor). So on a hypothetical QNN 3.x the check returns False and the test runs — which is the desired behavior here, just flagging that the helper is major-scoped by design.

Minor

  • Test now covers both use_16a8w and use_16a4w with subTest — good added coverage. Consider adding a short comment noting the QNN 2.48 requirement mirrors the backend change (the skip message "UT pass after QNN 2.48" covers it, so optional).

Overall this looks like a clean, well-motivated change. My main ask is confirming point #1 (that the static convert is folded so this is actually faster, not just more accurate) and clarifying the setattr asymmetry in #2.

@psiddh

Copy link
Copy Markdown
Contributor

One qq: The test is skipped below QNN 2.48. does that mean the new pcq-embedding pattern only works on 2.48+,
or does it still work on 2.37 and 2.48 is just where the accuracy test passes?

attr = getattr(attr, target_atom)
return attr

setattr(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should setattr match the getattr walk pattern here?

@psiddh

Copy link
Copy Markdown
Contributor

Also just for my understanding : The convert now runs over the whole weight table instead of just the gathered rows. Since the weight is static, is this convert constant-folded during QNN compile (so it's computed once and the runtime
graph is just the gather)? Just want to confirm it's a one-time AOT cost and not a per-inference convert of the full table.

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

One qq: The test is skipped below QNN 2.48. does that mean the new pcq-embedding pattern only works on 2.48+, or does it still work on 2.37 and 2.48 is just where the accuracy test passes?

Yes, this new optimize pattern is only supported after QNN 2.48. Otherwise, it will failed to compile

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Also just for my understanding : The convert now runs over the whole weight table instead of just the gathered rows. Since the weight is static, is this convert constant-folded during QNN compile (so it's computed once and the runtime graph is just the gather)? Just want to confirm it's a one-time AOT cost and not a per-inference convert of the full table.

Yes, it will be fused by QNN HTP during AOT.

@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_pcq_embedding branch from f749d49 to f105816CompareJuly 6, 2026 14:28
@psiddh

Copy link
Copy Markdown
Contributor

One qq: The test is skipped below QNN 2.48. does that mean the new pcq-embedding pattern only works on 2.48+, or does it still work on 2.37 and 2.48 is just where the accuracy test passes?

Yes, this new optimize pattern is only supported after QNN 2.48. Otherwise, it will failed to compile

Thanks for confirming. Since it fails to compile below 2.48 and the builder emits this pattern unconditionally, won't pcq embedding now break for anyone on our pinned 2.37? Should we either gate the builder on is_qnn_sdk_version_less_than("2.48") (fall back to the old pattern below 2.48), or bump the minimum QNN version to 2.48 and move CI accordingly? Otherwise the test skip hides a compile failure for 2.37 users

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Thanks for confirming. Since it fails to compile below 2.48 and the builder emits this pattern unconditionally, won't pcq embedding now break for anyone on our pinned 2.37? Should we either gate the builder on is_qnn_sdk_version_less_than("2.48") (fall back to the old pattern below 2.48), or bump the minimum QNN version to 2.48 and move CI accordingly? Otherwise the test skip hides a compile failure for 2.37 users

Thanks for flagging that. I agree — I keep the legacy pattern for pcq embedding to maintain backward compatibility.

@psiddh

Copy link
Copy Markdown
Contributor

Thanks for all the follow up changes.. lgtm

@psiddh
psiddh merged commit 7c3f2ed into pytorch:mainJul 7, 2026
191 of 193 checks passed
psiddh added a commit that referenced this pull request Jul 21, 2026
Summary:
D110960687 (#20686) rewrote `Embedding.define_node` to
select between the optimized and legacy pcq-embedding lowerings based on
`is_qnn_sdk_version_less_than("2.48")`. That helper resolves the SDK
version via
`get_sdk_build_id`, which builds a path from
`os.environ["QNN_SDK_ROOT"]`.
`define_node` runs during AOT partitioning (`is_node_supported`), and
AOT-only
environments do not necessarily have `QNN_SDK_ROOT` set. In that case
`os.path.join(os.environ.get("QNN_SDK_ROOT", None), ...)` raised
`TypeError: expected str, bytes or os.PathLike object, not NoneType`,
breaking
every QNN lowering that contains an embedding — including the internal
`test_dummy_llama_qnn_16a4w_aot_and_runtime`.
Fall back to the legacy embedding lowering (valid on all QNN versions)
when
`QNN_SDK_ROOT` is unavailable, so the version-gated optimization is only
taken
when the SDK version can actually be determined. Also make
`get_sdk_build_id`
raise a clear `EnvironmentError` instead of a cryptic `TypeError` when
the
environment variable is missing.
This diff was authored with Claude Code.
Differential Revision: D112944232
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@shewu-quic@psiddh
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Qualcomm AI Engine Direct - Optimize performance of pcq embedding - #20686

Merged
psiddh merged 3 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_pcq_embedding
Jul 7, 2026
Merged

Qualcomm AI Engine Direct - Optimize performance of pcq embedding#20686
psiddh merged 3 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_pcq_embedding

Conversation

@shewu-quic

Copy link
Copy Markdown
Collaborator

Summary:

  • Change pcq embedding pattern for backend optimization
    • Note that it is supported after QNN 2.48

Test plan

python3 backends/qualcomm/tests/test_qnn_delegate.py TestQNNQuantizedOperator.test_qnn_backend_embedding_per_channel --build_folder build-android --host {HOST} --device {DEVICE} --soc_model SM8850 -a {ARTIFACTS} 

@pytorch-bot

pytorch-botBot commented Jul 2, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20686

Note: Links to docs will display an error until the docs builds have been completed.

❌ 2 New Failures

As of commit a551ca3 with merge base f41d66d (image):

NEW FAILURES - The following jobs have failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jul 2, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_pcq_embedding branch from 231b11e to 11bd40cCompareJuly 2, 2026 04:05
@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Hi @psiddh,
This PR is to optimize performance for PCQ embedding on HTP.
Could you please have a look?
Thanks,
Hutton

@psiddh

Copy link
Copy Markdown
Contributor

@claude Can you review this PR ?

@claude

claudeBot commented Jul 6, 2026

Copy link
Copy Markdown

Claude finished @psiddh's task in 3m 23s —— View job


PR Review: Optimize performance of PCQ embedding

  • Read the PR diff and changed files
  • Analyze op_embedding.py restructure
  • Analyze test and llm_wrappers.py changes
  • Post review

Nice cleanup — the new pattern is clearly better than the old one. Here's my review.

Summary of the change

The core change in op_embedding.py inverts the op order for per-channel-quantized (PCQ) embeddings:

  • Before:Gather(per-channel weight) → intermediate output where the gather output was forced to a per-tensor approximation (scale = max(scales), zp = max(zero_points)), then a Convert requantized to the activation encoding. That max-scale approximation is exactly what caused the accuracy problem the test was skipped for ("Bad accuracy for HTP").
  • After:Convert(per-channel weight) → per-tensor activation-encoded weight table → Gather → output. The per-channel→per-tensor requant now happens on the weight table itself (properly, per-channel), and gather just indexes it. This removes the lossy max(scales) heuristic entirely, which is why the test can now be enabled.

This is a solid structural improvement, and dropping the node.meta[QCOM_QUANT_ATTRS] mutate-and-restore dance (node_quant_attrs/intermediate_quant_attrs) removes a fragile side effect. Removing the now-unused QCOM_QUANT_MAX/MIN/SCALE(S)/ZERO_POINT(S) imports is good hygiene.

Correctness notes (looks right)

  • The final output wrapper is now registered under node.name (op_embedding.py:96-103), so downstream consumers resolve the correct tensor. The Convert output is an intermediate _convert native tensor feeding into Gather. ✅
  • Gather input (converted weight) and Gather output share the same activation encoding via get_quant_encoding_conf(node, node) — correct, since gather only indexes and must not change encoding. ✅
  • op_wrapper_list ordering is correct: convert_op is appended before gather_op, so the convert executes first. ✅

Questions / things worth confirming

1. Memory/perf of converting the full weight table (op_embedding.py:71-89)
The Convert now runs over the entire embedding table (dims=weight_tensor.size(), i.e. vocab × hidden) instead of over the small gather output. For LLM-sized vocabularies this is a large tensor. Since the convert input is a QNN_TENSOR_TYPE_STATIC weight, I assume the intent is that HTP constant-folds this convert at compile time so only the gather runs at inference — which would make this a genuine perf win. Can you confirm the convert-of-static is folded (and doesn't materialize a full-size runtime NATIVE tensor each inference)? A one-line note in the code or PR would help future readers understand why this is faster.

2. llm_wrappers.py getattr/setattr asymmetry (llm_wrappers.py:785-797)
The new _get_attr correctly traverses dotted targets (e.g. static_llama.tok_embedding.weight) for the read side. But the write side still uses a flat setattr(unquantized_model, unquantized_node.target, ...). If unquantized_node.target can also be dotted, setattr will create a single literal attribute named "a.b.c" on the model rather than setting model.a.b.c. Is unquantized_node.target guaranteed to be flat here, while only quantized_node.target is nested? If not, the write path needs the same traversal.

3. Version gate
is_qnn_sdk_version_less_than("2.48") only compares within the same major version (current_major == target_major and current_minor < target_minor). So on a hypothetical QNN 3.x the check returns False and the test runs — which is the desired behavior here, just flagging that the helper is major-scoped by design.

Minor

  • Test now covers both use_16a8w and use_16a4w with subTest — good added coverage. Consider adding a short comment noting the QNN 2.48 requirement mirrors the backend change (the skip message "UT pass after QNN 2.48" covers it, so optional).

Overall this looks like a clean, well-motivated change. My main ask is confirming point #1 (that the static convert is folded so this is actually faster, not just more accurate) and clarifying the setattr asymmetry in #2.

@psiddh

Copy link
Copy Markdown
Contributor

One qq: The test is skipped below QNN 2.48. does that mean the new pcq-embedding pattern only works on 2.48+,
or does it still work on 2.37 and 2.48 is just where the accuracy test passes?

attr = getattr(attr, target_atom)
return attr

setattr(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should setattr match the getattr walk pattern here?

@psiddh

Copy link
Copy Markdown
Contributor

Also just for my understanding : The convert now runs over the whole weight table instead of just the gathered rows. Since the weight is static, is this convert constant-folded during QNN compile (so it's computed once and the runtime
graph is just the gather)? Just want to confirm it's a one-time AOT cost and not a per-inference convert of the full table.

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

One qq: The test is skipped below QNN 2.48. does that mean the new pcq-embedding pattern only works on 2.48+, or does it still work on 2.37 and 2.48 is just where the accuracy test passes?

Yes, this new optimize pattern is only supported after QNN 2.48. Otherwise, it will failed to compile

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Also just for my understanding : The convert now runs over the whole weight table instead of just the gathered rows. Since the weight is static, is this convert constant-folded during QNN compile (so it's computed once and the runtime graph is just the gather)? Just want to confirm it's a one-time AOT cost and not a per-inference convert of the full table.

Yes, it will be fused by QNN HTP during AOT.

@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_pcq_embedding branch from f749d49 to f105816CompareJuly 6, 2026 14:28
@psiddh

Copy link
Copy Markdown
Contributor

One qq: The test is skipped below QNN 2.48. does that mean the new pcq-embedding pattern only works on 2.48+, or does it still work on 2.37 and 2.48 is just where the accuracy test passes?

Yes, this new optimize pattern is only supported after QNN 2.48. Otherwise, it will failed to compile

Thanks for confirming. Since it fails to compile below 2.48 and the builder emits this pattern unconditionally, won't pcq embedding now break for anyone on our pinned 2.37? Should we either gate the builder on is_qnn_sdk_version_less_than("2.48") (fall back to the old pattern below 2.48), or bump the minimum QNN version to 2.48 and move CI accordingly? Otherwise the test skip hides a compile failure for 2.37 users

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Thanks for confirming. Since it fails to compile below 2.48 and the builder emits this pattern unconditionally, won't pcq embedding now break for anyone on our pinned 2.37? Should we either gate the builder on is_qnn_sdk_version_less_than("2.48") (fall back to the old pattern below 2.48), or bump the minimum QNN version to 2.48 and move CI accordingly? Otherwise the test skip hides a compile failure for 2.37 users

Thanks for flagging that. I agree — I keep the legacy pattern for pcq embedding to maintain backward compatibility.

@psiddh

Copy link
Copy Markdown
Contributor

Thanks for all the follow up changes.. lgtm

@psiddh
psiddh merged commit 7c3f2ed into pytorch:mainJul 7, 2026
191 of 193 checks passed
psiddh added a commit that referenced this pull request Jul 21, 2026
Summary:
D110960687 (#20686) rewrote `Embedding.define_node` to
select between the optimized and legacy pcq-embedding lowerings based on
`is_qnn_sdk_version_less_than("2.48")`. That helper resolves the SDK
version via
`get_sdk_build_id`, which builds a path from
`os.environ["QNN_SDK_ROOT"]`.
`define_node` runs during AOT partitioning (`is_node_supported`), and
AOT-only
environments do not necessarily have `QNN_SDK_ROOT` set. In that case
`os.path.join(os.environ.get("QNN_SDK_ROOT", None), ...)` raised
`TypeError: expected str, bytes or os.PathLike object, not NoneType`,
breaking
every QNN lowering that contains an embedding — including the internal
`test_dummy_llama_qnn_16a4w_aot_and_runtime`.
Fall back to the legacy embedding lowering (valid on all QNN versions)
when
`QNN_SDK_ROOT` is unavailable, so the version-gated optimization is only
taken
when the SDK version can actually be determined. Also make
`get_sdk_build_id`
raise a clear `EnvironmentError` instead of a cryptic `TypeError` when
the
environment variable is missing.
This diff was authored with Claude Code.
Differential Revision: D112944232
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@shewu-quic@psiddh
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Qualcomm AI Engine Direct - Optimize performance of pcq embedding - #20686

Merged
psiddh merged 3 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_pcq_embedding
Jul 7, 2026
Merged

Qualcomm AI Engine Direct - Optimize performance of pcq embedding#20686
psiddh merged 3 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_pcq_embedding

Conversation

@shewu-quic

Copy link
Copy Markdown
Collaborator

Summary:

  • Change pcq embedding pattern for backend optimization
    • Note that it is supported after QNN 2.48

Test plan

python3 backends/qualcomm/tests/test_qnn_delegate.py TestQNNQuantizedOperator.test_qnn_backend_embedding_per_channel --build_folder build-android --host {HOST} --device {DEVICE} --soc_model SM8850 -a {ARTIFACTS} 

@pytorch-bot

pytorch-botBot commented Jul 2, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20686

Note: Links to docs will display an error until the docs builds have been completed.

❌ 2 New Failures

As of commit a551ca3 with merge base f41d66d (image):

NEW FAILURES - The following jobs have failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jul 2, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_pcq_embedding branch from 231b11e to 11bd40cCompareJuly 2, 2026 04:05
@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Hi @psiddh,
This PR is to optimize performance for PCQ embedding on HTP.
Could you please have a look?
Thanks,
Hutton

@psiddh

Copy link
Copy Markdown
Contributor

@claude Can you review this PR ?

@claude

claudeBot commented Jul 6, 2026

Copy link
Copy Markdown

Claude finished @psiddh's task in 3m 23s —— View job


PR Review: Optimize performance of PCQ embedding

  • Read the PR diff and changed files
  • Analyze op_embedding.py restructure
  • Analyze test and llm_wrappers.py changes
  • Post review

Nice cleanup — the new pattern is clearly better than the old one. Here's my review.

Summary of the change

The core change in op_embedding.py inverts the op order for per-channel-quantized (PCQ) embeddings:

  • Before:Gather(per-channel weight) → intermediate output where the gather output was forced to a per-tensor approximation (scale = max(scales), zp = max(zero_points)), then a Convert requantized to the activation encoding. That max-scale approximation is exactly what caused the accuracy problem the test was skipped for ("Bad accuracy for HTP").
  • After:Convert(per-channel weight) → per-tensor activation-encoded weight table → Gather → output. The per-channel→per-tensor requant now happens on the weight table itself (properly, per-channel), and gather just indexes it. This removes the lossy max(scales) heuristic entirely, which is why the test can now be enabled.

This is a solid structural improvement, and dropping the node.meta[QCOM_QUANT_ATTRS] mutate-and-restore dance (node_quant_attrs/intermediate_quant_attrs) removes a fragile side effect. Removing the now-unused QCOM_QUANT_MAX/MIN/SCALE(S)/ZERO_POINT(S) imports is good hygiene.

Correctness notes (looks right)

  • The final output wrapper is now registered under node.name (op_embedding.py:96-103), so downstream consumers resolve the correct tensor. The Convert output is an intermediate _convert native tensor feeding into Gather. ✅
  • Gather input (converted weight) and Gather output share the same activation encoding via get_quant_encoding_conf(node, node) — correct, since gather only indexes and must not change encoding. ✅
  • op_wrapper_list ordering is correct: convert_op is appended before gather_op, so the convert executes first. ✅

Questions / things worth confirming

1. Memory/perf of converting the full weight table (op_embedding.py:71-89)
The Convert now runs over the entire embedding table (dims=weight_tensor.size(), i.e. vocab × hidden) instead of over the small gather output. For LLM-sized vocabularies this is a large tensor. Since the convert input is a QNN_TENSOR_TYPE_STATIC weight, I assume the intent is that HTP constant-folds this convert at compile time so only the gather runs at inference — which would make this a genuine perf win. Can you confirm the convert-of-static is folded (and doesn't materialize a full-size runtime NATIVE tensor each inference)? A one-line note in the code or PR would help future readers understand why this is faster.

2. llm_wrappers.py getattr/setattr asymmetry (llm_wrappers.py:785-797)
The new _get_attr correctly traverses dotted targets (e.g. static_llama.tok_embedding.weight) for the read side. But the write side still uses a flat setattr(unquantized_model, unquantized_node.target, ...). If unquantized_node.target can also be dotted, setattr will create a single literal attribute named "a.b.c" on the model rather than setting model.a.b.c. Is unquantized_node.target guaranteed to be flat here, while only quantized_node.target is nested? If not, the write path needs the same traversal.

3. Version gate
is_qnn_sdk_version_less_than("2.48") only compares within the same major version (current_major == target_major and current_minor < target_minor). So on a hypothetical QNN 3.x the check returns False and the test runs — which is the desired behavior here, just flagging that the helper is major-scoped by design.

Minor

  • Test now covers both use_16a8w and use_16a4w with subTest — good added coverage. Consider adding a short comment noting the QNN 2.48 requirement mirrors the backend change (the skip message "UT pass after QNN 2.48" covers it, so optional).

Overall this looks like a clean, well-motivated change. My main ask is confirming point #1 (that the static convert is folded so this is actually faster, not just more accurate) and clarifying the setattr asymmetry in #2.

@psiddh

Copy link
Copy Markdown
Contributor

One qq: The test is skipped below QNN 2.48. does that mean the new pcq-embedding pattern only works on 2.48+,
or does it still work on 2.37 and 2.48 is just where the accuracy test passes?

attr = getattr(attr, target_atom)
return attr

setattr(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should setattr match the getattr walk pattern here?

@psiddh

Copy link
Copy Markdown
Contributor

Also just for my understanding : The convert now runs over the whole weight table instead of just the gathered rows. Since the weight is static, is this convert constant-folded during QNN compile (so it's computed once and the runtime
graph is just the gather)? Just want to confirm it's a one-time AOT cost and not a per-inference convert of the full table.

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

One qq: The test is skipped below QNN 2.48. does that mean the new pcq-embedding pattern only works on 2.48+, or does it still work on 2.37 and 2.48 is just where the accuracy test passes?

Yes, this new optimize pattern is only supported after QNN 2.48. Otherwise, it will failed to compile

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Also just for my understanding : The convert now runs over the whole weight table instead of just the gathered rows. Since the weight is static, is this convert constant-folded during QNN compile (so it's computed once and the runtime graph is just the gather)? Just want to confirm it's a one-time AOT cost and not a per-inference convert of the full table.

Yes, it will be fused by QNN HTP during AOT.

@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_pcq_embedding branch from f749d49 to f105816CompareJuly 6, 2026 14:28
@psiddh

Copy link
Copy Markdown
Contributor

One qq: The test is skipped below QNN 2.48. does that mean the new pcq-embedding pattern only works on 2.48+, or does it still work on 2.37 and 2.48 is just where the accuracy test passes?

Yes, this new optimize pattern is only supported after QNN 2.48. Otherwise, it will failed to compile

Thanks for confirming. Since it fails to compile below 2.48 and the builder emits this pattern unconditionally, won't pcq embedding now break for anyone on our pinned 2.37? Should we either gate the builder on is_qnn_sdk_version_less_than("2.48") (fall back to the old pattern below 2.48), or bump the minimum QNN version to 2.48 and move CI accordingly? Otherwise the test skip hides a compile failure for 2.37 users

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Thanks for confirming. Since it fails to compile below 2.48 and the builder emits this pattern unconditionally, won't pcq embedding now break for anyone on our pinned 2.37? Should we either gate the builder on is_qnn_sdk_version_less_than("2.48") (fall back to the old pattern below 2.48), or bump the minimum QNN version to 2.48 and move CI accordingly? Otherwise the test skip hides a compile failure for 2.37 users

Thanks for flagging that. I agree — I keep the legacy pattern for pcq embedding to maintain backward compatibility.

@psiddh

Copy link
Copy Markdown
Contributor

Thanks for all the follow up changes.. lgtm

@psiddh
psiddh merged commit 7c3f2ed into pytorch:mainJul 7, 2026
191 of 193 checks passed
psiddh added a commit that referenced this pull request Jul 21, 2026
Summary:
D110960687 (#20686) rewrote `Embedding.define_node` to
select between the optimized and legacy pcq-embedding lowerings based on
`is_qnn_sdk_version_less_than("2.48")`. That helper resolves the SDK
version via
`get_sdk_build_id`, which builds a path from
`os.environ["QNN_SDK_ROOT"]`.
`define_node` runs during AOT partitioning (`is_node_supported`), and
AOT-only
environments do not necessarily have `QNN_SDK_ROOT` set. In that case
`os.path.join(os.environ.get("QNN_SDK_ROOT", None), ...)` raised
`TypeError: expected str, bytes or os.PathLike object, not NoneType`,
breaking
every QNN lowering that contains an embedding — including the internal
`test_dummy_llama_qnn_16a4w_aot_and_runtime`.
Fall back to the legacy embedding lowering (valid on all QNN versions)
when
`QNN_SDK_ROOT` is unavailable, so the version-gated optimization is only
taken
when the SDK version can actually be determined. Also make
`get_sdk_build_id`
raise a clear `EnvironmentError` instead of a cryptic `TypeError` when
the
environment variable is missing.
This diff was authored with Claude Code.
Differential Revision: D112944232
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@shewu-quic@psiddh
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Qualcomm AI Engine Direct - Optimize performance of pcq embedding - #20686

Merged
psiddh merged 3 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_pcq_embedding
Jul 7, 2026
Merged

Qualcomm AI Engine Direct - Optimize performance of pcq embedding#20686
psiddh merged 3 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_pcq_embedding

Conversation

@shewu-quic

Copy link
Copy Markdown
Collaborator

Summary:

  • Change pcq embedding pattern for backend optimization
    • Note that it is supported after QNN 2.48

Test plan

python3 backends/qualcomm/tests/test_qnn_delegate.py TestQNNQuantizedOperator.test_qnn_backend_embedding_per_channel --build_folder build-android --host {HOST} --device {DEVICE} --soc_model SM8850 -a {ARTIFACTS} 

@pytorch-bot

pytorch-botBot commented Jul 2, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20686

Note: Links to docs will display an error until the docs builds have been completed.

❌ 2 New Failures

As of commit a551ca3 with merge base f41d66d (image):

NEW FAILURES - The following jobs have failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jul 2, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_pcq_embedding branch from 231b11e to 11bd40cCompareJuly 2, 2026 04:05
@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Hi @psiddh,
This PR is to optimize performance for PCQ embedding on HTP.
Could you please have a look?
Thanks,
Hutton

@psiddh

Copy link
Copy Markdown
Contributor

@claude Can you review this PR ?

@claude

claudeBot commented Jul 6, 2026

Copy link
Copy Markdown

Claude finished @psiddh's task in 3m 23s —— View job


PR Review: Optimize performance of PCQ embedding

  • Read the PR diff and changed files
  • Analyze op_embedding.py restructure
  • Analyze test and llm_wrappers.py changes
  • Post review

Nice cleanup — the new pattern is clearly better than the old one. Here's my review.

Summary of the change

The core change in op_embedding.py inverts the op order for per-channel-quantized (PCQ) embeddings:

  • Before:Gather(per-channel weight) → intermediate output where the gather output was forced to a per-tensor approximation (scale = max(scales), zp = max(zero_points)), then a Convert requantized to the activation encoding. That max-scale approximation is exactly what caused the accuracy problem the test was skipped for ("Bad accuracy for HTP").
  • After:Convert(per-channel weight) → per-tensor activation-encoded weight table → Gather → output. The per-channel→per-tensor requant now happens on the weight table itself (properly, per-channel), and gather just indexes it. This removes the lossy max(scales) heuristic entirely, which is why the test can now be enabled.

This is a solid structural improvement, and dropping the node.meta[QCOM_QUANT_ATTRS] mutate-and-restore dance (node_quant_attrs/intermediate_quant_attrs) removes a fragile side effect. Removing the now-unused QCOM_QUANT_MAX/MIN/SCALE(S)/ZERO_POINT(S) imports is good hygiene.

Correctness notes (looks right)

  • The final output wrapper is now registered under node.name (op_embedding.py:96-103), so downstream consumers resolve the correct tensor. The Convert output is an intermediate _convert native tensor feeding into Gather. ✅
  • Gather input (converted weight) and Gather output share the same activation encoding via get_quant_encoding_conf(node, node) — correct, since gather only indexes and must not change encoding. ✅
  • op_wrapper_list ordering is correct: convert_op is appended before gather_op, so the convert executes first. ✅

Questions / things worth confirming

1. Memory/perf of converting the full weight table (op_embedding.py:71-89)
The Convert now runs over the entire embedding table (dims=weight_tensor.size(), i.e. vocab × hidden) instead of over the small gather output. For LLM-sized vocabularies this is a large tensor. Since the convert input is a QNN_TENSOR_TYPE_STATIC weight, I assume the intent is that HTP constant-folds this convert at compile time so only the gather runs at inference — which would make this a genuine perf win. Can you confirm the convert-of-static is folded (and doesn't materialize a full-size runtime NATIVE tensor each inference)? A one-line note in the code or PR would help future readers understand why this is faster.

2. llm_wrappers.py getattr/setattr asymmetry (llm_wrappers.py:785-797)
The new _get_attr correctly traverses dotted targets (e.g. static_llama.tok_embedding.weight) for the read side. But the write side still uses a flat setattr(unquantized_model, unquantized_node.target, ...). If unquantized_node.target can also be dotted, setattr will create a single literal attribute named "a.b.c" on the model rather than setting model.a.b.c. Is unquantized_node.target guaranteed to be flat here, while only quantized_node.target is nested? If not, the write path needs the same traversal.

3. Version gate
is_qnn_sdk_version_less_than("2.48") only compares within the same major version (current_major == target_major and current_minor < target_minor). So on a hypothetical QNN 3.x the check returns False and the test runs — which is the desired behavior here, just flagging that the helper is major-scoped by design.

Minor

  • Test now covers both use_16a8w and use_16a4w with subTest — good added coverage. Consider adding a short comment noting the QNN 2.48 requirement mirrors the backend change (the skip message "UT pass after QNN 2.48" covers it, so optional).

Overall this looks like a clean, well-motivated change. My main ask is confirming point #1 (that the static convert is folded so this is actually faster, not just more accurate) and clarifying the setattr asymmetry in #2.

@psiddh

Copy link
Copy Markdown
Contributor

One qq: The test is skipped below QNN 2.48. does that mean the new pcq-embedding pattern only works on 2.48+,
or does it still work on 2.37 and 2.48 is just where the accuracy test passes?

attr = getattr(attr, target_atom)
return attr

setattr(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should setattr match the getattr walk pattern here?

@psiddh

Copy link
Copy Markdown
Contributor

Also just for my understanding : The convert now runs over the whole weight table instead of just the gathered rows. Since the weight is static, is this convert constant-folded during QNN compile (so it's computed once and the runtime
graph is just the gather)? Just want to confirm it's a one-time AOT cost and not a per-inference convert of the full table.

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

One qq: The test is skipped below QNN 2.48. does that mean the new pcq-embedding pattern only works on 2.48+, or does it still work on 2.37 and 2.48 is just where the accuracy test passes?

Yes, this new optimize pattern is only supported after QNN 2.48. Otherwise, it will failed to compile

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Also just for my understanding : The convert now runs over the whole weight table instead of just the gathered rows. Since the weight is static, is this convert constant-folded during QNN compile (so it's computed once and the runtime graph is just the gather)? Just want to confirm it's a one-time AOT cost and not a per-inference convert of the full table.

Yes, it will be fused by QNN HTP during AOT.

@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_pcq_embedding branch from f749d49 to f105816CompareJuly 6, 2026 14:28
@psiddh

Copy link
Copy Markdown
Contributor

One qq: The test is skipped below QNN 2.48. does that mean the new pcq-embedding pattern only works on 2.48+, or does it still work on 2.37 and 2.48 is just where the accuracy test passes?

Yes, this new optimize pattern is only supported after QNN 2.48. Otherwise, it will failed to compile

Thanks for confirming. Since it fails to compile below 2.48 and the builder emits this pattern unconditionally, won't pcq embedding now break for anyone on our pinned 2.37? Should we either gate the builder on is_qnn_sdk_version_less_than("2.48") (fall back to the old pattern below 2.48), or bump the minimum QNN version to 2.48 and move CI accordingly? Otherwise the test skip hides a compile failure for 2.37 users

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Thanks for confirming. Since it fails to compile below 2.48 and the builder emits this pattern unconditionally, won't pcq embedding now break for anyone on our pinned 2.37? Should we either gate the builder on is_qnn_sdk_version_less_than("2.48") (fall back to the old pattern below 2.48), or bump the minimum QNN version to 2.48 and move CI accordingly? Otherwise the test skip hides a compile failure for 2.37 users

Thanks for flagging that. I agree — I keep the legacy pattern for pcq embedding to maintain backward compatibility.

@psiddh

Copy link
Copy Markdown
Contributor

Thanks for all the follow up changes.. lgtm

@psiddh
psiddh merged commit 7c3f2ed into pytorch:mainJul 7, 2026
191 of 193 checks passed
psiddh added a commit that referenced this pull request Jul 21, 2026
Summary:
D110960687 (#20686) rewrote `Embedding.define_node` to
select between the optimized and legacy pcq-embedding lowerings based on
`is_qnn_sdk_version_less_than("2.48")`. That helper resolves the SDK
version via
`get_sdk_build_id`, which builds a path from
`os.environ["QNN_SDK_ROOT"]`.
`define_node` runs during AOT partitioning (`is_node_supported`), and
AOT-only
environments do not necessarily have `QNN_SDK_ROOT` set. In that case
`os.path.join(os.environ.get("QNN_SDK_ROOT", None), ...)` raised
`TypeError: expected str, bytes or os.PathLike object, not NoneType`,
breaking
every QNN lowering that contains an embedding — including the internal
`test_dummy_llama_qnn_16a4w_aot_and_runtime`.
Fall back to the legacy embedding lowering (valid on all QNN versions)
when
`QNN_SDK_ROOT` is unavailable, so the version-gated optimization is only
taken
when the SDK version can actually be determined. Also make
`get_sdk_build_id`
raise a clear `EnvironmentError` instead of a cryptic `TypeError` when
the
environment variable is missing.
This diff was authored with Claude Code.
Differential Revision: D112944232
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@shewu-quic@psiddh
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Qualcomm AI Engine Direct - Optimize performance of pcq embedding - #20686

Merged
psiddh merged 3 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_pcq_embedding
Jul 7, 2026
Merged

Qualcomm AI Engine Direct - Optimize performance of pcq embedding#20686
psiddh merged 3 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_pcq_embedding

Conversation

@shewu-quic

Copy link
Copy Markdown
Collaborator

Summary:

  • Change pcq embedding pattern for backend optimization
    • Note that it is supported after QNN 2.48

Test plan

python3 backends/qualcomm/tests/test_qnn_delegate.py TestQNNQuantizedOperator.test_qnn_backend_embedding_per_channel --build_folder build-android --host {HOST} --device {DEVICE} --soc_model SM8850 -a {ARTIFACTS} 

@pytorch-bot

pytorch-botBot commented Jul 2, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20686

Note: Links to docs will display an error until the docs builds have been completed.

❌ 2 New Failures

As of commit a551ca3 with merge base f41d66d (image):

NEW FAILURES - The following jobs have failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jul 2, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_pcq_embedding branch from 231b11e to 11bd40cCompareJuly 2, 2026 04:05
@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Hi @psiddh,
This PR is to optimize performance for PCQ embedding on HTP.
Could you please have a look?
Thanks,
Hutton

@psiddh

Copy link
Copy Markdown
Contributor

@claude Can you review this PR ?

@claude

claudeBot commented Jul 6, 2026

Copy link
Copy Markdown

Claude finished @psiddh's task in 3m 23s —— View job


PR Review: Optimize performance of PCQ embedding

  • Read the PR diff and changed files
  • Analyze op_embedding.py restructure
  • Analyze test and llm_wrappers.py changes
  • Post review

Nice cleanup — the new pattern is clearly better than the old one. Here's my review.

Summary of the change

The core change in op_embedding.py inverts the op order for per-channel-quantized (PCQ) embeddings:

  • Before:Gather(per-channel weight) → intermediate output where the gather output was forced to a per-tensor approximation (scale = max(scales), zp = max(zero_points)), then a Convert requantized to the activation encoding. That max-scale approximation is exactly what caused the accuracy problem the test was skipped for ("Bad accuracy for HTP").
  • After:Convert(per-channel weight) → per-tensor activation-encoded weight table → Gather → output. The per-channel→per-tensor requant now happens on the weight table itself (properly, per-channel), and gather just indexes it. This removes the lossy max(scales) heuristic entirely, which is why the test can now be enabled.

This is a solid structural improvement, and dropping the node.meta[QCOM_QUANT_ATTRS] mutate-and-restore dance (node_quant_attrs/intermediate_quant_attrs) removes a fragile side effect. Removing the now-unused QCOM_QUANT_MAX/MIN/SCALE(S)/ZERO_POINT(S) imports is good hygiene.

Correctness notes (looks right)

  • The final output wrapper is now registered under node.name (op_embedding.py:96-103), so downstream consumers resolve the correct tensor. The Convert output is an intermediate _convert native tensor feeding into Gather. ✅
  • Gather input (converted weight) and Gather output share the same activation encoding via get_quant_encoding_conf(node, node) — correct, since gather only indexes and must not change encoding. ✅
  • op_wrapper_list ordering is correct: convert_op is appended before gather_op, so the convert executes first. ✅

Questions / things worth confirming

1. Memory/perf of converting the full weight table (op_embedding.py:71-89)
The Convert now runs over the entire embedding table (dims=weight_tensor.size(), i.e. vocab × hidden) instead of over the small gather output. For LLM-sized vocabularies this is a large tensor. Since the convert input is a QNN_TENSOR_TYPE_STATIC weight, I assume the intent is that HTP constant-folds this convert at compile time so only the gather runs at inference — which would make this a genuine perf win. Can you confirm the convert-of-static is folded (and doesn't materialize a full-size runtime NATIVE tensor each inference)? A one-line note in the code or PR would help future readers understand why this is faster.

2. llm_wrappers.py getattr/setattr asymmetry (llm_wrappers.py:785-797)
The new _get_attr correctly traverses dotted targets (e.g. static_llama.tok_embedding.weight) for the read side. But the write side still uses a flat setattr(unquantized_model, unquantized_node.target, ...). If unquantized_node.target can also be dotted, setattr will create a single literal attribute named "a.b.c" on the model rather than setting model.a.b.c. Is unquantized_node.target guaranteed to be flat here, while only quantized_node.target is nested? If not, the write path needs the same traversal.

3. Version gate
is_qnn_sdk_version_less_than("2.48") only compares within the same major version (current_major == target_major and current_minor < target_minor). So on a hypothetical QNN 3.x the check returns False and the test runs — which is the desired behavior here, just flagging that the helper is major-scoped by design.

Minor

  • Test now covers both use_16a8w and use_16a4w with subTest — good added coverage. Consider adding a short comment noting the QNN 2.48 requirement mirrors the backend change (the skip message "UT pass after QNN 2.48" covers it, so optional).

Overall this looks like a clean, well-motivated change. My main ask is confirming point #1 (that the static convert is folded so this is actually faster, not just more accurate) and clarifying the setattr asymmetry in #2.

@psiddh

Copy link
Copy Markdown
Contributor

One qq: The test is skipped below QNN 2.48. does that mean the new pcq-embedding pattern only works on 2.48+,
or does it still work on 2.37 and 2.48 is just where the accuracy test passes?

attr = getattr(attr, target_atom)
return attr

setattr(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should setattr match the getattr walk pattern here?

@psiddh

Copy link
Copy Markdown
Contributor

Also just for my understanding : The convert now runs over the whole weight table instead of just the gathered rows. Since the weight is static, is this convert constant-folded during QNN compile (so it's computed once and the runtime
graph is just the gather)? Just want to confirm it's a one-time AOT cost and not a per-inference convert of the full table.

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

One qq: The test is skipped below QNN 2.48. does that mean the new pcq-embedding pattern only works on 2.48+, or does it still work on 2.37 and 2.48 is just where the accuracy test passes?

Yes, this new optimize pattern is only supported after QNN 2.48. Otherwise, it will failed to compile

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Also just for my understanding : The convert now runs over the whole weight table instead of just the gathered rows. Since the weight is static, is this convert constant-folded during QNN compile (so it's computed once and the runtime graph is just the gather)? Just want to confirm it's a one-time AOT cost and not a per-inference convert of the full table.

Yes, it will be fused by QNN HTP during AOT.

@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_pcq_embedding branch from f749d49 to f105816CompareJuly 6, 2026 14:28
@psiddh

Copy link
Copy Markdown
Contributor

One qq: The test is skipped below QNN 2.48. does that mean the new pcq-embedding pattern only works on 2.48+, or does it still work on 2.37 and 2.48 is just where the accuracy test passes?

Yes, this new optimize pattern is only supported after QNN 2.48. Otherwise, it will failed to compile

Thanks for confirming. Since it fails to compile below 2.48 and the builder emits this pattern unconditionally, won't pcq embedding now break for anyone on our pinned 2.37? Should we either gate the builder on is_qnn_sdk_version_less_than("2.48") (fall back to the old pattern below 2.48), or bump the minimum QNN version to 2.48 and move CI accordingly? Otherwise the test skip hides a compile failure for 2.37 users

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Thanks for confirming. Since it fails to compile below 2.48 and the builder emits this pattern unconditionally, won't pcq embedding now break for anyone on our pinned 2.37? Should we either gate the builder on is_qnn_sdk_version_less_than("2.48") (fall back to the old pattern below 2.48), or bump the minimum QNN version to 2.48 and move CI accordingly? Otherwise the test skip hides a compile failure for 2.37 users

Thanks for flagging that. I agree — I keep the legacy pattern for pcq embedding to maintain backward compatibility.

@psiddh

Copy link
Copy Markdown
Contributor

Thanks for all the follow up changes.. lgtm

@psiddh
psiddh merged commit 7c3f2ed into pytorch:mainJul 7, 2026
191 of 193 checks passed
psiddh added a commit that referenced this pull request Jul 21, 2026
Summary:
D110960687 (#20686) rewrote `Embedding.define_node` to
select between the optimized and legacy pcq-embedding lowerings based on
`is_qnn_sdk_version_less_than("2.48")`. That helper resolves the SDK
version via
`get_sdk_build_id`, which builds a path from
`os.environ["QNN_SDK_ROOT"]`.
`define_node` runs during AOT partitioning (`is_node_supported`), and
AOT-only
environments do not necessarily have `QNN_SDK_ROOT` set. In that case
`os.path.join(os.environ.get("QNN_SDK_ROOT", None), ...)` raised
`TypeError: expected str, bytes or os.PathLike object, not NoneType`,
breaking
every QNN lowering that contains an embedding — including the internal
`test_dummy_llama_qnn_16a4w_aot_and_runtime`.
Fall back to the legacy embedding lowering (valid on all QNN versions)
when
`QNN_SDK_ROOT` is unavailable, so the version-gated optimization is only
taken
when the SDK version can actually be determined. Also make
`get_sdk_build_id`
raise a clear `EnvironmentError` instead of a cryptic `TypeError` when
the
environment variable is missing.
This diff was authored with Claude Code.
Differential Revision: D112944232
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@shewu-quic@psiddh