Qualcomm AI Engine Direct - Optimize the performance for AR-N model - #9079

Merged
cccclai merged 4 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_arn_model
Mar 13, 2025
Merged

Qualcomm AI Engine Direct - Optimize the performance for AR-N model#9079
cccclai merged 4 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_arn_model

Conversation

@shewu-quic

@shewu-quicshewu-quic commented Mar 10, 2025

Copy link
Copy Markdown
Collaborator

Summary:

  • Fix the bug of rms norm builder
  • Use HuggingFace version RoPE to improve the performance due to stride = 1 in StrideSlice Op
  • Modificate the axis order of the conv in qkv, feedforward and output
    • Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
    • New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)

Test Result:

  • Verify the output for story llama with smart mask, CL=128, prefill_ar_n=16, prompt="Once"
    Note that using Hugging Face RoPE will slightly affect accuracy
    • Original (mainline)
INFO:root:Results[0]:
Once upon a time, there was a little girl named Lily. She loved to play with her toys and her favorite toy was a big, red ball. One day, Lily's mom asked her to help her with the laundry. Lily was happy to help and she put all the clothes in the washing machine. After the clothes were washed, Lily's mom asked her to help her hang them up to dry. Lily saw a big, black rake and asked her mom what it was. Her mom told her it was a rake and that it helps to
  • Optimized (this PR)
INFO:root:Results[0]:
Once upon a time, there was a little girl named Lily. She loved to play with her toys and her favorite toy was a big, red ball. One day, Lily's mom asked her to help her with the laundry. Lily was happy to help and she put all the clothes in the washing machine. After the clothes were washed, Lily's mom asked her to help her hang them up to dry. Lily saw a big, black iron on the counter and asked her mom what it was for. Her mom explained that it was used to make clothes smooth
  • Verify the performance for llama 3.2 1B with shift pointer, CL=2048, prefill_ar_n=256
    • Original (mainline)
I 00:00:02.048851 executorch:runner.cpp:354] Prompt Processor: total 256 tokens (AR-256 * 1 iters)
I 00:00:36.606984 executorch:runner.cpp:456] Prompt Tokens: 256 Generated Tokens: 1791
I 00:00:36.607049 executorch:runner.cpp:462] Model Load Time: 2.012000 (seconds)
I 00:00:36.607062 executorch:runner.cpp:472] Total inference time: 34.592000 (seconds) Rate: 51.774977 (tokens/second)
I 00:00:36.607072 executorch:runner.cpp:480] Prompt evaluation:	0.293000 (seconds) Rate: 873.720137 (tokens/second)
I 00:00:36.607080 executorch:runner.cpp:491] Generated 1791 tokens:	34.299000 (seconds) Rate: 52.217266 (tokens/second)
I 00:00:36.607089 executorch:runner.cpp:499] Time to first generated token:	0.293000 (seconds)
I 00:00:36.607099 executorch:runner.cpp:506] Sampling time over 1791 tokens:	1.473000 (seconds)
  • Optimized (this PR)
I 00:00:01.827440 executorch:runner.cpp:354] Prompt Processor: total 256 tokens (AR-256 * 1 iters)
I 00:00:03.143673 executorch:runner.cpp:456] Prompt Tokens: 256 Generated Tokens: 64
I 00:00:03.143686 executorch:runner.cpp:462] Model Load Time: 1.791000 (seconds)
I 00:00:03.143698 executorch:runner.cpp:472] Total inference time: 1.350000 (seconds) Rate: 47.407407 (tokens/second)
I 00:00:03.143706 executorch:runner.cpp:480] Prompt evaluation:	0.126000 (seconds) Rate: 2031.746032 (tokens/second)
I 00:00:03.143715 executorch:runner.cpp:491] Generated 64 tokens:	1.224000 (seconds) Rate: 52.287582 (tokens/second)
I 00:00:03.143723 executorch:runner.cpp:499] Time to first generated token:	0.126000 (seconds)
I 00:00:03.143733 executorch:runner.cpp:506] Sampling time over 64 tokens:	0.058000 (seconds)

@pytorch-bot

pytorch-botBot commented Mar 10, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/9079

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Unrelated Failure

As of commit c94c0bd with merge base acae017 (image):

NEW FAILURE - The following job has failed:

BROKEN TRUNK - The following job failed but were present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 10, 2025
@shewu-quic

shewu-quic commented Mar 10, 2025

Copy link
Copy Markdown
CollaboratorAuthor

Hi @cccclai,
This PR aims to enhance the performance of the prompt processor (AR-N model).
The main improvement comes from using a stride slice operation with a stride of 1, which offers better performance. Consequently, we’ve switched to the Huggingface version of RoPE.
Could you help to take a look?

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.


def __init__(self):
super().__init__()
def __init__(self, edge_program: torch.export.ExportedProgram):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can follow #8505 to get rid of some recompose logic to reduce engineer effort there

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for your information. I will try it.

) -> torch.Tensor:
x_r, x_i = x[..., ::2], x[..., 1::2]

# Change to RoPE of huggingface version

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Which one is the huggingface version and why is it better?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The implementation of RoPE in huggingface process query and key with two half instead of interleaved way.
The main difference is stride in StrideSlice op. For interleaved way, stride is two which is not friendly for HTP backend about this memory handle.
Ref: huggingface/transformers#25199

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe add this comment to part of the code comment, just so others know the context.

@cccclaicccclai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The perf improvement looks awesome!!

@cccclai

Copy link
Copy Markdown
Contributor

There is still lint error, can you fix it?

@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_arn_model branch from b4d5a63 to fde0f80CompareMarch 12, 2025 03:51
@shewu-quic
shewu-quic requested a review from SS-JIA as a code ownerMarch 12, 2025 03:51
@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

There is still lint error, can you fix it?

Done. Thanks :)

@cccclaicccclai added the release notes: qualcomm Changes to the Qualcomm backend delegate label Mar 12, 2025
@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai

Copy link
Copy Markdown
Contributor

This seems need rebase

 Summary:
- Fix the bug of rms norm builder
- Use HuggingFace version RoPE to improve the performance due to
stride = 1 in StrideSlice Op
- Modificate the axis order of the conv in qkv, feedforward and output
- Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
- New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)
@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_arn_model branch from fde0f80 to c5c149cCompareMarch 12, 2025 23:12
@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai

Copy link
Copy Markdown
Contributor

Can you also share which part of the logic do the following optmization?

Modificate the axis order of the conv in qkv, feedforward and output
Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)

Is it the weight permutation or something else? Also good to have it as part of the code comment so it's easy to understand the intention.

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Can you also share which part of the logic do the following optmization?

Modificate the axis order of the conv in qkv, feedforward and output
Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)

Is it the weight permutation or something else? Also good to have it as part of the code comment so it's easy to understand the intention.

Got it. But this optimization is not quite general. Based on our experiments, the performance which sets sequence length to width dimension (1, 1, seq_len, CL) is better than the performance which sets sequence length to height dimension (1, seq_len, 1, CL) for the input axis order of the conv op. And another reason is that this change will be close with the structure of AI Hub version llama

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai
cccclai merged commit baf35d2 into pytorch:mainMar 13, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.release notes: qualcommChanges to the Qualcomm backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@shewu-quic@facebook-github-bot@cccclai
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Qualcomm AI Engine Direct - Optimize the performance for AR-N model - #9079

Merged
cccclai merged 4 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_arn_model
Mar 13, 2025
Merged

Qualcomm AI Engine Direct - Optimize the performance for AR-N model#9079
cccclai merged 4 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_arn_model

Conversation

@shewu-quic

@shewu-quicshewu-quic commented Mar 10, 2025

Copy link
Copy Markdown
Collaborator

Summary:

  • Fix the bug of rms norm builder
  • Use HuggingFace version RoPE to improve the performance due to stride = 1 in StrideSlice Op
  • Modificate the axis order of the conv in qkv, feedforward and output
    • Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
    • New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)

Test Result:

  • Verify the output for story llama with smart mask, CL=128, prefill_ar_n=16, prompt="Once"
    Note that using Hugging Face RoPE will slightly affect accuracy
    • Original (mainline)
INFO:root:Results[0]:
Once upon a time, there was a little girl named Lily. She loved to play with her toys and her favorite toy was a big, red ball. One day, Lily's mom asked her to help her with the laundry. Lily was happy to help and she put all the clothes in the washing machine. After the clothes were washed, Lily's mom asked her to help her hang them up to dry. Lily saw a big, black rake and asked her mom what it was. Her mom told her it was a rake and that it helps to
  • Optimized (this PR)
INFO:root:Results[0]:
Once upon a time, there was a little girl named Lily. She loved to play with her toys and her favorite toy was a big, red ball. One day, Lily's mom asked her to help her with the laundry. Lily was happy to help and she put all the clothes in the washing machine. After the clothes were washed, Lily's mom asked her to help her hang them up to dry. Lily saw a big, black iron on the counter and asked her mom what it was for. Her mom explained that it was used to make clothes smooth
  • Verify the performance for llama 3.2 1B with shift pointer, CL=2048, prefill_ar_n=256
    • Original (mainline)
I 00:00:02.048851 executorch:runner.cpp:354] Prompt Processor: total 256 tokens (AR-256 * 1 iters)
I 00:00:36.606984 executorch:runner.cpp:456] Prompt Tokens: 256 Generated Tokens: 1791
I 00:00:36.607049 executorch:runner.cpp:462] Model Load Time: 2.012000 (seconds)
I 00:00:36.607062 executorch:runner.cpp:472] Total inference time: 34.592000 (seconds) Rate: 51.774977 (tokens/second)
I 00:00:36.607072 executorch:runner.cpp:480] Prompt evaluation:	0.293000 (seconds) Rate: 873.720137 (tokens/second)
I 00:00:36.607080 executorch:runner.cpp:491] Generated 1791 tokens:	34.299000 (seconds) Rate: 52.217266 (tokens/second)
I 00:00:36.607089 executorch:runner.cpp:499] Time to first generated token:	0.293000 (seconds)
I 00:00:36.607099 executorch:runner.cpp:506] Sampling time over 1791 tokens:	1.473000 (seconds)
  • Optimized (this PR)
I 00:00:01.827440 executorch:runner.cpp:354] Prompt Processor: total 256 tokens (AR-256 * 1 iters)
I 00:00:03.143673 executorch:runner.cpp:456] Prompt Tokens: 256 Generated Tokens: 64
I 00:00:03.143686 executorch:runner.cpp:462] Model Load Time: 1.791000 (seconds)
I 00:00:03.143698 executorch:runner.cpp:472] Total inference time: 1.350000 (seconds) Rate: 47.407407 (tokens/second)
I 00:00:03.143706 executorch:runner.cpp:480] Prompt evaluation:	0.126000 (seconds) Rate: 2031.746032 (tokens/second)
I 00:00:03.143715 executorch:runner.cpp:491] Generated 64 tokens:	1.224000 (seconds) Rate: 52.287582 (tokens/second)
I 00:00:03.143723 executorch:runner.cpp:499] Time to first generated token:	0.126000 (seconds)
I 00:00:03.143733 executorch:runner.cpp:506] Sampling time over 64 tokens:	0.058000 (seconds)

@pytorch-bot

pytorch-botBot commented Mar 10, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/9079

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Unrelated Failure

As of commit c94c0bd with merge base acae017 (image):

NEW FAILURE - The following job has failed:

BROKEN TRUNK - The following job failed but were present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 10, 2025
@shewu-quic

shewu-quic commented Mar 10, 2025

Copy link
Copy Markdown
CollaboratorAuthor

Hi @cccclai,
This PR aims to enhance the performance of the prompt processor (AR-N model).
The main improvement comes from using a stride slice operation with a stride of 1, which offers better performance. Consequently, we’ve switched to the Huggingface version of RoPE.
Could you help to take a look?

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.


def __init__(self):
super().__init__()
def __init__(self, edge_program: torch.export.ExportedProgram):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can follow #8505 to get rid of some recompose logic to reduce engineer effort there

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for your information. I will try it.

) -> torch.Tensor:
x_r, x_i = x[..., ::2], x[..., 1::2]

# Change to RoPE of huggingface version

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Which one is the huggingface version and why is it better?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The implementation of RoPE in huggingface process query and key with two half instead of interleaved way.
The main difference is stride in StrideSlice op. For interleaved way, stride is two which is not friendly for HTP backend about this memory handle.
Ref: huggingface/transformers#25199

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe add this comment to part of the code comment, just so others know the context.

@cccclaicccclai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The perf improvement looks awesome!!

@cccclai

Copy link
Copy Markdown
Contributor

There is still lint error, can you fix it?

@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_arn_model branch from b4d5a63 to fde0f80CompareMarch 12, 2025 03:51
@shewu-quic
shewu-quic requested a review from SS-JIA as a code ownerMarch 12, 2025 03:51
@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

There is still lint error, can you fix it?

Done. Thanks :)

@cccclaicccclai added the release notes: qualcomm Changes to the Qualcomm backend delegate label Mar 12, 2025
@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai

Copy link
Copy Markdown
Contributor

This seems need rebase

 Summary:
- Fix the bug of rms norm builder
- Use HuggingFace version RoPE to improve the performance due to
stride = 1 in StrideSlice Op
- Modificate the axis order of the conv in qkv, feedforward and output
- Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
- New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)
@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_arn_model branch from fde0f80 to c5c149cCompareMarch 12, 2025 23:12
@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai

Copy link
Copy Markdown
Contributor

Can you also share which part of the logic do the following optmization?

Modificate the axis order of the conv in qkv, feedforward and output
Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)

Is it the weight permutation or something else? Also good to have it as part of the code comment so it's easy to understand the intention.

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Can you also share which part of the logic do the following optmization?

Modificate the axis order of the conv in qkv, feedforward and output
Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)

Is it the weight permutation or something else? Also good to have it as part of the code comment so it's easy to understand the intention.

Got it. But this optimization is not quite general. Based on our experiments, the performance which sets sequence length to width dimension (1, 1, seq_len, CL) is better than the performance which sets sequence length to height dimension (1, seq_len, 1, CL) for the input axis order of the conv op. And another reason is that this change will be close with the structure of AI Hub version llama

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai
cccclai merged commit baf35d2 into pytorch:mainMar 13, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.release notes: qualcommChanges to the Qualcomm backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@shewu-quic@facebook-github-bot@cccclai
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Qualcomm AI Engine Direct - Optimize the performance for AR-N model - #9079

Merged
cccclai merged 4 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_arn_model
Mar 13, 2025
Merged

Qualcomm AI Engine Direct - Optimize the performance for AR-N model#9079
cccclai merged 4 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_arn_model

Conversation

@shewu-quic

@shewu-quicshewu-quic commented Mar 10, 2025

Copy link
Copy Markdown
Collaborator

Summary:

  • Fix the bug of rms norm builder
  • Use HuggingFace version RoPE to improve the performance due to stride = 1 in StrideSlice Op
  • Modificate the axis order of the conv in qkv, feedforward and output
    • Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
    • New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)

Test Result:

  • Verify the output for story llama with smart mask, CL=128, prefill_ar_n=16, prompt="Once"
    Note that using Hugging Face RoPE will slightly affect accuracy
    • Original (mainline)
INFO:root:Results[0]:
Once upon a time, there was a little girl named Lily. She loved to play with her toys and her favorite toy was a big, red ball. One day, Lily's mom asked her to help her with the laundry. Lily was happy to help and she put all the clothes in the washing machine. After the clothes were washed, Lily's mom asked her to help her hang them up to dry. Lily saw a big, black rake and asked her mom what it was. Her mom told her it was a rake and that it helps to
  • Optimized (this PR)
INFO:root:Results[0]:
Once upon a time, there was a little girl named Lily. She loved to play with her toys and her favorite toy was a big, red ball. One day, Lily's mom asked her to help her with the laundry. Lily was happy to help and she put all the clothes in the washing machine. After the clothes were washed, Lily's mom asked her to help her hang them up to dry. Lily saw a big, black iron on the counter and asked her mom what it was for. Her mom explained that it was used to make clothes smooth
  • Verify the performance for llama 3.2 1B with shift pointer, CL=2048, prefill_ar_n=256
    • Original (mainline)
I 00:00:02.048851 executorch:runner.cpp:354] Prompt Processor: total 256 tokens (AR-256 * 1 iters)
I 00:00:36.606984 executorch:runner.cpp:456] Prompt Tokens: 256 Generated Tokens: 1791
I 00:00:36.607049 executorch:runner.cpp:462] Model Load Time: 2.012000 (seconds)
I 00:00:36.607062 executorch:runner.cpp:472] Total inference time: 34.592000 (seconds) Rate: 51.774977 (tokens/second)
I 00:00:36.607072 executorch:runner.cpp:480] Prompt evaluation:	0.293000 (seconds) Rate: 873.720137 (tokens/second)
I 00:00:36.607080 executorch:runner.cpp:491] Generated 1791 tokens:	34.299000 (seconds) Rate: 52.217266 (tokens/second)
I 00:00:36.607089 executorch:runner.cpp:499] Time to first generated token:	0.293000 (seconds)
I 00:00:36.607099 executorch:runner.cpp:506] Sampling time over 1791 tokens:	1.473000 (seconds)
  • Optimized (this PR)
I 00:00:01.827440 executorch:runner.cpp:354] Prompt Processor: total 256 tokens (AR-256 * 1 iters)
I 00:00:03.143673 executorch:runner.cpp:456] Prompt Tokens: 256 Generated Tokens: 64
I 00:00:03.143686 executorch:runner.cpp:462] Model Load Time: 1.791000 (seconds)
I 00:00:03.143698 executorch:runner.cpp:472] Total inference time: 1.350000 (seconds) Rate: 47.407407 (tokens/second)
I 00:00:03.143706 executorch:runner.cpp:480] Prompt evaluation:	0.126000 (seconds) Rate: 2031.746032 (tokens/second)
I 00:00:03.143715 executorch:runner.cpp:491] Generated 64 tokens:	1.224000 (seconds) Rate: 52.287582 (tokens/second)
I 00:00:03.143723 executorch:runner.cpp:499] Time to first generated token:	0.126000 (seconds)
I 00:00:03.143733 executorch:runner.cpp:506] Sampling time over 64 tokens:	0.058000 (seconds)

@pytorch-bot

pytorch-botBot commented Mar 10, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/9079

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Unrelated Failure

As of commit c94c0bd with merge base acae017 (image):

NEW FAILURE - The following job has failed:

BROKEN TRUNK - The following job failed but were present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 10, 2025
@shewu-quic

shewu-quic commented Mar 10, 2025

Copy link
Copy Markdown
CollaboratorAuthor

Hi @cccclai,
This PR aims to enhance the performance of the prompt processor (AR-N model).
The main improvement comes from using a stride slice operation with a stride of 1, which offers better performance. Consequently, we’ve switched to the Huggingface version of RoPE.
Could you help to take a look?

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.


def __init__(self):
super().__init__()
def __init__(self, edge_program: torch.export.ExportedProgram):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can follow #8505 to get rid of some recompose logic to reduce engineer effort there

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for your information. I will try it.

) -> torch.Tensor:
x_r, x_i = x[..., ::2], x[..., 1::2]

# Change to RoPE of huggingface version

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Which one is the huggingface version and why is it better?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The implementation of RoPE in huggingface process query and key with two half instead of interleaved way.
The main difference is stride in StrideSlice op. For interleaved way, stride is two which is not friendly for HTP backend about this memory handle.
Ref: huggingface/transformers#25199

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe add this comment to part of the code comment, just so others know the context.

@cccclaicccclai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The perf improvement looks awesome!!

@cccclai

Copy link
Copy Markdown
Contributor

There is still lint error, can you fix it?

@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_arn_model branch from b4d5a63 to fde0f80CompareMarch 12, 2025 03:51
@shewu-quic
shewu-quic requested a review from SS-JIA as a code ownerMarch 12, 2025 03:51
@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

There is still lint error, can you fix it?

Done. Thanks :)

@cccclaicccclai added the release notes: qualcomm Changes to the Qualcomm backend delegate label Mar 12, 2025
@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai

Copy link
Copy Markdown
Contributor

This seems need rebase

 Summary:
- Fix the bug of rms norm builder
- Use HuggingFace version RoPE to improve the performance due to
stride = 1 in StrideSlice Op
- Modificate the axis order of the conv in qkv, feedforward and output
- Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
- New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)
@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_arn_model branch from fde0f80 to c5c149cCompareMarch 12, 2025 23:12
@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai

Copy link
Copy Markdown
Contributor

Can you also share which part of the logic do the following optmization?

Modificate the axis order of the conv in qkv, feedforward and output
Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)

Is it the weight permutation or something else? Also good to have it as part of the code comment so it's easy to understand the intention.

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Can you also share which part of the logic do the following optmization?

Modificate the axis order of the conv in qkv, feedforward and output
Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)

Is it the weight permutation or something else? Also good to have it as part of the code comment so it's easy to understand the intention.

Got it. But this optimization is not quite general. Based on our experiments, the performance which sets sequence length to width dimension (1, 1, seq_len, CL) is better than the performance which sets sequence length to height dimension (1, seq_len, 1, CL) for the input axis order of the conv op. And another reason is that this change will be close with the structure of AI Hub version llama

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai
cccclai merged commit baf35d2 into pytorch:mainMar 13, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.release notes: qualcommChanges to the Qualcomm backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@shewu-quic@facebook-github-bot@cccclai
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Qualcomm AI Engine Direct - Optimize the performance for AR-N model - #9079

Merged
cccclai merged 4 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_arn_model
Mar 13, 2025
Merged

Qualcomm AI Engine Direct - Optimize the performance for AR-N model#9079
cccclai merged 4 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_arn_model

Conversation

@shewu-quic

@shewu-quicshewu-quic commented Mar 10, 2025

Copy link
Copy Markdown
Collaborator

Summary:

  • Fix the bug of rms norm builder
  • Use HuggingFace version RoPE to improve the performance due to stride = 1 in StrideSlice Op
  • Modificate the axis order of the conv in qkv, feedforward and output
    • Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
    • New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)

Test Result:

  • Verify the output for story llama with smart mask, CL=128, prefill_ar_n=16, prompt="Once"
    Note that using Hugging Face RoPE will slightly affect accuracy
    • Original (mainline)
INFO:root:Results[0]:
Once upon a time, there was a little girl named Lily. She loved to play with her toys and her favorite toy was a big, red ball. One day, Lily's mom asked her to help her with the laundry. Lily was happy to help and she put all the clothes in the washing machine. After the clothes were washed, Lily's mom asked her to help her hang them up to dry. Lily saw a big, black rake and asked her mom what it was. Her mom told her it was a rake and that it helps to
  • Optimized (this PR)
INFO:root:Results[0]:
Once upon a time, there was a little girl named Lily. She loved to play with her toys and her favorite toy was a big, red ball. One day, Lily's mom asked her to help her with the laundry. Lily was happy to help and she put all the clothes in the washing machine. After the clothes were washed, Lily's mom asked her to help her hang them up to dry. Lily saw a big, black iron on the counter and asked her mom what it was for. Her mom explained that it was used to make clothes smooth
  • Verify the performance for llama 3.2 1B with shift pointer, CL=2048, prefill_ar_n=256
    • Original (mainline)
I 00:00:02.048851 executorch:runner.cpp:354] Prompt Processor: total 256 tokens (AR-256 * 1 iters)
I 00:00:36.606984 executorch:runner.cpp:456] Prompt Tokens: 256 Generated Tokens: 1791
I 00:00:36.607049 executorch:runner.cpp:462] Model Load Time: 2.012000 (seconds)
I 00:00:36.607062 executorch:runner.cpp:472] Total inference time: 34.592000 (seconds) Rate: 51.774977 (tokens/second)
I 00:00:36.607072 executorch:runner.cpp:480] Prompt evaluation:	0.293000 (seconds) Rate: 873.720137 (tokens/second)
I 00:00:36.607080 executorch:runner.cpp:491] Generated 1791 tokens:	34.299000 (seconds) Rate: 52.217266 (tokens/second)
I 00:00:36.607089 executorch:runner.cpp:499] Time to first generated token:	0.293000 (seconds)
I 00:00:36.607099 executorch:runner.cpp:506] Sampling time over 1791 tokens:	1.473000 (seconds)
  • Optimized (this PR)
I 00:00:01.827440 executorch:runner.cpp:354] Prompt Processor: total 256 tokens (AR-256 * 1 iters)
I 00:00:03.143673 executorch:runner.cpp:456] Prompt Tokens: 256 Generated Tokens: 64
I 00:00:03.143686 executorch:runner.cpp:462] Model Load Time: 1.791000 (seconds)
I 00:00:03.143698 executorch:runner.cpp:472] Total inference time: 1.350000 (seconds) Rate: 47.407407 (tokens/second)
I 00:00:03.143706 executorch:runner.cpp:480] Prompt evaluation:	0.126000 (seconds) Rate: 2031.746032 (tokens/second)
I 00:00:03.143715 executorch:runner.cpp:491] Generated 64 tokens:	1.224000 (seconds) Rate: 52.287582 (tokens/second)
I 00:00:03.143723 executorch:runner.cpp:499] Time to first generated token:	0.126000 (seconds)
I 00:00:03.143733 executorch:runner.cpp:506] Sampling time over 64 tokens:	0.058000 (seconds)

@pytorch-bot

pytorch-botBot commented Mar 10, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/9079

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Unrelated Failure

As of commit c94c0bd with merge base acae017 (image):

NEW FAILURE - The following job has failed:

BROKEN TRUNK - The following job failed but were present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 10, 2025
@shewu-quic

shewu-quic commented Mar 10, 2025

Copy link
Copy Markdown
CollaboratorAuthor

Hi @cccclai,
This PR aims to enhance the performance of the prompt processor (AR-N model).
The main improvement comes from using a stride slice operation with a stride of 1, which offers better performance. Consequently, we’ve switched to the Huggingface version of RoPE.
Could you help to take a look?

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.


def __init__(self):
super().__init__()
def __init__(self, edge_program: torch.export.ExportedProgram):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can follow #8505 to get rid of some recompose logic to reduce engineer effort there

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for your information. I will try it.

) -> torch.Tensor:
x_r, x_i = x[..., ::2], x[..., 1::2]

# Change to RoPE of huggingface version

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Which one is the huggingface version and why is it better?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The implementation of RoPE in huggingface process query and key with two half instead of interleaved way.
The main difference is stride in StrideSlice op. For interleaved way, stride is two which is not friendly for HTP backend about this memory handle.
Ref: huggingface/transformers#25199

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe add this comment to part of the code comment, just so others know the context.

@cccclaicccclai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The perf improvement looks awesome!!

@cccclai

Copy link
Copy Markdown
Contributor

There is still lint error, can you fix it?

@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_arn_model branch from b4d5a63 to fde0f80CompareMarch 12, 2025 03:51
@shewu-quic
shewu-quic requested a review from SS-JIA as a code ownerMarch 12, 2025 03:51
@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

There is still lint error, can you fix it?

Done. Thanks :)

@cccclaicccclai added the release notes: qualcomm Changes to the Qualcomm backend delegate label Mar 12, 2025
@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai

Copy link
Copy Markdown
Contributor

This seems need rebase

 Summary:
- Fix the bug of rms norm builder
- Use HuggingFace version RoPE to improve the performance due to
stride = 1 in StrideSlice Op
- Modificate the axis order of the conv in qkv, feedforward and output
- Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
- New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)
@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_arn_model branch from fde0f80 to c5c149cCompareMarch 12, 2025 23:12
@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai

Copy link
Copy Markdown
Contributor

Can you also share which part of the logic do the following optmization?

Modificate the axis order of the conv in qkv, feedforward and output
Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)

Is it the weight permutation or something else? Also good to have it as part of the code comment so it's easy to understand the intention.

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Can you also share which part of the logic do the following optmization?

Modificate the axis order of the conv in qkv, feedforward and output
Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)

Is it the weight permutation or something else? Also good to have it as part of the code comment so it's easy to understand the intention.

Got it. But this optimization is not quite general. Based on our experiments, the performance which sets sequence length to width dimension (1, 1, seq_len, CL) is better than the performance which sets sequence length to height dimension (1, seq_len, 1, CL) for the input axis order of the conv op. And another reason is that this change will be close with the structure of AI Hub version llama

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai
cccclai merged commit baf35d2 into pytorch:mainMar 13, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.release notes: qualcommChanges to the Qualcomm backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@shewu-quic@facebook-github-bot@cccclai
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Qualcomm AI Engine Direct - Optimize the performance for AR-N model - #9079

Merged
cccclai merged 4 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_arn_model
Mar 13, 2025
Merged

Qualcomm AI Engine Direct - Optimize the performance for AR-N model#9079
cccclai merged 4 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_arn_model

Conversation

@shewu-quic

@shewu-quicshewu-quic commented Mar 10, 2025

Copy link
Copy Markdown
Collaborator

Summary:

  • Fix the bug of rms norm builder
  • Use HuggingFace version RoPE to improve the performance due to stride = 1 in StrideSlice Op
  • Modificate the axis order of the conv in qkv, feedforward and output
    • Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
    • New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)

Test Result:

  • Verify the output for story llama with smart mask, CL=128, prefill_ar_n=16, prompt="Once"
    Note that using Hugging Face RoPE will slightly affect accuracy
    • Original (mainline)
INFO:root:Results[0]:
Once upon a time, there was a little girl named Lily. She loved to play with her toys and her favorite toy was a big, red ball. One day, Lily's mom asked her to help her with the laundry. Lily was happy to help and she put all the clothes in the washing machine. After the clothes were washed, Lily's mom asked her to help her hang them up to dry. Lily saw a big, black rake and asked her mom what it was. Her mom told her it was a rake and that it helps to
  • Optimized (this PR)
INFO:root:Results[0]:
Once upon a time, there was a little girl named Lily. She loved to play with her toys and her favorite toy was a big, red ball. One day, Lily's mom asked her to help her with the laundry. Lily was happy to help and she put all the clothes in the washing machine. After the clothes were washed, Lily's mom asked her to help her hang them up to dry. Lily saw a big, black iron on the counter and asked her mom what it was for. Her mom explained that it was used to make clothes smooth
  • Verify the performance for llama 3.2 1B with shift pointer, CL=2048, prefill_ar_n=256
    • Original (mainline)
I 00:00:02.048851 executorch:runner.cpp:354] Prompt Processor: total 256 tokens (AR-256 * 1 iters)
I 00:00:36.606984 executorch:runner.cpp:456] Prompt Tokens: 256 Generated Tokens: 1791
I 00:00:36.607049 executorch:runner.cpp:462] Model Load Time: 2.012000 (seconds)
I 00:00:36.607062 executorch:runner.cpp:472] Total inference time: 34.592000 (seconds) Rate: 51.774977 (tokens/second)
I 00:00:36.607072 executorch:runner.cpp:480] Prompt evaluation:	0.293000 (seconds) Rate: 873.720137 (tokens/second)
I 00:00:36.607080 executorch:runner.cpp:491] Generated 1791 tokens:	34.299000 (seconds) Rate: 52.217266 (tokens/second)
I 00:00:36.607089 executorch:runner.cpp:499] Time to first generated token:	0.293000 (seconds)
I 00:00:36.607099 executorch:runner.cpp:506] Sampling time over 1791 tokens:	1.473000 (seconds)
  • Optimized (this PR)
I 00:00:01.827440 executorch:runner.cpp:354] Prompt Processor: total 256 tokens (AR-256 * 1 iters)
I 00:00:03.143673 executorch:runner.cpp:456] Prompt Tokens: 256 Generated Tokens: 64
I 00:00:03.143686 executorch:runner.cpp:462] Model Load Time: 1.791000 (seconds)
I 00:00:03.143698 executorch:runner.cpp:472] Total inference time: 1.350000 (seconds) Rate: 47.407407 (tokens/second)
I 00:00:03.143706 executorch:runner.cpp:480] Prompt evaluation:	0.126000 (seconds) Rate: 2031.746032 (tokens/second)
I 00:00:03.143715 executorch:runner.cpp:491] Generated 64 tokens:	1.224000 (seconds) Rate: 52.287582 (tokens/second)
I 00:00:03.143723 executorch:runner.cpp:499] Time to first generated token:	0.126000 (seconds)
I 00:00:03.143733 executorch:runner.cpp:506] Sampling time over 64 tokens:	0.058000 (seconds)

@pytorch-bot

pytorch-botBot commented Mar 10, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/9079

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Unrelated Failure

As of commit c94c0bd with merge base acae017 (image):

NEW FAILURE - The following job has failed:

BROKEN TRUNK - The following job failed but were present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 10, 2025
@shewu-quic

shewu-quic commented Mar 10, 2025

Copy link
Copy Markdown
CollaboratorAuthor

Hi @cccclai,
This PR aims to enhance the performance of the prompt processor (AR-N model).
The main improvement comes from using a stride slice operation with a stride of 1, which offers better performance. Consequently, we’ve switched to the Huggingface version of RoPE.
Could you help to take a look?

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.


def __init__(self):
super().__init__()
def __init__(self, edge_program: torch.export.ExportedProgram):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can follow #8505 to get rid of some recompose logic to reduce engineer effort there

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for your information. I will try it.

) -> torch.Tensor:
x_r, x_i = x[..., ::2], x[..., 1::2]

# Change to RoPE of huggingface version

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Which one is the huggingface version and why is it better?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The implementation of RoPE in huggingface process query and key with two half instead of interleaved way.
The main difference is stride in StrideSlice op. For interleaved way, stride is two which is not friendly for HTP backend about this memory handle.
Ref: huggingface/transformers#25199

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe add this comment to part of the code comment, just so others know the context.

@cccclaicccclai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The perf improvement looks awesome!!

@cccclai

Copy link
Copy Markdown
Contributor

There is still lint error, can you fix it?

@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_arn_model branch from b4d5a63 to fde0f80CompareMarch 12, 2025 03:51
@shewu-quic
shewu-quic requested a review from SS-JIA as a code ownerMarch 12, 2025 03:51
@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

There is still lint error, can you fix it?

Done. Thanks :)

@cccclaicccclai added the release notes: qualcomm Changes to the Qualcomm backend delegate label Mar 12, 2025
@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai

Copy link
Copy Markdown
Contributor

This seems need rebase

 Summary:
- Fix the bug of rms norm builder
- Use HuggingFace version RoPE to improve the performance due to
stride = 1 in StrideSlice Op
- Modificate the axis order of the conv in qkv, feedforward and output
- Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
- New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)
@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_arn_model branch from fde0f80 to c5c149cCompareMarch 12, 2025 23:12
@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai

Copy link
Copy Markdown
Contributor

Can you also share which part of the logic do the following optmization?

Modificate the axis order of the conv in qkv, feedforward and output
Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)

Is it the weight permutation or something else? Also good to have it as part of the code comment so it's easy to understand the intention.

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Can you also share which part of the logic do the following optmization?

Modificate the axis order of the conv in qkv, feedforward and output
Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)

Is it the weight permutation or something else? Also good to have it as part of the code comment so it's easy to understand the intention.

Got it. But this optimization is not quite general. Based on our experiments, the performance which sets sequence length to width dimension (1, 1, seq_len, CL) is better than the performance which sets sequence length to height dimension (1, seq_len, 1, CL) for the input axis order of the conv op. And another reason is that this change will be close with the structure of AI Hub version llama

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai
cccclai merged commit baf35d2 into pytorch:mainMar 13, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.release notes: qualcommChanges to the Qualcomm backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@shewu-quic@facebook-github-bot@cccclai
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Qualcomm AI Engine Direct - Optimize the performance for AR-N model - #9079

Merged
cccclai merged 4 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_arn_model
Mar 13, 2025
Merged

Qualcomm AI Engine Direct - Optimize the performance for AR-N model#9079
cccclai merged 4 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_arn_model

Conversation

@shewu-quic

@shewu-quicshewu-quic commented Mar 10, 2025

Copy link
Copy Markdown
Collaborator

Summary:

  • Fix the bug of rms norm builder
  • Use HuggingFace version RoPE to improve the performance due to stride = 1 in StrideSlice Op
  • Modificate the axis order of the conv in qkv, feedforward and output
    • Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
    • New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)

Test Result:

  • Verify the output for story llama with smart mask, CL=128, prefill_ar_n=16, prompt="Once"
    Note that using Hugging Face RoPE will slightly affect accuracy
    • Original (mainline)
INFO:root:Results[0]:
Once upon a time, there was a little girl named Lily. She loved to play with her toys and her favorite toy was a big, red ball. One day, Lily's mom asked her to help her with the laundry. Lily was happy to help and she put all the clothes in the washing machine. After the clothes were washed, Lily's mom asked her to help her hang them up to dry. Lily saw a big, black rake and asked her mom what it was. Her mom told her it was a rake and that it helps to
  • Optimized (this PR)
INFO:root:Results[0]:
Once upon a time, there was a little girl named Lily. She loved to play with her toys and her favorite toy was a big, red ball. One day, Lily's mom asked her to help her with the laundry. Lily was happy to help and she put all the clothes in the washing machine. After the clothes were washed, Lily's mom asked her to help her hang them up to dry. Lily saw a big, black iron on the counter and asked her mom what it was for. Her mom explained that it was used to make clothes smooth
  • Verify the performance for llama 3.2 1B with shift pointer, CL=2048, prefill_ar_n=256
    • Original (mainline)
I 00:00:02.048851 executorch:runner.cpp:354] Prompt Processor: total 256 tokens (AR-256 * 1 iters)
I 00:00:36.606984 executorch:runner.cpp:456] Prompt Tokens: 256 Generated Tokens: 1791
I 00:00:36.607049 executorch:runner.cpp:462] Model Load Time: 2.012000 (seconds)
I 00:00:36.607062 executorch:runner.cpp:472] Total inference time: 34.592000 (seconds) Rate: 51.774977 (tokens/second)
I 00:00:36.607072 executorch:runner.cpp:480] Prompt evaluation:	0.293000 (seconds) Rate: 873.720137 (tokens/second)
I 00:00:36.607080 executorch:runner.cpp:491] Generated 1791 tokens:	34.299000 (seconds) Rate: 52.217266 (tokens/second)
I 00:00:36.607089 executorch:runner.cpp:499] Time to first generated token:	0.293000 (seconds)
I 00:00:36.607099 executorch:runner.cpp:506] Sampling time over 1791 tokens:	1.473000 (seconds)
  • Optimized (this PR)
I 00:00:01.827440 executorch:runner.cpp:354] Prompt Processor: total 256 tokens (AR-256 * 1 iters)
I 00:00:03.143673 executorch:runner.cpp:456] Prompt Tokens: 256 Generated Tokens: 64
I 00:00:03.143686 executorch:runner.cpp:462] Model Load Time: 1.791000 (seconds)
I 00:00:03.143698 executorch:runner.cpp:472] Total inference time: 1.350000 (seconds) Rate: 47.407407 (tokens/second)
I 00:00:03.143706 executorch:runner.cpp:480] Prompt evaluation:	0.126000 (seconds) Rate: 2031.746032 (tokens/second)
I 00:00:03.143715 executorch:runner.cpp:491] Generated 64 tokens:	1.224000 (seconds) Rate: 52.287582 (tokens/second)
I 00:00:03.143723 executorch:runner.cpp:499] Time to first generated token:	0.126000 (seconds)
I 00:00:03.143733 executorch:runner.cpp:506] Sampling time over 64 tokens:	0.058000 (seconds)

@pytorch-bot

pytorch-botBot commented Mar 10, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/9079

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Unrelated Failure

As of commit c94c0bd with merge base acae017 (image):

NEW FAILURE - The following job has failed:

BROKEN TRUNK - The following job failed but were present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 10, 2025
@shewu-quic

shewu-quic commented Mar 10, 2025

Copy link
Copy Markdown
CollaboratorAuthor

Hi @cccclai,
This PR aims to enhance the performance of the prompt processor (AR-N model).
The main improvement comes from using a stride slice operation with a stride of 1, which offers better performance. Consequently, we’ve switched to the Huggingface version of RoPE.
Could you help to take a look?

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.


def __init__(self):
super().__init__()
def __init__(self, edge_program: torch.export.ExportedProgram):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can follow #8505 to get rid of some recompose logic to reduce engineer effort there

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for your information. I will try it.

) -> torch.Tensor:
x_r, x_i = x[..., ::2], x[..., 1::2]

# Change to RoPE of huggingface version

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Which one is the huggingface version and why is it better?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The implementation of RoPE in huggingface process query and key with two half instead of interleaved way.
The main difference is stride in StrideSlice op. For interleaved way, stride is two which is not friendly for HTP backend about this memory handle.
Ref: huggingface/transformers#25199

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe add this comment to part of the code comment, just so others know the context.

@cccclaicccclai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The perf improvement looks awesome!!

@cccclai

Copy link
Copy Markdown
Contributor

There is still lint error, can you fix it?

@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_arn_model branch from b4d5a63 to fde0f80CompareMarch 12, 2025 03:51
@shewu-quic
shewu-quic requested a review from SS-JIA as a code ownerMarch 12, 2025 03:51
@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

There is still lint error, can you fix it?

Done. Thanks :)

@cccclaicccclai added the release notes: qualcomm Changes to the Qualcomm backend delegate label Mar 12, 2025
@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai

Copy link
Copy Markdown
Contributor

This seems need rebase

 Summary:
- Fix the bug of rms norm builder
- Use HuggingFace version RoPE to improve the performance due to
stride = 1 in StrideSlice Op
- Modificate the axis order of the conv in qkv, feedforward and output
- Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
- New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)
@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_arn_model branch from fde0f80 to c5c149cCompareMarch 12, 2025 23:12
@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai

Copy link
Copy Markdown
Contributor

Can you also share which part of the logic do the following optmization?

Modificate the axis order of the conv in qkv, feedforward and output
Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)

Is it the weight permutation or something else? Also good to have it as part of the code comment so it's easy to understand the intention.

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Can you also share which part of the logic do the following optmization?

Modificate the axis order of the conv in qkv, feedforward and output
Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)

Is it the weight permutation or something else? Also good to have it as part of the code comment so it's easy to understand the intention.

Got it. But this optimization is not quite general. Based on our experiments, the performance which sets sequence length to width dimension (1, 1, seq_len, CL) is better than the performance which sets sequence length to height dimension (1, seq_len, 1, CL) for the input axis order of the conv op. And another reason is that this change will be close with the structure of AI Hub version llama

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai
cccclai merged commit baf35d2 into pytorch:mainMar 13, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.release notes: qualcommChanges to the Qualcomm backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@shewu-quic@facebook-github-bot@cccclai
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Qualcomm AI Engine Direct - Optimize the performance for AR-N model - #9079

Merged
cccclai merged 4 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_arn_model
Mar 13, 2025
Merged

Qualcomm AI Engine Direct - Optimize the performance for AR-N model#9079
cccclai merged 4 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_arn_model

Conversation

@shewu-quic

@shewu-quicshewu-quic commented Mar 10, 2025

Copy link
Copy Markdown
Collaborator

Summary:

  • Fix the bug of rms norm builder
  • Use HuggingFace version RoPE to improve the performance due to stride = 1 in StrideSlice Op
  • Modificate the axis order of the conv in qkv, feedforward and output
    • Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
    • New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)

Test Result:

  • Verify the output for story llama with smart mask, CL=128, prefill_ar_n=16, prompt="Once"
    Note that using Hugging Face RoPE will slightly affect accuracy
    • Original (mainline)
INFO:root:Results[0]:
Once upon a time, there was a little girl named Lily. She loved to play with her toys and her favorite toy was a big, red ball. One day, Lily's mom asked her to help her with the laundry. Lily was happy to help and she put all the clothes in the washing machine. After the clothes were washed, Lily's mom asked her to help her hang them up to dry. Lily saw a big, black rake and asked her mom what it was. Her mom told her it was a rake and that it helps to
  • Optimized (this PR)
INFO:root:Results[0]:
Once upon a time, there was a little girl named Lily. She loved to play with her toys and her favorite toy was a big, red ball. One day, Lily's mom asked her to help her with the laundry. Lily was happy to help and she put all the clothes in the washing machine. After the clothes were washed, Lily's mom asked her to help her hang them up to dry. Lily saw a big, black iron on the counter and asked her mom what it was for. Her mom explained that it was used to make clothes smooth
  • Verify the performance for llama 3.2 1B with shift pointer, CL=2048, prefill_ar_n=256
    • Original (mainline)
I 00:00:02.048851 executorch:runner.cpp:354] Prompt Processor: total 256 tokens (AR-256 * 1 iters)
I 00:00:36.606984 executorch:runner.cpp:456] Prompt Tokens: 256 Generated Tokens: 1791
I 00:00:36.607049 executorch:runner.cpp:462] Model Load Time: 2.012000 (seconds)
I 00:00:36.607062 executorch:runner.cpp:472] Total inference time: 34.592000 (seconds) Rate: 51.774977 (tokens/second)
I 00:00:36.607072 executorch:runner.cpp:480] Prompt evaluation:	0.293000 (seconds) Rate: 873.720137 (tokens/second)
I 00:00:36.607080 executorch:runner.cpp:491] Generated 1791 tokens:	34.299000 (seconds) Rate: 52.217266 (tokens/second)
I 00:00:36.607089 executorch:runner.cpp:499] Time to first generated token:	0.293000 (seconds)
I 00:00:36.607099 executorch:runner.cpp:506] Sampling time over 1791 tokens:	1.473000 (seconds)
  • Optimized (this PR)
I 00:00:01.827440 executorch:runner.cpp:354] Prompt Processor: total 256 tokens (AR-256 * 1 iters)
I 00:00:03.143673 executorch:runner.cpp:456] Prompt Tokens: 256 Generated Tokens: 64
I 00:00:03.143686 executorch:runner.cpp:462] Model Load Time: 1.791000 (seconds)
I 00:00:03.143698 executorch:runner.cpp:472] Total inference time: 1.350000 (seconds) Rate: 47.407407 (tokens/second)
I 00:00:03.143706 executorch:runner.cpp:480] Prompt evaluation:	0.126000 (seconds) Rate: 2031.746032 (tokens/second)
I 00:00:03.143715 executorch:runner.cpp:491] Generated 64 tokens:	1.224000 (seconds) Rate: 52.287582 (tokens/second)
I 00:00:03.143723 executorch:runner.cpp:499] Time to first generated token:	0.126000 (seconds)
I 00:00:03.143733 executorch:runner.cpp:506] Sampling time over 64 tokens:	0.058000 (seconds)

@pytorch-bot

pytorch-botBot commented Mar 10, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/9079

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Unrelated Failure

As of commit c94c0bd with merge base acae017 (image):

NEW FAILURE - The following job has failed:

BROKEN TRUNK - The following job failed but were present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 10, 2025
@shewu-quic

shewu-quic commented Mar 10, 2025

Copy link
Copy Markdown
CollaboratorAuthor

Hi @cccclai,
This PR aims to enhance the performance of the prompt processor (AR-N model).
The main improvement comes from using a stride slice operation with a stride of 1, which offers better performance. Consequently, we’ve switched to the Huggingface version of RoPE.
Could you help to take a look?

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.


def __init__(self):
super().__init__()
def __init__(self, edge_program: torch.export.ExportedProgram):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can follow #8505 to get rid of some recompose logic to reduce engineer effort there

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for your information. I will try it.

) -> torch.Tensor:
x_r, x_i = x[..., ::2], x[..., 1::2]

# Change to RoPE of huggingface version

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Which one is the huggingface version and why is it better?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The implementation of RoPE in huggingface process query and key with two half instead of interleaved way.
The main difference is stride in StrideSlice op. For interleaved way, stride is two which is not friendly for HTP backend about this memory handle.
Ref: huggingface/transformers#25199

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe add this comment to part of the code comment, just so others know the context.

@cccclaicccclai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The perf improvement looks awesome!!

@cccclai

Copy link
Copy Markdown
Contributor

There is still lint error, can you fix it?

@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_arn_model branch from b4d5a63 to fde0f80CompareMarch 12, 2025 03:51
@shewu-quic
shewu-quic requested a review from SS-JIA as a code ownerMarch 12, 2025 03:51
@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

There is still lint error, can you fix it?

Done. Thanks :)

@cccclaicccclai added the release notes: qualcomm Changes to the Qualcomm backend delegate label Mar 12, 2025
@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai

Copy link
Copy Markdown
Contributor

This seems need rebase

 Summary:
- Fix the bug of rms norm builder
- Use HuggingFace version RoPE to improve the performance due to
stride = 1 in StrideSlice Op
- Modificate the axis order of the conv in qkv, feedforward and output
- Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
- New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)
@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_arn_model branch from fde0f80 to c5c149cCompareMarch 12, 2025 23:12
@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai

Copy link
Copy Markdown
Contributor

Can you also share which part of the logic do the following optmization?

Modificate the axis order of the conv in qkv, feedforward and output
Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)

Is it the weight permutation or something else? Also good to have it as part of the code comment so it's easy to understand the intention.

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Can you also share which part of the logic do the following optmization?

Modificate the axis order of the conv in qkv, feedforward and output
Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)

Is it the weight permutation or something else? Also good to have it as part of the code comment so it's easy to understand the intention.

Got it. But this optimization is not quite general. Based on our experiments, the performance which sets sequence length to width dimension (1, 1, seq_len, CL) is better than the performance which sets sequence length to height dimension (1, seq_len, 1, CL) for the input axis order of the conv op. And another reason is that this change will be close with the structure of AI Hub version llama

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai
cccclai merged commit baf35d2 into pytorch:mainMar 13, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.release notes: qualcommChanges to the Qualcomm backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@shewu-quic@facebook-github-bot@cccclai
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Qualcomm AI Engine Direct - Optimize the performance for AR-N model - #9079

Merged
cccclai merged 4 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_arn_model
Mar 13, 2025
Merged

Qualcomm AI Engine Direct - Optimize the performance for AR-N model#9079
cccclai merged 4 commits into
pytorch:mainfrom
CodeLinaro:dev1/hutton/optimize_arn_model

Conversation

@shewu-quic

@shewu-quicshewu-quic commented Mar 10, 2025

Copy link
Copy Markdown
Collaborator

Summary:

  • Fix the bug of rms norm builder
  • Use HuggingFace version RoPE to improve the performance due to stride = 1 in StrideSlice Op
  • Modificate the axis order of the conv in qkv, feedforward and output
    • Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
    • New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)

Test Result:

  • Verify the output for story llama with smart mask, CL=128, prefill_ar_n=16, prompt="Once"
    Note that using Hugging Face RoPE will slightly affect accuracy
    • Original (mainline)
INFO:root:Results[0]:
Once upon a time, there was a little girl named Lily. She loved to play with her toys and her favorite toy was a big, red ball. One day, Lily's mom asked her to help her with the laundry. Lily was happy to help and she put all the clothes in the washing machine. After the clothes were washed, Lily's mom asked her to help her hang them up to dry. Lily saw a big, black rake and asked her mom what it was. Her mom told her it was a rake and that it helps to
  • Optimized (this PR)
INFO:root:Results[0]:
Once upon a time, there was a little girl named Lily. She loved to play with her toys and her favorite toy was a big, red ball. One day, Lily's mom asked her to help her with the laundry. Lily was happy to help and she put all the clothes in the washing machine. After the clothes were washed, Lily's mom asked her to help her hang them up to dry. Lily saw a big, black iron on the counter and asked her mom what it was for. Her mom explained that it was used to make clothes smooth
  • Verify the performance for llama 3.2 1B with shift pointer, CL=2048, prefill_ar_n=256
    • Original (mainline)
I 00:00:02.048851 executorch:runner.cpp:354] Prompt Processor: total 256 tokens (AR-256 * 1 iters)
I 00:00:36.606984 executorch:runner.cpp:456] Prompt Tokens: 256 Generated Tokens: 1791
I 00:00:36.607049 executorch:runner.cpp:462] Model Load Time: 2.012000 (seconds)
I 00:00:36.607062 executorch:runner.cpp:472] Total inference time: 34.592000 (seconds) Rate: 51.774977 (tokens/second)
I 00:00:36.607072 executorch:runner.cpp:480] Prompt evaluation:	0.293000 (seconds) Rate: 873.720137 (tokens/second)
I 00:00:36.607080 executorch:runner.cpp:491] Generated 1791 tokens:	34.299000 (seconds) Rate: 52.217266 (tokens/second)
I 00:00:36.607089 executorch:runner.cpp:499] Time to first generated token:	0.293000 (seconds)
I 00:00:36.607099 executorch:runner.cpp:506] Sampling time over 1791 tokens:	1.473000 (seconds)
  • Optimized (this PR)
I 00:00:01.827440 executorch:runner.cpp:354] Prompt Processor: total 256 tokens (AR-256 * 1 iters)
I 00:00:03.143673 executorch:runner.cpp:456] Prompt Tokens: 256 Generated Tokens: 64
I 00:00:03.143686 executorch:runner.cpp:462] Model Load Time: 1.791000 (seconds)
I 00:00:03.143698 executorch:runner.cpp:472] Total inference time: 1.350000 (seconds) Rate: 47.407407 (tokens/second)
I 00:00:03.143706 executorch:runner.cpp:480] Prompt evaluation:	0.126000 (seconds) Rate: 2031.746032 (tokens/second)
I 00:00:03.143715 executorch:runner.cpp:491] Generated 64 tokens:	1.224000 (seconds) Rate: 52.287582 (tokens/second)
I 00:00:03.143723 executorch:runner.cpp:499] Time to first generated token:	0.126000 (seconds)
I 00:00:03.143733 executorch:runner.cpp:506] Sampling time over 64 tokens:	0.058000 (seconds)

@pytorch-bot

pytorch-botBot commented Mar 10, 2025

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/9079

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Unrelated Failure

As of commit c94c0bd with merge base acae017 (image):

NEW FAILURE - The following job has failed:

BROKEN TRUNK - The following job failed but were present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@facebook-github-botfacebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Mar 10, 2025
@shewu-quic

shewu-quic commented Mar 10, 2025

Copy link
Copy Markdown
CollaboratorAuthor

Hi @cccclai,
This PR aims to enhance the performance of the prompt processor (AR-N model).
The main improvement comes from using a stride slice operation with a stride of 1, which offers better performance. Consequently, we’ve switched to the Huggingface version of RoPE.
Could you help to take a look?

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.


def __init__(self):
super().__init__()
def __init__(self, edge_program: torch.export.ExportedProgram):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can follow #8505 to get rid of some recompose logic to reduce engineer effort there

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for your information. I will try it.

) -> torch.Tensor:
x_r, x_i = x[..., ::2], x[..., 1::2]

# Change to RoPE of huggingface version

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Which one is the huggingface version and why is it better?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The implementation of RoPE in huggingface process query and key with two half instead of interleaved way.
The main difference is stride in StrideSlice op. For interleaved way, stride is two which is not friendly for HTP backend about this memory handle.
Ref: huggingface/transformers#25199

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe add this comment to part of the code comment, just so others know the context.

@cccclaicccclai left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The perf improvement looks awesome!!

@cccclai

Copy link
Copy Markdown
Contributor

There is still lint error, can you fix it?

@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_arn_model branch from b4d5a63 to fde0f80CompareMarch 12, 2025 03:51
@shewu-quic
shewu-quic requested a review from SS-JIA as a code ownerMarch 12, 2025 03:51
@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

There is still lint error, can you fix it?

Done. Thanks :)

@cccclaicccclai added the release notes: qualcomm Changes to the Qualcomm backend delegate label Mar 12, 2025
@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai

Copy link
Copy Markdown
Contributor

This seems need rebase

 Summary:
- Fix the bug of rms norm builder
- Use HuggingFace version RoPE to improve the performance due to
stride = 1 in StrideSlice Op
- Modificate the axis order of the conv in qkv, feedforward and output
- Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
- New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)
@shewu-quic
shewu-quicforce-pushed the dev1/hutton/optimize_arn_model branch from fde0f80 to c5c149cCompareMarch 12, 2025 23:12
@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai

Copy link
Copy Markdown
Contributor

Can you also share which part of the logic do the following optmization?

Modificate the axis order of the conv in qkv, feedforward and output
Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)

Is it the weight permutation or something else? Also good to have it as part of the code comment so it's easy to understand the intention.

@shewu-quic

Copy link
Copy Markdown
CollaboratorAuthor

Can you also share which part of the logic do the following optmization?

Modificate the axis order of the conv in qkv, feedforward and output
Original (AR:128, CL:2048): QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,2048,1)->QNN_Transpose (1,128,1,2048)->self.output-> QNN_Transpose(1,128,2048,1) -> QNN_Reshape (1,1,128,2048)
New: QNN_RmsNorm (1,1,128,2048) -> QNN_Reshape (1,128,1,2048)->QNN_Transpose (1,1,128,2048)->self.output-> QNN_Transpose(1,128,1,2048) -> QNN_Reshape (1,1,128,2048)

Is it the weight permutation or something else? Also good to have it as part of the code comment so it's easy to understand the intention.

Got it. But this optimization is not quite general. Based on our experiments, the performance which sets sequence length to width dimension (1, 1, seq_len, CL) is better than the performance which sets sequence length to height dimension (1, seq_len, 1, CL) for the input axis order of the conv op. And another reason is that this change will be close with the structure of AI Hub version llama

@facebook-github-bot

Copy link
Copy Markdown
Contributor

@cccclai has imported this pull request. If you are a Meta employee, you can view this diff on Phabricator.

@cccclai
cccclai merged commit baf35d2 into pytorch:mainMar 13, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.release notes: qualcommChanges to the Qualcomm backend delegate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@shewu-quic@facebook-github-bot@cccclai