[QNN ISSUE] Error Occurred When Running Llama-3.2-1B-Instruct Model with llama_main (QNN 8A8W Quantization) on Device #6546

Description

@crinex

🐛 Describe the bug

Dear @shewu-quic

I followed your instructions to convert the Llama-3.2-1B-Instruct model to a QNN backend.pte file. However, during the On Device process, I encountered an error as shown in Figure 1 below, and I would like you to help me check this issue.

The command below is what I used to generate the .pte file.

MODEL_DIR="Llama-3.2-1B-Instruct/original"
python -m examples.models.llama.export_llama --checkpoint "${MODEL_DIR}/consolidated.00.pth" -p "${MODEL_DIR}/params.json" -kv --disable_dynamic_shape --qnn --pt2e_quantize qnn_8a8w -d fp32 --metadata '{"get_bos_id":128000, "get_eos_ids":[128009, 128001]}' --num_sharding 8 --output_name="llama3_2_qnn_8a8w.pte"

Then, I ran the generated file on the device using the command below.

./llama_main_qnn --model_path /data/local/tmp/llama/llama3_2_qnn_8a8w.pte --tokenizer_path /data/local/tmp/llama/tokenizer.model --prompt "<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nYou are a funny chatbot.<|eot_id|><|start_header_id|>user<|end_header_id|>\n\nWhat is the capital of Korea?<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n" --seq_len 512 —temperature 0

Then, an error occurred as shown in Figure 1 below.

image
[Figure 1. The error occurring when running with llama_main (QNN version).]

I would appreciate it if you could help me identify the issue.
Thank you

Versions

PyTorch version: 2.6.0.dev20241019+cpu
Is debug build: False
CUDA used to build PyTorch: None
ROCM used to build PyTorch: N/A

OS: Ubuntu 22.04.5 LTS (x86_64)
GCC version: (Ubuntu 11.4.0-1ubuntu1~22.04) 11.4.0
Clang version: Could not collect
CMake version: version 3.30.5
Libc version: glibc-2.35

Python version: 3.10.12 (main, Sep 11 2024, 15:47:36) [GCC 11.4.0] (64-bit runtime)
Python platform: Linux-6.8.0-40-generic-x86_64-with-glibc2.35
Is CUDA available: False
CUDA runtime version: No CUDA
CUDA_MODULE_LOADING set to: N/A
GPU models and configuration: No CUDA
Nvidia driver version: No CUDA
cuDNN version: No CUDA
HIP runtime version: N/A
MIOpen runtime version: N/A
Is XNNPACK available: True

CPU:
Architecture: x86_64
CPU op-mode(s): 32-bit, 64-bit
Address sizes: 46 bits physical, 48 bits virtual
Byte Order: Little Endian
CPU(s): 24
On-line CPU(s) list: 0-23
Vendor ID: GenuineIntel
Model name: Intel(R) Core(TM) i9-10920X CPU @ 3.50GHz
CPU family: 6
Model: 85
Thread(s) per core: 2
Core(s) per socket: 12
Socket(s): 1
Stepping: 7
CPU max MHz: 4800.0000
CPU min MHz: 1200.0000
BogoMIPS: 6999.82
Flags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc art arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc cpuid aperfmperf pni pclmulqdq dtes64 monitor ds_cpl vmx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid dca sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm 3dnowprefetch cpuid_fault epb cat_l3 cdp_l3 ssbd mba ibrs ibpb stibp ibrs_enhanced tpr_shadow flexpriority ept vpid ept_ad fsgsbase tsc_adjust bmi1 avx2 smep bmi2 erms invpcid cqm mpx rdt_a avx512f avx512dq rdseed adx smap clflushopt clwb intel_pt avx512cd avx512bw avx512vl xsaveopt xsavec xgetbv1 xsaves cqm_llc cqm_occup_llc cqm_mbm_total cqm_mbm_local dtherm ida arat pln pts hwp hwp_act_window hwp_epp hwp_pkg_req vnmi avx512_vnni md_clear flush_l1d arch_capabilities
Virtualization: VT-x
L1d cache: 384 KiB (12 instances)
L1i cache: 384 KiB (12 instances)
L2 cache: 12 MiB (12 instances)
L3 cache: 19.3 MiB (1 instance)
NUMA node(s): 1
NUMA node0 CPU(s): 0-23
Vulnerability Gather data sampling: Mitigation; Microcode
Vulnerability Itlb multihit: KVM: Mitigation: VMX disabled
Vulnerability L1tf: Not affected
Vulnerability Mds: Not affected
Vulnerability Meltdown: Not affected
Vulnerability Mmio stale data: Mitigation; Clear CPU buffers; SMT vulnerable
Vulnerability Reg file data sampling: Not affected
Vulnerability Retbleed: Mitigation; Enhanced IBRS
Vulnerability Spec rstack overflow: Not affected
Vulnerability Spec store bypass: Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization
Vulnerability Spectre v2: Mitigation; Enhanced / Automatic IBRS; IBPB conditional; RSB filling; PBRSB-eIBRS SW sequence; BHI SW loop, KVM SW loop
Vulnerability Srbds: Not affected
Vulnerability Tsx async abort: Mitigation; TSX disabled

Versions of relevant libraries:
[pip3] executorch==0.5.0a0+16b633b
[pip3] numpy==1.26.4
[pip3] torch==2.6.0.dev20241019+cpu
[pip3] torchao==0.5.0+git0916b5b2
[pip3] torchaudio==2.5.0.dev20241019+cpu
[pip3] torchsr==1.0.4
[pip3] torchvision==0.20.0.dev20241019+cpu
[conda] Could not collect

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    module: qnnIssues related to Qualcomm's QNN delegate and code under backends/qualcomm/partner: qualcommFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Qualcomm

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all \u003cpre\u003e\u003ccode\u003e blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks"); } } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); } })(); (function(){ try { var __m = "github.com"; var __re = new RegExp('^' + "github\\.com" + '
      Skip to content

      [QNN ISSUE] Error Occurred When Running Llama-3.2-1B-Instruct Model with llama_main (QNN 8A8W Quantization) on Device #6546

      Description

      @crinex

      🐛 Describe the bug

      Dear @shewu-quic

      I followed your instructions to convert the Llama-3.2-1B-Instruct model to a QNN backend.pte file. However, during the On Device process, I encountered an error as shown in Figure 1 below, and I would like you to help me check this issue.

      The command below is what I used to generate the .pte file.

      MODEL_DIR="Llama-3.2-1B-Instruct/original"
      python -m examples.models.llama.export_llama --checkpoint "${MODEL_DIR}/consolidated.00.pth" -p "${MODEL_DIR}/params.json" -kv --disable_dynamic_shape --qnn --pt2e_quantize qnn_8a8w -d fp32 --metadata '{"get_bos_id":128000, "get_eos_ids":[128009, 128001]}' --num_sharding 8 --output_name="llama3_2_qnn_8a8w.pte"

      Then, I ran the generated file on the device using the command below.

      ./llama_main_qnn --model_path /data/local/tmp/llama/llama3_2_qnn_8a8w.pte --tokenizer_path /data/local/tmp/llama/tokenizer.model --prompt "<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nYou are a funny chatbot.<|eot_id|><|start_header_id|>user<|end_header_id|>\n\nWhat is the capital of Korea?<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n" --seq_len 512 —temperature 0

      Then, an error occurred as shown in Figure 1 below.

      image
      [Figure 1. The error occurring when running with llama_main (QNN version).]

      I would appreciate it if you could help me identify the issue.
      Thank you

      Versions

      PyTorch version: 2.6.0.dev20241019+cpu
      Is debug build: False
      CUDA used to build PyTorch: None
      ROCM used to build PyTorch: N/A

      OS: Ubuntu 22.04.5 LTS (x86_64)
      GCC version: (Ubuntu 11.4.0-1ubuntu1~22.04) 11.4.0
      Clang version: Could not collect
      CMake version: version 3.30.5
      Libc version: glibc-2.35

      Python version: 3.10.12 (main, Sep 11 2024, 15:47:36) [GCC 11.4.0] (64-bit runtime)
      Python platform: Linux-6.8.0-40-generic-x86_64-with-glibc2.35
      Is CUDA available: False
      CUDA runtime version: No CUDA
      CUDA_MODULE_LOADING set to: N/A
      GPU models and configuration: No CUDA
      Nvidia driver version: No CUDA
      cuDNN version: No CUDA
      HIP runtime version: N/A
      MIOpen runtime version: N/A
      Is XNNPACK available: True

      CPU:
      Architecture: x86_64
      CPU op-mode(s): 32-bit, 64-bit
      Address sizes: 46 bits physical, 48 bits virtual
      Byte Order: Little Endian
      CPU(s): 24
      On-line CPU(s) list: 0-23
      Vendor ID: GenuineIntel
      Model name: Intel(R) Core(TM) i9-10920X CPU @ 3.50GHz
      CPU family: 6
      Model: 85
      Thread(s) per core: 2
      Core(s) per socket: 12
      Socket(s): 1
      Stepping: 7
      CPU max MHz: 4800.0000
      CPU min MHz: 1200.0000
      BogoMIPS: 6999.82
      Flags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc art arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc cpuid aperfmperf pni pclmulqdq dtes64 monitor ds_cpl vmx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid dca sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm 3dnowprefetch cpuid_fault epb cat_l3 cdp_l3 ssbd mba ibrs ibpb stibp ibrs_enhanced tpr_shadow flexpriority ept vpid ept_ad fsgsbase tsc_adjust bmi1 avx2 smep bmi2 erms invpcid cqm mpx rdt_a avx512f avx512dq rdseed adx smap clflushopt clwb intel_pt avx512cd avx512bw avx512vl xsaveopt xsavec xgetbv1 xsaves cqm_llc cqm_occup_llc cqm_mbm_total cqm_mbm_local dtherm ida arat pln pts hwp hwp_act_window hwp_epp hwp_pkg_req vnmi avx512_vnni md_clear flush_l1d arch_capabilities
      Virtualization: VT-x
      L1d cache: 384 KiB (12 instances)
      L1i cache: 384 KiB (12 instances)
      L2 cache: 12 MiB (12 instances)
      L3 cache: 19.3 MiB (1 instance)
      NUMA node(s): 1
      NUMA node0 CPU(s): 0-23
      Vulnerability Gather data sampling: Mitigation; Microcode
      Vulnerability Itlb multihit: KVM: Mitigation: VMX disabled
      Vulnerability L1tf: Not affected
      Vulnerability Mds: Not affected
      Vulnerability Meltdown: Not affected
      Vulnerability Mmio stale data: Mitigation; Clear CPU buffers; SMT vulnerable
      Vulnerability Reg file data sampling: Not affected
      Vulnerability Retbleed: Mitigation; Enhanced IBRS
      Vulnerability Spec rstack overflow: Not affected
      Vulnerability Spec store bypass: Mitigation; Speculative Store Bypass disabled via prctl
      Vulnerability Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization
      Vulnerability Spectre v2: Mitigation; Enhanced / Automatic IBRS; IBPB conditional; RSB filling; PBRSB-eIBRS SW sequence; BHI SW loop, KVM SW loop
      Vulnerability Srbds: Not affected
      Vulnerability Tsx async abort: Mitigation; TSX disabled

      Versions of relevant libraries:
      [pip3] executorch==0.5.0a0+16b633b
      [pip3] numpy==1.26.4
      [pip3] torch==2.6.0.dev20241019+cpu
      [pip3] torchao==0.5.0+git0916b5b2
      [pip3] torchaudio==2.5.0.dev20241019+cpu
      [pip3] torchsr==1.0.4
      [pip3] torchvision==0.20.0.dev20241019+cpu
      [conda] Could not collect

      Activity

      Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

      Metadata

      Metadata

      Assignees

      No one assigned

        Labels

        module: qnnIssues related to Qualcomm's QNN delegate and code under backends/qualcomm/partner: qualcommFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Qualcomm

        Type

        No type

        Projects

        No projects

          Milestone

          No milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
          Skip to content

          [QNN ISSUE] Error Occurred When Running Llama-3.2-1B-Instruct Model with llama_main (QNN 8A8W Quantization) on Device #6546

          Description

          @crinex

          🐛 Describe the bug

          Dear @shewu-quic

          I followed your instructions to convert the Llama-3.2-1B-Instruct model to a QNN backend.pte file. However, during the On Device process, I encountered an error as shown in Figure 1 below, and I would like you to help me check this issue.

          The command below is what I used to generate the .pte file.

          MODEL_DIR="Llama-3.2-1B-Instruct/original"
          python -m examples.models.llama.export_llama --checkpoint "${MODEL_DIR}/consolidated.00.pth" -p "${MODEL_DIR}/params.json" -kv --disable_dynamic_shape --qnn --pt2e_quantize qnn_8a8w -d fp32 --metadata '{"get_bos_id":128000, "get_eos_ids":[128009, 128001]}' --num_sharding 8 --output_name="llama3_2_qnn_8a8w.pte"

          Then, I ran the generated file on the device using the command below.

          ./llama_main_qnn --model_path /data/local/tmp/llama/llama3_2_qnn_8a8w.pte --tokenizer_path /data/local/tmp/llama/tokenizer.model --prompt "<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nYou are a funny chatbot.<|eot_id|><|start_header_id|>user<|end_header_id|>\n\nWhat is the capital of Korea?<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n" --seq_len 512 —temperature 0

          Then, an error occurred as shown in Figure 1 below.

          image
          [Figure 1. The error occurring when running with llama_main (QNN version).]

          I would appreciate it if you could help me identify the issue.
          Thank you

          Versions

          PyTorch version: 2.6.0.dev20241019+cpu
          Is debug build: False
          CUDA used to build PyTorch: None
          ROCM used to build PyTorch: N/A

          OS: Ubuntu 22.04.5 LTS (x86_64)
          GCC version: (Ubuntu 11.4.0-1ubuntu1~22.04) 11.4.0
          Clang version: Could not collect
          CMake version: version 3.30.5
          Libc version: glibc-2.35

          Python version: 3.10.12 (main, Sep 11 2024, 15:47:36) [GCC 11.4.0] (64-bit runtime)
          Python platform: Linux-6.8.0-40-generic-x86_64-with-glibc2.35
          Is CUDA available: False
          CUDA runtime version: No CUDA
          CUDA_MODULE_LOADING set to: N/A
          GPU models and configuration: No CUDA
          Nvidia driver version: No CUDA
          cuDNN version: No CUDA
          HIP runtime version: N/A
          MIOpen runtime version: N/A
          Is XNNPACK available: True

          CPU:
          Architecture: x86_64
          CPU op-mode(s): 32-bit, 64-bit
          Address sizes: 46 bits physical, 48 bits virtual
          Byte Order: Little Endian
          CPU(s): 24
          On-line CPU(s) list: 0-23
          Vendor ID: GenuineIntel
          Model name: Intel(R) Core(TM) i9-10920X CPU @ 3.50GHz
          CPU family: 6
          Model: 85
          Thread(s) per core: 2
          Core(s) per socket: 12
          Socket(s): 1
          Stepping: 7
          CPU max MHz: 4800.0000
          CPU min MHz: 1200.0000
          BogoMIPS: 6999.82
          Flags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc art arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc cpuid aperfmperf pni pclmulqdq dtes64 monitor ds_cpl vmx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid dca sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm 3dnowprefetch cpuid_fault epb cat_l3 cdp_l3 ssbd mba ibrs ibpb stibp ibrs_enhanced tpr_shadow flexpriority ept vpid ept_ad fsgsbase tsc_adjust bmi1 avx2 smep bmi2 erms invpcid cqm mpx rdt_a avx512f avx512dq rdseed adx smap clflushopt clwb intel_pt avx512cd avx512bw avx512vl xsaveopt xsavec xgetbv1 xsaves cqm_llc cqm_occup_llc cqm_mbm_total cqm_mbm_local dtherm ida arat pln pts hwp hwp_act_window hwp_epp hwp_pkg_req vnmi avx512_vnni md_clear flush_l1d arch_capabilities
          Virtualization: VT-x
          L1d cache: 384 KiB (12 instances)
          L1i cache: 384 KiB (12 instances)
          L2 cache: 12 MiB (12 instances)
          L3 cache: 19.3 MiB (1 instance)
          NUMA node(s): 1
          NUMA node0 CPU(s): 0-23
          Vulnerability Gather data sampling: Mitigation; Microcode
          Vulnerability Itlb multihit: KVM: Mitigation: VMX disabled
          Vulnerability L1tf: Not affected
          Vulnerability Mds: Not affected
          Vulnerability Meltdown: Not affected
          Vulnerability Mmio stale data: Mitigation; Clear CPU buffers; SMT vulnerable
          Vulnerability Reg file data sampling: Not affected
          Vulnerability Retbleed: Mitigation; Enhanced IBRS
          Vulnerability Spec rstack overflow: Not affected
          Vulnerability Spec store bypass: Mitigation; Speculative Store Bypass disabled via prctl
          Vulnerability Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization
          Vulnerability Spectre v2: Mitigation; Enhanced / Automatic IBRS; IBPB conditional; RSB filling; PBRSB-eIBRS SW sequence; BHI SW loop, KVM SW loop
          Vulnerability Srbds: Not affected
          Vulnerability Tsx async abort: Mitigation; TSX disabled

          Versions of relevant libraries:
          [pip3] executorch==0.5.0a0+16b633b
          [pip3] numpy==1.26.4
          [pip3] torch==2.6.0.dev20241019+cpu
          [pip3] torchao==0.5.0+git0916b5b2
          [pip3] torchaudio==2.5.0.dev20241019+cpu
          [pip3] torchsr==1.0.4
          [pip3] torchvision==0.20.0.dev20241019+cpu
          [conda] Could not collect

          Activity

          Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

          Metadata

          Metadata

          Assignees

          No one assigned

            Labels

            module: qnnIssues related to Qualcomm's QNN delegate and code under backends/qualcomm/partner: qualcommFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Qualcomm

            Type

            No type

            Projects

            No projects

              Milestone

              No milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length \u003e 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
              Skip to content

              [QNN ISSUE] Error Occurred When Running Llama-3.2-1B-Instruct Model with llama_main (QNN 8A8W Quantization) on Device #6546

              Description

              @crinex

              🐛 Describe the bug

              Dear @shewu-quic

              I followed your instructions to convert the Llama-3.2-1B-Instruct model to a QNN backend.pte file. However, during the On Device process, I encountered an error as shown in Figure 1 below, and I would like you to help me check this issue.

              The command below is what I used to generate the .pte file.

              MODEL_DIR="Llama-3.2-1B-Instruct/original"
              python -m examples.models.llama.export_llama --checkpoint "${MODEL_DIR}/consolidated.00.pth" -p "${MODEL_DIR}/params.json" -kv --disable_dynamic_shape --qnn --pt2e_quantize qnn_8a8w -d fp32 --metadata '{"get_bos_id":128000, "get_eos_ids":[128009, 128001]}' --num_sharding 8 --output_name="llama3_2_qnn_8a8w.pte"

              Then, I ran the generated file on the device using the command below.

              ./llama_main_qnn --model_path /data/local/tmp/llama/llama3_2_qnn_8a8w.pte --tokenizer_path /data/local/tmp/llama/tokenizer.model --prompt "<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nYou are a funny chatbot.<|eot_id|><|start_header_id|>user<|end_header_id|>\n\nWhat is the capital of Korea?<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n" --seq_len 512 —temperature 0

              Then, an error occurred as shown in Figure 1 below.

              image
              [Figure 1. The error occurring when running with llama_main (QNN version).]

              I would appreciate it if you could help me identify the issue.
              Thank you

              Versions

              PyTorch version: 2.6.0.dev20241019+cpu
              Is debug build: False
              CUDA used to build PyTorch: None
              ROCM used to build PyTorch: N/A

              OS: Ubuntu 22.04.5 LTS (x86_64)
              GCC version: (Ubuntu 11.4.0-1ubuntu1~22.04) 11.4.0
              Clang version: Could not collect
              CMake version: version 3.30.5
              Libc version: glibc-2.35

              Python version: 3.10.12 (main, Sep 11 2024, 15:47:36) [GCC 11.4.0] (64-bit runtime)
              Python platform: Linux-6.8.0-40-generic-x86_64-with-glibc2.35
              Is CUDA available: False
              CUDA runtime version: No CUDA
              CUDA_MODULE_LOADING set to: N/A
              GPU models and configuration: No CUDA
              Nvidia driver version: No CUDA
              cuDNN version: No CUDA
              HIP runtime version: N/A
              MIOpen runtime version: N/A
              Is XNNPACK available: True

              CPU:
              Architecture: x86_64
              CPU op-mode(s): 32-bit, 64-bit
              Address sizes: 46 bits physical, 48 bits virtual
              Byte Order: Little Endian
              CPU(s): 24
              On-line CPU(s) list: 0-23
              Vendor ID: GenuineIntel
              Model name: Intel(R) Core(TM) i9-10920X CPU @ 3.50GHz
              CPU family: 6
              Model: 85
              Thread(s) per core: 2
              Core(s) per socket: 12
              Socket(s): 1
              Stepping: 7
              CPU max MHz: 4800.0000
              CPU min MHz: 1200.0000
              BogoMIPS: 6999.82
              Flags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc art arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc cpuid aperfmperf pni pclmulqdq dtes64 monitor ds_cpl vmx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid dca sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm 3dnowprefetch cpuid_fault epb cat_l3 cdp_l3 ssbd mba ibrs ibpb stibp ibrs_enhanced tpr_shadow flexpriority ept vpid ept_ad fsgsbase tsc_adjust bmi1 avx2 smep bmi2 erms invpcid cqm mpx rdt_a avx512f avx512dq rdseed adx smap clflushopt clwb intel_pt avx512cd avx512bw avx512vl xsaveopt xsavec xgetbv1 xsaves cqm_llc cqm_occup_llc cqm_mbm_total cqm_mbm_local dtherm ida arat pln pts hwp hwp_act_window hwp_epp hwp_pkg_req vnmi avx512_vnni md_clear flush_l1d arch_capabilities
              Virtualization: VT-x
              L1d cache: 384 KiB (12 instances)
              L1i cache: 384 KiB (12 instances)
              L2 cache: 12 MiB (12 instances)
              L3 cache: 19.3 MiB (1 instance)
              NUMA node(s): 1
              NUMA node0 CPU(s): 0-23
              Vulnerability Gather data sampling: Mitigation; Microcode
              Vulnerability Itlb multihit: KVM: Mitigation: VMX disabled
              Vulnerability L1tf: Not affected
              Vulnerability Mds: Not affected
              Vulnerability Meltdown: Not affected
              Vulnerability Mmio stale data: Mitigation; Clear CPU buffers; SMT vulnerable
              Vulnerability Reg file data sampling: Not affected
              Vulnerability Retbleed: Mitigation; Enhanced IBRS
              Vulnerability Spec rstack overflow: Not affected
              Vulnerability Spec store bypass: Mitigation; Speculative Store Bypass disabled via prctl
              Vulnerability Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization
              Vulnerability Spectre v2: Mitigation; Enhanced / Automatic IBRS; IBPB conditional; RSB filling; PBRSB-eIBRS SW sequence; BHI SW loop, KVM SW loop
              Vulnerability Srbds: Not affected
              Vulnerability Tsx async abort: Mitigation; TSX disabled

              Versions of relevant libraries:
              [pip3] executorch==0.5.0a0+16b633b
              [pip3] numpy==1.26.4
              [pip3] torch==2.6.0.dev20241019+cpu
              [pip3] torchao==0.5.0+git0916b5b2
              [pip3] torchaudio==2.5.0.dev20241019+cpu
              [pip3] torchsr==1.0.4
              [pip3] torchvision==0.20.0.dev20241019+cpu
              [conda] Could not collect

              Activity

              Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

              Metadata

              Metadata

              Assignees

              No one assigned

                Labels

                module: qnnIssues related to Qualcomm's QNN delegate and code under backends/qualcomm/partner: qualcommFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Qualcomm

                Type

                No type

                Projects

                No projects

                  Milestone

                  No milestone

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions

                  , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
                  Skip to content

                  [QNN ISSUE] Error Occurred When Running Llama-3.2-1B-Instruct Model with llama_main (QNN 8A8W Quantization) on Device #6546

                  Description

                  @crinex

                  🐛 Describe the bug

                  Dear @shewu-quic

                  I followed your instructions to convert the Llama-3.2-1B-Instruct model to a QNN backend.pte file. However, during the On Device process, I encountered an error as shown in Figure 1 below, and I would like you to help me check this issue.

                  The command below is what I used to generate the .pte file.

                  MODEL_DIR="Llama-3.2-1B-Instruct/original"
                  python -m examples.models.llama.export_llama --checkpoint "${MODEL_DIR}/consolidated.00.pth" -p "${MODEL_DIR}/params.json" -kv --disable_dynamic_shape --qnn --pt2e_quantize qnn_8a8w -d fp32 --metadata '{"get_bos_id":128000, "get_eos_ids":[128009, 128001]}' --num_sharding 8 --output_name="llama3_2_qnn_8a8w.pte"

                  Then, I ran the generated file on the device using the command below.

                  ./llama_main_qnn --model_path /data/local/tmp/llama/llama3_2_qnn_8a8w.pte --tokenizer_path /data/local/tmp/llama/tokenizer.model --prompt "<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nYou are a funny chatbot.<|eot_id|><|start_header_id|>user<|end_header_id|>\n\nWhat is the capital of Korea?<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n" --seq_len 512 —temperature 0

                  Then, an error occurred as shown in Figure 1 below.

                  image
                  [Figure 1. The error occurring when running with llama_main (QNN version).]

                  I would appreciate it if you could help me identify the issue.
                  Thank you

                  Versions

                  PyTorch version: 2.6.0.dev20241019+cpu
                  Is debug build: False
                  CUDA used to build PyTorch: None
                  ROCM used to build PyTorch: N/A

                  OS: Ubuntu 22.04.5 LTS (x86_64)
                  GCC version: (Ubuntu 11.4.0-1ubuntu1~22.04) 11.4.0
                  Clang version: Could not collect
                  CMake version: version 3.30.5
                  Libc version: glibc-2.35

                  Python version: 3.10.12 (main, Sep 11 2024, 15:47:36) [GCC 11.4.0] (64-bit runtime)
                  Python platform: Linux-6.8.0-40-generic-x86_64-with-glibc2.35
                  Is CUDA available: False
                  CUDA runtime version: No CUDA
                  CUDA_MODULE_LOADING set to: N/A
                  GPU models and configuration: No CUDA
                  Nvidia driver version: No CUDA
                  cuDNN version: No CUDA
                  HIP runtime version: N/A
                  MIOpen runtime version: N/A
                  Is XNNPACK available: True

                  CPU:
                  Architecture: x86_64
                  CPU op-mode(s): 32-bit, 64-bit
                  Address sizes: 46 bits physical, 48 bits virtual
                  Byte Order: Little Endian
                  CPU(s): 24
                  On-line CPU(s) list: 0-23
                  Vendor ID: GenuineIntel
                  Model name: Intel(R) Core(TM) i9-10920X CPU @ 3.50GHz
                  CPU family: 6
                  Model: 85
                  Thread(s) per core: 2
                  Core(s) per socket: 12
                  Socket(s): 1
                  Stepping: 7
                  CPU max MHz: 4800.0000
                  CPU min MHz: 1200.0000
                  BogoMIPS: 6999.82
                  Flags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc art arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc cpuid aperfmperf pni pclmulqdq dtes64 monitor ds_cpl vmx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid dca sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm 3dnowprefetch cpuid_fault epb cat_l3 cdp_l3 ssbd mba ibrs ibpb stibp ibrs_enhanced tpr_shadow flexpriority ept vpid ept_ad fsgsbase tsc_adjust bmi1 avx2 smep bmi2 erms invpcid cqm mpx rdt_a avx512f avx512dq rdseed adx smap clflushopt clwb intel_pt avx512cd avx512bw avx512vl xsaveopt xsavec xgetbv1 xsaves cqm_llc cqm_occup_llc cqm_mbm_total cqm_mbm_local dtherm ida arat pln pts hwp hwp_act_window hwp_epp hwp_pkg_req vnmi avx512_vnni md_clear flush_l1d arch_capabilities
                  Virtualization: VT-x
                  L1d cache: 384 KiB (12 instances)
                  L1i cache: 384 KiB (12 instances)
                  L2 cache: 12 MiB (12 instances)
                  L3 cache: 19.3 MiB (1 instance)
                  NUMA node(s): 1
                  NUMA node0 CPU(s): 0-23
                  Vulnerability Gather data sampling: Mitigation; Microcode
                  Vulnerability Itlb multihit: KVM: Mitigation: VMX disabled
                  Vulnerability L1tf: Not affected
                  Vulnerability Mds: Not affected
                  Vulnerability Meltdown: Not affected
                  Vulnerability Mmio stale data: Mitigation; Clear CPU buffers; SMT vulnerable
                  Vulnerability Reg file data sampling: Not affected
                  Vulnerability Retbleed: Mitigation; Enhanced IBRS
                  Vulnerability Spec rstack overflow: Not affected
                  Vulnerability Spec store bypass: Mitigation; Speculative Store Bypass disabled via prctl
                  Vulnerability Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization
                  Vulnerability Spectre v2: Mitigation; Enhanced / Automatic IBRS; IBPB conditional; RSB filling; PBRSB-eIBRS SW sequence; BHI SW loop, KVM SW loop
                  Vulnerability Srbds: Not affected
                  Vulnerability Tsx async abort: Mitigation; TSX disabled

                  Versions of relevant libraries:
                  [pip3] executorch==0.5.0a0+16b633b
                  [pip3] numpy==1.26.4
                  [pip3] torch==2.6.0.dev20241019+cpu
                  [pip3] torchao==0.5.0+git0916b5b2
                  [pip3] torchaudio==2.5.0.dev20241019+cpu
                  [pip3] torchsr==1.0.4
                  [pip3] torchvision==0.20.0.dev20241019+cpu
                  [conda] Could not collect

                  Activity

                  Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                  Metadata

                  Metadata

                  Assignees

                  No one assigned

                    Labels

                    module: qnnIssues related to Qualcomm's QNN delegate and code under backends/qualcomm/partner: qualcommFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Qualcomm

                    Type

                    No type

                    Projects

                    No projects

                      Milestone

                      No milestone

                      Relationships

                      None yet

                      Development

                      No branches or pull requests

                      Issue actions

                      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                      Skip to content

                      [QNN ISSUE] Error Occurred When Running Llama-3.2-1B-Instruct Model with llama_main (QNN 8A8W Quantization) on Device #6546

                      Description

                      @crinex

                      🐛 Describe the bug

                      Dear @shewu-quic

                      I followed your instructions to convert the Llama-3.2-1B-Instruct model to a QNN backend.pte file. However, during the On Device process, I encountered an error as shown in Figure 1 below, and I would like you to help me check this issue.

                      The command below is what I used to generate the .pte file.

                      MODEL_DIR="Llama-3.2-1B-Instruct/original"
                      python -m examples.models.llama.export_llama --checkpoint "${MODEL_DIR}/consolidated.00.pth" -p "${MODEL_DIR}/params.json" -kv --disable_dynamic_shape --qnn --pt2e_quantize qnn_8a8w -d fp32 --metadata '{"get_bos_id":128000, "get_eos_ids":[128009, 128001]}' --num_sharding 8 --output_name="llama3_2_qnn_8a8w.pte"

                      Then, I ran the generated file on the device using the command below.

                      ./llama_main_qnn --model_path /data/local/tmp/llama/llama3_2_qnn_8a8w.pte --tokenizer_path /data/local/tmp/llama/tokenizer.model --prompt "<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nYou are a funny chatbot.<|eot_id|><|start_header_id|>user<|end_header_id|>\n\nWhat is the capital of Korea?<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n" --seq_len 512 —temperature 0

                      Then, an error occurred as shown in Figure 1 below.

                      image
                      [Figure 1. The error occurring when running with llama_main (QNN version).]

                      I would appreciate it if you could help me identify the issue.
                      Thank you

                      Versions

                      PyTorch version: 2.6.0.dev20241019+cpu
                      Is debug build: False
                      CUDA used to build PyTorch: None
                      ROCM used to build PyTorch: N/A

                      OS: Ubuntu 22.04.5 LTS (x86_64)
                      GCC version: (Ubuntu 11.4.0-1ubuntu1~22.04) 11.4.0
                      Clang version: Could not collect
                      CMake version: version 3.30.5
                      Libc version: glibc-2.35

                      Python version: 3.10.12 (main, Sep 11 2024, 15:47:36) [GCC 11.4.0] (64-bit runtime)
                      Python platform: Linux-6.8.0-40-generic-x86_64-with-glibc2.35
                      Is CUDA available: False
                      CUDA runtime version: No CUDA
                      CUDA_MODULE_LOADING set to: N/A
                      GPU models and configuration: No CUDA
                      Nvidia driver version: No CUDA
                      cuDNN version: No CUDA
                      HIP runtime version: N/A
                      MIOpen runtime version: N/A
                      Is XNNPACK available: True

                      CPU:
                      Architecture: x86_64
                      CPU op-mode(s): 32-bit, 64-bit
                      Address sizes: 46 bits physical, 48 bits virtual
                      Byte Order: Little Endian
                      CPU(s): 24
                      On-line CPU(s) list: 0-23
                      Vendor ID: GenuineIntel
                      Model name: Intel(R) Core(TM) i9-10920X CPU @ 3.50GHz
                      CPU family: 6
                      Model: 85
                      Thread(s) per core: 2
                      Core(s) per socket: 12
                      Socket(s): 1
                      Stepping: 7
                      CPU max MHz: 4800.0000
                      CPU min MHz: 1200.0000
                      BogoMIPS: 6999.82
                      Flags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc art arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc cpuid aperfmperf pni pclmulqdq dtes64 monitor ds_cpl vmx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid dca sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm 3dnowprefetch cpuid_fault epb cat_l3 cdp_l3 ssbd mba ibrs ibpb stibp ibrs_enhanced tpr_shadow flexpriority ept vpid ept_ad fsgsbase tsc_adjust bmi1 avx2 smep bmi2 erms invpcid cqm mpx rdt_a avx512f avx512dq rdseed adx smap clflushopt clwb intel_pt avx512cd avx512bw avx512vl xsaveopt xsavec xgetbv1 xsaves cqm_llc cqm_occup_llc cqm_mbm_total cqm_mbm_local dtherm ida arat pln pts hwp hwp_act_window hwp_epp hwp_pkg_req vnmi avx512_vnni md_clear flush_l1d arch_capabilities
                      Virtualization: VT-x
                      L1d cache: 384 KiB (12 instances)
                      L1i cache: 384 KiB (12 instances)
                      L2 cache: 12 MiB (12 instances)
                      L3 cache: 19.3 MiB (1 instance)
                      NUMA node(s): 1
                      NUMA node0 CPU(s): 0-23
                      Vulnerability Gather data sampling: Mitigation; Microcode
                      Vulnerability Itlb multihit: KVM: Mitigation: VMX disabled
                      Vulnerability L1tf: Not affected
                      Vulnerability Mds: Not affected
                      Vulnerability Meltdown: Not affected
                      Vulnerability Mmio stale data: Mitigation; Clear CPU buffers; SMT vulnerable
                      Vulnerability Reg file data sampling: Not affected
                      Vulnerability Retbleed: Mitigation; Enhanced IBRS
                      Vulnerability Spec rstack overflow: Not affected
                      Vulnerability Spec store bypass: Mitigation; Speculative Store Bypass disabled via prctl
                      Vulnerability Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization
                      Vulnerability Spectre v2: Mitigation; Enhanced / Automatic IBRS; IBPB conditional; RSB filling; PBRSB-eIBRS SW sequence; BHI SW loop, KVM SW loop
                      Vulnerability Srbds: Not affected
                      Vulnerability Tsx async abort: Mitigation; TSX disabled

                      Versions of relevant libraries:
                      [pip3] executorch==0.5.0a0+16b633b
                      [pip3] numpy==1.26.4
                      [pip3] torch==2.6.0.dev20241019+cpu
                      [pip3] torchao==0.5.0+git0916b5b2
                      [pip3] torchaudio==2.5.0.dev20241019+cpu
                      [pip3] torchsr==1.0.4
                      [pip3] torchvision==0.20.0.dev20241019+cpu
                      [conda] Could not collect

                      Activity

                      Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                      Metadata

                      Metadata

                      Assignees

                      No one assigned

                        Labels

                        module: qnnIssues related to Qualcomm's QNN delegate and code under backends/qualcomm/partner: qualcommFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Qualcomm

                        Type

                        No type

                        Projects

                        No projects

                          Milestone

                          No milestone

                          Relationships

                          None yet

                          Development

                          No branches or pull requests

                          Issue actions

                          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                          Skip to content

                          [QNN ISSUE] Error Occurred When Running Llama-3.2-1B-Instruct Model with llama_main (QNN 8A8W Quantization) on Device #6546

                          Description

                          @crinex

                          🐛 Describe the bug

                          Dear @shewu-quic

                          I followed your instructions to convert the Llama-3.2-1B-Instruct model to a QNN backend.pte file. However, during the On Device process, I encountered an error as shown in Figure 1 below, and I would like you to help me check this issue.

                          The command below is what I used to generate the .pte file.

                          MODEL_DIR="Llama-3.2-1B-Instruct/original"
                          python -m examples.models.llama.export_llama --checkpoint "${MODEL_DIR}/consolidated.00.pth" -p "${MODEL_DIR}/params.json" -kv --disable_dynamic_shape --qnn --pt2e_quantize qnn_8a8w -d fp32 --metadata '{"get_bos_id":128000, "get_eos_ids":[128009, 128001]}' --num_sharding 8 --output_name="llama3_2_qnn_8a8w.pte"

                          Then, I ran the generated file on the device using the command below.

                          ./llama_main_qnn --model_path /data/local/tmp/llama/llama3_2_qnn_8a8w.pte --tokenizer_path /data/local/tmp/llama/tokenizer.model --prompt "<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nYou are a funny chatbot.<|eot_id|><|start_header_id|>user<|end_header_id|>\n\nWhat is the capital of Korea?<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n" --seq_len 512 —temperature 0

                          Then, an error occurred as shown in Figure 1 below.

                          image
                          [Figure 1. The error occurring when running with llama_main (QNN version).]

                          I would appreciate it if you could help me identify the issue.
                          Thank you

                          Versions

                          PyTorch version: 2.6.0.dev20241019+cpu
                          Is debug build: False
                          CUDA used to build PyTorch: None
                          ROCM used to build PyTorch: N/A

                          OS: Ubuntu 22.04.5 LTS (x86_64)
                          GCC version: (Ubuntu 11.4.0-1ubuntu1~22.04) 11.4.0
                          Clang version: Could not collect
                          CMake version: version 3.30.5
                          Libc version: glibc-2.35

                          Python version: 3.10.12 (main, Sep 11 2024, 15:47:36) [GCC 11.4.0] (64-bit runtime)
                          Python platform: Linux-6.8.0-40-generic-x86_64-with-glibc2.35
                          Is CUDA available: False
                          CUDA runtime version: No CUDA
                          CUDA_MODULE_LOADING set to: N/A
                          GPU models and configuration: No CUDA
                          Nvidia driver version: No CUDA
                          cuDNN version: No CUDA
                          HIP runtime version: N/A
                          MIOpen runtime version: N/A
                          Is XNNPACK available: True

                          CPU:
                          Architecture: x86_64
                          CPU op-mode(s): 32-bit, 64-bit
                          Address sizes: 46 bits physical, 48 bits virtual
                          Byte Order: Little Endian
                          CPU(s): 24
                          On-line CPU(s) list: 0-23
                          Vendor ID: GenuineIntel
                          Model name: Intel(R) Core(TM) i9-10920X CPU @ 3.50GHz
                          CPU family: 6
                          Model: 85
                          Thread(s) per core: 2
                          Core(s) per socket: 12
                          Socket(s): 1
                          Stepping: 7
                          CPU max MHz: 4800.0000
                          CPU min MHz: 1200.0000
                          BogoMIPS: 6999.82
                          Flags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc art arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc cpuid aperfmperf pni pclmulqdq dtes64 monitor ds_cpl vmx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid dca sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm 3dnowprefetch cpuid_fault epb cat_l3 cdp_l3 ssbd mba ibrs ibpb stibp ibrs_enhanced tpr_shadow flexpriority ept vpid ept_ad fsgsbase tsc_adjust bmi1 avx2 smep bmi2 erms invpcid cqm mpx rdt_a avx512f avx512dq rdseed adx smap clflushopt clwb intel_pt avx512cd avx512bw avx512vl xsaveopt xsavec xgetbv1 xsaves cqm_llc cqm_occup_llc cqm_mbm_total cqm_mbm_local dtherm ida arat pln pts hwp hwp_act_window hwp_epp hwp_pkg_req vnmi avx512_vnni md_clear flush_l1d arch_capabilities
                          Virtualization: VT-x
                          L1d cache: 384 KiB (12 instances)
                          L1i cache: 384 KiB (12 instances)
                          L2 cache: 12 MiB (12 instances)
                          L3 cache: 19.3 MiB (1 instance)
                          NUMA node(s): 1
                          NUMA node0 CPU(s): 0-23
                          Vulnerability Gather data sampling: Mitigation; Microcode
                          Vulnerability Itlb multihit: KVM: Mitigation: VMX disabled
                          Vulnerability L1tf: Not affected
                          Vulnerability Mds: Not affected
                          Vulnerability Meltdown: Not affected
                          Vulnerability Mmio stale data: Mitigation; Clear CPU buffers; SMT vulnerable
                          Vulnerability Reg file data sampling: Not affected
                          Vulnerability Retbleed: Mitigation; Enhanced IBRS
                          Vulnerability Spec rstack overflow: Not affected
                          Vulnerability Spec store bypass: Mitigation; Speculative Store Bypass disabled via prctl
                          Vulnerability Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization
                          Vulnerability Spectre v2: Mitigation; Enhanced / Automatic IBRS; IBPB conditional; RSB filling; PBRSB-eIBRS SW sequence; BHI SW loop, KVM SW loop
                          Vulnerability Srbds: Not affected
                          Vulnerability Tsx async abort: Mitigation; TSX disabled

                          Versions of relevant libraries:
                          [pip3] executorch==0.5.0a0+16b633b
                          [pip3] numpy==1.26.4
                          [pip3] torch==2.6.0.dev20241019+cpu
                          [pip3] torchao==0.5.0+git0916b5b2
                          [pip3] torchaudio==2.5.0.dev20241019+cpu
                          [pip3] torchsr==1.0.4
                          [pip3] torchvision==0.20.0.dev20241019+cpu
                          [conda] Could not collect

                          Activity

                          Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                          Metadata

                          Metadata

                          Assignees

                          No one assigned

                            Labels

                            module: qnnIssues related to Qualcomm's QNN delegate and code under backends/qualcomm/partner: qualcommFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Qualcomm

                            Type

                            No type

                            Projects

                            No projects

                              Milestone

                              No milestone

                              Relationships

                              None yet

                              Development

                              No branches or pull requests

                              Issue actions

                              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
                              Skip to content

                              [QNN ISSUE] Error Occurred When Running Llama-3.2-1B-Instruct Model with llama_main (QNN 8A8W Quantization) on Device #6546

                              Description

                              @crinex

                              🐛 Describe the bug

                              Dear @shewu-quic

                              I followed your instructions to convert the Llama-3.2-1B-Instruct model to a QNN backend.pte file. However, during the On Device process, I encountered an error as shown in Figure 1 below, and I would like you to help me check this issue.

                              The command below is what I used to generate the .pte file.

                              MODEL_DIR="Llama-3.2-1B-Instruct/original"
                              python -m examples.models.llama.export_llama --checkpoint "${MODEL_DIR}/consolidated.00.pth" -p "${MODEL_DIR}/params.json" -kv --disable_dynamic_shape --qnn --pt2e_quantize qnn_8a8w -d fp32 --metadata '{"get_bos_id":128000, "get_eos_ids":[128009, 128001]}' --num_sharding 8 --output_name="llama3_2_qnn_8a8w.pte"

                              Then, I ran the generated file on the device using the command below.

                              ./llama_main_qnn --model_path /data/local/tmp/llama/llama3_2_qnn_8a8w.pte --tokenizer_path /data/local/tmp/llama/tokenizer.model --prompt "<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nYou are a funny chatbot.<|eot_id|><|start_header_id|>user<|end_header_id|>\n\nWhat is the capital of Korea?<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n" --seq_len 512 —temperature 0

                              Then, an error occurred as shown in Figure 1 below.

                              image
                              [Figure 1. The error occurring when running with llama_main (QNN version).]

                              I would appreciate it if you could help me identify the issue.
                              Thank you

                              Versions

                              PyTorch version: 2.6.0.dev20241019+cpu
                              Is debug build: False
                              CUDA used to build PyTorch: None
                              ROCM used to build PyTorch: N/A

                              OS: Ubuntu 22.04.5 LTS (x86_64)
                              GCC version: (Ubuntu 11.4.0-1ubuntu1~22.04) 11.4.0
                              Clang version: Could not collect
                              CMake version: version 3.30.5
                              Libc version: glibc-2.35

                              Python version: 3.10.12 (main, Sep 11 2024, 15:47:36) [GCC 11.4.0] (64-bit runtime)
                              Python platform: Linux-6.8.0-40-generic-x86_64-with-glibc2.35
                              Is CUDA available: False
                              CUDA runtime version: No CUDA
                              CUDA_MODULE_LOADING set to: N/A
                              GPU models and configuration: No CUDA
                              Nvidia driver version: No CUDA
                              cuDNN version: No CUDA
                              HIP runtime version: N/A
                              MIOpen runtime version: N/A
                              Is XNNPACK available: True

                              CPU:
                              Architecture: x86_64
                              CPU op-mode(s): 32-bit, 64-bit
                              Address sizes: 46 bits physical, 48 bits virtual
                              Byte Order: Little Endian
                              CPU(s): 24
                              On-line CPU(s) list: 0-23
                              Vendor ID: GenuineIntel
                              Model name: Intel(R) Core(TM) i9-10920X CPU @ 3.50GHz
                              CPU family: 6
                              Model: 85
                              Thread(s) per core: 2
                              Core(s) per socket: 12
                              Socket(s): 1
                              Stepping: 7
                              CPU max MHz: 4800.0000
                              CPU min MHz: 1200.0000
                              BogoMIPS: 6999.82
                              Flags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc art arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc cpuid aperfmperf pni pclmulqdq dtes64 monitor ds_cpl vmx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid dca sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm 3dnowprefetch cpuid_fault epb cat_l3 cdp_l3 ssbd mba ibrs ibpb stibp ibrs_enhanced tpr_shadow flexpriority ept vpid ept_ad fsgsbase tsc_adjust bmi1 avx2 smep bmi2 erms invpcid cqm mpx rdt_a avx512f avx512dq rdseed adx smap clflushopt clwb intel_pt avx512cd avx512bw avx512vl xsaveopt xsavec xgetbv1 xsaves cqm_llc cqm_occup_llc cqm_mbm_total cqm_mbm_local dtherm ida arat pln pts hwp hwp_act_window hwp_epp hwp_pkg_req vnmi avx512_vnni md_clear flush_l1d arch_capabilities
                              Virtualization: VT-x
                              L1d cache: 384 KiB (12 instances)
                              L1i cache: 384 KiB (12 instances)
                              L2 cache: 12 MiB (12 instances)
                              L3 cache: 19.3 MiB (1 instance)
                              NUMA node(s): 1
                              NUMA node0 CPU(s): 0-23
                              Vulnerability Gather data sampling: Mitigation; Microcode
                              Vulnerability Itlb multihit: KVM: Mitigation: VMX disabled
                              Vulnerability L1tf: Not affected
                              Vulnerability Mds: Not affected
                              Vulnerability Meltdown: Not affected
                              Vulnerability Mmio stale data: Mitigation; Clear CPU buffers; SMT vulnerable
                              Vulnerability Reg file data sampling: Not affected
                              Vulnerability Retbleed: Mitigation; Enhanced IBRS
                              Vulnerability Spec rstack overflow: Not affected
                              Vulnerability Spec store bypass: Mitigation; Speculative Store Bypass disabled via prctl
                              Vulnerability Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization
                              Vulnerability Spectre v2: Mitigation; Enhanced / Automatic IBRS; IBPB conditional; RSB filling; PBRSB-eIBRS SW sequence; BHI SW loop, KVM SW loop
                              Vulnerability Srbds: Not affected
                              Vulnerability Tsx async abort: Mitigation; TSX disabled

                              Versions of relevant libraries:
                              [pip3] executorch==0.5.0a0+16b633b
                              [pip3] numpy==1.26.4
                              [pip3] torch==2.6.0.dev20241019+cpu
                              [pip3] torchao==0.5.0+git0916b5b2
                              [pip3] torchaudio==2.5.0.dev20241019+cpu
                              [pip3] torchsr==1.0.4
                              [pip3] torchvision==0.20.0.dev20241019+cpu
                              [conda] Could not collect

                              Activity

                              Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                              Metadata

                              Metadata

                              Assignees

                              No one assigned

                                Labels

                                module: qnnIssues related to Qualcomm's QNN delegate and code under backends/qualcomm/partner: qualcommFor backend delegation, kernels, demo, etc. from the 3rd-party partner, Qualcomm

                                Type

                                No type

                                Projects

                                No projects

                                  Milestone

                                  No milestone

                                  Relationships

                                  None yet

                                  Development

                                  No branches or pull requests

                                  Issue actions