Uh oh!
There was an error while loading. Please reload this page.
Qualcomm AI Engine Direct - [LLM QAT] LLM Quant-Aware Distillation (QAD) - #21036
Conversation
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21036
Note: Links to docs will display an error until the docs builds have been completed. ❌ 1 Cancelled JobAs of commit 54a691b with merge base bc2833f ( CANCELLED JOB - The following job was cancelled. Please retry:
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
@psiddh Hi, With QAT, we can further push quantization down to W4 PCQ, and potentially even lower precisions (W2). Below is a comparison between PTQ and QAT on Please have a look, thanks! Experiment Results: W4 Decoder PCQ (QAT vs PTQ)QAT (W4 PCQ decoder)python examples/qualcomm/oss_scripts/llama/llama.py --build_folder build-android --device ${SERIAL_NUM} --soc_model ${SOC_MODEL} --decoder_model smollm2_135m --model_mode hybrid --max_seq_len 1024 --prompt "What dose it mean to edit written content?" --qat --train_hf_dataset "HuggingFaceTB/smol-smoltalk" --train_hf_limit 1000 --calib_hf_dataset "HuggingFaceTB/smol-smoltalk" --calib_hf_limit 1000Result[INFO 2026-07-20 12:10:52,191 llama.py:290] Device Inference Results[0]:
<|im_start|>user
What dose it mean to edit written content?<|im_end|><|im_start|>assistant
1. **Editing for clarity**: When you edit your writing, you aim to make it clear and concise, ensuring that your message is easy for readers to understand and retain. This involves cutting unnecessary words, phrases, or sentences, and rephrasing or reorganizing your content to make it more readable.
2. **Editing forgrammar**: Grammar is another crucial aspect of writing. When you edit your writing, you check for errorsin grammar, syntax, and punctuation, which can make your writing more polished and error-free.
3. **Editing forstyle**: The final stage of editing is where you refine your writing to make it more engaging, clear, and engaging. This involves using language that is engaging, yet clear, and engagingin the way you want it to be.
4. **Editing forimpact**: The final stage of editing is where you aim to make your writing more impactful. This means that you strive to make your message clear, concise, and impactful, and to convey itin a way that resonates with your readers.
5. **Editing forconsistency**: When you edit your writing, you ensure that it is consistentin terms of tone, style, and voice. This involves using the same language, vocabulary, and structure to create a consistent tone, which is essential for effective communication.
To get started, take a step back and look at your writing. Read it aloud, if possible, to get a sense of the tone, style, and voice you're using. Then, read your writing aloud to get a sense of the rhythm, cadence, and flow.When you're on the right track, you can begin to refine your writing and make it more engaging and effective.<|im_end|>PTQ (W4 PCQ decoder)python examples/qualcomm/oss_scripts/llama/llama.py --build_folder build-android --device ${SERIAL_NUM} --soc_model ${SOC_MODEL} --decoder_model smollm2_135m --model_mode hybrid --max_seq_len 1024 --prompt "What dose it mean to edit written content?" --calib_hf_dataset "HuggingFaceTB/smol-smoltalk" --calib_hf_limit 1000Result[INFO 2026-07-20 15:12:57,609 llama.py:290] Device Inference Results[0]:
<|im_start|>user
What dose it mean to edit written content?<|im_end|><|im_start|>assistant
1. A sentence is usually written in the first person, i.e., "He said, she said" or "She said, she said."
2. A paragraph is usually written in the third person, i.e., "She said, he said" or "She said, she said"
3. A couple is usually written in the second person, i.e., "She said, he said" or "He said, she said"
4. A little is usually written in the third person, i.e., "She said, he said" or "She said, she said"
5. A little is usually written in the second person, i.e., "She said, she said" or "She said, she said"
6. A little is usually written in the second person, i.e., "She said, she said" or "She said, she said"
7. A little is usually written in the second person, i.e., "She said, she said" or "She said, she said"
8. A little is usually written in the second person, i.e., "She said, she said" or "She said, she said"
9. A little is usually written in the second person, i.e., "She said, she said" or "She said, she said"
10. A little is usually written in the second person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 2nd person, i.e., "He said, she said" or "She said, she said"
Please note that the above examples are from the 2nd person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 1st person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 2nd person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 1st person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 2nd person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 1st person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 2nd person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 2nd person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 1st person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 2nd person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 1st person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 2nd person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 1st person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 2nd person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 2nd person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 1st person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 2nd person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 1st person, i.e., "She said, she said" or "She said, she said"
Please note that the above examples are from the 2nd person, i.e., " |
DannyYuyang-quic
commented
Jul 20, 2026
@pytorchbot label "release notes: qualcomm" |
psiddh
commented
Aug 11, 2026
@claude review |
psiddh
commented
Aug 11, 2026
Claude finished @psiddh's task in 5m 52s —— View job Claude Code ReviewReviewed the full diff ( 🔴 Blocking1. CI test passes Either register a 🟠 Correctness2. foriinrange(n-1):
ifi<len(assistant_mask) andassistant_mask[i]:
labels[i-1] =tokens[i]When 3. Backend qconfig change affects existing QAT callers, not just this feature 🟡 Minor / Questions4. Operator-precedence readability in the if (
self.mode==Mode.CALIBRATEandself.control_args.quant_recipe_suggestionoruse_qat
):This parses as 5. 6. 7. 8. Stale docstring in Nits
Overall the design is solid and the abstractions are clean. Item #1 blocks the test from passing and #2/#3 are the substantive correctness items to resolve before merge. |
psiddh
commented
Aug 11, 2026
@DannyYuyang-quic can you resolve teh merge conflicts ? |
…for LLMs Training: - Add BaseTrainer, Trainer (CE), KDTrainer (knowledge distillation) - Add CrossEntropyLoss, KLDivergenceLoss, linear warm-up cosine LR scheduler - Add TrainingArgs dataclass (epochs, lr, alpha, temperature, grad_accum_steps, warmup_ratio, lr_config) Data pipeline: - Add DecoderDatasetBuilder.from_hf_source for HuggingFace SFT chat datasets - Add build_qat_dataloaders: explicit calib/train split - Add LLMTrainingCollator, make_causal_labels, make_conversation_labels - Add DataConfig train fields: train_tasks, train_hf_dataset, train_hf_limit Quantization Strategy: - Add QATStrategy: PTQ calibration pass followed by KDTrainer/Trainer fine-tuning - Branch TextDecoder.quantize on --qat: prepare_qat_pt2e + move_exported_model_to_train vs prepare_pt2e - Select qat_recipe over quant_recipe when --qat is active Quant recipe: - Add StaticLLMQATRecipe base class as a example - Add Smollm2QATQuantRecipe; add qat_recipe field to LLMModelConfig - Add 16a8w QAT qconfig Fix: - SeqMSE: unwrap FakeQuantize wrapper before extracting observer Testing: - Add test_static_llm_qat: assert QAT PPL < PTQ PPL on smollm2_135m
2587988 to
54a691bCompareDannyYuyang-quic
commented
Aug 12, 2026
@psiddh I've rebased the branch and addressed code review comments. Please take a look, thanks! |
Uh oh!
There was an error while loading. Please reload this page.
Summary
Training:
Data pipeline:
Quantization Strategy:
Quant recipe:
Fix:
CI Testing:
README:
examples/qualcomm/oss_scripts/llama/quantization_guidance.mdE2E script:
Result
Test plan
cc: @shewu-quic@haowhsu-quic@winskuo-quic@psiddh@abhinaykukkadapu