[ExecuTorch][WebGPU] Add Llama K16 online causal attention - #21132

Merged
meta-codesync[bot] merged 11 commits into
gh/JCNTH/108/basefrom
gh/JCNTH/108/head
Aug 7, 2026
Merged

[ExecuTorch][WebGPU] Add Llama K16 online causal attention#21132
meta-codesync[bot] merged 11 commits into
gh/JCNTH/108/basefrom
gh/JCNTH/108/head

Conversation

@JCNTH

@JCNTHJCNTH commented Jul 22, 2026

Copy link
Copy Markdown
Contributor

Stack from ghstack (oldest at bottom):

Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.

Key changes:

  • runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
    header): the online-softmax single-pass causal kernel.
  • Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
    handling, and fallback to the materialized route.
    @exported-using-ghexport

Differential Revision: D113171738

Differential Revision: D113171738

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jul 22, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21132

Note: Links to docs will display an error until the docs builds have been completed.

❌ 48 New Failures, 28 Pending

As of commit 2a2acde with merge base 28a7fac (image):

NEW FAILURES - The following jobs have failed:

  • Cadence Build & Test / cpu-test / test-aot / test-aot (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CCF4:2285B6:9FA64A4:215D387A:6A760B4F and timestamp 2026-08-07 16:43:59 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • Cadence Build & Test / cpu-test / test-ops / test-ops (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 8BA4:3FD34A:8647060:1C0B45CC:6A760B50 and timestamp 2026-08-07 16:44:01 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-arm-backend-no-driver (test_pytest_models_tosa) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-backend-no-driver (test_pytest_ops_no_target) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-backend-no-driver (test_pytest_ops_tosa) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-backend-no-driver (test_run_tosa) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-cortex-m-size-test (bare_metal) / linux-job (gh)
    Error: Input required and not supplied: path
  • pull / test-arm-cortex-m-size-test (zephyr-preset) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-binary-size-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B5A2:358D5F:85A25A6:1C262AF0:6A760BB2 and timestamp 2026-08-07 16:45:38 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-custom-ops-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B798:F25C:37BAAEE:BD4C352:6A760B2F and timestamp 2026-08-07 16:43:27 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-llama-runner-qnn-linux (fp32, qnn_16a16w, qnn) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CC30:37837C:8300DD0:1B7CBF24:6A760B84 and timestamp 2026-08-07 16:44:52 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-llama-runner-qnn-linux (fp32, qnn_8a8w, qnn) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B3E8:A8704:36FE3E5:BB0E486:6A760B75 and timestamp 2026-08-07 16:44:37 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add_mul, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID E702:6D448:834B732:1B9D6EB9:6A760B5D and timestamp 2026-08-07 16:44:13 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add_mul, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 87EA:2AB583:90C6DC9:1E698C05:6A760B62 and timestamp 2026-08-07 16:44:18 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID C2F4:28EBC1:4A7761F:FCA651C:6A760B75 and timestamp 2026-08-07 16:44:37 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CB94:957DF:4FA5958:110A0219:6A760B52 and timestamp 2026-08-07 16:44:02 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (emformer_transcribe, portable, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (emformer_transcribe, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 83B4:3FD34A:864D9D5:1C0CA3F0:6A760B5C and timestamp 2026-08-07 16:44:13 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (ic3, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID D922:36966A:36B5B74:BABE9FD:6A760B45 and timestamp 2026-08-07 16:43:49 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (ic3, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CA72:2ED0C1:4D21445:104EF054:6A760B66 and timestamp 2026-08-07 16:44:22 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (linear, portable, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (linear, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 8B64:157945:882A215:1C9637F8:6A760B50 and timestamp 2026-08-07 16:44:00 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (mobilebert, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CF54:1677CB:8842365:1C76D99B:6A760B6F and timestamp 2026-08-07 16:44:31 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (mobilebert, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID D83E:381F55:881BD75:1CADDB76:6A760B71 and timestamp 2026-08-07 16:44:33 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (mv2, portable, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (mv2, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AA84:D4276:501F9C8:10EA32D9:6A760B37 and timestamp 2026-08-07 16:43:35 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (resnet18, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AA0A:1B80FE:80E622D:1B172D42:6A760B52 and timestamp 2026-08-07 16:44:02 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (resnet18, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (resnet50, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B4B4:211490:85F4316:1C169E46:6A760B4E and timestamp 2026-08-07 16:43:58 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (resnet50, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 9992:3EBE49:88F199D:1CBF5B20:6A760B8A and timestamp 2026-08-07 16:44:58 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (mv3, portable, buck2, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID A9CE:2ABE79:4C6DDAA:10299B38:6A760B84 and timestamp 2026-08-07 16:44:52 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (mv3, portable, cmake, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID BF6A:4723F:36B7699:BA082FF:6A760B86 and timestamp 2026-08-07 16:44:54 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (mv3, xnnpack-quantization-delegation, buck2, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux-basic (mv3, xnnpack-quantization-delegation, cmake, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux-basic (vit, portable, buck2, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AA0E:3E7121:3585EA5:B62E226:6A760B78 and timestamp 2026-08-07 16:44:40 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (vit, portable, cmake, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID C9E2:3BE4F5:82CA081:1B7F9ED6:6A760B9D and timestamp 2026-08-07 16:45:17 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (vit, xnnpack-quantization-delegation, buck2, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux-basic (vit, xnnpack-quantization-delegation, cmake, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID D634:A8704:3707F6E:BB2F5D0:6A760B92 and timestamp 2026-08-07 16:45:06 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-moshi-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 9904:36BA82:39BFDAA:C47C542:6A760B45 and timestamp 2026-08-07 16:43:49 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-openvino-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID EA74:56AB6:82F1981:1B7814B3:6A760B8D and timestamp 2026-08-07 16:45:01 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-qnn-buck-build-linux / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-qnn-models-linux (dl3) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AB08:4723F:369EC95:B9B468A:6A760B3F and timestamp 2026-08-07 16:43:43 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-qnn-models-linux (mv2) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AB9E:1BCB3:382B40F:BEF5392:6A760B38 and timestamp 2026-08-07 16:43:36 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-qnn-models-linux (mv3) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 8236:1DE33:4E0A78A:109376F6:6A760B43 and timestamp 2026-08-07 16:43:47 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-quantized-aot-lib-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID E4B2:1677CB:88532AD:1C7A6BD9:6A760B8F and timestamp 2026-08-07 16:45:03 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-setup-linux-gcc / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 9D18:359DC9:4B34578:FEF937B:6A760B43 and timestamp 2026-08-07 16:43:47 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-sqnr-static-llm-qnn-linux (smollm2_135m) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 85E2:3B7C4D:3A51F9D:C610992:6A760B48 and timestamp 2026-08-07 16:43:52 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-static-llama-qnn-linux (stories_110m) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 949A:36966A:36B7B0E:BAC553A:6A760B4A and timestamp 2026-08-07 16:43:54 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jul 22, 2026
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
 fbsource master
[ghstack-poisoned]
@meta-codesync
meta-codesyncBot merged commit c1b38bc into gh/JCNTH/108/baseAug 7, 2026
132 of 182 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JCNTH/108/head branch August 7, 2026 17:20
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@JCNTH@psiddh
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

[ExecuTorch][WebGPU] Add Llama K16 online causal attention - #21132

Merged
meta-codesync[bot] merged 11 commits into
gh/JCNTH/108/basefrom
gh/JCNTH/108/head
Aug 7, 2026
Merged

[ExecuTorch][WebGPU] Add Llama K16 online causal attention#21132
meta-codesync[bot] merged 11 commits into
gh/JCNTH/108/basefrom
gh/JCNTH/108/head

Conversation

@JCNTH

@JCNTHJCNTH commented Jul 22, 2026

Copy link
Copy Markdown
Contributor

Stack from ghstack (oldest at bottom):

Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.

Key changes:

  • runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
    header): the online-softmax single-pass causal kernel.
  • Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
    handling, and fallback to the materialized route.
    @exported-using-ghexport

Differential Revision: D113171738

Differential Revision: D113171738

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jul 22, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21132

Note: Links to docs will display an error until the docs builds have been completed.

❌ 48 New Failures, 28 Pending

As of commit 2a2acde with merge base 28a7fac (image):

NEW FAILURES - The following jobs have failed:

  • Cadence Build & Test / cpu-test / test-aot / test-aot (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CCF4:2285B6:9FA64A4:215D387A:6A760B4F and timestamp 2026-08-07 16:43:59 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • Cadence Build & Test / cpu-test / test-ops / test-ops (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 8BA4:3FD34A:8647060:1C0B45CC:6A760B50 and timestamp 2026-08-07 16:44:01 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-arm-backend-no-driver (test_pytest_models_tosa) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-backend-no-driver (test_pytest_ops_no_target) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-backend-no-driver (test_pytest_ops_tosa) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-backend-no-driver (test_run_tosa) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-cortex-m-size-test (bare_metal) / linux-job (gh)
    Error: Input required and not supplied: path
  • pull / test-arm-cortex-m-size-test (zephyr-preset) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-binary-size-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B5A2:358D5F:85A25A6:1C262AF0:6A760BB2 and timestamp 2026-08-07 16:45:38 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-custom-ops-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B798:F25C:37BAAEE:BD4C352:6A760B2F and timestamp 2026-08-07 16:43:27 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-llama-runner-qnn-linux (fp32, qnn_16a16w, qnn) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CC30:37837C:8300DD0:1B7CBF24:6A760B84 and timestamp 2026-08-07 16:44:52 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-llama-runner-qnn-linux (fp32, qnn_8a8w, qnn) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B3E8:A8704:36FE3E5:BB0E486:6A760B75 and timestamp 2026-08-07 16:44:37 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add_mul, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID E702:6D448:834B732:1B9D6EB9:6A760B5D and timestamp 2026-08-07 16:44:13 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add_mul, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 87EA:2AB583:90C6DC9:1E698C05:6A760B62 and timestamp 2026-08-07 16:44:18 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID C2F4:28EBC1:4A7761F:FCA651C:6A760B75 and timestamp 2026-08-07 16:44:37 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CB94:957DF:4FA5958:110A0219:6A760B52 and timestamp 2026-08-07 16:44:02 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (emformer_transcribe, portable, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (emformer_transcribe, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 83B4:3FD34A:864D9D5:1C0CA3F0:6A760B5C and timestamp 2026-08-07 16:44:13 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (ic3, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID D922:36966A:36B5B74:BABE9FD:6A760B45 and timestamp 2026-08-07 16:43:49 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (ic3, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CA72:2ED0C1:4D21445:104EF054:6A760B66 and timestamp 2026-08-07 16:44:22 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (linear, portable, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (linear, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 8B64:157945:882A215:1C9637F8:6A760B50 and timestamp 2026-08-07 16:44:00 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (mobilebert, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CF54:1677CB:8842365:1C76D99B:6A760B6F and timestamp 2026-08-07 16:44:31 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (mobilebert, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID D83E:381F55:881BD75:1CADDB76:6A760B71 and timestamp 2026-08-07 16:44:33 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (mv2, portable, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (mv2, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AA84:D4276:501F9C8:10EA32D9:6A760B37 and timestamp 2026-08-07 16:43:35 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (resnet18, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AA0A:1B80FE:80E622D:1B172D42:6A760B52 and timestamp 2026-08-07 16:44:02 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (resnet18, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (resnet50, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B4B4:211490:85F4316:1C169E46:6A760B4E and timestamp 2026-08-07 16:43:58 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (resnet50, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 9992:3EBE49:88F199D:1CBF5B20:6A760B8A and timestamp 2026-08-07 16:44:58 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (mv3, portable, buck2, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID A9CE:2ABE79:4C6DDAA:10299B38:6A760B84 and timestamp 2026-08-07 16:44:52 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (mv3, portable, cmake, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID BF6A:4723F:36B7699:BA082FF:6A760B86 and timestamp 2026-08-07 16:44:54 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (mv3, xnnpack-quantization-delegation, buck2, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux-basic (mv3, xnnpack-quantization-delegation, cmake, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux-basic (vit, portable, buck2, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AA0E:3E7121:3585EA5:B62E226:6A760B78 and timestamp 2026-08-07 16:44:40 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (vit, portable, cmake, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID C9E2:3BE4F5:82CA081:1B7F9ED6:6A760B9D and timestamp 2026-08-07 16:45:17 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (vit, xnnpack-quantization-delegation, buck2, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux-basic (vit, xnnpack-quantization-delegation, cmake, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID D634:A8704:3707F6E:BB2F5D0:6A760B92 and timestamp 2026-08-07 16:45:06 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-moshi-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 9904:36BA82:39BFDAA:C47C542:6A760B45 and timestamp 2026-08-07 16:43:49 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-openvino-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID EA74:56AB6:82F1981:1B7814B3:6A760B8D and timestamp 2026-08-07 16:45:01 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-qnn-buck-build-linux / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-qnn-models-linux (dl3) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AB08:4723F:369EC95:B9B468A:6A760B3F and timestamp 2026-08-07 16:43:43 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-qnn-models-linux (mv2) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AB9E:1BCB3:382B40F:BEF5392:6A760B38 and timestamp 2026-08-07 16:43:36 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-qnn-models-linux (mv3) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 8236:1DE33:4E0A78A:109376F6:6A760B43 and timestamp 2026-08-07 16:43:47 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-quantized-aot-lib-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID E4B2:1677CB:88532AD:1C7A6BD9:6A760B8F and timestamp 2026-08-07 16:45:03 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-setup-linux-gcc / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 9D18:359DC9:4B34578:FEF937B:6A760B43 and timestamp 2026-08-07 16:43:47 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-sqnr-static-llm-qnn-linux (smollm2_135m) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 85E2:3B7C4D:3A51F9D:C610992:6A760B48 and timestamp 2026-08-07 16:43:52 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-static-llama-qnn-linux (stories_110m) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 949A:36966A:36B7B0E:BAC553A:6A760B4A and timestamp 2026-08-07 16:43:54 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jul 22, 2026
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
 fbsource master
[ghstack-poisoned]
@meta-codesync
meta-codesyncBot merged commit c1b38bc into gh/JCNTH/108/baseAug 7, 2026
132 of 182 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JCNTH/108/head branch August 7, 2026 17:20
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@JCNTH@psiddh
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[ExecuTorch][WebGPU] Add Llama K16 online causal attention - #21132

Merged
meta-codesync[bot] merged 11 commits into
gh/JCNTH/108/basefrom
gh/JCNTH/108/head
Aug 7, 2026
Merged

[ExecuTorch][WebGPU] Add Llama K16 online causal attention#21132
meta-codesync[bot] merged 11 commits into
gh/JCNTH/108/basefrom
gh/JCNTH/108/head

Conversation

@JCNTH

@JCNTHJCNTH commented Jul 22, 2026

Copy link
Copy Markdown
Contributor

Stack from ghstack (oldest at bottom):

Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.

Key changes:

  • runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
    header): the online-softmax single-pass causal kernel.
  • Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
    handling, and fallback to the materialized route.
    @exported-using-ghexport

Differential Revision: D113171738

Differential Revision: D113171738

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jul 22, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21132

Note: Links to docs will display an error until the docs builds have been completed.

❌ 48 New Failures, 28 Pending

As of commit 2a2acde with merge base 28a7fac (image):

NEW FAILURES - The following jobs have failed:

  • Cadence Build & Test / cpu-test / test-aot / test-aot (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CCF4:2285B6:9FA64A4:215D387A:6A760B4F and timestamp 2026-08-07 16:43:59 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • Cadence Build & Test / cpu-test / test-ops / test-ops (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 8BA4:3FD34A:8647060:1C0B45CC:6A760B50 and timestamp 2026-08-07 16:44:01 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-arm-backend-no-driver (test_pytest_models_tosa) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-backend-no-driver (test_pytest_ops_no_target) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-backend-no-driver (test_pytest_ops_tosa) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-backend-no-driver (test_run_tosa) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-cortex-m-size-test (bare_metal) / linux-job (gh)
    Error: Input required and not supplied: path
  • pull / test-arm-cortex-m-size-test (zephyr-preset) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-binary-size-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B5A2:358D5F:85A25A6:1C262AF0:6A760BB2 and timestamp 2026-08-07 16:45:38 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-custom-ops-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B798:F25C:37BAAEE:BD4C352:6A760B2F and timestamp 2026-08-07 16:43:27 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-llama-runner-qnn-linux (fp32, qnn_16a16w, qnn) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CC30:37837C:8300DD0:1B7CBF24:6A760B84 and timestamp 2026-08-07 16:44:52 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-llama-runner-qnn-linux (fp32, qnn_8a8w, qnn) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B3E8:A8704:36FE3E5:BB0E486:6A760B75 and timestamp 2026-08-07 16:44:37 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add_mul, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID E702:6D448:834B732:1B9D6EB9:6A760B5D and timestamp 2026-08-07 16:44:13 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add_mul, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 87EA:2AB583:90C6DC9:1E698C05:6A760B62 and timestamp 2026-08-07 16:44:18 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID C2F4:28EBC1:4A7761F:FCA651C:6A760B75 and timestamp 2026-08-07 16:44:37 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CB94:957DF:4FA5958:110A0219:6A760B52 and timestamp 2026-08-07 16:44:02 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (emformer_transcribe, portable, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (emformer_transcribe, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 83B4:3FD34A:864D9D5:1C0CA3F0:6A760B5C and timestamp 2026-08-07 16:44:13 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (ic3, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID D922:36966A:36B5B74:BABE9FD:6A760B45 and timestamp 2026-08-07 16:43:49 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (ic3, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CA72:2ED0C1:4D21445:104EF054:6A760B66 and timestamp 2026-08-07 16:44:22 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (linear, portable, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (linear, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 8B64:157945:882A215:1C9637F8:6A760B50 and timestamp 2026-08-07 16:44:00 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (mobilebert, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CF54:1677CB:8842365:1C76D99B:6A760B6F and timestamp 2026-08-07 16:44:31 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (mobilebert, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID D83E:381F55:881BD75:1CADDB76:6A760B71 and timestamp 2026-08-07 16:44:33 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (mv2, portable, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (mv2, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AA84:D4276:501F9C8:10EA32D9:6A760B37 and timestamp 2026-08-07 16:43:35 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (resnet18, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AA0A:1B80FE:80E622D:1B172D42:6A760B52 and timestamp 2026-08-07 16:44:02 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (resnet18, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (resnet50, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B4B4:211490:85F4316:1C169E46:6A760B4E and timestamp 2026-08-07 16:43:58 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (resnet50, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 9992:3EBE49:88F199D:1CBF5B20:6A760B8A and timestamp 2026-08-07 16:44:58 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (mv3, portable, buck2, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID A9CE:2ABE79:4C6DDAA:10299B38:6A760B84 and timestamp 2026-08-07 16:44:52 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (mv3, portable, cmake, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID BF6A:4723F:36B7699:BA082FF:6A760B86 and timestamp 2026-08-07 16:44:54 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (mv3, xnnpack-quantization-delegation, buck2, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux-basic (mv3, xnnpack-quantization-delegation, cmake, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux-basic (vit, portable, buck2, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AA0E:3E7121:3585EA5:B62E226:6A760B78 and timestamp 2026-08-07 16:44:40 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (vit, portable, cmake, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID C9E2:3BE4F5:82CA081:1B7F9ED6:6A760B9D and timestamp 2026-08-07 16:45:17 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (vit, xnnpack-quantization-delegation, buck2, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux-basic (vit, xnnpack-quantization-delegation, cmake, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID D634:A8704:3707F6E:BB2F5D0:6A760B92 and timestamp 2026-08-07 16:45:06 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-moshi-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 9904:36BA82:39BFDAA:C47C542:6A760B45 and timestamp 2026-08-07 16:43:49 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-openvino-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID EA74:56AB6:82F1981:1B7814B3:6A760B8D and timestamp 2026-08-07 16:45:01 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-qnn-buck-build-linux / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-qnn-models-linux (dl3) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AB08:4723F:369EC95:B9B468A:6A760B3F and timestamp 2026-08-07 16:43:43 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-qnn-models-linux (mv2) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AB9E:1BCB3:382B40F:BEF5392:6A760B38 and timestamp 2026-08-07 16:43:36 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-qnn-models-linux (mv3) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 8236:1DE33:4E0A78A:109376F6:6A760B43 and timestamp 2026-08-07 16:43:47 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-quantized-aot-lib-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID E4B2:1677CB:88532AD:1C7A6BD9:6A760B8F and timestamp 2026-08-07 16:45:03 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-setup-linux-gcc / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 9D18:359DC9:4B34578:FEF937B:6A760B43 and timestamp 2026-08-07 16:43:47 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-sqnr-static-llm-qnn-linux (smollm2_135m) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 85E2:3B7C4D:3A51F9D:C610992:6A760B48 and timestamp 2026-08-07 16:43:52 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-static-llama-qnn-linux (stories_110m) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 949A:36966A:36B7B0E:BAC553A:6A760B4A and timestamp 2026-08-07 16:43:54 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jul 22, 2026
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
 fbsource master
[ghstack-poisoned]
@meta-codesync
meta-codesyncBot merged commit c1b38bc into gh/JCNTH/108/baseAug 7, 2026
132 of 182 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JCNTH/108/head branch August 7, 2026 17:20
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@JCNTH@psiddh
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[ExecuTorch][WebGPU] Add Llama K16 online causal attention - #21132

Merged
meta-codesync[bot] merged 11 commits into
gh/JCNTH/108/basefrom
gh/JCNTH/108/head
Aug 7, 2026
Merged

[ExecuTorch][WebGPU] Add Llama K16 online causal attention#21132
meta-codesync[bot] merged 11 commits into
gh/JCNTH/108/basefrom
gh/JCNTH/108/head

Conversation

@JCNTH

@JCNTHJCNTH commented Jul 22, 2026

Copy link
Copy Markdown
Contributor

Stack from ghstack (oldest at bottom):

Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.

Key changes:

  • runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
    header): the online-softmax single-pass causal kernel.
  • Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
    handling, and fallback to the materialized route.
    @exported-using-ghexport

Differential Revision: D113171738

Differential Revision: D113171738

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jul 22, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21132

Note: Links to docs will display an error until the docs builds have been completed.

❌ 48 New Failures, 28 Pending

As of commit 2a2acde with merge base 28a7fac (image):

NEW FAILURES - The following jobs have failed:

  • Cadence Build & Test / cpu-test / test-aot / test-aot (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CCF4:2285B6:9FA64A4:215D387A:6A760B4F and timestamp 2026-08-07 16:43:59 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • Cadence Build & Test / cpu-test / test-ops / test-ops (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 8BA4:3FD34A:8647060:1C0B45CC:6A760B50 and timestamp 2026-08-07 16:44:01 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-arm-backend-no-driver (test_pytest_models_tosa) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-backend-no-driver (test_pytest_ops_no_target) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-backend-no-driver (test_pytest_ops_tosa) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-backend-no-driver (test_run_tosa) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-cortex-m-size-test (bare_metal) / linux-job (gh)
    Error: Input required and not supplied: path
  • pull / test-arm-cortex-m-size-test (zephyr-preset) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-binary-size-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B5A2:358D5F:85A25A6:1C262AF0:6A760BB2 and timestamp 2026-08-07 16:45:38 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-custom-ops-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B798:F25C:37BAAEE:BD4C352:6A760B2F and timestamp 2026-08-07 16:43:27 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-llama-runner-qnn-linux (fp32, qnn_16a16w, qnn) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CC30:37837C:8300DD0:1B7CBF24:6A760B84 and timestamp 2026-08-07 16:44:52 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-llama-runner-qnn-linux (fp32, qnn_8a8w, qnn) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B3E8:A8704:36FE3E5:BB0E486:6A760B75 and timestamp 2026-08-07 16:44:37 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add_mul, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID E702:6D448:834B732:1B9D6EB9:6A760B5D and timestamp 2026-08-07 16:44:13 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add_mul, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 87EA:2AB583:90C6DC9:1E698C05:6A760B62 and timestamp 2026-08-07 16:44:18 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID C2F4:28EBC1:4A7761F:FCA651C:6A760B75 and timestamp 2026-08-07 16:44:37 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CB94:957DF:4FA5958:110A0219:6A760B52 and timestamp 2026-08-07 16:44:02 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (emformer_transcribe, portable, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (emformer_transcribe, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 83B4:3FD34A:864D9D5:1C0CA3F0:6A760B5C and timestamp 2026-08-07 16:44:13 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (ic3, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID D922:36966A:36B5B74:BABE9FD:6A760B45 and timestamp 2026-08-07 16:43:49 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (ic3, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CA72:2ED0C1:4D21445:104EF054:6A760B66 and timestamp 2026-08-07 16:44:22 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (linear, portable, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (linear, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 8B64:157945:882A215:1C9637F8:6A760B50 and timestamp 2026-08-07 16:44:00 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (mobilebert, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CF54:1677CB:8842365:1C76D99B:6A760B6F and timestamp 2026-08-07 16:44:31 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (mobilebert, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID D83E:381F55:881BD75:1CADDB76:6A760B71 and timestamp 2026-08-07 16:44:33 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (mv2, portable, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (mv2, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AA84:D4276:501F9C8:10EA32D9:6A760B37 and timestamp 2026-08-07 16:43:35 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (resnet18, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AA0A:1B80FE:80E622D:1B172D42:6A760B52 and timestamp 2026-08-07 16:44:02 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (resnet18, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (resnet50, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B4B4:211490:85F4316:1C169E46:6A760B4E and timestamp 2026-08-07 16:43:58 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (resnet50, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 9992:3EBE49:88F199D:1CBF5B20:6A760B8A and timestamp 2026-08-07 16:44:58 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (mv3, portable, buck2, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID A9CE:2ABE79:4C6DDAA:10299B38:6A760B84 and timestamp 2026-08-07 16:44:52 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (mv3, portable, cmake, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID BF6A:4723F:36B7699:BA082FF:6A760B86 and timestamp 2026-08-07 16:44:54 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (mv3, xnnpack-quantization-delegation, buck2, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux-basic (mv3, xnnpack-quantization-delegation, cmake, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux-basic (vit, portable, buck2, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AA0E:3E7121:3585EA5:B62E226:6A760B78 and timestamp 2026-08-07 16:44:40 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (vit, portable, cmake, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID C9E2:3BE4F5:82CA081:1B7F9ED6:6A760B9D and timestamp 2026-08-07 16:45:17 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (vit, xnnpack-quantization-delegation, buck2, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux-basic (vit, xnnpack-quantization-delegation, cmake, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID D634:A8704:3707F6E:BB2F5D0:6A760B92 and timestamp 2026-08-07 16:45:06 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-moshi-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 9904:36BA82:39BFDAA:C47C542:6A760B45 and timestamp 2026-08-07 16:43:49 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-openvino-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID EA74:56AB6:82F1981:1B7814B3:6A760B8D and timestamp 2026-08-07 16:45:01 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-qnn-buck-build-linux / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-qnn-models-linux (dl3) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AB08:4723F:369EC95:B9B468A:6A760B3F and timestamp 2026-08-07 16:43:43 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-qnn-models-linux (mv2) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AB9E:1BCB3:382B40F:BEF5392:6A760B38 and timestamp 2026-08-07 16:43:36 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-qnn-models-linux (mv3) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 8236:1DE33:4E0A78A:109376F6:6A760B43 and timestamp 2026-08-07 16:43:47 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-quantized-aot-lib-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID E4B2:1677CB:88532AD:1C7A6BD9:6A760B8F and timestamp 2026-08-07 16:45:03 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-setup-linux-gcc / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 9D18:359DC9:4B34578:FEF937B:6A760B43 and timestamp 2026-08-07 16:43:47 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-sqnr-static-llm-qnn-linux (smollm2_135m) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 85E2:3B7C4D:3A51F9D:C610992:6A760B48 and timestamp 2026-08-07 16:43:52 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-static-llama-qnn-linux (stories_110m) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 949A:36966A:36B7B0E:BAC553A:6A760B4A and timestamp 2026-08-07 16:43:54 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jul 22, 2026
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
 fbsource master
[ghstack-poisoned]
@meta-codesync
meta-codesyncBot merged commit c1b38bc into gh/JCNTH/108/baseAug 7, 2026
132 of 182 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JCNTH/108/head branch August 7, 2026 17:20
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@JCNTH@psiddh
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

[ExecuTorch][WebGPU] Add Llama K16 online causal attention - #21132

Merged
meta-codesync[bot] merged 11 commits into
gh/JCNTH/108/basefrom
gh/JCNTH/108/head
Aug 7, 2026
Merged

[ExecuTorch][WebGPU] Add Llama K16 online causal attention#21132
meta-codesync[bot] merged 11 commits into
gh/JCNTH/108/basefrom
gh/JCNTH/108/head

Conversation

@JCNTH

@JCNTHJCNTH commented Jul 22, 2026

Copy link
Copy Markdown
Contributor

Stack from ghstack (oldest at bottom):

Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.

Key changes:

  • runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
    header): the online-softmax single-pass causal kernel.
  • Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
    handling, and fallback to the materialized route.
    @exported-using-ghexport

Differential Revision: D113171738

Differential Revision: D113171738

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jul 22, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21132

Note: Links to docs will display an error until the docs builds have been completed.

❌ 48 New Failures, 28 Pending

As of commit 2a2acde with merge base 28a7fac (image):

NEW FAILURES - The following jobs have failed:

  • Cadence Build & Test / cpu-test / test-aot / test-aot (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CCF4:2285B6:9FA64A4:215D387A:6A760B4F and timestamp 2026-08-07 16:43:59 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • Cadence Build & Test / cpu-test / test-ops / test-ops (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 8BA4:3FD34A:8647060:1C0B45CC:6A760B50 and timestamp 2026-08-07 16:44:01 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-arm-backend-no-driver (test_pytest_models_tosa) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-backend-no-driver (test_pytest_ops_no_target) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-backend-no-driver (test_pytest_ops_tosa) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-backend-no-driver (test_run_tosa) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-cortex-m-size-test (bare_metal) / linux-job (gh)
    Error: Input required and not supplied: path
  • pull / test-arm-cortex-m-size-test (zephyr-preset) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-binary-size-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B5A2:358D5F:85A25A6:1C262AF0:6A760BB2 and timestamp 2026-08-07 16:45:38 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-custom-ops-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B798:F25C:37BAAEE:BD4C352:6A760B2F and timestamp 2026-08-07 16:43:27 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-llama-runner-qnn-linux (fp32, qnn_16a16w, qnn) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CC30:37837C:8300DD0:1B7CBF24:6A760B84 and timestamp 2026-08-07 16:44:52 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-llama-runner-qnn-linux (fp32, qnn_8a8w, qnn) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B3E8:A8704:36FE3E5:BB0E486:6A760B75 and timestamp 2026-08-07 16:44:37 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add_mul, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID E702:6D448:834B732:1B9D6EB9:6A760B5D and timestamp 2026-08-07 16:44:13 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add_mul, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 87EA:2AB583:90C6DC9:1E698C05:6A760B62 and timestamp 2026-08-07 16:44:18 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID C2F4:28EBC1:4A7761F:FCA651C:6A760B75 and timestamp 2026-08-07 16:44:37 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CB94:957DF:4FA5958:110A0219:6A760B52 and timestamp 2026-08-07 16:44:02 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (emformer_transcribe, portable, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (emformer_transcribe, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 83B4:3FD34A:864D9D5:1C0CA3F0:6A760B5C and timestamp 2026-08-07 16:44:13 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (ic3, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID D922:36966A:36B5B74:BABE9FD:6A760B45 and timestamp 2026-08-07 16:43:49 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (ic3, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CA72:2ED0C1:4D21445:104EF054:6A760B66 and timestamp 2026-08-07 16:44:22 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (linear, portable, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (linear, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 8B64:157945:882A215:1C9637F8:6A760B50 and timestamp 2026-08-07 16:44:00 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (mobilebert, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CF54:1677CB:8842365:1C76D99B:6A760B6F and timestamp 2026-08-07 16:44:31 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (mobilebert, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID D83E:381F55:881BD75:1CADDB76:6A760B71 and timestamp 2026-08-07 16:44:33 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (mv2, portable, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (mv2, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AA84:D4276:501F9C8:10EA32D9:6A760B37 and timestamp 2026-08-07 16:43:35 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (resnet18, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AA0A:1B80FE:80E622D:1B172D42:6A760B52 and timestamp 2026-08-07 16:44:02 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (resnet18, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (resnet50, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B4B4:211490:85F4316:1C169E46:6A760B4E and timestamp 2026-08-07 16:43:58 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (resnet50, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 9992:3EBE49:88F199D:1CBF5B20:6A760B8A and timestamp 2026-08-07 16:44:58 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (mv3, portable, buck2, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID A9CE:2ABE79:4C6DDAA:10299B38:6A760B84 and timestamp 2026-08-07 16:44:52 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (mv3, portable, cmake, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID BF6A:4723F:36B7699:BA082FF:6A760B86 and timestamp 2026-08-07 16:44:54 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (mv3, xnnpack-quantization-delegation, buck2, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux-basic (mv3, xnnpack-quantization-delegation, cmake, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux-basic (vit, portable, buck2, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AA0E:3E7121:3585EA5:B62E226:6A760B78 and timestamp 2026-08-07 16:44:40 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (vit, portable, cmake, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID C9E2:3BE4F5:82CA081:1B7F9ED6:6A760B9D and timestamp 2026-08-07 16:45:17 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (vit, xnnpack-quantization-delegation, buck2, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux-basic (vit, xnnpack-quantization-delegation, cmake, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID D634:A8704:3707F6E:BB2F5D0:6A760B92 and timestamp 2026-08-07 16:45:06 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-moshi-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 9904:36BA82:39BFDAA:C47C542:6A760B45 and timestamp 2026-08-07 16:43:49 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-openvino-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID EA74:56AB6:82F1981:1B7814B3:6A760B8D and timestamp 2026-08-07 16:45:01 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-qnn-buck-build-linux / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-qnn-models-linux (dl3) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AB08:4723F:369EC95:B9B468A:6A760B3F and timestamp 2026-08-07 16:43:43 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-qnn-models-linux (mv2) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AB9E:1BCB3:382B40F:BEF5392:6A760B38 and timestamp 2026-08-07 16:43:36 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-qnn-models-linux (mv3) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 8236:1DE33:4E0A78A:109376F6:6A760B43 and timestamp 2026-08-07 16:43:47 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-quantized-aot-lib-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID E4B2:1677CB:88532AD:1C7A6BD9:6A760B8F and timestamp 2026-08-07 16:45:03 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-setup-linux-gcc / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 9D18:359DC9:4B34578:FEF937B:6A760B43 and timestamp 2026-08-07 16:43:47 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-sqnr-static-llm-qnn-linux (smollm2_135m) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 85E2:3B7C4D:3A51F9D:C610992:6A760B48 and timestamp 2026-08-07 16:43:52 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-static-llama-qnn-linux (stories_110m) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 949A:36966A:36B7B0E:BAC553A:6A760B4A and timestamp 2026-08-07 16:43:54 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jul 22, 2026
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
 fbsource master
[ghstack-poisoned]
@meta-codesync
meta-codesyncBot merged commit c1b38bc into gh/JCNTH/108/baseAug 7, 2026
132 of 182 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JCNTH/108/head branch August 7, 2026 17:20
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@JCNTH@psiddh
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[ExecuTorch][WebGPU] Add Llama K16 online causal attention - #21132

Merged
meta-codesync[bot] merged 11 commits into
gh/JCNTH/108/basefrom
gh/JCNTH/108/head
Aug 7, 2026
Merged

[ExecuTorch][WebGPU] Add Llama K16 online causal attention#21132
meta-codesync[bot] merged 11 commits into
gh/JCNTH/108/basefrom
gh/JCNTH/108/head

Conversation

@JCNTH

@JCNTHJCNTH commented Jul 22, 2026

Copy link
Copy Markdown
Contributor

Stack from ghstack (oldest at bottom):

Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.

Key changes:

  • runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
    header): the online-softmax single-pass causal kernel.
  • Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
    handling, and fallback to the materialized route.
    @exported-using-ghexport

Differential Revision: D113171738

Differential Revision: D113171738

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jul 22, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21132

Note: Links to docs will display an error until the docs builds have been completed.

❌ 48 New Failures, 28 Pending

As of commit 2a2acde with merge base 28a7fac (image):

NEW FAILURES - The following jobs have failed:

  • Cadence Build & Test / cpu-test / test-aot / test-aot (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CCF4:2285B6:9FA64A4:215D387A:6A760B4F and timestamp 2026-08-07 16:43:59 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • Cadence Build & Test / cpu-test / test-ops / test-ops (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 8BA4:3FD34A:8647060:1C0B45CC:6A760B50 and timestamp 2026-08-07 16:44:01 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-arm-backend-no-driver (test_pytest_models_tosa) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-backend-no-driver (test_pytest_ops_no_target) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-backend-no-driver (test_pytest_ops_tosa) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-backend-no-driver (test_run_tosa) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-cortex-m-size-test (bare_metal) / linux-job (gh)
    Error: Input required and not supplied: path
  • pull / test-arm-cortex-m-size-test (zephyr-preset) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-binary-size-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B5A2:358D5F:85A25A6:1C262AF0:6A760BB2 and timestamp 2026-08-07 16:45:38 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-custom-ops-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B798:F25C:37BAAEE:BD4C352:6A760B2F and timestamp 2026-08-07 16:43:27 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-llama-runner-qnn-linux (fp32, qnn_16a16w, qnn) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CC30:37837C:8300DD0:1B7CBF24:6A760B84 and timestamp 2026-08-07 16:44:52 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-llama-runner-qnn-linux (fp32, qnn_8a8w, qnn) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B3E8:A8704:36FE3E5:BB0E486:6A760B75 and timestamp 2026-08-07 16:44:37 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add_mul, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID E702:6D448:834B732:1B9D6EB9:6A760B5D and timestamp 2026-08-07 16:44:13 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add_mul, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 87EA:2AB583:90C6DC9:1E698C05:6A760B62 and timestamp 2026-08-07 16:44:18 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID C2F4:28EBC1:4A7761F:FCA651C:6A760B75 and timestamp 2026-08-07 16:44:37 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CB94:957DF:4FA5958:110A0219:6A760B52 and timestamp 2026-08-07 16:44:02 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (emformer_transcribe, portable, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (emformer_transcribe, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 83B4:3FD34A:864D9D5:1C0CA3F0:6A760B5C and timestamp 2026-08-07 16:44:13 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (ic3, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID D922:36966A:36B5B74:BABE9FD:6A760B45 and timestamp 2026-08-07 16:43:49 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (ic3, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CA72:2ED0C1:4D21445:104EF054:6A760B66 and timestamp 2026-08-07 16:44:22 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (linear, portable, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (linear, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 8B64:157945:882A215:1C9637F8:6A760B50 and timestamp 2026-08-07 16:44:00 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (mobilebert, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CF54:1677CB:8842365:1C76D99B:6A760B6F and timestamp 2026-08-07 16:44:31 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (mobilebert, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID D83E:381F55:881BD75:1CADDB76:6A760B71 and timestamp 2026-08-07 16:44:33 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (mv2, portable, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (mv2, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AA84:D4276:501F9C8:10EA32D9:6A760B37 and timestamp 2026-08-07 16:43:35 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (resnet18, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AA0A:1B80FE:80E622D:1B172D42:6A760B52 and timestamp 2026-08-07 16:44:02 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (resnet18, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (resnet50, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B4B4:211490:85F4316:1C169E46:6A760B4E and timestamp 2026-08-07 16:43:58 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (resnet50, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 9992:3EBE49:88F199D:1CBF5B20:6A760B8A and timestamp 2026-08-07 16:44:58 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (mv3, portable, buck2, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID A9CE:2ABE79:4C6DDAA:10299B38:6A760B84 and timestamp 2026-08-07 16:44:52 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (mv3, portable, cmake, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID BF6A:4723F:36B7699:BA082FF:6A760B86 and timestamp 2026-08-07 16:44:54 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (mv3, xnnpack-quantization-delegation, buck2, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux-basic (mv3, xnnpack-quantization-delegation, cmake, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux-basic (vit, portable, buck2, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AA0E:3E7121:3585EA5:B62E226:6A760B78 and timestamp 2026-08-07 16:44:40 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (vit, portable, cmake, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID C9E2:3BE4F5:82CA081:1B7F9ED6:6A760B9D and timestamp 2026-08-07 16:45:17 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (vit, xnnpack-quantization-delegation, buck2, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux-basic (vit, xnnpack-quantization-delegation, cmake, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID D634:A8704:3707F6E:BB2F5D0:6A760B92 and timestamp 2026-08-07 16:45:06 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-moshi-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 9904:36BA82:39BFDAA:C47C542:6A760B45 and timestamp 2026-08-07 16:43:49 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-openvino-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID EA74:56AB6:82F1981:1B7814B3:6A760B8D and timestamp 2026-08-07 16:45:01 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-qnn-buck-build-linux / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-qnn-models-linux (dl3) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AB08:4723F:369EC95:B9B468A:6A760B3F and timestamp 2026-08-07 16:43:43 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-qnn-models-linux (mv2) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AB9E:1BCB3:382B40F:BEF5392:6A760B38 and timestamp 2026-08-07 16:43:36 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-qnn-models-linux (mv3) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 8236:1DE33:4E0A78A:109376F6:6A760B43 and timestamp 2026-08-07 16:43:47 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-quantized-aot-lib-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID E4B2:1677CB:88532AD:1C7A6BD9:6A760B8F and timestamp 2026-08-07 16:45:03 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-setup-linux-gcc / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 9D18:359DC9:4B34578:FEF937B:6A760B43 and timestamp 2026-08-07 16:43:47 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-sqnr-static-llm-qnn-linux (smollm2_135m) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 85E2:3B7C4D:3A51F9D:C610992:6A760B48 and timestamp 2026-08-07 16:43:52 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-static-llama-qnn-linux (stories_110m) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 949A:36966A:36B7B0E:BAC553A:6A760B4A and timestamp 2026-08-07 16:43:54 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jul 22, 2026
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
 fbsource master
[ghstack-poisoned]
@meta-codesync
meta-codesyncBot merged commit c1b38bc into gh/JCNTH/108/baseAug 7, 2026
132 of 182 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JCNTH/108/head branch August 7, 2026 17:20
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@JCNTH@psiddh
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[ExecuTorch][WebGPU] Add Llama K16 online causal attention - #21132

Merged
meta-codesync[bot] merged 11 commits into
gh/JCNTH/108/basefrom
gh/JCNTH/108/head
Aug 7, 2026
Merged

[ExecuTorch][WebGPU] Add Llama K16 online causal attention#21132
meta-codesync[bot] merged 11 commits into
gh/JCNTH/108/basefrom
gh/JCNTH/108/head

Conversation

@JCNTH

@JCNTHJCNTH commented Jul 22, 2026

Copy link
Copy Markdown
Contributor

Stack from ghstack (oldest at bottom):

Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.

Key changes:

  • runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
    header): the online-softmax single-pass causal kernel.
  • Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
    handling, and fallback to the materialized route.
    @exported-using-ghexport

Differential Revision: D113171738

Differential Revision: D113171738

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jul 22, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21132

Note: Links to docs will display an error until the docs builds have been completed.

❌ 48 New Failures, 28 Pending

As of commit 2a2acde with merge base 28a7fac (image):

NEW FAILURES - The following jobs have failed:

  • Cadence Build & Test / cpu-test / test-aot / test-aot (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CCF4:2285B6:9FA64A4:215D387A:6A760B4F and timestamp 2026-08-07 16:43:59 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • Cadence Build & Test / cpu-test / test-ops / test-ops (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 8BA4:3FD34A:8647060:1C0B45CC:6A760B50 and timestamp 2026-08-07 16:44:01 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-arm-backend-no-driver (test_pytest_models_tosa) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-backend-no-driver (test_pytest_ops_no_target) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-backend-no-driver (test_pytest_ops_tosa) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-backend-no-driver (test_run_tosa) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-cortex-m-size-test (bare_metal) / linux-job (gh)
    Error: Input required and not supplied: path
  • pull / test-arm-cortex-m-size-test (zephyr-preset) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-binary-size-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B5A2:358D5F:85A25A6:1C262AF0:6A760BB2 and timestamp 2026-08-07 16:45:38 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-custom-ops-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B798:F25C:37BAAEE:BD4C352:6A760B2F and timestamp 2026-08-07 16:43:27 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-llama-runner-qnn-linux (fp32, qnn_16a16w, qnn) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CC30:37837C:8300DD0:1B7CBF24:6A760B84 and timestamp 2026-08-07 16:44:52 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-llama-runner-qnn-linux (fp32, qnn_8a8w, qnn) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B3E8:A8704:36FE3E5:BB0E486:6A760B75 and timestamp 2026-08-07 16:44:37 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add_mul, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID E702:6D448:834B732:1B9D6EB9:6A760B5D and timestamp 2026-08-07 16:44:13 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add_mul, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 87EA:2AB583:90C6DC9:1E698C05:6A760B62 and timestamp 2026-08-07 16:44:18 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID C2F4:28EBC1:4A7761F:FCA651C:6A760B75 and timestamp 2026-08-07 16:44:37 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CB94:957DF:4FA5958:110A0219:6A760B52 and timestamp 2026-08-07 16:44:02 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (emformer_transcribe, portable, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (emformer_transcribe, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 83B4:3FD34A:864D9D5:1C0CA3F0:6A760B5C and timestamp 2026-08-07 16:44:13 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (ic3, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID D922:36966A:36B5B74:BABE9FD:6A760B45 and timestamp 2026-08-07 16:43:49 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (ic3, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CA72:2ED0C1:4D21445:104EF054:6A760B66 and timestamp 2026-08-07 16:44:22 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (linear, portable, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (linear, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 8B64:157945:882A215:1C9637F8:6A760B50 and timestamp 2026-08-07 16:44:00 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (mobilebert, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CF54:1677CB:8842365:1C76D99B:6A760B6F and timestamp 2026-08-07 16:44:31 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (mobilebert, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID D83E:381F55:881BD75:1CADDB76:6A760B71 and timestamp 2026-08-07 16:44:33 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (mv2, portable, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (mv2, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AA84:D4276:501F9C8:10EA32D9:6A760B37 and timestamp 2026-08-07 16:43:35 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (resnet18, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AA0A:1B80FE:80E622D:1B172D42:6A760B52 and timestamp 2026-08-07 16:44:02 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (resnet18, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (resnet50, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B4B4:211490:85F4316:1C169E46:6A760B4E and timestamp 2026-08-07 16:43:58 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (resnet50, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 9992:3EBE49:88F199D:1CBF5B20:6A760B8A and timestamp 2026-08-07 16:44:58 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (mv3, portable, buck2, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID A9CE:2ABE79:4C6DDAA:10299B38:6A760B84 and timestamp 2026-08-07 16:44:52 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (mv3, portable, cmake, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID BF6A:4723F:36B7699:BA082FF:6A760B86 and timestamp 2026-08-07 16:44:54 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (mv3, xnnpack-quantization-delegation, buck2, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux-basic (mv3, xnnpack-quantization-delegation, cmake, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux-basic (vit, portable, buck2, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AA0E:3E7121:3585EA5:B62E226:6A760B78 and timestamp 2026-08-07 16:44:40 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (vit, portable, cmake, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID C9E2:3BE4F5:82CA081:1B7F9ED6:6A760B9D and timestamp 2026-08-07 16:45:17 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (vit, xnnpack-quantization-delegation, buck2, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux-basic (vit, xnnpack-quantization-delegation, cmake, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID D634:A8704:3707F6E:BB2F5D0:6A760B92 and timestamp 2026-08-07 16:45:06 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-moshi-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 9904:36BA82:39BFDAA:C47C542:6A760B45 and timestamp 2026-08-07 16:43:49 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-openvino-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID EA74:56AB6:82F1981:1B7814B3:6A760B8D and timestamp 2026-08-07 16:45:01 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-qnn-buck-build-linux / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-qnn-models-linux (dl3) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AB08:4723F:369EC95:B9B468A:6A760B3F and timestamp 2026-08-07 16:43:43 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-qnn-models-linux (mv2) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AB9E:1BCB3:382B40F:BEF5392:6A760B38 and timestamp 2026-08-07 16:43:36 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-qnn-models-linux (mv3) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 8236:1DE33:4E0A78A:109376F6:6A760B43 and timestamp 2026-08-07 16:43:47 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-quantized-aot-lib-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID E4B2:1677CB:88532AD:1C7A6BD9:6A760B8F and timestamp 2026-08-07 16:45:03 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-setup-linux-gcc / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 9D18:359DC9:4B34578:FEF937B:6A760B43 and timestamp 2026-08-07 16:43:47 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-sqnr-static-llm-qnn-linux (smollm2_135m) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 85E2:3B7C4D:3A51F9D:C610992:6A760B48 and timestamp 2026-08-07 16:43:52 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-static-llama-qnn-linux (stories_110m) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 949A:36966A:36B7B0E:BAC553A:6A760B4A and timestamp 2026-08-07 16:43:54 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jul 22, 2026
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
 fbsource master
[ghstack-poisoned]
@meta-codesync
meta-codesyncBot merged commit c1b38bc into gh/JCNTH/108/baseAug 7, 2026
132 of 182 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JCNTH/108/head branch August 7, 2026 17:20
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@JCNTH@psiddh
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

[ExecuTorch][WebGPU] Add Llama K16 online causal attention - #21132

Merged
meta-codesync[bot] merged 11 commits into
gh/JCNTH/108/basefrom
gh/JCNTH/108/head
Aug 7, 2026
Merged

[ExecuTorch][WebGPU] Add Llama K16 online causal attention#21132
meta-codesync[bot] merged 11 commits into
gh/JCNTH/108/basefrom
gh/JCNTH/108/head

Conversation

@JCNTH

@JCNTHJCNTH commented Jul 22, 2026

Copy link
Copy Markdown
Contributor

Stack from ghstack (oldest at bottom):

Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.

Key changes:

  • runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
    header): the online-softmax single-pass causal kernel.
  • Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
    handling, and fallback to the materialized route.
    @exported-using-ghexport

Differential Revision: D113171738

Differential Revision: D113171738

[ghstack-poisoned]
@pytorch-bot

pytorch-botBot commented Jul 22, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21132

Note: Links to docs will display an error until the docs builds have been completed.

❌ 48 New Failures, 28 Pending

As of commit 2a2acde with merge base 28a7fac (image):

NEW FAILURES - The following jobs have failed:

  • Cadence Build & Test / cpu-test / test-aot / test-aot (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CCF4:2285B6:9FA64A4:215D387A:6A760B4F and timestamp 2026-08-07 16:43:59 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • Cadence Build & Test / cpu-test / test-ops / test-ops (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 8BA4:3FD34A:8647060:1C0B45CC:6A760B50 and timestamp 2026-08-07 16:44:01 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-arm-backend-no-driver (test_pytest_models_tosa) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-backend-no-driver (test_pytest_ops_no_target) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-backend-no-driver (test_pytest_ops_tosa) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-backend-no-driver (test_run_tosa) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-arm-cortex-m-size-test (bare_metal) / linux-job (gh)
    Error: Input required and not supplied: path
  • pull / test-arm-cortex-m-size-test (zephyr-preset) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-binary-size-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B5A2:358D5F:85A25A6:1C262AF0:6A760BB2 and timestamp 2026-08-07 16:45:38 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-custom-ops-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B798:F25C:37BAAEE:BD4C352:6A760B2F and timestamp 2026-08-07 16:43:27 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-llama-runner-qnn-linux (fp32, qnn_16a16w, qnn) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CC30:37837C:8300DD0:1B7CBF24:6A760B84 and timestamp 2026-08-07 16:44:52 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-llama-runner-qnn-linux (fp32, qnn_8a8w, qnn) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B3E8:A8704:36FE3E5:BB0E486:6A760B75 and timestamp 2026-08-07 16:44:37 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add_mul, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID E702:6D448:834B732:1B9D6EB9:6A760B5D and timestamp 2026-08-07 16:44:13 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add_mul, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 87EA:2AB583:90C6DC9:1E698C05:6A760B62 and timestamp 2026-08-07 16:44:18 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID C2F4:28EBC1:4A7761F:FCA651C:6A760B75 and timestamp 2026-08-07 16:44:37 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (add, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CB94:957DF:4FA5958:110A0219:6A760B52 and timestamp 2026-08-07 16:44:02 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (emformer_transcribe, portable, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (emformer_transcribe, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 83B4:3FD34A:864D9D5:1C0CA3F0:6A760B5C and timestamp 2026-08-07 16:44:13 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (ic3, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID D922:36966A:36B5B74:BABE9FD:6A760B45 and timestamp 2026-08-07 16:43:49 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (ic3, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CA72:2ED0C1:4D21445:104EF054:6A760B66 and timestamp 2026-08-07 16:44:22 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (linear, portable, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (linear, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 8B64:157945:882A215:1C9637F8:6A760B50 and timestamp 2026-08-07 16:44:00 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (mobilebert, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID CF54:1677CB:8842365:1C76D99B:6A760B6F and timestamp 2026-08-07 16:44:31 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (mobilebert, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID D83E:381F55:881BD75:1CADDB76:6A760B71 and timestamp 2026-08-07 16:44:33 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (mv2, portable, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (mv2, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AA84:D4276:501F9C8:10EA32D9:6A760B37 and timestamp 2026-08-07 16:43:35 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (resnet18, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AA0A:1B80FE:80E622D:1B172D42:6A760B52 and timestamp 2026-08-07 16:44:02 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (resnet18, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux (resnet50, portable, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID B4B4:211490:85F4316:1C169E46:6A760B4E and timestamp 2026-08-07 16:43:58 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux (resnet50, xnnpack-quantization-delegation, linux.2xlarge) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 9992:3EBE49:88F199D:1CBF5B20:6A760B8A and timestamp 2026-08-07 16:44:58 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (mv3, portable, buck2, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID A9CE:2ABE79:4C6DDAA:10299B38:6A760B84 and timestamp 2026-08-07 16:44:52 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (mv3, portable, cmake, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID BF6A:4723F:36B7699:BA082FF:6A760B86 and timestamp 2026-08-07 16:44:54 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (mv3, xnnpack-quantization-delegation, buck2, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux-basic (mv3, xnnpack-quantization-delegation, cmake, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux-basic (vit, portable, buck2, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AA0E:3E7121:3585EA5:B62E226:6A760B78 and timestamp 2026-08-07 16:44:40 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (vit, portable, cmake, linux.2xlarge, executorch-ubuntu-22.04-clang12) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID C9E2:3BE4F5:82CA081:1B7F9ED6:6A760B9D and timestamp 2026-08-07 16:45:17 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-models-linux-basic (vit, xnnpack-quantization-delegation, buck2, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-models-linux-basic (vit, xnnpack-quantization-delegation, cmake, linux.2xlarge, executorch-u... / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID D634:A8704:3707F6E:BB2F5D0:6A760B92 and timestamp 2026-08-07 16:45:06 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-moshi-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 9904:36BA82:39BFDAA:C47C542:6A760B45 and timestamp 2026-08-07 16:43:49 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-openvino-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID EA74:56AB6:82F1981:1B7814B3:6A760B8D and timestamp 2026-08-07 16:45:01 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-qnn-buck-build-linux / linux-job (gh)
    ##[error]An error occurred trying to start process '/usr/bin/bash' with working directory '/home/ec2-user/actions-runner/_work/executorch/executorch/pytorch/executorch'. No such file or directory
  • pull / test-qnn-models-linux (dl3) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AB08:4723F:369EC95:B9B468A:6A760B3F and timestamp 2026-08-07 16:43:43 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-qnn-models-linux (mv2) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID AB9E:1BCB3:382B40F:BEF5392:6A760B38 and timestamp 2026-08-07 16:43:36 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-qnn-models-linux (mv3) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 8236:1DE33:4E0A78A:109376F6:6A760B43 and timestamp 2026-08-07 16:43:47 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-quantized-aot-lib-linux / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID E4B2:1677CB:88532AD:1C7A6BD9:6A760B8F and timestamp 2026-08-07 16:45:03 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-setup-linux-gcc / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 9D18:359DC9:4B34578:FEF937B:6A760B43 and timestamp 2026-08-07 16:43:47 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-sqnr-static-llm-qnn-linux (smollm2_135m) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 85E2:3B7C4D:3A51F9D:C610992:6A760B48 and timestamp 2026-08-07 16:43:52 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting
  • pull / test-static-llama-qnn-linux (stories_110m) / linux-job (gh)
    ##[error]API rate limit exceeded for installation. If you reach out to GitHub Support for help, please include the request ID 949A:36966A:36B7B0E:BAC553A:6A760B4A and timestamp 2026-08-07 16:43:54 UTC. For more on scraping GitHub and how it may affect your rights, please review our Terms of Service (https://docs.github.com/en/site-policy/github-terms/github-terms-of-service) - https://docs.github.com/en/rest/using-the-rest-api/getting-started-with-the-rest-api#rate-limiting

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-clameta-claBot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jul 22, 2026
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
 fbsource master
[ghstack-poisoned]
@meta-codesync
meta-codesyncBot merged commit c1b38bc into gh/JCNTH/108/baseAug 7, 2026
132 of 182 checks passed
@meta-codesync
meta-codesyncBot deleted the gh/JCNTH/108/head branch August 7, 2026 17:20
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
JCNTH added a commit that referenced this pull request Aug 7, 2026
Pull Request resolved: #21132
Materializing the full attention-weight matrix for Llama prefill is memory- and
bandwidth-heavy and does not scale to longer sequences. This adds a single-pass
online-softmax K16 causal-attention kernel for the exact Llama Hq32/Hkv8/G4/D64
geometry with fp16 KV, guarded by explicit adapter limits, so accepted prefill
shapes through S512 never materialize attention weights. S1 keeps the existing
FlashDecoding path, and unsupported geometry, storage, or capabilities fall back
to the materialized implementation. No Vulkan analogue (WebGPU-specific): the
Vulkan backend has only a materialized attention (compute the weights, then a
separate multi-pass softmax), with no online-softmax kernel.
Key changes:
- runtime/ops/sdpa/streaming_attention_k16_causal_bound.wgsl (+ generated
header): the online-softmax single-pass causal kernel.
- Sdpa.cpp: exact Llama geometry and capability guards, S512 transition
handling, and fallback to the materialized route.
ghstack-source-id: 411961446
@exported-using-ghexport
Differential Revision: [D113171738](https://our.internmc.facebook.com/intern/diff/D113171738/)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA SignedThis label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@JCNTH@psiddh