Uh oh!
There was an error while loading. Please reload this page.
[ExecuTorch][WebGPU] Replace broad QKV fusion with BK64 kernel - #21131
Conversation
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21131
Note: Links to docs will display an error until the docs builds have been completed. ❌ 26 New Failures, 30 PendingAs of commit 72b6031 with merge base 28a7fac ( NEW FAILURES - The following jobs have failed:
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
This PR needs a |
fbsource master [ghstack-poisoned]
Uh oh!
There was an error while loading. Please reload this page.
Pull Request resolved: #21131 The previously landed broad QKV fusion applied too widely and did not match the BK64 schedule now used for the ordinary projections. This corrective diff replaces it with a capability- and geometry-qualified BK64 kernel that fuses the exact Llama Q/K/V projection triple at live M128, M508, and M512, packing the constant weights and scales once and scattering the result into three distinct planner-safe outputs. Outside the accepted shapes it switches atomically back to the ordinary Steel and bicol routes, and it deletes the obsolete broad shader and header so only one QKV path remains. Mirrors Vulkan xplat/executorch/backends/vulkan/runtime/graph/ops/glsl/q4gsw_linear_gemm__w_4x8.glsl for the per-projection GEMM; the three-output fusion itself is WebGPU-specific. Key changes: - Adds runtime/ops/quantized_linear/q4gsw_qkv_bk64.wgsl (+ generated header); removes the obsolete q4gsw_linear_gemm_qkv_fused shader and header. - WebGPUGraph.{cpp,h}: geometry and capability predicate, one-time constant packing, distinct Q/K/V outputs, and atomic fallback to Steel/bicol. ghstack-source-id: 411961443 @exported-using-ghexport Differential Revision: [D113171749](https://our.internmc.facebook.com/intern/diff/D113171749/)
Pull Request resolved: #21131 The previously landed broad QKV fusion applied too widely and did not match the BK64 schedule now used for the ordinary projections. This corrective diff replaces it with a capability- and geometry-qualified BK64 kernel that fuses the exact Llama Q/K/V projection triple at live M128, M508, and M512, packing the constant weights and scales once and scattering the result into three distinct planner-safe outputs. Outside the accepted shapes it switches atomically back to the ordinary Steel and bicol routes, and it deletes the obsolete broad shader and header so only one QKV path remains. Mirrors Vulkan xplat/executorch/backends/vulkan/runtime/graph/ops/glsl/q4gsw_linear_gemm__w_4x8.glsl for the per-projection GEMM; the three-output fusion itself is WebGPU-specific. Key changes: - Adds runtime/ops/quantized_linear/q4gsw_qkv_bk64.wgsl (+ generated header); removes the obsolete q4gsw_linear_gemm_qkv_fused shader and header. - WebGPUGraph.{cpp,h}: geometry and capability predicate, one-time constant packing, distinct Q/K/V outputs, and atomic fallback to Steel/bicol. ghstack-source-id: 411961443 @exported-using-ghexport Differential Revision: [D113171749](https://our.internmc.facebook.com/intern/diff/D113171749/)
Pull Request resolved: #21131 The previously landed broad QKV fusion applied too widely and did not match the BK64 schedule now used for the ordinary projections. This corrective diff replaces it with a capability- and geometry-qualified BK64 kernel that fuses the exact Llama Q/K/V projection triple at live M128, M508, and M512, packing the constant weights and scales once and scattering the result into three distinct planner-safe outputs. Outside the accepted shapes it switches atomically back to the ordinary Steel and bicol routes, and it deletes the obsolete broad shader and header so only one QKV path remains. Mirrors Vulkan xplat/executorch/backends/vulkan/runtime/graph/ops/glsl/q4gsw_linear_gemm__w_4x8.glsl for the per-projection GEMM; the three-output fusion itself is WebGPU-specific. Key changes: - Adds runtime/ops/quantized_linear/q4gsw_qkv_bk64.wgsl (+ generated header); removes the obsolete q4gsw_linear_gemm_qkv_fused shader and header. - WebGPUGraph.{cpp,h}: geometry and capability predicate, one-time constant packing, distinct Q/K/V outputs, and atomic fallback to Steel/bicol. ghstack-source-id: 411961443 @exported-using-ghexport Differential Revision: [D113171749](https://our.internmc.facebook.com/intern/diff/D113171749/)
Pull Request resolved: #21131 The previously landed broad QKV fusion applied too widely and did not match the BK64 schedule now used for the ordinary projections. This corrective diff replaces it with a capability- and geometry-qualified BK64 kernel that fuses the exact Llama Q/K/V projection triple at live M128, M508, and M512, packing the constant weights and scales once and scattering the result into three distinct planner-safe outputs. Outside the accepted shapes it switches atomically back to the ordinary Steel and bicol routes, and it deletes the obsolete broad shader and header so only one QKV path remains. Mirrors Vulkan xplat/executorch/backends/vulkan/runtime/graph/ops/glsl/q4gsw_linear_gemm__w_4x8.glsl for the per-projection GEMM; the three-output fusion itself is WebGPU-specific. Key changes: - Adds runtime/ops/quantized_linear/q4gsw_qkv_bk64.wgsl (+ generated header); removes the obsolete q4gsw_linear_gemm_qkv_fused shader and header. - WebGPUGraph.{cpp,h}: geometry and capability predicate, one-time constant packing, distinct Q/K/V outputs, and atomic fallback to Steel/bicol. ghstack-source-id: 411961443 @exported-using-ghexport Differential Revision: [D113171749](https://our.internmc.facebook.com/intern/diff/D113171749/)
Stack from ghstack (oldest at bottom):
The previously landed broad QKV fusion applied too widely and did not match the
BK64 schedule now used for the ordinary projections. This corrective diff
replaces it with a capability- and geometry-qualified BK64 kernel that fuses the
exact Llama Q/K/V projection triple at live M128, M508, and M512, packing the
constant weights and scales once and scattering the result into three distinct
planner-safe outputs. Outside the accepted shapes it switches atomically back
to the ordinary Steel and bicol routes, and it deletes the obsolete broad
shader and header so only one QKV path remains. Mirrors Vulkan
xplat/executorch/backends/vulkan/runtime/graph/ops/glsl/q4gsw_linear_gemm__w_4x8.glsl
for the per-projection GEMM; the three-output fusion itself is WebGPU-specific.
Key changes:
removes the obsolete q4gsw_linear_gemm_qkv_fused shader and header.
packing, distinct Q/K/V outputs, and atomic fallback to Steel/bicol.
@exported-using-ghexport
Differential Revision: D113171749
Differential Revision: D113171749