Uh oh!
There was an error while loading. Please reload this page.
forked from ggml-org/llama.cpp
- Notifications
You must be signed in to change notification settings - Fork 8
Pull requests: AMD-Ecosystem/llama.cpp
Author
Uh oh!
There was an error while loading. Please reload this page.
Label
Uh oh!
There was an error while loading. Please reload this page.
Projects
Uh oh!
There was an error while loading. Please reload this page.
Milestones
Uh oh!
There was an error while loading. Please reload this page.
Reviews
Assignee
Assigned to nobodyLoading
Uh oh!
There was an error while loading. Please reload this page.
Sort
Pull requests list
CUDA: dispatch the chunked gated_delta_net in CU mode when it fills the device
#79
opened Jul 30, 2026 by
roberteg16Loading…
llama: let the backend pick n_ubatch, and default to 2048 on RDNA3.5
#77
opened Jul 29, 2026 by
roberteg16
•
Draft
3 tasks done
tools: add mmq-tune, an MMQ tile-width autotuning harness
#76
opened Jul 29, 2026 by
roberteg16
•
Draft
3 tasks done
ggml-cuda: f32 tall-skinny GEMM kernel for gfx1151 (RDNA3.5)
#74
opened Jul 28, 2026 by
roberteg16
•
Draft
feat(cuda): fuse activations and residual add into mmv f/q epilogues
#67
opened Jul 24, 2026 by
roberteg16
•
Draft
tests: add MoE MMQ benchmark with routing-distribution generator
#62
opened Jul 20, 2026 by
roberteg16
•
Draft
3 tasks
gfx1151 (Strix Halo): fuse attn_k+v into single MMVQ dispatch
#59
opened Jul 17, 2026 by
jeffli-xilinxLoading…
7 tasks done
ggml-cuda: GEMM weight row padding + one-time K-padded f16 dequant for prefill
#57
opened Jul 17, 2026 by
roberteg16
•
Draft
4 of 5 tasks
ProTip!
What’s not been updated in a month: updated:<2026-07-13.