Uh oh!
There was an error while loading. Please reload this page.
- Notifications
You must be signed in to change notification settings - Fork 799
Pull requests: NVIDIA/TransformerEngine
Author
Uh oh!
There was an error while loading. Please reload this page.
Label
Uh oh!
There was an error while loading. Please reload this page.
Projects
Uh oh!
There was an error while loading. Please reload this page.
Milestones
Uh oh!
There was an error while loading. Please reload this page.
Reviews
Assignee
Assigned to nobodyLoading
Uh oh!
There was an error while loading. Please reload this page.
Sort
Pull requests list
mxfp8: add swizzled-scale fast path for cast-only quantization
2.19
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3338
opened Aug 10, 2026 by
WanZzzzzzContributorLoading…
13 tasks
[common] Improved performance of Group MXFP8 kernels
#3337
opened Aug 10, 2026 by
Oleg-GoncharovCollaboratorLoading…
6 of 13 tasks
[PyTorch] Fine-grained recipe docs
documentation
Improvements or additions to documentation
#3336
opened Aug 10, 2026 by
negvetCollaboratorLoading…
13 tasks
Add a backwards linear function to be used with the fused mla q up-proj
2.19
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3330
opened Aug 7, 2026 by
chaseblockContributorLoading…
13 tasks
[PyTorch] Fix deferred initialization in fusible ops
#3327
opened Aug 7, 2026 by
deneraCollaboratorLoading…
8 of 13 tasks
Refactor GroupedLinear quantization dispatch
#3326
opened Aug 7, 2026 by
negvetCollaboratorLoading…
13 tasks
[PyTorch] Enable NVFP4 row-scaled (per-token) backward for GroupedLinear
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3324
opened Aug 7, 2026 by
cael-lingContributorLoading…
1 of 13 tasks
[Pytorch] Enable TE Op to consume extra_outputs from a previously run Op in TE Sequential
#3320
opened Aug 5, 2026 by
vthumbe1503CollaboratorLoading…
13 tasks
[PyTorch] Advance FusedAdam step counter for empty param groups
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3318
opened Aug 5, 2026 by
adityasingh2400Loading…
[Common][PyTorch] Fuse the RHT into grouped NVFP4 quantize on non-SM100 architectures
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3317
opened Aug 5, 2026 by
davidkny22ContributorLoading…
6 of 13 tasks
[Common/PyTorch] Grouped weighted-SwiGLU MXFP8 kernel
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3315
opened Aug 4, 2026 by
cael-lingContributorLoading…
3 of 13 tasks
Log when thd with dropout falls to the composite cuDNN engine
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3313
opened Aug 4, 2026 by
bzantiumLoading…
[CI] Publish GB200 aarch64 wheel artifacts
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3311
opened Aug 4, 2026 by
bvolpatoLoading…
5 of 13 tasks
nvrtc MXFP8 kernels
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3302
opened Aug 3, 2026 by
CarlosGomes98Contributor
•
Draft
13 tasks
NVRTC NVFP4 quantization kernels
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3301
opened Aug 3, 2026 by
CarlosGomes98ContributorLoading…
8 of 13 tasks
Add NVFP4 RHT Support for SM120 and SM121
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3300
opened Aug 3, 2026 by
new-TonyWangLoading…
[PyTorch] Scope the quantized-param caching flag to its own graph capture
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3298
opened Aug 1, 2026 by
xiuhu17ContributorLoading…
6 tasks done
[Pytorch][Common] Option to Disable 2nd level Scale in NVFP4
#3297
opened Jul 31, 2026 by
vthumbe1503CollaboratorLoading…
13 tasks
[PTQ] Store FP32 global scaling factors (absmax or scale_inv) for all quantized activations and weights.
org-contribution
#3296
opened Jul 31, 2026 by
cspadesMemberLoading…
7 of 13 tasks
[Common] Fix pointer arithmatic to generate correct LDS/STS instructions
#3291
opened Jul 30, 2026 by
kainzhongCollaboratorLoading…
8 of 13 tasks
Fix cuDNN SDPA score_mod cache collision, sm_12x arch gates, per-device plan cache
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3289
opened Jul 30, 2026 by
YangXu1990uiucLoading…
[PyTorch][torch.compile] Support for DotProductAttention on flash and unfused backends
#3286
opened Jul 30, 2026 by
pggPLCollaboratorLoading…
9 tasks done
PreviousNext
ProTip!
Type gp on any issue or pull request to go back to the pull request listing page.