Skip to content

Pull requests: NVIDIA/TransformerEngine

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobodyLoading
Sort

Pull requests list

mxfp8: add swizzled-scale fast path for cast-only quantization 2.19 community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3338 opened Aug 10, 2026 by WanZzzzzzContributorLoading…
13 tasks
[common] Improved performance of Group MXFP8 kernels
#3337 opened Aug 10, 2026 by Oleg-GoncharovCollaboratorLoading…
6 of 13 tasks
[PyTorch] Fine-grained recipe docs documentation Improvements or additions to documentation
#3336 opened Aug 10, 2026 by negvetCollaboratorLoading…
13 tasks
Add a backwards linear function to be used with the fused mla q up-proj 2.19 community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3330 opened Aug 7, 2026 by chaseblockContributorLoading…
13 tasks
[PyTorch] Fix deferred initialization in fusible ops
#3327 opened Aug 7, 2026 by deneraCollaboratorLoading…
8 of 13 tasks
Refactor GroupedLinear quantization dispatch
#3326 opened Aug 7, 2026 by negvetCollaboratorLoading…
13 tasks
Prototype NVFP4 with FP8 UE5M3 block scales 2.19 enhancement New feature or request
#3325 opened Aug 7, 2026 by timmoon10Member Draft
5 of 13 tasks
[PyTorch] Enable NVFP4 row-scaled (per-token) backward for GroupedLinear community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3324 opened Aug 7, 2026 by cael-lingContributorLoading…
1 of 13 tasks
[PyTorch] Advance FusedAdam step counter for empty param groups community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3318 opened Aug 5, 2026 by adityasingh2400Loading…
[Common][PyTorch] Fuse the RHT into grouped NVFP4 quantize on non-SM100 architectures community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3317 opened Aug 5, 2026 by davidkny22ContributorLoading…
6 of 13 tasks
[Common/PyTorch] Grouped weighted-SwiGLU MXFP8 kernel community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3315 opened Aug 4, 2026 by cael-lingContributorLoading…
3 of 13 tasks
Log when thd with dropout falls to the composite cuDNN engine community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3313 opened Aug 4, 2026 by bzantiumLoading…
[CI] Publish GB200 aarch64 wheel artifacts community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3311 opened Aug 4, 2026 by bvolpatoLoading…
5 of 13 tasks
nvrtc MXFP8 kernels community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3302 opened Aug 3, 2026 by CarlosGomes98Contributor Draft
13 tasks
NVRTC NVFP4 quantization kernels community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3301 opened Aug 3, 2026 by CarlosGomes98ContributorLoading…
8 of 13 tasks
Add NVFP4 RHT Support for SM120 and SM121 community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3300 opened Aug 3, 2026 by new-TonyWangLoading…
[PyTorch] Scope the quantized-param caching flag to its own graph capture community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3298 opened Aug 1, 2026 by xiuhu17ContributorLoading…
6 tasks done
[Pytorch][Common] Option to Disable 2nd level Scale in NVFP4
#3297 opened Jul 31, 2026 by vthumbe1503CollaboratorLoading…
13 tasks
[Common] Fix pointer arithmatic to generate correct LDS/STS instructions
#3291 opened Jul 30, 2026 by kainzhongCollaboratorLoading…
8 of 13 tasks
Fix cuDNN SDPA score_mod cache collision, sm_12x arch gates, per-device plan cache community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3289 opened Jul 30, 2026 by YangXu1990uiucLoading…
[PyTorch][torch.compile] Support for DotProductAttention on flash and unfused backends
#3286 opened Jul 30, 2026 by pggPLCollaboratorLoading…
9 tasks done
ProTip! Type gp on any issue or pull request to go back to the pull request listing page.