Skip to content

Pull requests: NVIDIA/TransformerEngine

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

Single group mxfp8 grouped mlp community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3267 opened Jul 27, 2026 by sraman-rgb Contributor Loading…
13 tasks
[PyTorch] Fix shape and size() for columnwise-only quantized tensors
#3266 opened Jul 27, 2026 by pggPL Collaborator Loading…
7 of 13 tasks
[PyTorch] Add architecture gate to NVFP4 split_quantize RHT path community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3265 opened Jul 27, 2026 by davidkny22 Loading…
6 of 13 tasks
[PyTorch] MXFP4 weight QAT on MXFP8 and FP8 block-scaling recipes community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3264 opened Jul 27, 2026 by xiuhu17 Contributor Draft
8 tasks done
[common] Fix UE8M0 code 0 (2^-127) and code 255 (NaN) expansion in ptx::exp2f community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3262 opened Jul 25, 2026 by xiuhu17 Contributor Loading…
8 tasks done
[Common][PyTorch] Fuse NVFP4 quantization launches for MoE experts community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3261 opened Jul 25, 2026 by liujshi Loading…
7 of 13 tasks
[Add] Add FlashAttention 4 context parallel support. community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3259 opened Jul 25, 2026 by Baibaifan Contributor Loading…
test(pytorch): cover QuantizedTensor view NotImplementedError community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3257 opened Jul 25, 2026 by andrewwhitecdw Contributor Loading…
Acquire the GIL in lazy init_extension community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3255 opened Jul 24, 2026 by xiuhu17 Contributor Loading…
5 tasks done
Use a temp directory for CMake build
#3254 opened Jul 24, 2026 by fheinecke Collaborator Draft
3 of 13 tasks
Add CUDA wheels as build-time requirements for build isolation
#3253 opened Jul 24, 2026 by fheinecke Collaborator Draft
3 of 13 tasks
Enable runtime resolution of CUDA header path for NVRTC
#3252 opened Jul 24, 2026 by fheinecke Collaborator Draft
3 of 13 tasks
[CI] Improve build time dependency resolution
#3251 opened Jul 24, 2026 by fheinecke Collaborator Draft
4 of 13 tasks
[CI] Pin JAX image to 2026-07-21
#3250 opened Jul 24, 2026 by fheinecke Collaborator Loading…
4 of 13 tasks
Opt in trusted FA CI checkpoints to pickle loading
#3247 opened Jul 24, 2026 by sudhakarsingh27 Member Loading…
6 of 9 tasks
[Pytorch] Add support for row-wise quanted input for grouped gemm
#3244 opened Jul 23, 2026 by YangFei1990 Collaborator Loading…
8 of 13 tasks
Activation + GroupedLinear Fusion for MOE and other MOE optimizations 2.18
#3238 opened Jul 22, 2026 by vthumbe1503 Collaborator Loading…
13 tasks
[All] Bump minimum supported cuDNN version to 9.11
#3236 opened Jul 22, 2026 by cyanguwa Collaborator Loading…
8 of 13 tasks
Add stream-ordered CP gradient return primitive community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3235 opened Jul 22, 2026 by foraxe Loading…
5 of 13 tasks
[PyTorch] Optionally release columnwise copy of frozen FP8 block-scaled weights after dgrad community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3233 opened Jul 22, 2026 by 1tex Loading…
8 of 13 tasks
Improve device-init grouped linear module with single grouped weight support community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3224 opened Jul 20, 2026 by zhongbozhu Collaborator Loading…
13 tasks
[Common] Experimental CuTeDSL MXFP4 backend community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3223 opened Jul 20, 2026 by janekb04 Collaborator Draft
2 of 13 tasks
Work around intermittent SM120 FP8 gradient corruption in RTC cast-transpose community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3215 opened Jul 15, 2026 by AlbertYang514 Loading…
Explore error correction for pertoken recipe community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3214 opened Jul 15, 2026 by YigongQin Contributor Draft
13 tasks
ProTip! Mix and match filters to narrow down what you’re looking for.