-
Notifications
You must be signed in to change notification settings - Fork 504
All issues
Issue creation is restricted in this repository
- #3754 · anwithk opened
on May 8, 2026
Issues
is:issue state:open
is:issue state:open
Search results
[bug] gpt_step does not pass loss_mask to the model, so the MTP loss (and its gradient) is computed over padding
area:trainingTraining loop, callbacks, and runtime integrationTraining loop, callbacks, and runtime integrationbugSomething isn't workingSomething isn't workingneeds-triageNew item needs classification and ownershipNew item needs classification and ownershipStatus: Open.#6103 In NVIDIA-NeMo/Megatron-Bridge;[bug] SafeTensorsStateSource._resolve_path ignores the requested revision; offline caches pinned to a commit sha fail with LocalEntryNotFoundError
area:ckptCheckpoint conversion, loading, export, and save pathsCheckpoint conversion, loading, export, and save pathsbugSomething isn't workingSomething isn't workingneeds-triageNew item needs classification and ownershipNew item needs classification and ownershipStatus: Open.#6104 In NVIDIA-NeMo/Megatron-Bridge;[model] Add support for Qwen3.8-Flash-Next
area:modelModel implementations and HF bridge logicModel implementations and HF bridge logicfeatureNew capabilities, enhancements, or enablement workNew capabilities, enhancements, or enablement workneeds-triageNew item needs classification and ownershipNew item needs classification and ownershipStatus: Open.#6060 In NVIDIA-NeMo/Megatron-Bridge;[feature] Enable model-specific custom TaskEncoders in the Energon data path
area:dataDataset builders, preprocessing, and samplersDataset builders, preprocessing, and samplersfeatureNew capabilities, enhancements, or enablement workNew capabilities, enhancements, or enablement workwaiting-on-customerWaiting on the original author to respondWaiting on the original author to respondStatus: Open.#6048 In NVIDIA-NeMo/Megatron-Bridge;[bug] NemotronH-56B 256-GPU BF16 perf recipes keep the 64-GPU global_batch_size (192), so config validation fails at DP=128
area:perfPerformance optimizations and benchmarkingPerformance optimizations and benchmarkingbugSomething isn't workingSomething isn't workingneeds-triageNew item needs classification and ownershipNew item needs classification and ownershipStatus: Open.#6024 In NVIDIA-NeMo/Megatron-Bridge;moe_a2a_overlap=false cannot disable A2A overlap when the recipe enables it
area:perfPerformance optimizations and benchmarkingPerformance optimizations and benchmarkingbugSomething isn't workingSomething isn't workingneeds-triageNew item needs classification and ownershipNew item needs classification and ownershipStatus: Open.#6006 In NVIDIA-NeMo/Megatron-Bridge;[ckpt] GPU export overwrites GLM-5.2
max_position_embeddings(1048576 → training seq_length 8192) in the exportedconfig.json(related: #5153)area:ckptCheckpoint conversion, loading, export, and save pathsCheckpoint conversion, loading, export, and save pathsbugSomething isn't workingSomething isn't workingneeds-triageNew item needs classification and ownershipNew item needs classification and ownershipStatus: Open.#5994 In NVIDIA-NeMo/Megatron-Bridge;nemo:26.08 aarch64 image: PyTorch built against cuDNN 9.23.0 but only libcudnn 9.21.1 is installed —
torch.backends.cudnn.version()raisesarea:buildDependencies, packaging, images, and environment setupDependencies, packaging, images, and environment setupbugSomething isn't workingSomething isn't workingneeds-triageNew item needs classification and ownershipNew item needs classification and ownershipStatus: Open.#5993 In NVIDIA-NeMo/Megatron-Bridge;[training] flex dispatcher downgrades an explicitly requested
deepeptoalltoallon GB200 (warning only, no error) — clarify GB200 policy or fail fastarea:trainingTraining loop, callbacks, and runtime integrationTraining loop, callbacks, and runtime integrationbugSomething isn't workingSomething isn't workingneeds-triageNew item needs classification and ownershipNew item needs classification and ownershipStatus: Open.#5992 In NVIDIA-NeMo/Megatron-Bridge;[bug] temporary_distributed_context has a singleton TCP rendezvous race
area:ckptCheckpoint conversion, loading, export, and save pathsCheckpoint conversion, loading, export, and save pathsbugSomething isn't workingSomething isn't workingneeds-triageNew item needs classification and ownershipNew item needs classification and ownershipStatus: Open.#5967 In NVIDIA-NeMo/Megatron-Bridge;[bug] Transformer Engine fused cross-entropy zeroes MTP acceptance metrics
area:trainingTraining loop, callbacks, and runtime integrationTraining loop, callbacks, and runtime integrationbugSomething isn't workingSomething isn't workingwaiting-on-maintainersWaiting on maintainers to respondWaiting on maintainers to respondStatus: Open.#5947 In NVIDIA-NeMo/Megatron-Bridge;[ALERT] Access to NVIDIA's Self-Hosted Runners is Expiring
ciCI, automation, test queue, or workflow infrastructure workCI, automation, test queue, or workflow infrastructure workneeds-triageNew item needs classification and ownershipNew item needs classification and ownershipStatus: Open.#5886 In NVIDIA-NeMo/Megatron-Bridge;