NVIDIA AI has announced an expansion of its Megatron Core framework, adding support for advanced optimizers including Muon, as well as research-oriented optimizers such as MOP and REKLS. The move is aimed at improving the efficiency of large-scale model training and giving developers more options when optimizing modern AI workloads.
Focused on large-model training
According to the announcement, the update is designed for training models at the scale of Kimi K2 and Qwen3 30B. NVIDIA emphasized that standard data parallelism alone is no longer sufficient for efficient training at this level. As models continue to grow in size and complexity, more specialized optimization techniques are becoming increasingly important across the AI infrastructure stack.
Broader optimizer options for developers
By integrating Muon alongside MOP and REKLS, Megatron Core expands the set of training approaches available to researchers and engineers. In practice, this could allow teams to test different optimization strategies depending on the model architecture and training objective, with potential implications for convergence behavior, stability, and compute utilization.
No benchmark details yet
NVIDIA did not provide specific benchmark results or implementation details in the announcement. That means the real-world impact of these new optimizer integrations—including speed gains, training efficiency improvements, or deployment trade-offs—remains unclear for now. Even so, the update underlines a broader industry trend: low-level training optimization is becoming a key battleground as competition around large AI models intensifies.

