AMD unveils fully open-source Instella-MoE, a 16B-parameter model aimed at top-tier open models

AMD unveils fully open-source Instella-MoE, a 16B-parameter model aimed at top-tier open models

N
News Editor
2026-07-25 03:18:25
AMD on July 25 introduced Instella-MoE, a fully open-source mixture-of-experts language model with 16 billion total parameters and 2.8 billion active parameters per token. The company said the model was trained from scratch entirely on its own AMD Instinct MI300X and MI325X GPUs, using the ROCm software stack, and incorporates architectural designs including Gated Multi-head Latent Attention, or Gated MLA, and FarSkip-Collective to improve training and inference efficiency. According to AMD, the Instella-MoE-16B-A3B base model posted an average score of 76.7 in performance tests, placing it among the leading fully open-source models and ahead of models such as SmolLM3-3B and OLMo-3-7B. AMD also said the model can compete with larger systems while activating only 2.8 billion parameters. Instella-MoE supports 64K-token long-context processing and has gone through a full training pipeline that includes pretraining, mid-training, long-context extension, supervised fine-tuning, direct preference optimization, and reinforcement learning. AMD said it is releasing the full model weights, training configurations, data mix, intermediate checkpoints, and inference code.
AMDInstella-MoEOpen-source AIMoEROCmMI300XMI325XLanguage Models

AMD on July 25 announced Instella-MoE, a fully open-source mixture-of-experts model with 16 billion total parameters and 2.8 billion active parameters per token. The company said the model delivers leading performance among open-source language models of a similar size.

According to AMD, Instella-MoE was trained from scratch entirely on AMD Instinct MI300X and MI325X GPUs using the ROCm software stack. It also uses architectural changes including Gated Multi-head Latent Attention, or Gated MLA, and FarSkip-Collective to improve training and inference efficiency.

Benchmark results and model specifications

AMD said performance tests showed the Instella-MoE-16B-A3B base model reached an average score of 76.7, putting it at the front of the fully open-source field and ahead of models including SmolLM3-3B and OLMo-3-7B. The company added that the model can compete with larger systems while activating only 2.8 billion parameters.

Instella-MoE supports 64K-token long-context processing. AMD said the model has completed a full training pipeline covering pretraining, mid-training, long-context extension, supervised fine-tuning, direct preference optimization, and reinforcement learning.

Weights, checkpoints, and code released

AMD is also releasing all Instella-MoE model weights, training configurations, data mixture details, intermediate checkpoints, and inference code. The company said the release is intended to support open AI model research and reproducibility.

AMD said Instella-MoE demonstrates the ability to train large-scale MoE models on AMD hardware and an open software ecosystem. The company added that it will keep working on larger open-source language models with stronger reasoning capabilities and higher efficiency.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
100

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.