AMD on July 25 announced Instella-MoE, a fully open-source mixture-of-experts model with 16 billion total parameters and 2.8 billion active parameters per token. The company said the model delivers leading performance among open-source language models of a similar size.
According to AMD, Instella-MoE was trained from scratch entirely on AMD Instinct MI300X and MI325X GPUs using the ROCm software stack. It also uses architectural changes including Gated Multi-head Latent Attention, or Gated MLA, and FarSkip-Collective to improve training and inference efficiency.
Benchmark results and model specifications
AMD said performance tests showed the Instella-MoE-16B-A3B base model reached an average score of 76.7, putting it at the front of the fully open-source field and ahead of models including SmolLM3-3B and OLMo-3-7B. The company added that the model can compete with larger systems while activating only 2.8 billion parameters.
Instella-MoE supports 64K-token long-context processing. AMD said the model has completed a full training pipeline covering pretraining, mid-training, long-context extension, supervised fine-tuning, direct preference optimization, and reinforcement learning.
Weights, checkpoints, and code released
AMD is also releasing all Instella-MoE model weights, training configurations, data mixture details, intermediate checkpoints, and inference code. The company said the release is intended to support open AI model research and reproducibility.
AMD said Instella-MoE demonstrates the ability to train large-scale MoE models on AMD hardware and an open software ecosystem. The company added that it will keep working on larger open-source language models with stronger reasoning capabilities and higher efficiency.

