JetBrains has released Mellum2.1, an open-source model built for coding agents and fast subagents. The company describes it as a 12 billion-parameter mixture-of-experts reasoning model with 2.5 billion active parameters per token. According to the announcement, Mellum2.1 outperformed its predecessor Mellum2 in 15 of 17 benchmarks and ranked ahead of Qwen3.5-9B on several coding evaluations, including LiveCodeBench v6. JetBrains published the model on Hugging Face under the Apache 2.0 license and said the architecture remains consistent with Mellum2. The main upgrade came from reinforcement learning training conducted in real software environments. JetBrains said the model is small enough to be self-hosted and can explore code repositories, edit files, and inspect its own changes. The company added that users can run it on their own GPUs through vLLM or SGLang. Its speed tests were based on a single NVIDIA H200. Mellum2.1 supports a 131,072-token context window and is aimed at agentic work, general reasoning assistants, and private self-hosted deployments, according to MarkTechPost as cited by Techub.
JetBrains has released Mellum2.1, an open-source model designed for coding agents and fast subagents. The company said Mellum2.1 is a 12 billion-parameter mixture-of-experts reasoning model with 2.5 billion active parameters per token.
On benchmarks, Mellum2.1 beat its predecessor Mellum2 in 15 of 17 tests. It also ranked ahead of Qwen3.5-9B on several coding benchmarks, including LiveCodeBench v6.
Mellum2.1 lands on Hugging Face under Apache 2.0
The model has been released on Hugging Face under the Apache 2.0 license. Its architecture remains the same as Mellum2, while the main upgrade came from reinforcement learning training carried out in real software environments, JetBrains said.
According to the company, that makes Mellum2.1 a small model that can be self-hosted and used to explore code repositories, edit files, and inspect its own changes.
Built for self-hosted deployment
JetBrains said Mellum2.1 can run on users' own GPUs through vLLM or SGLang. Its speed tests were based on one NVIDIA H200. The model has a 131,072-token context window and is aimed at three use cases: agentic work, general reasoning assistants, and private self-hosted deployment.
The update was cited by MarkTechPost and published in a Techub news brief.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.