Xiaomi has launched the MiMo-V2.5 family, merging text, image, audio, and video processing into a single system. The Pro version delivers top-tier benchmark scores, resolving 57.2% of tasks on SWE-bench Pro — more than double the typical average of around 25%. The company claims it now matches Claude Opus 4.6 and GPT-5.4 across most coding and agent benchmarks.
Unified Multimodal Architecture
The earlier MiMo-V2-Pro handled text and code only; multimodal features required a separate, weaker model. V2.5 eliminates that divide. Users can upload a photo for analysis, step through a video tutorial, or extract action items from a recorded meeting — all within the same interface. Xiaomi calls the Pro version “a major leap from MiMo-V2-Pro in general agentic capabilities, complex software engineering, and long-horizon tasks.”
Benchmark Performance and Positioning
On SWE-bench Pro, the Pro model scores 57.2%, far above the 25% average. Results on τ3-bench and ClawEval place it near leading models. However, on the tougher reasoning test Humanity's Last Exam, it scores 48.0% versus GPT-5.4's 58.7%. The base model targets everyday use, running at 100-150 tokens per second and priced at $0.40 per million input tokens and $2.00 per million output. The Pro version runs at 60-80 tokens per second and costs $1.00 input, $3.00 output. Both support a 1M-token context window, handling large datasets or extended conversations.
Token Efficiency and Pricing Strategy
Efficiency stands out as a key differentiator. Xiaomi says the Pro model uses 42% fewer tokens than Kimi K2.6 for similar results, while the base model consumes nearly half the tokens of comparable systems. For developers operating at scale, lower token usage directly cuts costs. The launch also brings updated pricing: Xiaomi removed extra charges for using the full 1M-token context window and reset user credits. Models are available via the MiMo API, while AI Studio access remains limited. Free access through the Hermes agentic AI tool has helped boost early adoption.
Rapid Iteration and Ecosystem Growth
The release cadence has been steady. Xiaomi introduced MiMo-V2-Flash in late 2025, followed by V2-Pro, Omni, and TTS models in March, and now the V2.5 series. Lei Jun announced an $8.7B AI investment over three years, and deployment has visibly accelerated. Platform data confirms momentum: Xiaomi models accounted for roughly 21% of OpenRouter traffic as of early April, with usage surging over 42% in a single week. That growth followed a period of free access through the Hermes tool, expanding visibility. Xiaomi said future models will focus on “deeper reasoning, tighter tool integration, and richer real-world grounding,” hinting that another release could come sooner than expected.

