Xiaomi’s MiMo halts two public RL training runs after spending $3.47 million

Xiaomi’s MiMo halts two public RL training runs after spending $3.47 million

N
News Editor
2026-09-21 09:55:57
Xiaomi has stopped both publicly streamed reinforcement learning training runs for its MiMo model, with Pro and Flash each completing 30 training steps as their last full run. Combined spending reached $3.4747 million, including $2.6207 million for Pro and $854,000 for Flash. The gains were notable: on DeepSWE v1.1, Pro rose from 58.41 at step 1 to 72.57, while Flash climbed from 48.67 to 65.68. Still, both curves saw clear pullbacks during training rather than moving up in a straight line. Other benchmark panels also improved, with Pro’s internal coding evaluation increasing from 57.54 to 65.43 and AutomationBench rising from 45.2 to 53.1. The process also ran into operational issues. Pro hit a GPU OOM caused by uneven expert load, and both the training cluster and evaluator experienced disconnects and restarts. Flash had to restart from step 15 because of an infrastructure error. Xiaomi also later filtered out tasks that had become too easy for Pro and removed the cyber dataset from later Pro training after abnormal rollout patterns appeared.

Xiaomi has halted both publicly streamed reinforcement learning training runs for MiMo. Pro and Flash each reached 30 training steps as their last fully completed run.

$3.4747 million spent across the two runs

The two runs cost a combined $3.4747 million. Of that total, Pro accounted for $2.6207 million and Flash for $854,000.

Scores improved, but the path was not linear

The spending produced visible gains. On DeepSWE v1.1, Pro improved from 58.41 at step 1 to 72.57, a gain of 14.16 points. Flash rose from 48.67 to 65.68, up 17.01 points.

At the same time, both training curves showed clear mid-run pullbacks, rather than a steady climb as training continued.

Two other benchmark panels also moved higher. Pro’s internal code evaluation rose from 57.54 to 65.43, while AutomationBench increased from 45.2 to 53.1.

Training was disrupted by several issues

The process was not smooth. Pro at one point triggered a GPU OOM because of uneven expert load, and the training cluster and evaluator also went through disconnections and restarts. Flash had to rerun from step 15 after an infrastructure error.

Later in the process, Xiaomi filtered out tasks that had become too easy for Pro. It also removed the cyber dataset from subsequent Pro training after abnormal rollout patterns appeared.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.