Xiaomi says MiMo-V2.6 repeated tool calls were caused by RL design, fixes issue for about $90,000

Xiaomi says MiMo-V2.6 repeated tool calls were caused by RL design, fixes issue for about $90,000

N
News Editor
2026-09-28 07:53:13
Xiaomi’s MiMo team said it has identified and fixed a repeated tool-calling issue that appeared after the launch of MiMo-V2.6. In some cases, the model would call the same or highly similar tools over and over, consuming context without making progress on the task. In OpenCode, the share of responses with repeated tool calls at one point reached 1.02% for Flash and 0.54% for Pro. According to the team, the problem came from reinforcement learning reward design. Training heavily rewarded whether the final task was completed correctly, but did not penalize inefficient behavior enough during the process. Under the original rule, penalties applied only when tool calls in a single round exceeded 32, which left repeated calls below that threshold unpunished. Xiaomi said this behavior became more pronounced as RL scaled up. Instead of a full retraining plan that would have required lowering the penalty threshold to 8, rerunning about 20 MixRL steps, and spending an estimated $2.31 million, Xiaomi trained a dedicated RL teacher to correct repeated calls. The fix used 12 steps and about 7,000 samples, then merged the capability back into Pro and Flash through MOPD. Xiaomi said the full repair cost about $90,000, roughly 4% of the full retraining option, while other major benchmarks were largely unchanged.

Xiaomi’s MiMo team has published a postmortem on the repeated tool-calling issue that surfaced after MiMo-V2.6 went live. The team said the model would sometimes call the same tool, or highly similar tools, again and again, consuming context while making no progress on the task.

In OpenCode, the share of responses that showed repeated tool calls at one point reached 1.02% for Flash and 0.54% for Pro.

The issue was traced to reinforcement learning reward design

Xiaomi said the root cause was in the reward design used for reinforcement learning. Training mainly rewarded whether the task was completed correctly in the end, while inefficient behavior during the process was not penalized enough.

Under the original rule, a penalty was triggered only when tool calls in a single round exceeded 32. Repeated calls below that threshold received no penalty at all. As RL scale increased, Xiaomi said, that bad habit was reinforced rather than reduced.

After replaying training checkpoints, the team found that in Flash, the share of abnormal samples with more than 10 calls in a single round rose from 11.1% at step 0 to 24.6% at step 20.

A full retraining route was estimated at $2.31 million

Xiaomi said the most direct fix would have been to lower the penalty threshold from 32 to 8 and rerun about 20 MixRL steps. That option, however, was expected to cost $2.31 million.

Xiaomi chose a targeted RL teacher instead

Rather than fully retraining the model, Xiaomi said it trained a dedicated RL teacher focused on correcting repeated calls. The run used 12 steps and about 7,000 samples, and the resulting capability was merged back into Pro and Flash through MOPD.

The company said the full repair cost about $90,000, or roughly 4% of the full retraining plan, while other major benchmarks were largely unchanged.

Updated weights and API are already live

Xiaomi said the fixed MiMo-V2.6-Pro-MOPD and Flash-MOPD weights have been released, and the API has already been switched to the new version without changing the call name.

The company also said it will reset the remaining quota for MiMo Desktop users in the current billing cycle.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.