Xiaomi MiMo lead Luofuli said after the release of MiMo-V2.6 that the research innovation and engineering difficulty behind the model exceeded DeepSeek R1, a project she previously worked on. She also explained why MiMo uses both MixRL and MOPD in its reinforcement learning setup. According to her description, MixRL trains verifiable tasks such as coding, general agent work, vision, and cybersecurity together in the same reinforcement learning round. MOPD is used for tasks that are very long, hard to verify, or judged with more subjective rewards. Those tasks are trained separately first, and the resulting capabilities are then merged back into the main model. Luofuli added that tasks such as games and 3D runs take much longer to execute and are difficult to score automatically. If they are placed in the same RL round as coding and similar tasks, training slows down noticeably. MiMo therefore trains those tasks separately and uses MOPD to fold the learned abilities back into the core model.
Xiaomi MiMo lead Luofuli said after the release of MiMo-V2.6 that the model’s research innovation and engineering challenges had, in her view, surpassed DeepSeek R1, a project she previously participated in.
She also explained why MiMo uses both MixRL and MOPD. MixRL, she said, combines verifiable tasks including coding, general agent work, vision, and cybersecurity into the same round of reinforcement learning training.
MOPD is used for tasks that are ultra-long, difficult to verify, or scored with relatively subjective rewards. Those tasks are trained separately first, and the capabilities learned there are later integrated back into the main model.
Luofuli said game and 3D tasks take too long to run in a single pass and are also hard to evaluate automatically. If they are trained in the same RL round as coding and similar tasks, the overall training process slows down significantly. For that reason, MiMo trains those tasks separately and then uses MOPD to merge the learned capabilities back into the main model.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.