Luo Fuli: The R&D challenges of MiMo-V2.6 have surpassed those of DeepSeek R1 that I have been involved in before.
2026-09-22 11:57:44
According to CoinMeta, on-chain analyst Luo Fuli stated that the development challenges of MiMo-V2.6 have surpassed those of DeepSeek R1 in which she participated. After the release of MiMo-V2.6, she further explained the innovations and engineering challenges of this round of reinforcement learning. Luo Fuli mentioned that MiMo uses both mixrl and mopd simultaneously, while mixrl combines verifiable tasks such as code, general agent, vision, and network security in the same round of reinforcement learning training. mopd, on the other hand, deals with extremely long, difficult-to-verify, or subjectively rewarding tasks; these are trained separately before being integrated back into the main model. She pointed out that games and 3D tasks have longer execution times and it is difficult to automatically determine right from wrong, therefore MiMo trains such tasks separately before integrating their capabilities through mopd.
Bullish 0
Bearish 0
Source:Internet
This content is for market information only and does not constitute investment advice.