NaiveAI, founded by Tsinghua University associate professor Dai Jifeng, has released Naive-N0.5-Flash and opened its model weights under the MIT license. The model is based on Xiaomi’s MiMo-V2.5 Base and keeps the original mixture-of-experts, or MoE, architecture while changing the attention design. According to the team, global attention was replaced with sliding-window attention and DeepSeek sparse attention, followed by continued training on 3.25 trillion tokens. The updated model natively supports a 1 million-token context window, which the company said cuts compute costs for very long text processing. NaiveAI also said AI systems were directly involved in development, handling coding, experiments, result analysis and later optimization, while human researchers set direction and made key decisions. Using its NaiveRT inference system as an example, the team said it completed 151 optimization rounds over six days, with 63 adopted. The company also disclosed a peak inference speed of 2,122 tok/s under a specific setup, while saying standard mode runs at about 50 tok/s per user.
NaiveAI, founded by Tsinghua University associate professor Dai Jifeng, has released Naive-N0.5-Flash and opened the model weights under the MIT license. The model is built on Xiaomi’s MiMo-V2.5 Base.
On the technical side, NaiveAI kept the original mixture-of-experts (MoE) architecture and focused its changes on the attention mechanism. The team replaced global attention with sliding-window attention and DeepSeek sparse attention, then continued training on 3.25 trillion tokens. After the modification, the model natively supports a 1 million-token context window, which reduces compute required for handling very long text.
NaiveAI also said AI was directly involved in the research and development process. It handled coding, ran experiments, analyzed results and carried out later optimization, while researchers were responsible for setting direction and making key decisions.
Using the NaiveRT inference system as an example, the team said it completed 151 optimization experiments over six days, and 63 of those rounds were adopted. The company disclosed a peak inference speed of 2,122 tok/s, but said the result came under specific test conditions: eight GPUs were used, thinking mode was turned off, input processing time was excluded, and the figure reflected the best one-second result out of 41 requests. In standard mode, the official figure was about 50 tok/s per user.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.