Kimi Researcher Challenges NaiveAI Claim That 2,000 tok/s Shows "AI Building AI"

Kimi Researcher Challenges NaiveAI Claim That 2,000 tok/s Shows "AI Building AI"

N
News Editor
2026-09-28 08:35:57
NaiveAI, founded by Tsinghua University associate professor Dai Jifeng, has released its first model, N0.5-Flash, and promoted it with the claim that "AI participated in developing AI." The company said the model was involved in architecture design, code writing, and inference optimization, while reporting a peak single-stream inference speed of 2,122 tok/s. That framing was quickly challenged by Yang Xinyu, a researcher at Moonshot AI, the company behind Kimi. Yang said reaching 2,000 tok/s is entirely feasible on its own, but argued that NaiveAI used a weak comparison baseline. He noted that the company compared its result against SGLang at a little over 300 tok/s, while Xiaomi had previously pushed the larger MiMo-V2.5-Pro past 1,000 tok/s. Yang also disputed NaiveAI's characterization of its architectural changes. He said N0.5-Flash was modified from Xiaomi's MiMo-V2.5 Base by replacing global attention with DSA and ultimately using a DSA + GQA setup. In his view, that looks more like an engineering trade-off built on an existing base model, one that may also reduce GPU utilization, rather than a major architecture breakthrough. He added that NaiveAI has shown a model can be made faster through continued training and engineering optimization, but has not shown that "AI participation in R&D" produced a breakthrough beyond what human teams can achieve.

NaiveAI has released its first model, N0.5-Flash, and framed the launch around the idea that "AI participated in developing AI." The company said the model took part in architecture design, code writing, and inference optimization, and reported a peak single-stream inference speed of 2,122 tok/s.

That claim was later questioned in public by Yang Xinyu, a researcher at Moonshot AI, the company behind Kimi. Yang said 2,000 tok/s is entirely achievable by itself, but argued that NaiveAI chose a weak comparison baseline. He said the company used SGLang at a little over 300 tok/s as its reference point, while Xiaomi had previously pushed the larger MiMo-V2.5-Pro to more than 1,000 tok/s. With a lower baseline, the resulting performance improvement naturally looks larger.

Yang also rejected NaiveAI's description of the work as an architecture innovation. Based on the disclosed details, N0.5-Flash was adapted from Xiaomi's MiMo-V2.5 Base, replacing the original global attention mechanism with DSA and ending up with a DSA + GQA design.

In Yang's view, that is closer to an engineering trade-off built on an existing base model. He also said it would reduce GPU utilization, making it hard to call the change a major architecture breakthrough. His conclusion was that NaiveAI has shown a model can be sped up through continued training and engineering optimization, but so far has not shown that "AI participation in R&D" itself delivered a breakthrough beyond what human teams can achieve.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
100

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.