Artificial Analysis has completed an independent test of DeepSeek V4.1 Flash, giving the model a composite intelligence score of 40. That puts it ahead of DeepSeek’s own V4 Pro 0813, which scored 36, but still below Kimi K3 at 44 and GLM-5.3 at 45. By that third-party ranking, DeepSeek has not reclaimed the top spot among domestic models.
The price gap was far wider than the score gap. DeepSeek V4.1 Flash cost an average of $0.27 to complete one Intelligence Index task, while Kimi K3 and GLM-5.3 were both around $2, a difference of about 7x. Output speed also stood out at 197 tokens per second, compared with roughly 36 tokens per second for Kimi K3 and about 58 tokens per second for GLM-5.3.
Its strongest area remained agent performance. On AutomationBench-AA, V4.1 Flash scored 69%, matching GPT-6 Astra and beating GLM-5.3’s 62%. It also reached 84% on the long-context AA-LCR test. One drawback noted in the test was verbosity: in AA tasks, the model produced about 89,000 tokens per task on average, 62% more than V4 Pro, which reduced part of its pricing advantage.
Artificial Analysis has completed an independent evaluation of DeepSeek V4.1 Flash. The model posted a composite intelligence score of 40, ahead of DeepSeek’s own V4 Pro 0813 at 36 but still behind Kimi K3 at 44 and GLM-5.3 at 45. On that third-party table, DeepSeek has yet to retake the top position among domestic models.
Cost and speed stood out
The pricing gap was especially sharp. DeepSeek V4.1 Flash spent an average of $0.27 to complete one Intelligence Index task. Kimi K3 and GLM-5.3 were both around $2, leaving V4.1 Flash at roughly one-seventh of the cost.
Its output speed reached 197 tokens per second, well above Kimi K3 at about 36 tokens per second and GLM-5.3 at about 58 tokens per second.
Agent capability remained a strength
Agent performance continued to be the model’s strongest area. On AutomationBench-AA, V4.1 Flash scored 69%, matching GPT-6 Astra and surpassing GLM-5.3’s 62%. It also recorded 84% on the AA-LCR long-context test.
Verbosity cut into part of the price edge
The main weakness highlighted in the test was how much the model tends to generate. In AA testing, V4.1 Flash produced about 89,000 tokens per task on average, 62% more than V4 Pro. That higher output volume consumed part of its cost advantage.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.