DeepSeek V4 Flash tops rankings, but posts only a 53.8% pass rate in independent test
DeepSeek V4 Flash reached the top of a ranking, but its showing in an independent evaluation was far weaker. According to Techub, citing CryptoBriefing, Composio tested the model on complex agent tasks and found it completed only 53.8% of them. The evaluation covered 30 multi-step workflows and included 240 runs across real-world tools such as Gmail, GitHub, and Slack. Results also showed wide performance gaps across agent frameworks. Pi Agent delivered the best outcome, completing 20 tasks. The report added that DeepSeek V4 Flash has an input token price as low as $0.14 per million, roughly one-tenth of comparable products, but said the high failure rate could create extra computing and debugging costs. The findings point to a gap between leaderboard standing and execution in tool-based, multi-step workflows.








