Kimi open-sources PerceptionBench as no model tops 60% accuracy in visual perception test

Kimi open-sources PerceptionBench as no model tops 60% accuracy in visual perception test

N
News Editor
2026-07-28 16:02:00
Kimi has released PerceptionBench, an open-source benchmark designed to measure visual perception in multimodal large language models by breaking the task into 10 atomic capabilities. The benchmark covers areas including visual relations, counting, attributes, depth and 3D, localization, comparison, fine-grained recognition, context integration, OCR, and hallucination detection. According to PANews, the dataset was built from model failure cases collected across 42 existing evaluation sets and contains 3,000 manually verified questions. Each question is designed to test only one visual skill and does not require reasoning or external knowledge. Results across 16 leading multimodal models showed that none achieved an overall accuracy above 60%. GPT-5.6-Sol ranked first with 59.7%, followed by Kimi K3 at 58.5%, Claude-Fable-5 at 57.2%, Gemini-3.1-Pro at 56.2%, and GPT-5.5 at 55.8%. The report said hallucination remained the weakest area across models, indicating that core visual perception performance still has significant room for improvement.
KimiPerceptionBenchmultimodal modelsvisual perceptionAI benchmarkGPT-5.6-Soltechnology

Kimi has open-sourced PerceptionBench, a benchmark for evaluating visual perception in multimodal large language models by splitting the task into 10 atomic capabilities.

How the benchmark is structured

PerceptionBench covers visual relations, counting, attributes, depth and 3D, localization, comparison, fine-grained recognition, context integration, OCR, and hallucination detection.

The benchmark was built from model failure cases found in 42 existing evaluation datasets and includes 3,000 manually verified questions. Each question tests a single visual capability and does not require reasoning or external knowledge.

Results across 16 models

The evaluation showed that none of the 16 leading multimodal models posted an overall accuracy above 60%. GPT-5.6-Sol ranked first at 59.7%, followed by Kimi K3 at 58.5%, Claude-Fable-5 at 57.2%, Gemini-3.1-Pro at 56.2%, and GPT-5.5 at 55.8%.

What the report found

The report said hallucination remained the weakest capability across models, leaving substantial room for improvement in overall visual perception.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.