Scale Labs study finds 25 multimodal AI models still trail humans on everyday common-sense cues

Scale Labs study finds 25 multimodal AI models still trail humans on everyday common-sense cues

N
News Editor
2026-10-08 11:32:17
Scale Labs, a research unit under Scale AI, and Elorian have released a new benchmark called Humanity’s Sixth Sense, or HSS, aimed at testing whether AI systems can infer unstated information from images and videos. The benchmark includes 522 open-ended questions built from 288 images and 234 videos, covering situations such as whether two cars can fit through a gap, why a woman chasing a bus suddenly slows down, and who holds more social power in a scene. The team evaluated 25 multimodal models and compared them with 20 human participants. Humans posted a 93.1% accuracy rate, while the top-performing model, GPT-6 Astra, reached 53.6% even under its highest reasoning setting. GPT-6.1 Sol and Claude Opus 5.5 scored 46.6% and 44.6%, and the median score across all models was 30.9%. Researchers reviewed 8,573 model failures and said 94% were tied to missed visual cues, object misidentification, or an inability to infer implicit relationships in a scene. Only about 5% were labeled as logic errors. Social understanding was especially weak, and video questions were generally harder than image questions. The study also said more reasoning tokens did not consistently improve results.

Scale Labs, a research organization under Scale AI, and Elorian have introduced a new benchmark called Humanity’s Sixth Sense, or HSS, to test whether AI can detect implicit information in images and videos.

The benchmark contains 522 open-ended questions drawn from 288 images and 234 videos. The prompts focus on everyday common-sense judgments, including whether two cars can pass through a space, why a woman chasing a bus suddenly slows down, and who has more social authority in a given setting.

The researchers tested 25 multimodal models and compared their results with answers from 20 human participants. Humans ranked first with 93.1% accuracy. GPT-6 Astra, the top model in the test, scored 53.6% even with its maximum reasoning setting enabled. GPT-6.1 Sol and Claude Opus 5.5 posted 46.6% and 44.6%, while the median score across all models was just 30.9%.

The team examined 8,573 model failures and found that 94% were linked to missing key clues, misidentifying objects, or failing to infer implicit relationships in a scene. Only about 5% were classified as logic-reasoning mistakes.

Social understanding stood out as a major weakness. Among the 25 models tested, 21 performed worst in that category. Video-based questions were also generally harder than image-based ones.

The study said adding more reasoning effort did not always help. On average, models used about 4,000 reasoning tokens per question, and on some tasks, longer reasoning led to worse scores. The researchers also tried using agents to zoom in, crop, and repeatedly inspect the visual material. Scores improved, but the gap with human performance remained clear.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
100

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.