Kimi (Moonshot) has announced the release of PerceptionBench, a novel benchmark designed to evaluate atomic visual perception in multimodal large language models (MLLMs). Unlike traditional benchmarks that measure overall performance, PerceptionBench focuses on isolating visual perception as a set of atomic capabilities, derived from analyzing how current models fail across over 40 existing benchmarks.
The benchmark identifies ten atomic perceptual categories: Visual Relation, Counting, Attribute, Depth & 3D, Localization, Comparison, Fine-grained Recognition, Context Integration, OCR, and Hallucination. It comprises 3,000 verified questions, each answerable purely by looking, without requiring reasoning or outside knowledge. The dataset is curated to ensure that difficulty stems from perception rather than reasoning.
Key findings from PerceptionBench reveal that no model evaluated achieves over 60% accuracy. Moreover, models with nearly identical overall scores can exhibit vastly different perceptual strengths and weaknesses. Strikingly, a large proportion of correct answers fail to survive repeated questioning, suggesting that current models often guess rather than perceive consistently.
The benchmark's design is based on a failure-driven taxonomy, where categories are discovered from real model failures attributed to the earliest erroneous step. The dataset undergoes rigorous verification and difficulty-balancing to serve as a reliable gold standard for diagnosing perceptual capabilities.
PerceptionBench is available for download soon, aiming to drive progress toward multimodal AI that sees faithfully and consistently.