Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Kimi (Moonshot)

Kimi Releases PerceptionBench: A New Benchmark for Atomic Visual Perception in Multimodal AI

AI By Crimson AI Moonshot / Kimi 25 July 2026 · 08:17 32 views
Share: X Telegram

Kimi (Moonshot) unveils PerceptionBench, a benchmark designed to isolate and evaluate atomic visual perception capabilities in multimodal large language models, revealing that no model exceeds 60% accuracy and many correct answers are inconsistent.

Kimi Releases PerceptionBench: A New Benchmark for Atomic Visual Perception in Multimodal AI

Key points

Kimi (Moonshot) has announced the release of PerceptionBench, a novel benchmark designed to evaluate atomic visual perception in multimodal large language models (MLLMs). Unlike traditional benchmarks that measure overall performance, PerceptionBench focuses on isolating visual perception as a set of atomic capabilities, derived from analyzing how current models fail across over 40 existing benchmarks.

The benchmark identifies ten atomic perceptual categories: Visual Relation, Counting, Attribute, Depth & 3D, Localization, Comparison, Fine-grained Recognition, Context Integration, OCR, and Hallucination. It comprises 3,000 verified questions, each answerable purely by looking, without requiring reasoning or outside knowledge. The dataset is curated to ensure that difficulty stems from perception rather than reasoning.

Key findings from PerceptionBench reveal that no model evaluated achieves over 60% accuracy. Moreover, models with nearly identical overall scores can exhibit vastly different perceptual strengths and weaknesses. Strikingly, a large proportion of correct answers fail to survive repeated questioning, suggesting that current models often guess rather than perceive consistently.

The benchmark's design is based on a failure-driven taxonomy, where categories are discovered from real model failures attributed to the earliest erroneous step. The dataset undergoes rigorous verification and difficulty-balancing to serve as a reliable gold standard for diagnosing perceptual capabilities.

PerceptionBench is available for download soon, aiming to drive progress toward multimodal AI that sees faithfully and consistently.

Source
Kimi (Moonshot) · Moonshot / Kimi
Related news
Research paper
Kimi (Moonshot) 17 Aug 2026

Moonshot AI Unveils Kimi K2.5: Open-Source Visual Agentic Intelligence with Swarm Capabilities

Moonshot AI introduces Kimi K2.5, a native multimodal open-source model with advanced coding and vision, featuring a self-directed...

22
Research paper
Kimi (Moonshot) 17 Aug 2026

Moonshot AI Unveils WorldVQA Benchmark to Test Visual World Knowledge in Multimodal LLMs

Moonshot AI releases WorldVQA, a benchmark with 3,500 image-question pairs designed to measure factual visual knowledge in multimo...

18
Research paper
Kimi (Moonshot) 17 Aug 2026

Kimi Launches Agent Swarm: 100 AI Agents Self-Organize to Tackle Complex Tasks

Kimi (Moonshot) unveils Agent Swarm, a research preview that lets K2.5 deploy up to 100 parallel sub-agents that self-organize int...

16