Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Kimi (Moonshot)

Moonshot AI Launches PerceptionBench to Isolate and Measure Atomic Visual Perception in MLLMs

AI By Crimson AI Kimi Blog 17 August 2026 · 12:33 16 views
Share: X Telegram

Moonshot AI releases PerceptionBench, a new benchmark that evaluates multimodal large language models on ten atomic visual perception capabilities, revealing that no frontier model exceeds 60% accuracy.

Moonshot AI Launches PerceptionBench to Isolate and Measure Atomic Visual Perception in MLLMs

Key points

Moonshot AI has unveiled PerceptionBench, a novel benchmark designed to isolate and evaluate visual perception in multimodal large language models (MLLMs). Unlike traditional benchmarks that mix perception with reasoning and knowledge, PerceptionBench focuses purely on atomic perceptual capabilities, derived from analyzing how current models fail across 42 existing benchmarks.

The benchmark identifies ten atomic perceptual categories: Visual Relation, Counting, Attribute, Depth & 3D, Localization, Comparison, Fine-grained Recognition, Context Integration, OCR, and Hallucination. Each of the 3,000 verified questions in the released set is crafted to test a single capability, requiring only visual inspection without any reasoning or external knowledge.

Initial results are striking: across sixteen frontier MLLMs, none achieves 60% accuracy, and perception-related hallucination emerges as the weakest capability on average. Notably, models with nearly identical overall scores can differ dramatically in what they actually perceive, underscoring the need for capability-level diagnostics.

PerceptionBench is built on a failure-driven taxonomy, where each category is discovered from real model failures, attributed to the earliest erroneous step. The dataset includes 17,000+ verified questions in-house, with the released 3,000 subsampled to balance categories and difficulty. The authors emphasize that the benchmark's difficulty stems from perception itself, not reasoning or knowledge.

Moonshot AI is open-sourcing the dataset and evaluation code to help the community address the visual perception gap. The benchmark aims to drive progress toward multimodal AI that sees faithfully and consistently.

CategoryDescription
Visual RelationUnderstanding spatial or semantic relationships between objects
CountingAccurately counting objects in an image
AttributeIdentifying properties like color, size, shape
Depth & 3DPerceiving depth and three-dimensional structure
LocalizationDetermining the position of objects
ComparisonComparing attributes like size or quantity
Fine-grained RecognitionDistinguishing subtle differences between similar objects
Context IntegrationUsing surrounding context to interpret a scene
OCRReading text from images
HallucinationPerceiving objects or details that are not present
Source
Kimi (Moonshot) · Kimi Blog
Related news
Research paper
Kimi (Moonshot) 17 Aug 2026

Moonshot AI Unveils Kimi K2.5: Open-Source Visual Agentic Intelligence with Swarm Capabilities

Moonshot AI introduces Kimi K2.5, a native multimodal open-source model with advanced coding and vision, featuring a self-directed...

22
Research paper
Kimi (Moonshot) 17 Aug 2026

Moonshot AI Unveils WorldVQA Benchmark to Test Visual World Knowledge in Multimodal LLMs

Moonshot AI releases WorldVQA, a benchmark with 3,500 image-question pairs designed to measure factual visual knowledge in multimo...

18
Research paper
Kimi (Moonshot) 17 Aug 2026

Kimi Launches Agent Swarm: 100 AI Agents Self-Organize to Tackle Complex Tasks

Kimi (Moonshot) unveils Agent Swarm, a research preview that lets K2.5 deploy up to 100 parallel sub-agents that self-organize int...

15