Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

PerceptionBench: New Benchmark Reveals MLLMs Struggle with Atomic Visual Perception

AI By Crimson AI Hugging Face Papers 29 July 2026 · 00:00 11 views
Share: X Telegram

Hugging Face researchers introduce PerceptionBench, a benchmark isolating ten atomic visual perception capabilities. Tests on 16 frontier MLLMs show no model exceeds 60% accuracy, highlighting persistent perception gaps.

PerceptionBench: New Benchmark Reveals MLLMs Struggle with Atomic Visual Perception

Key points

Researchers at Hugging Face have released PerceptionBench, a new benchmark designed to isolate and evaluate the atomic visual perception abilities of Multimodal Large Language Models (MLLMs). The work addresses a critical flaw in existing evaluations: they often conflate perception errors with failures in reasoning or domain knowledge, making it difficult to pinpoint where models truly struggle.

The team adopted a bottom-up approach by analyzing the earliest failure points in responses from frontier MLLMs across 42 existing benchmarks. This analysis yielded an error taxonomy whose perception branch defines ten atomic perceptual capabilities. Using this taxonomy, they constructed 3,000 verified questions, each isolating a single capability with short, unambiguous answers. Difficulty is derived purely from perception, not reasoning or knowledge.

Results from testing sixteen frontier MLLMs reveal that atomic visual perception remains largely unsolved. No model achieved more than 60% accuracy, and perception-related hallucination emerged as the weakest capability on average. Notably, models with similar overall scores often exhibited sharply divergent capability profiles, underscoring the need for fine-grained diagnostic tools.

PerceptionBench provides a capability-level standard for measuring and diagnosing the visual perception boundaries of MLLMs, offering a clearer path toward improving their foundational visual understanding.

Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1