Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

ActiveVision Benchmark Reveals MLLMs Collapse on Active Observation Tasks

AI By Crimson AI Hugging Face Papers 23 July 2026 · 00:00 10 views
Share: X Telegram

A new benchmark, ActiveVision, tests whether multimodal large language models (MLLMs) can perform active visual observation—repeatedly seeking new evidence. Humans solve 96.1% of tasks, while the best model, GPT-5.5, scores only 10.6%, and Claude Fable 5 achieves just 3.5%.

ActiveVision Benchmark Reveals MLLMs Collapse on Active Observation Tasks

Key points

Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued that this active observation is essential for a wide range of tasks. Whether today's multimodal large language models (MLLMs) exercise active observation is an empirical question that current vision-language benchmarks do not answer.

Researchers introduce ActiveVision, a benchmark that makes active observation measurable for MLLMs. It comprises 17 tasks across 3 categories, designed to force repeated visual perception rather than a single static description. Frontier MLLMs collapse on ActiveVision: the highest-scoring model evaluated, GPT-5.5 at the highest exposed reasoning-effort tier, solves only 10.6% of items and scores zero on 11 of the 17 tasks. Even Claude Fable 5, despite topping most reasoning and coding leaderboards, solves just 3.5%, far behind three human participants who average 96.1%.

Furthermore, much of the gap persists even when models write and run their own vision code: such code is unreliable on realistic imagery, and catching its failures itself requires the active perception the models lack. Together, these results indicate that current MLLMs lack robust active visual observation, motivating architectures and training objectives that close the perception-reasoning loop.

For more details, visit the project website, code repository, and dataset on Hugging Face. The paper is available on arXiv.

Model / ParticipantAccuracy (%)
Human (average of 3)96.1
GPT-5.5 (highest reasoning-effort tier)10.6
Claude Fable 53.5
Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1