Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

ProVisE: New Benchmark Lets Image-Generation Models Show Spatial Intelligence in Pixels

AI By Crimson AI Hugging Face Papers 24 July 2026 · 00:00 9 views
Share: X Telegram

Researchers introduce ProVisE, a framework that evaluates image-generation models on spatial tasks by letting them answer directly in pixels, and SpatialGen-Bench, a 470-sample benchmark. Results show pixel-based models excel at direct spatial expression while text VLMs lead in compositional reasoning.

ProVisE: New Benchmark Lets Image-Generation Models Show Spatial Intelligence in Pixels

Key points

A new research paper from Hugging Face, titled Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text, tackles a fundamental mismatch in how spatial reasoning is evaluated. Current benchmarks typically require text or coordinate outputs, which disadvantages image-generation models that naturally express spatial concepts through visual marks.

The authors propose ProVisE (Protocolized Visual Evaluation), a benchmark-agnostic framework that elicits protocol-constrained visual answers from image-generation models and parses them into structured predictions compatible with original metrics. ProVisE includes an Agentic builder that constructs and validates task-specific protocols for new benchmarks.

To support this evaluation, the team introduces SpatialGen-Bench, a curated diagnostic benchmark of 470 samples across 14 spatial subtasks, four capability levels, and diverse answer forms. They evaluate 31 text-output VLMs and image-generation models in a unified setting and validate Agentic protocol construction on six external spatial benchmarks.

Key findings reveal complementary strengths: image-generation models are competitive when spatial answers can be externalized directly in pixel space, while text-output VLMs retain a clear advantage in compositional spatial reasoning. Notably, GPT Image 2 correctly solves 37% of spatial cases missed by GPT-5.4, highlighting the value of pixel-based expression.

The paper establishes a metric-compatible testbed for studying spatial cognition in image-generation models, bridging the gap between visual and textual reasoning interfaces.

BenchmarkSamplesSubtasksCapability Levels
SpatialGen-Bench470144
Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1