Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

SpatialCLI: Teaching VLMs to Reason With Spatial Tools, Then Internalize Them

AI By Crimson AI Hugging Face Papers 1 August 2026 · 00:00 19 views
Share: X Telegram

Hugging Face researchers introduce SpatialCLI, a three-stage framework that boosts spatial reasoning in vision-language models by first using specialist tools and then internalizing those capabilities for tool-free inference, achieving dramatic gains on the MindCube benchmark.

SpatialCLI: Teaching VLMs to Reason With Spatial Tools, Then Internalize Them

Key points

Vision-language models (VLMs) are increasingly deployed in embodied agents, where they must interpret visual inputs, reason about spatial relationships, and make task-level decisions. However, a fundamental mismatch persists: general VLMs excel at high-level reasoning but often miss critical visual details, while specialist vision models capture those details but cannot translate them into task-level decisions.

To bridge this gap, researchers from Hugging Face propose SpatialCLI, a framework that teaches VLMs to reason with spatial tools and progressively internalize the specialist perceptual capabilities those tools provide. The approach unfolds in three stages: Call, Learn, and Internalize.

The team also introduces SpatialCLI-Bench, a benchmark comprising 516 examples for compositional perception across localization, segmentation, depth, and pose estimation.

On the MindCube benchmark, SpatialCLI elevates Qwen3-VL-8B-Instruct from 29.3% to 84.6% accuracy when tools are available, surpassing GPT-5.6 Sol with tools (72.1%). Remarkably, after internalization, the model retains 73.8% accuracy even without tools, demonstrating the effectiveness of the framework in transferring specialist capabilities into the model itself.

ModelWith ToolsWithout Tools
Qwen3-VL-8B-Instruct (baseline)29.3%-
Qwen3-VL-8B-Instruct + SpatialCLI84.6%73.8%
GPT-5.6 Sol (with tools)72.1%-
Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1