Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding

AI By Crimson AI Hugging Face Papers 8 August 2026 · 00:00 21 views
Share: X Telegram

Hugging Face researchers introduce SmartMage, a unified multimodal LLM that dynamically selects task-relevant modalities for 3D scene understanding, achieving state-of-the-art results across five benchmarks.

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding

Key points

Understanding 3D scenes is crucial for embodied intelligence, requiring joint reasoning over visual and geometric cues. However, existing Multimodal Large Language Models (MLLMs) often rely on fixed modality combinations, which can introduce semantic noise from irrelevant modalities and underutilize informative ones, leading to wasted computation and diluted reasoning.

To address this, researchers from Hugging Face propose SmartMage, a unified MLLM that dynamically orchestrates heterogeneous modalities for semantic-aware 3D scene understanding. SmartMage incorporates two key modules: a Semantic-guided Modality Adaptive RouTng (SMART) module that selects task-relevant modalities using semantic priors, text-modality alignment, and modality quality; and a Modality-Aware Gating Expert (MAGE) module that uses modality priors to guide expert activation, fostering adaptive specialization in multimodal reasoning.

Empirically, SmartMage achieves state-of-the-art performance across five 3D scene understanding benchmarks and attains competitive results on RGB-only video understanding benchmarks. In the diagnostic benchmark ScanFacet, tasks are divided into fine-grained semantic categories, enabling analysis of modality combinations preferred by each semantic type. The observed modality-semantic patterns provide further evidence of SmartMage's effectiveness.

For more details, visit the project page: https://yuecheong.github.io/SmartMage/.

Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1