Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

Xiaomi-Robotics-1: A VLA Foundation Model Trained on Over 100K Hours of Real-World Data

AI By Crimson AI Hugging Face Papers 20 July 2026 · 00:00 11 views
Share: X Telegram

Xiaomi introduces Xiaomi-Robotics-1, a vision-language-action foundation model pre-trained on over 100,000 hours of real-world manipulation trajectories, achieving state-of-the-art results on multiple benchmarks and enabling efficient adaptation to new tasks with minimal data.

Xiaomi-Robotics-1: A VLA Foundation Model Trained on Over 100K Hours of Real-World Data

Key points

Xiaomi has unveiled Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model designed to scale robot learning through large-scale pre-training on diverse, embodiment-free manipulation data. The model is pre-trained on over 100,000 hours of real-world trajectories collected via UMI devices, covering more than 1,700 scenarios across homes, commercial spaces, industrial sites, and outdoor environments.

A key innovation is a scalable auto-labeling pipeline that uses a vision-language model to annotate trajectory clips with natural language descriptions of state transitions. This converts unstructured manipulation data into training examples that link visual observations, language goals, and actions, enabling the model to learn broad and generalizable action-generation capabilities.

After pre-training, Xiaomi-Robotics-1 undergoes post-training to align with physical robot embodiments using over 7,200 hours of in-house real-robot data, and to follow natural-language instructions. The aligned model can perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, with performance improving as pre-training data or model size increases.

In fine-tuning experiments on four real-world tasks (phone packing, printer refilling, laundry loading, box packing), Xiaomi-Robotics-1 achieved a 75% overall success rate with fewer than 10 hours of demonstrations per task, compared to 40% for π0.5 under the same data budget. With fewer than 40 hours per task, the success rate reached even higher levels.

On simulation benchmarks, Xiaomi-Robotics-1 sets new state-of-the-art results: 74.5% on RoboCasa, 57.4% on RoboCasa365 (surpassing previous best of 46.6%), 59.1% on VLABench, and 13.93% on RoboDojo (outperforming prior SOTA of 13.07). The model demonstrates strong scaling behavior and generalization across environments, tasks, and embodiments.

BenchmarkXiaomi-Robotics-1Previous SOTA
RoboCasa74.5%
RoboCasa36557.4%46.6%
VLABench59.1%
RoboDojo13.93%13.07%
Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1