Xiaomi has unveiled Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model designed to scale robot learning through large-scale pre-training on diverse, embodiment-free manipulation data. The model is pre-trained on over 100,000 hours of real-world trajectories collected via UMI devices, covering more than 1,700 scenarios across homes, commercial spaces, industrial sites, and outdoor environments.
A key innovation is a scalable auto-labeling pipeline that uses a vision-language model to annotate trajectory clips with natural language descriptions of state transitions. This converts unstructured manipulation data into training examples that link visual observations, language goals, and actions, enabling the model to learn broad and generalizable action-generation capabilities.
After pre-training, Xiaomi-Robotics-1 undergoes post-training to align with physical robot embodiments using over 7,200 hours of in-house real-robot data, and to follow natural-language instructions. The aligned model can perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, with performance improving as pre-training data or model size increases.
In fine-tuning experiments on four real-world tasks (phone packing, printer refilling, laundry loading, box packing), Xiaomi-Robotics-1 achieved a 75% overall success rate with fewer than 10 hours of demonstrations per task, compared to 40% for π0.5 under the same data budget. With fewer than 40 hours per task, the success rate reached even higher levels.
On simulation benchmarks, Xiaomi-Robotics-1 sets new state-of-the-art results: 74.5% on RoboCasa, 57.4% on RoboCasa365 (surpassing previous best of 46.6%), 59.1% on VLABench, and 13.93% on RoboDojo (outperforming prior SOTA of 13.07). The model demonstrates strong scaling behavior and generalization across environments, tasks, and embodiments.