Researchers from Hugging Face and collaborators have introduced HiFi-UMI, a portable, robot-free data production system designed to generate high-fidelity manipulation demonstrations. The system achieves 3 mm workspace-local end-effector accuracy, sub-40 microsecond cross-sensor synchronization, and ultra-wide six-view sensing, all without external tracking infrastructure.
The key innovation is that HiFi-UMI data alone can be used for post-training, eliminating the need for any real-robot teleoperation data in that phase. Across three backbones—StarVLA-QwenPI, OpenPI-π0.5, and LingBot-VA—policies post-trained solely on HiFi-UMI demonstrations matched in-domain robot teleoperation, with success-rate differences of -2.5, +3.1, and -0.6 percentage points respectively.
The strongest policy achieved 85% success on a precision insertion task, even though no HiFi-UMI demonstration was collected in the evaluation scene. Pre-training on 4,000 hours from the same corpus reduced action error on ten unseen tasks by 41% and improved real-robot success by an additional 18.1 percentage points on StarVLA-QwenPI.
The team is releasing HiFi-UMI-2K, a dataset of 2,000 hours and over 482,000 replayable demonstrations across 110+ scenes, under CC BY 4.0 license. This resource is intended to serve as a large-scale, high-fidelity foundation for the robot-learning community.