Embodied intelligence—the ability of AI to perceive and act in the physical world—faces a critical data bottleneck. Existing datasets often fragment the rich, simultaneous experience of perception, motion, and interaction across different viewpoints or modalities, leaving the full perception-action loop only partially observed. To address this, researchers at Hugging Face have introduced the Ambient Capture Engine (ACE), a human-centric data engine that transforms ordinary home environments into spatially calibrated, temporally synchronized recording studios.
ACE operates at two complementary scales: a table-scale configuration that captures fine-grained hand-object manipulation, and a room-scale configuration that records whole-body motion, locomotion, and interactions across a furnished home. The system unifies egocentric and multi-view exocentric video, full-body and articulated hand motion, object geometry with 6-DoF trajectories, audio, and tactile signals into a single multisensory stream.
Using ACE, the team built ACE-Data-0, a dataset comprising 150 hours and 17 million video frames across 200 task categories, performed by 50 participants in two environments, totaling 75,000 interaction episodes. The data spans atomic manipulation, long-horizon household activity chains, and human-scene interaction, while preserving natural behavioral variation by using goal-level instructions rather than step-by-step guidance.
The researchers also introduce a hierarchical benchmark that progresses from signals to scene components and then to interactions. Evaluations of state-of-the-art methods reveal substantial gaps under contact, occlusion, egomotion, and long temporal horizons, highlighting the dataset's value as a challenging testbed. ACE-Data-0 provides synchronized human demonstrations with aligned perceptual, kinematic, and contact supervision, offering a scalable foundation for imitation learning, world models, vision-language-action systems, and embodied AI.