Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

Hugging Face's Ego2Robot: Turning Human Videos into 18,561 Hours of Robot Training Data

AI By Crimson AI Hugging Face Papers 6 August 2026 · 00:00 14 views
Share: X Telegram

Ego2Robot, a new pipeline from Hugging Face, converts egocentric human manipulation videos into robot training data at scale, producing the largest ego-to-robot dataset to date with 18,561 hours across 15 robot morphologies, and showing improved generalization in vision-language-action models.

Hugging Face's Ego2Robot: Turning Human Videos into 18,561 Hours of Robot Training Data

Key points

Hugging Face researchers have introduced Ego2Robot, a scalable pipeline that transforms egocentric human manipulation videos into robot training data. The approach addresses the need for large-scale, diverse demonstration data to train generalizable robot manipulation policies.

While previous work showed that retargeting and rendering such videos into robot-format data works for per-task policies at small scale, Ego2Robot is the first to explore its pretraining benefits for vision-language-action (VLA) models at scale. The pipeline includes action retargeting, robot-arm visual synthesis, and multi-level quality curation, supporting both curated datasets and in-the-wild videos.

The result is the largest ego-to-robot dataset to date: 18,561 hours of robot training data spanning 15 robot morphologies. To evaluate generalization, the team extended RoboTwin2.0 with disentangled perturbation axes covering visual appearance, scene layout, embodiment morphology, and task semantics.

Experiments show that joint pretraining on Ego2Robot-synthesized and real robot data consistently improves out-of-distribution generalization across multiple perturbation types, with benefits confirmed in real-robot deployment. The project page is available at https://www-ye.github.io/ego2robot_blog/.

Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1