Collecting large-scale robot manipulation data is costly and difficult to scale, while abundant egocentric human videos offer rich behavioral experiences. However, transferring these experiences to robots is challenging due to differences between human hands and robotic end-effectors. Recent video world models promise to synthesize robot-centric videos from human observations, but their cross-embodiment capabilities remain largely unexplored.
To address this, researchers at Hugging Face introduce H2R-Bench, a benchmark for evaluating cross-embodiment human-to-robot manipulation video generation. Each instance includes a human demonstration video, target embodiment constraints, and source-grounded annotations covering task goals, action events, functional contacts, and object responses.
The benchmark assesses generated videos across five dimensions: goal-state completion, action-event completion, functional contact transfer, embodiment correctness, and general video quality. The team benchmarked eleven state-of-the-art video generation models across six manipulation families and two robot embodiments.
Results show that current video world models remain limited in human-to-robot transfer, with even leading models often failing in embodiment consistency, functional interaction, and task execution. H2R-Bench provides a systematic diagnostic framework to evaluate whether these models can bridge the human-to-robot embodiment gap and convert human observations into robot-centric training resources.