Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

Hugging Face Unveils DreamX-Phi 1.0: A World Model for Faithful Robot Video Prediction

AI By Crimson AI Hugging Face Papers 14 August 2026 · 00:00 11 views
Share: X Telegram

DreamX-Phi 1.0, an action-conditioned video world model for robotic manipulation, predicts future observations from language instructions and action sequences, achieving top ranks in the WorldArena 2.0 Challenge.

Hugging Face Unveils DreamX-Phi 1.0: A World Model for Faithful Robot Video Prediction

Key points

Researchers from Hugging Face have introduced DreamX-Phi 1.0, an action-conditioned video world model designed for robotic manipulation. The model takes an observed frame, a language instruction, and a prescribed action sequence—comprising end-effector poses and gripper states—to predict the resulting future observations.

To ensure predictions are faithful, the model incorporates per-arm SE(3) transformations into its attention mechanism using PRoPE-style geometric encoding. This preserves arm identity and rigid-motion structure, preventing issues like moving the wrong arm or losing track of manipulated objects.

For scene-level geometry, the model adds a lightweight depth branch and uses SAM3 masks with a frozen V-JEPA teacher to maintain object consistency during grasping. Additionally, a distribution-matching distillation technique compresses the multi-step generator into a few-step student for efficient deployment.

At the time of writing, DreamX-Phi 1.0 achieves first place on Track 1 and second place on Track 2 of the WorldArena 2.0 Challenge. The model weights and inference code will be publicly available on GitHub after the challenge concludes.

Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

0
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

0
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

0