Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

DyPES-VLA: A New Cross-Embodiment VLA Model with Shared Dynamics Priors

AI By Crimson AI Hugging Face Papers 8 August 2026 · 00:00 17 views
Share: X Telegram

Researchers propose DyPES-VLA, a cross-embodiment VLA model that learns shared dynamics priors via future prediction and uses an embodiment-specific Mixture-of-Experts action head to control different robots natively, achieving state-of-the-art results on multiple benchmarks.

DyPES-VLA: A New Cross-Embodiment VLA Model with Shared Dynamics Priors

Key points

Vision-Language-Action (VLA) models have become a powerful approach for robot manipulation, but training a single generalist policy that works across heterogeneous robot embodiments remains a significant challenge. Existing methods often underutilize shared dynamics priors from diverse visual and interaction data, limiting cross-embodiment transfer. They also require extensive manual preprocessing to align different action spaces into a common format.

To address these issues, researchers introduce DyPES-VLA, a cross-embodiment VLA that learns shared Dynamics Priors and Embodiment-Specific control. The model first trains a vision-language model (VLM) with a future-prediction objective on cross-embodiment data, enabling the shared query representation to capture object motion, contact, and interaction-induced scene changes. This allows the model to learn dynamics priors that are transferable across different robots.

Second, an embodiment-specific Mixture-of-Experts (MoE) action head translates these shared dynamics priors into executable controls directly in each embodiment's native action space. This head shares attention layers to capture common temporal action structures, while its embodiment-specific feed-forward experts handle unique kinematic constraints and control semantics. This design eliminates the need for manual pre-alignment of heterogeneous actions.

As a generalist policy, DyPES-VLA achieves state-of-the-art performance across simulation and real-world evaluations, reaching 98.0% success on LIBERO, 59.25% on RoboCasa-GR1, and 89.02% on RoboTwin 2.0. The results demonstrate the effectiveness of learning shared dynamics priors and embodiment-specific control for cross-embodiment manipulation.

BenchmarkSuccess Rate
LIBERO98.0%
RoboCasa-GR159.25%
RoboTwin 2.089.02%
Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1