Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

Hugging Face Unveils WCM: A World Critic Model to Fix Value Estimation in VLA Reinforcement Learning

AI By Crimson AI Hugging Face Papers 4 August 2026 · 00:00 9 views
Share: X Telegram

Researchers at Hugging Face propose the World Critic Model (WCM), a lightweight LeJEPA-based architecture that jointly predicts future latent states and estimates values, addressing the partial observability bottleneck in Vision-Language-Action reinforcement learning. WCM achieves state-of-the-art results across 149 tasks and four benchmarks, with strong generalization gains.

Hugging Face Unveils WCM: A World Critic Model to Fix Value Estimation in VLA Reinforcement Learning

Key points

Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown promise for robotic manipulation, but a fundamental bottleneck persists: value estimation under partial observability. Existing critic models typically rely on single-frame observations or single-frame VLM backbone latents, which fail to capture the temporal dynamics essential for accurate decision-making in robot control.

To address this, researchers at Hugging Face introduce the World Critic Model (WCM), built on a lightweight LeJEPA architecture. WCM jointly predicts future latent states and estimates values, explicitly training the critic's representation to capture temporal dynamics rather than merely regressing scalar returns. This predictive state representation unifies world modeling with critic learning, enabling the model to understand how the environment evolves.

WCM integrates seamlessly into both on-policy and off-policy training pipelines and is compatible with state-of-the-art VLA backbones including Pi0, Pi0.5, and OpenVLA-OFT. Extensive experiments on 149 tasks across four benchmarks demonstrate that WCM consistently achieves state-of-the-art performance in both in-distribution and out-of-distribution settings, with particularly strong generalization gains.

The team further validated WCM on seven real-world manipulation tasks using OpenVLA-OFT and Pi0.5 with off-policy RL, confirming stable deployment across diverse settings. The paper and project page are available for deeper exploration.

Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1