Research Papers
RL^2-VLA: Adaptive Steering Boosts VLA Robots in Out-of-Domain Tasks
Researchers introduce RL^2, an adaptive inference-time steering framework that applies reinforcement learning on VLA latents only...
ODEWorld: Continuous-Time World Modeling via Physical-Time Flow
Researchers introduce ODEWorld, a continuous-time latent world model that learns an ODE-based velocity field in physical time, ena...
Hugging Face Researchers Introduce EVR: A New Reward Model for Consistent Multi-Reference Image Editing
A new paper from Hugging Face proposes the Multi-dimensional Evaluation-Verification Reward (EVR) to improve multi-reference image...
ExtractBench: New Benchmark Measures Accuracy, Cost, and Grounding in Enterprise Document Extraction
Hugging Face researchers introduce ExtractBench, the first benchmark to jointly score value accuracy, record completeness, groundi...
CriPO: Self-Distillation Fixes Two Hidden Failure Modes in Rubric-Based RL
Hugging Face researchers introduce Criterion-Distilled Policy Optimization (CriPO), an on-policy framework that tackles both Unexp...
Hugging Face Researchers Introduce CAPA: A Benchmark for Personalized Ambiguity Adaptation in Coding Assistants
A new benchmark, CAPA, evaluates how well coding assistants adapt to recurring, user-specific ambiguities across sessions, aiming...
SAF-OPD: A Stable Advantage Fusion Framework to Combine RLVR and On-Policy Distillation
Researchers propose Stable Advantage Fusion (SAF) to combine reinforcement learning with verifiable rewards (RLVR) and on-policy d...
N0-TWAM: First Large-Scale Tactile World-Action Model Predicts Touch and Vision for Contact-Rich Robots
Hugging Face researchers introduce N0-TWAM, a tactile-native world-action model that jointly predicts future vision and contact, t...
Weak-to-Strong On-Policy Distillation: Boosting LLMs with Weaker Teachers
A new Hugging Face research paper introduces Weak-to-Strong On-Policy Distillation (W2S-OPD), a method that improves a strong stud...
Mental World Modeling: A New Framework for Predicting Human Decisions
Researchers introduce Mental World Modeling (MWM), a framework that integrates hidden mental states into world models, showing tha...
Hugging Face Study Reveals Scaling Laws for Text Conditioning in Visual Generation
New research from Hugging Face uncovers that diffusion loss scales with structured language in prompts, leading to a system that o...
Meshy T2: Flow Matching Enables Fast Native 3D Mesh Generation
Hugging Face researchers introduce Meshy T2, a flow-matching framework that generates native polygonal meshes with artist-style to...