Research Papers
Hugging Face Researchers Introduce VAD: A Counterfactual Approach to Visual Distillation
A new paper from Hugging Face introduces Visual Attribution Distillation (VAD), a counterfactual algorithm that isolates visually...
Hugging Face Unveils SwanTale: A Unified Model for Multi-Speaker Speech and Audio Generation
SwanTale, a new model from Hugging Face, unifies zero-shot and instruction-based multi-speaker speech and audio generation, achiev...
LongHorizon-Harness: A New Framework to Keep AI Agents on Track in Long Tasks
Researchers introduce LongHorizon-Harness, a task-state management framework that decouples execution from state tracking, boostin...
Hugging Face Research: CSCR Reallocates Token Credit to Improve Long-CoT Reasoning
A new paper proposes Counterfactual Sensitivity Credit Reallocation (CSCR), a simple extension of GRPO that reduces credit for hig...
EMBL AI Librarian: A Natural-Language Knowledge Layer for Life-Science Agents
Hugging Face researchers unveil EMBL AI Librarian, a knowledge layer that lets life-science AI agents query Europe PMC in natural...
RL^2-VLA: Adaptive Steering Boosts VLA Robots in Out-of-Domain Tasks
Researchers introduce RL^2, an adaptive inference-time steering framework that applies reinforcement learning on VLA latents only...
ODEWorld: Continuous-Time World Modeling via Physical-Time Flow
Researchers introduce ODEWorld, a continuous-time latent world model that learns an ODE-based velocity field in physical time, ena...
Hugging Face Researchers Introduce EVR: A New Reward Model for Consistent Multi-Reference Image Editing
A new paper from Hugging Face proposes the Multi-dimensional Evaluation-Verification Reward (EVR) to improve multi-reference image...
ExtractBench: New Benchmark Measures Accuracy, Cost, and Grounding in Enterprise Document Extraction
Hugging Face researchers introduce ExtractBench, the first benchmark to jointly score value accuracy, record completeness, groundi...
CriPO: Self-Distillation Fixes Two Hidden Failure Modes in Rubric-Based RL
Hugging Face researchers introduce Criterion-Distilled Policy Optimization (CriPO), an on-policy framework that tackles both Unexp...
Hugging Face Researchers Introduce CAPA: A Benchmark for Personalized Ambiguity Adaptation in Coding Assistants
A new benchmark, CAPA, evaluates how well coding assistants adapt to recurring, user-specific ambiguities across sessions, aiming...
SAF-OPD: A Stable Advantage Fusion Framework to Combine RLVR and On-Policy Distillation
Researchers propose Stable Advantage Fusion (SAF) to combine reinforcement learning with verifiable rewards (RLVR) and on-policy d...