Research Papers
PRM-as-a-Judge 1.5: New Toolkit for Fine-Grained Robot Process Evaluation
Hugging Face researchers introduce PRM-as-a-Judge 1.5, a toolkit that evaluates embodied robotic models beyond binary success rate...
Hugging Face Researchers Unveil LOPD: A New Self-Distillation Method That Learns Its Own Teaching Context
Latent On-Policy Self-Distillation (LOPD) makes the teacher's privileged context learnable end-to-end, outperforming existing meth...
DFM Mimir v1: Open 1B Model Sets Danish SOTA Using Only Permissible Data
Hugging Face unveils Mimir v1, a 1-billion-parameter Hierarchical Reasoning Model trained solely on permissible data, delivering c...
Hugging Face's Marionette: A New World Model That Predicts 3D States Instead of Pixels
Marionette, a novel world model for interactive games, explicitly predicts 3D articulated states, uses a fixed renderer for geomet...
SimpleOPD: A Tokenizer-Agnostic Approach to Distill Long-Context Reasoning into Smaller Models
Researchers introduce SimpleOPD, a method for on-policy distillation from long-context reasoning teachers to short-context student...
Hugging Face Unveils Apodex Discovery: A Framework for Verifiable AI-Driven Scientific Investigation
Apodex Discovery introduces a framework for building and evaluating 'discoverative' AI systems that pursue extended, verifiable in...
Hugging Face Research: Self-Supervised Distillation Boosts Small VLMs Without Privileged Data
A new self-supervised method, S2VOPD, improves small vision-language models by distilling from original images into strongly augme...
CPI-Bench: New Benchmark Aims to Better Evaluate Real-World Image Editing Models
Hugging Face researchers introduce CPI-Bench, a comprehensive benchmark for real-world image editing that covers multi-image tasks...
Hugging Face Unveils HumanTracker: A Human-Aligned Benchmark for Humanoid Motion Tracking
HumanTracker is a large-scale benchmark with a preference-aligned metric (HumanScore) that evaluates humanoid motion tracking base...
Hugging Face Unveils MobileMem: A Benchmark for Year-Scale On-Device Memory
MobileMem is a new benchmark and framework from Hugging Face for evaluating on-device long-term memory using year-scale, multimoda...
Beyond Final Scores: New Evaluation Framework Reveals AI Agents Are Engineering Optimizers, Not Autonomous Researchers
A systematic evaluation of seven frontier models across 36 long-horizon tasks reveals that AI agents excel at engineering optimiza...
Hugging Face Unveils Mobius-v0: Decoupling Knowledge and Reasoning for Efficient AI
Mobius-v0, a new foundation model architecture from Hugging Face, separates global memory from iterative reasoning modules, achiev...