Research Papers
Hugging Face Unveils DME: A Two-Stage Multimodal Embedding Model for Billion-Scale Search
Douyin's DME combines contrastive pre-training with training-only reasoning and reconstruction to achieve state-of-the-art results...
Hugging Face Study: Offline Top-K Distillation Cuts Memory, Boosts Throughput
A new paper from Hugging Face shows that caching teacher logits and using a chunked KL loss can make knowledge distillation for sm...
Small Cognitive Models Match Giants In-Distribution but Scale Better Out-of-Distribution
New research shows that small language models fine-tuned on human behavioral data can match a 70B baseline in-distribution, but la...
Referential Dangling: A Hidden Failure Mode in Hard Prompt Compression
New research from Hugging Face reveals that hard prompt compression methods often delete the context needed to interpret retained...
Fine-Tuned Activation Oracles Develop Concept-Specific Blind Spots, Study Finds
New research from Hugging Face reveals that fine-tuning activation oracles on a subject model that hides a concept makes them sele...
DCAS: Decoupling Scaffold Planning to Make CLI Agents Generalize
A new interception layer, DCAS, decouples planning from scaffold-specific training, enabling fine-tuned CLI coding agents to gener...
MatrAIx: Simulating the World with 8.3 Billion Persona Agents
Researchers from Harvard, MIT, and major AI labs introduce MatrAIx, a population-scale evaluation infrastructure with 8.3 billion...
SFT Conflicts, RL Coexists: New Study Explains Multi-Task Learning in LLMs
A new paper from Hugging Face reveals that supervised fine-tuning (SFT) suffers from task conflicts in multi-task training, while...
Beyond Scaling: New Methods for Designing Effective Multimodal Agent Training Environments
A new study from Hugging Face challenges the assumption that simply scaling up multimodal environments improves agent training, pr...
StreamArena: A New Benchmark for Hour-Long Interactive Video Understanding
Hugging Face researchers introduce StreamArena, a benchmark for hour-scale interactive video understanding, and StreamMind, a two-...
Hugging Face Researchers Unveil WorldTrace to Fix Memory in Video World Models
A new training-free framework, WorldTrace, addresses visual persistence failures in interactive video world models by keeping comp...
Uncertainty-Aware World Model Improves Aerial Image-Goal Navigation
Hugging Face researchers introduce UA-NWM, an uncertainty-aware latent world model that scores trajectories via conditional out-of...