Research Papers
MatrAIx: Simulating the World with 8.3 Billion Persona Agents
Researchers from Harvard, MIT, and major AI labs introduce MatrAIx, a population-scale evaluation infrastructure with 8.3 billion...
SFT Conflicts, RL Coexists: New Study Explains Multi-Task Learning in LLMs
A new paper from Hugging Face reveals that supervised fine-tuning (SFT) suffers from task conflicts in multi-task training, while...
Beyond Scaling: New Methods for Designing Effective Multimodal Agent Training Environments
A new study from Hugging Face challenges the assumption that simply scaling up multimodal environments improves agent training, pr...
StreamArena: A New Benchmark for Hour-Long Interactive Video Understanding
Hugging Face researchers introduce StreamArena, a benchmark for hour-scale interactive video understanding, and StreamMind, a two-...
Hugging Face Researchers Unveil WorldTrace to Fix Memory in Video World Models
A new training-free framework, WorldTrace, addresses visual persistence failures in interactive video world models by keeping comp...
Uncertainty-Aware World Model Improves Aerial Image-Goal Navigation
Hugging Face researchers introduce UA-NWM, an uncertainty-aware latent world model that scores trajectories via conditional out-of...
AudioRubrics: Self-Evolving Rubrics Boost Audio Reasoning in RL
Hugging Face researchers introduce AudioRubrics, a reinforcement learning framework that uses self-evolving, audio-grounded rubric...
YOLO-PEFT: A Structure-Aware Framework for Parameter-Efficient Fine-Tuning of Real-Time Detectors
Researchers propose YOLO-PEFT, a framework that treats adapter placement as an auditable constraint-planning problem, outperformin...
SimWAM: A Simple World Action Model for Efficient End-to-End Autonomous Driving
Hugging Face researchers introduce SimWAM, a World Action Model that uses video generation only as a training signal, achieving 91...
GaussianSelector: Scribble-Based 3D Object Selection Without Retraining
Hugging Face researchers introduce GaussianSelector, a training-free framework that lets users select complete 3D objects from spa...
Task-Conditional Flow Matching: A New SOTA for Multilingual Embedding Adaptation
Researchers propose Task-Conditional Flow Matching (TCFM), a framework that adapts multilingual embedding models by selectively ap...
FactorJEPA: A New World Model for Crowded Global South Urban Scenes
Hugging Face researchers introduce FactorJEPA, a world model that decomposes future prediction into layout, agent, and interaction...