Latest stories
AudioRubrics: Self-Evolving Rubrics Boost Audio Reasoning in RL
Hugging Face researchers introduce AudioRubrics, a reinforcement learning framework that uses self-evolving, audio-grounded rubric...
YOLO-PEFT: A Structure-Aware Framework for Parameter-Efficient Fine-Tuning of Real-Time Detectors
Researchers propose YOLO-PEFT, a framework that treats adapter placement as an auditable constraint-planning problem, outperformin...
SimWAM: A Simple World Action Model for Efficient End-to-End Autonomous Driving
Hugging Face researchers introduce SimWAM, a World Action Model that uses video generation only as a training signal, achieving 91...
GaussianSelector: Scribble-Based 3D Object Selection Without Retraining
Hugging Face researchers introduce GaussianSelector, a training-free framework that lets users select complete 3D objects from spa...
Task-Conditional Flow Matching: A New SOTA for Multilingual Embedding Adaptation
Researchers propose Task-Conditional Flow Matching (TCFM), a framework that adapts multilingual embedding models by selectively ap...
FactorJEPA: A New World Model for Crowded Global South Urban Scenes
Hugging Face researchers introduce FactorJEPA, a world model that decomposes future prediction into layout, agent, and interaction...
Weights vs. Skills: New Survey Maps the Road to Self-Improving Robots
A comprehensive survey from Hugging Face researchers organizes robot learning around two competing paradigms: frozen-weight polici...
Hugging Face Unveils MASS: A New World Model for Scalable Multiplayer Simulation
Researchers at Hugging Face propose MASS, a world model that decouples shared world state from view rendering, enabling consistent...
ContextMaster: Real-Time Multi-Shot Video Creation with Sparse Context Routing
Hugging Face researchers introduce ContextMaster, a unified model that enables real-time interactive multi-shot video creation by...
CalibForge: Using Solver Behavior to Calibrate Terminal Tasks for Agent Training
Hugging Face researchers introduce CalibForge, a system that calibrates terminal tasks using solver feedback to create a 'learnabl...
EffectLearner: New AI Framework Removes Objects and Their Effects in Video
Hugging Face researchers introduce EffectLearner, a semantic-reasoning framework for video object removal that also eliminates obj...
SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding
Hugging Face researchers introduce SmartMage, a unified multimodal LLM that dynamically selects task-relevant modalities for 3D sc...