Research Papers
UniME-R1: Learning from Failures with Retrieval-Centric Chain-of-Thought for Multimodal Retrieval
Hugging Face researchers introduce UniME-R1, an embedder-adviser framework that generates Retrieval-Centric Chain-of-Thought (RC-C...
WorldClaw: Agentic Framework Generates Large-Scale 3D Worlds from Text
Hugging Face researchers introduce WorldClaw, an agentic coarse-to-fine framework that generates large-scale, editable 3D open wor...
AgentOPSD: Recursive Self-Distillation Boosts Agentic RL Credit Assignment
Hugging Face researchers introduce AgentOPSD, a critic-free recursive method for turn-level credit assignment in agentic reinforce...
Google DeepMind's WeatherNext AI Gives Cyclone Forecasters an Extra Day of Warning
Google DeepMind's WeatherNext AI model, detailed in Nature, achieves state-of-the-art accuracy in cyclone track, intensity, and wi...
WorldCycle: Self-Verifiable RL Cuts Drift in Video World Models by 44%
Hugging Face researchers introduce WorldCycle, a self-verifiable reinforcement learning framework that uses reversible action cycl...
FocusMem: A New Latent Memory Framework for GUI Agents
Hugging Face researchers introduce FocusMem, a latent memory interface that separates content, readout, and trust to improve GUI a...
Study: VLM Agents' Spatial Memory Goes Stale, Causing Safety Failures
A new empirical study reveals that memory-augmented VLM agents often fail to detect when their spatial memory is stale, leading to...
AVE-Compass: New Benchmark Puts Audio-Visual Video Editing to the Test
Researchers introduce AVE-Compass, a comprehensive benchmark with 145 videos and 2,688 checklist items, revealing that current mod...
Hugging Face Researchers Unveil RSTG: Selective Distillation Boosts RL Fine-Tuning of LLMs
A new paper from Hugging Face introduces RSTG, a method that selectively applies teacher distillation to improve reinforcement lea...
LG AI Research Unveils K-EXAONE 2.0: A 750B-Parameter Open-Weight MoE Model
LG AI Research has released K-EXAONE 2.0, an open-weight multilingual foundation model with 750B total parameters and 37B activate...
HelloWorld: A Video World Model That Lets You Interact with On-Screen Characters
Hugging Face researchers introduce HelloWorld, a video world model that enables users to prompt in-world characters to respond to...
Hugging Face's Ego2Robot: Turning Human Videos into 18,561 Hours of Robot Training Data
Ego2Robot, a new pipeline from Hugging Face, converts egocentric human manipulation videos into robot training data at scale, prod...