Research Papers
GST-Bench: New Benchmark Reveals VLMs' Global Spatial Awareness Gap
Hugging Face researchers introduce GST-Bench, a video-based benchmark for global spatial awareness, showing that the best zero-sho...
Hugging Face Researchers Shrink MEG Speech Decoder 20x While Boosting Interpretability
A new study from Hugging Face presents a compact MEG-to-speech retrieval model that is 20 times smaller than prior systems yet ach...
OSReward: Standardizing Evaluation of VLM Judges for Computer-Use Agents
A new benchmark, OSReward, systematically evaluates VLM judges on computer-use agent trajectories, revealing a systematic leniency...
Hugging Face Study: On-Policy Delta Distillation Boosts Multilingual Math Reasoning
A new paper from Hugging Face explores On-Policy Delta Distillation (OPD²) for math reasoning in English, Korean, and Japanese, sh...
EnvACE: Training LLM Agents by Rehearsing World Dynamics, Not Interacting with Environments
Hugging Face researchers propose EnvACE, a reinforcement learning method that lets LLM agents rehearse environment responses inter...
ChronoVision: New Framework Boosts Temporal Reasoning in Multimodal AI
Hugging Face researchers introduce ChronoVision, a multimodal framework that improves temporal reasoning by reconstructing latent...
UniME-R1: Learning from Failures with Retrieval-Centric Chain-of-Thought for Multimodal Retrieval
Hugging Face researchers introduce UniME-R1, an embedder-adviser framework that generates Retrieval-Centric Chain-of-Thought (RC-C...
WorldClaw: Agentic Framework Generates Large-Scale 3D Worlds from Text
Hugging Face researchers introduce WorldClaw, an agentic coarse-to-fine framework that generates large-scale, editable 3D open wor...
AgentOPSD: Recursive Self-Distillation Boosts Agentic RL Credit Assignment
Hugging Face researchers introduce AgentOPSD, a critic-free recursive method for turn-level credit assignment in agentic reinforce...
WorldCycle: Self-Verifiable RL Cuts Drift in Video World Models by 44%
Hugging Face researchers introduce WorldCycle, a self-verifiable reinforcement learning framework that uses reversible action cycl...
FocusMem: A New Latent Memory Framework for GUI Agents
Hugging Face researchers introduce FocusMem, a latent memory interface that separates content, readout, and trust to improve GUI a...
Study: VLM Agents' Spatial Memory Goes Stale, Causing Safety Failures
A new empirical study reveals that memory-augmented VLM agents often fail to detect when their spatial memory is stale, leading to...