Research Papers
Hugging Face Study Dissects Lossy Verification in Speculative Decoding, Revealing Failure Modes
A new paper from Hugging Face provides a principled analysis of lossy verification in speculative decoding, classifying methods in...
ReToken: A Single Learnable Token Boosts Vision-Language Models for Visual Retrieval
ReToken introduces a single learnable embedding that acts as an explicit retrieval target, selecting sparse query-relevant visual...
Filesystem Memory for LLM Agents: First Systematic Study Finds Organization Cuts Costs but Not Accuracy
A new study from Hugging Face researchers provides the first systematic exploration of filesystem-based memory for LLM agents, rev...
Multi-Head Attention Residuals: A New Routing Mechanism for Transformers
Hugging Face researchers introduce Multi-Head Attention Residuals (MHAR), a zero-parameter enhancement to attention residuals that...
Echoverse: Evolving Synthetic Environments Boost Computer-Use Agents from 36.5% to 67.1%
Microsoft Research introduces Echoverse, a pipeline that compiles specifications into stateful synthetic applications with grounde...
LedgerMind: A Provenance-Constrained Framework to Boost Faithfulness in Multimodal Agents
LedgerMind introduces a provenance-constrained state machine for multimodal agents, using a Structured Evidence Ledger to ensure g...
OpenAI Unveils Breakthroughs in Mathematics and Theoretical Computer Science
OpenAI announces new results on long-standing open problems in mathematics and theoretical computer science, covering geometry, cr...
Chimera: A Hybrid Diffusion Backbone That Scales Efficiently to Long-Context Visual Generation
Hugging Face researchers introduce Chimera, a hybrid visual diffusion backbone that combines linear-complexity attention with spar...
β-OPSD: A Principled Generalization of On-Policy Self-Distillation for Reasoning Models
Researchers introduce β-OPSD, a generalization of on-policy self-distillation that turns a fixed KL penalty into a tunable paramet...
See2Think: New Benchmark Probes Whether Multimodal Models Truly Use Visual Reasoning States
Hugging Face researchers introduce See2Think, a unified evaluation framework with a 1,200-problem benchmark and Visual Action-of-T...
SpatialCLI: Teaching VLMs to Reason With Spatial Tools, Then Internalize Them
Hugging Face researchers introduce SpatialCLI, a three-stage framework that boosts spatial reasoning in vision-language models by...
RefCaptioner: New Framework Grounds Video Captions to Multiple Reference Images
Hugging Face researchers introduce RefCaptioner, a two-stage post-training framework for multi-reference image-grounded video capt...