Research Papers
Hugging Face Unveils KVAE: A New Family of Multimodal Tokenizers for Generative Models
KVAE introduces a family of tokenizers for audio, image, and video, designed for latent diffusion models, achieving competitive or...
PaDoc: Parallel Decoding Framework Speeds Up Document Parsing While Preserving Full-Page Context
Hugging Face researchers introduce PaDoc, a layout-grounded document parser that enables parallel decoding of layout and content b...
Deterministic Screen-Activity Compiler Turns Agent Memory into Auditable, Replayable Frames
A new zero-model pipeline compiles passively captured screen activity into typed, byte-identical memory frames, cutting context si...
Hugging Face Unveils DataSpace: A New Benchmark for Verifiable Data Agents in Complex Workspaces
Hugging Face introduces DataSpace, a benchmark with 410 cross-language tasks and 7,439 artifacts (15.01 GB) across six formats, ch...
Hugging Face Paper Outlines Blueprint for 'Economic World Models' as Generative AI Engines
A new paper from Hugging Face proposes a six-level capability ladder for building Economic World Models (EWMs) — generative simula...
New Benchmark HarnessOpt-Bench Measures How Well LLMs Optimize Agent Harnesses
Researchers introduce HarnessOpt-Bench, a benchmark for evaluating LLMs' ability to optimize agent harnesses under budgeted, stoch...
GST-Bench: New Benchmark Reveals VLMs' Global Spatial Awareness Gap
Hugging Face researchers introduce GST-Bench, a video-based benchmark for global spatial awareness, showing that the best zero-sho...
Hugging Face Researchers Shrink MEG Speech Decoder 20x While Boosting Interpretability
A new study from Hugging Face presents a compact MEG-to-speech retrieval model that is 20 times smaller than prior systems yet ach...
OSReward: Standardizing Evaluation of VLM Judges for Computer-Use Agents
A new benchmark, OSReward, systematically evaluates VLM judges on computer-use agent trajectories, revealing a systematic leniency...
Hugging Face Study: On-Policy Delta Distillation Boosts Multilingual Math Reasoning
A new paper from Hugging Face explores On-Policy Delta Distillation (OPD²) for math reasoning in English, Korean, and Japanese, sh...
EnvACE: Training LLM Agents by Rehearsing World Dynamics, Not Interacting with Environments
Hugging Face researchers propose EnvACE, a reinforcement learning method that lets LLM agents rehearse environment responses inter...
ChronoVision: New Framework Boosts Temporal Reasoning in Multimodal AI
Hugging Face researchers introduce ChronoVision, a multimodal framework that improves temporal reasoning by reconstructing latent...