Research Papers
AI4AI at Test-Time: Strong Models Boost Weak Ones Without Retraining
New research from Hugging Face shows that stronger AI models can build inference-time harnesses that nearly double the performance...
OpenART: A New Arena for Scaling AI Agent Red Teaming via Evolving Environments
Hugging Face researchers introduce OpenART, an open-ended arena with over 10,000 stateful scenarios, and EMHA, an attack policy th...
OpenAI: Enterprises Shift from AI Assistance to Autonomous Execution
New OpenAI research shows enterprises are moving from using AI for assistance to deploying agentic AI with ChatGPT and Codex, with...
TSDS-Toolbox: A Unified Framework for Measuring Time-Series Dataset Similarity
Hugging Face researchers introduce TSDS-Toolbox, a unified, extensible framework for reproducible comparison of time-series datase...
360CityArena: New Benchmark Shows Huge Gap in AI Urban Navigation
A new photorealistic benchmark built from 360-degree videos of Tokyo's Akihabara district reveals that even the best AI agents per...
New SPIEval Benchmark Exposes Major Gaps in LLM Mobile Assistants
SPIEval, a new human-curated benchmark, evaluates LLMs as mobile assistants handling scattered personal data, revealing that even...
DistilVDR: Compact 524M Visual Document Retriever Achieves Near-Teacher Accuracy via Dual-Student Distillation
Hugging Face researchers introduce DistilVDR, a 524M-parameter visual document retriever distilled from an 8B teacher using cosine...
InSight-doc: Adaptive Visual Perception Cuts Hallucinations and Latency in Long-Document AI
Hugging Face researchers introduce InSight-doc, an agentic framework that dynamically adjusts visual resolution during reasoning,...
Early Pruning Boosts Efficiency in Deep Research Agents, Study Finds
A new study from Hugging Face shows that pruning context early in deep research agents yields the largest efficiency gains, reduci...
UniMoMo: Compressing Recommendation MoE Models with Functional Expert Merging
Hugging Face researchers introduce UniMoMo, a post-training compression framework that converts large recommendation MoE models in...
Hugging Face Researchers Boost Multilingual Translation with Reference-Free RL
A new open-source model family, MiLMMT-46-v1.0, uses reference-free reinforcement learning and checkpoint interpolation to surpass...
Decoding-Level Taboo: A New Stress Test Exposes LLM Fragility Off the Beaten Path
Researchers introduce Decoding-Level Taboo, a runtime logit-space stress test that forces LLMs off their nominal generation paths,...