Research Papers
Evo-Bench: New Benchmark Tests Whether LLMs Can Improve Their Own Agent Harness
Hugging Face researchers introduce Evo-Bench, the first benchmark designed to isolate and evaluate language models' ability to aut...
Business Arena: New Benchmark Reveals LLM Agents Struggle with Realistic Business Operations
A new benchmark from Hugging Face and Accio evaluates LLM agents in a realistic cross-border shop, finding a ninefold gap in perfo...
Interpretability Scales with Capability in New Training-Time Approach
A new paper from Hugging Face shows that making interpretability a training constraint yields scalable, disentangled representatio...
OasisKV: Boosting LLM Throughput by Prefetching Sparse KV Caches Beyond HBM
OasisKV, a new memory-centric inference system from Hugging Face researchers, stores full KV caches in cheaper memory tiers and us...
Hugging Face Research: Three-Stage Framework Boosts Follow-Up Edit Suggestions in Image Conversations
A new multimodal framework from Hugging Face improves follow-up edit suggestions in image-creation conversations, reducing visual...
Researchers Expose Flaw Allowing Theft of Hidden Reasoning from Major AI APIs
A new study reveals that encrypted reasoning traces from proprietary LLMs can be intercepted and decrypted by injecting them into...
Ouroboros: Self-Developing Coding Agent Sets New Benchmarks in Terminal-Bench and OSWorld
Hugging Face researchers unveil Ouroboros, a self-improving coding agent harness that evolves its own tools and prompts, achieving...
Unsupervised Self-Distillation: LLMs Learn from Their Own Majority Votes
A new Hugging Face paper introduces U-OPSD, a method that lets LLMs self-distill without external supervision by using majority-vo...
Hugging Face's BDH-CQ: 150M Model Sets New Cost-Accuracy Frontier on ARC-AGI-1
A new 150M-parameter reasoning model, BDH-CQ, combines in-context learning with recurrent latent reasoning to achieve 29.5% pass@2...
SPOT: A New Distillation Method to Boost Reasoning in Student Models
Researchers introduce SPOT, a novel on-policy distillation technique that uses sparse probing and outcome calibration to improve r...
Hugging Face Unveils Motif 3: A 314B-Parameter MoE Model with Novel Attention
Hugging Face introduces Motif 3, a 314B-parameter Mixture-of-Experts model with 13.2B active parameters, featuring Grouped Differe...
Sci-VBench: New Benchmark Tests AI Video Generation's Scientific Reasoning
Hugging Face researchers introduce Sci-VBench, a benchmark with 1,253 expert-annotated examples across 60 scientific subjects, rev...