Research Papers
GDPevo: New Benchmark Tests AI Agents' Self-Evolution on Real Business Tasks
Hugging Face researchers introduce GDPevo, the first benchmark for evaluating agent self-evolution on GDP-related enterprise workf...
Skill Entropy: A New Metric and Training Signal for Long-Horizon Reasoning in LLMs
Researchers introduce Skill Entropy, a measure of cross-skill switching difficulty, and Skill^2-Bench, a benchmark spanning 558 sk...
Hugging Face's RST Framework Generates 37K Terminal Tasks at $0.05 Each
A new recursive synthesis framework from Hugging Face produces 37,484 long-horizon terminal-agent tasks at roughly $0.05 per task,...
ABSeeker: New Training Method Boosts Small AI Search Agents to Rival 30B Models
Hugging Face researchers introduce ABSeeker, a framework that uses answer-backtracked credit assignment to train long-horizon sear...
NOLLI Benchmark Reveals Korean AI Gaps in Jamo Execution, Not Language
A new procedurally generated puzzle benchmark, NOLLI, diagnoses where English-Korean performance gaps arise in AI models, finding...
SA-OPD: New Framework Filters Spurious Teacher Signals in On-Policy Distillation
Researchers propose SA-OPD, a spurious-signal-aware framework for on-policy distillation that filters misleading token-level teach...
Hugging Face Unveils OneDayAgent: A Long-Horizon Harness for Autonomous Agents
Hugging Face researchers present OneDayAgent, a harness that manages long-horizon, cross-environment tasks for LLM agents, achievi...
Hugging Face Study Unveils 'Physics' of Multimodal Pretraining, Cuts Compute by 95%
A new Hugging Face paper systematically explores multimodal pretraining, revealing four key insights into knowledge flow, modality...
Study: LLMs Fabricate User Profiles in 41.6% of Claims; Self-Monitoring Misleads
A new benchmark, MirageBench, reveals that all 12 tested LLMs over-infer user attributes in 35-49% of claims, and that self-report...
ToolArtist: A Fully Agentic Multimodal Model for Open-World Image Generation
Hugging Face researchers introduce ToolArtist, a unified multimodal model that orchestrates reasoning, external tool use, and nati...
MiniWorld: A Lightweight, Reproducible Framework for Training Video World Models from Scratch
Hugging Face researchers introduce MiniWorld, a reproducible framework for training streaming video world models from scratch on a...
GROVE: A Training-Free Memory Framework for Wearable AI Assistants
Hugging Face researchers introduce GROVE, a training-free framework that grows a single temporally stratified memory from streamin...