Research Papers
AVE-Compass: New Benchmark Puts Audio-Visual Video Editing to the Test
Researchers introduce AVE-Compass, a comprehensive benchmark with 145 videos and 2,688 checklist items, revealing that current mod...
Hugging Face Researchers Unveil RSTG: Selective Distillation Boosts RL Fine-Tuning of LLMs
A new paper from Hugging Face introduces RSTG, a method that selectively applies teacher distillation to improve reinforcement lea...
LG AI Research Unveils K-EXAONE 2.0: A 750B-Parameter Open-Weight MoE Model
LG AI Research has released K-EXAONE 2.0, an open-weight multilingual foundation model with 750B total parameters and 37B activate...
HelloWorld: A Video World Model That Lets You Interact with On-Screen Characters
Hugging Face researchers introduce HelloWorld, a video world model that enables users to prompt in-world characters to respond to...
Hugging Face's Ego2Robot: Turning Human Videos into 18,561 Hours of Robot Training Data
Ego2Robot, a new pipeline from Hugging Face, converts egocentric human manipulation videos into robot training data at scale, prod...
GDPevo: New Benchmark Tests AI Agents' Self-Evolution on Real Business Tasks
Hugging Face researchers introduce GDPevo, the first benchmark for evaluating agent self-evolution on GDP-related enterprise workf...
Skill Entropy: A New Metric and Training Signal for Long-Horizon Reasoning in LLMs
Researchers introduce Skill Entropy, a measure of cross-skill switching difficulty, and Skill^2-Bench, a benchmark spanning 558 sk...
Hugging Face's RST Framework Generates 37K Terminal Tasks at $0.05 Each
A new recursive synthesis framework from Hugging Face produces 37,484 long-horizon terminal-agent tasks at roughly $0.05 per task,...
ABSeeker: New Training Method Boosts Small AI Search Agents to Rival 30B Models
Hugging Face researchers introduce ABSeeker, a framework that uses answer-backtracked credit assignment to train long-horizon sear...
NOLLI Benchmark Reveals Korean AI Gaps in Jamo Execution, Not Language
A new procedurally generated puzzle benchmark, NOLLI, diagnoses where English-Korean performance gaps arise in AI models, finding...
SA-OPD: New Framework Filters Spurious Teacher Signals in On-Policy Distillation
Researchers propose SA-OPD, a spurious-signal-aware framework for on-policy distillation that filters misleading token-level teach...
Hugging Face Unveils OneDayAgent: A Long-Horizon Harness for Autonomous Agents
Hugging Face researchers present OneDayAgent, a harness that manages long-horizon, cross-environment tasks for LLM agents, achievi...