Latest stories
OpenAI and APA Launch Three-Year Partnership for Youth Mental Health AI
OpenAI and the American Psychological Association (APA) have announced a three-year collaboration to develop guidance, resources,...
WorldCycle: Self-Verifiable RL Cuts Drift in Video World Models by 44%
Hugging Face researchers introduce WorldCycle, a self-verifiable reinforcement learning framework that uses reversible action cycl...
FocusMem: A New Latent Memory Framework for GUI Agents
Hugging Face researchers introduce FocusMem, a latent memory interface that separates content, readout, and trust to improve GUI a...
Study: VLM Agents' Spatial Memory Goes Stale, Causing Safety Failures
A new empirical study reveals that memory-augmented VLM agents often fail to detect when their spatial memory is stale, leading to...
AVE-Compass: New Benchmark Puts Audio-Visual Video Editing to the Test
Researchers introduce AVE-Compass, a comprehensive benchmark with 145 videos and 2,688 checklist items, revealing that current mod...
Hugging Face Researchers Unveil RSTG: Selective Distillation Boosts RL Fine-Tuning of LLMs
A new paper from Hugging Face introduces RSTG, a method that selectively applies teacher distillation to improve reinforcement lea...
LG AI Research Unveils K-EXAONE 2.0: A 750B-Parameter Open-Weight MoE Model
LG AI Research has released K-EXAONE 2.0, an open-weight multilingual foundation model with 750B total parameters and 37B activate...
HelloWorld: A Video World Model That Lets You Interact with On-Screen Characters
Hugging Face researchers introduce HelloWorld, a video world model that enables users to prompt in-world characters to respond to...
Hugging Face's Ego2Robot: Turning Human Videos into 18,561 Hours of Robot Training Data
Ego2Robot, a new pipeline from Hugging Face, converts egocentric human manipulation videos into robot training data at scale, prod...
GDPevo: New Benchmark Tests AI Agents' Self-Evolution on Real Business Tasks
Hugging Face researchers introduce GDPevo, the first benchmark for evaluating agent self-evolution on GDP-related enterprise workf...
Skill Entropy: A New Metric and Training Signal for Long-Horizon Reasoning in LLMs
Researchers introduce Skill Entropy, a measure of cross-skill switching difficulty, and Skill^2-Bench, a benchmark spanning 558 sk...
Hugging Face's RST Framework Generates 37K Terminal Tasks at $0.05 Each
A new recursive synthesis framework from Hugging Face produces 37,484 long-horizon terminal-agent tasks at roughly $0.05 per task,...