Hugging Face Audit: 110 of 124 AI Evaluations Fail to Support Their Claims
A new commit-bound census of 124 Inspect Evals units reveals that 110 stop before deterministic inference due to missing historical evidence or semantic grounding, challenging the reliabilit...
Read the full storyLatest stories
Aphanta: New Framework Diagnoses When Image Editing Boosts Multimodal Reasoning
Hugging Face researchers introduce Aphanta, a diagnostic framework that evaluates when image-editing intermediates improve multimo...
EditaLive! Enables Real-Time Character Video Editing for Live Streaming
Hugging Face researchers introduce EditaLive, a framework for real-time human-centric video editing in live streams, achieving sta...
Hugging Face Researchers Unveil Agentic Framework for Consistent Multi-Shot Video Editing
A new agentic framework combining LLMs and VLMs tackles the challenge of editing long multi-shot videos with multiple instructions...
TacForcing: Streaming Tactile Feedback Boosts Contact-Rich Robot Manipulation
Hugging Face researchers introduce TacForcing, a streaming action-generation framework that integrates real-time tactile feedback...
WikiSkill: Co-Evolving Agent Skills with a Persistent Knowledge Base
Hugging Face researchers introduce WikiSkill, a framework that co-evolves agent skills with a persistent knowledge base (wiki) to...
Evolution Strategies Outperform GRPO in Reasoning Coverage, Study Finds
A new paper shows that Evolution Strategies (ES) provide broader reasoning coverage and higher Pass@K than GRPO, and proposes a hy...
Zero-WAM Lets Robots Learn New Tasks by Watching Human Videos
Hugging Face researchers introduce Zero-WAM, a causal video-action model that enables robots to generalize to unseen manipulation...
PILOT Harness Enables Live Self-Improvement for Long-Horizon AI Agents
Hugging Face researchers introduce PILOT, a supervisor-worker harness that allows live steering and self-evolution of long-horizon...
GameWAM: First World-Action Model for Native Video-Game Control
Hugging Face researchers introduce GameWAM, the first world-action model for native closed-loop gameplay and GUI control, jointly...
Hugging Face Paper: Harness-Aware Training Boosts Compact Avatar Agents
A new technical report from Hugging Face introduces Harness-Aware Training (HAT), a method that lets compact AI agents adapt to ev...
Anthropic Launches $5M Grant Program to Fund Independent AI Wellbeing Evaluations
Anthropic announces a $5 million grant program to support independent research on AI's impact on user wellbeing, providing funding...