Latest stories
Hugging Face Audit: 110 of 124 AI Evaluations Fail to Support Their Claims
A new commit-bound census of 124 Inspect Evals units reveals that 110 stop before deterministic inference due to missing historica...
Aphanta: New Framework Diagnoses When Image Editing Boosts Multimodal Reasoning
Hugging Face researchers introduce Aphanta, a diagnostic framework that evaluates when image-editing intermediates improve multimo...
EditaLive! Enables Real-Time Character Video Editing for Live Streaming
Hugging Face researchers introduce EditaLive, a framework for real-time human-centric video editing in live streams, achieving sta...
Hugging Face Researchers Unveil Agentic Framework for Consistent Multi-Shot Video Editing
A new agentic framework combining LLMs and VLMs tackles the challenge of editing long multi-shot videos with multiple instructions...
TacForcing: Streaming Tactile Feedback Boosts Contact-Rich Robot Manipulation
Hugging Face researchers introduce TacForcing, a streaming action-generation framework that integrates real-time tactile feedback...
WikiSkill: Co-Evolving Agent Skills with a Persistent Knowledge Base
Hugging Face researchers introduce WikiSkill, a framework that co-evolves agent skills with a persistent knowledge base (wiki) to...
Evolution Strategies Outperform GRPO in Reasoning Coverage, Study Finds
A new paper shows that Evolution Strategies (ES) provide broader reasoning coverage and higher Pass@K than GRPO, and proposes a hy...
Zero-WAM Lets Robots Learn New Tasks by Watching Human Videos
Hugging Face researchers introduce Zero-WAM, a causal video-action model that enables robots to generalize to unseen manipulation...
PILOT Harness Enables Live Self-Improvement for Long-Horizon AI Agents
Hugging Face researchers introduce PILOT, a supervisor-worker harness that allows live steering and self-evolution of long-horizon...
GameWAM: First World-Action Model for Native Video-Game Control
Hugging Face researchers introduce GameWAM, the first world-action model for native closed-loop gameplay and GUI control, jointly...
Hugging Face Paper: Harness-Aware Training Boosts Compact Avatar Agents
A new technical report from Hugging Face introduces Harness-Aware Training (HAT), a method that lets compact AI agents adapt to ev...
CaSKG: Counterfactual-Causal Skill Graphs Boost LLM Agent Retrieval
A new framework from Hugging Face researchers calibrates skill relations using counterfactual-causal graphs, improving retrieval a...