Research Papers
StateAct: Program State, Not Pixels, Boosts Computer-Use Agents by 9x Cheaper
Hugging Face's StateAct uses program state instead of screenshots for computer-use agents, achieving state-of-the-art on OSWorld 2...
JarvisHub: Open-Source Canvas-Native Harness for Long-Horizon Multimodal Creative Agents
Hugging Face introduces JarvisHub, an open-source harness that reimagines creative AI as a canvas-native agent system, enabling lo...
MAPD: Multi-Agent Protocol Distillation Bridges Proprietary-to-Open-Source Gap in Agentic Search
A new framework, Multi-Agent Protocol Distillation (MAPD), combines distillation and reinforcement learning to transfer reasoning...
Hugging Face Unveils Kimi K3: A 2.8T Parameter Open-Source AI Model
Hugging Face introduces Kimi K3, a 2.8 trillion parameter Mixture-of-Experts model with 104 billion activated parameters, native v...
Hugging Face Survey Unifies Progress Reward Modeling for Robotic Learning
A new comprehensive survey from Hugging Face provides a unified framework for progress reward modeling in robotics, addressing the...
Multimodal Speaker Verification Threatens Anonymization, Study Finds
A new study shows that aggregating audio, prosodic, and linguistic cues across multiple anonymized utterances can significantly im...
VisCo: Hugging Face Researchers Use LLMs as Intrinsic Encoders for Visual Token Compression
VisCo introduces a training-efficient self-compression framework that reuses a pretrained VLM as an intrinsic compressor, achievin...
Spectral Alignment (SPA): A Lightweight Fix for Exposure Bias in Diffusion Models
Researchers propose Spectral Alignment (SPA), a guidance-based method that calibrates the power spectrum of intermediate predictio...
Training-Free Method Solves Revisit Inconsistency in Autoregressive Video Generation
Researchers propose a training-free approach that uses 3D engine correspondences to maintain consistent appearance when autoregres...
ID-V2V: Netflix and Hugging Face Introduce Identity-Preserving Video Restylization
A new research paper from Netflix and Hugging Face, to appear at SIGGRAPH Asia 2026, presents ID-V2V, a video-to-video framework t...
Multi-Head Latent Control: Lightweight Layer Enables Smarter LLM Agent Decisions
Hugging Face researchers introduce Multi-Head Latent Control, a lightweight layer that reads hidden states from frozen LLMs to pro...
O-VAD: Training-Free Agentic Framework Outperforms Frontier VLMs in Industrial Video Anomaly Detection
Researchers introduce O-VAD, a training-free agentic framework that uses object-centric tracking and reasoning to detect anomalies...