Latest stories
PlayWorld Benchmark Puts World Models to the Test with Agent Players
Hugging Face researchers introduce PlayWorld, a benchmark that uses multimodal agent players to evaluate interactive video world m...
Rhetoric Can Hack AI Peer Review: New Study Reveals Structured Sensitivity
A new study from Hugging Face shows that rhetorical framing biases AI review scores in structured ways, with evidence framing and...
Hugging Face Unveils Evoke: An Interactive World Model with Persistent Memory for Endless Video Generation
Evoke, a new interactive world model from Hugging Face, uses external persistent memory and a redesigned long-horizon teacher to e...
Hugging Face Unveils Spatial Memory Agent: Self-Evolving Spatial Reasoning for Frozen VLMs
A new framework from Hugging Face enables frozen vision-language models to improve spatial reasoning through self-evolution, witho...
AutoDesign: Meta-Harness Optimization Boosts Long-Horizon Agentic Design
Hugging Face researchers introduce AutoDesign, a framework that uses a meta-harness optimizer to recursively improve a code agent...
DarwinX: Evolving Agent Harnesses via Natural Selection Boosts Benchmarks
Hugging Face researchers introduce DarwinX, a method that evolves agent harnesses through population selection with frozen models,...
Hugging Face Unveils Intern-S2-Preview: A Scientific Agentic Foundation Model Series
Intern-S2-Preview integrates multimodal pre-training, multi-task reinforcement learning, and memory-augmented extensions to suppor...
Hugging Face Unveils LLMRouter: A Unified Infrastructure for LLM Routing
LLMRouter formalizes LLM routing as a sequential decision process, offering a unified benchmark (xRouteBench) and modular infrastr...
Hugging Face Unveils DreamX-Phi 1.0: A World Model for Faithful Robot Video Prediction
DreamX-Phi 1.0, an action-conditioned video world model for robotic manipulation, predicts future observations from language instr...
Simplax: A New Augmentation for Uniform Discrete Diffusion Models
Researchers introduce Simplax, an exact Dirichlet-categorical augmentation that enriches uniform discrete diffusion without alteri...
Hugging Face Researchers Unveil GazeAnywhere: Promptable Gaze Estimation with Text or Visual Cues
A new end-to-end transformer model, GazeAnywhere, enables gaze target estimation using natural language or visual prompts, elimina...
Hugging Face Researchers Launch MBA-Bench: A Multimodal Benchmark for Business Ideation Agents
A new benchmark, MBA-Bench, evaluates multimodal AI agents on business ideation, with proposed models MBA-b and MBA-k outperformin...