Research Papers
Hugging Face Unveils LLMRouter: A Unified Infrastructure for LLM Routing
LLMRouter formalizes LLM routing as a sequential decision process, offering a unified benchmark (xRouteBench) and modular infrastr...
Hugging Face Unveils DreamX-Phi 1.0: A World Model for Faithful Robot Video Prediction
DreamX-Phi 1.0, an action-conditioned video world model for robotic manipulation, predicts future observations from language instr...
Simplax: A New Augmentation for Uniform Discrete Diffusion Models
Researchers introduce Simplax, an exact Dirichlet-categorical augmentation that enriches uniform discrete diffusion without alteri...
Hugging Face Researchers Unveil GazeAnywhere: Promptable Gaze Estimation with Text or Visual Cues
A new end-to-end transformer model, GazeAnywhere, enables gaze target estimation using natural language or visual prompts, elimina...
Hugging Face Researchers Launch MBA-Bench: A Multimodal Benchmark for Business Ideation Agents
A new benchmark, MBA-Bench, evaluates multimodal AI agents on business ideation, with proposed models MBA-b and MBA-k outperformin...
Persistent Project Worlds Enable Autonomous Software Evolution, New EvoX Genesis Approach Shows
A new research paper introduces EvoX Genesis, which organizes long-horizon software development around a persistent project rather...
Hugging Face Study Exposes 'Illusion' in Visual Tool-Use by Multimodal LLMs
A new causal audit reveals that visual tool-use in multimodal LLMs often fails to causally influence answers, despite aggregate ac...
Hugging Face Researchers Unveil LDR: A Video World Model That Extrapolates Physics Beyond Training Data
A new paper from Hugging Face introduces Latent Dynamics Reasoning (LDR), a video world model that integrates kinematic dynamics i...
Hugging Face Researchers Unveil Closed-Loop Framework for Video Reflection Removal
A new closed-loop framework combines physics-based video synthesis, diffusion-based dereflection, and a dedicated benchmark, achie...
ToolHazard: New Framework Scales Adversarial Testing for LLM Agents Against Indirect Prompt Injections
Researchers introduce ToolHazard, a scalable framework that automatically synthesizes adversarial environments to expose and mitig...
New Benchmark Puts AI Coding Agents to the Test as Autonomous World-Model Researchers
Hugging Face researchers introduce AutoWorldModel-Bench, a closed-loop benchmark that evaluates frontier coding agents on open-end...
SHAPER: Train-Free Framework Evolves Embodied Agents via Skill-Harness Optimization
Hugging Face researchers introduce SHAPER, a train-free framework that improves embodied agents by evolving reusable skills and a...