Research Papers
SHAPER: Train-Free Framework Evolves Embodied Agents via Skill-Harness Optimization
Hugging Face researchers introduce SHAPER, a train-free framework that improves embodied agents by evolving reusable skills and a...
SkillZip: New Graph Compression Method Boosts LLM Agent Skill Libraries by 3.46x
Hugging Face researchers introduce SkillZip, a contract-preserving graph compression framework that compresses procedural skills i...
Mechanist: Autonomous AI System Uncovers and Controls the Mechanisms of Intelligence
Hugging Face researchers introduce Mechanist, an agentic system that autonomously discovers and controls the mechanisms underlying...
Self-Geometry: Test-Time Adaptation Boosts 3D Vision Foundation Models Without Ground Truth
A new plug-and-play method, Self-Geometry, enforces explicit multi-view geometric constraints during test-time adaptation, improvi...
StateFlow: A Persistent 3D World State for Controllable Previsualization
StateFlow introduces a persistent 3D world state to enable iterative, controllable previsualization for film and game design, offe...
New Benchmark Reveals LLM Agents Struggle to Keep Stories Consistent Over Long Interactions
Researchers introduce NCP-Bench, a benchmark of 100 narrative environments, to evaluate long-horizon consistency in interactive st...
Hugging Face Unveils Spark-to-Paper: A Composable Skill for End-to-End Research Paper Generation
Spark-to-Paper, a new system from Hugging Face, generates complete research papers inside coding assistants using 13 composable sk...
AI4AI at Test-Time: Strong Models Boost Weak Ones Without Retraining
New research from Hugging Face shows that stronger AI models can build inference-time harnesses that nearly double the performance...
OpenART: A New Arena for Scaling AI Agent Red Teaming via Evolving Environments
Hugging Face researchers introduce OpenART, an open-ended arena with over 10,000 stateful scenarios, and EMHA, an attack policy th...
TSDS-Toolbox: A Unified Framework for Measuring Time-Series Dataset Similarity
Hugging Face researchers introduce TSDS-Toolbox, a unified, extensible framework for reproducible comparison of time-series datase...
360CityArena: New Benchmark Shows Huge Gap in AI Urban Navigation
A new photorealistic benchmark built from 360-degree videos of Tokyo's Akihabara district reveals that even the best AI agents per...
New SPIEval Benchmark Exposes Major Gaps in LLM Mobile Assistants
SPIEval, a new human-curated benchmark, evaluates LLMs as mobile assistants handling scattered personal data, revealing that even...