Latest stories
UrbanGround: New Sandbox Reveals Limits of AI Agents in Real-Scale City Navigation
Researchers introduce UrbanGround, a realistic 3D replica of Hong Kong, to test whether multimodal AI agents can sustain navigatio...
PAWBench: Probabilistic Alignment in World Models Remains Elusive
A new benchmark, PAWBench, reveals that current video generation models fail to match reference behavior distributions, highlighti...
Hugging Face Proposes RLHEV: Using Game Engines as Verifiable Data Engines for World Models
A new paper from Hugging Face argues that scaling world models requires more than just more video data and compute. It proposes Re...
JIT-Agent: On-the-Fly Harness Synthesis Boosts Off-the-Shelf LLMs Beyond GPT-5.6
Hugging Face researchers introduce JIT-Agent, a trainable model that generates adaptive agent harnesses on the fly for any off-the...
WarpSAC: Regime-Aware Off-Policy RL Boosts Scalability and Sim-to-Real Transfer
Hugging Face researchers introduce WarpSAC, a family of off-policy RL algorithms that adapt stabilizers to data availability, impr...
FrontierChallenge: AI Agents Fail to Complete Scientific Workflows, New Benchmark Shows
A new benchmark, FrontierChallenge, reveals that even the best AI agents complete only 20.6% of end-to-end scientific workflows, d...
VoiceMem: A Dual-Brain Memory Architecture for Real-Time, Emotionally Aware Speech AI
Hugging Face researchers introduce VoiceMem, a streaming dual-brain memory system for speech language models that boosts retrieval...
VGI-Bench: New Benchmark Probes Visual Reasoning in Video Generation Models
Hugging Face researchers introduce VGI-Bench, a benchmark with 27 tasks and 810 instances to evaluate visual reasoning in video ge...
Hugging Face Unveils VBVR-Pro: A Scalable Testbed for Native Visual Reasoning
VBVR-Pro is a closed-loop testbed that makes native visual reasoning trainable, verifiable, and controllable, with 300 procedural...
Dual Nature of Generalization in On-Policy Distillation of LLMs
A new study reveals that on-policy distillation transfers reasoning behaviors rather than answers, with generalization strongly ti...
InfinityEdit: A Lightweight Adapter for Unbounded Video Editing
Researchers introduce InfinityEdit, a lightweight adapter that enables continuous, unbounded video editing by extending edits to f...
EviRank: Training-Free Multimodal Image Re-ranking via Structured Evidence
Hugging Face researchers introduce EviRank, a training-free method that reformulates multimodal image re-ranking as semantic const...