Latest stories
β-OPSD: A Principled Generalization of On-Policy Self-Distillation for Reasoning Models
Researchers introduce β-OPSD, a generalization of on-policy self-distillation that turns a fixed KL penalty into a tunable paramet...
See2Think: New Benchmark Probes Whether Multimodal Models Truly Use Visual Reasoning States
Hugging Face researchers introduce See2Think, a unified evaluation framework with a 1,200-problem benchmark and Visual Action-of-T...
SpatialCLI: Teaching VLMs to Reason With Spatial Tools, Then Internalize Them
Hugging Face researchers introduce SpatialCLI, a three-stage framework that boosts spatial reasoning in vision-language models by...
RefCaptioner: New Framework Grounds Video Captions to Multiple Reference Images
Hugging Face researchers introduce RefCaptioner, a two-stage post-training framework for multi-reference image-grounded video capt...
Hugging Face Introduces PALATE: A Person-Aligned Benchmark for Evaluating Role-Playing Agents
A new benchmark from Hugging Face, PALATE, uses user simulators to evaluate role-playing agents in multi-turn conversations, addre...
Explorative Modeling: A New Pretraining Axis and End-to-End Generation
Researchers introduce Explorative Modeling, a paradigm that factors the training loop instead of the generation procedure, improvi...
INTACT: A Search-Free JEPA That Maps Intent Directly to Actions
Hugging Face researchers introduce INTACT, an end-to-end JEPA that learns intent-to-action mapping without test-time search, achie...
Σ-Mem: A New Reliability Memory That Helps LLM Multi-Agent Systems Adapt Without Retraining
Researchers introduce Σ-Mem, an online reliability memory that records which agents are trustworthy under different conditions, en...
MemHarness: Reconstructing Memory, Not Replaying It, Boosts LLM Agents
A new framework called MemHarness teaches LLM agents to reconstruct past experiences to fit the current context, outperforming sta...
LLMs Take on Algorithmic Trading: New Framework Outperforms Baselines in Parent-Order Execution
A new study introduces PACE, a hierarchical framework that uses large language models for parent-order execution, outperforming tr...
ShadowDancer: A New Method to Control Video World Models with Any Action
Hugging Face researchers introduce ShadowDancer, a method that learns unified dynamics representations from video pairs with resam...
ACE-Data-0: Turning Homes into Embodied AI Data Engines
Hugging Face researchers introduce ACE, a human-centric data engine that captures synchronized multisensory data in real homes, an...