Research Papers
Hugging Face Introduces PALATE: A Person-Aligned Benchmark for Evaluating Role-Playing Agents
A new benchmark from Hugging Face, PALATE, uses user simulators to evaluate role-playing agents in multi-turn conversations, addre...
Explorative Modeling: A New Pretraining Axis and End-to-End Generation
Researchers introduce Explorative Modeling, a paradigm that factors the training loop instead of the generation procedure, improvi...
INTACT: A Search-Free JEPA That Maps Intent Directly to Actions
Hugging Face researchers introduce INTACT, an end-to-end JEPA that learns intent-to-action mapping without test-time search, achie...
Σ-Mem: A New Reliability Memory That Helps LLM Multi-Agent Systems Adapt Without Retraining
Researchers introduce Σ-Mem, an online reliability memory that records which agents are trustworthy under different conditions, en...
MemHarness: Reconstructing Memory, Not Replaying It, Boosts LLM Agents
A new framework called MemHarness teaches LLM agents to reconstruct past experiences to fit the current context, outperforming sta...
LLMs Take on Algorithmic Trading: New Framework Outperforms Baselines in Parent-Order Execution
A new study introduces PACE, a hierarchical framework that uses large language models for parent-order execution, outperforming tr...
ShadowDancer: A New Method to Control Video World Models with Any Action
Hugging Face researchers introduce ShadowDancer, a method that learns unified dynamics representations from video pairs with resam...
ACE-Data-0: Turning Homes into Embodied AI Data Engines
Hugging Face researchers introduce ACE, a human-centric data engine that captures synchronized multisensory data in real homes, an...
MPIE-Bench: New Benchmark Exposes Anatomical Flaws in Multi-Person Image Editing
Hugging Face researchers introduce MPIE-Bench, a 2,500-sample benchmark that reveals persistent anatomical and geometric errors in...
Flux-OPD: New On-Policy Distillation Method Uses Evolving Contexts to Improve Open-Ended LLM Training
Researchers propose Flux-OPD, an on-policy distillation paradigm that leverages evolving contexts as in-training supervision to ca...
Memory Decoder at Scale: 6.9B Parametric Memory Boosts Small Models Past 12B Baselines
Researchers scale parametric long-term memory to 6.9B parameters and 300B tokens, showing that pairing a small backbone with a lar...
VideoCoCo: Using Blender Code as Chain-of-Thought for Physically Consistent Video Generation
Hugging Face researchers introduce VideoCoCo, an agentic dual-engine framework that uses executable Blender code as a process-leve...