Research Papers
τ₀-VLA: World-Model-Guided Test-Time Search Boosts Long-Horizon Robot Manipulation
Hugging Face researchers introduce τ₀-VLA, a hierarchical robot foundation model that uses world-model-guided test-time computatio...
IAR: A Three-Stage Post-Training Framework for Retrieval-Free Document Knowledge Internalization
Hugging Face researchers introduce IAR (Inject, Align, Recover), a staged post-training method that lets LLMs internalize document...
Reasoning in Greek: SFT Shifts Language, RL Fixes Defects, Accuracy Misses Both
A new study on fine-tuning MoE models to reason in Greek finds accuracy is a noisy, blind metric; SFT builds fluency while RLVR re...
FlashPrefill V2: Block-Sparse Attention Speeds Up Long-Context LLM Serving by Up to 47x
Hugging Face researchers present FlashPrefill V2, an optimized sparse attention operator for the prefill phase of long-context LLM...
Repo0: A New Framework for Zero-to-All Code Generation with Dual-DAG Architecture
Hugging Face researchers introduce Repo0, a framework that generates complete software repositories from natural-language requirem...
SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback
SkillEvo introduces a framework for continuous agent skill improvement by converting multi-turn user simulations into feedback gen...
New SWE-bench Science Benchmark Reveals Why Coding Agents Fail at Scientific Software Repair
Researchers introduce SWE-bench Science, a benchmark of 119 scientific software engineering tasks, showing that even top agents sc...
ForgeWM: Progressive Causal Training Yields Fast, Controllable Video World Models
Hugging Face researchers introduce ForgeWM, a progressive framework that distills bidirectional video generators into few-step int...
MemTrapBench: New Benchmark Exposes How Memory Can Mislead LLMs
Hugging Face researchers introduce MemTrapBench, a benchmark revealing that retrieved memories can distort LLM reasoning and belie...
WithEveryone: New Framework Generates Group Images with Up to 10 Identities
Hugging Face researchers introduce WithEveryone, a unified framework for identity-preserving group image generation that grounds i...
FACET: New Framework for Synthesizing Executable Terminal Tasks for Agent Training
Hugging Face researchers introduce FACET, a framework that synthesizes high-quality terminal tasks by preserving source intent and...
Hugging Face Unveils 4DAnyone: Turning Casual Videos into 4D Human Reconstructions
4DAnyone reconstructs 4D humans from monocular video by generating multiview-consistent videos and lifting them into 4D Gaussian S...