Latest stories
CoToGrasp: New Framework Synthesizes Dexterous Grasps Conditioned on Contact Topologies
Researchers introduce CoToGrasp, a generative framework that synthesizes stable, diverse grasps conditioned on specific contact to...
GOAG: Object-Agnostic Grasp Planner Achieves 86.93% Success on MultiDex
Researchers introduce GOAG, a deep generative grasp planner that learns gripper-specific contact surfaces to generalize to unseen...
QuoteBench: Matched Scores Mask Command-Path Failures in LLM Coding Agents
A new benchmark, QuoteBench, reveals that execution-boundary parsing errors significantly reduce LLM coding agent success, and dis...
Chain-of-Experience: A New Paradigm for Continual LLM Improvement at Test Time
A new study introduces Chain-of-Experience (CoE), a test-time feedback loop that lets LLMs learn from iterative experience, outper...
LLMs and Embedding Models Tie on Quality, but Cost Gap Is Huge
A new study finds that LLMs and dedicated embedding models perform nearly identically across tasks, but LLMs cost up to 1,431x mor...
τ₀-VLA: World-Model-Guided Test-Time Search Boosts Long-Horizon Robot Manipulation
Hugging Face researchers introduce τ₀-VLA, a hierarchical robot foundation model that uses world-model-guided test-time computatio...
IAR: A Three-Stage Post-Training Framework for Retrieval-Free Document Knowledge Internalization
Hugging Face researchers introduce IAR (Inject, Align, Recover), a staged post-training method that lets LLMs internalize document...
Reasoning in Greek: SFT Shifts Language, RL Fixes Defects, Accuracy Misses Both
A new study on fine-tuning MoE models to reason in Greek finds accuracy is a noisy, blind metric; SFT builds fluency while RLVR re...
FlashPrefill V2: Block-Sparse Attention Speeds Up Long-Context LLM Serving by Up to 47x
Hugging Face researchers present FlashPrefill V2, an optimized sparse attention operator for the prefill phase of long-context LLM...
Repo0: A New Framework for Zero-to-All Code Generation with Dual-DAG Architecture
Hugging Face researchers introduce Repo0, a framework that generates complete software repositories from natural-language requirem...
SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback
SkillEvo introduces a framework for continuous agent skill improvement by converting multi-turn user simulations into feedback gen...
New SWE-bench Science Benchmark Reveals Why Coding Agents Fail at Scientific Software Repair
Researchers introduce SWE-bench Science, a benchmark of 119 scientific software engineering tasks, showing that even top agents sc...