Latest stories
WarpSAC: Regime-Aware Off-Policy RL Boosts Scalability and Sim-to-Real Transfer
Hugging Face researchers introduce WarpSAC, a family of off-policy RL algorithms that adapt stabilizers to data availability, impr...
FrontierChallenge: AI Agents Fail to Complete Scientific Workflows, New Benchmark Shows
A new benchmark, FrontierChallenge, reveals that even the best AI agents complete only 20.6% of end-to-end scientific workflows, d...
VoiceMem: A Dual-Brain Memory Architecture for Real-Time, Emotionally Aware Speech AI
Hugging Face researchers introduce VoiceMem, a streaming dual-brain memory system for speech language models that boosts retrieval...
VGI-Bench: New Benchmark Probes Visual Reasoning in Video Generation Models
Hugging Face researchers introduce VGI-Bench, a benchmark with 27 tasks and 810 instances to evaluate visual reasoning in video ge...
Hugging Face Unveils VBVR-Pro: A Scalable Testbed for Native Visual Reasoning
VBVR-Pro is a closed-loop testbed that makes native visual reasoning trainable, verifiable, and controllable, with 300 procedural...
Dual Nature of Generalization in On-Policy Distillation of LLMs
A new study reveals that on-policy distillation transfers reasoning behaviors rather than answers, with generalization strongly ti...
InfinityEdit: A Lightweight Adapter for Unbounded Video Editing
Researchers introduce InfinityEdit, a lightweight adapter that enables continuous, unbounded video editing by extending edits to f...
EviRank: Training-Free Multimodal Image Re-ranking via Structured Evidence
Hugging Face researchers introduce EviRank, a training-free method that reformulates multimodal image re-ranking as semantic const...
Graph Engineering: A New Paradigm for Coordinating LLM Agents into System Intelligence
A new survey from Hugging Face introduces Graph Engineering, a paradigm that uses dynamic graph structures to organize multi-agent...
New Benchmark Reveals Omni-LLMs Struggle as Real-Time Video Assistants
Hugging Face researchers introduce OmniAssistBench, a benchmark for evaluating omni-modal LLMs as interactive video assistants. Re...
Hugging Face Proposes Compute-Efficient Hyperparameter Transfer for Large-Scale MoE Models
A new framework from Hugging Face predicts optimal learning rates for large Mixture-of-Experts models by scaling across widths and...
FlowEvo: Self-Evolving Agents Co-Develop Workflows and Skills at Inference Time
FlowEvo, a training-free framework from Hugging Face, enables LLM agents to co-evolve reusable skills and workflows during inferen...