Latest stories
Graph Engineering: A New Paradigm for Coordinating LLM Agents into System Intelligence
A new survey from Hugging Face introduces Graph Engineering, a paradigm that uses dynamic graph structures to organize multi-agent...
New Benchmark Reveals Omni-LLMs Struggle as Real-Time Video Assistants
Hugging Face researchers introduce OmniAssistBench, a benchmark for evaluating omni-modal LLMs as interactive video assistants. Re...
Hugging Face Proposes Compute-Efficient Hyperparameter Transfer for Large-Scale MoE Models
A new framework from Hugging Face predicts optimal learning rates for large Mixture-of-Experts models by scaling across widths and...
FlowEvo: Self-Evolving Agents Co-Develop Workflows and Skills at Inference Time
FlowEvo, a training-free framework from Hugging Face, enables LLM agents to co-evolve reusable skills and workflows during inferen...
NAPE: A Minimalist Causal Transformer for Scalable Audio Self-Supervised Learning
Researchers introduce NAPE, a self-supervised audio learning framework that uses causal Transformers to predict next spectrogram p...
PolicyGuide: A New Framework to Enforce Policy Compliance in LLM Agents Across Entire Workflows
Hugging Face researchers introduce PolicyGuide, a framework that compiles domain policies into workflow graphs and uses a proactiv...
Hierarchical Self-Improvement: Evolving Agent Harnesses for Frozen LLMs
A new framework from HKUST lets a frozen LLM evolve its own task-specific execution harness and evolution strategy, yielding signi...
New NARU Benchmark Tests AI's Grasp of Japanese Long-Form Video Narratives
Researchers introduce NARU, a benchmark with 1,481 questions across 155 Japanese videos (146.8 hours) to evaluate narrative evolut...
TinyCast: A 146K-Parameter Zero-Shot Forecaster That Runs on Embedded Devices
Hugging Face researchers introduce TinyCast, an attention-free zero-shot forecaster with only 146,505 parameters that computes per...
New Study Quantifies Benchmark Optimization in ASR Models, Revealing Inflated Scores
A new paper from Hume AI introduces a methodology to quantify benchmark optimization in ASR models, showing that top-scoring model...
EXIMO: A Three-Stage Method to Efficiently Fine-Tune VLA Robot Policies
Researchers propose EXIMO, a novel algorithm that combines VLM-guided exploration, imitation learning, and residual off-policy RL...
From Atari to EVE Online: Google DeepMind Marks 15 Years of AI Research in Games
Google DeepMind reflects on 15 years of AI research in games, from Atari to EVE Online, highlighting milestones like AlphaGo and S...