Research Papers
NAPE: A Minimalist Causal Transformer for Scalable Audio Self-Supervised Learning
Researchers introduce NAPE, a self-supervised audio learning framework that uses causal Transformers to predict next spectrogram p...
PolicyGuide: A New Framework to Enforce Policy Compliance in LLM Agents Across Entire Workflows
Hugging Face researchers introduce PolicyGuide, a framework that compiles domain policies into workflow graphs and uses a proactiv...
Hierarchical Self-Improvement: Evolving Agent Harnesses for Frozen LLMs
A new framework from HKUST lets a frozen LLM evolve its own task-specific execution harness and evolution strategy, yielding signi...
New NARU Benchmark Tests AI's Grasp of Japanese Long-Form Video Narratives
Researchers introduce NARU, a benchmark with 1,481 questions across 155 Japanese videos (146.8 hours) to evaluate narrative evolut...
TinyCast: A 146K-Parameter Zero-Shot Forecaster That Runs on Embedded Devices
Hugging Face researchers introduce TinyCast, an attention-free zero-shot forecaster with only 146,505 parameters that computes per...
New Study Quantifies Benchmark Optimization in ASR Models, Revealing Inflated Scores
A new paper from Hume AI introduces a methodology to quantify benchmark optimization in ASR models, showing that top-scoring model...
EXIMO: A Three-Stage Method to Efficiently Fine-Tune VLA Robot Policies
Researchers propose EXIMO, a novel algorithm that combines VLM-guided exploration, imitation learning, and residual off-policy RL...
CoToGrasp: New Framework Synthesizes Dexterous Grasps Conditioned on Contact Topologies
Researchers introduce CoToGrasp, a generative framework that synthesizes stable, diverse grasps conditioned on specific contact to...
GOAG: Object-Agnostic Grasp Planner Achieves 86.93% Success on MultiDex
Researchers introduce GOAG, a deep generative grasp planner that learns gripper-specific contact surfaces to generalize to unseen...
QuoteBench: Matched Scores Mask Command-Path Failures in LLM Coding Agents
A new benchmark, QuoteBench, reveals that execution-boundary parsing errors significantly reduce LLM coding agent success, and dis...
Chain-of-Experience: A New Paradigm for Continual LLM Improvement at Test Time
A new study introduces Chain-of-Experience (CoE), a test-time feedback loop that lets LLMs learn from iterative experience, outper...
LLMs and Embedding Models Tie on Quality, but Cost Gap Is Huge
A new study finds that LLMs and dedicated embedding models perform nearly identically across tasks, but LLMs cost up to 1,431x mor...