Research Papers
SA-MRPO: A Saturation-Aware Approach to Multi-Reward RL for Language Models
Researchers introduce Saturation Aware Advantage Reweighting (SA-MRPO), a method that standardizes each reward objective independe...
StateM Runtime Boosts GPT-5.6 to 95.3% Accuracy on Terminal-Bench 2.1
A new runtime system called StateM improves long-horizon agent performance without changing model weights, achieving 95.3% raw acc...
GenRouter: Adaptive Routing Cuts Image-Gen Costs by 95%
Hugging Face researchers introduce GenRouter, a unified routing framework that adaptively assigns prompts to optimal agentic image...
Hugging Face Proposes ACID-Compliant Framework for Reliable LLM Agents
A new research paper from Hugging Face introduces 'agentic transactions,' reinterpreting ACID database guarantees for LLM agents,...
UI-Mate: Open-Weight GUI Agent Sets New Benchmarks with In-Context Demonstrations
Hugging Face researchers introduce UI-Mate, a foundation GUI agent that combines environment-grounded training with in-context dem...
Hugging Face Unveils ClawGym II: Black-Box RL Framework for Agent Harness Optimization
A new black-box reinforcement learning framework from Hugging Face enables stable, scalable optimization of general agents through...
HarnessEval-W: Agentic Framework for Transparent World Model Evaluation
Hugging Face researchers introduce HarnessEval-W, an agentified evaluation pipeline that decomposes world-model assessments into v...
VibeWorlding: RL-Trained Open-Source Agents Outperform Closed-Source in 3D World Building
Hugging Face researchers introduce VibeWorlding, a unified framework that benchmarks and trains multimodal agents to construct 3D...
Moonshot AI Unveils Kimi K2.5: Open-Source Visual Agentic Intelligence with Swarm Capabilities
Moonshot AI introduces Kimi K2.5, a native multimodal open-source model with advanced coding and vision, featuring a self-directed...
Moonshot AI Unveils WorldVQA Benchmark to Test Visual World Knowledge in Multimodal LLMs
Moonshot AI releases WorldVQA, a benchmark with 3,500 image-question pairs designed to measure factual visual knowledge in multimo...
Kimi Launches Agent Swarm: 100 AI Agents Self-Organize to Tackle Complex Tasks
Kimi (Moonshot) unveils Agent Swarm, a research preview that lets K2.5 deploy up to 100 parallel sub-agents that self-organize int...
Moonshot AI Launches PerceptionBench to Isolate and Measure Atomic Visual Perception in MLLMs
Moonshot AI releases PerceptionBench, a new benchmark that evaluates multimodal large language models on ten atomic visual percept...