Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

OSReward: Standardizing Evaluation of VLM Judges for Computer-Use Agents

AI By Crimson AI Hugging Face Papers 7 August 2026 · 00:00 14 views
Share: X Telegram

A new benchmark, OSReward, systematically evaluates VLM judges on computer-use agent trajectories, revealing a systematic leniency bias and a cost-accuracy trade-off. The authors release OS-Shepherd, open reward models that match commercial judges at 30-60% lower cost.

OSReward: Standardizing Evaluation of VLM Judges for Computer-Use Agents

Key points

Computer-using agents (CUAs) are rapidly advancing, but evaluating whether they successfully complete tasks remains a challenge. A CUA trajectory records the agent's actions, states, and reasoning, and verifying task fulfillment is crucial for evaluation, data curation, and reinforcement learning. Human verification is not scalable, so the field increasingly relies on vision-language models (VLMs) as judges. However, the reliability of these VLM judges has been largely unexamined.

To address this, researchers introduce OSReward, a realistic, high-quality benchmark that evaluates VLM judges on CUA trajectories. The trajectories come from diverse agent backbones executing human-verified instructions across platforms (web, Windows, Ubuntu, and mobile), with ground-truth verdicts established through multi-stage human annotation. The benchmark includes OSReward-Hard for challenging cases and OSReward-Multi for fine-grained efficiency and alignment scoring.

The most comprehensive evaluation of VLM judges to date reveals that even state-of-the-art models fall short of an ideal judge, exhibiting a systematic leniency bias that mislabels failed runs as successes. The few reliable models are too expensive for large-scale use, while affordable open models lag significantly behind.

To close this gap, the team constructs and releases OS-Shepherd-100K, an open corpus of reasoning-annotated trajectory judgments. They train OS-Shepherd (9B and 35B), open reward models that supply low-cost, stable, and reliable reward signals, matching commercial judges at 30-60% lower cost. The code, benchmark, dataset, and model checkpoints are publicly available.

Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1