Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

New Benchmark OmegaUse-OfficeVal Tests LLM Agents on Cost-Effective Office Tasks

AI By Crimson AI Hugging Face Papers 30 July 2026 · 00:00 25 views
Share: X Telegram

Hugging Face researchers introduce OmegaUse-OfficeVal, a benchmark for evaluating LLM agents on long-horizon office-suite tasks with economic grounding, comparing agent performance against human labor costs and quality.

New Benchmark OmegaUse-OfficeVal Tests LLM Agents on Cost-Effective Office Tasks

Key points

Researchers at Hugging Face have unveiled OmegaUse-OfficeVal, a new benchmark designed to evaluate large language model (LLM) agents on long-horizon office-suite tasks with a focus on economic efficiency. The benchmark addresses a gap in existing evaluations, which often overlook the cost-effectiveness of AI agents in real-world office workflows.

The benchmark comprises 100 tasks derived from office-suite requests proposed by practitioners, adapted through a privacy-preserving process. On average, each task requires 2.32 hours of human labor to complete, providing a realistic baseline for comparison.

A key feature of OmegaUse-OfficeVal is its economic grounding: each task is paired with two economic signals—human labor time and a task price proxy. These signals enable direct comparisons between human costs and LLM inference costs, as well as value-weighted evaluation, allowing for a nuanced assessment of agent performance.

To ensure stable evaluation, the researchers developed code-based verifiers from fine-grained rubrics. They evaluated several frontier LLMs alongside a human baseline. The results indicate that while all evaluated LLMs are substantially cheaper and faster than human workers, they have not yet approached human-level deliverable quality.

The code and dataset are fully open-sourced, with more information available on the project website: https://omegause-officeval.github.io.

Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1