Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

Tencent WorkBuddy Bench: A Contamination-Resistant Multi-Domain Coding-Agent Benchmark

AI By Crimson AI Hugging Face Papers 25 July 2026 · 00:00 19 views
Share: X Telegram

Tencent introduces WorkBuddy Bench, a multi-domain benchmark for coding agents with a contamination-resistant task construction method, covering Code, Web, Office, and Security domains.

Tencent WorkBuddy Bench: A Contamination-Resistant Multi-Domain Coding-Agent Benchmark

Key points

Tencent has unveiled WorkBuddy Bench, a novel multi-domain evaluation suite designed to assess coding agents across four distinct work domains: Code, Web, Office, and Security. The benchmark aims to provide a more realistic and contamination-resistant assessment of agent capabilities.

Unlike traditional benchmarks that adapt public issue text, WorkBuddy Bench constructs each task by reverse-engineering real commits, pull requests, or business scenarios. These are rewritten as short, colloquial, role-played requests, ensuring that the task prompt cannot be recovered by searching the original issue online. This approach, combined with dataset versioning, provides contamination resistance without relying on secrecy.

The benchmark is fully open-source, releasing task directories, environment images, evaluation harnesses, tests, and reference solutions. This allows third parties to re-run each task and inspect its content, ensuring reproducibility and auditability. All tasks are packaged in a uniform directory format and run under a reproducible protocol on two agent harnesses: CodeBuddy Code and Claude Code.

The four subsets probe complementary facets of real work: repository-level engineering, front-end development, office and business workflows, and red-/blue-team security. Each subset uses a different scoring instrument, so scores are not comparable across subsets, and the suite reports no overall average. A cross-model leaderboard across several model families is provided.

DomainDescriptionScoring Instrument
CodeRepository-level engineeringCustom per task
WebFront-end developmentCustom per task
OfficeOffice and business workflowsCustom per task
SecurityRed-/blue-team securityCustom per task
Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1