Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

UI-Mate: Open-Weight GUI Agent Sets New Benchmarks with In-Context Demonstrations

AI By Crimson AI Hugging Face Papers 18 August 2026 · 00:00 21 views
Share: X Telegram

Hugging Face researchers introduce UI-Mate, a foundation GUI agent that combines environment-grounded training with in-context demonstration learning, achieving state-of-the-art results on computer-use benchmarks and significantly improving long-horizon task reliability.

UI-Mate: Open-Weight GUI Agent Sets New Benchmarks with In-Context Demonstrations

Key points

Hugging Face researchers have unveiled UI-Mate, a foundation GUI agent designed to automate complex digital tasks with greater reliability. The system integrates an environment-grounded training stack with in-context demonstration learning, addressing key challenges such as scarce training data, ambiguous prompts, and unreliable execution.

The training stack uses a closed-loop data engine that automates task generation, environment construction, rollout, filtering, capability balancing, supervised fine-tuning, and online reinforcement learning across massively parallel environments. This is complemented by a mechanism that transforms multimodal demonstrations into flexible subtask-level workflows, allowing the agent to follow relevant steps and re-plan from the live interface.

UI-Mate sets a new open-weight state of the art on general computer-use benchmarks, scoring 77.0% on OSWorld-Verified and 66.2% on WindowsAgentArena. On the new OSWorkerBench benchmark—100 long-horizon office tasks across 41 applications—it achieves 41.0% strict success and 76.9% progress, outperforming its Qwen3.6-27B base by 17.7 and 24.5 points respectively.

Notably, in the 33-task self-demo subset, a single demonstration raises strict success from 17.2% to 35.4% and progress from 67.9% to 81.1%, underscoring the value of in-context demonstrations for long-horizon reliability. The project page is available at https://ui-mate.github.io.

BenchmarkUI-Mate-27BQwen3.6-27B (Base)Delta
OSWorld-Verified77.0%--
WindowsAgentArena66.2%--
OSWorkerBench (Strict Success)41.0%23.3%+17.7
OSWorkerBench (Progress)76.9%52.4%+24.5
Self-Demo Subset (Strict Success)35.4% (with demo)17.2% (without demo)+18.2
Self-Demo Subset (Progress)81.1% (with demo)67.9% (without demo)+13.2
Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 29 Aug 2026

Hugging Face Audit: 110 of 124 AI Evaluations Fail to Support Their Claims

A new commit-bound census of 124 Inspect Evals units reveals that 110 stop before deterministic inference due to missing historica...

4
Research paper
Hugging Face 29 Aug 2026

Aphanta: New Framework Diagnoses When Image Editing Boosts Multimodal Reasoning

Hugging Face researchers introduce Aphanta, a diagnostic framework that evaluates when image-editing intermediates improve multimo...

5
Research paper
Hugging Face 29 Aug 2026

EditaLive! Enables Real-Time Character Video Editing for Live Streaming

Hugging Face researchers introduce EditaLive, a framework for real-time human-centric video editing in live streams, achieving sta...

4