Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

New SPIEval Benchmark Exposes Major Gaps in LLM Mobile Assistants

AI By Crimson AI Hugging Face Papers 12 August 2026 · 00:00 7 views
Share: X Telegram

SPIEval, a new human-curated benchmark, evaluates LLMs as mobile assistants handling scattered personal data, revealing that even the best model achieves only 57.3% accuracy and that most failures stem from poor information localization.

New SPIEval Benchmark Exposes Major Gaps in LLM Mobile Assistants

Key points

Large language models (LLMs) are increasingly used as mobile assistants, but their ability to leverage personal information scattered across multiple apps remains poorly understood. To address this, researchers have introduced SPIEval, a human-curated benchmark designed to evaluate LLMs on tasks that require reasoning, disambiguation, integration, preference inference, and multi-intent decomposition.

SPIEval comprises 250 tasks spanning 4,335 personal records distributed across 10 apps, with multi-turn interaction supported through 21 tools. The benchmark is designed to feature diverse scenarios, challenging tasks, scattered information, controllable environments, and verifiable outcomes.

Evaluation of nine representative LLMs reveals substantial room for improvement. The best-performing model, GPT-5.5 (xhigh), achieved only 57.3% accuracy, while the weakest scored just 16.4%. Further analysis shows that 79% of failures stem from inaccurate information localization, as models often commit to plausible but incorrect information instead of continuing retrieval for verification.

Additionally, fewer than 2% of retrieval actions employ advanced search methods, and there is substantial variation in search efficiency across models. These findings expose fundamental limitations of current LLM-based mobile assistants and motivate future research. Data and code are available at https://huggingface.co/datasets/Junjie-Ye/SPIEval.

ModelAccuracy
GPT-5.5 (xhigh)57.3%
Weakest model16.4%
Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

0
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1