Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

360CityArena: New Benchmark Shows Huge Gap in AI Urban Navigation

AI By Crimson AI Hugging Face Papers 12 August 2026 · 00:00 9 views
Share: X Telegram

A new photorealistic benchmark built from 360-degree videos of Tokyo's Akihabara district reveals that even the best AI agents perform far below human levels in urban navigation and spatial reasoning.

360CityArena: New Benchmark Shows Huge Gap in AI Urban Navigation

Key points

Researchers have introduced 360CityArena, a new benchmark designed to evaluate the urban exploration capabilities of embodied AI agents in a photorealistic virtual environment. The benchmark is constructed from 602 real-world 360-degree video segments covering 85 streets in the Akihabara district of Tokyo, Japan, providing a realistic and complex setting for testing navigation and spatial reasoning.

The benchmark includes 175 meticulously human-crafted tasks across three categories: Environment Understanding, Path Reasoning, and Spatial Reasoning. These tasks cover fundamental abilities such as localization, landmark search, path planning, and relational spatial reasoning, enabling a comprehensive assessment of an agent's ability to operate in realistic urban scenes.

In evaluations using state-of-the-art large multimodal model (LMM)-based agents, the results show a significant performance gap. The strongest model tested, Gemini 2.5 Flash, achieved only 17.1% accuracy, compared to 77.3% for human participants. This stark difference highlights the substantial challenges that remain in city-scale embodied navigation and reasoning.

The authors argue that existing outdoor benchmarks often lack photorealism or complexity, making them insufficient for real-world applications. 360CityArena aims to fill this gap by providing a necessary and challenging testbed for photorealistic urban-district navigation and spatial reasoning, pushing the field toward more capable and robust embodied agents.

MetricHumanGemini 2.5 Flash
Accuracy77.3%17.1%
Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1