Researchers have introduced 360CityArena, a new benchmark designed to evaluate the urban exploration capabilities of embodied AI agents in a photorealistic virtual environment. The benchmark is constructed from 602 real-world 360-degree video segments covering 85 streets in the Akihabara district of Tokyo, Japan, providing a realistic and complex setting for testing navigation and spatial reasoning.
The benchmark includes 175 meticulously human-crafted tasks across three categories: Environment Understanding, Path Reasoning, and Spatial Reasoning. These tasks cover fundamental abilities such as localization, landmark search, path planning, and relational spatial reasoning, enabling a comprehensive assessment of an agent's ability to operate in realistic urban scenes.
In evaluations using state-of-the-art large multimodal model (LMM)-based agents, the results show a significant performance gap. The strongest model tested, Gemini 2.5 Flash, achieved only 17.1% accuracy, compared to 77.3% for human participants. This stark difference highlights the substantial challenges that remain in city-scale embodied navigation and reasoning.
The authors argue that existing outdoor benchmarks often lack photorealism or complexity, making them insufficient for real-world applications. 360CityArena aims to fill this gap by providing a necessary and challenging testbed for photorealistic urban-district navigation and spatial reasoning, pushing the field toward more capable and robust embodied agents.