In a new research paper, scientists from Hugging Face and collaborators propose UrbanGround, a first-of-its-kind sandbox built on a georegistered 3D replica of Hong Kong. The environment allows multimodal large language models (MLLMs) to interact with a realistic city from a first-person perspective, complete with interactive maps and closed-loop control.
The study systematically evaluates whether current AI agents can turn local perception into reliable navigation. While MLLMs demonstrate strong atomic abilities in visual recognition and short-range spatial reasoning, the research reveals a critical gap: these skills do not compose into sustained goal-directed behavior over extended exploration.
As agents navigate longer distances, small spatial errors accumulate, leading to failures that are difficult to recover from. Dynamic changes, such as road closures and moving pedestrians, pose particular challenges, highlighting the fragility of current systems in complex urban environments.
UrbanGround is available as web and native builds, along with evaluation code and tasks. The researchers hope it will serve as a playground for studying how to bridge the gap between strong local perception and reliable city-scale agency.