UrbanGround is a sandbox that tests how well multimodal large language model agents can translate local street‑view perception into reliable action within a physically realistic replica of Hong Kong. The platform offers closed‑loop first‑person interaction and an interactive map, allowing agents to navigate the 3D city and answer spatial questions. The study evaluates agents across three research questions—scene grounding, navigation over increasing distances, and robustness to route changes—revealing that while agents excel at visual recognition and short‑range reasoning, they struggle with sustained goal‑directed behavior and pedestrian‑aware movement.
By Tianjie Ju, Zheng Wu, Yueqing Sun, Yuhan Cui, Bobo Li, Shengqiong Wu, Pengzhou Cheng, Haodong Zhao, Zongru Wu, Xinbei Ma, Doris Zhang, Kunling Li, Mong-Li Lee, Wynne Hsu, Hao Fei, Qi Gu, Gongshen Liu, Zhuosheng Zhang
Robots deployed in delivery, campus, and emergency-response settings often need to navigate from buildings to streets within a single continuous episode. Existing benchmarks usually evaluate indoor and outdoor navigation separately, and many abstract away robot execution, leaving exit finding, boundary traversal, adaptation, and kinodynamic failures underexplored.
arXiv:2608.29483v1 Announce Type: cross
Abstract: Modern Vision-Language Models (VLMs) perform well above the human baseline in image geolocalization, a task critically important in disaster response...
By Arka Mukherjee, Soham Roy, Kartikeya Trivedi, Shreya Ghosh
arXiv:2606. 09669v1 Announce Type: new Abstract: Spatial reasoning is a foundational capability for multimodal large language models (MLLMs) to perceive and operate within the physical world.
By Hongcheng Gao, Hailong Qu, Jingyi Tang, Jiahao Wang, Zihao Huang, Hengkang Qiao, Shihong Huang, Junming Yang, Yi Li, Hongyixuan Yuan, Wenjie Li, Bohan Zeng, Wenbo Li, Bo Wang, Jianhui Liu, Olive Huang, Haoyang Huang, Wentao Zhang, Guoqing Huang, Nan Duan, Yinpeng Dong
arXiv:2510.23576v2 Announce Type: replace-cross
Abstract: Urban micromobility applications, such as delivery robots, demand reliable navigation across large-scale urban environments while following l...
By Anqi Li, Zhiyong Wang, Jiazhao Zhang, Minghan Li, Yunpeng Qi, Zhibo Chen, Zhizheng Zhang, He Wang
arXiv:2606. 06147v1 Announce Type: new Abstract: End-to-end Vision-Language-Action (VLA) models have shown promise in UAV navigation.
By Shengtao Zheng, Kai Li, Weichen Zhang, Yu Meng, Chen Gao, Xinlei Chen, Yong Li, Xiao-Ping Zhang