GeoAgent: Evaluating VLM Geolocalization Through Embodied Navigation
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
UrbanGround is a sandbox that tests how well multimodal large language model agents can translate local street‑view perception into reliable action within a physically realistic replica of Hong Kong. The platform offers closed‑loop first‑person interaction and an interactive map, allowing agents to navigate the 3D city and answer spatial questions. The study evaluates agents across three research questions—scene grounding, navigation over increasing distances, and robustness to route changes—revealing that while agents excel at visual recognition and short‑range reasoning, they struggle with sustained goal‑directed behavior and pedestrian‑aware movement.
arXiv:2608.29880v1 Announce Type: new Abstract: Open-world geo-localization requires models to reason over ambiguous visual cues through multi-step reasoning and external knowledge grounding. While r...
GTPred is a new benchmark for geo‑temporal prediction that evaluates multi‑modal large language models (MLLMs) on 370 images taken across 120 years worldwide. It assesses predictions by matching both the year and a hierarchical location sequence, and includes annotated reasoning chains to test intermediate reasoning. Experiments on 15 MLLMs show that while visual perception is strong, models still lack world knowledge and geo‑temporal reasoning, and that adding temporal data improves location inference.
arXiv:2606. 06147v1 Announce Type: new Abstract: End-to-end Vision-Language-Action (VLA) models have shown promise in UAV navigation.
arXiv:2606. 07172v1 Announce Type: cross Abstract: Geospatial understanding is a critical yet underexplored dimension in the development of machine learning systems for tasks such as image geolocation and spatial reasoning.
arXiv:2608. 08814v1 Announce Type: cross Abstract: We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a photorealistic environment constructed from 360-degree videos.