arXiv Machine Learning By Xuhui Lin, Stephen Law, Nanjiang Chen, Kunyao Li, Tao Yang

A 3D Isovist World Model -- Revealing a City's Unseen Geometry and Its Emergent Cross-City Signature

Read the original on arXiv Machine Learning →

arXiv:2606. 03609v1 Announce Type: cross Abstract: Embodied agents that navigate cities rely on world models that predict how their surroundings will change as they move.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computer Vision
Aug 28

UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City

UrbanGround is a sandbox that tests how well multimodal large language model agents can translate local street‑view perception into reliable action within a physically realistic replica of Hong Kong. The platform offers closed‑loop first‑person interaction and an interactive map, allowing agents to navigate the 3D city and answer spatial questions. The study evaluates agents across three research questions—scene grounding, navigation over increasing distances, and robustness to route changes—revealing that while agents excel at visual recognition and short‑range reasoning, they struggle with sustained goal‑directed behavior and pedestrian‑aware movement.

By Tianjie Ju, Zheng Wu, Yueqing Sun, Yuhan Cui, Bobo Li, Shengqiong Wu, Pengzhou Cheng, Haodong Zhao, Zongru Wu, Xinbei Ma, Doris Zhang, Kunling Li, Mong-Li Lee, Wynne Hsu, Hao Fei, Qi Gu, Gongshen Liu, Zhuosheng Zhang
arXiv Computation and Language
4d ago

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

arXiv:2609.38177v1 Announce Type: cross Abstract: Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs...

By Jaewoo Jung, Hyeonseo Yu, Honggyu An, Jisang Han, Mungyeom Kim, Minkyeong Jeon, Heeseong Shin, Wonjun Moon, Federico Tombari, Daniel Barath, Marc Pollefeys, Seungryong Kim, Sunghwan Hong
arXiv Computer Vision
Sep 18

SceneTeract: Probing and Improving Agent-Aware Activity Reasoning in 3D Indoor Scenes

SceneTeract is a verification interface that separates semantic action understanding from physical feasibility in indoor 3D scenes. It decomposes activities into atomic actions and performs explicit geometric checks to determine executability, providing diagnostic traces for failures. The system reveals widespread functional and accessibility issues in synthetic scenes, shows that existing VLMs over‑predict action feasibility, and improves VLM performance through post‑training with verifier feedback, with benefits that generalize to real‑world scenes.

By L\'eopold Maillard, Francis Engelmann, Tom Durand, Boxiao Pan, Yang You, Leonidas Guibas, Maks Ovsjanikov