arXiv:2608.23720v1 Announce Type: new
Abstract: World models have demonstrated significant potential for perceiving and simulating complex environments. Despite their strong performance, the fundamen...
By Wenhow Li (The Hong Kong University of Science and Technology), Chengwei MA (The Hong Kong University of Science and Technology), Hui Xiong (The Hong Kong University of Science and Technology), Ying-Cong Chen (The Hong Kong University of Science and Technology), Lei Zhang (The Hong Kong University of Science and Technology)
Current VLM evaluations often conflate language priors with genuine spatial reasoning. To address this, we introduce CRISP, a novel structural-diagnostic evaluation paradigm that assesses visual spatial intelligence through consistency, the alignment between implicit perception and explicit reasoning.
arXiv:2606. 17536v1 Announce Type: cross Abstract: Generative world models for autonomous driving face two unresolved tensions: heterogeneous control injection, where free-form language, HD-maps, trajectories, and camera poses reside in incompatible representational spaces, and post-hoc cross-view fusion, where per-camera latents fail to encode global 3-D geometry.
By Zijie Meng, Yufei Liu, Chengqian Ma, Zhiyu Li, Jiyuan Liu, Wenhua Nie, Bingcai Wei, Shuqin Chen, Weichen Xu, Jiquan Yuan, Miao Zhang
Vision-language-action (VLA) models aim to map multimodal inputs to robot actions. However, most existing approaches struggle to cover complex dynamic scenarios due to treating all visual tokens uniformly and reasoning with human-selected factors, which lack mechanisms to emphasize task-critical evidence and ignore underlying factors.
arXiv:2606. 26535v1 Announce Type: cross Abstract: Current VLM evaluations often conflate language priors with genuine spatial reasoning.
By Zhixing Li, Yinan Yu
arXiv:2607. 08182v1 Announce Type: cross Abstract: Vision-language-action (VLA) models aim to map multimodal inputs to robot actions.
By Qi Lyu, Baicheng Liu, Xudong Wang, Jiahua Dong, Lianqing Liu, Zhi Han
arXiv:2605. 30705v2 Announce Type: replace-cross Abstract: Geometry-aware generative models and novel view synthesis approaches have shown strong potential in visual fidelity and consistency.
By Sunghyun Kim, Jaehoon Hahm, Jeongwoo Shin, Joonseok Lee
arXiv:2607. 12382v1 Announce Type: new Abstract: How can an agent build a structured map of its world from nothing but an ongoing sequence of raw sensory input and its own movements, especially when natural variation means exact sensory patterns rarely repeat?
By Arash Nikzad, Sasan Sarbishegi, Ali Dasmeh, Muhammad Asif, Parsa Gharavi, Erik Husom, Sagar Sen, Andrew B. Lehr, Olivier Penacchio, Ana Clemente, Tristan M. St\"ober
arXiv:2512. 12225v3 Announce Type: replace Abstract: Developing artificial agents that unify representation, memory, adaptation, and prediction remains a fundamental challenge in artificial intelligence.
By Laha Ale
arXiv:2603. 12231v2 Announce Type: replace Abstract: Learning good representations is essential for latent planning with world models.
By Ying Wang, Oumayma Bounou, Gaoyue Zhou, Randall Balestriero, Tim G. J. Rudner, Yann LeCun, Mengye Ren
arXiv:2602. 24181v2 Announce Type: replace-cross Abstract: Pre-trained vision encoders like DINOv2 have demonstrated exceptional performance on unimodal tasks.
By Rishabh Kabra, Maks Ovsjanikov, Drew A. Hudson, Ye Xia, Skanda Koppula, Andre Araujo, Joao Carreira, Niloy J. Mitra
arXiv:2608.30388v1 Announce Type: cross
Abstract: Cross-view video representation learning aims to capture viewpoint-invariant action semantics despite substantial appearance changes across egocentri...
By Youngchae Chee, Hosu Lee, Sungjune Park, Junho Kim, Yong Man Ro