arXiv:2609.36458v1 Announce Type: new
Abstract: Semantic-preserving transformations can induce substantial motion in learned representations, while small changes may strongly affect model predictions...
By Abdullah All Tanvir, Xin Zhong
arXiv:2602.02611v2 Announce Type: replace
Abstract: A prevailing paradigm in modern representation learning is the map-first approach, in which a representation map is learned from reconstruction, em...
By David Vigouroux (ANITI, IMT Atlantique - DSD, LaTIM), Lucas Drumetz (IMT Atlantique - MEE, Lab-STICC\_OSE, ODYSSEY), Ronan Fablet (IMT Atlantique - MEE, Lab-STICC\_OSE, ODYSSEY), Fran\c{c}ois Rousseau (IMT Atlantique - DSD, LaTIM)
arXiv:2606. 00124v1 Announce Type: cross Abstract: Positional embeddings (PEs) in Vision Transformers (ViTs) are known to impact performance and robustness, but their role in shaping internal spatial representations is not well understood.
By Mahmoud Mannes
Metric questions about video require vision-language models to use supplied real-world references to convert visual measurements into physical units. Yet we find that current models use this scale inf...
arXiv:2607. 02386v1 Announce Type: cross Abstract: While Vision Transformers have achieved remarkable success across computer vision and language applications, the geometric evolution of their internal representations throughout training remains insufficiently understood.
By Kaustubh Kapil, Kishor P. Upla
arXiv:2609.38285v1 Announce Type: cross
Abstract: Vision-language models (VLMs) can contradict themselves across views of the same spatial relation and fail to respond when that relation changes. Add...
By Hongbo Wang, Zihan Lin, Wenkui Yang, Shiran Ge, Yuang Ai, Jie Cao, Huaibo Huang, Ran He
The paper introduces EquiSD, a label‑free training method that exploits scale equivariance to improve metric grounding in vision‑language models. By projecting model predictions onto a scale‑equivariant family and fine‑tuning on the resulting targets, EquiSD boosts a 3B model’s median response slope from 0.66 to 0.94 and raises mean relative accuracy by 9.2 points across simulated scales, with positive transfer to real QuantiPhy videos.
By Kaizhen Tan, Yang Feng, Heqing Du, Siru Tao, Xin Xu, Hanzhe Hong
arXiv:2606.08918v2 Announce Type: replace
Abstract: Worldwide image geo-localization aims to determine where on Earth a single image was captured. However, visually similar scenes may lie thousands o...
By Junchao Cui, Xuanzi Ma, Wenqi Shi, Nan Wu, Biru Zhu, Xiangyang Luo
arXiv:2606. 05833v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) excel at 2D semantic understanding but lack intrinsic 3D awareness, resulting in representations that fail to maintain geometric and spatial consistency across video frames.
By Haibo Wang, Lifu Huang
arXiv:2609.06880v1 Announce Type: cross
Abstract: Reasoning over language instructions in embodied tasks such as robotics often requires understanding spatial relations from a speaker's situated pers...
By Mimo Shirasaka, Haochen Zhang, Yonatan Bisk
arXiv:2605. 16713v2 Announce Type: replace-cross Abstract: Modern Vision-Language Models (VLMs) achieve strong semantic recognition, yet remain brittle on elementary spatial relations such as left of, on, behind, and between.
By Renjie Gu, Kaichen Zhou, Yan Luo, Mengyu Wang
arXiv:2605. 20448v2 Announce Type: replace-cross Abstract: Vision-language models reliably name objects in a scene, but do they represent the 3D layout those objects inhabit?
By Animesh Maheshwari, Divyansh Sahu, Nishit Verma