Most computer-vision systems organize visual input toward a predefined interpretation, such as semantic categories, prompted regions, learned object-like representations, or a single spatial partition...
Hierarchical 3D scene graphs are a promising representation for high-level spatial reasoning in autonomous mobile platforms. However, existing extraction frameworks typically rely on purely local visual clustering or strict geometric heuristics, such as wall-separated rooms, which fail in open-plan or arbitrarily-structured environments.
The paper introduces GeoSim, a four‑level framework for analyzing how vision‑language models (VLMs) represent low‑level vision tasks. It evaluates hidden‑layer representations across 24 tasks and two VLM paradigms—autoregressive models and diffusion transformers—using global similarity, local geometry, sparse feature decomposition, and topological verification. The study uncovers the organizing principles of low‑level visual representations and highlights their limitations in cross‑task and cross‑model agreement, offering an interpretability lens for assessing latent transferability and diagnosing model‑specific issues.
By Shao-Jun Xia, Huixin Zhang, Zhen Lei, Anlan Sun, Yuner Zhang, Xiaoyang Chen
arXiv:2608.30233v1 Announce Type: new
Abstract: Generalized Visual Grounding (GVG) task aims to localize targets in an image based on referring expressions, extends the classical visual grounding par...
By Kaiyan Lei, Xu-Yao Zhang
arXiv:2608. 10195v1 Announce Type: cross Abstract: Human vision organizes what it sees into wholes: same-colored points group into series, similar marks cohere into categories, and shapes complete into recognizable objects.
By Sudhanva Manjunath Athreya, Sai Phani Kumar Malladi
Generalized Visual Grounding (GVG) task aims to localize targets in an image based on referring expressions, extends the classical visual grounding paradigm by integrating multi-target and non-target...