Most computer-vision systems organize visual input toward a predefined interpretation, such as semantic categories, prompted regions, learned object-like representations, or a single spatial partition...
Hierarchical 3D scene graphs are a promising representation for high-level spatial reasoning in autonomous mobile platforms. However, existing extraction frameworks typically rely on purely local visual clustering or strict geometric heuristics, such as wall-separated rooms, which fail in open-plan or arbitrarily-structured environments.
The paper introduces GeoSim, a four‑level framework for analyzing how vision‑language models (VLMs) represent low‑level vision tasks. It evaluates hidden‑layer representations across 24 tasks and two VLM paradigms—autoregressive models and diffusion transformers—using global similarity, local geometry, sparse feature decomposition, and topological verification. The study uncovers the organizing principles of low‑level visual representations and highlights their limitations in cross‑task and cross‑model agreement, offering an interpretability lens for assessing latent transferability and diagnosing model‑specific issues.
By Shao-Jun Xia, Huixin Zhang, Zhen Lei, Anlan Sun, Yuner Zhang, Xiaoyang Chen
arXiv:2608.30233v1 Announce Type: new
Abstract: Generalized Visual Grounding (GVG) task aims to localize targets in an image based on referring expressions, extends the classical visual grounding par...
By Kaiyan Lei, Xu-Yao Zhang
arXiv:2608. 10195v1 Announce Type: cross Abstract: Human vision organizes what it sees into wholes: same-colored points group into series, similar marks cohere into categories, and shapes complete into recognizable objects.
By Sudhanva Manjunath Athreya, Sai Phani Kumar Malladi
Generalized Visual Grounding (GVG) task aims to localize targets in an image based on referring expressions, extends the classical visual grounding paradigm by integrating multi-target and non-target...
We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the native text and image generation spaces of a unified multimodal model, without task-specific architectures. Under this formulation, SenseNova-Vision uses natural-language instructions and optional visual prompts to specify tasks, target regions or views, and decoding conventions, and generates responses as text for symbolic outputs, images for dense spatial predictions, or mixed text-and-image outputs for compositional tasks.
GrabVG is a visual grounding framework for UAV imagery that tackles the challenges of small, densely packed, and visually similar objects by separating the task into preattentive hypothesis search and graph-attentive feature binding. It first generates a compact set of reliable object hypotheses using distillation-guided proposal induction and text-aware filtering, then constructs a sparse graph where language-guided visual cues and inter-instance topological relationships are jointly bound and propagated via graph attention. Experiments on AerialVG and AerialSense demonstrate that GrabVG achieves a strong accuracy–speed trade‑off, reaching 67.31% and 80.34% Acc@0.5 and outperforming baselines by 10.55 and 8.76 percentage points.
GrabVG is a visual grounding framework for UAV imagery that tackles the challenges of small, densely packed, and visually similar objects. It splits the task into preattentive hypothesis search and graph‑attentive feature binding, using distillation‑guided proposals and a sparse graph to capture intra‑ and inter‑instance relationships. Experiments on AerialVG and AerialSense show that GrabVG achieves higher accuracy and speed, outperforming baselines by significant margins.
By Chaowei Wang, Yan Di, Jingjun Sun, Baozhe Liu, Jiaxu Tian, Yuheng Li, Guangqian Guo, Shan Gao
arXiv:2610.09440v1 Announce Type: new
Abstract: Pre-trained vision encoders contain layer-wise visual representations that differ in spatial granularity, semantic abstraction, and sensitivity to loca...
By Jeonghwan Kim, Sofia Stoica, Jiwan Chung, Ansel Blume, Hyeonjeong Ha, Zhenhailong Wang, Xin Luna Dong, Heng Ji
arXiv:2607. 05798v1 Announce Type: cross Abstract: Recent advancements in Multimodal Large Language Models (MLLMs) have evolved from static perception to interleaved visual-language reasoning, often referred to as ``thinking with images''.
By Yake Wei, Yuan Wang, Fengyun Rao, Jing Lyu, Di Hu
Open-vocabulary remote sensing segmentation has recently emerged as a promising paradigm that enables pixel-level recognition of arbitrary categories specified by natural language, including classes unseen during training. However, geospatial domain shifts caused by heterogeneous regions, spatial resolutions, and acquisition platforms weaken visual-text matching and limit cross-dataset generalization.