arXiv:2606. 07861v1 Announce Type: cross Abstract: Recent vision-language models (VLMs) excel at multimodal understanding and reasoning, yet their fine-grained visual perception remains underexplored.
By Lujun Li, Lama Sleem, Niccolo Gentile, Yangjie Xu, Yewei Song, Wenbo Wu, Radu State
Large vision-language models can recognize the objects and attributes in a crowded scene yet assign an attribute to the wrong same-class instance. Generic visual-question-answering accuracy marks the...
The paper introduces a perception interface that separates vision from language in vision‑language models. A frozen perception stack detects objects, a deterministic semantic serializer converts the perceived state into text, and a standard text‑only large language model (LLM) answers questions. Experiments on a campus‑robot benchmark show that this serialized interface outperforms a zero‑shot VLM of the same language‑model size, especially as the language model shrinks, and that the advantage persists under paraphrase and different supervision regimes.
By Cong Xu, Ravi Sankar
The paper shows that chain‑of‑thought (CoT) instructions can distort evaluation of vision‑language models (VLMs) when a scorer reads answer‑label logits before the model generates a rationale. On ScienceQA, Qwen2.5‑VL‑7B’s accuracy falls from 80.76% to 45.48% under this CoT‑prefix scoring, and most predictions incorrectly pick the first option. Linear probes and free generation recover most of the lost accuracy, indicating that the answer information remains in the hidden states but is missed by the early readout. The authors explain the mismatch with vocabulary and layer diagnostics, noting that probability mass shifts toward continuation tokens while answer information stays linearly accessible in later layers. The effect varies across datasets and models, but the study demonstrates that CoT‑prefix scoring can misrepresent model knowledge unless the requested and scored outputs are aligned.
By Zeyan Li, Siyuan Qiu, Jianfeng Xu
arXiv:2607. 15565v1 Announce Type: cross Abstract: Where should the question go in a vision-language model (VLM) prompt: before the image or after it?
By Rakshanda Hassan Abhinandan, John Galeotti, Deva Ramanan, Gautam Rajendrakumar Gare
The paper introduces a tool‑augmented framework that enhances a small Vision‑Language Model (Qwen3.5‑4B) with geometric tools—3D object detection, metric depth estimation, and deterministic solvers for distance, size, and bearing—to improve metric spatial reasoning. By moving metric computation from the model’s weights into explicit solvers, the approach achieves significant gains on ReVSI‑Bench tasks, notably increasing absolute distance accuracy from 0.46 to 0.74 MRA and relative direction accuracy from 25.9% to 73.4%. The modular design allows swapping in different detectors, enabling a clear separation between perception and reasoning errors, and the model can autonomously sequence the tools to match a scripted pipeline on most tasks.
By Kai Glantz, Clemens Grange
arXiv:2609.38777v1 Announce Type: new
Abstract: A central goal of vision-language model (VLM) distillation is to transfer both the teacher's language capabilities and its visual understanding. Howeve...
By Yuanhao Sun, Huawei Ji, Jiaxin Ding, Luoyi Fu, Xinbing Wang
arXiv:2607. 13305v1 Announce Type: cross Abstract: Benchmark accuracy in video large language models (LLMs) is often treated as evidence of visual understanding.
By Jae Joong Lee
arXiv:2607. 12815v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting is widely used as a test-time scaling strategy for Vision-Language Models (VLMs), but it remains unclear what is extended when VLMs generate longer reasoning traces.
By Hiroto Osaka, Shohei Taniguchi, Gouki Minegishi, Kai Yamashita, Masahiro Suzuki, Yutaka Matsuo
arXiv:2608. 16805v1 Announce Type: cross Abstract: Large vision-language models can recognize the objects and attributes in a crowded scene yet assign an attribute to the wrong same-class instance.
By Yuanzhi Xu, Qian Gao, Jun Fan, Guohui Ding, Zhenyu Yang, Yuteng Xiao, Sixue Lin
arXiv:2608.16081v2 Announce Type: replace
Abstract: Open-weight and frontier vision-language models (VLMs) perform well on general image understanding, but their ability to interpret fine-grained han...
By Taegang Kim, Saleh Afroogh, Junfeng Jiao
arXiv:2609.13308v1 Announce Type: cross
Abstract: A companion evaluation found that naming the target part in a manipulation prompt increased action accuracy by 0.32-0.63 across eight vision-language...
By Sarthak Sattigeri