arXiv Computation and Language By Tianhao Niu, Qingfu Zhu, Wanxiang Che

Inside VLM Chart Reading: Tracing Value Reading from Vertical Bar Charts Across Space and Depth

Read the original on arXiv Computation and Language →

The paper investigates how vision‑language models (VLMs) read exact values from vertical bar charts. Using controlled counterfactual activation patching on Qwen2.5VL‑7B‑Instruct and InternVL3.5‑8B, the authors find that the bar‑top region contributes more to answer accuracy than the bar body, and that legend and series states lose recoverability earlier than geometry and axis‑scale states. They also show that prompt‑series positions mediate legend influence and that the two models differ in how they combine geometry and scale information from separate image donors.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computer Vision
Sep 23

Reading Right, Answering Wrong: How Visual Configuration Changes Affect Evidence Use in VLMs

Vision‑language models (VLMs) can lose accuracy when images are resized, even with minimal changes. The study shows that such small visual configuration changes—like tiling or token arrangement—cause more correctness flips across multiple checkpoints and benchmarks. Interestingly, in many cases the models still read the correct answer but fail to use it, and attention interventions reveal that configuration shifts weaken the use of readable information. By guiding models with field cues and their own transcriptions, the authors correct 97.2% of these errors.

By Dingyang Lin, Yingfeng Luo, Chenglong Wang, Chenwei Zhu, Anxiang Ma, Jingbo Zhu, Tong Xiao
arXiv Computer Vision
Sep 3

Who Drives the Probability Game of VLMs? A Temporal Causal Drive Evaluation Framework

The paper introduces a causal and temporal evaluation framework for vision‑language models (VLMs) that tracks how visual input, question text, and generated prefixes influence autoregressive decoding. It defines three step‑indexed causal‑drive metrics—Visual Causal Drive (VCD), Question Causal Drive (QCD), and Prefix Causal Drive (PCD)—using a Structural Causal Model and interventions. Experiments on Qwen3‑VL‑8B‑Instruct and other datasets show a shift from early question and visual guidance to increased reliance on generated prefixes, and demonstrate that QCD and PCD reduce recovery error and improve bias detection.

By Shuyao Xiao, Shengling Wang, Haoyu Niu, Ke Chao, Changwei Xu, Xinran Duan, Chaoyong Jiang
Hugging Face Trending Papers
Jul 14

Visual Access Boundaries in Vision-Language Model Reasoning

Chain-of-Thought (CoT) prompting is widely used as a test-time scaling strategy for Vision-Language Models (VLMs), but it remains unclear what is extended when VLMs generate longer reasoning traces. We ask whether CoT requires continued access to image tokens, or whether it mainly operates over visual information already made available earlier in the forward pass.

arXiv Computer Vision
Sep 21

From Retrieval to Recognition:How Vision--Language Models Become OCR Specialists

The paper investigates how general vision‑language models (VLMs) develop specialized optical character recognition (OCR) capabilities. By applying a causal intervention protocol, the authors identify sparse, stable OCR‑head sets in several VLMs and show that these heads largely overlap with textual retrieval/copy heads found in general VLMs. The study concludes that full‑sequence OCR functions as a dense multimodal copy‑and‑paste mechanism, and that when a VLM is fine‑tuned for OCR, it largely preserves the same head identities while redistributing their functional and causal strengths.

By Yuanxiang Huangfu, Hanmeng Zhong, Linqing Chen, Jeffrey Tiong Jee Hui