arXiv Computation and Language

Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification

The study investigates why multimodal large language models (VLMs) perform better at verifying scientific claims when evidence is presented as a table rather than a chart, despite both formats containing the same data. Using layer‑wise linear probing and attention analysis on three open‑weight VLMs, the authors find that chart information is indeed encoded in intermediate representations but never reaches the prediction layer, a gap absent for tables. Attention patterns reveal that this disconnect manifests differently across model families, suggesting the issue lies in how encoded visual data is utilized at prediction time rather than in the encoding process itself.

arXiv AI
Aug 28

DEEPCHART: How Far are LLMs from Faithful Data-Science Chart Generation?

DEEPCHART is a new benchmark that evaluates large language models (LLMs) on faithful data‑science chart generation. It contains 1,482 expert‑annotated instances from scientific papers, financial filings, and ecosystem reports, and assesses chart creation through an Extract–Reason–Visualize pipeline. Experiments show that while LLMs can produce visually plausible charts, they frequently hallucinate data at the extraction and reasoning stages, especially in long, noisy, and multimodal contexts.

By Jiahui tang, Kuicai Dong, Dexun Li, Hongchao Gu, Haocheng Yu, Wei Han, Chen Zhang, Yong Liu, Hao Wang, Enhong Chen
arXiv AI
2d ago

LUMOS: Tracing Parametric Knowledge from Training Data to Behavioral Outputs in LLMs

LUMOS is a diagnostic framework that tracks how knowledge in large language models (LLMs) originates from training data and manifests in outputs, using the fully transparent OLMo 2 corpus. The study finds that while models encode rare facts with high separability (84%), they often fail to express them behaviorally (54%), and self‑reflection accuracy drops sharply on unseen content. These results show that incorporating the training‑data axis into evaluation turns speculative claims into verifiable evidence, suggesting it should become a standard part of LLM knowledge assessment.

By Seoyeon Ye, Gayoung Kim, Jiyoung Hong, Sookyung Kim, Hyunsoo Cho
arXiv Computer Vision
Sep 22

Pay More Attention To Text In High-Resolution MLLMs

The paper introduces EviSpec, a training‑free compiler that generates complementary evidence specifications to improve high‑resolution multimodal large language models (MLLMs). By explicitly guiding visual search with structured evidence specifications, EviSpec achieves significant relative gains—up to 14.8% over random evidence—across five MLLMs and three benchmarks, and also sets new state‑of‑the‑art results on VQA and hallucination‑focused tasks.

By Zhongkuan Mao, Wenzhuo Zhao, Xianjie Liu, Yidong Wang, Zhao Gao, Ronghao Xian, Yao Jiang, Yi Zhang, Liangjian Wen, Keren Fu
arXiv Computer Vision
Aug 25

Investigating Relational Reasoning in VLMs

arXiv:2608.23518v1 Announce Type: new Abstract: Vision-Language Models (VLMs) achieve strong performance in visual reasoning tasks, but it remains unclear whether they understand visual relations, or...

By Adhithya Laxman Ravi Shankar Geetha, Aulia Kharis Rakhmasari, Haleema Ramzan, Xander Yap
arXiv AI
Aug 14

How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures

arXiv:2608. 13267v1 Announce Type: cross Abstract: Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty (how they behave when visual evidence is missing or misleading).

By Paul Osemudiame Oamen, Owusu-Banahene Osei, Ananya Mukherjee, Christian Greisinger, Steffen Eger, Pius Onobhayedo, Wei Zhao
arXiv AI
Aug 26

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes

The paper introduces CAIT, a benchmark of 400 synthetic scenes featuring counter‑intuitive actions that challenge multimodal large language models (MLLMs). Human participants and proprietary models like Claude and Gemini perform well, but standard open‑source instruction‑tuned MLLMs fail, largely due to a strong language prior that overrides contradictory visual evidence. The study shows that Chain‑of‑Thought reasoning can help but introduces new issues, while targeted fine‑tuning and structured prompting can reduce reliance on language priors and improve visual grounding.

By Chen Ling, Tongwei Zhang, Hanqian Li, Nai Ding
arXiv AI
Jul 28

Chart Deception in Vision-Language Models: From Vulnerability to Mitigation

arXiv:2607. 22600v1 Announce Type: new Abstract: Information visualizations are widely used to communicate patterns, trends, and outliers, yet deceptive design choices-such as truncated or inverted axes, distorted aspect ratios, inappropriate encodings, and misleading color mappings-can systematically alter interpretation while preserving the underlying data.

By Ridwan Mahbub, Mohammed Saidul Islam, Md Tahmid Rahman Laskar, Mizanur Rahman, Mir Tafseer Nayeem, Enamul Hoque