arXiv Computation and Language By Sunisth Kumar, Xanh Ho, Tim Schopf, Andre Greiner-Petter, Florian Boudin, Akiko Aizawa

Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification

Read the original on arXiv Computation and Language →

The study investigates why multimodal large language models (VLMs) perform better at verifying scientific claims when evidence is presented as a table rather than a chart, despite both formats containing the same data. Using layer‑wise linear probing and attention analysis on three open‑weight VLMs, the authors find that chart information is indeed encoded in intermediate representations but never reaches the prediction layer, a gap absent for tables. Attention patterns reveal that this disconnect manifests differently across model families, suggesting the issue lies in how encoded visual data is utilized at prediction time rather than in the encoding process itself.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Aug 28

DEEPCHART: How Far are LLMs from Faithful Data-Science Chart Generation?

DEEPCHART is a new benchmark that evaluates large language models (LLMs) on faithful data‑science chart generation. It contains 1,482 expert‑annotated instances from scientific papers, financial filings, and ecosystem reports, and assesses chart creation through an Extract–Reason–Visualize pipeline. Experiments show that while LLMs can produce visually plausible charts, they frequently hallucinate data at the extraction and reasoning stages, especially in long, noisy, and multimodal contexts.

By Jiahui tang, Kuicai Dong, Dexun Li, Hongchao Gu, Haocheng Yu, Wei Han, Chen Zhang, Yong Liu, Hao Wang, Enhong Chen
arXiv AI
2d ago

LUMOS: Tracing Parametric Knowledge from Training Data to Behavioral Outputs in LLMs

LUMOS is a diagnostic framework that tracks how knowledge in large language models (LLMs) originates from training data and manifests in outputs, using the fully transparent OLMo 2 corpus. The study finds that while models encode rare facts with high separability (84%), they often fail to express them behaviorally (54%), and self‑reflection accuracy drops sharply on unseen content. These results show that incorporating the training‑data axis into evaluation turns speculative claims into verifiable evidence, suggesting it should become a standard part of LLM knowledge assessment.

By Seoyeon Ye, Gayoung Kim, Jiyoung Hong, Sookyung Kim, Hyunsoo Cho
arXiv Computer Vision
Sep 22

Pay More Attention To Text In High-Resolution MLLMs

The paper introduces EviSpec, a training‑free compiler that generates complementary evidence specifications to improve high‑resolution multimodal large language models (MLLMs). By explicitly guiding visual search with structured evidence specifications, EviSpec achieves significant relative gains—up to 14.8% over random evidence—across five MLLMs and three benchmarks, and also sets new state‑of‑the‑art results on VQA and hallucination‑focused tasks.

By Zhongkuan Mao, Wenzhuo Zhao, Xianjie Liu, Yidong Wang, Zhao Gao, Ronghao Xian, Yao Jiang, Yi Zhang, Liangjian Wen, Keren Fu