arXiv Computer Vision

Structure-Token Evidence-Anchored Reasoning for Scientific Chart Understanding

arXiv AI
Aug 28

DEEPCHART: How Far are LLMs from Faithful Data-Science Chart Generation?

DEEPCHART is a new benchmark that evaluates large language models (LLMs) on faithful data‑science chart generation. It contains 1,482 expert‑annotated instances from scientific papers, financial filings, and ecosystem reports, and assesses chart creation through an Extract–Reason–Visualize pipeline. Experiments show that while LLMs can produce visually plausible charts, they frequently hallucinate data at the extraction and reasoning stages, especially in long, noisy, and multimodal contexts.

By Jiahui tang, Kuicai Dong, Dexun Li, Hongchao Gu, Haocheng Yu, Wei Han, Chen Zhang, Yong Liu, Hao Wang, Enhong Chen
arXiv Computer Vision
Sep 22

Monitorable Chart Reasoning Agents via Verifiable Process Rewards

The paper introduces Chart‑RVR, a reinforcement‑learning framework that trains chart‑reasoning agents to produce monitorable, verifiable outputs. It decomposes reasoning into three auditable stages—Structure, Evidence, and Derivation—allowing stakeholders to trace how the model reads the chart, extracts data, and computes the answer. Experiments on six benchmarks show that Chart‑RVR matches or exceeds state‑of‑the‑art accuracy while delivering rationales that are far more verifiable and evidence‑grounded than existing methods.

By Sanchit Sinha, Oana Frunza, Kashif Rasul, Aidong Zhang
arXiv Computer Vision
Sep 22

Pay More Attention To Text In High-Resolution MLLMs

The paper introduces EviSpec, a training‑free compiler that generates complementary evidence specifications to improve high‑resolution multimodal large language models (MLLMs). By explicitly guiding visual search with structured evidence specifications, EviSpec achieves significant relative gains—up to 14.8% over random evidence—across five MLLMs and three benchmarks, and also sets new state‑of‑the‑art results on VQA and hallucination‑focused tasks.

By Zhongkuan Mao, Wenzhuo Zhao, Xianjie Liu, Yidong Wang, Zhao Gao, Ronghao Xian, Yao Jiang, Yi Zhang, Liangjian Wen, Keren Fu
arXiv AI
Jun 10

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering

arXiv:2510. 04514v3 Announce Type: replace Abstract: Recent multimodal LLMs have shown promise in chart-based visual question answering, but their performance declines sharply on unannotated charts-those requiring precise visual interpretation rather than relying on textual shortcuts.

By Rachneet Kaur, Nishan Srishankar, Zhen Zeng, Sumitra Ganesh, Manuela Veloso
arXiv AI
6d ago

Can Linguistic Reasoning Vectors Enhance Multimodal Reasoning Ability?

The paper introduces LIFT, a lightweight vector‑intervention technique that transfers reasoning capability from a base large language model (LLM) to a vision‑language model (VLM) without retraining the VLM backbone. LIFT defines Reasoning Vectors as differences in hidden states between a reasoning path with an explicit trace and a solver path without it, and injects these vectors into the VLM’s language‑side activations. Experiments on two VLMs across six reasoning benchmarks show that vectors derived from the base LLM consistently outperform those derived from the aligned VLM, indicating that the base LLM is a more effective source for recovering degraded reasoning. "whyItMatters":"The study demonstrates that a simple, frozen‑backbone intervention can partially restore reasoning abilities in multimodal models, highlighting the value of leveraging the original language model’s reasoning power."

By Ziyi Wang, Li Li, Aolin Zhou, Yankun Shen, Chonghan Liu, Shuxia Lin, Xu Yang