arXiv AI

ChartDensity-Bench: Benchmarking MLLMs for Numerical Data Reconstruction under Visual Density

ChartDensity-Bench is a benchmark designed to evaluate multimodal large language models (MLLMs) on their ability to reconstruct structured numerical data from scientific charts that vary in visual density. The benchmark uses charts paired with source-level ground-truth data and systematically changes the number of simultaneously presented charts (k = 1, 3, 6, 9) to assess how density affects reconstruction performance. A multi‑dimensional evaluation framework measures structural reliability, reconstruction completeness, parseability, and numerical fidelity, revealing that numerical reconstruction generally worsens as visual density increases, with varying degrees of degradation across different models.

arXiv AI
Aug 28

DEEPCHART: How Far are LLMs from Faithful Data-Science Chart Generation?

DEEPCHART is a new benchmark that evaluates large language models (LLMs) on faithful data‑science chart generation. It contains 1,482 expert‑annotated instances from scientific papers, financial filings, and ecosystem reports, and assesses chart creation through an Extract–Reason–Visualize pipeline. Experiments show that while LLMs can produce visually plausible charts, they frequently hallucinate data at the extraction and reasoning stages, especially in long, noisy, and multimodal contexts.

By Jiahui tang, Kuicai Dong, Dexun Li, Hongchao Gu, Haocheng Yu, Wei Han, Chen Zhang, Yong Liu, Hao Wang, Enhong Chen
arXiv AI
Sep 15

ChartAnno: Benchmarking Multimodal Large Language Models for Chart Annotation Generation

arXiv:2608.03464v2 Announce Type: replace Abstract: Annotations are essential to communicative visualization, helping explain data, emphasize key findings, and guide attention. While multimodal large...

By Zhenghan Chen, Zekai Shao, Lidan Tan, Xin Lin, Xingchen Zeng, Yi Shan, Ziyue Lin, Xiaoliang Fu, Xinyuan Liu, Yuetong Guo, Fen Wang, Bongshin Lee, Siming Chen
arXiv AI
Aug 5

ChartAnno: Evaluating MLLMs for Chart Annotation Generation

arXiv:2608. 03464v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have made significant progress in chart understanding, generation, and editing, but their ability to annotate existing charts remains underexplored.

By Zhenghan Chen, Zekai Shao, Lidan Tan, Xin Lin, Xingchen Zeng, Yi Shan, Ziyue Lin, Xiaoliang Fu, Xinyuan Liu, Yuetong Guo, Fen Wang, Bongshin Lee, Siming Chen
Hugging Face Trending Papers
Aug 4

ChartAnno: Evaluating MLLMs for Chart Annotation Generation

Multimodal large language models (MLLMs) have made significant progress in chart understanding, generation, and editing, but their ability to annotate existing charts remains underexplored. Annotating charts is a common yet challenging communicative task, requiring models to infer intended messages, interpret chart semantics, and place appropriate textual or graphical elements.

Hugging Face Trending Papers
Jul 29

SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context

Scientific images are the core elements of presenting experimental conclusions, elaborating system architecture, and supporting comparative arguments in scientific papers. However, existing image quality assessment (IQA) methods are predominantly designed for natural photographs or AI-generated content, which cannot be directly applied to scientific papers.