arXiv:2608.03464v2 Announce Type: replace
Abstract: Annotations are essential to communicative visualization, helping explain data, emphasize key findings, and guide attention. While multimodal large...
By Zhenghan Chen, Zekai Shao, Lidan Tan, Xin Lin, Xingchen Zeng, Yi Shan, Ziyue Lin, Xiaoliang Fu, Xinyuan Liu, Yuetong Guo, Fen Wang, Bongshin Lee, Siming Chen
arXiv:2609.24172v1 Announce Type: new
Abstract: Legends are fundamental to chart understanding, as reliable interpretation requires correctly binding legend entries to corresponding visual marks. Whi...
By Xinnuo Zhang, Zhike Tang, Jing Xu, Haoyuan Zhao, Weikai Yang
arXiv:2606. 27383v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as research assistants, yet it remains unclear whether they can calibrate research takeaways to the strength and scope of the supporting evidence.
By Yu Fu, Yongqi Kang, Yong Zhao
arXiv:2605. 21347v3 Announce Type: replace Abstract: Diagnosing failures in LLM agents remains largely manual.
By Akshay Manglik, Apaar Shanker, Kaustubh Deshpande, Jason Qin, Yash Maurya, Veronica Chatrath, Vijay S. Kalmath, Levi Lentz, Yuan Xue
VisAudit is a new benchmark that tests multimodal agents on visual diagnosis, repair, and verification tasks. It presents agents with rendered charts and auxiliary evidence—such as source data, intended summaries, and code—to iteratively detect defects, modify the visualization, and confirm successful repairs. The benchmark includes 1,900 flawed charts across 21 types and 10 flaw categories, plus 300 correct charts, and shows that current models recover only about 47.4% of flawed charts autonomously.
By Shicheng Liu, Adam Kahirov, Qi Zhang, Zhimin Hu, Song Wang, Junhong Lin, Julian Shun, Yada Zhu
arXiv:2607. 16122v1 Announce Type: new Abstract: Evaluations should do more than measure a models current performance.
By Vipul Gupta, Zihao Wang, Razvan-Gabriel Dumitru, MohammadHossein Rezaei, Aakash Sabharwal, Yunzhong He
The paper introduces LayerWiseBench, a benchmark that evaluates visual language models on layer-wise chart understanding and editing. It focuses on three core concepts—layer attribution, layer binding, and visibility ordering—by pairing rendered charts with spatially aligned per-layer RGBA assets and functional role labels. The benchmark includes 2,800 charts, 7,329 understanding questions, and 53,791 editing variants, revealing that models excel at attribution and binding but struggle with visibility ordering, especially when editing overlapping components.
arXiv:2609.08657v1 Announce Type: cross
Abstract: Charts are structured visual compositions whose elements have distinct functional roles, semantic correspondences, and visibility relations. This str...
By Xiaochuan Zhong, Yifan Hou, Chenxi Pang, Shaobo Cui
arXiv:2608. 04719v1 Announce Type: new Abstract: Agent evaluations tell us that a model picked the wrong tool, but rarely why.
By Atul Anand, Sourav Chattaraj
arXiv:2606. 00384v1 Announce Type: new Abstract: Fitting quantitative models to data is a central step in scientific workflows, yet it remains one of the least automated.
By William Rudman, Abhishek Divekar, Kanishk Jain, Sebastian Joseph, Stella S. R. Offner, Matthew Lease, Kyle Mahowald, Greg Durrett, Junyi Jessy Li
arXiv:2603. 29139v2 Announce Type: replace Abstract: Recent advances in large language models (LLMs) have enabled agentic systems to translate natural-language intent into executable scientific visualization (SciVis) tasks.
By Kuangshi Ai, Haichao Miao, Kaiyuan Tang, Nathaniel Gorski, Jianxin Sun, Guoxi Liu, Helgi I. Ingolfsson, David Lenz, Hanqi Guo, Hongfeng Yu, Teja Leburu, Michael Molash, Bei Wang, Tom Peterka, Chaoli Wang, Shusen Liu
ATP‑Bench proposes a new benchmark for evaluating agentic tool planning in multimodal large language models (MLLMs) that generate interleaved text-and-image responses. The benchmark contains 7,702 QA pairs, including 1,592 visual‑question‑answer pairs, across eight categories and 25 visual‑critical intents, all verified by humans. A Multi‑Agent MLLM‑as‑a‑Judge (MAM) system is introduced to assess tool‑call precision, missed opportunities, and overall response quality without relying on ground‑truth references.
By Yinuo Liu, Zi Qian, Heng Zhou, Jiahao Zhang, Yajie Zhang, Zhihang Li, Mengyu Zhou, Erchao Zhao, Xiaoxi Jiang, Guanjun Jiang