arXiv Machine Learning By Shicheng Liu, Adam Kahirov, Qi Zhang, Zhimin Hu, Song Wang, Junhong Lin, Julian Shun, Yada Zhu

VisAudit: Evaluating Multimodal Agents for Visual Diagnosis and Repair

Read the original on arXiv Machine Learning →

VisAudit is a new benchmark that tests multimodal agents on visual diagnosis, repair, and verification tasks. It presents agents with rendered charts and auxiliary evidence—such as source data, intended summaries, and code—to iteratively detect defects, modify the visualization, and confirm successful repairs. The benchmark includes 1,900 flawed charts across 21 types and 10 flaw categories, plus 300 correct charts, and shows that current models recover only about 47.4% of flawed charts autonomously.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 17

Lexara-RF: Reference-Free Metrics for Evaluating Conversational Visual Analytics Agents

Lexara-RF introduces reference‑free metrics for evaluating conversational visual analytics agents that generate visualizations and natural‑language explanations. The framework uses only the prompt, data, and model response to score outputs, applying 13 metrics derived from visualization design theory and Gricean principles as consistency, intent‑alignment, and design validity checks. In tests against a human‑rated corpus, Lexara‑RF matches reference‑based methods, outperforms surface‑similarity NLG baselines, and accurately identifies structurally grounded failures.

By Srishti Palani, Vidya Setlur
arXiv AI
Jun 30

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents

arXiv:2603. 29139v2 Announce Type: replace Abstract: Recent advances in large language models (LLMs) have enabled agentic systems to translate natural-language intent into executable scientific visualization (SciVis) tasks.

By Kuangshi Ai, Haichao Miao, Kaiyuan Tang, Nathaniel Gorski, Jianxin Sun, Guoxi Liu, Helgi I. Ingolfsson, David Lenz, Hanqi Guo, Hongfeng Yu, Teja Leburu, Michael Molash, Bei Wang, Tom Peterka, Chaoli Wang, Shusen Liu
arXiv Computation and Language
Sep 1

Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence

The paper surveys Multimodal Code Intelligence, focusing on tasks where code is generated, edited, refined, or reasoned about under visually grounded inputs such as screenshots, charts, and videos. It categorizes the field by the role of code—rendered artifact, editable structure, intermediate reasoning trace, or executable tool interface—and organizes benchmarks into four domains: Graphical User Interface, Scientific Visualization, Structured Graphics, and Frontier Tasks and Frameworks. The authors argue that reliable evaluation must include evidence of semantics and interaction beyond visual fidelity, and propose four verification-centered research directions to advance the field toward evidence-grounded executable systems.

By Xuanle Zhao, Qiushi Sun, Jingyu Xiao, Xuexin Liu, Haoyue Yang, Qiaosheng Chen, Xianzhen Luo, Jing Huang, Yufeng Zhong, Lei Chen, Shuai Fu, Zhenlin Wei, Jinhe Bi, Lei Jiang, Haibo Qiu, Siqi Yang, Peng Shi, Jian Hu, Zhixiong Zeng
arXiv AI
Jul 17

MM-IssueLoc: A Controlled Benchmark for Evaluating Visual Evidence in Multimodal Repository-Level Issue Localization

arXiv:2607. 15205v1 Announce Type: cross Abstract: Real repository issues routinely include visual evidence such as screenshots, error dialogs, rendered UI states, and logs, yet repository-level issue localization is evaluated mostly as a text-only task.

By Shaoxiong Zhan, Shi Hu, Boyu Feng, Hai Lin, Andrew Gong, Zhengda Zhou, Jiaying Zhou, Yunyun Hou, Hao Su, Hai-Tao Zheng