arXiv:2606. 05525v1 Announce Type: new Abstract: Recent advances in agentic visualization have enabled the translation of natural language into executable scientific visualization (SciVis) workflows.
By Kuangshi Ai, Haichao Miao, Kaiyuan Tang, Shusen Liu, Chaoli Wang
OSWorld-Science is a benchmark and evaluation environment for computer-using agents that use visual language models (VLMs) to perform scientific software tasks. It includes 12 VLMs and 146 high-quality tasks across domains such as molecular drawing, pathology image analysis, statistical computing, and physical simulation, with artifact-based evaluation and a harness that logs interactions and supports model comparison. The benchmark was developed through expert proposals and iterative human–AI co‑design, and results show that current VLMs still struggle with key scientific questions, offering insights into factors like language, reasoning, and context length.
By Dingyuan Dai, Heli Qi, Lei Liu, Yinxi Li, Baiding Chen, Zijun Dou, Qingcheng Zeng, Qi Kang, Oliver Sun, Eric Wang, Bo Zhou, Haixin Wang, Yufan Du, Shi Bo, Ruihan Lin, Mengqi Yuan, Dunjie Lu, Steven Dillmann, Yiming Shi, Tina Su, Amy Xin, Minghao Liu, Xi Wang, Xu Huang, Ge Zhang, Pengyu Nie, Zhen Yang, Jie Tang, Juanzi Li, Weihao Xuan, Tianyu Liu
arXiv:2606. 26614v1 Announce Type: cross Abstract: Large language model (LLM) agents enable natural language interaction for scientific visualization (SciVis).
By Kuangshi Ai, Patrick Phuoc Do, Chaoli Wang
MatToolBench is a new benchmark that evaluates multimodal GUI agents on professional materials science software. It contains 204 tasks across 10 tools in three modalities—GUI operation, OriginPro scripting, and code-based database queries—executed inside a Windows 11 VM. The benchmark offers fine-grained, expert-decomposed scoring and a high-performing multimodal judge for aesthetic assessment, revealing that strong general benchmark performance does not transfer to scientific workflows.
By Mei Wu, Rui Xie, Runyu Zhang, Yuqiang Li, Tianfan Fu, Bo Chen, Kai Yu, Xin Chen, Lu Chen
arXiv:2604. 27996v3 Announce Type: replace Abstract: This paper examines how large language model (LLM) agents perform on scientific visualization (SciVis) tasks that require generating visualization workflows from natural-language instructions.
By Jackson Vonderhorst, Kuangshi Ai, Haichao Miao, Shusen Liu, Chaoli Wang
arXiv:2603. 01421v3 Announce Type: replace Abstract: While large language models accelerate scientific discovery, existing agents face severe limitations in adaptability, domain generalization, and multimodal scalability, often struggling to autonomously process raw, domain-specific experimental data.
By Ke Lin, Owais Aijaz, Yilin Lu, Yiyang Luo, Xuehang Guo, Preslav Nakov
arXiv:2606. 12736v1 Announce Type: new Abstract: AI agents are increasingly being developed to accelerate scientific discovery, yet their practical capabilities in real research settings remain poorly understood.
By Tianyu Liu, Allen Xin Wang, Antonia Panescu, Lisa Xinyi Chen, Wenxin Long, Xinyu Wei, Yueqian Jing, Ziyao Zeng, Jihang Chen, Sihan Jiang, Ziqing Wang, Siyi Gu, Siyu Chen, Xinyang Hu, Haoran Shao, Leqi Xu, Wangjie Zheng, Zhiyuan Cao, Ada Fang, Botao Yu, Kunyang Sun, Rex Ying, Arman Cohan, Qingyu Chen, Lingzhou Xue, Kaize Ding, Yuanqi Du, Wengong Jin, Zhuoran Yang, Marinka Zitnik, James Zou, Hua Xu, Hongyu Zhao
arXiv:2609.06192v1 Announce Type: new
Abstract: Scientific coding agents produce interdependent code, results, figures, and claims, yet evaluating final
outputs alone does not establish whether the...
By Bowen Liu, Shuo Nie, Bodong Du, Xiaomeng Li
arXiv:2607. 15176v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-centric and provide limited evidence of understanding of scientific visualization (SciVis).
By Patrick Phuoc Do, Chau M. Ta, Chaoli Wang
arXiv:2608. 03451v1 Announce Type: new Abstract: Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattered across databases, structured files, long documents, and multimedia.
By Boyan Li, Zhuowen Liang, Yupeng Xie, Xiaotian Lin, Tianqi Luo, Xinyu Liu, Yizhang Zhu, Zhangyang Peng, Yuan Li, Zhengxuan Zhang, Jiayi Zhang, Nan Tang, Guoliang Li, Yuyu Luo
SciDocBench is a workflow-centered benchmark for scientific document understanding that includes 124 expert-authored questions across seven capability groups and 19 subtasks in five scientific domains. Each question is evaluated under four conditions—English or Chinese, all-images-first or interleaved document representations—resulting in 496 evaluation instances. The benchmark is paired with SciDocIR, a typed evidence-graph representation, and SciDocDataset, a collection of 15K fine-tuning and 8K reinforcement-learning samples, forming an evaluation-to-training framework for scientific-document assistants.
By Shenxi Wu, Yuhong Liu, Haosong Zhang, Tongjin Zou, Yanxun Zhang, Gaochang Chen, Dun Liang, Jiaqi Wang, Zhecan James Wang, Yuhang Zang, Dahua Lin
arXiv:2606. 00384v1 Announce Type: new Abstract: Fitting quantitative models to data is a central step in scientific workflows, yet it remains one of the least automated.
By William Rudman, Abhishek Divekar, Kanishk Jain, Sebastian Joseph, Stella S. R. Offner, Matthew Lease, Kyle Mahowald, Greg Durrett, Junyi Jessy Li