arXiv:2606. 05525v1 Announce Type: new Abstract: Recent advances in agentic visualization have enabled the translation of natural language into executable scientific visualization (SciVis) workflows.
By Kuangshi Ai, Haichao Miao, Kaiyuan Tang, Shusen Liu, Chaoli Wang
OSWorld-Science is a benchmark and evaluation environment for computer-using agents that use visual language models (VLMs) to perform scientific software tasks. It includes 12 VLMs and 146 high-quality tasks across domains such as molecular drawing, pathology image analysis, statistical computing, and physical simulation, with artifact-based evaluation and a harness that logs interactions and supports model comparison. The benchmark was developed through expert proposals and iterative human–AI co‑design, and results show that current VLMs still struggle with key scientific questions, offering insights into factors like language, reasoning, and context length.
By Dingyuan Dai, Heli Qi, Lei Liu, Yinxi Li, Baiding Chen, Zijun Dou, Qingcheng Zeng, Qi Kang, Oliver Sun, Eric Wang, Bo Zhou, Haixin Wang, Yufan Du, Shi Bo, Ruihan Lin, Mengqi Yuan, Dunjie Lu, Steven Dillmann, Yiming Shi, Tina Su, Amy Xin, Minghao Liu, Xi Wang, Xu Huang, Ge Zhang, Pengyu Nie, Zhen Yang, Jie Tang, Juanzi Li, Weihao Xuan, Tianyu Liu
arXiv:2606. 26614v1 Announce Type: cross Abstract: Large language model (LLM) agents enable natural language interaction for scientific visualization (SciVis).
By Kuangshi Ai, Patrick Phuoc Do, Chaoli Wang
MatToolBench is a new benchmark that evaluates multimodal GUI agents on professional materials science software. It contains 204 tasks across 10 tools in three modalities—GUI operation, OriginPro scripting, and code-based database queries—executed inside a Windows 11 VM. The benchmark offers fine-grained, expert-decomposed scoring and a high-performing multimodal judge for aesthetic assessment, revealing that strong general benchmark performance does not transfer to scientific workflows.
By Mei Wu, Rui Xie, Runyu Zhang, Yuqiang Li, Tianfan Fu, Bo Chen, Kai Yu, Xin Chen, Lu Chen
arXiv:2604. 27996v3 Announce Type: replace Abstract: This paper examines how large language model (LLM) agents perform on scientific visualization (SciVis) tasks that require generating visualization workflows from natural-language instructions.
By Jackson Vonderhorst, Kuangshi Ai, Haichao Miao, Shusen Liu, Chaoli Wang
arXiv:2603. 01421v3 Announce Type: replace Abstract: While large language models accelerate scientific discovery, existing agents face severe limitations in adaptability, domain generalization, and multimodal scalability, often struggling to autonomously process raw, domain-specific experimental data.
By Ke Lin, Owais Aijaz, Yilin Lu, Yiyang Luo, Xuehang Guo, Preslav Nakov