CompMat-Bench is a new benchmark comprising 94 tasks drawn from recent computational materials science studies, designed to evaluate AI agents on realistic research steps without requiring costly simulations during testing. The benchmark pre‑reproduces inputs and outputs to provide ground truth, allowing agents to be graded with fixed rules rather than an LLM judge. It supports both single tasks and multi‑step workflows, with varying levels of methodological guidance, and shows that while agents can achieve high pass rates on individual tasks, performance drops in longer workflows or with reduced guidance, often due to scientific rather than software errors.
By Chenmu Zhang, Levi Felix, Jun-Jie Zhang, Xingfu Li, Xuelian Jiang, Tao Jiang, Subhendu Mishra, Xixi Qin, Boris Yakobson
arXiv:2605. 26179v2 Announce Type: replace-cross Abstract: Density functional theory (DFT) serves as the basis for computational discovery in materials science and chemistry, yet each calculation demands extensive human effort: adjusting algorithms when convergence stalls, revising plans when unexpected physics emerges, and inserting steps as intermediate results reshape the problem.
By Penghui Yang, Zhonghan Zhang, Yue Li, Xinrun Wang, Yanchen Deng, Yuhao Lu, Bijun Tang, Zheng Liu, Bo An
El Agente Potente is an agentic system that integrates typed execution graphs and a coding mode to facilitate machine‑learning interatomic potential (MLIP) driven atomistic simulations. Typed execution graphs offer structured, provenance‑aware workflows where large language models handle planning and routing while deterministic Python code performs scientific computation and validation. The coding agent builds customized workflows for tasks needing procedural flexibility, invoking existing Potente functions for supported calculations. The system is demonstrated across materials discovery, energy‑landscape exploration, adsorption, and catalytic reaction workflows, with benchmarks on reproducibility and LLM token cost.
By Tsz Wai Ko, Jiaru Bai, Thomas Swanick, Yeonghun Kang, Changhyeok Choi, Angelina Qihong Jiang, Aiwei Yin, Varinia Bernales, Al\'an Aspuru-Guzik
arXiv:2603. 15952v2 Announce Type: replace Abstract: Large language models (LLMs) are capable of emulating reasoning and using tools, creating opportunities for autonomous agents that execute complex scientific tasks.
By Jacopo Teneggi, S. M. Bargeen A. Turzo, Tanya Marwah, Alberto Bietti, P. Douglas Renfrew, Vikram Khipple Mulligan, Siavash Golkar
arXiv:2608. 03501v1 Announce Type: new Abstract: AI for Research (AI4Research) leverages AI to automate and improve scientific workflows.
By Zejun Liu, Jian Wu, Ru Peng, Yuliang Ji, Dongyuan Li, Renhe Jiang, Yue Zhang
arXiv:2507. 14267v2 Announce Type: replace Abstract: Large language model (LLM) agents can execute long-horizon scientific workflows, but their numerical outputs are difficult to trust: agents lose context, game verification checks, and can produce large volumes of plausible yet invalid results.
By Ziqi Wang, Hongshuo Huang, Hancheng Zhao, Changwen Xu, Shang Zhu, Jan Janssen, Venkatasubramanian Viswanathan