The article reviews methods for assessing large language model (LLM) based AI agents in materials synthesis, focusing on their integration with experimental tools. It outlines evaluation strategies—including knowledge, reasoning, tool‑use, and closed‑loop benchmarks—and applies them to atomic layer deposition (ALD) as a case study. A practical framework for evaluating LLMs in this context is also presented.
By Angel Yanguas-Gil
El Agente Potente is an agentic system that integrates typed execution graphs and a coding mode to facilitate machine‑learning interatomic potential (MLIP) driven atomistic simulations. Typed execution graphs offer structured, provenance‑aware workflows where large language models handle planning and routing while deterministic Python code performs scientific computation and validation. The coding agent builds customized workflows for tasks needing procedural flexibility, invoking existing Potente functions for supported calculations. The system is demonstrated across materials discovery, energy‑landscape exploration, adsorption, and catalytic reaction workflows, with benchmarks on reproducibility and LLM token cost.
By Tsz Wai Ko, Jiaru Bai, Thomas Swanick, Yeonghun Kang, Changhyeok Choi, Angelina Qihong Jiang, Aiwei Yin, Varinia Bernales, Al\'an Aspuru-Guzik
arXiv:2608. 11224v1 Announce Type: new Abstract: Materials research advances through accumulated experience - scripts that work, protocols that are trusted, warnings attached to failed calculations or experiments, and judgement that links a new question to an old result.
By Siyu Liu, Bo Hu, Beilin Ye, He Cao, David J. Srolovitz, Tongqi Wen
CompMat-Bench is a new benchmark comprising 94 tasks drawn from recent computational materials science studies, designed to evaluate AI agents on realistic research steps without requiring costly simulations during testing. The benchmark pre‑reproduces inputs and outputs to provide ground truth, allowing agents to be graded with fixed rules rather than an LLM judge. It supports both single tasks and multi‑step workflows, with varying levels of methodological guidance, and shows that while agents can achieve high pass rates on individual tasks, performance drops in longer workflows or with reduced guidance, often due to scientific rather than software errors.
By Chenmu Zhang, Levi Felix, Jun-Jie Zhang, Xingfu Li, Xuelian Jiang, Tao Jiang, Subhendu Mishra, Xixi Qin, Boris Yakobson
arXiv:2607. 11526v1 Announce Type: cross Abstract: Material property prediction (MPP) infers key properties from chemical composition and structure, accelerating the discovery and optimization of novel materials.
By Hongxiao Li, Wanling Gao
arXiv:2512. 11935v2 Announce Type: replace Abstract: Agentic AI systems increasingly connect large language models (LLMs) to external scientific tools, yet whether and when tool access improves prediction accuracy remains uncharacterized.
By Jaehyung Lee, Justin Ely, Kent Zhang, Akshaya Ajith, Charles Rhys Campbell, Kamal Choudhary