arXiv AI

Evaluating LLM-based AI agents integrated with materials synthesis tools: the case of atomic layer deposition

The article reviews methods for assessing large language model (LLM) based AI agents in materials synthesis, focusing on their integration with experimental tools. It outlines evaluation strategies—including knowledge, reasoning, tool‑use, and closed‑loop benchmarks—and applies them to atomic layer deposition (ALD) as a case study. A practical framework for evaluating LLMs in this context is also presented.

arXiv AI
Aug 18

ALKEMIE Agent: an autonomous platform for computational materials design

arXiv:2608. 15776v1 Announce Type: cross Abstract: Despite the powerful multi-scale modeling methods and high-throughput infrastructures established in the materials community, real material computation workflows remain fragmented and heavily manual, requiring researchers to constantly bridge software tools, data analysis, and intermediate decisions.

By Hongfu Huang, Yuzhe Li, Ao Xu, Bo Liu, Changrui Wang, Kan Tang, Ning Yang, Shengxian Liu, Hanyu Liu, Pengpeng Zhang, Linggang Zhu, Fengkai Liu, Yichen Lu, Tong Zhao, Naihua Miao, Jian Zhou, Zhimei Sun
arXiv AI
2d ago

Hypothesis-Driven Autonomous Materials Synthesis with Multimodal LLM Agents

SynAgent is a framework that uses large language model agents to run autonomous experiments while building an explicit, revisable understanding of the synthesis process. Unlike traditional black‑box optimizers, SynAgent generates analysis skills on the fly and reasons multimodally over data such as X‑ray diffraction patterns and electron micrographs. In an 18‑experiment campaign on LiCoO₂ thin‑film deposition, SynAgent produced highly crystalline films and uncovered a sharp temperature threshold and optimal growth window (650–690 °C) for crystallization.

By Izumi Takahara, Kazunori Nishio, Akira Aiba, Shigeru Kobayashi, Takao Nakajima, Taro Hitosugi, Teruyasu Mizoguchi
arXiv AI
Sep 2

Agentic programs: an emerging form of scientific software in computational materials science

The article introduces the concept of agentic programs—scientific software that blends deterministic algorithms with bounded large‑language‑model (LLM) judgment, task‑specific verification, episodic maturation, and full delegation in production. It argues that recent LLM‑based agents enable this new form of computational materials science software. The authors illustrate the idea with DeMARS, an agentic program designed to build atomistic models from experimentally measured disordered crystal structures.

By Yunsung Lim, Haekwan Jeon, Jaesun Kim, Jisu Kim, Seungwu Han
arXiv AI
4d ago

El Agente Potente: High-Throughput Agentic Atomistic Simulations

El Agente Potente is an agentic system that integrates typed execution graphs and a coding mode to facilitate machine‑learning interatomic potential (MLIP) driven atomistic simulations. Typed execution graphs offer structured, provenance‑aware workflows where large language models handle planning and routing while deterministic Python code performs scientific computation and validation. The coding agent builds customized workflows for tasks needing procedural flexibility, invoking existing Potente functions for supported calculations. The system is demonstrated across materials discovery, energy‑landscape exploration, adsorption, and catalytic reaction workflows, with benchmarks on reproducibility and LLM token cost.

By Tsz Wai Ko, Jiaru Bai, Thomas Swanick, Yeonghun Kang, Changhyeok Choi, Angelina Qihong Jiang, Aiwei Yin, Varinia Bernales, Al\'an Aspuru-Guzik
arXiv AI
Sep 3

Can Coding Agents Reproduce Findings in Computational Materials Science?

The paper introduces AutoMat, a benchmark designed to test large language model (LLM) coding agents on their ability to reproduce claims from computational materials science. AutoMat presents three challenges: reconstructing underspecified procedures, navigating specialized toolchains, and assessing whether the evidence supports a claim. Experiments show that current LLM agents achieve low success rates, with the best setting reaching only 53%, and failures stem mainly from incomplete procedures, methodological deviations, and execution fragility.

By Ziyang Huang, Yi Cao, Ali K. Shargh, Jing Luo, Ruidong Mei, Mohd Zaki, Zhan Liu, Wyatt Bunstine, William Jurayj, Somdatta Goswami, Tyrel McQueen, Michael Shields, Jaafar El-Awady, Paulette Clancy, Benjamin Van Durme, Nicholas Andrews, William Walden, Daniel Khashabi
arXiv AI
Aug 24

An LLM agent for end-to-end computational materials discovery

MAESTRO is a large language model agent that automates the full screening pipeline for metal‑organic frameworks (MOFs). It parses extensive MOF literature, links publications to crystal structures, curates a computation‑ready database, and then applies a progressively more expensive computational strategy to identify promising candidates. The identified materials for wet flue gas separation come from unrelated studies, demonstrating the agent’s ability to uncover high‑performance materials across domains.

By Chen Yuntong, Huang Ju, Liu Yu, Zhao Dan, Sun Mingqi, Ju Chentian, Liu Yanbing, Huang Lijiang, Zhao Guobin
arXiv AI
Aug 28

From Atomic to Agentic: Towards Interpretable Evaluation of LLMs' Agentic Mathematical Capabilities

The paper introduces a new benchmark that evaluates large language models (LLMs) on their agentic mathematical reasoning rather than just final answers. It aligns problem‑solving behaviors with a taxonomy of reusable mathematical atomic capabilities and includes planning, action, and feedback tasks in both textual and multimodal settings. Experiments show that models with similar end‑to‑end accuracy can have very different agentic profiles, highlighting the importance of process‑level evaluation.

By Jiayi Kuang, Yinghui Li, Yunze Song, Keyu Chen, Zhifeng Shen, Yangning Li, Yidong Wang, Di Yin, Ruizhi Qiao, Xing Sun, Kai Jin, Ying Shen, Liang Lin, Philip S. Yu